Ir para detalhes do preprintIr para avaliações PREreview

Avaliações PREreview de (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs

1 PREreview

  1. Avaliação PREreview de Karmendra Pandey

    ## Summary

    This paper makes the case that LLM APIs need their own regression-testing discipline, and backs it with an exploratory case study: toxicity detection across five GPT-3.5 models (text-davinci-002 through gpt-3.5-turbo-instruct, spanning 18 months), two datasets (Civil Comments, GitHub…

    Ler a avaliação PREreview de Karmendra Pandey