Aller directement aux détails du preprintAller directement aux PREreviews

PREreviews de (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs

1 PREreview

  1. PREreview par Karmendra Pandey

    ## Summary

    This paper makes the case that LLM APIs need their own regression-testing discipline, and backs it with an exploratory case study: toxicity detection across five GPT-3.5 models (text-davinci-002 through gpt-3.5-turbo-instruct, spanning 18 months), two datasets (Civil Comments, GitHub…

    Lire la PREreview de Karmendra Pandey