Ir para a Avaliação PREreview

Avalilação PREreview de (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs

Publicado
DOI
10.5281/zenodo.23214938
Licença
CC0 1.0

## Summary

This paper makes the case that LLM APIs need their own regression-testing discipline, and backs it with an exploratory case study: toxicity detection across five GPT-3.5 models (text-davinci-002 through gpt-3.5-turbo-instruct, spanning 18 months), two datasets (Civil Comments, GitHub Discussions), and four prompting strategies (simple, instruction-last, detailed, few-shot). The headline results: 58.8% of prompt+model combinations lose accuracy across API updates (70.2% of those by more than 5 points); 55% of updates help some prompts while hurting others, so the "best" prompt changes from version to version; 10.9% of individual predictions regress even when aggregate accuracy improves; and regressions concentrate on data slices — 90% land on toxic discussions, disproportionately politics-triggered, code-targeting, and severe ones. From this the authors derive three rethinks for LLM regression testing: slice-level granularity (a new correctness notion), versioning prompts alongside models, and handling non-determinism. I cite this paper in my own work on behavioral regression testing for AI agents as the academic home of the "why" argument, so I read it as someone building directly on it.

## Strengths

The "prompt is also a version" insight is the paper's most original contribution. Prior work on evolving ML APIs treated the model as the moving part. This paper shows the moving part is the pair: the same silent update that drops one prompt by 9.6% lifts another by 5.1%, and the detailed prompt that was best on every model until the last one suddenly trails few-shot by 8.7%. That non-uniformity result is genuinely new, well-evidenced, and it reframes prompt engineering from a one-time activity into a versioned artifact. Anyone running agents in production — where the "prompt" is a whole system of templates, tool definitions, and routing policy — should read this as a warning about a larger blast radius. The slice-level regression finding has real consequences: regressions landing disproportionately on politics-triggered and severe toxicity is a fairness and safety observation, not just a testing one. Aggregate accuracy improving while a slice degrades is exactly the failure mode that ships incidents. The limitations section is honest, and the paper resists overselling an exploratory study as a solution.

## Major comments

1. The headline numbers are single-trial at temperature 0, but the paper itself establishes that temperature-0 outputs are non-deterministic (citing Ouyang et al.). If predictions can flip between identical runs, then the 58.8% combination-level regression rate and the 10.9% per-prediction flip rate inherit run-to-run variance that is never quantified. The paper's own discussion warns that flakiness must be designed into LLM regression testing — but the case study's measurements don't report any repeat-run stability. Re-running even a stratified subset 3-5x would bound whether the headline figures are signal or partly sampling noise.

2. The confidence analysis rests on a rough proxy. Entropy is estimated from n=20 samples at t=0.7 because APIs don't expose probabilities — a reasonable workaround, but 20 samples is a noisy estimate of a distribution, and the striking "63.8% of regressions happen at entropy=0" claim depends on it. The paper finds models differ in self-consistency yet doesn't connect this back: a model whose calibration shifts under update arguably needs calibration regression as its own test dimension.

3. Slice discovery is the load-bearing recommendation with the thinnest scaffolding. The slice finding depends on author-provided metadata the paper admits "may not be available for many datasets." Slices discovered on the old model may not be the slices that regress on the new one — discovering slices predictive of future regressions is the harder problem practitioners actually need.

4. No cost dimension. Re-running a slice-level suite across prompts x models on every silent API update costs real money; the economics of the proposed continuous re-validation deserve at least a mention.

5. Prompt versioning is asserted more than designed. Given the central claim that prompts are first-class versions, a sketch of the versioning data model would strengthen the vision: what exactly gets versioned, and what does a "prompt diff" that explains a regression look like?

## Minor comments

- The deprecation angle deserves more foreground: 4 of the 5 studied models were scheduled for deprecation, so this is about forced migrations, the sharper practical trigger for regression testing.

- The two measurement regimes (temp 0 for accuracy, t=0.7 for entropy) should be cross-referenced on the same inputs.

- Figure 1's prompt-rank flip (8.7%) is nearly as quotable as the 58.8% — it deserves a sentence in the abstract.

## Overall

A strong, timely vision paper with genuine empirical grounding. The prompt-as-version insight is novel and well-evidenced, and the slice-level findings have consequences beyond testing. My major comments ask for stability bounds on the headline numbers, a harder look at the slice-discovery chicken-and-egg problem, and more design detail on the versioning machinery. I am building directly on this work and would be glad to see it extended along these lines.

---

AI disclosure: An AI assistant helped draft this review from my notes; I read the paper in full, wrote the technical judgments, and approved the final text.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.