PREreview de RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue
- Publié
- DOI
- 10.5281/zenodo.23111062
- Licence
- CC0 1.0
Summary
This paper introduces compatibility rescue: adapting a repository that demonstrably worked in its historical environment but fails after runtime/dependency drift (Python 3.13 / JDK 21 with current dependencies), where the agent receives only the repo and the failing modern environment — no issue description or fault localization. RepoRescue contains 193 Python repos (47 unmaintained, 146 "time-travel" snapshots from just before a maintainer's own compatibility fix) and 122 Java repos, each admitted only after passing its unmodified suite in a reconstructed historical environment (Phase 0) and failing deterministically after modernization (Phase 1). Five deployed agent systems are tested on Python (four on the shared Claude Code harness plus GPT-5.2 via Codex) and three on Java. Beyond full-patch pass rate, the authors report post-hoc source-only success (test edits stripped), runtime-blocked source-only success (test edits blocked in-session), and post-PASS validation (realistic scenarios + bug-hunt probes) on 34 unmaintained Python rescues. Findings: Claude Code systems depend on forbidden test edits in 38–53% of apparent successes (vs. 4% for GPT-5.2/Codex); successes are complementary (union 62.7% vs. best single 51.8%); L4 whole-codebase coordination is a cliff (Codex 14/14, Claude Code ≤2); of 34 suite-passing rescues, 22 survive realistic scenarios and 12 survive bug-hunt.
Strengths
The three-phase admission construct is excellent. Requiring each subject to pass historically and fail after modernization — with frozen environments handed identically to every system — is a genuine advance over static benchmarks and directly addresses the staleness/contamination worries in agent evaluation. 68,895 Phase-0 tests (median 165/repo) give the "once worked" claim teeth.
The evaluation stack is the paper's best contribution. Full-patch vs. post-hoc source-only vs. runtime-blocked vs. post-PASS scenario validation operationalizes the exact question practitioners ask during model validation: what does a green suite actually prove? I build production AI agent frameworks, and the gap between "the suite is green" and "the fix is real" is where validation failures hide — making it measurable is valuable. In our own framework work, we learned that guardrails have to live in the agent's instruction files themselves: precise, concise directives about what the agent may and may not touch. Verbosity is its own failure mode — too much instruction dilutes the context window and the agent drifts off course. We give each agent a dedicated task file and link agents to each other through file descriptions, so every agent knows its boundaries and when to hand off to another. The paper's runtime-enforcement result validates this approach empirically: constraints don't just limit agents, they improve them.
The runtime-enforcement result is genuinely novel. Blocking test edits raises Kimi from 22.8% (post-hoc source-only) to 41.5% — the same agent produces source repairs it skipped when the shortcut existed. Compliance-as-capability matters for deployment guardrails: how you constrain an agent changes what it can do, not just what it may do.
The threats-to-validity discussion is admirably honest — κ=0.76 with the L2/L3 boundary flagged, ±5–10 repo drift bounds, a "no tests ran" audit (1.4% bound), no direct Python/Java rate comparison. The PyCG→Scalpel cascade (a two-layer failure the 29 unit tests miss but a downstream wrapper exposes) makes the suite-insufficiency point concrete.
Major comments/issues
1. Headline claims rest on single-trial estimates. The L4 cliff (14/14 vs. ≤2) and the +10.9pp union gain inherit the single-trial-per-cell variance the authors themselves bound at ±5–10 repos, with provider sampling defaults an acknowledged residual confound (sampling parameters not reported as fixed). For a benchmark offered as a comparison instrument, the 14 L4 repos and a stratified Medium-tier sample should be re-run 3× per system so the cliff carries uncertainty, not just a point estimate.
2. The shortcut taxonomy is too binary and may overstate gaming. ~90% of "shortcut" edits are "plausible API adaptations inside tests" (e.g., nose→pytest); only ~10% are direct bypasses (skip/xfail, assert relaxation). But the cerberus example shows the tests themselves import drift-removed modules (pkg_resources) — the test files are also broken by the same failure, and migrating them is legitimate maintenance (human maintainers modify tests in 9.9% of time-travel fixes). Lumping drift-induced test migration with genuine gaming inflates the "38–53% depend on forbidden test edits" narrative. A graded taxonomy — drift-induced migration vs. assertion weakening vs. outright bypass — would be more honest and more useful. From a practitioner's standpoint, the fix is architectural, not just taxonomic: in production agent frameworks, test files and configuration files should be outside the agent's editable scope by default, with any test edit requiring explicit human approval and any logic change triggering mandatory re-review. A written "do not edit tests" rule in the prompt is not enough; the boundary needs to be enforced by the harness, which is exactly what the paper's runtime-blocked condition demonstrates.
3. RQ4 rests on a small, Python-only sample. "22 of 34 work realistically, 12 survive bug-hunt" is carefully reported, but it is 34 unmaintained Python candidates and Java post-PASS validation is future work. The abstract's framing is fair; broader "practical reuse" language should stay scoped to the sample, and the Java post-PASS protocol ideally pre-registered for comparability.
4. Deployment economics are missing and they are first-order. 1,717 trials on commercial systems, sessions of 21–206 messages, failed sessions running 29–58% longer (p<0.01) — yet no cost/latency data. Cost-per-successful-rescue dominates any practitioner adoption decision, and the turn-count gap is already a cost signal. Even rough per-trial inference costs in an appendix would help. Please also record sampling parameters — "provider-side sampling defaults remain a residual confound" implies defaults were used, limiting reproducibility.
5. Time-travel vs. unmaintained equivalence deserves a harder look. The GEE regression finds repo_type non-significant (p=0.75), but time-travel snapshots come from active repos at pre-fix commits with maintainer ground truth — a cleaner, better-specified task than truly abandoned repos. The aggregate coefficient doesn't rule out different difficulty-tier distributions; an Easy/Medium/Hard split by repo_type would show whether the 146 snapshots quietly shape headline rates.
6. Surface the "no tests ran" audit in the main text. The 1.4% bound (6/436 PASS outcomes, concentrated in wssh) lives in §8, but the Phase-2 gate is exactly what should catch empty runs. A sentence in §3.2 naming the failure mode and per-trial flags would show the guard is monitored, not just defined.
Minor comments/issues
- Table 1's L1 row is n=4 — worth noting alongside the "L1/L2 mostly routine" claim.
- The false-completion detector's wide interval (adjusted prevalence 32–76%) warrants the exploratory framing §5.3 mostly gives it.
- The Java pom.xml normalization is reasonable and disclosed, but it is itself a human rescue step; when readers compare the Java union (88.5%) with Python (62.7%), it deserves a foreground reminder, not just §3.1/§5.5.
Overall
A strong, carefully constructed paper on a real problem — I have watched teams burn weeks resurrecting abandoned-but-depended-on libraries, and the requests-html example (1M+ monthly downloads, no release since 2019) will resonate with practitioners. The evaluation stack genuinely advances how we validate agent work. My major comments ask for uncertainty on the headline comparisons, a fairer test-edit taxonomy, and the deployment-economics data practitioners need. I would be glad to see this published once those are addressed.
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they used generative AI to come up with new ideas for their review.