Skip to PREreview

PREreview of Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus

Published
DOI
10.5281/zenodo.23293074
License
CC0 1.0

Thank you for this paper. You take TRACE's public corpus of 590 harness-triggered compaction boundaries on AppWorld, each replayed under the pre-compaction context (PRE) and under the summary (POST), and ask whether what the agent did before a boundary predicts the burden of its next actions - calls that error or repeat a call already made (Sections 1-2). The placement contrast is a wide null, and the "has-written" label mostly tracks trajectory phase (Section 3). The best extension trigger reaches held-out AUROC 0.66 against 0.72 for a same-boundary replicate (Section 4). I liked that Appendix A dates every amendment and reports the gates that were not met.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.

1. On what a blocked compaction costs - The frontier counts compaction opportunities with equal weight, because the release has no token counts (Section 4, Appendix A), and the offline value prices only the next k actions (Section 5). For me the other side of this trade is the bill. In my own logs (722 sessions, DOI 10.5281/zenodo.22759216) more than 85% of the modeled cost was work with context (cache reads plus cache writes), and sessions longer than 200 calls (8% of sessions) gave more than 90% of the money. The cost is modeled from a price list, not from invoices. So a blocked compaction has a price too - the PRE context stays in every later call. When corpora ship token counts, could the x-axis weight each opportunity by the tokens a compaction would remove? That would also allow the token-budget comparison that the abstract says cannot be evaluated on this release.

2. On when to clear - First compactions carry more burden than later ones (+0.274 [+0.084, +0.474] under task fixed effects, Section 5 and Appendix D), and the stump's shorter-history leaf has the higher positive-burden rate in all five folds (Appendix E). I see a similar shape in a different setting (in my runs the session is cleared between tasks, not summarised). In 36 of my own coding-agent runs under six clearing policies, cost followed a U-curve with the minimum at clearing every 3 tasks (plateau from 3 to 6), and clearing after every task cost about a third more (+33.5%, DOI 10.5281/zenodo.22759217). The comparison point was chosen after the runs, and in the every-task arm the agent wrote more tests, so part of that gap may be extra work. Could you add a plain rule on compaction index and prefix steps only as one more comparator on the frontier? If it lands near the stump, a harness designer gets the same advice without a fitted model.

3. On what counts as harm - I like that you judge compaction by the next actions and not only by task success. In my runs the test gate returned 4,086 tests and 0 failures across all 36 runs, while cost between conditions differed by 33.5%. Tests per run ranged from 97 to 127 because the agent wrote the tests itself, so "all green" told me little. For your best model the refetch channel is more predictable than the blocked one (held-out AUROC 0.65 against 0.56, Appendix F), and in the Section 2 example the refetch is the agent re-issuing a call it already made after a NameError. Could you report how many refetches are followed by a valid step? That would show how much of the burden is a re-read and how much is a stall.

4. On data - You do not redistribute the corpus (it has no licence file), and the analysis record is available on request (Section 5). Section 5 also lists what corpora should ship: ordered actions, token counts, summary text and more replays. My U-curve run data is public on Hugging Face (DOI 10.57967/hf/10366) and on OSF (DOI 10.17605/OSF.IO/5QTWY), so anyone can recompute it. Could you post the frozen protocol files with their SHA-256 hashes and the per-boundary out-of-fold scores? Readers could then check the frontiers without writing to you.

Thank you for the careful write-up.

Competing interests

The review cites my own technical reports and open data (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.57967/hf/10366). I have no connection to the authors of the preprint.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.