Skip to PREreview

PREreview of AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Published
DOI
10.5281/zenodo.23111148
License
CC BY 4.0

You give a coding agent a compact() action (it replaces the earlier history with a working-state summary the agent writes itself, and the original task stays). During data collection a judge (GPT-5.5-Codex) reviews each step of the base agent (when to compact, what the summary keeps, what the agent does next), and the corrected output is executed instead of the original. The 1,052 corrected trajectories go to SFT, then RL with GRPO on SWE-Gym, where the only reward is whether the final patch passes the tests. With Qwen3-Coder-30B-A3B-Instruct it reaches 39.6% on SWE-bench Verified and 24.5% on SWE-PolyBench Verified (9.2 and 5.0 points over base). I liked the "summary ignored" ablation in Section 3.3 - the same checkpoint does worse with compact() calls skipped, so the gain is not only from training.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. My four comments are about when to compact, cost, checking the summaries, and the size of the differences.

1. On when to compact - at 256K no trajectory reaches the forced-compaction threshold, and AutoCompact-SFT still beats the full-history Base at every budget (Section 3.3, Figure 3a). Figure 4a shows compact() on 44.3% of tasks after SFT and 58.5% after RL, but not how many times it fires per trajectory, or at which stage. In my own experiment (36 coding-agent runs, six clearing policies, DOI 10.5281/zenodo.22759217) clearing after every task was about a third more expensive (+33.5%, p = 0.002) than clearing every 3 tasks. I chose that comparison point after the runs. So I would like to see compactions per trajectory, the stage (after localization, after the edit, after tests), and how pass rate and cost change with the count. I think more tasks with compaction is not good by itself.

2. On cost - Section 3.1 estimates cost from token usage with Alibaba Cloud Model Studio prices (cached tokens at 20% of the input rate). Figure 3 gives pass rate per budget, but not the split of spending. In my measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216) more than 85% of modeled cost was context work - cache reads plus cache writes. On 168 sessions 94.4% of paid tokens were re-reading context already sent (contextburn, DOI 10.5281/zenodo.22712985). Since each call to compact() rewrites the context prefix (Section 2.3), for each budget and method please show spending on cached input, fresh input, output and summary writing. This would show what each compaction does to cached versus fresh input and where the saving comes from.

3. On checking the summaries - the rates in Figure 4b-c (key state omitted 3.1% to 0.2%, no next action 8.2% to 2.2%) come from keyword screening plus random manual spot checks (Section 3.3). Self-consistency is shown on one example only (django-13809, Figure 5), where the SFT summary records an unresolved syntax error but proposes to finish. In my guides I write that a model sounds equally sure when it is right and when it is wrong. I also write that a second fresh agent, told to treat the work as wrong until shown otherwise, finds what the first one missed. So I suggest to give each summary to a separate fresh model with one task - find a next action that contradicts the recorded state. Please report this rate for SFT and RL, and say how many summaries were checked by hand.

4. On the size of the differences - all results are averaged over three runs (Section 3.1), but Table 1 gives single numbers without spread. AutoCompact-SFT is ahead of SWE-Compressor by 1.2 and 1.6 points, and you read this as online correction being better than offline insertion (Section 3.2). In my own experiment I ran six repeats per condition and reported a p-value with Holm correction (DOI 10.5281/zenodo.22759217). The run data is public (Hugging Face, DOI 10.57967/hf/10366), so anyone can recompute it. I would give the spread across runs (or confidence intervals) at least for these two gaps before the conclusion. If possible please also release the corrected trajectories.

Thank you for a clear and useful paper.

Competing interests

Yes: the text cites the author's own technical reports and data (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985, 10.57967/hf/10366); no connection to the article's authors.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.