Comentarios
Escribir un comentarioNo se han publicado comentarios aún.
## Summary
This paper tackles a genuinely important problem in production agent systems: context management as a *learned policy* rather than a length-triggered fallback. The authors augment a coding agent (Qwen3-Coder-30B-A3B-Instruct, terminal REPL scaffold) with a `compact()` action and train it in two stages: (1) SFT on 1,052 judge-corrected trajectories collected from 379 SWE-rebench tasks, where GPT-5.5-Codex acts as an online judge correcting compaction timing, working-state summaries, and post-compaction actions, with corrections executed in the environment before being used as training data; (2) GRPO-based RL on SWE-Gym with binary task-success rewards, no KL penalty, no entropy bonus, and no compaction-specific reward shaping. The headline result is +9.2% / +5.0% absolute pass-rate gains over the base model on SWE-bench Verified (39.6%) and SWE-PolyBench Verified (24.5%), with gains holding across per-task inference budgets from $0.10 to $4.00. The paper's most distinctive claim is that supervising *what the agent does after compaction* — not just when to compact or what to write in the summary — is the missing piece in prior work, and the online-correction data collection method is the mechanism that delivers it.
## Strengths
1. **The problem framing is correct and under-studied.** The introduction's distinction — compaction tied to *task progress* rather than *context length* — matches production experience: stale exploration (dead hypotheses, verbose tool outputs) degrades agent behavior long before any window limit is hit. The Figure 3(a) result, where proactive compaction helps even in the 256K regime where nothing overflows, is the paper's most practically significant finding.
2. **The data-collection design is the real contribution.** Correcting outputs *before* execution so the trajectory continues from corrected decisions (Section 2.2) is meaningfully different from post-hoc annotation or offline insertion (SWE-Compressor). The AutoCompact-SFT vs. SWE-Compressor comparison (+1.2% / +1.6% with comparable data) is a fair, well-scoped ablation supporting this.
3. **"Summary self-consistency" is a useful new lens.** The Section 3.4 case study (django-13809) — a summary can satisfy coverage checks while proposing a next action incompatible with its own recorded state — names a failure mode I have seen in production agent summaries but rarely seen formalized. That SFT imitates summaries while RL constrains them through downstream success is a clean, honest explanation of the mechanism.
4. **Budget-framed evaluation is the right framing.** Reporting pass rate against dollars-per-task, with explicit token pricing (Alibaba Cloud Model Studio, cached tokens at 20%), treats inference economics as a first-class metric. Too few agent papers do this.
## Major concerns
**1. The judge is load-bearing but unvalidated.** GPT-5.5-Codex reviews every proposed action during data collection and rewrites flawed outputs — yet the paper reports no judge agreement rates, no human audit of corrections, and not even the *correction rate* (what fraction of proposed actions were rewritten?). The 24%/53%/23% split across correction types describes the collected data, not the judge's reliability. This matters because the SFT stage imitates judge-written outputs: to the extent the judge's compaction preferences are idiosyncratic, AutoCompact-SFT distills GPT-5.5-Codex's priors as much as it learns compaction. A modest human audit (e.g., 100 sampled corrections rated by the authors) would substantially strengthen the paper.
**2. RL credit assignment is unresolved.** The binary task-success reward is shared across all trajectory segments, so compaction decisions receive the same advantage signal as coding actions (§2.3). The summary-ignored ablation (§3.3) shows executing compaction helps, but it cannot separate "better compaction policy" from "RL made it a better coder" — RL on SWE-Gym is independently known to improve coding (cf. SWE-RL, cited in the paper). The 7.4% RL-over-SFT gain could be largely general coding improvement. The clean control the paper needs is RL *without* the `compact()` action available: if that closes most of the gap, the compaction-specific story weakens considerably.
**3. The main table confounds trigger type with context regime.** In Table 1, the length-triggered baselines (Fixed Compaction, CompactionRL) are evaluated at a 16K forced-compaction threshold while everything else runs at 256K. The claim that "length-triggered compaction yields limited gains" is therefore confounded — Fixed Compaction underperforms Base partly because it operates in a far more constrained regime. The fair 16K-vs-16K comparison exists only in Figure 3(c) (Base vs. AutoCompact, shared fallback). The headline +9.2% should be reported alongside the 16K-regime numbers in the main table, not relegated to a budget-curve subplot.
**4. The 16K regime is the production reality; the paper leads with 256K.** The authors' own limitations section notes that Codex and Claude Code compact automatically at the window limit — i.e., production *is* the length-triggered fallback setting. The 256K "never overflows" regime, where the headline gains are measured, is the less realistic deployment scenario. A practitioner reading this wants the constrained-regime numbers front and center, plus guidance on how learned proactive compaction interacts with harness-level auto-compaction (left as future work).
**5. No cost/latency accounting for compaction itself.** The budget analysis prices task tokens but never decomposes the overhead of the mechanism being proposed: how many output tokens does the average `compact()` call cost, and what is the *net* token delta per solved task versus Base? Relatedly, the training bill is unreported — 1,052 trajectories corrected online by GPT-5.5-Codex represents substantial judge inference cost. For a paper whose §3.3 leads with cost efficiency, the cost of the method itself should be quantified.
**6. Generalization is untested beyond bug-fix tasks.** All training and evaluation tasks are SWE-bench-style code repair (SWE-rebench, SWE-Gym, SWE-bench Verified, SWE-PolyBench Verified). "Long-horizon coding agents" in the title really means "bug-fix agents" in the experiments. There is no evidence on multi-file feature development, long-horizon data analysis, or tool-heavy research tasks where the working state to preserve looks very different from a bug-localization summary.
**7. Statistical reporting is thin.** Results are "averaged over three runs" with no variance or error bars anywhere — relevant when several baseline gaps are 1–2%. The Figure 4 summary-quality metrics rest on "keyword-based screening, supplemented by random manual spot checks," but the screening keywords, the spot-check sample size, and any inter-rater agreement are unspecified.
## Minor concerns
- The RL training uses 32K-token sequences while evaluation runs at 256K (acknowledged in limitations). Compaction *timing* learned under a 32K horizon may not transfer to 256K timing decisions; the paper could at least discuss this.
- Concurrent work SWE-MeM (Gao et al., 2026) is cited in related work but never empirically compared, despite addressing the same problem with RL.
- The $0.10–$4.00 budget conclusions depend on a single vendor's pricing (Alibaba Cloud) and the 20%-of-input cached-token assumption; a sensitivity note would help.
- The "AI Use Statement" discloses generative AI use for "trajectory correction for training" — given Major concern #1 about the judge, it would be worth clarifying whether the GPT-5.5-Codex judge outputs count as the "trajectory correction" referred to here, since those outputs are training data, not just editing assistance.
## Practitioner perspective
In production agentic systems, context is the dominant cost driver and the dominant failure driver at the same time — every stale tool output in the window is both money spent and a distraction injected. This paper's core instinct is right: the industry's default (compact when the window fills) compacts at the worst possible moment, mid-stage, when the evidence the agent still needs is most at risk. The "compact on task progress" principle, and especially the emphasis on *continuation behavior after compaction*, matches what breaks in real deployments: I have repeatedly seen agents generate reasonable summaries and then ignore them, re-running exploration the summary already recorded.
What would make this actionable for a practitioner: (a) the 16K-regime numbers as the headline, since that mirrors every major coding harness; (b) a failure-mode analysis of compaction itself — when the summary silently drops a critical fact, does the agent recover or corrupt silently? There is no measurement of compaction-induced errors anywhere in the paper; (c) the net token economics including the summary-generation overhead, not just task tokens. The paper gestures at (a) and leaves (b) and (c) untouched.
## Overall assessment
This is a solid, well-executed paper whose best idea — online judge-corrected data collection that supervises post-compaction behavior, not just compaction decisions — is genuinely novel relative to the cited baselines, and whose budget-framed evaluation sets a good example. I would recommend it with revisions: validate the judge (even a small human audit), add the RL-without-compaction control, report the 16K numbers in the main table rather than only in a subplot, and quantify the token overhead of compaction itself. None of these requires new large-scale training runs except the control in #2, which reuses the existing RL setup minus the action.
**AI disclosure:** I used an AI assistant to help draft the initial text of this review; I added my own practitioner assessment, verified the claims against the paper, and approved the final version.
The author declares that they have no competing interests.
The author declares that they used generative AI to come up with new ideas for their review.
No se han publicado comentarios aún.