PREreview de FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
- Publié
- DOI
- 10.5281/zenodo.23057016
- Licence
- CC BY 4.0
The authors present FOCUS, a training-free way to compress an agent's history at test time. A draft model writes plan sketches and the spans they cite most often are kept. They report peak context down by up to 48% and task success up to 8.9 points over uncompressed execution. I liked that it works on whole spans, not tokens (Section 4.1), so tool calls, code and error messages are not cut in half. I also liked the defensive verification step (Section 4.4) - it keeps a failed tool call or rejected API request even if no plan cites it. And the cost tables are honest - with gpt-4.1 as agent and draft the cost per task is essentially equal to no compression (0.245 and 0.249 USD).
I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. My four comments are about cost with caching, how often to compress, the defensive check, and repeats.
1. On cost with caching - Appendix A.9 and Table 13 price every input token the same, $2.00 per 1M for gpt-4.1. Without compression the history only grows at the end, so earlier calls share the same prefix. When FOCUS compresses, it rebuilds the trace, so the next call starts from a new prefix. In my own measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216) more than 85% of modeled cost was cache reads plus cache writes, and providers price those differently from fresh input. In my experiment on 168 sessions 94.4% of paid tokens were re-reading context already sent (DOI 10.5281/zenodo.22712985). I suggest reporting Table 13 also with cached-input pricing, or at least counting per task how many calls come right after a compression. Then readers can see if the 41.18 to 37.82 USD saving holds when caching is on.
2. On how often to compress - In Figure 3 (AppWorld) a smaller delta_mem triggers compression earlier and more often, and moderate thresholds give the best trade-off between accuracy and peak tokens. The figure shows accuracy and peak tokens, not API cost, and Tables 5, 12 and 13 use one fixed threshold. Every compression adds N draft-model calls (N = 3 by default, Appendix A.9). In my own 36-run experiment with six context-clearing policies, the cost had a U-shape, with the minimum at clearing every 3 tasks (clearing after every task was about a third more expensive, DOI 10.5281/zenodo.22759217). (My caveat: that comparison point was chosen after the runs.) So I would add total API cost (agent plus draft) for each delta_mem setting in Figure 3, because compressing too often may cost more, not less.
3. On the defensive check - in Section 4.4 the same draft model that wrote the plan sketches then reviews the spans that were not cited. In Appendix A.2.5 you say it helps most when the draft model is overconfident in discarding a span (AppWorld goes from 56.5% to 64.9%). In my guides I write that a model sounds equally sure when it is right and when it is wrong, so I think the same model may miss its own overconfidence. Maybe try the verification step with a different model than the one that wrote the plans (a second, fresh agent finds what the first missed). And add a check that does not depend on any model - log for each dropped span whether the main agent later repeated a failed action or asked again for the same information.
4. On repeats - in Appendix A.9 the main agent runs at temperature 0.0 with seed 42, and the tables report one number per cell without spread. Appendix A.2.5 mentions run-to-run variance of about plus or minus 1.4% std on about 100-task splits. The headline AppWorld gain is 56.0% (no compression) to 64.9% (FOCUS-D), Table 1. In my context-clearing experiment I ran six repeats per policy and put the run data in the open (Hugging Face dataset, DOI 10.57967/hf/10366). I suggest reporting several seeds with the spread for the main tables and releasing the agent trajectories, so others can check the gains.
Thank you for a clear and useful paper.
Competing interests
Yes: the text cites the author's own technical reports (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985) and his open dataset (DOI 10.57967/hf/10366); no connection to the preprint's authors.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.