Saltar a PREreview

PREreview del Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback

Publicado
DOI
10.5281/zenodo.23282285
Licencia
CC0 1.0

Thank you for this paper. You model the always-loaded context file (AGENTS.md and similar) as an assortment of instructions competing for limited attention, each with a setup cost per session (Section 2). You get a bound on the optimal file size that does not depend on the number of candidates (Proposition 1), show that appending every instruction with positive standalone value can be arbitrarily worse than choosing a subset (Corollary 1), and prove that deleting what the agent ignores is blind to value (Proposition 3). In the size sweep the best tested file stays at 10, 5 and 10 items for Haiku and 40, 20 and 40 for Sonnet as the pool grows from 200 to 3,200 (Section 5). I liked that Appendix D lists every deviation from the plan.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.

1. On what one session costs - The abstract says each loaded token "is charged again in every later round of the session". In Section 2 it is charged once per session and later turns "hit the prompt cache", so the setup cost κC(S) does not depend on session length. In my own logs cached context is not a small part of the bill. Over 722 sessions (DOI 10.5281/zenodo.22759216) more than 85% of the modeled cost was cache reads plus cache writes, and sessions longer than 200 calls (8% of sessions) gave more than 90% of the money. The cost is modeled from a price list, not an invoice. In 168 other sessions 94.4% of billed tokens were re-reading context already sent (DOI 10.5281/zenodo.22712985) - a share of tokens, not of money. Could the setup cost be one cache write plus a cache read per call?

2. On deleting - Proposition 3 fits with how I write these files. In my guides I keep the file short (up to about 200 lines) and test each line with one question: if I remove it, will the agent start making mistakes? That asks about value, not compliance, but has the censoring you describe - the answer shows up only when a mistake does. In the same guides, procedures needed only now and then go into skills, which load on demand. Could the model get an on-demand tier that pays its tokens only in sessions where the instruction is relevant (probability p_i)?

3. On what the experiment measures - The size sweep uses ManyIFEval writing tasks, the file is the whole system prompt, there are no tools and prompt caching is off (Appendix D). Most items are rewritten as keyword rules, and the score "measures instruction compliance and not the overall quality of the response" (Section 5). Your introduction's example is different - "always use conda" breaks a repository that uses uv, so a followed rule has a negative gain. In my guides I write that every task needs its own check the agent can run, because without it "looks done" is the only signal it has. A coding-agent version of the sweep could use such a check as the gain. Did any run include an instruction with negative value, so that Proposition A.2 can be seen in data?

4. On repeats - Your sweep has two replicates per cell that differ in item order, and Appendix D.1 shows that order matters (at b = 800 of the largest pool the first fifth of the file is followed about twice as often as the rest). Would more orders change the measured optima n*(N)? The setup cost is also paid each time a session starts, so how often an operator clears the session matters. In my own 36 coding-agent runs, clearing after every task cost about a third more (+33.5%) than clearing every three tasks (DOI 10.5281/zenodo.22759217). The comparison point was chosen after the runs and the agent wrote more tests in that arm, so it may not transfer. My run data is public on Hugging Face (DOI 10.57967/hf/10366). Per-cell results of your sweep, without the item text, would let readers recompute the optima.

Thank you for a clear and careful paper.

Competing interests

Yes: the review cites the author's own technical reports and dataset (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22712985, 10.5281/zenodo.22759217, 10.57967/hf/10366). No connection to the preprint's authors.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.