Commentaires
Écrire un commentaireAucun commentaire n’a encore été publié.
Thank you for this paper. You ran five models through three production harnesses (Claude Code, mini-SWE-agent and OpenCode) on SWE-bench Verified, and reran the same configurations to see how much a score moves when nothing changes but the run (Section 3). On the 45 hardest tasks, swapping the harness flips 13% of tasks - the same as rerunning the same harness (Section 4.2). The one effect that clears the noise is a loss, with OpenCode trailing by up to 9 points (Section 4.3). And cost per task differs by up to 3x, set mostly by the preamble each harness pays again on every step (16,581 tokens in Claude Code, 829 in mini-SWE-agent) times the number of steps (Section 4.4, Table A4). I liked the recommendation in Section 5 to report a second run of the same cell next to any harness gap.
I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.
1. On the bill - Section 4.4 says 96 to 99% of input tokens are cache hits, and the cache price scales the bill but does not reorder it. In my own 722 agent sessions (DOI 10.5281/zenodo.22759216) more than 85% of the modeled cost was work with context, meaning cache reads plus cache writes. These ran on a subscription, so the cost comes from a price list (the same caveat as your Opus column, Table A5). In a separate set of 168 sessions, 94.4% of billed tokens were re-reading context already sent (DOI 10.5281/zenodo.22712985) - a share of tokens, not of money. Where a provider prices cache writes separately, could you report cache-write tokens as a fourth line, and mark which steps followed a compaction (a compaction rewrites the cached prefix)? My IETF Internet-Draft (draft-arsentev-agent-run-metrics-00, an individual draft, not a standard) proposes a common format for agent run cost records, and your three numbers - first call, growth per step, step count - are what such a record should carry.
2. On compaction - Claude Code compacts in a fifth to a third of its trials (Section 4.4). On the Qwen arms its window was set to 110k, but on the vendor arms it kept the default and compacted much later, at a median of 168k prompt tokens on GLM-5.3-Flash (Table A1, Table A4 note d). In my own experiment (36 coding-agent runs, six context-clearing policies with six repeats each, DOI 10.5281/zenodo.22759217), clearing the context after every task cost about a third more (+33.5%) than clearing every three tasks. So for me, when to compact is a cost setting of its own. Could you show the cost per trial split by trials with and without a compaction?
3. On the runs where the model stopped on its own - under OpenCode the model stops sooner: on Qwen3.6-35B-A3B over the pool, a median of 28 model calls against 71 under mini-SWE-agent and 55 under Claude Code, and the data don't say why (Section 4.3). In my guides I write that every task needs its own check the agent can run. Without it, "looks done" is the only signal the agent has, and a model sounds equally sure when it is right and when it is wrong. So I wonder about the failures where the model stopped on its own (Figure 3). From the tool-call sequences, did the agent run the repository's tests before it stopped? Splitting these failures by whether any test was run might show part of what makes the model stop sooner under OpenCode.
4. On reruns and the data - I agree with the second-run recommendation in Section 5. In my 36-run experiment every policy had six repeats, and the run data is public on Hugging Face (DOI 10.57967/hf/10366), so anyone can recompute it. The trajectories will be released with the paper (Reproducibility statement). Could you deposit them with a DOI and include per-step token counts split by cache hit, miss and output? Then readers could reprice the bill under other price lists, as you did yourselves in Appendix G.
Thank you for a clear and useful paper.
Yes: the text cites the author's own technical reports (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985), dataset (DOI 10.57967/hf/10366) and IETF Internet-Draft draft-arsentev-agent-run-metrics-00; no connection to the article's authors.
The author declares that they did not use generative AI to come up with new ideas for their review.
Aucun commentaire n’a encore été publié.