Ir para a Avaliação PREreview

Avalilação PREreview de Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

Publicado
DOI
10.5281/zenodo.23011395
Licença
CC BY 4.0

This paper looks at 1,200 trajectories from Claude Code and Mini-SWE-Agent on SWE-bench Verified. Three cost-inefficient behaviors (subsumed retrieval, similar script generation, test re-execution) show up in 79 to 98% of tasks and take up to 22.75% of task cost (Table 1). What I liked - tasks split by creation date, baselines run three times before an effect is called robust, and the negative CodeGraph result reported openly.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents and I measure what they cost me, so maybe my practical side is useful here. My four comments: where the cost sits, subagents and re-reads, re-running tests, and the skill set size.

1. Regarding where the cost sits - in Section 4.1, Figure 6, task cost moves more than behavior-attributed cost, slopes 2.83 on Verified-200 and 1.33 on Pro-100. One reason you give is that a shorter trajectory means less repeated cache reading of earlier context (the other is blocking unflagged actions). In 722 of my own sessions (150,902 model calls) more than 85% of the modeled spend was cache reads plus cache writes (sessions I ran myself, cost modeled from list prices, not bills; technical report, not peer-reviewed, DOI 10.5281/zenodo.22759216). In a separate 168-session run 94.4% of paid tokens were re-reads of context already sent (DOI 10.5281/zenodo.22712985, tokens, not the bill). So my numbers agree with your first explanation. Maybe splitting per-trajectory cost into cache reads, cache writes, fresh input and output, for baseline and DevSkills, would show how much of the 2.83 slope is shorter re-reading and how much is unflagged actions.

2. About subagents and re-reads - Table 2 shows cross-agent subsumed retrieval only in Claude Code, 50.15% of its subsumed retrieval (subagents return summaries, not the code, so the main agent reads it again). In my guides I advise giving heavy reading to subagents so the main conversation stays clean. Your paper shows both sides of it - the re-read is the price of delegation, and in Table 4, with CodeGraph, Haiku 4.5 subagent calls drop to zero. About the same tokens move to Sonnet 4.6 (3 times the token price) and cost goes up 8.30% on Verified-200 and 12.19% on Pro-100. I'd report the cost of cross-agent re-reads next to what delegation saved in the same trajectories to show the net effect.

3. On re-running tests - in my work every task needs a check the agent can run itself, otherwise "looks done" is its only signal, and the model sounds just as sure when it's wrong. Per Section 3.3, reruns hit 49.67 to 83.00% of tasks and cost up to 5.39% (causes: gaps in repo-specific test knowledge, truncated output, stalled progress). The Section 4.3 skills tell the agent to understand the test harness and rerun only when the code changes or new evidence calls for it. With them Claude Code lost 1.50 Pass@1 points on Verified-200. A rule that cuts reruns deserves its own check - were the lost tasks the ones where reruns were cut?

4. About the size of the skill set - seven general developer principles against 23 to 41 concrete, scenario-specific synthesized rules. Developer skills cut cost robustly in six of eight settings (7.88 to 41.73%), synthesized ones in three of eight (8.86 to 22.32%). Both sets are preloaded into the system prompt (Trace2Skill, appendix A.4), so they work like an always-on instruction file, and length and abstraction change together. And the developer skills are still concrete in form ("Reuse Context": check whether content is already in context before reading it). I keep my agent file short, up to about 200 lines, because a bloated file gets followed worse. My test - if I remove a line, does the agent start making mistakes? If not, it goes. In my experience concrete rules get followed more reliably than vague ones. An ablation holding length fixed (say, synthesized rules pruned to seven) would separate the two effects.

I couldn't find a statement on releasing the trajectories, the behavior detectors or the synthesized skill sets. That would let others run the Table 1 analysis on their own agents. Thanks for a careful, useful study.

Competing interests

Yes: the text cites the author's own technical reports (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22712985); no connection to the article's authors.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.