Saltar a PREreview

PREreview del TokenCast: Forecasting Token Consumption During LLM Agent Execution

Publicado
DOI
10.5281/zenodo.23037648
Licencia
CC BY 4.0

The authors claim that TokenCast forecasts token use while an agent is running, and the forecast is refreshed with no extra LLM calls. MAE is 14.5% lower than the strongest comparator on average over 96 combinations. I liked that the paper says openly where it loses (24 combinations, all at Task Start and Call Start). The calibrated intervals (82.0% coverage at Task Start against 52.7% for Self-Prediction) and the ablations are also useful.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents and I measure what they cost me, so maybe my practical side is useful here. My four comments are about tokens versus money, the expensive tail, context compaction, and the budget replay.

1. On tokens and money - In Section 3.1 the target is provider-accounted input and output tokens. For GPT-5.4 on SWE-bench Verified the input is 99% of the tokens (Appendix C.1). But the paper counts a cached input token, a cache write token and a fresh input token the same. I measured 722 agent sessions (technical report, DOI 10.5281/zenodo.22759216) and more than 85% of the modeled cost was context work - cache reads plus cache writes. Providers price those differently from fresh input. I also wrote an IETF Internet-Draft (draft-arsentev-agent-run-metrics-00) with a common format for agent run cost records, because vendors report run usage in incompatible forms (an individual draft, not a standard). I suggest splitting the input into cache read, cache write and uncached input, in the segment representation or at least in the reported traces. Then the forecast can be read in money, not only in tokens.

2. On the expensive tail - In my own 722 agent sessions about 3% of sessions produced 80% of the cost. Sessions longer than 200 calls (8% of sessions) produced more than 90% of the money. In your Table 6 the median SWE-bench Verified run has about 20 calls, and runs are capped at 500 calls (Appendix C.1). The metrics give every task equal weight (Section 4.1), and Appendix D.4.4 already tests the longest band of tasks (MAE rises by at most 7.8%). A budget decision matters most on the few runs that eat the budget. So please also report the error (for example WAPE) on the most expensive runs, such as the top 10% by consumption.

3. On compaction - in Appendix C.1 the context compaction, result pruning and subagents are turned off in DeepSeek Harness, and in OpenHands the condenser is disabled. Eq. (1) composes segments through the net input-length change g, so a segment where context shrinks would have a negative g. In my experiment on 168 sessions 94.4% of paid tokens were re-reading context that had already been sent (contextburn, DOI 10.5281/zenodo.22712985). In my 36-run experiment on clearing the agent's context the cost followed a U-shaped curve (the minimum was at clearing every 3 tasks, and clearing after every task was about a third more expensive, DOI 10.5281/zenodo.22759217). I suggest to add a set of runs with the condenser or compaction switched on, so the composition identity is tested on segments with negative g (as many production agents run).

4. On the budget replay - in Section 4.3 you report 21.3% fewer tokens (288 GPT-5.4 runs, seven budgets) while matching fixed-budget trace completion. A run is trace-complete when it reaches its recorded terminal state, and you say the replay does not measure task resolution. A run can also end without submitting or by hitting a cap (Appendix C.1), so trace completion counts a failed run the same as a solved one. I think part of the 21.3% could come from stopping runs that would have solved the task. In my work every task needs a check the agent can run itself, otherwise "looks done" is the only signal. The patches are already scored, so please report the resolved rate of the replayed runs next to trace completion (for TokenCast and for the fixed budget).

The code is public, but I could not find a statement in the preprint on releasing the 11,712 execution traces. Releasing them with per-request usage would let others test their forecasters on the same data. Thank you for a careful and useful study.

Competing interests

Yes: the text cites the author's own technical reports (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985) and his individual IETF Internet-Draft draft-arsentev-agent-run-metrics-00; no connection to the preprint's authors.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.