Aller directement à la PREreview

PREreview de Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery

Publié
DOI
10.5281/zenodo.23257662
Licence
CC0 1.0

Thank you for this paper. You open-coded 200 CyberGym traces from four agents, each run as a baseline and with four existing efficiency methods (Section 2). You found that code localization and understanding plus vulnerability reasoning and trigger design take 60.4% of tokens and 68.7% of the attributed failure weight (Section 3.2). Only 39 of 160 matched pairs (24.4%) keep success at a lower total cost (Section 3.3). You then built AVRI, an interface around a persistent Bidirectional Evidence Trace, which on 20 separate tasks lowers total cost by 18.0% for Codex and 23.7% for OpenCode with success unchanged (Section 5, Table 3). I liked that the auxiliary model's bill is counted inside each method's cost (Figure 5).

I read it as a practitioner who runs coding agents every day, not as a security researcher. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.

1. On what a token share means - Section 3.2 says the stage shares "measure billed usage during each activity, including historical context, and do not directly measure redundant reading or new reasoning". A late turn labeled as reasoning still pays for all the code read earlier, so the 43.1% for stage L may say more about when a stage happens than about what it costs. In my own 722 sessions (DOI 10.5281/zenodo.22759216) more than 85% of the modeled cost was cache reads plus cache writes, though these ran on a subscription, so the cost is modeled from a price list, not an invoice. In 168 other sessions, 94.4% of billed tokens were re-reading context already sent (DOI 10.5281/zenodo.22712985), a share of tokens, not of money. Could you split each stage in Figure 1 into fresh input, cached input and output, and show the same figure in USD?

2. On cutting history - Finding 1 says simple truncation in Cybench and EnIGMA "may discard useful context", and Section 3.1 adds that the aggregate results do not isolate this effect. Also "context recovery" was assigned to only two of 160 pairs (Sections 2.5 and 3.3). In my own small experiment, 36 coding-agent runs, clearing the context after every task cost about a third more (+33.5%) than clearing every three tasks (DOI 10.5281/zenodo.22759217). The comparison point was chosen after the runs, the agent wrote more tests in that arm, and the tasks were ordinary coding, so it may not transfer. How many truncations or compactions happened per trace for each agent? And in the Table 3 runs, how many times was saved BET evidence actually recovered after a compaction?

3. On runs that never tested anything - 38 traces have no observed PoC execution feedback, and all of them are unsuccessful (Section 3.2). That is more than half of the 67 failed traces, yet in Figure 2 "validation and feedback iteration" and "task goals and stopping" get a small share of the failure weight. In my guides I write that every task needs its own check the agent can run. Without it, "looks done" is the only signal the agent has, and a model sounds equally sure when it is right and when it is wrong. So I liked that BET keeps a hypothesis apart from an observation ("a static derivation is not an observed crash", Section 4). For those 38 traces, how did the run end - the 7,200-second timeout, the agent stopping on its own, or the environment?

4. On single runs - you state the limit yourselves. Section 6 says one run per configuration and task does not establish robustness to run-to-run variation. For OpenCode, Table 3 gives $4.38 for the baseline against $3.34 for AVRI, and a difference of about one dollar over 20 tasks is hard to read without repeats. The Table 3 caption says "One selected run per configuration". What does "selected" mean here - was there more than one run per cell? Could you repeat the baseline and AVRI cells a few times on a subset of tasks, and deposit the artifact with a DOI? In my own 36-run experiment each policy had six repeats, and the run data is public on Hugging Face (DOI 10.57967/hf/10366).

Thank you for a clear and useful paper.

Competing interests

The review cites the reviewer's own technical reports (DOI 10.5281/zenodo.22759216, DOI 10.5281/zenodo.22759217, DOI 10.5281/zenodo.22712985) and the reviewer's own public dataset (DOI 10.57967/hf/10366). No connection to the preprint's authors.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.