Comentarios
Escribir un comentarioNo se han publicado comentarios aún.
Regarding your paper - you built a framework to check how terminal agents verify their own work. You take the first complete candidate solution in each trajectory, replay it in a fresh environment, run the official evaluator and compare that with the agent's own check (Section 2.2). Ten agents on TerminalBench2.1, three runs each. Agents verify in 99.53% of eligible trajectories, but only 61.43% of incorrect candidates are detected and 49.36% of detected errors are repaired. Then SCVD - the student makes the candidate, a stronger teacher (GLM-5.2) verifies and repairs from the same state, and only that continuation is trained on. I liked that replaying the candidate separates "the check said fine" from "the solution was fine".
I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. My four comments are about what the agent checks with, the "no error" signal, the LLM judges, and cost.
1. On what the agent checks with - Section 2.1 lists self-generated tests, inspecting artifacts and querying system state, but Table 5 gives error detection per model, not per kind of check. Section 5 cites studies saying agent-generated tests often give weak evidence. In my own experiment (36 coding-agent runs, DOI 10.5281/zenodo.22759217) the test gate returned 4,086 tests and 0 failures, while cost differed by 33.5% between conditions. The agent wrote the tests itself, so tests per run went from 97 to 127. It wrote the criterion and then met it. I wrote about this on Qeios (DOI 10.32388/0BV3Z8). I suggest splitting the error detection rate and the "no error" signal by what the agent checked with - its own tests, tests or files already in the task, compiler or runtime output, or just reading files. Then readers can see whether missed errors pile up where the agent graded its own work.
2. On the "no error" signal - in Finding 2 a passing check means a correct candidate only 51.52% of the time on average (VPR). In my guides I write that a model sounds equally sure when it is right and when it is wrong. I also write that a second, fresh agent told to disprove the work ("treat it as wrong until shown otherwise") finds what the first one missed, because independent agents rarely make the same mistake. As a baseline without any training, I would hand the student's candidate to a separate fresh agent whose only task is to show it is wrong, and report its detection rate next to the agent's own. This would show how much of the SCVD gain could come from a second pair of eyes.
3. On the LLM judges - in Appendix B.1 three judges (DeepSeek-V4-Pro-0813, GLM-5.2, Kimi-K3) label the candidate boundary and the verification outcome, calibrated on 20 human-reviewed trajectories. Three-way exact agreement is 74.29% for the boundary and 85.16% for the verification semantics, and 106 of 2,669 trajectories needed the tie-breaker. In my guides I write that when independent agents disagree, the disagreement points to the weak spot - and a round of checking lowers risk but does not remove it. So I would report the diagnostic metrics separately for trajectories where all three judges agreed and where they did not, and enlarge the human-checked set beyond 20 (small next to 2,669).
4. On cost - in Section 4.5 and Table 7 SCVD cuts agent turns by 12.0-35.6% and total tokens by 0.3-22.4% versus Base, while generated tokens grow 2.6-4.1 times. I agree that fewer rounds should mean less re-reading. In my measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216) more than 85% of modeled cost was context work (cache reads plus cache writes). In my experiment on 168 sessions 94.4% of paid tokens were re-reading context already sent (contextburn, DOI 10.5281/zenodo.22712985). Table 7 does not separate cached from fresh input and gives no price, and providers price those and output differently. So I would split input into cached and fresh and add a dollar cost under one public price list - one token total can hide which way the cost moves.
Thank you for a clear and useful paper.
Yes: the text cites the author's own technical reports and article (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985, 10.32388/0BV3Z8); no connection to the article's authors.
The author declares that they did not use generative AI to come up with new ideas for their review.
No se han publicado comentarios aún.