PREreview del Discriminating Fixture Coverage in Agent-Infrastructure Verification Suites
- Publicado
- DOI
- 10.5281/zenodo.23159191
- Licencia
- CC BY 4.0
Thank you for this paper. You take a client-side projection layer for multi-session agent chat, write a reference implementation, an all-defects variant and first-order mutants, and ask what the usual validation of an invariant suite is worth. The eleven-check suite passes the reference and fails the all-defects variant on eight checks, and still a mutant without event-identity deduplication survives all eleven (Section 5.1). After you froze and hashed the twelve-check suite, it killed 5 of 10 mutants specified by an outside reader (Section 5.2). Three survivors were never activated and two were masked at the oracle (Section 5.3). Then seven input dimensions were registered in advance, and one fixture per uncovered dimension killed all five, with the oracles reused unchanged (Section 5.4). I liked that you keep 5/10 and 10/10 apart and call the second one a repair result, not a held-out estimate.
I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.
1. On what the report records - Section 6 says a suite that has never fired is compatible with a suite that cannot fire, and that a pass/fail report can't tell the two failure modes apart. I wrote about the same problem on Qeios (DOI 10.32388/0BV3Z8). A set of checks where each check could fail can still prove nothing, if no check changes state between the conditions being compared. That is a property of the set, not of one check, so a field of a single check can't express it. Your Section 5.2 result on invariant families (H1 caught only by C-08, a convergence check) points the same way. Could the paper suggest what a suite report should carry next to pass/fail? For example the post-freeze kill rate with the suite hash, and for each check whether any fixture activated a difference at all. I edit v0.1 of a conformance reporting format in a W3C community group, and a short list like this from you would help that kind of work.
2. On who wrote the checks and the code - Section 1 says the authors wrote the reference, the mutants and the suite, and the Scope paragraph (Section 6) names an implementation the authors did not write as a next step. In my own experiment (36 coding-agent runs, DOI 10.5281/zenodo.22759217) the test gate returned 4,086 tests and 0 failures, while cost differed by 33.5% between conditions. The agent wrote its own tests (97 to 127 per run). So it wrote the criterion and then met it, and the green result told me nothing about the difference. I'd add one sentence on whether any of the reference, the fixtures or the checks were produced with an AI coding tool. If yes, the extension to an implementation you did not write matters even more.
3. On the outside reader - Section 5.2 says the ten mutants came from an adversarial reader who saw the system description and the check titles, but not the fixtures. Was this a person or a model, and what instruction did the reader get? In my guides I write that a second, fresh agent told to disprove the work ("treat it as wrong until shown otherwise") finds what the first one missed, because independent agents rarely make the same mistake. If a fresh agent with that instruction can play the reader, the second challenge round you plan could be repeated after every change of the suite, with the number of new survivors per round reported.
4. On the artifact - the abstract says you give the artifact, with the frozen hash, the registered predictions, the mutants and the run logs. In the arXiv version I didn't find a link or the hash value itself. Could you print both in the paper, ideally with a DOI? My own run data is public on Hugging Face (DOI 10.57967/hf/10366), so anyone can recompute it, and I think a frozen hash is only checkable if a reader can see it next to the 5/10 number.
Thank you for a careful paper.
Competing interests
Yes: the review cites the author's own technical reports and posts (DOI 10.5281/zenodo.22759217, 10.32388/0BV3Z8, 10.57967/hf/10366). The author edits v0.1 of a conformance reporting format in the W3C Agent Conformance and Benchmarking Community Group. No connection to the preprint's authors.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.