PREreview of When Does a Second Model Help? Cross-Model Review in LLM Verification
- Published
- DOI
- 10.5281/zenodo.23132634
- License
- CC BY 4.0
Thank you for this paper. You took 30 Korean-language artifacts (code modules, tutorials and presentation scripts, generated by Claude Opus 4.6) with 150 planted errors, and ran 900 review sessions in 10 conditions. The top cross-model reviewer (GPT-5.4) is not significantly different from same-model review in a fresh session, and you note this is not equivalence. At two review calls, one same-model plus one cross-model review matches 56.7% of planted errors against 42.7% for two same-model reviews, but not significantly more than two top-tier cross-model reviews. I liked the audit of session records before the analysis. One baseline run was dropped because most of its findings appear verbatim in the script that wrote the files, and all-session values are still reported. And the paper says plainly what each test does not show.
I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.
1. On the reviewer's task - each reviewer returns a structured list of findings (Section 4.1, Procedure), but the review prompt itself is not quoted. Could you publish it? In my guides I write that a second, fresh agent told to disprove the work ("treat it as wrong until shown otherwise") finds what the first one missed. The two approaches overlap only partly (Jaccard 41.2%, Section 5.2). Could you add one condition where the same-model fresh reviewer gets that adversarial instruction, at the same number of calls? Then we would see how much of the cross-model gain a change of stance alone gives.
2. On the false positives - in Section 5.8, 187 of 1,995 cross-model false positives were classified as model bias, and in 136 of them the reviewer flagged generator-specific features or tools as nonexistent. The classification is keyword-based and not validated. Also, Appendix A counts a finding that correctly identifies a pre-existing (not planted) defect as a false positive too. In my guides I write that a model sounds equally sure when it is right and when it is wrong. Two suggestions. Report the severity label (Section 4.1) separately for true and false positives, so readers can see whether it helps a human decide which findings to check first. I expect it may not help much. And check a random sample of false positives by hand, to see how many are real pre-existing defects.
3. On who wrote the answer key - according to Appendix A, the errors were injected by the generator model itself (Claude Opus 4.6), and the ground-truth records were written at injection time. The same model did the same-model reviews (Section 3.2). In Section 5.4, CCR has F1 40.7% against 37.2% for XMR-GPT on code, and your conjectured explanation is a shared understanding of programming patterns. I'd add a second reading. The advantage may partly come from the same model having written the planted errors. In my own experiment (36 runs, DOI 10.5281/zenodo.22759217) the agent wrote its own tests, and the gate returned 4,086 tests and 0 failures. The agent wrote the criterion and then met it itself, so all-green told me nothing. A small set of artifacts with errors injected by a different model or by a person would separate the two explanations.
4. On cost and open data - Section 2.5 matches review calls, not compute. In my measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216), more than 85% of modeled cost was context work (cache reads plus cache writes). So I'd add input and output tokens per review session for each condition. Information restriction also changes the input, so then Section 6 can be read as cost per found error, not per call. The Limitations also say CCR runs differ widely (56, 37 and 29 matched errors) and the headline pairing is among the higher run combinations. In my own experiment I also chose the comparison point after the runs, and said so in the report. My run data is public on Hugging Face (DOI 10.57967/hf/10366). I'd suggest posting the per-session records publicly with a DOI instead of on request, so the set-level numbers can be recomputed by anyone.
Thank you for a careful paper.
Competing interests
Yes: the text cites the author's own technical reports and dataset (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.57967/hf/10366); no connection to the preprint's author.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.