Skip to PREreview

PREreview of A Case Study in Assuring AI-Written Software

Published
DOI
10.5281/zenodo.23207006
License
CC BY 4.0

Thank you for this paper. It is a retrospective case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training, using two structured interviews and six months of project records (Section 2). Section 3 finds asymmetric delegation: she delegated implementation and much of the technical review, but not the authority over what becomes durable project context and what enters production. Section 4 finds that the checks became another fallible software layer, for example backup monitors that did not establish that a usable backup existed, and an audit that stopped running without producing a failure. I liked that the paper reports the failures of its own checks instead of only the wins (Section 4).

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so the operator's position is close to mine and maybe my practical side is useful here. Four comments.

1. On the fresh reviewer - Section 2 says each change was "reviewed by a separate agent in a fresh context", but it doesn't say what that reviewer was asked to do. Was the brief to confirm the change, or to try to break it? In my guides I write that a second, fresh agent finds what the first one missed when its brief is to disprove, meaning it treats the work as wrong until proven otherwise. A round like this lowers the risk but doesn't remove it, which matches Section 3 ("adding another reviewer moved rather than removed part of the oversight problem"). For the multi-model council, was disagreement between the models recorded and used as a signal of where to look, or did the operator only read the recommendations?

2. On tests that pass - In Section 4 the operator recreates targeted failures and confirms that the corresponding checks fail, which the paper links to mutation testing. That is a good check for each test. My question is about the suite. In my Qeios article "A Suite That Does Not Discriminate" (DOI 10.32388/0BV3Z8, a preprint with open review) I argue that a suite can consist of honest checks, each of which could fail, and still prove nothing if no check changed state between the conditions being compared. A per-check field cannot express that. In my experiment (36 coding-agent runs, DOI 10.5281/zenodo.22759217) the gate returned 4,086 tests and 0 failures, while cost differed by 33.5% between the cheapest and the most expensive policy. The agent wrote the tests itself, so it wrote the criterion and then met it. Was there a suite-level record of which checks ever changed their verdict over the six months, and which never did? And who wrote the seeded failures: the same agent that wrote the tests, a fresh one, or the operator herself?

3. On the reassuring signal - Section 4 says that "In each case, a reassuring signal supported a stronger claim than the evidence justified." In my guides on arsentev.ai I write that an AI model sounds equally sure when it is right and when it is wrong, and that every task needs its own check that the agent can run. So for me the four concerns in Section 4 read as one thing - a check needs its own check on the outcome, not on the event. For each of the four incidents, could the paper say how it was found (by her, by an agent, by accident) and for how long the signal had stayed green before that? Also, was the expected number of checks for the pass-rate report written down anywhere before the run?

4. On the rulebook - In Section 2 the primary coding agent re-reads a versioned rulebook, task and issue lists and skill files at the start of each session, and agents can propose lessons that get approved and stored. In my guides I write that the instruction file (CLAUDE.md or AGENTS.md) should stay short, up to about 200 lines, because a bloated file is followed worse. In my practice, concrete rules are followed more reliably than vague ones, and what is needed only now and then goes into skills. How large did the rulebook get over six months? Were lessons ever removed or merged, or only added? Did she see old rules being skipped as the file grew?

Thank you for writing this case up.

Competing interests

The review cites the author's own work (DOI 10.32388/0BV3Z8 and DOI 10.5281/zenodo.22759217). The author has no connection to the preprint's authors.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.