Comments
Write a commentNo comments have been published yet.
This paper looks at what happens when we trust an AI coding agent’s work without independently checking what actually changed. The authors show a case where a protected file was changed, but GTB still returned VERIFIED because the file was left out of the verification process.
I thought one of the strongest parts of the paper was the distinction between stopping an agent from doing something and actually being able to prove that it didn’t do it. The paper also does a good job of showing the attack in practice, explaining why the verification failed, and suggesting ways to make the process more reliable.
The experiment focuses on GTB, so testing other coding-agent or verification systems would help show whether the issue is specific to GTB or appears more broadly.
The proposed fixes have not yet been independently tested. Testing them against the same attack, along with a few variations, would make the results more convincing.
Including the exact software versions and experimental environment would make the experiment easier to reproduce.
A simple diagram showing the normal verification path and where the modified file gets excluded would make the main failure easier to understand.
It would also be useful to state how many times the attack was repeated and whether the same result occurred consistently
The author declares that they have no competing interests.
The author declares that they did not use generative AI to come up with new ideas for their review.
No comments have been published yet.