AI Models Excel at Orchestration but Falter at Biological Judgment: Findings from an Agentic Gene Annotation Study
- Publicado
- Servidor
- bioRxiv
- DOI
- 10.64898/2026.09.23.753765
Large language model (LLM) agents are increasingly used both to direct biological analyses and to interpret their results. The core functions of agents-workflow control and biological adjudication-are often combined within the same agent and evaluated end-to-end, making it difficult to determine whether a model that is useful in one role is also reliable in the other. In this study, Genome Skeptic, a bacterial gene-annotation framework was developed in which all sequence-level measurements are generated by deterministic bioinformatics tools, while an AI controller can request follow-up analyses only from a predefined registry. In a prospectively locked, blinded cohort of 20 genomes balanced for tet(A)/tet(B) presence and absence, GPT-5.6 Sol acting as the workflow controller achieved the same accuracy as fixed and exhaustive strategies (18/20 correct) while requiring 32% and 47% fewer follow-up analyses, respectively. This was consistent with a 60-case prospective benchmark, in which markedly different evidence-acquisition policies produced different final evidence states but identical endpoint calls in all 60 cases. When the same Sol model was instead given final decision authority over the exact evidence state used by the deterministic adjudicator, accuracy fell from 18/20 to 11/20 (exact McNemar P = 0.016), specificity fell from 1.00 to 0.30, and the model corrected none of the deterministic errors. All nine endpoint disagreements shifted negative calls to unresolved rather than identifying missed positives. The same model that efficiently decided what to analyse next performed worse when asked to decide what the evidence meant. In this task, LLM performed well in orchestration and not biological adjudication.