Skip to main content

Write a PREreview

AI Models Excel at Orchestration but Falter at Biological Judgment: Findings from an Agentic Gene Annotation Study

Posted
Server
bioRxiv
DOI
10.64898/2026.09.23.753765

Large language model (LLM) agents are increasingly used both to direct biological analyses and to interpret their results. The core functions of agents-workflow control and biological adjudication-are often combined within the same agent and evaluated end-to-end, making it difficult to determine whether a model that is useful in one role is also reliable in the other. In this study, Genome Skeptic, a bacterial gene-annotation framework was developed in which all sequence-level measurements are generated by deterministic bioinformatics tools, while an AI controller can request follow-up analyses only from a predefined registry. In a prospectively locked, blinded cohort of 20 genomes balanced for tet(A)/tet(B) presence and absence, GPT-5.6 Sol acting as the workflow controller achieved the same accuracy as fixed and exhaustive strategies (18/20 correct) while requiring 32% and 47% fewer follow-up analyses, respectively. This was consistent with a 60-case prospective benchmark, in which markedly different evidence-acquisition policies produced different final evidence states but identical endpoint calls in all 60 cases. When the same Sol model was instead given final decision authority over the exact evidence state used by the deterministic adjudicator, accuracy fell from 18/20 to 11/20 (exact McNemar P = 0.016), specificity fell from 1.00 to 0.30, and the model corrected none of the deterministic errors. All nine endpoint disagreements shifted negative calls to unresolved rather than identifying missed positives. The same model that efficiently decided what to analyse next performed worse when asked to decide what the evidence meant. In this task, LLM performed well in orchestration and not biological adjudication.

You can write a PREreview of AI Models Excel at Orchestration but Falter at Biological Judgment: Findings from an Agentic Gene Annotation Study. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now