Skip to PREreview

PREreview of Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent

Published
DOI
10.5281/zenodo.23197406
License
CC0 1.0

## Summary

This paper studies Leni, a production enterprise AI business-analyst agent, and asks a question most vendor papers avoid: where does the system's measured reliability actually come from? Rather than reporting a single headline score, the author decomposes the uplift over the bare frontier base model across three public benchmarks chosen to stress different failure modes — SpreadsheetBench Verified (silent computation error), BullshitBench v2 (premise confabulation), and the GAIA validation split (cascade error over long tool chains). The full system improves by +11.0 percentage points on SpreadsheetBench (91.25% vs 80.25%, n=400, p<0.001), +7 to +10 points on BullshitBench (n=100), and roughly +15 points on GAIA validation (75.2% pass@1, n=165).

The central finding is a decomposition: most of the uplift comes from scaffolding, routing, and specialist models rather than from the verification loop itself, whose isolated contribution is small (+1.5 points on SpreadsheetBench) but positionally decisive — it converts tasks at the top of the score distribution, the difference between mid-leaderboard and near-top. The author instruments the deterministic loop end to end, producing an empirical verifier confusion matrix (catch rate about 0.20, fix rate 0.75, no false-alarm regressions), and formalizes a compounding-reliability model with imperfect verification. Specialist-swap ablations suggest the observer matters as much as the loop: replacing the small post-trained verifier with the generating frontier model eliminates most rescues. A valid-premise control (100 expert-level questions, zero over-rejections) bounds the firewall's false-positive rate.

## Strengths

1. Honesty as a feature. The paper states which cross-system comparisons are statistically resolvable and which are not, and — remarkably — corrects its own company's previously published GAIA headline after re-grading every stored trajectory with the official scorer. That kind of self-correction in a vendor paper is rare and builds trust.

2. The decomposition is the right question for practitioners. Anyone who has shipped an agent knows the leaderboard number tells you nothing about what to build. Breaking the +11 points into scaffolding (+9.5) versus verification (+1.5) tells an engineering team where to spend effort.

3. The independent-observer finding rings true. My experience with self-critique in production matches this: a model checking its own work mostly agrees with itself. The specialist-swap result (rescues drop from 6 tasks to 2 when the frontier model verifies its own artifact) is the paper's most actionable insight.

4. A usable reliability model. Equation (1), p' = p(1−fb) + (1−p)cr, makes the verifier's error profile a first-class quantity. The condition under which a loop *reduces* reliability ((1−p)cr > pfb) is exactly the kind of guardrail a production deployment needs.

## Major comments

1. This is a vendor evaluating its own system. The limitations section says so, which I respect — but the decomposition rests on one production system and one team's scaffolding. Independent replication on a different stack is needed before treating "most uplift comes from scaffolding, not verification" as a general result.

2. No cost or latency accounting. Verification loops, specialist models, and multi-pass scaffolding multiply inference cost per task, yet the paper never reports dollars-per-task, token budgets, or latency. A production reader cannot decide whether the +1.5 points from the loop is worth the compute — the biggest gap in an otherwise production-minded paper.

3. GAIA validation split only (n=165). The author is candid about the corrected score, but the corrected pass@1 still sits on the validation split with a hidden-test-set submission listed as future work. The decomposition claims should be flagged as preliminary on GAIA.

4. Catch rate 0.20 deserves more attention. Four out of five true errors slip through the deterministic loop. The compounding model is elegant, but its sensitivity to the (c, r, f) estimates is never explored — if the catch rate drops to 0.10 on a different domain, does the loop still pay for itself?

## Minor comments

1. BullshitBench is a community benchmark (n=100) — the paper weights conclusions accordingly, which is appropriate.

2. The judge panel spans three providers but two configurations are Claude-based with a Claude judge; the paper flags judge bias, but a sensitivity check with the Claude judge excluded would strengthen this.

## Overall assessment

Recommend with revisions. The decomposition, the honest uncertainty accounting, and the independent-observer result are genuinely valuable to anyone building production agents. I would ask for per-task cost/latency reporting and a clearer flag on the GAIA preliminary status before treating this as settled.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.