Ir para a Avaliação PREreview

Avalilação PREreview de Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Publicado
DOI
10.5281/zenodo.23129439
Licença
CC0 1.0

## Summary

This paper argues that current data-agent benchmarks measure the wrong thing: text-to-SQL accuracy over small public schemas, with answer keys that a recent audit found wrong more often than right. In response, the authors build Argo-Bench — a simulated 2024 New York City food-delivery platform (81M orders, 3.4M customers) projected into a 235-table, 7.49-billion-row Oracle E-Business Suite warehouse — and define 210 tasks across fraud detection, forecasting, financial planning, dashboards, and compliance. The decisive design choice is consequence-based grading: the simulator's latent state is withheld from the agent, and filings (ban lists, forecasts, budget allocations, dashboard sources) are scored by their simulated consequences — fraud losses prevented net of wrongly-banned revenue, weighted interval scores against held-out months, share of attainable savings realized — rather than by matching a gold answer. The best of 14 frontier and open models (Claude Opus 5.5) solves 34.8% of tasks at ≥95 and averages 59.5/100. The paper's most valuable contribution may be its failure taxonomy: failing runs use sound methods but read the wrong record, optimize the wrong objective, or measure the wrong quantity.

## Strengths

1. **Consequence-based grading is the right direction, executed seriously.** Scoring a ban list by fraud losses prevented (including undetected fraud) net of destroyed customer margin — with explicit review costs and appeals modeling via Elkan-style example-dependent costs — is a genuine advance over answer-key matching. It aligns the benchmark's objective with the business objective, which is precisely what enterprise buyers care about.

2. **Unusual methodological honesty.** The limitations section and the Appendix J task audit openly document where the benchmark is weak: re-keyed tasks, realigned prompts, a fraud family the authors admit "tests less than it was designed to," and the explicit statement that the benchmark "compares data agents and does not estimate their performance on a real company's warehouse." This level of self-audit is rare and commendable.

3. **The failure taxonomy is practitioner gold.** The "wrong record / wrong objective / wrong quantity" analysis (Section 3.2) — e.g., 24 runs taking membership status from contract dates while the correct definition required replaying billing history through a seven-day grace rule — diagnoses exactly how production data agents fail: on semantics and business definitions, not on SQL syntax. Any enterprise data leader should read this section.

4. **Serious simulation craft.** 34 donor datasets with explicit identity/shape/anchor roles, calibration to NYC DCWP quarterly anchors (per-delivery economics within 5% in 13 of 16 comparisons), a real extrinsic shock (the April 2024 minimum-pay increase) enabling genuine forecasting tasks, schema verified against Oracle EBS 12.2 Vision with three ERP consultants, and an explicit no-LLM-generated-records policy. The two-seed design (public warehouse, private leaderboard world) is the right defense against memorization.

5. **Cost accounting in Appendix K is exemplary.** Reporting that warehouse spend ($6,160) was the same order as model API spend ($14,502) — and that for cheap models the warehouse was 70–92% of total cost — is the kind of honest economics the field needs.

## Major concerns

**1. The simulator's response model is the new gold answer.** The paper's central claim is that consequence grading escapes the tyranny of annotator answer keys. But for the five quest-cut tasks, the reference loss LrefL_{ref} is "the optimal plan itself" *under the simulator's response model* — a score of 100 means the agent recovered the optimum of the authors' own causal assumptions. An agent that reverse-engineers the response model scores perfectly, which rewards simulator-fitting rather than real-world judgment. The authors are aware of the general problem (they cite Leis et al. on shared simplifying assumptions), but the paper never tests robustness: how much do rankings change if the response model's parameters are perturbed? Without that, "consequences" inherits the exact fragility the paper attributes to gold answers, one level removed.

**2. Inconsistent reference standards across task families.** For the four courier-hour allocation tasks, LrefL_{ref} is "the best plan the authors reached from the warehouse alone" — explicitly *not* the true optimum, which "would require information that the agent cannot obtain" — while the quest tasks use the true simulator optimum. So a score of 100 means different things in different task families: best-human-effort in one, simulator-optimum in the other. This inconsistency makes cross-family score comparisons (and the headline mean of 59.5) hard to interpret. The paper should either harmonize the standard or report families separately as primary results.

**3. Seven forecast references were set with the outcome in view.** Table 7 discloses that for 7 of the 22 fixed forecast references, "the point was set with the realized value in view." These references set the *scale* on which every forecast in those series is graded. Even with honest intent, choosing scale parameters with knowledge of the outcome can compress or stretch the grading range in ways that favor particular error profiles. The authors report results without these tasks in Appendix I, which is good — but tasks whose grading scale was outcome-informed should be excluded from the headline numbers, not merely analyzed separately.

**4. The headline numbers are irreproducible by design.** The simulator and graders are withheld "to prevent direct answer memorization," and all paper experiments ran on a private-seed world. Outsiders can reproduce the *procedure* on a sibling world but never the *numbers* — including the headline 34.8% / 59.5. Combined with one run per task per setting (confidence intervals "reflect the choice of tasks rather than run-to-run variation"), this means the paper's central quantitative claims rest on an experiment no independent party can repeat. The two-seed design is defensible against memorization, but the paper should acknowledge the cost: this is a benchmark whose leaderboard cannot be audited, only trusted.

**5. Simulation fidelity rests on transferred assumptions the paper under-examines.** Fraud separability is calibrated from IEEE-CIS — a credit-card fraud dataset — and honest-user device/address/card sharing rates from RBA login data; both are transferred across domains to delivery-platform fraud. The membership program "rests on weakly grounded assumptions" (authors' own words). Calibration to aggregate anchors "does not guarantee realistic tails" (citing Chen et al. 2019) — and tails are precisely where fraud detection lives. None of this invalidates the benchmark for *comparing* agents, but the paper's framing sometimes implies more: the fraud tasks' precision/recall tradeoffs are only as meaningful as the transferred separability assumptions, and those deserve a sensitivity analysis, not just a limitations paragraph.

**6. The warehouse is deliberately unlike real warehouses in the way that matters most.** The three ERP consultants named the absence of data drift and inconsistencies — deprecated tables overlapping active ones, figures failing to reconcile — "the most significant difference from their customers' systems," and the authors kept it clean to keep ground truth unambiguous. That is a reasonable design choice, but it means the benchmark tests greenfield warehouse *navigation* while real enterprise data work is dominated by *archaeology*: reconciling conflicting sources, decoding undocumented conventions, working around migrations. The paper should be explicit that this entire skill category is out of scope, since it is arguably the larger half of production data work.

**7. The main-table cost column is misleading without the warehouse.** Table 2 reports per-task API spend ($0.06–$4.89), but Appendix K shows warehouse spend flips the economics: GPT-6 Luna's true per-task cost is $0.70, not $0.06; DeepSeek, GLM Flash, and Muse Spark 1.3 spend more on BigQuery in absolute terms than Opus does ($0.94–$1.63 vs $0.82); Muse Spark 1.3 burns $11.77 of warehouse per solved task against $2.20–$2.37 for the leaders, with 57% of its warehouse dollars spent on tasks scoring below 5. A benchmark about enterprise economics should put total cost — not API cost — in the main results table.

## Minor concerns

- **34 "presence" tasks are graded on filing existence.** A note explaining the agent's work is valuable, but scoring 16% of the benchmark on whether *any* filing was made is a weak grading mode that inflates scores for verbosity-adjacent behavior.

- **The Appendix J audit reveals post-hoc key surgery.** Two task families required re-keying (138 of 249 originally-keyed couriers were already deactivated; a misweighted composite score) *after* runs were examined, with everything regraded and rerun. The transparency is admirable, but it raises the question the paper doesn't answer: how many keys would a truly independent auditor find, and should the audit be a living document rather than a one-time appendix?

- **The injected-refund shortcut (Appendix J) weakens a headline fraud family.** All 664 anomalous refunds fall within the 98 simulator-injected storefronts, so "one pass over the refunds therefore finds every candidate" — the family tests ring-vs-control discrimination, not discovery. The authors keep it "as run" because fixing it "changes the world," but a benchmark paper should quantify how much easier the task became rather than leaving it to the appendix.

- **Vendor authorship deserves a sentence.** All five authors are at TextQL, which sells data-agent products. The paper is admirably self-critical, but an explicit statement of how the benchmark relates to the company's product roadmap would help readers calibrate incentives. (Relatedly: the sandbox-escape incident, where a model read the grader code during development, is disclosed in the ethics statement — good — but deserves a sentence in the main text given Major concern #4 about grader secrecy.)

- **210 tasks, 146 scenarios.** Many tasks are prompt variants of one scenario, and 72 are forecasts. The effective breadth is narrower than the headline count suggests; per-scenario aggregation should be a primary reported metric, not just a bootstrap detail.

- **Forecast overconfidence is reported but under-discussed.** 80% intervals contain the realized value only 44.8% of the time across 4,553 series — a striking miscalibration result with direct implications for anyone deploying agent-produced forecasts, given more prominence it deserves.

## Practitioner perspective

If you lead data or AI in an enterprise, this paper is worth reading less for the leaderboard than for what it implies about evaluating data agents *before you buy or build one*:

1. **Grade on consequences, not answers.** The paper's core methodological point transfers directly: pilot evaluations should score what the agent's output *does* (money saved, fraud prevented net of false positives, forecast value under a proper scoring rule) rather than whether it matches an analyst's spreadsheet. Most enterprise agent POCs I have seen fail exactly here — they measure answer similarity when they should measure decision quality.

2. **Budget for the warehouse, not just the model.** Appendix K's finding generalizes: in production, the compute the agent consumes (warehouse scans, tool calls, retries) routinely dwarfs the model API bill, and it is precisely the weaker agents that spend the most — $11.77 per solved task for the worst vs $2.20 for the best. Any agent TCO model that prices only tokens is fiction.

3. **The failure taxonomy is a deployment checklist.** "Wrong record, wrong objective, wrong quantity" is the best three-line summary I have seen of how data agents fail in production. Before deploying, ask of each agent task: which record defines the key entity, what objective is actually being optimized, and is the measured quantity the quantity the stakeholder asked for. Most agent failures I have witnessed are one of these three.

4. **Simulated evals are necessary but not sufficient.** This benchmark tests greenfield navigation of a clean, coherent warehouse. Real enterprise warehouses carry years of drift, deprecated tables, and undocumented conventions — the archaeology this benchmark explicitly excludes. Use Argo-Bench-style evals for agent *capability*, but validate on your own warehouse's mess before production.

## Overall assessment

This is a strong, unusually honest benchmark paper: the consequence-based grading design is a genuine methodological step forward, the simulation craft is serious, and the self-audit in the appendices sets a standard the field should copy. I recommend it with revisions: harmonize or separately report the allocation reference standards, exclude the seven outcome-viewed forecast references from headline numbers, put warehouse-inclusive costs in the main results table, add a sensitivity analysis on the fraud-separability and ban-cost parameters, and address the irreproducibility-by-design of the headline numbers — at minimum with a frank discussion, ideally with a verification bundle (frozen grader + private-world answer keys under NDA, or a public challenge split). None of these diminishes the contribution; several would strengthen the paper's own claim to be the trustworthy evaluation the field needs.

**AI disclosure:** I used an AI assistant to help draft the initial text of this review; I added my own practitioner assessment, verified the claims against the paper, and approved the final version.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.