Skip to PREreview

PREreview of Orchestrating Intelligence: Confidence-Aware Routing for Efficient Multi-Agent Collaboration across Multi-Scale Models

Published
DOI
10.5281/zenodo.23197278
License
CC0 1.0

## Summary

OI-MAS tackles a cost problem anyone running multi-agent systems feels immediately: existing frameworks deploy one large LLM uniformly across every agent role, so even trivial subtasks burn frontier-model tokens. The paper's "conductor" jointly decides, per reasoning turn, which agent roles to activate (Role Router) and which backbone from a heterogeneous pool (Qwen2.5-3B/7B, Llama3.1-8B, Llama3.1-70B) each role gets (Model Router), optimizing an RL objective that weights a cost penalty by a confidence signal — average token log-probability of the generated sequence. High confidence enforces a stronger cost penalty (stay small); low confidence relaxes it (escalate).

On GSM8K, MATH, MedQA, GPQA, and MBPP, OI-MAS beats multi-agent baselines by up to 12.88% accuracy while cutting cost 17–79%. An out-of-distribution test (policy trained on MBPP, evaluated on HumanEval) reaches 91.46% pass@1 at less than half the next-best baseline's cost. Wall-clock latency on GPQA is 23.12s versus 36–39s for baselines. The routing-behavior analysis is nice: model selection shifts toward larger backbones with MATH difficulty level, generative roles (Generator, GeneratorCoT) escalate more than post-processing roles, and the Refiner shows a bimodal easy/hard split. The limitations section is admirably honest about memory, concurrency, and safety gaps.

## Strengths

1. Joint role-model routing is the right framing. Prior work routes either agents (who acts) or models (how much power), but never both per step. The two-stage conductor — plan the functionality, then size the capacity — is a clean decomposition that maps directly onto how production teams think about agent budgets.

2. The routing-behavior analysis is genuinely informative. The role-level capacity regimes (generators escalate, ensemblers don't, refiners are bimodal) give practitioners a prior for where to spend model budget. This is more useful than the headline accuracy number.

3. Ablations are well designed. Removing the model router, the cost term, and the confidence weighting separately shows each component earns its place — the confidence term in particular, since dropping it degrades accuracy on both MedQA and MBPP.

4. Honest limitations. The authors state plainly that safety is not a design objective and that the cost balance may not hold under highly concurrent large-scale deployment. More papers should do this.

## Major comments

1. The "cost" is synthetic, not measured. Inference cost comes from a token-based pricing scheme in Appendix C, while all experiments run on owned A100s with vLLM. On a real API bill — or a real self-hosted fleet — the numbers shift, and the conductor's own inference overhead (two routing networks plus confidence computation at every turn) does not appear in the accounting at all. For a paper whose headline claim is cost reduction, measured dollars or measured token counts including router overhead are needed.

2. Confidence is average log-probability at temperature 0, with no calibration check. The entire confidence-aware mechanism rests on this signal, yet there is no analysis of whether low-confidence states are actually the states where escalation helps, versus merely high-entropy outputs. The "two traps" framing in the introduction (overconfident small-model use causing mission failure) is never empirically tested.

3. No tool-use benchmarks. The motivation is production multi-agent systems, but evaluation is static Q&A and math — no tool calls, no retries, no growing context. In real agent deployments, cost is dominated by tool-call loops and context accumulation, exactly the dynamics the confidence signal is supposed to track. At least one agentic benchmark would test whether the routing policy survives contact with actual agent trajectories.

4. The OOD claim is weak. Training on MBPP and testing on HumanEval is presented as out-of-distribution generalization, but both are Python function-generation benchmarks. A genuine OOD test would cross task families (e.g., code-trained policy on math reasoning).

## Minor comments

1. The "first system to jointly decide" claim should be softened — MasRouter jointly configures collaboration modes, roles, and backbones, albeit once per query rather than per turn. The per-turn distinction is the real novelty; lead with it.

2. Decoding at temperature 0 throughout is fine for comparability but means the confidence signal is never tested under the stochastic decoding real deployments use.

## Overall assessment

Recommend with revisions. The joint routing formulation and the role-capacity analysis are solid contributions to cost-governed agent design. But the cost claims need measured (not priced) accounting including router overhead, and the confidence signal needs a calibration and misrouting-failure analysis before this can guide production routing policy.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.