Saltar a PREreview

PREreview del You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

Publicado
DOI
10.5281/zenodo.23147062
Licencia
CC0 1.0

## Summary

This paper identifies a genuinely under-appreciated routing axis: after a model router picks Llama-3.3-70B, the client must still choose *which provider serves it* — and the price list is a bad guide. Measuring live endpoints across 6 open models, multiple providers, and three waves over 43 days, the authors find price predicts latency (median Spearman −0.61) but not accuracy (+0.05) or availability (0.00); feasibility is task-selective (one Llama-3.3-70B deployment is near-normal on knowledge tasks but catastrophically degraded on multi-step reasoning); and the map drifts (routes changed in 4/17 comparable cells over days 0–13). Their measured-map policy — cheapest provider within 5 points of best measured accuracy with >90% availability — yields median 50% savings vs. the premium provider at matched quality. FACET, their online router, certifies per-(provider×task) feasibility facets before serving cheap endpoints, fails safe to an anchor, and monitors certified facets with a slip detector. A 36-hour live run (Llama-3.3-70B, 12 providers, 538 queries) cut average served price from $0.924/M to $0.210/M tokens — 63.7% serving-cost reduction (57.1% incl. probes) at 88.3% accuracy.

## Strengths

1. **The decision axis is real and well-positioned.** Same weights, different quantization/kernels/batching, wildly different service — anyone running multi-provider open-weight inference has felt this. Appendix B's Table 7 draws the line cleanly against FrugalGPT/RouteLLM/CARROT/MixLLM: those pick the model; this picks the server underneath, and the two layers compose (Section 5).

2. **Unusually honest measurement reporting.** Three-wave drift study, thorough Appendix A limitations (attribution-agnostic claims; pinned probing can induce rate limits; single aggregator and client location; price–latency is associational, not causal). The task-selective failure finding alone is worth the paper.

3. **FACET is framed as insurance, not optimization.** Separating certification (admission) from monitoring (drift) with component ablations (Tables 3–4), plus an explicit break-even analysis of the safety premium (~17× an ordinary query vs. SW-UCB, Appendix E), is the right framing: you pay for the anchor to avoid serving a catastrophic mine.

## Major concerns

1. **The live deployment never tests the core mechanism.** The single 36-hour run is caveated: "no provider quality collapse occurred during this window, so the experiment validates cold-start certification and live migration rather than slip detection." Slip detection is FACET's raison d'être, and it is exercised only in exact-replay — with Appendix A admitting the ablation and bias stress tests all use a single Llama-3.3-70B/GSM8K replay cell. The headline experiment should show the detector catching a real slip, on multiple cells, live.

2. **Threshold fragility is buried in the appendices.** The 50% saving rests on δ=0.05, nmin=12, 90% availability — but the out-of-sample below-floor rate swings from 0% to 22.2% across threshold choices (Appendix H), and the confidence-based analysis cuts median savings from 55% to 46% (Appendix U). The main text should carry uncertainty-quantified savings, not point-estimate victory margins.

3. **The trust boundary reintroduces the cost the paper claims to remove.** Systematic evaluator bias "can corrupt certification unless ground-truth probes or audits provide an independent quality signal" — but ground-truth labels for arbitrary production queries are exactly what doesn't exist. And the anchor must be maintained: in practice it's the premium provider you were trying to escape. Who pays for the anchor and the gold audits is left unquantified.

## Practitioner perspective

The task-selective failure finding is the one I'd pin to every team's wall: provider safety is a (provider × task × time) measurement, not a label. But the production math that matters — anchor cost + probe cost + gold-audit cost vs. measured-map saving — is only half quantified. Before adopting FACET, I'd want the slip-detection demo on live traffic and the full cost of the trust infrastructure in one table.

## Overall assessment

An important paper with novel framing, careful measurement, and rare honesty about limitations. Recommend with revisions: demonstrate slip detection live (or across multiple replay cells), move uncertainty-quantified savings into the main text, and quantify who pays for the anchor and ground-truth probes.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.