Harness Tokenomics: A Router for the Enterprise Agentic Control Plane
- Posted
- Server
- arXiv
- DOI
- 10.48550/arxiv.2609.28919
Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees. What started as pilots with a few hundred seats is now scaling to tens of thousands. Enterprises rarely build these harnesses and usually adopt the ones model vendors bundle with their seats and APIs, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought at it. Enterprises that keep a proprietary or untuned harness at its defaults inherit these choices and their bill. We build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. One user turn is many requests over a prompt cache that belongs to one model, so the router moves work only where no running conversation has to rebuild its cache. It routes at session start, in side lanes and at subagent launch. From the price sheet we derive when a mid-task switch pays back. We also find a crossover, in which the highest-priced model costs less than the next tier on long tool-heavy sessions, and repricing about 10,000 real sessions from public datasets confirms it. In an emulated enterprise of 10,000 seats with user behaviour taken from these datasets, the router recovers 13 to 21% of model spend at Anthropic's list prices of 21 September 2026, \$3.3M to \$5.1M a year. The paper also maps the risks across twenty-one harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within, starting now, with a ladder for deciding later whether to own the harness.