Skip to PREreview

PREreview of Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

Published
DOI
10.5281/zenodo.23197319
License
CC0 1.0

## Summary

This paper names a problem every agent deployer has felt but few have measured: open-source agent runtimes like OpenClaw expose every tool to every session by default — a summarization task gets shell execution, subagent spawning, and credential access, a 15x over-provision the authors call the capability overprovisioning problem. They propose AgentWarden (titled "Aethelgard" in the abstract), a governance middleware with three parts: a Capability Governor that scopes which tools the agent is even told about per session (via AGENTS.md injection plus a deny list), a Safety Router that intercepts every tool call before execution using a hybrid rule-based plus fine-tuned LLM classifier, and a PPO policy trained on the audit log to learn the minimum viable tool set per task type. The Skill Economy Ratio (SER) — fraction of exposed tools actually invoked — formalizes least privilege as a measurable property.

On a live OpenClaw deployment, the system blocks 26.2% of 447 intercepted tool calls (exec and sessions_spawn blocked at 100%) and neutralizes 92% of 100 adversarial tasks. On AgentDojo (97 tasks, 949 injection cases) it holds benign utility at 89.7% (−1.0pp) while reducing attack success rate by only 1.4pp overall. The paper documents a genuinely interesting finding: prompt injection that succeeds at the LLM level can still fail at the infrastructure level. The limitations section is unusually candid about PPO collapse, classifier data scarcity, and what the system does not cover.

## Strengths

1. The core insight is correct and underappreciated. "An agent cannot misuse a tool it does not know exists" — scoping capability awareness, not just execution permission, is the right layer for least privilege, and SER gives the field a metric to argue about instead of vibes.

2. Honest about the threat-surface boundary. The paper clearly states that the router governs tool calls while important_instructions-style attacks operate at the LLM reasoning boundary — and the AgentDojo results demonstrate it. This kind of scoped honesty is exactly what practitioners need to place a defense correctly.

3. The infrastructure-vs-model finding is valuable. Showing that qwen2.5:7b generates dangerous calls the infrastructure catches, while DeepSeek-chat refuses at the model level, makes the defense-in-depth thesis concrete rather than slogan-level.

4. Multi-runtime validation. Testing the middleware across OpenClaw, DeepAgents, and Hermes wire formats (including XML tool calls) shows this is built as infrastructure, not a demo.

## Major comments

1. The PPO policy collapsed in production testing. Both PPO checkpoints degenerated to single-tool policies (predicting only memory_search for all task types) — a known artifact of SER-maximizing training in synthetic environments. The paper is honest about this, but it means the headline "learned policy" does not actually work outside simulation; the YAML fallback did the real work. The learned-governance claim should be downgraded until the policy is retrained on real session data, as the authors themselves plan.

2. The AgentDojo results undercut the headline numbers. On the independent benchmark, ASR moves from 90.3% to 88.9% — a 1.4pp reduction. The 100% block rates come from the authors' own 100 hand-crafted adversarial tasks. Both are reported, which I respect, but the abstract leans on the friendlier numbers.

3. The classifier is not production-ready by the authors' own account. The fine-tuned router trained on 273 examples with a 95/5 class imbalance, and the base Qwen2.5-1.5B "performs poorly" without fine-tuning. A production-quality classifier needs 2,000–5,000 balanced examples (their estimate). The ~800ms–2s LLM-stage latency per tool call also compounds in long agent loops — a per-task latency analysis is missing.

4. Name and structure inconsistencies. The abstract calls the system "Aethelgard" and "four layer"; the body says "AgentWarden" and "three-layer." For a security paper, where precision is the product, this sloppiness undermines confidence.

5. The router inspects tool calls, not reasoning text. The paper documents a reverse-shell case where the model returned instructions as text with no tool call and the router saw nothing. This is a fundamental architectural boundary, correctly identified — but it means the system's protection assumes the dangerous action goes through a tool call, which a clever agent or prompt injection need not do.

## Minor comments

1. The evaluation dataset is "available upon request" rather than public — for a security paper making empirical claims, public release (sanitized) would materially strengthen trust.

2. The DeepAgents LocalShellBackend bypass finding is important and well documented; it deserves a line in the abstract's scope statement.

## Overall assessment

Recommend with major revisions. The capability-scoping framing, SER metric, and infrastructure-level enforcement are sound contributions that the field needs. But the collapsed RL policy, the thin classifier training data, and the gap between hand-crafted adversarial results and AgentDojo performance mean the "learned" and "adaptive" claims currently outrun the evidence. Fix the naming inconsistencies, fence the claims to what the production tests actually show, and this becomes a paper I would cite when designing agent deployments.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.