Skip to PREreview

PREreview of Sandlock: Confining AI Agent Code with Unprivileged Linux Primitives

Published
DOI
10.5281/zenodo.23197347
License
CC0 1.0

## Summary

Sandlock is a Rust process sandbox aimed squarely at the AI-agent workstation: short-lived, frequent, untrusted code execution (model-generated shell commands, package scripts, MCP tool servers, third-party plugins) on developer machines where root is unavailable and container startup costs are visible. The central design is a split enforcement model: static, input-independent policy compiles into kernel-enforced rules (Landlock for filesystem/ports/IPC scope, seccomp-bpf for syscall filtering), while a narrow supervisor handles runtime-dependent decisions via seccomp user notification — including TOCTOU-safe inspection of execve arguments, HTTP-level (method/host/path) network policy, and copy-on-write filesystem effects with commit/abort/dry-run semantics. A pipeline operator composes stages with heterogeneous confinements, so the stage exposed to untrusted content gets no network while the networked stage never sees private data — kernel-enforced capability separation for prompt-injection-resistant decompositions like the dual-LLM pattern.

Evaluation is explicitly preliminary: ~5ms startup overhead (~6ms wall, 44x faster than Docker on their setup), bare-metal Redis throughput within measurement noise, lower p99 latency than Docker, and ~1,900 COW-forks/s. Ten deterministic deny/allow checks from the threat model all behave as specified. The paper is open source, was presented at the ASPLOS 2026 Agentic OS Workshop, and is admirably clear about its boundaries: unprivileged-by-design limits, multi-tenant workloads still needing microVMs, and network effects not being rolled back.

## Strengths

1. The static-vs-runtime split is the right abstraction. Anyone who has written agent sandboxing policy knows the pain: some rules are known upfront, some depend on syscall-time values. Making the split explicit — and failing closed when a decision can't be made without a race — is the kind of design judgment that survives contact with production.

2. Pipeline composition addresses the real threat. The paper correctly identifies that the dangerous agent failures come from untrusted data, not just untrusted code — the "lethal trifecta" of private data, untrusted content, and external communication. Per-stage confinement with kernel enforcement is a genuine substrate contribution for dual-LLM/CaMeL-style decompositions.

3. HTTP-level network policy is genuinely useful. Method/host/path control is the right granularity for agent network policy (an agent that may call api.openai.com should not reach attacker hosts through the same allowed IP). The honest caveat — HTTPS needs a sandbox CA, otherwise pinned-endpoint allowlisting — is exactly the deployment detail practitioners need.

4. TOCTOU handling shows systems maturity. The policy_fn mechanism is designed against the classic syscall-interposition hazards (Garfinkel), and the paper's two caveats on stage composition (capabilities vs. pipe content; sound decomposition remains the author's responsibility) show the authors know what their tool does not prove.

## Major comments

1. The evaluation is microbenchmarks, not agent workloads. Redis throughput and /bin/echo startup show the sandbox is cheap; they say nothing about what policy_fn adds to per-tool-call latency in a real agent loop, where dozens of short-lived commands run per minute and every execve triggers supervisor round-trips. The paper needs at least one realistic agent workload (e.g., an SWE-agent-style trajectory) measuring end-to-end task latency with and without Sandlock.

2. No adversarial evaluation. The effectiveness section is ten deterministic deny/allow checks — necessary but not sufficient. A sandbox paper should include escape-attempt testing: TOCTOU races against policy_fn, supervisor DoS via syscall storms, and the seccomp-notification fd-passing paths. The design anticipates these; the evaluation should exercise them.

3. The Docker comparison mixes baselines. Benchmarks use rootful Docker (the "common deployment") while the feature table compares against rootless Docker (the unprivileged competitor). The 44x startup claim should be measured against the same baseline the no-root claim is made against — rootless Docker — or the mismatch should be prominently flagged.

4. Resource caps without cgroups are cooperative. RLIMIT-based caps are kernel-enforced per process, but the threat model scopes out pressure on shared kernel resources. The paper states the boundary; it should also say what happens in practice when an agent's dependency tree legitimately needs more than the cap — fail-closed kills the task, and operators need guidance on sizing.

## Minor comments

1. Compatibility is demonstrated on python3/make/node/pytest — a compatibility study across the long tail of agent tooling is listed as future work and genuinely needed before wide adoption.

2. The MCP tool-server retrofit mode (wrap each invocation as a confined child under a per-tool policy) is the most deployment-relevant paragraph in the paper and deserves more than a subsection — it is the adoption path.

## Overall assessment

Recommend with revisions. The design is principled, the threat model is honest, and the pipeline-composition story is the strongest unprivileged answer I've seen to capability separation for agents. But a systems paper needs a systems evaluation: one real agent workload, adversarial testing of the supervisor paths, and a consistent Docker baseline. The current evaluation proves the sandbox is cheap, not that it is sufficient.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.