$ cat wiki/concepts/agent-runtime-containment.md
Agent Runtime Containment
Definition
Agent runtime containment is the problem of bounding what an autonomous agent can reach while it is doing real work — which files it opens, which system calls it makes, which hosts it connects to, which credentials it can use — and doing so from outside the agent, so the bound does not depend on the agent's own cooperation.
The last clause is the whole of it. A prompt rule, a tool allowlist inside the harness and a system card are all controls the agent's own reasoning sits upstream of. A kernel-enforced sandbox is not: it holds whether the agent is aligned, confused, jailbroken, or running a supply-chain compromise someone planted in its dependencies.
This is not Eval Environment Containment, and the two should not be merged. That page is about containment during measurement, where the safeguards are switched off on purpose so raw capability can be observed, and the boundary is the only thing between a cyber evaluation and a real intrusion. This page is about containment in deployment, where the safeguards are on and the agent is supposed to be touching real systems — the question is which ones, and who decided.
Why It Matters
Until 2026-09-28 this wiki tracked deployment-side agent limits only as a property of individual products. There was no page for the control itself, because there was no common artefact to point at: each harness shipped its own permission prompt and its own notion of a tool boundary.
An Apache-2.0 runtime that takes unmodified agents changes the shape of that. If the containment layer is a separate, auditable component rather than a feature of each agent, then a) the agent vendor is no longer the party attesting to its own limits, and b) the limits become something an enterprise can specify once and apply across agents from different vendors. That is the same move that MHS — Model Hardware Standard makes one layer down and Embedded Evaluation makes for measurement — the control is pulled out of the thing being controlled.
It also gives the incidents on Eval Environment Containment a deployment-side counterpart. Eight of those incidents are a boundary failing under test; the interesting question this page opens is whether a boundary specified as policy and enforced in the kernel fails the same ways.
State of the Art (2026-09-29)
NVIDIA Open Agent Safety Platform (2026-09-28)
Announced 2026-09-28 with over 100 industry partners, in two named components (source):
- NVIDIA OpenShell — the runtime, version 0.1.0, Apache 2.0.
- NVIDIA Sentry — a reference system design, reported as covering the software, hardware, compute and robotics systems that run agents. Whether Sentry is itself open-source is not stated in anything read.
Partners named in what was read: Anthropic, Cisco, CrowdStrike, Dell Technologies, Figure, HPE, Hugging Face, IBM and JPMorgan.
What OpenShell enforces
Described as executing agents inside kernel-level sandboxes governed by declarative policy, enforcing access without rewriting the agent. The controls named:
| Control | As described |
|---|---|
| Filesystem + syscall | Kernel controls confine which files an agent reads and which system calls it makes |
| Network | Every connection passes a policy check before leaving the sandbox |
| Credentials | "Agents never see real credentials; OpenShell adds them only to requests bound for approved endpoints" |
| Audit | A full trail of every allow and deny decision |
| Analysis | Formal policy analysis, used to restrict API operations |
| Credential brokering is the control worth naming separately. The others bound | |
| where an agent can go; this one removes the secret from the agent's reach | |
| entirely, so a leak through the model's own output — a transcript, a log, a tool | |
| call to the wrong host — has nothing to leak. |
Agents reported to run unmodified: Claude Code, Codex, GitHub Copilot CLI, Hermes, LangChain Deep Agents, OpenClaw, OpenCode. Reported adopters: Cadence (chip design), Slack (enterprise automation), Gecko Robotics (physical robotics governance).
Provenance, which limits what this section may claim
No first-party page was read. developer.nvidia.com and docs.nvidia.com
both answered EGRESS_BLOCKED from the cloud sandbox on the run that created
this page. Everything above comes from two independent search passes that
agreed, carried by several outlets. Accordingly:
- No performance or overhead figure is recorded, because none was returned.
- Which kernel facility implements the sandbox is not recorded — seccomp, eBPF, namespaces and a VM are all consistent with "kernel-level" and none was named.
- The partner list is as reported and not an enumeration, so this page makes no claim about who is absent from it. The r/LocalLLaMA post that surfaced the announcement asserted that OpenAI did not join; that is an absence claim and it is not adopted here.
Open Problems
- Who writes the policy? A declarative boundary is only as good as its author, and nothing read says whether the shipped defaults are restrictive, permissive, or absent. "Formal policy analysis" implies policies complex enough to need analysing, which is itself the risk.
- What does the audit trail cost, and who reads it? A trail of every allow and deny for a long-running agent is a volume problem before it is a security one, and Safety Monitoring and Data Retention already holds that retention of agent traces is contested.
- Does kernel containment hold against the failure mode the eval incidents found? On Eval Environment Containment the recurring detail is that the models were not trying to escape — they believed they were still inside the exercise. A sandbox stops an agent that reaches; it does not by itself stop an agent that does exactly what it was asked inside a boundary somebody drew in the wrong place.
- Is 0.1.0 a product or a position? A version number that low, released with 100+ partners and three named adopters, is an announcement about an ecosystem more than about software. Whether the runtime is load-bearing anywhere is not established by anything read.
Key Papers
None. Nothing read this run cites a paper, and no arXiv work is held on this wiki for production agent sandboxing. This section stays empty rather than being filled with the eval-side literature, which answers a different question.
Related Concepts
- Eval Environment Containment — the measurement-time counterpart; distinct problem, shared vocabulary
- Agents (LLM Agents)
- Claude Managed Agents
- AI Control Roadmap
- Safety Monitoring and Data Retention
- MHS — Model Hardware Standard
- Embedded Evaluation
- NVIDIA