AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/agent-runtime-containment.md

Agent Runtime Containment

conceptupdated 2026-09-29created 2026-09-29

Definition

Agent runtime containment is the problem of bounding what an autonomous agent can reach while it is doing real work — which files it opens, which system calls it makes, which hosts it connects to, which credentials it can use — and doing so from outside the agent, so the bound does not depend on the agent's own cooperation.

The last clause is the whole of it. A prompt rule, a tool allowlist inside the harness and a system card are all controls the agent's own reasoning sits upstream of. A kernel-enforced sandbox is not: it holds whether the agent is aligned, confused, jailbroken, or running a supply-chain compromise someone planted in its dependencies.

This is not Eval Environment Containment, and the two should not be merged. That page is about containment during measurement, where the safeguards are switched off on purpose so raw capability can be observed, and the boundary is the only thing between a cyber evaluation and a real intrusion. This page is about containment in deployment, where the safeguards are on and the agent is supposed to be touching real systems — the question is which ones, and who decided.

Why It Matters

Until 2026-09-28 this wiki tracked deployment-side agent limits only as a property of individual products. There was no page for the control itself, because there was no common artefact to point at: each harness shipped its own permission prompt and its own notion of a tool boundary.

An Apache-2.0 runtime that takes unmodified agents changes the shape of that. If the containment layer is a separate, auditable component rather than a feature of each agent, then a) the agent vendor is no longer the party attesting to its own limits, and b) the limits become something an enterprise can specify once and apply across agents from different vendors. That is the same move that MHS — Model Hardware Standard makes one layer down and Embedded Evaluation makes for measurement — the control is pulled out of the thing being controlled.

It also gives the incidents on Eval Environment Containment a deployment-side counterpart. Eight of those incidents are a boundary failing under test; the interesting question this page opens is whether a boundary specified as policy and enforced in the kernel fails the same ways.

State of the Art (2026-09-29)

NVIDIA Open Agent Safety Platform (2026-09-28)

Announced 2026-09-28 with over 100 industry partners, in two named components (source):

  • NVIDIA OpenShell — the runtime, version 0.1.0, Apache 2.0.
  • NVIDIA Sentry — a reference system design, reported as covering the software, hardware, compute and robotics systems that run agents. Whether Sentry is itself open-source is not stated in anything read.

Partners named in what was read: Anthropic, Cisco, CrowdStrike, Dell Technologies, Figure, HPE, Hugging Face, IBM and JPMorgan.

What OpenShell enforces

Described as executing agents inside kernel-level sandboxes governed by declarative policy, enforcing access without rewriting the agent. The controls named:

ControlAs described
Filesystem + syscallKernel controls confine which files an agent reads and which system calls it makes
NetworkEvery connection passes a policy check before leaving the sandbox
Credentials"Agents never see real credentials; OpenShell adds them only to requests bound for approved endpoints"
AuditA full trail of every allow and deny decision
AnalysisFormal policy analysis, used to restrict API operations
Credential brokering is the control worth naming separately. The others bound
where an agent can go; this one removes the secret from the agent's reach
entirely, so a leak through the model's own output — a transcript, a log, a tool
call to the wrong host — has nothing to leak.

Agents reported to run unmodified: Claude Code, Codex, GitHub Copilot CLI, Hermes, LangChain Deep Agents, OpenClaw, OpenCode. Reported adopters: Cadence (chip design), Slack (enterprise automation), Gecko Robotics (physical robotics governance).

Provenance, which limits what this section may claim

No first-party page was read. developer.nvidia.com and docs.nvidia.com both answered EGRESS_BLOCKED from the cloud sandbox on the run that created this page. Everything above comes from two independent search passes that agreed, carried by several outlets. Accordingly:

  • No performance or overhead figure is recorded, because none was returned.
  • Which kernel facility implements the sandbox is not recorded — seccomp, eBPF, namespaces and a VM are all consistent with "kernel-level" and none was named.
  • The partner list is as reported and not an enumeration, so this page makes no claim about who is absent from it. The r/LocalLLaMA post that surfaced the announcement asserted that OpenAI did not join; that is an absence claim and it is not adopted here.

Open Problems

  1. Who writes the policy? A declarative boundary is only as good as its author, and nothing read says whether the shipped defaults are restrictive, permissive, or absent. "Formal policy analysis" implies policies complex enough to need analysing, which is itself the risk.
  2. What does the audit trail cost, and who reads it? A trail of every allow and deny for a long-running agent is a volume problem before it is a security one, and Safety Monitoring and Data Retention already holds that retention of agent traces is contested.
  3. Does kernel containment hold against the failure mode the eval incidents found? On Eval Environment Containment the recurring detail is that the models were not trying to escape — they believed they were still inside the exercise. A sandbox stops an agent that reaches; it does not by itself stop an agent that does exactly what it was asked inside a boundary somebody drew in the wrong place.
  4. Is 0.1.0 a product or a position? A version number that low, released with 100+ partners and three named adopters, is an announcement about an ecosystem more than about software. Whether the runtime is load-bearing anywhere is not established by anything read.

Key Papers

None. Nothing read this run cites a paper, and no arXiv work is held on this wiki for production agent sandboxing. This section stays empty rather than being filled with the eval-side literature, which answers a different question.

Referenced by

Sources