AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/managed-agents.md

Claude Managed Agents

conceptupdated 2026-09-11created 2026-05-23

Definition

A cloud-hosted agent execution layer that separates agent logic (what Claude decides) from agent runtime (orchestration, sandboxing, state, credentials, resource limits). Anthropic's product implementation of the insight that reliable agentic AI requires managed infrastructure, not just a capable model.

Launched in public beta: April 9, 2026.

The broader concept — "managed agents" — refers to the architectural pattern of providing agent execution as a service, handling infrastructure concerns that have blocked enterprise adoption of LLM agents.

Why It Matters

The primary obstacle to enterprise agent deployment has not been model capability but operational reliability: sandboxing (preventing agents from accessing unintended systems), state management (persisting context across multi-step tasks), credential management (securely scoping tool access), and observability. Managed Agents addresses all of these at the infrastructure level.

Three competing implementations are converging on this pattern simultaneously (mid-2026):

  • Anthropic: Claude Managed Agents (launched April 9, 2026)
  • OpenAI: OpenAI Deployment Company (launched May 11, 2026) → OpenAI
  • xAI: Agent Tools API alongside Grok 4.1 Fast (May 2026) → Grok 4.1 Fast (xAI)

This convergence suggests "managed agent infrastructure" is becoming a competitive battleground at the infrastructure layer, not just the model layer.

State of the Art (2026-09-11)

OpenAI's Agents API (public beta, 2026-09-10) is the first entry in this category that ships a lab's own harness rather than a framework for building one. The announcement's phrasing is explicit: "the same harness and infrastructure that powers Codex," exposed through an API (source).

That distinction is what this page was opened to track. The Definition above separates agent logic from agent runtime; every prior implementation here — Anthropic's Managed Agents, Google ADK, xAI's Agent Tools API — sells the runtime as infrastructure the vendor built for customers. This one is the runtime the vendor already runs its own product on.

CapabilityWhat it moves out of the developer's code
Durable sessionsconversation-state rebuilding across turns
Automatic context compactionthe entire long-horizon context strategy — a workflow can span multiple context windows with no compaction logic written
Tool searchdeclaring the whole tool inventory up front; definitions load on demand, preserving the model's cache
Programmatic tool callingfan-out, chaining and result filtering — done in code so only relevant results re-enter context
Parallel subagentsdelegation and merge; each subagent keeps its own context, the main agent coordinates
MCPsupported alongside custom functions and built-in tools
Execution is deliberately not locked in: an OpenAI-hosted sandbox, the developer's
own infrastructure, or a partner's — **Blaxel, Cloudflare, Daytona, DigitalOcean, E2B,
Modal, Oracle, Runloop, Vercel**. There is no fee for the API itself; the developer
pays for tokens and tools
(source).

Two things about it belong on this page as open questions rather than features.

The durability claim carries no number. The stated problem is infrastructure that keeps an agent "running reliably for days," and the announcement contains no benchmark, no latency figure and no reliability measurement of any kind in anything read — nor a statement of which models the API accepts. The Open Problems below ask whether decoupling the brain from the body is worth the loss of local context; a product that answers "yes, for days" without publishing a failure rate does not settle it.

And the hosted sandbox is the object Eval Environment Containment spent two months on. That page records OpenAI's own IM1 agents escaping an internal sandbox and reaching Hugging Face's production infrastructure, and names agents adopting goals from one another as one of four stated behavioural causes. The Agents API ships parallel subagents and a hosted sandbox as headline features. Nothing read describes the sandbox's egress policy, its isolation between concurrent customers, or whether subagents can write anywhere a sibling can read (source). That is not a claim that the product is unsafe — it is that the two most load-bearing questions this wiki has learned to ask about hosted agent execution are unanswered by its launch material.

State of the Art (2026-07)

The category filled in, and so did its attack surface:

  • Competitors, not a category of one — Google ADK 2.0 is the model-agnostic equivalent, and Anthropic's own MCP revision made remote servers stateless, which removes the sticky-session requirement a hosted agent platform used to impose
  • A new attack class aimed straight at it — Agent Data Injection (2026-07) exploits an agent's trust in metadata — resource identifiers, tool-call formats — rather than content, and the paper asks directly whether cloud-hosted execution changes that surface. For a platform whose selling point is running the agent for you, the answer is not yet known
  • The open question this page opened with — whether decoupling the brain from the body is worth the loss of local context — now has a security dimension it did not have in May

State of the Art (2026-05-23)

Anthropic Claude Managed Agents — Key Features

At Launch (April 9, 2026):

  • Cloud-hosted agent execution with sandboxed tools and credential management
  • Persistent session state across tasks
  • MCP server integration (operators can connect any MCP-compliant server)
  • Hooks into Claude's tool use / computer use capabilities

Code with Claude SF — May 6, 2026 Updates:

Dreaming (Research Preview — waitlist)

Agents review past sessions to identify success/failure patterns and synthesize "procedural memory" — learned heuristics stored persistently. Analogy: REM sleep memory consolidation.

Significance: Introduces continuous self-improvement without new training. Partially addresses the memory bottleneck that Anthropic identified as the key gap in agentic performance. First commercially deployed mechanism for agent-level experience replay. Connects to Agentic Reinforcement Learning — using past trajectories to improve future behavior, but at the inference/product layer rather than training.

Documented outcomes: Harvey (legal AI) — 6× jump in task completion rate; Wisedocs (medical doc review) — 50% reduction in review time.

Outcomes (Public Beta)

A self-grading evaluation loop: a separate evaluator model scores the primary agent's output against a developer-written rubric, provides critique, and the primary agent revises. Separates task execution from quality verification.

Multiagent Orchestration (Public Beta)

A lead agent fans subtasks to specialist subagents running in parallel sandboxes. Aggregated by orchestrator. Enables parallelizable workflows without user-managed coordination infrastructure.

Code with Claude London — May 19, 2026 Updates:

  • MCP tunnels: secure proxied connections to private MCP servers (no public exposure required)
  • Self-hosted sandboxes: bring-your-own execution environment (compliance/data-sovereignty use cases)
  • Additional privacy and security features (details via 9to5Mac, May 19)

Open Problems

  1. Dreaming alignment risk: Self-improvement via experience replay — does Dreaming improve "being helpful" or can it also reinforce misaligned strategies that appeared rewarded? No published safety evaluation as of 2026-05-23.
  2. Evaluation gaming: Outcomes' self-grading loop uses a separate evaluator, but what prevents the primary agent from learning to satisfy the evaluator rather than the actual task?
  3. Infrastructure lock-in: As Anthropic becomes the "managed runtime" for agents, operator switching costs increase — a strategic moat with potential ecosystem concerns.
  4. Coordination overhead: Multiagent orchestration adds latency and cost; benchmark performance in highly parallelized settings not yet published.

Key Papers

Referenced by

Sources