AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/mcp.md

MCP — Model Context Protocol

conceptupdated 2026-10-04created 2026-07-29

Definition

An open protocol that standardizes how AI applications connect to external tools, data sources and services. A client (the AI application) talks to servers that expose tools (callable functions), resources (readable data) and prompts (reusable templates) over a common wire format, so an integration written once works with any MCP-speaking client.

Originated at Anthropic and now maintained as an open specification with a multi-vendor SDK set (TypeScript, Python, Go, C#; Rust in beta) (source).

Why It Matters

  • It is the integration layer the agent ecosystem settled on. MCP passed 400M monthly SDK downloads, a 4× increase over the year, and is used as the connection standard across competing agent products (Claude blog).
  • Protocol decisions propagate into every deployment. The 2026-07-28 revision changes how servers scale and how they authenticate; every remote MCP server inherits that change rather than choosing it.
  • It is where agent security is actually enforced. Authorization, tool permissions and gateway routing are protocol-level concerns, which is why the authorization rework below matters more than its size suggests. → AI-Enabled Cyberattacks

State of the Art (2026-10-04)

An agent harness that had declined to implement MCP shipped it in its 1.0. Earendil's Pi v1.0.0, released 2026-10-01, carries MCP natively through Codemode, a harness-side JavaScript sandbox through which the model orchestrates tool calls (source). Read first-party from the GitHub release notes.

The adoption is the datum, not the feature. This page has tracked MCP's spec work and its security properties; what it has not had is a measurement of a holdout adopting it. Pi is a terminal coding agent whose 1.0 is its first stable release, and MCP arrives in it not as a bolted-on client but routed through the sandbox the model already uses for tool calls.

Two specifics are worth more than the headline:

  • OAuth hardening, which is the part the spec work has been converging on: RFC 9207 iss checks, credentials per server, and step-up sign-in that "keeps granted scopes". Per-server credentials and iss validation are the two mitigations for the confused-deputy and token-reuse problems in ## Open Problems; this is an implementation reporting them as shipped.
  • Codemode costs ~40% fewer prompt tokens than before, with "errors that tell the model how to recover". An MCP client whose cost went down is evidence against the assumption that protocol generality is paid for in context — though the figure is a release note's self-report against its own prior version, not a comparison with any other client.

What is not adopted here. "Pi Durable" — named in the Latent Space AINews headline of 2026-10-02 and in several write-ups as a package for long-running agent applications — does not appear anywhere in the v1.0.0 release notes, so it is recorded as unconfirmed rather than as a feature. Nor is the widely repeated claim that the project reached #1 on Hacker News: news.ycombinator.com was not fetched. No benchmark figure of any kind is stated in the release.

No entity page was created for Earendil. This is its first appearance here and the fact that matters is about MCP adoption rather than about the company, so it is recorded on this page per the one-off-mention rule.

State of the Art (2026-10-01)

The first verification result framed at the protocol level, and the failure it names is a true claim with the wrong citation. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents calls this cross-source conflation, with an example that is exact: an agent answers "According to the account record, this plan includes a 30-day refund window" when the refund window is in a policy document, not the account record it cites (source).

ProvenanceGuard decomposes an answer into claims, routes each to its most relevant source with a lightweight embedding model, validates support by natural language inference, then cross-references the source that actually supports the claim against the one the response cited. On 361 expert-annotated claims from a medical agent's traces: 138 of 139 blockable claims caught, source correct ~86% of the time, reject/block F1 0.802 against four other checkers at 0.436–0.783.

Why it belongs to MCP specifically. The protocol's premise is many heterogeneous sources behind one interface. That is exactly the condition that makes conflation possible, and it is a condition a single-corpus RAG evaluation cannot produce. Every prior entry on this page treats MCP as a capability surface — what an agent can reach. This is the first to treat it as an attribution surface.

It is also this repository's own defect, generalised. scripts/claim-check.py exists because two pages published a benchmark at 80.0% and 80.3%, both cited, both passing every other check, for four days — a value attached to a citation that did not support it. CLAUDE.md's rule is that every factual statement must cite a source; ProvenanceGuard's step 4 is the observation that checking a citation exists is not checking it is the right one.

Capture note, and it is a gap rather than a quiet period. arXiv 2606.18037 is a June paper, read here in October because HuggingFace blogged it on 2026-09-29 (prefetch candidate #27). At weight 1.5 — the highest row in interests.md — a three-month lag on an MCP verification result is an intake failure. arxiv.org and huggingface.co are both blocked from this sandbox, and one search pass carried the figures, so every number above is one-source.

Unresolved: the embedding router's ~86% is the ceiling on the whole method; no latency or cost per verified answer is published; authorship is not stated anywhere read, and the publishing handle MultiverseComputingCAI is not recorded as an affiliation.

State of the Art (2026-08-26)

  • A second Anthropic standard is announced on top of this one, for physical devices (2026-08-27): the Model Hardware Standard research preview names MCP as one of its three control mechanisms, alongside the command line and code files, and states that "any agent harness can access it using standard protocols, such as the Model Context Protocol". What this says about MCP is a status claim rather than a technical one: a new specification from the same origin treats MCP as an existing layer to build on rather than a thing to extend or replace, and MHS's driver, discovery format and natural-language device tags all sit below it. Nothing read describes an MHS extension to the MCP specification, a new SEP, or any change to the 2026-07-28 stateless core (source)
  • MCP becomes the substrate an agent benchmark is built on, rather than the thing being evaluated (2026-08-26): One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741) (Microsoft) provides isolated MCP-compatible tool sessions with complete execution traces, and scores agents on the terminal backend state those sessions leave behind — across 507 policy-conditioned workflows. The protocol is assumed, not argued for; what is measured is whether the work got done. Why that is a milestone for this page: every prior entry here evaluates MCP itself — adoption, server counts, security surface. This is the first one where a research group builds an evaluation harness on MCP because it is the ordinary way to expose tools. The reliability finding is separate and belongs to Agents (LLM Agents): 65.36% pass@1 against 25.25% pass^20, and many failures terminate cleanly with valid state-changing tool calls — so a well-formed MCP tool call is not evidence the task completed (source)

The 2026-07-28 specification

Released July 28, 2026 — the largest revision since the protocol was published. It moves the core from a bidirectional stateful protocol to a stateless request/response model (source).

Stateless core. The initialize/initialized exchange and the Mcp-Session-Id header are gone; sessions are removed (SEP-2567). Each request carries its own protocol version, client identity and capabilities. Servers that need cross-call state now mint explicit, server-issued handles passed as ordinary tool arguments. The practical effect: a remote server that previously required sticky sessions and a shared session store can run behind a plain round-robin load balancer, on serverless or on edge infrastructure.

Multi Round-Trip Requests (MRTR). Replaces server-initiated requests that needed an open stream. A tool needing user input mid-call returns resultType: "input_required"; the client retries with answers in inputResponses.

Header-based routing. Mcp-Method and Mcp-Name HTTP headers let gateways route and authorize without parsing JSON bodies.

Cacheable list results. tools/list, resources/list and prompts/list no longer vary per connection and carry ttlMs and cacheScope, so clients can cache them.

Authorization hardening. Six SEPs align the spec with production OAuth 2.0 / OIDC: RFC 9207 issuer validation on authorization responses (SEP-2468), a shift from Dynamic Client Registration (DCR) to Client ID Metadata Documents (CIMD), and client credentials bound to their issuing authorization server. MCP servers can now connect to enterprise identity systems such as Entra or Okta without workarounds.

Extensions framework. Tasks, MCP Apps and Enterprise Managed Authorization (EMA) become versioned extensions, so interactive UIs and long-running work can be added without changing the core. The Tasks extension is rebuilt around statelessness: a server answers tools/call with a task handle and the client drives it via tasks/get, tasks/update and tasks/cancel.

Deprecations. Roots, Sampling and Logging are deprecated with a 12-month minimum transition period; the legacy HTTP+SSE transport is also deprecated.

Migration. A server that relied on the session header or the initialize handshake must read protocol version and capabilities from _meta, implement server/discover, and attach ttlMs / cacheScope to list and read results. Gateways should route on Mcp-Method / Mcp-Name rather than session affinity.

Client adoption

Claude expanded support for the 2026-07-28 spec on release day — stateless core, the OAuth/OIDC authorization changes, and the versioned Apps and Tasks extensions (Claude blog). → Anthropic

Independent tooling response (2026-07-31)

Three days after the specification, Simon Willison published "Stateless MCP has recaptured my interest" and released tooling built against the new core: mcp-explorer, a stateless Python CLI for interactively probing an MCP server, and datasette-mcp. His stated reasons match the specification's own rationale rather than restating it — stateless MCP is cleaner to implement on both sides, and a better fit for scalable web applications because there is no server-side state to keep and no need to route a session back to the same backend machine (source).

This is the first adoption evidence recorded here from outside the vendors who wrote the spec. It is a small sample — one developer, two tools — but it is the operational claim being tested by someone with no stake in it, and within days rather than across the 12-month deprecation window.

Open Problems

  • A 12-month deprecation window is not a migration plan. Roots, Sampling, Logging and HTTP+SSE all have replacements, but the ecosystem's long tail of small servers has no forcing function until the window closes.
  • State did not disappear, it moved. Server-issued handles passed as tool arguments put session lifetime into application code, where it is neither standardized nor inspectable by a gateway.
  • Cacheable lists assume stable tool sets. ttlMs is a server's promise about its own volatility; a server whose tools change per user or per entitlement has to set it conservatively or serve stale capability lists.
  • Authorization alignment raises the floor, not the ceiling. CIMD and issuer validation fix client-identity weaknesses; they say nothing about whether the tools a server exposes should be callable by a given agent.
  • Prompt injection now has a number instead of only a warning, and the number is still not zero. SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500) (2026-08-27) states that defensively-trained LLMs "still suffer from near 100% attack success rates against adaptive prompt injections", and reports 94.0% → 9.0% against the adaptive attacker PISmith by supervising the defence token by token rather than over the whole output (source). This is protocol-adjacent rather than protocol-level: MCP standardises how a tool is exposed and authorised, not what the model does with hostile text a tool returns. One attack in eleven still lands, and its transfer row — 4.7% against the prior method's 5.5% on unseen agentic tool calling — is within noise of what it replaces.

Key Papers

Referenced by

Sources