AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/embodied-agents.md

Embodied Agents

Definition

Agents that act through perception and actuation in the physical world or in simulated environments. Unlike pure text-output LLMs, they require a multi-modal perception API + actuation API + skill library structure.

Why It Matters

  • Matches agents (1.5x) × intersects with robotics — adjacent to personal interests
  • The next frontier beyond applying LLMs to software
  • Key 2026 currents: NVIDIA Project GR00T, DeepMind robotics research

State of the Art (2026-10-02)

A robot exposure index puts capability far ahead of cost, and the gap is measured in decades rather than in benchmarks.

Anthropic's What work can robots do? — Russell Legate-Yang and Maxim Massenkoff — scored roughly 7,594 physical job tasks from O*NET (~900 occupations, ~19,000 tasks) into four exposure tiers (E0–E3), keyed to how much environmental control each task requires (source).

The two headline numbers point opposite ways, which is the finding:

  • Robots can perform 74% of physical tasks, comprising 34% of US working hours. With LLMs included, 80% of job tasks by working time are exposed.
  • Robots are cost-competitive with people for 0.3% of job tasks. At the historical 3% annual price decline, reaching 10% takes 40 years.

So for this page the binding constraint on embodiment is not the capability work every result above tracks — it is unit economics, and the stated timescale is long enough that a capability jump would have to change the cost curve, not the capability curve, to matter. The authors say so directly: "Capabilities and costs are the biggest impediments to robot adoption", and they put cost first in effect if not in word order.

The distributional findings are the part a capability page usually drops. The top exposed occupation, taxi drivers, scores 2.2 on a 0–3 scale; packers/packagers employment is down 22% since 2015; and highly exposed workers are 20 percentage points less likely to be female and 55 percentage points less likely to hold a bachelor's degree. Jobs more exposed to robots saw greater wage and employment declines over the past 50 years — a backward-looking validation of the index rather than a forecast.

Caveats, stated by the authors. "Predicting the pace of robot advances is difficult"; the exposure estimates "require many judgment calls"; O*NET task statements are "often terse, and omit details"; the cost scenarios "apply the same cost decline...to every task"; and AI-powered robots "could leapfrog our scale and do work they cannot today" — which would invalidate the 40-year figure rather than adjust it. No model version is stated for the Claude that did the scoring, so the index is not reproducible against a named model.

State of the Art (2026-09-28)

Three papers recorded here without pages of their own, per the one-off-mention rule — each is a component result inside an existing line of work on this page, and none carries a claim that needs a page to hold it.

  • MemBodied — the week's third memory paper, and the only one about a body. 2609.28256 gives a Vision-Language-Action policy a fixed-size episodic memory with two parts: an associative state recording interactions across policy calls, and an episode anchor holding a compact representation of the initial scene. The policy conditions on the memory instead of on retained past observations, which is the point — keeping observations in context works but costs ever-growing context and inference latency. On five RMBench tasks requiring memory it reaches 7.81× the mean success rate of a stateless policy and 2.98× vanilla recurrent memory, beats the strongest memory-augmented baseline by 1.3× with 10× fewer added parameters, and on the fully observable LIBERO-Long suite reaches 90.6%, +5.4% over stateless π₀. Why it is recorded here: this wiki captured two memory-architecture papers in the preceding days — JitMem deferring curation because the query is unknown, SpeakerMem-R1 arguing attribution must be built at write time — and both are about text. MemBodied is the same argument where the history is physical state, and it lands on the side of fixed-size compression: the associative state is written once per call and never grows. The 7.81× is against a stateless baseline on tasks selected for requiring memory, which is close to the largest number such a comparison can produce; the 90.6% / +5.4% on a fully observable suite is the more transferable figure and is much smaller (source)

  • Spatial-Interactor — interaction trajectories as supervision for state transitions. 2609.23038 argues that current spatial training for VLMs asks static questions about attributes and relations and therefore gives almost no supervision for state transitions, while an interaction trajectory intrinsically pairs observation → action → observation. It trains on a three-level curriculum — L1 passive world-state transitions, L2 active self-state transitions, L3 long-horizon trajectories — over LSI-108K, built from simulated and real trajectories, with SFT on L1/L2 and On-Policy Distillation on L3 where a teacher given segment-level transition descriptions supervises the student's on-policy chain of thought. Reported consistent gains across multiple VLMs and spatial benchmarks. No absolute numbers appear in the abstract — no benchmark score, no baseline, no dataset split — so the size of the effect is unread, which is why it is a mention and not a page (source)

  • PackLab — an MLLM beating geometric heuristics at bin packing, reported without a number. 2609.23784 supplies a physics-based simulator for generating packing trajectories and checking physical outcomes, a packing-specialised MLLM that jointly selects the object and predicts the placement in closed loop, and a benchmark at multiple difficulty levels. Reported to beat conventional packing heuristics, RL policies and general-purpose MLLMs on average across object sets and container configurations; code, model, dataset and benchmark are released at github.com/Correr-Zhou/PackLab. "On average" is the whole result text — no success rate, no per-difficulty breakdown and no baseline figures, so the claim is recorded and not quantified (source)

State of the Art (2026-09-06)

  • RoboTok — retrieval, not collection, as the answer to the demonstration bottleneck (2026-09-02, recorded 09-06) — 2609.03199 treats the web as the demonstration corpus: given a query human manipulation video, it retrieves manipulation-relevant human demonstrations from internet video to train dexterous policies. The representation is a latent motion space learned from 3D hand trajectories in estimated actor-centred reference frames, which lets behaviours be compared across camera viewpoint, scene appearance and actor occlusion while staying compact enough for continual indexing at internet scale. Reported to beat existing robot-data retrieval approaches on retrieval benchmarks and on downstream policy success. Why it is recorded here without a page: the argument is the same one EgoScale made from the other end — human video is the scalable supervision source — but where EgoScale trained on 20,000+ hours indiscriminately, this selects from an open corpus per task, which is the cheaper claim and the one that keeps working as the corpus grows. No absolute numbers appear in the abstract: no retrieval precision, no success rates, no corpus size, and no comparison against teleoperation-collected data at equal budget, so the size of the effect is unread (source)

  • The measurement, not the policy, is what two papers move on 2026-08-23 — and both say task success is the wrong number — SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation (arXiv:2608.18701) pairs policy-visible visuo-tactile observations with evaluator-only finite-element ground truth over 4,000 expert demonstrations at 20 Hz, and defines a Deformation-aware Success Rate that counts a rollout only if the task completes and peak deformation stays in tolerance. Across Diffusion Policy, π₀.₅ and FastWAM, all 12 in-distribution configurations contain successful rollouts that violate the tolerance — 0.7–24% of each configuration's successes. Its second finding is the one that should change designs: adding touch raises success in all six out-of-distribution policy–suite comparisons but DSR in only five, and in-distribution the benefit is mixed — "making touch available does not by itself ensure effective multimodal fusion." In parallel, Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning (arXiv:2608.18746) names decision-metric alignment for JEPA-style latent world models — whether Euclidean distance to a goal latent ranks plans by real progress — and shows it does not follow from representation quality: DA-LeWM converges faster and succeeds more online while probe scores remain similar. Two diagnostics are proposed (Plan-Real Spearman, CEM-stage Spearman) with encoder distortion, terminal rollout error and candidate margins as the controlling quantities. Both are the same correction from opposite ends of the stack: an independent channel the policy cannot see (FEM state), and a metric that measures the planner's actual decision rather than the representation feeding it. Neither reports absolute numbers in its abstract (source)

  • Zetta ζ — the harness closes the loop, and the policy is frozen (2026-08-17, recorded 08-21) — Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590) evolves code-based runtime critics and recovery skills online while the base policy stays untouched, through three timescale-separated loops: action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Reported LIBERO-Pro 90.8% and RoboCasa 93.6% with an 11.1× inference speedup — both stated as SOTA "under our current rollout budget", a qualifier kept because the budget is not published and both figures are scoped to it. Its premise is a frequency argument rather than a capability one: existing embodied harnesses are open-loop, reflecting only after an episode completes, and physical interaction needs decisions at a frequency beyond today's large agentic models — so the fast loop goes into generated code and the slow model writes it. This is the DarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545) pattern (evolve the harness, freeze the model) crossing from software into physical control, and it is the one place in that cluster where the argument for a harness is not about benchmark headroom. The "Aha Moments" and zero-shot skill transfer are asserted with no count and no target task (source)

  • Google DeepMind Gemini Robotics 2 (2026-07-28 / 2026-07-30) — three models: a VLA controlling a full humanoid (legs, torso, arms, multi-finger hands) under one learned policy, an embodied-reasoning VLM, and an on-device VLA. Whole-body manipulation on the Apollo 2 humanoid: 76.3% picking from a shelf, 45.7% from the floor. Fine-motor dexterity on the 22-DoF SharpaWave hand: 92% unscrewing a light bulb, 44% tying a trash bag, 40% sealing a ziplock, 36% screwing a bulb in, 32% dustpan. On-Device 2 adapts to a new embodiment in hours from fewer than 200 demonstrations. Released with ASIMOV-Agentic, an open benchmark for whether the reasoning layer refuses dangerous commands from the acting layer, judges physical feasibility, and asks for human help when uncertain. Significance: the dexterity spread is the state of the field in one table — removing a bulb is near-solved, inserting one is not, and deformable objects sit at the bottom. Note also which layer shipped: the reasoning model is public via the Gemini API, both acting models are early-access only. → Gemini Robotics 2, Gemini Robotics ER 2 (source)

  • GRANT / ORS3D — scheduling as the embodied objective (AAAI 2026 Oral; HF Daily #1 on 2026-07-30) — arXiv 2511.19430 defines ORS3D, where the agent minimizes total completion time by running parallelizable subtasks concurrently (clean the sink while the microwave runs), and trains an embodied multimodal LLM with a scheduling token on ORS3D-60K (60K composite tasks, 4K scenes). +30.53% Time Efficiency, +1.38% 3D grounding. Significance: the asymmetry is the point — grounding barely moves, scheduling moves a lot. As per-action reliability improves, the remaining headroom in a multi-step physical task is ordering, not perception. → Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT) (source)

  • Mistral Robostral Navigate (2026-07-08) — 8B robotics navigation model; single RGB camera + language prompts only; no depth sensors or LiDAR. Achieves 76.6% R2R-CE validation unseen — new SOTA (+9.7 pts over prior best single-camera; +4.5 pts over best depth/multi-camera). Sim-only training with CISPO online RL (+3.2%). Pointing-based prediction (image coordinates → target + orientation): naturally robust to camera intrinsic changes. Hardware-agnostic. First Mistral physical AI product. Significance: demonstrates that sim-to-real transfer in navigation is now strong enough to beat real-robot-trained depth-sensor systems using only RGB. Lowers hardware bar for deployment. → Robostral Navigate (source)

  • NVIDIA ENPIRE — Agentic Robot Self-Improvement on Real Hardware (2026-06-17) published ENPIRE: a fleet of 8 real robots autonomously runs its own research loop (read papers → hypothesize → run trials → verify → rewrite code) — zero humans in the loop. 99% pass@8 on contact-rich tasks. New finding: physical scaling law — 8 parallel robots improve policies superlinearly faster than fewer robots. The first demonstration that the LLM-era data-parallel scaling paradigm transfers to real-world physical manipulation research, not just skill acquisition. Closes the loop on EgoScale: EgoScale → acquire skills from egocentric video; ENPIRE → autonomously improve them. → ENPIRE: Agentic Robot Policy Self-Improvement in the Real World, Jim Fan (source) (arXiv)

  • NVIDIA EgoScale — Dexterous Humanoid from Egocentric Video (2026-06-07) ⚡ NEW — Jim Fan's team trained a humanoid with 22-DoF dexterous hands entirely from 20,000+ hours of egocentric human video (no robot teleoperation). Tasks: model car assembly, syringe operation, card sorting, shirt folding. Key result: log-linear scaling law (R² = 0.998) — human video volume → action prediction loss → real-robot success. First strong empirical evidence that the LLM data-scaling paradigm transfers directly to dexterous physical manipulation. Open-sourced: weights, code, dataset, eval set, whitepaper. → Jim Fan (source)

  • NVIDIA Cosmos 3 (2026-06-01) ⚡ NEW — The world's first fully open omnimodal physical AI foundation model. 64B (Super), Nano, Edge. MoT architecture: unifies world generation + physical reasoning + action prediction. #1 open model on R-Bench. MIT license. Cosmos Coalition (Agile Robots, Runway, Black Forest Labs, etc.). → Cosmos 3 Super (source)

  • Gemini Robotics-ER 1.6 (2026-04-15) — Google DeepMind. Gauge-reading accuracy 23%→93% (agentic vision pipeline). Boston Dynamics Spot collaboration. Crosses the threshold for autonomous inspection on industrial sites. Available via the Gemini API. → Gemini Robotics ER 1.6 (source)

  • CaP-X (2026-04) — Open-sourced by NVIDIA's Jim Fan team. "Vibe agents alive in the physical world." Robot arms + humanoids, auto-synthesize skill libraries. (source)

  • Project GR00T — NVIDIA humanoid robot foundation model

  • DrEureka — LLM writes robot skill training code (Jim Fan team)

  • Voyager — Lifelong learning agent in the Minecraft environment

Architecture Patterns

  • No-gradient orchestration: The LLM acts as the "prefrontal cortex" for high-level reasoning, while lower-level control is delegated to code-gen
  • Skill library: Learned/generated skills are accumulated into a retrievable library
  • Perception + actuation API separation — abstracting the environment

Open Problems

Key Sources

Open Debates

  • LLM-orchestrated structure vs end-to-end neural control — Jim Fan/NVIDIA strongly favor the former. Part of academia argues for the latter (e.g., the Sergey Levine group).

Referenced by

Sources