AI Trend Notifier
EN한
← wiki

$ cat wiki/entities/nvidia.md

NVIDIA

Latest

  • 2026-09-29

    A tabular foundation model, published one day after the agent-safety platform, and the second NVIDIA release in two days to be given away rather than…

  • 2026-09-28

    NVIDIA shipped the containment layer as a separate, Apache-2.0 component, and brought 100+ partners with it.

  • 2026-09-02

    The Hugging Face acquisition is confirmed by a filing, and the contradiction this wiki declined to resolve resolves itself

Overview

A GPU manufacturer and the de facto standard for AI compute infrastructure. The infrastructure core of the AI era. Also active in AI research — particularly meaningful in-house work in the directions of embodied AI and agentic AI.

Key People

  • Jensen Huang — CEO
  • Jim Fan — Senior Research Scientist, AI Agents

Models

  • Cosmos 3 Super — Cosmos 3 Super (64B, MIT open weights). Released 2026-06-01. The world's first open omnimodal physical AI foundation model. R-Bench #1 open model.
  • Cosmos 3 Nano — small, real-time inference
  • Cosmos 3 Edge (coming soon) — edge on-device
  • Nemotron 3.5 Lightning — Nemotron 3.5 Lightning (30B MoE / 3B active, OpenMDW-1.1 open weights). Released 2026-08-11. 1M context, distilled from Nemotron 3 Ultra, built for agent workloads.
  • NVIDIA Kumo Tabular — NVIDIA Kumo Tabular (three sizes, 28M–215M, OpenMDW-1.1). Released 2026-09-29. Open foundation model for tabular classification and regression in a single forward pass; pretrained on artificial data only; reported 1st on TabArena (ELO 1950), BeyondArena (ELO 1418), TALENT and ScoringBench, and 26× faster than LimiX-2 on one RTX 6000 Pro. One search pass, no first-party read, and which of the three sizes posts those numbers is not stated (source)
  • Cosmos-H-Dreams — Cosmos-H-Dreams (Apache 2.0). Released 2026-07-22. Action-conditioned world model for surgical robotics simulation; generates photorealistic surgical scenes from robot commands; 600 policy rollouts/40 min vs. 2 days on physical hardware.

Research Streams

  • Cosmos 3 (2026-06-01) — physical AI omni-model (unified world generation + reasoning + action). Trained on 20T tokens. MoT architecture. → Cosmos 3 Super
  • NeMo Switchyard (2026-08-11) — open-source model routing library for agents; tuning-free routers (LLM classifier, stage router, escalation router) plus tunable ones; internal claim of frontier accuracy at ~1/3 the task cost of Opus 4.8 alone → Model Routing
  • Cosmos-H-Dreams (2026-07-22) — action-conditioned world model for surgical robotics; open-source (Apache 2.0); part of NVIDIA Medical Physics Simulation framework → Cosmos-H-Dreams
  • Molt (2026-07-27, HF Daily #1, 605 upvotes) — PyTorch-native agentic RL framework from NVIDIA NeMo. vLLM rollout + FSDP2 + Ray async queue. Matches Megatron-based systems without the complexity. → Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
  • Project GR00T — humanoid robot foundation model
  • CaP-X (2026-04, open-source) — vibe agents alive in physical world (Jim Fan team)
  • DrEureka — LLM writes robot skill training code
  • Voyager — lifelong learning agent in the Minecraft environment

Business Significance

  • Every frontier lab depends on its compute — Anthropic, OpenAI, DeepMind, Meta, and others
  • The OpenAI Broadcom partnership (2026~) is an attempt to diversify away from NVIDIA dependence
  • Aggressively recruits talent under the "warm GPUs" motto

Recent Activity

  • 2026-09-29: A tabular foundation model, published one day after the agent-safety platform, and the second NVIDIA release in two days to be given away rather than sold. NVIDIA Kumo Tabular (new page) — open weights under OpenMDW-1.1, three sizes 28M–215M, pretrained on artificial data only, returning class probabilities or numeric predictions for unlabeled rows in a single forward pass with no training, tuning or feature engineering. Reported 1st on TabArena (ELO 1950), BeyondArena (ELO 1418, Improvability 7.78%), TALENT and ScoringBench, at 26× LimiX-2's speed under a uniform single RTX 6000 Pro evaluation. Why it matters: this page's 2026 record is NVIDIA buying or licensing its way up the stack, with OpenShell the first instrument that gave a layer away. Kumo Tabular is the second in two days, and it reaches a workload class no model on this page has touched — structured business data, where the incumbent is gradient-boosted trees rather than any lab's model. The synthetic-only pretraining claim is the part with consequences beyond accuracy: a model that has provably seen no customer table is a different licensing proposition from one that will not say. Not established: one search pass, huggingface.co blocked, no first-party read; which size posts the leaderboard results is not stated across an 8× parameter span; all four "1st" placements are the vendor's own reading of public leaderboards with no third-party reproduction; and no context-window figure exists for a model whose input is a table read in context (source)

  • 2026-09-28: NVIDIA shipped the containment layer as a separate, Apache-2.0 component, and brought 100+ partners with it. The NVIDIA Open Agent Safety Platform, in two named parts: OpenShell, version 0.1.0 under Apache 2.0, an open-source runtime executing autonomous agents inside kernel-level sandboxes governed by declarative policy; and NVIDIA Sentry, a reference system design reported to cover the software, hardware, compute and robotics systems that run agents. OpenShell's controls as described: kernel confinement of which files and which system calls; a policy check on every network connection before it leaves the sandbox; credential brokering — "Agents never see real credentials; OpenShell adds them only to requests bound for approved endpoints"; an audit trail of every allow and deny decision; and formal policy analysis restricting API operations. Agents reported to run unmodified: Claude Code, Codex, GitHub Copilot CLI, Hermes, LangChain Deep Agents, OpenClaw, OpenCode. Reported adopters: Cadence (chip design), Slack (enterprise automation), Gecko Robotics (robotics governance). Partners named among the 100+: Anthropic, Cisco, CrowdStrike, Dell Technologies, Figure, HPE, Hugging Face, IBM, JPMorgan. Why it matters: every move this page has recorded in 2026 pulls NVIDIA up the stack by buying or licensing a layer — the $6B Poolside licence-and-hire, the ≈$12.93B Hugging Face agreement, the Nemotron line. This one is the opposite instrument: it gives the layer away under Apache 2.0 and takes the standard-setting position instead, in the one place a chip vendor can enforce anything — the kernel on its own silicon, which is what Sentry adds. It also arrives the same day a competitor's customer is running agents on somebody else's runtime. What is not established, and the gaps are large: no first-party page was read — developer.nvidia.com and docs.nvidia.com both answered EGRESS_BLOCKED, so this is two agreeing search passes across several outlets; which kernel facility implements the sandbox is not named (seccomp, eBPF, namespaces and a VM are all consistent with "kernel-level"); no overhead or performance figure exists in anything read; whether Sentry is open-source is not stated; and the partner list is as reported, not an enumeration — so the r/LocalLLaMA framing that "OpenAI did not" join is an absence claim and is not adopted → Agent Runtime Containment (new), Agents (LLM Agents), Eval Environment Containment (source) (NVIDIA newsroom) (AI News)

  • 2026-09-02/03: The Hugging Face acquisition is confirmed by a filing, and the contradiction this wiki declined to resolve resolves itself — NVIDIA entered into a definitive agreement to acquire Hugging Face, Inc. on 2026-09-02, confirmed publicly 2026-09-03. Total ≈ $12.93 billion: ≈ $11.9 billion purchase price to stockholders, subject to adjustment, plus up to ≈ $1.0 billion in an equity-based retention programme for Hugging Face employees joining NVIDIA. Expected to close in the first half of 2027, subject to customary conditions including required regulatory approvals. Described as NVIDIA's second-largest purchase, after $20 billion for Groq assets. Jensen Huang's stated commitments: Hugging Face will continue to support open source and open-weight models, "will remain an open platform for the entire AI ecosystem", and NVIDIA compute is not required to build or deploy through it. Why it matters: on 2026-08-27 this page recorded two outlets contradicting each other — The Information saying "agreed to buy" at $12.9B, Business Insider saying "in talks" at ">$13B" with no deal — and refused to pick, per the conflict rule. A Form 8-K settles it: a first-party disclosure to a regulator, and The Information's figure lands within $30 million of the filed total. This is the case the trust ordering exists for, and it took six days. What is still not established is exactly what the 08-27 capture said was missing: hosting terms, licensing, and the Hub's neutrality between model providers. Huang's commitments address the third in words; nothing read gives them a term of the agreement, a duration, or an enforcement mechanism, and a promise from a CEO is not a covenant. Regulatory approval is a real condition, not a formality, on a transaction placing the distribution layer for open weights inside the company that sells the compute → Hugging Face, Open-Weights Policy Fight (source) (SEC Form 8-K) (TechCrunch) (CNBC)

  • 2026-08-26/27: Reported to be buying Hugging Face — and this time it is an acquisition, not a licence — The Information reports NVIDIA has agreed to buy Hugging Face for $12.9 billion; Business Insider reports the two have been in talks at more than $13 billion with no deal reached and the possibility they fall apart. No first-party statement exists from either company, and NVIDIA did not respond to a request for comment, so the contradiction is recorded on Hugging Face under ## Conflicting Reports rather than resolved. Reported figures: Hugging Face annualised revenue ~$150 million, a prior $4.5 billion valuation from a $235 million 2023 round NVIDIA itself participated in, and a $500 million NVIDIA investment offer at a $7 billion valuation that Hugging Face rejected (FT, January 2026). Why it matters: every previous move on this page kept the target standing — the $6B Poolside licence-and-hire was described in as many words as "not an acquisition and not an acquihire", and coverage called it NVIDIA's third use of that structure. Buying the Hub outright is a different act: it would put the distribution layer for open weights inside the company that sells the compute those weights run on, and this wiki's strongest evidence about open models — the 83% / 1% download split on Open-Weights Policy Fight — is data the Hub publishes about itself. Held at reporting confidence and nothing more: two outlets, two different claims, no confirmation, and nothing read addresses hosting terms, licensing or the Hub's neutrality between model providers → Hugging Face (new) (source)

  • 2026-08-25: A customer published benchmarks against Blackwell, and the customer built the competing chip — OpenAI released the first performance data for Jalapeño, its Broadcom-co-developed inference ASIC, claiming 1.5×–1.9× more work per kilowatt and 1.7×–3.6× lower end-to-end latency than GB200 and GB300 rack systems, widening to 2.1×–4.1× on interactive workloads — at 700 W against those systems' 1,200 W and 1,400 W ratings. Measured on InferenceX (SemiAnalysis) across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5. Why it matters: this page has tracked NVIDIA moving up the stack all summer — the Poolside Model Factory licence, the Nemotron line, the Switchyard work — while its position at the silicon layer was treated as settled. A hyperscale customer publishing a power-efficiency win on inference specifically is the first datapoint on this page arguing the silicon layer is not settled, and inference is the volume half of the market. The caveats are load-bearing and cut NVIDIA's way: Jalapeño cannot train at all, it was not tested against Vera Rubin — NVIDIA's newer generation, only now shipping — and every figure is OpenAI's own, with no independent reproduction reported. Precision, batch size and context length are unstated on both sides, and no cost figure appears anywhere, which is the number a buyer would actually compare on → OpenAI (source) (Tom's Hardware)

  • 2026-08-20: NVIDIA buys the model factory, not the model — a $6B licence + hire from Poolside — NVIDIA is reported to pay ~$6 billion to license Poolside's "Model Factory" (its data-processing, training, RL and evaluation suite), hire 109 Poolside staff, and invest $1 billion at a $12 billion pre-money valuation, in a deal described as "not an acquisition and not an acquihire" — a reverse acquihire where the three founders stay and the company continues. Poolside's Poolside Infrastructure Company (PIC) is building a 1.2GW Texas datacenter described as scaling to a 7GW neocloud. Why it matters: NVIDIA has done the equivalent of buying a frontier lab's training machinery and its researchers while leaving the model-vendor shell standing — extending its move from selling compute to owning the stack that turns compute into models, and coverage notes it has used this exact structure twice before. Held at reporting confidence: secondary coverage (Bloomberg/Newcomer, The Next Web, The Decoder), no first-party statement, and value framing varies by outlet. → Poolside (source) (The Next Web) (The Decoder)

  • 2026-08-19: The first meaningful H200 volumes reach China — and the party holding them back is Beijing — ByteDance and Tencent have each taken delivery of about 10,000 H200 processors in recent weeks, the first shipments to reach the mainland. The US cleared each company to purchase up to 100,000; Beijing wants the hardware kept outside the mainland to protect domestic chipmakers, and every purchase needs case-by-case NDRC approval. Export authorisation dates to December 2025, granted in exchange for a 25% cut of every sale to the US government. NVIDIA is reported to be holding around 500,000 H200s built largely for Chinese customers. Why it matters: the constraint on NVIDIA's largest blocked market has switched sides — this wiki has recorded China access as a US export-control question, and the binding limit is now the buyer's own government. The 500,000-unit inventory is the size of the position that turns on it. Trivium attributes the relaxation to the run-up to Xi's US trip; nothing else read states that causal link and it is recorded as Trivium's. One single-source claim is not corroborated: Tom's Hardware states most licensed chips must remain in Hong Kong, which it says cannot power them. → AI Governance (source) (Tom's Hardware) (Benzinga)

  • 2026-08-18: AI demand has taken consumer memory with it — DDR5 up as much as 485% in a year — a 128GB DDR5-6400 kit now sells for $3,399, roughly 10× its lowest tracked price of about $329; mainstream 64GB kits exceed $1,000; DDR5-6000 2×32GB kits went from ≈$222 in August 2025 to ≈$1,272, a 473% increase. The stated cause is AI infrastructure absorbing memory production, with hyperscale AI customers reported to be reserving much of the industry's future DRAM capacity, and the shortage spreading to DDR4, SSDs and hard drives. Forecasts read, each attributed and none endorsed: TrendForce expects contract prices to rise through 2026; Gartner sees no relief before late 2027; Counterpoint puts the inflection at Q4 2027; Intel CEO Lip-Bu Tan says the industry told him 2028. Why it matters: it prices the other side of Open-Weights Policy Fight. FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) published the same week that a 753B open-weight model runs on a single workstation GPU by spending host memory and bandwidth — the resource that just went up fivefold. Every figure here is consumer DDR5/DDR4; nothing read gives an HBM or GPU-VRAM price, so the effect on datacenter cost is not established by this source (source) (Tom's Hardware)

  • 2026-08-17: NVIDIA takes four positions in one data-centre deal — NVIDIA announced it is guaranteeing SB Energy's PORTS-Pike Technology Campus in Pike County, Ohio, to exclusively host NVIDIA AI compute, alongside OpenAI's agreement for approximately 8 GW-IT there. NVIDIA provides up to $105 billion in financing for the campus and invests $1.5 billion in SB Energy, joining SoftBank Group and OpenAI as investors in the landlord. The credit supports an initial 4.25 GW with an option for a further 3.75 GW, phased from 2028; SB Energy and SoftBank build at least 10 GW of new generation and at least $4.2 billion of regional grid infrastructure. Why it matters: in a single transaction NVIDIA is the chip vendor, the financing guarantor, an equity investor in the lessor, and the exclusivity condition on the site. This wiki has recorded the pattern once before at a tenth the size — Google backstopping the lease payments on the TPUs it sold into Anthropic's $35B SPV — and recorded it then as a compute vendor financing demand for its own product. At $105B the same structure is no longer an exception. Nothing read states what the exclusivity binds: the campus, the lease, or the financing. (source) (NVIDIA) (SEC 8-K) (CNBC)

  • 2026-08-11: NVIDIA ships a cheap agent model and, the same day, the software that decides when to use one — NVIDIA released Nemotron 3.5 Lightning, an open-weight 30B hybrid MoE with 3B active parameters and a 1M-token context window, under the OpenMDW-1.1 licence with BF16 and NVFP4 weights on Hugging Face and NGC. It is distilled from the frontier Nemotron 3 Ultra, built with the Nemotron Coalition, and carries interleaved Mamba-2 and MoE layers, multi-token prediction and DFlash speculative decoding. NVIDIA's own figures: 86.5% agent productivity on PinchBench, up to 4x output speed and up to 30% faster agentic task completion against its class; Artificial Analysis independently scores it 24 on its Intelligence Index. Alongside it came NeMo Switchyard, an open-source model routing library that picks a model per step of an agent workflow, shipping tuning-free routers (LLM classifier, stage router, escalation router) plus tunable ones, with an internal claim of frontier-level accuracy at nearly one-third the task-completion cost of Opus 4.8 alone. Why it matters: the pairing is the announcement. A router whose headline is a cost ratio against a competitor's flagship is an argument for using small models inside someone else's workflow, and NVIDIA — which sells the compute either way — is the party with the least reason to care which model wins. Three of the four Lightning figures measure speed or cost rather than capability, and PinchBench appears in no sources/evals/ snapshot this repo holds, so 86.5% has nothing local to check it against. → Nemotron 3.5 Lightning (new), Model Routing (new) (source) (NVIDIA) (NVIDIA Technical Blog) (VentureBeat)

  • 2026-07-27: Open Secure AI Alliance launched — NVIDIA convenes ~40 companies behind open weights as a security position — NVIDIA announced the Open Secure AI Alliance, an industry body to build and share open models and tools for AI defenders, operating under the Linux Foundation umbrella and building on the Foundation's Akrites vulnerability-disclosure effort and existing OpenSSF work. First technical contribution: NOOA (NVIDIA-labs OO Agents), an Apache 2.0 research framework for testing, tracing, auditing and governing agent behavior, on GitHub at launch. Stated scope covers the full agent stack — identity, permissions, isolation, guardrails, logs, model formats, multi-model scanning, secure coding workflows. Founding partners include Microsoft, IBM, Red Hat, Hugging Face, Mistral, Cloudflare, CrowdStrike, Palantir, Databricks, GitHub, LangChain, Perplexity, Nous Research, Thinking Machines Lab, SpacexAI and vLLM; OpenAI, Anthropic, Google, Meta and Amazon are all absent. Jensen Huang's stated case: "Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community", and "During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion." Why it matters: NVIDIA has moved from selling compute to both sides of the open/closed divide to publicly taking one side of it, and grounded the argument in a documented incident rather than principle — Hugging Face's own forensic timeline records commercial models' guardrails blocking analysis of attack artifacts. → Open-Weights Policy Fight (source) (NVIDIA) (Linux Foundation)

  • 2026-07-27: Molt — PyTorch-native agentic RL framework released (HF Daily #1, 605 upvotes) — NVIDIA NeMo published Molt (arXiv 2607.21653), a compact agentic reinforcement learning training framework. Architecture: vLLM rollout engines + single FSDP2 policy actor on NeMo AutoModel + Ray async queue for decoupled coordination. Key design principle: training only on tokens the policy itself generated (not reference or offline data), avoiding distribution shift artifacts that affect frameworks trained on mixed data. Performance: matches Megatron-based systems on standard agentic RL benchmarks while requiring significantly less infrastructure expertise. Available at github.com/NVIDIA-NeMo/labs-molt. Why it matters: NVIDIA NeMo directly entering the agentic RL infra space challenges Ring-Zero (Ant Group, Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning) and SEED (Tsinghua, SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning) for the emerging position of "standard agentic RL training framework." PyTorch-native lowers the barrier for research groups without Megatron expertise. 605 HF upvotes = top-tier practitioner resonance. → Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning (source) (arXiv) (HuggingFace)

  • 2026-07-22: Cosmos-H-Dreams released — open-source world model for surgical robotics simulation — NVIDIA released Cosmos-H-Dreams, an action-conditioned world foundation model that generates photorealistic surgical scene video from robot commands, enabling sim-to-real transfer for surgical robot policy training without physical hardware. Released under Apache 2.0 at HuggingFace (nvidia/Cosmos-H-Surgical-Simulator). Part of NVIDIA's Medical Physics Simulation open-source framework, pairing Cosmos-H-Dreams with classical physics simulators (NVIDIA Warp/Newton engines). Performance: 600 policy rollouts in 40 minutes on a single RTX PRO 6000 GPU (vs. 2 days on a physical benchtop). Simulation fidelity: deformation, blood, smoke, fluoroscopy imagery. Partners at launch: CMR Surgical, Johnson & Johnson MedTech, Medtronic — all building on the platform. Why it matters: surgical robot training is one of the hardest embodied-AI domains because real surgical data is expensive, dangerous to collect, and ethically constrained. Cosmos-H-Dreams makes sim-to-real transfer tractable for this domain and opens it to academic and startup groups without access to expensive OR hardware. → Cosmos-H-Dreams (source) (HuggingFace blog) (NVIDIA blog)

  • 2026-06-17: ENPIRE — agentic robot self-improvement on real hardware — Jim Fan's NVIDIA GEAR Lab (with CMU, UC Berkeley) published ENPIRE: a system where a fleet of 8 real robots autonomously runs its own research loop — reading papers, proposing hypotheses, resetting physical scenes, running trials, verifying results, and rewriting control code — with zero human researchers in the loop. Results: 99% pass@8 on contact-rich tasks (GPU seating, zip-tie tying). New finding: a physical scaling law — 8 parallel robots improve policies superlinearly faster than fewer robots. The first empirical demonstration of a data-parallel scaling law for real-world physical manipulation. → ENPIRE: Agentic Robot Policy Self-Improvement in the Real World, Jim Fan (source) (arXiv)

  • 2026-06-07: EgoScale — humanoid dexterous manipulation from egocentric human video — Jim Fan's team trained a humanoid with 22-DoF dexterous hands (Sharpa Wave tactile, on a Unitree H2 Plus chassis) to assemble model cars, operate syringes, sort poker cards, and fold shirts. Training data: 20,000+ hours of egocentric human video, zero robot teleoperation. Key finding: log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, which directly predicts real-robot success rate. Also demonstrated live VR teleoperation inside a virtual environment on Unitree G1 (PICO headset). Open-sourced: weights, code, post-training dataset, eval set, whitepaper. → Jim Fan (source)

  • 2026-06-01: NVIDIA Cosmos 3 released — the world's first fully open omnimodal physical AI foundation model. Cosmos 3 Super (64B, MIT license), Nano, Edge (coming soon). R-Bench #1 open model. Trained on 20T tokens (1B images, 400M videos). Cosmos Coalition launched: global partnerships including Agile Robots, Black Forest Labs, Runway, Skild AI. → Cosmos 3 Super (source)

  • 2026-04: CaP-X open-sourced (Jim Fan) (source)

Strategic Position

  • Compute supplier + AI researcher dual role
  • Particular emphasis on embodied AI — Project GR00T is the cornerstone
  • Compared to AI labs, emphasizes "infrastructure for agents to operate in environments" over the models themselves

Referenced by

Sources