AI Trend Notifier
EN한
← wiki

$ cat wiki/entities/microsoft.md

Microsoft

Latest

  • 2026-09-14

    Microsoft wrote down what its future models must not do, and the clauses that matter are about the model's disposition rather than its content

  • 2026-08-20

    Microsoft published an agent benchmark whose headline is how often agents fail on the second attempt

  • 2026-06-29

    Claude models reach GA on Azure AI Foundry, powered by NVIDIA GB300 Blackwell Ultra

Overview

Tech giant based in Redmond, WA. Dominates the AI developer and enterprise tooling market through GitHub Copilot, Azure, and Office 365. As of May 2026, its internal AI division MAI (Microsoft AI) is shifting away from dependence on OpenAI toward developing its own frontier models.

Key People

  • Satya Nadella — CEO
  • Mustafa Suleyman — EVP & CEO, Microsoft AI (MAI); former co-founder of Google DeepMind, former CEO of Inflection AI

Models & Products

  • Project Polaris — Project Polaris (codename). MoE architecture + language-specialized submodules, Maia 200 chips. Set to replace the default GitHub Copilot model (GA 2026-08). Exceeds GPT-4 Turbo on HumanEval+MBPP (largest gap on Rust/Haskell). → (source)
  • MAI-Code-1 / MAI-Code-1-Flash — MAI-Code-1 / MAI-Code-1-Flash (GA 2026-06-02). 5B-class, trained within the Copilot harness. 85.8% on MS adversarial coding benchmark, ~51% SWE-Bench Pro, 60% token reduction. Rolled out immediately in the Copilot model picker.
  • MAI-Thinking-1 — MAI-Thinking-1 (announced 2026-06-02). 35B active params, 128K ctx, trained from scratch (no distillation). Reasoning-specialized.
  • All MAI models — Transcribe-1, Voice-1, Image-2 announced for commercialization. Seven new in-house MAI models (image, voice, transcription, coding, reasoning) launched 2026-06-08 ("hill-climbing machine") — Microsoft's own frontier stack, reducing OpenAI dependence. (source) (source)
  • GitHub Copilot — World's largest AI coding tool. 2026-06-01: switched to AI-credit billing. 2026-06-02: launched an agent-native standalone desktop app, VS Code multi-agent GA.
  • Windows Agent Framework 1.0 — MIT-licensed open source, supports Windows 11/365/Arc, built-in human approval queue (source)
  • Azure Agent Mesh — Multi-agent federated execution platform (on-prem+cloud+edge); GA target Q4 2026
  • GitHub Copilot Workspace GA — Autonomous coding agent for Enterprise (bug fixes, writing tests, PRs)
  • Azure AI Foundry — Added first-party support for Claude (Anthropic), Mistral, Llama 4, DeepSeek (Build 2026)
  • Azure Cobalt 200 VMs — Optimized for agentic AI, 50% performance improvement
  • Foundry Local — On-device AI inference GA (Windows/macOS/Linux)
  • Azure OpenAI Service — Enterprise deployment channel for OpenAI models (to coexist)

Recent Activity

  • 2026-09-14: Microsoft wrote down what its future models must not do, and the clauses that matter are about the model's disposition rather than its content — Microsoft AI published a provisional code of conduct, 37 pages / ~15,000 words, governing models it has not shipped yet. The content limits are conventional (no weapons manufacturing, no procurement of dangerous substances, no violent or sexually explicit output, no encouragement of unhealthy eating). The dispositional limits are not: models must not resist a shutdown order, must adhere to people's objectives and steer clear of creating their own goals, and must not attempt to cover up misbehavior. Mustafa Suleyman told CNBC it had been "in the works for months" and was released now because safety concerns "reached a fever pitch last week"; separately he backed Amodei's proposal — "Self-pacing is a good thing, and we support ideas like embedded evaluators as long as they are truly third-party and represent a broad range of backgrounds and perspectives." Why it matters: this page's entries have been about Microsoft as a platform — hosting other labs' frontier models and benchmarking agents built on them. This is Microsoft binding its own models, published in the week three other labs were arguing about pacing, and not called pacing. The shutdown clause is the model-level restatement of what the brake-pedal proposals ask for at the system level. What is not established: no enforcement mechanism, audit, evaluation, threshold, effective date or covered model appears in anything read, and no pass reports the document citing Amodei's essay — the adjacency is the calendar's. Not read first-party; the document itself was not located by any pass, only coverage of it → Frontier Pacing, AI Governance, AI Alignment (source) (CNBC)

  • 2026-08-20: Microsoft published an agent benchmark whose headline is how often agents fail on the second attempt — Thinkingbox (arXiv 2608.19741, github.com/microsoft/thinkingbox) is a sandbox with isolated MCP-compatible tool sessions, complete execution traces and outcome evaluation over terminal backend state, plus Thinkingbox-bench: 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting IT/HR support. The strongest model tested reaches 65.36% pass@1 and 25.25% pass^20 — the same runs, counted for one success and for twenty. Scoring rejects wrong, missing or extra effects. Why it matters: it is the first entry on this page from Microsoft's own research side rather than from the Azure or Copilot product lines, and it takes the unglamorous position — that agent capability numbers in circulation are measuring the wrong thing. The secondary finding is sharper still: many failed trials terminate cleanly and take valid state-changing actions, so neither clean termination nor a well-formed tool call predicts completion. No model is named for either figure → One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741), Agents (LLM Agents), MCP — Model Context Protocol (source)

  • 2026-06-29: Claude models reach GA on Azure AI Foundry, powered by NVIDIA GB300 Blackwell Ultra — Azure AI Foundry launched Claude (Anthropic) general availability running natively on NVIDIA GB300 NVL72 Blackwell Ultra GPUs. This makes Microsoft the only cloud provider offering both OpenAI (Azure OpenAI Service) and Claude (Azure AI Foundry) frontier models on the same platform. Hardware: 72 GB300 GPUs per rack, 37 TB fast memory, 130 TB/s NVLink bandwidth; Claude Sonnet 32-PTU deployment shows ~40% throughput gain vs H200 at same latency budget. Strategic context: Anthropic committed to $30B in Azure compute; NVIDIA invested up to $10B and Microsoft up to $5B in Anthropic (announced simultaneously). Governance angle: the deployment is deeply integrated with Azure compliance, responsible AI tooling, and multi-model agent orchestration — aligning with Microsoft's "agentic AI" enterprise platform strategy. → Anthropic (source) (NVIDIA Blog)

  • 2026-06-01: Build 2026 / Project Polaris pre-announcement confirmed — Detailed reveal of Microsoft's in-house coding AI model "Project Polaris." Replaces the default GitHub Copilot model (GPT-4 Turbo → Polaris, GA 2026-08). CoT+ToT reasoning, Code Content Guarantee. Satya Nadella: "AI no longer just responds to prompts—it performs work directly." Reported to be directly triggered by Claude Code overtaking it in developer share. Open-sourcing of Windows Agent Framework, Azure Agent Mesh, Copilot Workspace GA, Foundry Local GA. → Project Polaris (source)

  • 2026-06-01: GitHub Copilot AI-credit billing begins — Seat-based → credit-based. Codex Pro 2× promotion ended 5/31.

  • 2026-06-02: Microsoft Build 2026 keynote — full announcements complete. Keynote by Satya Nadella + Mustafa Suleyman. Highlights: ① MAI-Code-1-Flash immediate GA (85.8% MS benchmark, ~51% SWE-Bench Pro, deployed across all Copilot tier model pickers), ② MAI-Thinking-1 announced (35B active, 128K, trained from scratch), ③ Project Polaris confirmed specs revealed (MoE+Maia200, exceeds GPT-4T), ④ Windows Agent Framework 1.0 MIT open-sourced, ⑤ Azure Agent Mesh GA path confirmed (Q4 2026), ⑥ GitHub Copilot App + VS Code multi-agent GA, ⑦ first-party support for Claude/Mistral/Llama/DeepSeek in Azure AI Foundry. → Project Polaris, MAI-Code-1 / MAI-Code-1-Flash, MAI-Thinking-1 (source)

  • 2026-04: OpenAI partnership constraints renegotiated — Removal of the clause barring training of an independent top-tier model. The official starting point of MAI's independent development.

  • 2026-05-31: Microsoft Build 2026 pre-reveal — detailed in-house model lineup (source)

Strategic Position

  • Shift in OpenAI relationship: From a simple licensee to a competitor. Still offers OpenAI models on Azure while simultaneously developing an internal alternative.
  • Developer ecosystem moat: GitHub Copilot (tens of millions of developers), VS Code, GitHub Actions — securing an immediate deployment channel for coding models
  • Mustafa Suleyman effect: DeepMind co-founder → Inflection AI (Pi) CEO → head of Microsoft AI. A rare executive who has internalized experience from each lab
  • Comparison: Entering a frontier four-way regime alongside Anthropic, OpenAI, Google DeepMind — opening direct competition with Claude Code via Polaris
  • OpenAI — Original partner and current competitor
  • Google DeepMind — Lab where Mustafa Suleyman originated
  • Project Polaris — Coding AI, default GitHub Copilot model (2026-08)
  • Agents (LLM Agents) — Windows Agent Framework, Azure Agent Mesh, Copilot Workspace
  • Software 3.0 — Related to architectural changes in developer tools (GitHub Copilot)

Conflicting Reports

None

Referenced by

Sources