$ cat wiki/entities/ai2.md
Ai2 (Allen Institute for AI)
Latest
- 2026-10-01
Olmo-core 3 — the open part is the training stack, not a model.
- 2026-08-07
TutorMoments — a tutoring benchmark that scores holding back, not just helping
Overview
Ai2, the Allen Institute for AI, is a research institute that publishes open models, datasets and benchmarks. This page was created on 2026-08-09 on the release of TutorMoments, and its contents are limited to what that release and its coverage stated (source).
Ai2's founding, funding, headcount and its wider model lines were not established on the run that created this page — the sandbox could not fetch any primary page that day (seventh consecutive day of blocked egress), so nothing beyond the TutorMoments release is recorded here. This is a stub created because the institute is a recurring publisher of open evaluation work that this wiki had no page for, not because one benchmark warrants an entity.
Key People
unknown — no named individual appeared in any source read on the run that created
this page.
Models & Products
- TutorMoments — open, replay-based benchmark for language-model tutors, preview released 2026-08-07 (see Recent Activity). No wiki page; recorded here rather than as a standalone page, per the one-off-mention rule.
Recent Activity
-
2026-10-01: Olmo-core 3 — the open part is the training stack, not a model. Ai2 released a redesigned mixture-of-experts training infrastructure under Apache-2.0 at
github.com/allenai/OLMo-core, aimed at scaling MoEs into the trillion-parameter range. It replaces FSDP with distributed data parallelism, keeping experts resident on GPUs, and adds rowwise expert parallelism, GPU-resident routing, grouped GEMM and MXFP8 low-precision support (source). Reported figures: a 47B-parameter MoE on NVIDIA B300 GPUs at 52,000 tokens/sec/GPU against 19,400 previously; with MXFP8, throughput ~21% above the BF16 baseline while peak active memory fell from 103 GiB to 95 GiB (four B300s, work distributed uniformly across experts); and a 1.2-trillion-parameter model trained across 512 GPUs at a sustained 858 TFLOP/s per accelerator. Stated forward plan: the next-generation Olmo will be an MoE, on Ai2's largest dataset and longest context window — no date, parameter count or benchmark for it.allenai.organsweredEGRESS_BLOCKED— newly recorded as blocked for this pipeline — andhuggingface.co, which carries the same post as prefetch candidate #59, is blocked by standing policy and was not attempted. So this is two agreeing search passes, not a first-party read, and the GPU count for the 52,000-vs-19,400 comparison is not adopted: one pass said eight B300s, the other described the MXFP8 benchmark on four, and the two may be different benchmarks. Not stated anywhere read: total compute or cost for the 1.2T run, whether those weights will be released, any loss or quality figure, or any hardware other than NVIDIA B300 — so the 858 TFLOP/s is a throughput result and not evidence that a 1.2T Olmo exists as a model -
2026-08-07: TutorMoments — a tutoring benchmark that scores holding back, not just helping — Ai2 released a preview of TutorMoments, an open, replay-based benchmark measuring whether a language-model tutor correctly chooses between helping a student and letting the student reason. TutorMoments-Preview is 462 de-identified, text-only transcripts of real one-on-one US maths tutoring with students in grades 2–7. Ai2's argument is that existing tutoring benchmarks reward a single fixed behaviour — never revealing the answer, or always offering a hint — "without accounting for whether that was the right move for where the student actually was in their understanding"; being replay-based lets TutorMoments judge an action against the actual state of the session. Preliminary results are reported as models tending to over-help. No per-model scores, leaderboard or numeric results were established from anything read. → Reasoning Models (source) (Ai2) (Hugging Face)
Strategic Position
Ai2's lever is the part of the stack other labs keep. The two artefacts this page holds are an open replay-based benchmark (TutorMoments, 2026-08-07) and an Apache-2.0 MoE training stack (Olmo-core 3, 2026-10-01) — an evaluation and an infrastructure, not a frontier model (source). Olmo-core 3 makes the position explicit: it is released so that others can train their own MoEs, adapt it to different hardware, and experiment with routing and parallelism, which is a bid for the layer beneath the models rather than for the model leaderboard. The weight-releasing labs this wiki tracks publish weights; Ai2 here publishes the thing that produced them.
The limit is equally visible from the same source. The 1.2-trillion-parameter run across 512 GPUs is reported as a throughput result — 858 TFLOP/s per accelerator — with no loss or quality figure and no statement that the weights will be released, and the next-generation Olmo has no date, parameter count or benchmark. So the capability claim is about what the stack can drive, and whether Ai2 converts it into a competitive open model is unestablished here.
This section was an annotated placeholder from the run that created this page until 2026-10-02, on the stated grounds that nothing read spoke to Ai2's positioning. Olmo-core 3 does, and is the whole basis for the above.
Related
- Reasoning Models — the over-helping result is about when a model should decline to supply reasoning