AI Trend Notifier
EN한
← wiki

$ cat wiki/models/gemini-omni.md

Gemini Omni

modelupdated 2026-08-28created 2026-05-19

Spec

AttributeValue
DeveloperGoogle DeepMind
Releasednot yet
Announced2026-05-19 (Google I/O 2026)
Context windowunknown (no detailed technical specs released at I/O)
Pricingunknown
Licenseunknown
AvailabilityNot yet available; expected via Gemini app and API (rollout timeline TBD)
TypeMultimodal video generation model
InputAny: image + audio + video + text
OutputVideo
This page is the I/O announcement and stays one. The Omni line has since
shipped: Gemini Omni 1.1 Flash was released 2026-08-27 with a
per-second price and a stated availability surface. The Released: not yet above
is still correct for the model announced at I/O — nothing read states that the
I/O-announced Gemini Omni and the shipped Omni Flash line are the same artefact
(source).

Open item, recorded rather than guessed: the Omni Flash line shipped at some point between 2026-05-19 and 2026-08-27, and this wiki captured neither that launch nor the 1.0 release. Nothing in the 08-27 capture dates it. The gap is carried on Google DeepMind.

Key Differentiator vs. Veo

VeoGemini Omni
InputTextAny (image + audio + video + text)
GroundingGeneral trainingGemini's real-world knowledge base
Use caseText-to-video creationMulti-modal video editing & synthesis
Gemini Omni represents "a leap forward in world understanding, multimodality and editing" — Google's framing at I/O.

Significance

The multimodal video generation space (Sora, Runway, Kling, Veo) is heating up significantly. Omni's any-input approach signals Google moving beyond text-to-video toward world model-grounded video synthesis — closer to Karpathy's "world simulator" framing than pure generative art tools.

Sources

Referenced by

Sources