$ cat wiki/models/embeddinggemma-2.md
EmbeddingGemma 2
Compared with
- Mistral Large 4
- Kolibri-1
- GPT-6.1 Sol
- Claude Sonnet 5.5
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Ling-3.0-tiny
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- Gemini 4 Argon
- Claude Opus 5.5
- GPT-6 Luna
- GPT-6 Sol
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Grok Build
- Claude Opus 4.7
- Muse Spark
An on-device multimodal embedding model from Google DeepMind, released on 2026-10-06. It maps text, code, images, audio and video into a shared embedding space. (source)
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-10-06 |
| Announced | 2026-10-06 |
| Context window | 8K tokens |
| Pricing | unknown |
| License | Apache 2.0 |
| Availability | Weights on Hugging Face and Kaggle; local inference |
| All established specifications come from Google's announcement. No hosted API price is stated there. (source) |
Release Date
Google announced the release on 2026-10-06 and linked downloadable weights. Model Garden availability is described as forthcoming. The announcement identifies Gemma 4 as the architectural basis. (source)
Benchmarks
Google reports MTEB Code improving from 68.76 for EmbeddingGemma to 78.68, a 9.92-point gain. It also claims leading performance among sub-1B multimodal embedders; the captured article gives no numeric audio or image table, so no such scores are asserted here. (source)
Use Cases
The 740M full model comprises 270M text parameters and optional 170M vision and 300M audio encoders. Output vectors can shrink from 768 dimensions to 512, 256 or 128. Google reports quantized weight memory of approximately 191MB for text and 567MB for the full model on a Pixel 11 Pro; these are device-specific weight-memory figures, not a universal application-memory budget. (source)
The context supports up to 5.5 minutes of audio, 29 images, or 58 video frames, including interleaved inputs. Suggested uses include private local retrieval, semantic media search and routing. (source)
Compared To
The predecessor comparison is task-specific: MTEB Code and context length are reported, but the article does not establish a single aggregate improvement across every modality. This model generates embeddings rather than conversational answers. (source)