AI Trend Notifier
EN한
← wiki

$ cat wiki/models/embeddinggemma-2.md

EmbeddingGemma 2

Compared with

An on-device multimodal embedding model from Google DeepMind, released on 2026-10-06. It maps text, code, images, audio and video into a shared embedding space. (source)

Spec

AttributeValue
DeveloperGoogle DeepMind
Released2026-10-06
Announced2026-10-06
Context window8K tokens
Pricingunknown
LicenseApache 2.0
AvailabilityWeights on Hugging Face and Kaggle; local inference
All established specifications come from Google's announcement. No hosted API price is stated there. (source)

Release Date

Google announced the release on 2026-10-06 and linked downloadable weights. Model Garden availability is described as forthcoming. The announcement identifies Gemma 4 as the architectural basis. (source)

Benchmarks

Google reports MTEB Code improving from 68.76 for EmbeddingGemma to 78.68, a 9.92-point gain. It also claims leading performance among sub-1B multimodal embedders; the captured article gives no numeric audio or image table, so no such scores are asserted here. (source)

Use Cases

The 740M full model comprises 270M text parameters and optional 170M vision and 300M audio encoders. Output vectors can shrink from 768 dimensions to 512, 256 or 128. Google reports quantized weight memory of approximately 191MB for text and 567MB for the full model on a Pixel 11 Pro; these are device-specific weight-memory figures, not a universal application-memory budget. (source)

The context supports up to 5.5 minutes of audio, 29 images, or 58 video frames, including interleaved inputs. Suggested uses include private local retrieval, semantic media search and routing. (source)

Compared To

The predecessor comparison is task-specific: MTEB Code and context length are reported, but the article does not establish a single aggregate improvement across every modality. This model generates embeddings rather than conversational answers. (source)

Referenced by

Sources