DeepMind releases EmbeddingGemma 2, a 740M open embedder
Google DeepMind released EmbeddingGemma 2, a 740M Apache 2.0 model that embeds text, code, images, audio, and video on device.
Google DeepMind on 6 October 2026 released EmbeddingGemma 2, a 740-million-parameter Apache 2.0 multimodal embedding model built on Gemma 4, with weights on Hugging Face and Kaggle. It maps text, code, images, audio, and video into one space, uses an 8K-token context, and supports Matryoshka vectors from 768 down to 128 dimensions. The MTEB Code score is 78.68, a 9.92-point rise from EmbeddingGemma's 68.76. Text-only workloads use about 270 million parameters and about 191 MB of quantized RAM; optional vision and audio encoders are about 170 million and 300 million parameters, and the full quantized model uses about 567 MB on a Pixel 11 Pro. The Decoder alone reports local inference of about 20 to 70 milliseconds and says Google calls the model the most compact of its kind and claims it beats rivals up to twice its size, while Hacker News says it leads sub-1B multimodal embedders. Those competitive claims are not stated the same way in DeepMind's post, and The Decoder's 191 MB RAM figure does not separate text-only use from the 567 MB full model reported by DeepMind and Hacker News.
- Google DeepMind released EmbeddingGemma 2 on 2026-10-06: 740 million parameters, Apache 2.0, built on Gemma 4, weights on Hugging Face and Kaggle.
- It embeds text, code, images, audio, and video in one space, with an 8K-token context and Matryoshka vectors from 768 down to 128 dimensions.
- MTEB Code score is 78.68, up 9.92 points from EmbeddingGemma's 68.76.
- Text-only use is about 270 million parameters; optional encoders are about 170 million (vision) and 300 million (audio).
- On a Pixel 11 Pro, quantized RAM is about 191 MB for text-only and about 567 MB for the full model.
- The Decoder reports about 20–70 ms local latency and a Google claim of beating rivals up to twice its size; Hacker News says it leads sub-1B multimodal embedders.
Coverage timelineoldest first · each row is one article
- · 2d agoEmbeddingGemma 2: An open, lightweight multimodal embedding model
Hacker News · AI· 66
Google DeepMind released EmbeddingGemma 2, a 740M open multimodal embedding model for on-device search.
- · 2d agoGoogle DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
MarkTechPost· 64
Google DeepMind released EmbeddingGemma 2, a 740M open multimodal embedding model under Apache 2.0.
- · 2d ago