DeepMind releases EmbeddingGemma 2, a 740M open embedder
Google DeepMind released EmbeddingGemma 2, a 740M Apache 2.0 model that embeds text, code, images, audio, and video.
Google DeepMind released EmbeddingGemma 2 on 6 October 2026, a 740-million-parameter Apache 2.0 multimodal embedding model built on Gemma 4, with weights on Hugging Face and Kaggle. Unsloth also shipped Dynamic 3.0 GGUF weights that trended at number 29. The model maps text, code, images, video, and audio into one 768-dimensional space with an 8,192-token context; text use is about 270 million parameters, with optional 170-million-parameter vision and 300-million-parameter audio encoders, and Matryoshka truncation to 512, 256, or 128 dimensions, which Unsloth says can cut storage by up to 6x. Full-precision MTEB Code is 78.68, up 9.92 points from 68.76; Unsloth calls that baseline EmbeddingGemma 1 and also cites multilingual MTEB 61.36 and more than 100 languages, while Hacker News and DeepMind name the baseline EmbeddingGemma and say it leads sub-1B multimodal embedders on several benchmarks. Quantized weights need about 191 MB RAM for text-only and about 567 MB for the full model on a Pixel 11 Pro; The Decoder instead cites roughly 191 MB, local latency of about 20 to 70 milliseconds, a 270-million-parameter text-only variant, and Google's claim that it is the most compact model of its kind and outperforms rivals up to twice its size. A 9 October Hugging Face card for an unrelated AtomicChat Qwen-Image encoder is not part of this release.
- Google DeepMind released EmbeddingGemma 2 on 6 October 2026: 740 million parameters, Apache 2.0, built on Gemma 4, with weights on Hugging Face and Kaggle.
- Unsloth published Dynamic 3.0 GGUF weights that trended at number 29 on Hugging Face the same day.
- It maps text, code, images, video, and audio into one 768-dimensional space with an 8,192-token context; the text stack is about 270 million parameters, with optional 170-million-parameter vision and 300-million-parameter audio encoders.
- Matryoshka truncation supports 512, 256, and 128 dimensions; Unsloth says that can cut vector storage by up to 6x.
- Full-precision MTEB Code is 78.68, up 9.92 points from 68.76; Unsloth calls the baseline EmbeddingGemma 1 and also cites multilingual MTEB 61.36 and 100-plus languages, while Hacker News and DeepMind name the baseline EmbeddingGemma.
- On a Pixel 11 Pro, quantized weights need about 191 MB RAM for text-only and about 567 MB for the full model; The Decoder cites roughly 191 MB and local latency of about 20 to 70 milliseconds.
- The Decoder reports Google's claim that it is the most compact model of its kind and outperforms rivals up to twice its size, and describes a 270-million-parameter text-only variant.
Coverage timelineoldest first · each row is one article
- · 4d agounsloth/embeddinggemma-2-GGUF — new model trending #29 on Hugging Face
Hugging Face trending models· 52
Unsloth released GGUF weights for Google DeepMind's EmbeddingGemma 2, a 740M multimodal embedding model.
- · 4d agoEmbeddingGemma 2: An open, lightweight multimodal embedding model
Hacker News · AI· 66
Google DeepMind released EmbeddingGemma 2, a 740M open multimodal embedding model for on-device search.
- · 4d ago