Unsloth released GGUF weights for Google DeepMind's EmbeddingGemma 2, a 740M multimodal embedding model.
Unsloth published Dynamic 3.0 GGUF weights for Google DeepMind's EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0. It maps text (including code), images, video, and audio into one 768-dimensional space, with optional 170M vision and 300M audio encoders around a 270M text stack. Context is 8,192 tokens, with Matryoshka truncation to 128, 256, or 512 dimensions for up to 6x less vector storage. Full-precision scores include multilingual MTEB 61.36 and code MTEB 78.68, versus 68.76 for EmbeddingGemma 1.
740M-parameter model unifies text, image, video, and audio in a 768-d space.
Text stack is about 270M; vision 170M and audio 300M encoders load optionally.
Matryoshka truncation supports 128, 256, 512, and 768 dimensions, cutting storage up to 6x.
Code MTEB mean rises to 78.68 from EmbeddingGemma 1's 68.76.
Full article2,804 words · extracted from huggingface.co · click to collapse
# Read our How to [Run EmbeddingGemma 2 Guide!](https://unsloth.ai/docs/models/embeddinggemma-2)
<div>
<p style="margin: 0 0 0px 0; margin-top: 0px;">
<em><a href="https://unsloth.ai/docs/basics/dynamic-3.0-ggufs">Unsloth Dynamic 3.0</a> achieves superior accuracy & outperforms other leading quants.</em>
**EmbeddingGemma 2** is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
* **Native multimodality:** Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
* **Multilinguality and code:** EmbeddingGemma 2 understands 100+ languages, and achieves a \~14% improvement on code tasks relative to its predecessor.
* **Flexible footprint:** Combines a 270M parameter text backbone (130M transformer \+ 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
* **Matryoshka Representation Learning (MRL):** Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a **6x reduction** in vector storage costs with minimal impact on quality.
* **Context length:** 8K token context window, capable of processing minutes of audio or video.
* **Task-steered representations:** Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).
EmbeddingGemma 2 was evaluated across text, code, vision, visual document, video, and audio embedding benchmarks. All results reported below use the full-precision checkpoint.
For optimal embedding quality and runtime efficiency, follow these configurations and best practices:
### 1. Task Instruction Prefixes
EmbeddingGemma 2 is trained with short task instruction prefixes prepended to text inputs. Using the right prefix improves quality; omitting it still works but reduces precision. Prefixes apply to text only. Pass images, video, and audio without any prefix.
Documents with a real title should be formatted as `title: {title} | text: {content}`. Use `title: none` when no title is available.
**Prefix Notation & Usage**
We offer two types of task prefixes, depending on how embeddings are used in the task. There are two types of tasks:
* **Asymmetric Tasks (e.g. retrieval):** Use a query prefix for queries and a document prefix for corpus items.
* **Symmetric Tasks (e.g. classification, similarity)**: Apply the same task prefix to all inputs being compared.
| Use Case | Task Type | Prompt Name | Query Task Instruction | Document Task Instruction *(use `none` if no title)* |
**Please note:** `prompt_name="Document"` applies `title: none`; titled documents must still be formatted manually, like `model.encode(f"title: {title} | text: {document}")`
### 2. Selective Encoder Loading
The vision and audio encoders are independent components. To reduce memory consumption when deploying text-only or single-modality pipelines, disable unused modality encoders via SentenceTransformer’s `config_kwargs`:
**Please note:** configuring EmbeddingGemma 2 to load with omission of some encoders differs among model libraries; please refer to the appropriate documentation for more information on this.
| Active Modalities | `config_kwargs` | Effective Size |