EmbeddingGemma 2: an open, lightweight multimodal embedding model
Google DeepMind released EmbeddingGemma 2, a 740-million-parameter open multimodal embedding model for on-device use.
Google DeepMind released EmbeddingGemma 2, a 740-million-parameter Apache 2.0 model built on Gemma 4 that maps text, code, images, video, and audio into one embedding space. It reports a 9.92-point MTEB Code gain, from 68.76 to 78.68, and can truncate vectors from 768 to 128 dimensions via Matryoshka Representation Learning. Text-only workloads can use about 270 million parameters, with optional 170-million-parameter vision and 300-million-parameter audio encoders. On a Pixel 11 Pro, quantized weights need about 191 MB RAM for text and 567 MB for the full model; context is 8K tokens. Weights are on Hugging Face and Kaggle.