EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2, an open and lightweight multimodal embedding model, has been launched by Google AI. This model extends beyond text to unify code, images, video, and audio in a single embedding space. Developed on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license, EmbeddingGemma 2 features 740 million parameters, making it ideal for on-device inference.
Its capabilities include locating specific video clips from voice memos or searching through hours of audio based on text queries, all processed by a single model.
EmbeddingGemma 2 retains the strong multilingual text performance of the original EmbeddingGemma model while achieving a remarkable 9.92-point improvement in code performance, rising from 68.76 to 78.68 in MTEB Code. This enhancement makes it particularly suitable for indexing local codebases, semantic code search, and coding agent retrieval. The model also sets new standards in quality-per-parameter for sub-1B models across image, video, documents, and audio, outperforming some specialist models more than twice its size.
The new model brings powerful capabilities to edge hardware, ensuring data privacy, reducing pipeline latency, and enabling developers to create cross-modal search and retrieval systems that operate offline. When used in conjunction with generative models like Gemma 4, EmbeddingGemma 2 facilitates on-device RAG pipelines capable of understanding complex multimodal data.
Its compatibility with Gemma 4 extends to sharing the same text tokenizer and audio encoder, allowing developers to run both models in a unified pipeline with a lower combined memory footprint.
Google AI has collaborated with various partners to ensure EmbeddingGemma 2 is readily available for developers. Detailed resources, including a developer guide, documentation, and guides for inference and fine-tuning, are available to assist in implementing on-device search and RAG systems with LiteRT.
Written by urgent.news from Google Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.