Google expands EmbeddingGemma beyond text to images, audio and video
Google LLC today released EmbeddingGemma 2, an open multimodal embedding model small enough to run on a smartphone. The release takes the EmbeddingGemma line beyond text, which was all the first version handled when Google introduced it in September 2025. Images, audio and video now share one embedding space with text. An app built on […] The post Google expands EmbeddingGemma beyond text to…
Google has expanded its EmbeddingGemma model to include support for images, audio, and video, in addition to text. The new EmbeddingGemma 2 version, released today, can run on a smartphone and allows apps to find matching moments between voice memos and videos or between text and images without needing to transfer data off the device.
Google DeepMind engineers Sahil Dua and Henrique Schechter Vera said the initial response was "beyond our expectations," with the model being downloaded more than 20 million times. The new version is twice the size of the original, with 740 million parameters, and is built on the Gemma 4 architecture. The vision and audio encoders contribute to the larger size.
The 270 million-parameter text core requires about 191 megabytes of memory, but an app can reduce this by using a quantized build. By employing Matryoshka Representation Learning, developers can further reduce the memory footprint of embeddings by up to six times. EmbeddingGemma 2 scores 78.68 on the code section of the Massive Text Embedding Benchmark, outperforming the original version by nearly 10 points.
Google highlights the model's performance as leading among multimodal embedding models under 1 billion parameters and outperforming some specialist models more than twice its size on image, video, and audio tasks. The model shares a text tokenizer and an audio encoder with Gemma 4, allowing on-device retrieval-augmented generation setups to run both models with reduced memory usage.
EmbeddingGemma 2 is available for download from Hugging Face Inc. and Google's Kaggle under an Apache 2.0 license that permits commercial use, and will soon be part of Google's AI Edge Foresight meeting app and AI Edge Gallery demo app.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.