{
  "id": 12701554,
  "title": "Google DeepMind Launches EmbeddingGemma 2 for On-Device Multimodal Search",
  "url": "https://urgent.news/2026/10/07/google-deepmind-launches-embeddinggemma-2-for-on-device-multimodal",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T13:50:38.000Z",
  "source": {
    "name": "TechRepublic",
    "slug": "techrepublic",
    "url": "https://www.techrepublic.com/article/news-google-embeddinggemma-2-on-device-multimodal-search/"
  },
  "original_language": "en",
  "account": "Google DeepMind has introduced EmbeddingGemma 2, a 740 million-parameter open-weight model that enables on-device text, image, video, and audio search. This model aims to allow developers to build multimodal search and retrieval features directly on consumer devices, reducing reliance on cloud AI and lowering memory requirements. The first EmbeddingGemma has surpassed 20 million downloads and is used for local search and retrieval-augmented generation (RAG). The second-generation model supports text, vision, and audio components, with the latter two adding 170 million and 300 million parameters, respectively. Quantized EmbeddingGemma 2 uses only about 191MB of RAM for text-only processing on a Pixel 11 Pro, while the full multimodal model uses around 567MB. The model's memory efficiency is attributed to Matryoshka Representation Learning, which reduces the number of dimensions in its vectors. EmbeddingGemma 2 has an 8,192-token context window and outperforms its predecessor in benchmarks, including MTEB Code, which scored 78.68 compared to 68.76 for the first model. Google demonstrates potential applications, such as Video Moments Finder, which indexes local video and audio to surface relevant timestamps, and Foresight, which enhances productivity by enriching notes and retrieving information from local data without sending it to the cloud. The developer stack includes MediaPipe Tasks for multimodal preprocessing and retrieval, and LiteRT for deployment and acceleration across various hardware. The most significant change is that personal media can now be searched using natural language without prior uploading, reducing latency and improving privacy-sensitive applications. However, the model lacks safety tuning, and performance varies across languages. EmbeddingGemma 2 signifies a shift toward local multimodal retrieval on everyday hardware, offering developers another architecture choice for processing personal data.",
  "summary": "Google DeepMind’s EmbeddingGemma 2 brings text, image, video and audio search to consumer devices with lower memory and less reliance on cloud AI. The post Google DeepMind Launches EmbeddingGemma 2 for On-Device Multimodal Search appeared first on TechRepublic .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}