Google CEO Sundar Pichai announces EmbeddingGemma 2 model; says it is company’s first open, natively multimodal embedding model
Google has launched EmbeddingGemma 2, an open multimodal embedding model for devices.
Google has launched EmbeddingGemma 2, an open multimodal embedding model for devices. The launch comes a year after the company introduced EmbeddingGemma for developers last year. For those unaware, an embedding model is a machine learning tool that turns complex data like words, sentences, images, or audio into a list of numbers called a vector. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.Announcing the new model, Google CEO Sundar Pichai said: “Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency.- our first open, natively multimodal embedding model- handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor- ideal for offline, privacy-first RAG when paired with Gemma 4- outperforms some specialist models more than twice its sizeWeights available now on Hugging Face.”EmbeddingGemma 2: Key capabilitiesEmbeddingGemma 2 is designed to provide a single, compact open model released under the Apache 2.0 license, delivering exceptional multimodal performance for its size. Based on Gemma 4, this sub-1B model maps text, code, images, video, and audio into a unified 768-dimensional space. Its modular architecture can scale from 270M parameters for text and code up to 740M parameters for all modalities. Key capabilities of the EmbeddingGemma include:Native multimodal retrieval: Unlocks search across text, code, images, video, and audio directly in a shared 768-dimensional vector space.Superior code understanding: Significantly outperforms EmbeddingGemma 1 on code search and technical retrieval, making it ideal for local codebase indexing and agentic code search.Modular memory footprint: Selectively load only the encoders you need at runtime: 270M (text/code), 440M (text + vision), 570M (text + audio), or 740M (full multimodal), all projecting into the same compatible vector space.Flexible vector storage: Matryoshka Representation Learning (MRL) enables dynamic truncation from 768 dimensions down to 128 dimensions, cutting vector database storage requirements while retaining much of the original quality. For instance, at 256 dimensions, most of the full quality of the original embedding on text and code is retained and about 95% on image, video, and speech retrievalHow it worksGoogle says that all inputs are processed through the shared backbone and their resulting embeddings occupy the same dimensional space:Text and code (270M base): An adapted Gemma 4 decoder with an 8,192-token context window.Vision (+170M): A vision encoder that processes images, visual documents (PDFs, slides, charts), and video frames.Audio (+300M): A dedicated speech and sound encoder that directly ingests raw audio.Modular loading: You only load what you need. Run text-only at 270M parameters, add vision for 440M parameters, or load the full multimodal model at 740M parameters.Shared tokenizer & audio encoder with Gemma 4: When paired with Gemma 4 in an on-device RAG pipeline, both models share the same text tokenizer and audio encoder architecture, reducing the total memory footprint.You use AI every day. Now get your AI Quotient. Take the AIQ test.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.