Google has released EmbeddingGemma 2, a 740-million-parameter model that maps text, code, images, audio and video into a shared embedding space. The aim is to let developers search across media types on local hardware, such as finding a video clip with a voice memo or searching audio recordings with a text query.
According to Google's launch announcement, the model uses the Gemma 4 architecture and carries an Apache 2.0 license. It extends the first EmbeddingGemma's text-focused role into multimodal retrieval. Google says the earlier model has passed 20 million downloads, though that adoption figure does not measure the new model's deployment performance.
Load the encoders a workload needs
EmbeddingGemma 2 separates its text, vision and audio components. Google lists 270 million parameters for text-only work, an optional 170-million-parameter vision encoder and a 300-million-parameter audio encoder. A text-search application therefore does not have to load every media component.
With quantization, Google reports about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro. Those are reported measurements for a specified device and configuration, rather than a promise that an entire retrieval application will fit into that amount of memory. Developers also have to budget for their stored information and the other models they run.
The model supports an 8K-token context window. Google describes that capacity as accommodating up to 5.5 minutes of audio, 29 images or 58 video frames, including interleaved combinations. These are input examples for the available context, not unlimited recording or video lengths.
Smaller vectors reduce storage requirements
EmbeddingGemma 2 uses Matryoshka Representation Learning to let developers shorten its 768-dimensional vectors to 512, 256 or 128 dimensions. Google says this can reduce vector storage and memory usage by up to six times. The choice gives a local application a way to trade vector size against its retrieval needs without switching to a different model.
That deployment concern also appears in NVIDIA's Nemotron 3 Embed collection, which supports selectable embedding dimensions in one of its checkpoints. The models differ in size and deployment targets, but both descriptions make vector storage an explicit part of retrieval design.
Google reports a 9.92-point improvement on MTEB Code, increasing the score from 68.76 to 78.68. It also says the model retains strong multilingual text performance and leads among sub-billion-parameter multimodal embedders on several evaluations. These are the company's reported results; Franklin has not reproduced the benchmark runs.
Retrieval and generation remain separate jobs
The model produces embeddings for finding relevant information. Google describes pairing it with Gemma 4 to build retrieval-augmented generation pipelines, with EmbeddingGemma 2 locating local files and Gemma 4 supplying contextual reasoning. Their shared text tokenizer and audio encoder can lower the combined memory footprint, according to Google.
The examples make the separation concrete. Google AI Edge Gallery includes Instant Media Search and Video Moments Finder, while the Foresight app pairs local retrieval with reasoning. The launch also points to a MediaPipe Decision Task API for classification and routing based on multimodal context.
Local embedding generation can keep that stage of a pipeline on the device and allow offline retrieval. An application's wider privacy behavior still depends on how it handles retrieved material and whether other components send data elsewhere.
Deployment choices include browsers and local runtimes
Google says weights are available through Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming later. It lists MediaPipe and LiteRT for device applications, transformers.js or WebGPU for browsers, and several local serving tools including MLX, llama.cpp and Ollama.
A useful first trial would use the application's own mixed-media collection and representative queries. Developers can compare text-only and full multimodal configurations, then test shorter vectors against the matches their users need. The launch offers those configuration choices without making them a substitute for workload-specific evaluation.