Skip to content
Breaking:

Google Releases EmbeddingGemma 2 for On-Device Multimodal Search

The 740-million-parameter open model maps text, code, images, audio, and video locally under an Apache 2.0 license.

By The Company Wire3 min read
Share
Google — Google Releases EmbeddingGemma 2 for On-Device Multimodal Search
Google — Google Releases EmbeddingGemma 2 for On-Device Multimodal Search. Photo: The Next Web.

Google has launched EmbeddingGemma 2, an open 740-million-parameter multimodal embedding model engineered to index and retrieve text, code, images, audio, and video directly on local hardware without sending data to external cloud servers, as reported by The Next Web (https://thenextweb.com/news/embeddinggemma-2-on-device-europe).

Released under the Apache 2.0 license, the model projects all supported media types into a shared 768-dimensional vector space. To manage memory constraints on edge devices, the architecture is modular: 270 million parameters handle text operations, while an optional 170-million-parameter vision encoder and a 300-million-parameter audio encoder can be loaded into memory only when needed.

On a Pixel 11 Pro, text-only workloads require approximately 191MB of RAM, while running all five modalities consumes around 567MB. By processing embeddings entirely on-device, systems can perform semantic search and retrieval workflows without routing user data across external networks.

EmbeddingGemma 2 builds on Google's Gemma 4 family of open-weight models released in April. The new model shares a text tokenizer and audio encoder with Gemma 4, allowing both architectures to run concurrently on devices such as smartphones and Raspberry Pi boards with reduced combined memory overhead.

The updated model supports a context window of 8,192 tokens across every modality, representing a fourfold increase over the original EmbeddingGemma. According to Google, this context length translates to roughly 29 static images, 58 video frames, or five and a half minutes of audio. Developers can also truncate embeddings from 768 dimensions down to 128 dimensions, which Google claims reduces storage requirements up to sixfold.

Google noted several constraints in its technical release documentation. EmbeddingGemma 2 includes no post-training safety tuning or output moderation layers, relying instead on filtering applied to its training data. While the model supports more than 100 languages, Google stated that performance varies across different languages. Google also reported that the wider Gemma family has surpassed one billion total downloads, with the first-generation EmbeddingGemma model representing 20 million of that total.

Sources

  1. The Next Web

Company: Google

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.