Overview
EmbeddingGemma 2 is the multimodal embedding model Google DeepMind released in October 2026. Built on the Gemma 4 architecture with 740 million parameters and open weights under Apache 2.0, what sets it apart is not what it generates but what it puts in one place: text, code, images, video frames and audio all map into a single embedding space. That means searches like finding a video clip with a sentence, or a document with a picture, can be answered directly by vector distance instead of requiring a transcription or captioning pass first.
The 740M size is deliberate. It belongs to the Google AI Edge line, aimed at running inside browsers and on-device, so the retrieval layer does not have to make a round trip to a server every time. For local knowledge bases, offline search and on-device RAG, that size matters more than a leaderboard position.
Key Features
- Five modalities in one space: Text, code, images, video frames and audio are encoded into the same vector space, so cross-modal similarity can be compared directly and retrieval no longer needs to flatten everything into text first
- 740M parameters, on-device scale: Kept at 740 million parameters, small enough with quantization to run in a browser or on-device, so retrieval does not depend on a server round trip and offline scenarios stay viable
- Apache 2.0 open weights: Weights ship under Apache 2.0, so commercial use, modification and self-hosting are all permitted, with no quota or vendor lock-in
- Built on Gemma 4: Constructed on the Gemma 4 architecture, placing it in the same ecosystem as the generative models in the series and TranslateGemma, so tooling and deployment patterns carry over
- Part of the Google AI Edge ecosystem: Released as a member of the AI Edge line for on-device and browser scenarios, connecting with Google's existing on-device inference tooling
- Made for retrieval, not generation: It outputs vectors rather than text. The usage pattern is indexing and recall: encode candidates once and store them, then encode only the query and let the vector store do the rest
Use Cases
- Local knowledge bases and offline search that index documents, screenshots and meeting recordings together
- On-device RAG, letting apps on a phone or in a browser recall their own material without uploading it
- Cross-modal asset management, searching image libraries and video clips with a written description
- Code retrieval, putting natural language questions and a codebase into one index to find functions and snippets directly
Pros
- Text, images, video, audio and code share one space, so cross-modal retrieval needs no intermediate step
- 740M parameters make browser and on-device deployment a real option
- Apache 2.0 licensed, with no obstacle to commercial use or self-hosting
- Based on Gemma 4, sharing an ecosystem with the rest of the series and its toolchain
- It emits vectors, so after indexing once the query cost is very low
- Open weights mean you can fine-tune on your own corpus instead of settling for a generic model
Pricing
The weights are released under Apache 2.0, free to download and self-host with no per-call fee. Your cost is the inference you run: on-device work uses the user's hardware, and server-side deployment bills on whatever machines or inference service you choose.
Summary
EmbeddingGemma 2 solves a specific problem. Getting mixed media into a searchable index used to mean transcribing images and audio into text first and searching over that, a step that is slow and lossy. With five modalities mapped into one embedding space, that step disappears and the distance between a question and a video frame can be computed directly.
What makes it worth watching is size rather than scale. At 740M parameters it can live on device, which is decisive for two kinds of work: anything where data must not leave the device, such as local archives and personal material on a phone, and anything where a network round trip per query is unacceptable, since encoding locally is faster and steadier than calling an API.
Two things are worth settling before you build on it. The benefit of a shared space only materializes if the index is built sensibly: video has to be sampled by frame or segment, audio has to be chunked, and those engineering decisions affect results more than the choice of model. And Apache 2.0 gives you room to adapt, so if your corpus is highly specialized, fine-tuning on your own data usually beats swapping in a larger general model.
It is not a general detector and not a generative model. It does one thing, turning content into comparable vectors, and doing that well is what gives the retrieval layer above it room to work.
Version History
- EmbeddingGemma 2 release (2026-10-06): Google DeepMind released EmbeddingGemma 2, an open weights multimodal embedding model with 740M parameters built on the Gemma 4 architecture and licensed under Apache 2.0. It maps text, code, images, video frames and audio into a single embedding space so cross-modal retrieval works by direct vector distance without converting everything to text first, and it is sized for browser and on-device deployment as part of the Google AI Edge ecosystem