Google releases EmbeddingGemma 2, a 740M multimodal model for phones
Text, code, images, video and audio now land in a single 768-dimensional search space in Google DeepMind’s EmbeddingGemma 2, released on Tuesday, 6 October. The open-weight model carries 740 million parameters, is built on the Gemma 4 architecture and ships under Apache 2.0.
· Originally published by ontime+ · Last verified: 6 Oct 2026 (Nicole Jeffrey)

Key Points
- Google DeepMind released EmbeddingGemma 2 on 6 October, mapping text, code, images, video and audio together.
- The 740-million-parameter model runs on phones and laptops under a commercially permissive Apache 2.0 licence.
- It extends Google's on-device search work beyond text, shifting safety responsibility onto application developers.
The latest:
Text, code, images, video and audio now land in a single 768-dimensional search space in Google DeepMind’s EmbeddingGemma 2, released on Tuesday, 6 October. The open-weight model carries 740 million parameters, is built on the Gemma 4 architecture and ships under Apache 2.0. Google said developers can use it to locate a video clip from a voice memo, or search hours of audio with a text query.
Details:
- The architecture: According to Google’s developer guide, the model is modular and does not have to be loaded whole. The text and code core holds 270 million parameters, a vision encoder adds 170 million and an audio encoder adds 300 million. Every combination outputs vectors in the same space.
- Why that matters: Because the spaces match, a query embedded with the text-only setup can be matched directly against documents embedded with the full multimodal model, according to Google’s developer guide. Embeddings convert content into numbers capturing meaning, the mechanism behind search tools and retrieval-augmented generation systems.
- Context window: All inputs share an 8,192-token context window, four times the original model’s. That covers up to 29 images, 58 video frames or 5.5 minutes of audio in a single input, well beyond the first text-only release.
- The benchmark: Investing.com reported the model scored 78.68 on the Massive Text Embedding Benchmark’s code test, a gain of 9.92 points over its predecessor’s 68.76. Google has not published a comparable multimodal benchmark figure alongside that code result.
- Compression trade-off: Matryoshka Representation Learning lets developers shorten vectors to as few as 128 dimensions, cutting storage up to six times. Google’s developer guide says 256-dimensional vectors retain roughly 95% of full quality on image, video and speech retrieval; at 128 dimensions, multimodal quality falls to around 75%.
- The safety gap: The model card states EmbeddingGemma 2 has not undergone post-training safety tuning, according to Unite.AI. That leaves safeguards to be built at the application level by whoever deploys it, rather than being carried in the released weights.
- Memory footprint: With quantisation on a Pixel 11 Pro, the text-only weights use about 191 megabytes of active memory and the full multimodal model about 567 megabytes, Google said. Those figures underpin the claim that the system runs on consumer phones.
- The demo apps: Seeking Alpha reported the Google AI Edge Gallery app will add two demos: Instant Media Search, for searching local photos and videos by description, and Video Moments Finder, which locates scenes such as a dog catching a frisbee without first transcribing audio.
- Desktop push: Google also introduced Google AI Edge Foresight for Mac, an experimental app that searches meeting transcripts, notes, images and documents stored locally, extending the on-device retrieval approach from phones to desktop workflows.
- Availability: The weights are on Hugging Face and Kaggle now. Google says the model reaches ML Kit for Android in the coming weeks, with hardware acceleration where supported, while availability in the Gemini Enterprise Agent Platform Model Garden is listed only as coming soon.
Background:
The first EmbeddingGemma, a 308-million-parameter text-only model built on Gemma 3, launched in September 2025. Google says it has been downloaded more than 20 million times, a base the company is now extending from text into images, video and audio on the same hardware class.
Between the lines:
The design choices point at a specific bet: shared embedding spaces across modality combinations, a six-fold storage cut via shortened vectors and a 567-megabyte full-model footprint all target developers who want retrieval to happen locally rather than in the cloud. The absence of post-training safety tuning sharpens the trade-off, moving responsibility for guardrails onto whoever ships the application.
What’s next
Watch for the ML Kit for Android rollout in the coming weeks and hardware-acceleration support on compatible devices, plus a confirmed date for Gemini Enterprise Agent Platform Model Garden availability and any published multimodal benchmark results.