Google releases EmbeddingGemma 2, a 740-million-parameter model running on phones
An open-weight embedding model small enough to run on a Pixel phone is now in developers’ hands. Google DeepMind released EmbeddingGemma 2 on October 6, 2026, under the Apache 2.0 license, which permits commercial use.
· Originally published by ontime+ · Last verified: 9 Oct 2026 (Khaled Aziz)

Key Points
- Google DeepMind released EmbeddingGemma 2, an open-weight multimodal embedding model, on October 6, 2026.
- It maps text, code, images, video and audio into one shared 768-dimensional vector space.
- Apache 2.0 licensing and on-device operation let developers build search without cloud servers.
The latest:
An open-weight embedding model small enough to run on a Pixel phone is now in developers’ hands. Google DeepMind released EmbeddingGemma 2 on October 6, 2026, under the Apache 2.0 license, which permits commercial use. The company describes it as a lightweight multimodal model that converts text, code, images, video and audio into comparable vectors. Weights are posted on Hugging Face and Kaggle.
Details:
- The architecture: According to Google DeepMind, the model carries 740 million parameters in total, built around a 270-million-parameter core backbone handling text and code. Vision and audio encoders are optional additions on top of that core, meaning developers can load only the modalities their application actually needs.
- The shared space: The model maps every input format into a single 768-dimensional vector space, per the company. That design allows an image, an audio clip and a line of text to be compared directly against one another, which is the technical basis for cross-format search and retrieval inside a single application.
- The memory figure: In its text-only configuration, the model uses roughly 191MB of active RAM on Pixel devices, Google DeepMind said. That footprint is the central claim behind the on-device pitch: embedding workloads that normally require a server call can run locally on a phone or laptop.
- The license: EmbeddingGemma 2 is distributed under Apache 2.0, one of the most permissive open-source licenses, which allows commercial deployment without a separate agreement. The release is described as open-weight, meaning the trained parameters themselves are downloadable rather than accessible only through an API.
- Language coverage: The model supports multilingual text across more than 100 languages, according to the company, alongside code as an input type. That coverage is aimed at retrieval and semantic search applications that must handle mixed-language corpora without a separate model per language.
- Distribution: Model weights are available on Hugging Face and Kaggle, the two main public repositories for open machine-learning models. Google DeepMind published the release through its official blog in a post titled “EmbeddingGemma 2: an open, lightweight multimodal embedding model” dated October 2026.
- What was not stated: The announcement did not include benchmark comparisons against competing embedding models, nor did it specify latency figures for the vision and audio encoders on mobile hardware. No pricing or hosted-service tier was announced alongside the open weights.
Background:
Embedding models turn content into numerical vectors so machines can measure similarity. They are the retrieval layer behind semantic search and the systems that feed external documents to chatbots, a job normally handled by paid cloud APIs.
Between the lines:
The combination that matters here is the 191MB memory figure and the Apache 2.0 license. Together they remove both the technical and the legal barriers to running retrieval entirely on a device, which means a developer can ship a search feature without per-query API costs and without user data leaving the handset. The optional-encoder design points the same way: load only what the device can carry.
What’s next
Watch for independent benchmark results against established embedding models, developer reports on vision and audio encoder performance on mobile hardware, and whether Google pairs the open weights with a hosted commercial tier.