EmbeddingGemma 2 puts multimodal search on a phone
Google DeepMind's EmbeddingGemma 2 maps text, code, images, video and audio into one vector space with 740M parameters. Text-only use needs about 191 MB of RAM.
4 min read

By the numbers
- total parameters across three encoders
- 740M
- embedding dimensions, truncatable to 128
- 768
- RAM for text-only use on a Pixel 11 Pro
- ~191 MB
- MTEB Code score, up from 68.76
- 78.68
Google DeepMind released EmbeddingGemma 2 on October 6, 2026, an open model that turns text, code, images, video and audio into vectors in one shared space. It has 740 million parameters and ships under the Apache 2.0 license, according to Google's developer guide. The text-only version runs in about 191 MB of memory on a phone. That makes it practical to search a user's photos, voice notes and documents without sending them to a server.
What an embedding model does
An embedding model turns a piece of content into a list of numbers, called a vector. Similar content ends up with similar vectors. Search then becomes a matter of finding the nearest vectors to a query.
This is the engine behind RAG, or retrieval-augmented generation, where an app finds relevant documents before asking a language model to answer. EmbeddingGemma 2 puts every kind of media in the same space. A text query like "dog on a beach" can therefore find a photo, a video frame or an audio clip directly.
The edge guide describes it as a model that "maps text, images, video frames, and audio into a single, unified vector space."
How it is built
The model is made of stackable encoders that share one embedding space:
| Part | Parameters |
|---|---|
| Text and code encoder | 270M |
| Vision encoder | 170M |
| Audio encoder | 300M |
| Total | 740M |
Stackable means a text-only app can load just the 270M core, and add vision or audio only if it needs them.
The output is a 768-dimension vector. It can be cut down to 512, 256 or 128 dimensions using Matryoshka Representation Learning, a training method that packs the most important information into the front of each vector. Google says 128 dimensions cuts storage by 6x. "Most of the full quality of the original embedding on text and code is retained and about 95% on image, video, and speech retrieval," the developer guide says.
The context window is 8,192 tokens, four times the first version's. Unite.AI lists what fits in one call:
- 29 images, at 280 tokens each
- 58 video frames, at 140 tokens each
- 5.5 minutes of audio, at 25 tokens per second
Unite.AI adds that the model supports more than 100 languages and has a training data cutoff of January 2025.
Benchmarks and speed
| Test | EmbeddingGemma 2 | Source |
|---|---|---|
| MTEB Code | 78.68 (EmbeddingGemma 1: 68.76) | Unite.AI, Crypto Briefing |
| MTEB multilingual | 61.36 | Unite.AI |
| MMEB v2 image retrieval | 57.28 | Unite.AI |
| MMEB v2 video retrieval | 50.67 | Unite.AI |
| MSEB audio | 69.54 | Unite.AI |
The code score is the biggest jump, a gain of 9.92 points. Google's guide describes it as 14% higher than the first version.
On hardware, Google reports about 191 MB of RAM for text-only use and about 567 MB for full multimodal use on a Pixel 11 Pro, with quantization. The edge guide measured image embedding at 37.3 milliseconds on a MacBook M5 Pro GPU, or 26.9 images a second.
The same guide says the model can sort an input among 500 options in under 100 milliseconds, with no training needed. That puts it close to the new class of decision models such as Cloudflare's Clef.
Where to get it
The weights are on Hugging Face as google/embeddinggemma-2, and on Kaggle and the LiteRT Community. Crypto Briefing lists support in sentence-transformers, Hugging Face transformers, LiteRT and MediaPipe, MLX and Ollama. Google also ships pre-quantized LiteRT bundles as .litertlm files. The model also powers AI Edge Foresight, an experimental Mac app that takes meeting notes offline.
What this means for developers
- Move search onto the device. At about 191 MB for text, the model fits in a mobile app. Private content can be indexed locally, which Unite.AI notes means personal files "don't have to be uploaded to a server just to become searchable."
- Use truncation to cut storage. If your vector database is large, test 256 or 128 dimensions. Google reports a 6x saving at 128 with most text quality kept.
- Re-embed if you change models. Vectors from EmbeddingGemma 1 and 2 are not interchangeable. Plan a full re-index rather than mixing them.
- Try it for code search. A 9.92-point gain on MTEB Code makes it worth testing for repository search and code RAG.
- Load only the encoders you need. A text-only feature does not need the 470M parameters of vision and audio encoders. Smaller loads mean faster start-up on phones.
Run it against your current embedding model on your own queries first. The public benchmarks show the direction, and your own data shows the real gain.
Sources
- EmbeddingGemma 2: The Developer Guide - Google Developers Blog
- Bring multimodal semantic search to the edge with EmbeddingGemma 2 - Google Developers Blog
- DeepMind Debuts EmbeddingGemma 2, Mapping Five Modalities Into One Space - Unite.AI
- Google releases EmbeddingGemma 2, a 740 million parameter open-weight model under Apache 2.0 - Crypto Briefing
Related articles

Liquid AI LFM2.5-VL-DSpark speeds up its 3B vision model
Liquid AI's 280M-parameter LFM2.5-VL-DSpark draft model decodes its 3B vision model up to 3.13x faster on an M5 Max, with identical output.

Google AI Edge Foresight takes meeting notes offline
Google AI Edge Foresight, an experimental Mac app, turns shorthand into meeting notes on-device with Gemma 4 and the 740M-parameter EmbeddingGemma 2.

Strands Decider 2B makes agent choices in 115 ms
AWS's Strands Decider 2B is an open-source model that picks from a list instead of writing text. It answers in 115 ms on an RTX 3090 and scores 72.3% on JevBench.
The weekly digest
One email every Friday with the week's top stories from all six desks: AI, dev tools, coding, the tech industry, startups and what's next. Free, no spam.