For AI agents: the public content index is available at https://we0.ai/llms.txt, and the English article bundle is available at https://we0.ai/llms-full.txt.
For AI agents: the complete content index is available at https://we0.ai/llms.txt, the full English article bundle is available at https://we0.ai/llms-full.txt, and this page is available as Markdown at https://we0.ai/articles/embeddinggemma-2-740m-multimodal-embeddin-2078d39c.md.
Google has released **EmbeddingGemma 2**, a compact multimodal embedding model designed to bring semantic search directly onto phones, lapto...

Google has released EmbeddingGemma 2, a compact multimodal embedding model designed to bring semantic search directly onto phones, laptops, browsers, and other edge devices.
The model was announced on October 6, 2026. Its headline specifications are unusually small for what it can do:
740M total parameters
768-dimensional unified embedding space
8K context window
100+ languages
~191MB active RAM for text-only weights on Pixel 11 Pro
~567MB active RAM for the full multimodal model on Pixel 11 Pro
EmbeddingGemma 2 can encode text, code, images, video, audio, and mixed combinations of those modalities into one shared vector space.
That means a natural-language query can retrieve a photo, an audio clip, or a moment inside a video without first requiring each file to be converted into text.
Google CEO Sundar Pichai described it as Google's first open, natively multimodal embedding model.

The practical difference is simple: search that previously depended on cloud APIs or separate image, audio, and text models can now run locally with a single compact model.
Embedding models are often invisible to end users, but they sit behind many modern search and RAG systems.
Their job is to turn content into numerical vectors:
content
→ embedding vector
→ position in a semantic space
Items with similar meaning are placed closer together. A search system then converts the user's query into another vector and retrieves the closest matches.
The first EmbeddingGemma focused on text. EmbeddingGemma 2 expands the idea into a unified multimodal space.
Google's model card says the new model maps text, images, video, and audio into the same 768-dimensional embedding space.
So a photo of a cat, the word "cat," and an audio recording of a cat can all be semantically related even though their raw formats are completely different.
This enables searches such as:
text query → matching images
text query → matching audio
text query → matching video moments
audio query → matching video
image query → related images or media
A single piece of content can also contain multiple modalities.
Google gives an example of a product page for trail shoes that contains text description, product photos, and a video showing grip on wet rocks. EmbeddingGemma 2 can represent the combined content with one embedding and match it against a query such as “waterproof shoes suitable for trail running.”
Shortly after launch, Hugging Face engineer Victor M demonstrated the model running directly in a browser.
According to the source article, he typed:
birds singing
and the result grid re-ranked immediately. The top results included both bird photos and audio clips showing bird-call waveforms.
The reported query took roughly 22 milliseconds. No server-side model call was involved and no external API was required.

That figure is a developer-reported browser test, not an official Google latency guarantee.
The broader point is nevertheless supported by Google's own deployment guidance: EmbeddingGemma 2 is designed for local execution through environments such as LiteRT, MediaPipe, WebGPU, transformers.js, MLX, llama.cpp, Ollama, and other edge-friendly runtimes.
EmbeddingGemma 2 uses an 8,192-token context window, four times the size of the previous EmbeddingGemma.
Google says a single-modality input can contain approximately:
5.5 minutes of audio
29 images
58 video frames
or interleaved mixtures of modalities.
The model also supports more than 100 languages.
That gives local search systems much more flexibility. Instead of indexing only short text snippets, an application can represent richer pieces of content containing longer passages, images, video frames, or audio segments.
Google compared EmbeddingGemma 2 with several similarly sized embedding systems.
The source article emphasizes the strongest differences in visual retrieval.
Google's model card reports:
| Benchmark | EmbeddingGemma 2 |
|---|---|
| MTEB Multilingual v2 | 61.36 |
| MTEB Code v1 | 78.68 |
| MIEB Lite | 64.64 |
| MMEB v2 Image | 57.28 |
| MMEB v2 Visual Document | 67.84 |
| MMEB v2 Video | 50.67 |
| MSEB Retrieval | 69.54 |

The source highlights a comparison against Jina v5 Omni-Nano:
MMEB v2 Image
EmbeddingGemma 2: 57.3
Jina v5 Omni-Nano: 31.6
MMEB v2 Video
EmbeddingGemma 2: 50.7
Jina v5 Omni-Nano: 31.2
These figures come from Google's published evaluation table.
As always, benchmark results should be interpreted within the specific evaluation setup rather than as a universal ranking for every production search workload.
The 740M total parameter count is split into modules.
Google's model card lists:
Text:
270M parameters total
130M backbone
140M embedder
Vision encoder:
170M parameters
Audio encoder:
300M parameters
Developers do not have to load all three components.
| Active Modalities | Effective Parameter Size |
|---|---|
| Text only | 270M |
| Text + image | 440M |
| Text + audio | 570M |
| Full multimodal | 740M |
This modularity matters on edge devices. If an app only searches text and code, there is no reason to keep the audio or vision encoder in memory.
Google reports that, after quantization on a Pixel 11 Pro, the model can run with approximately:
Text-only active RAM:
~191MB
Full multimodal active RAM:
~567MB
That is the basis for the source article's “under 600MB” headline.
The model uses quantization-aware training and supports INT4 and INT8 deployment options.
This is significant because multimodal search pipelines have traditionally required several separate components:
image encoder
+
speech recognition or audio encoder
+
text embedding model
+
extra preprocessing
EmbeddingGemma 2 collapses much of that into one unified architecture.
The native output dimension is 768.
EmbeddingGemma 2 also supports Matryoshka Representation Learning, allowing vectors to be truncated to:
512 dimensions
256 dimensions
128 dimensions
Google says this can reduce local vector-database storage by up to 6x relative to the full 768-dimensional representation.
Google AI Edge's deployment blog describes certain local index/storage configurations as achieving reductions of up to 8x, depending on how the representation and storage pipeline are configured.
The model card adds an important implementation detail: truncated vectors should be re-normalized before cosine-similarity search.
If query and document vectors use different dimensions, they also cannot be compared directly.
Google has published several reference applications around the model.
Google AI Edge Gallery's Instant Media Search lets users search local photos and videos using natural language or a sample image.
The app:
No internet connection is required for the embedding computation.
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
Video Moments Finder lets users search for specific moments inside local video files.
Examples from Google include queries such as:
kids laughing
dog catching a frisbee
person blowing out birthday candles
The app indexes visual and audio chunks and returns matching timestamps.
Google also launched AI Edge Foresight for Mac.
Foresight uses EmbeddingGemma 2 together with Gemma 4 for local retrieval and contextual reasoning.
Google says it can index meeting transcripts, search private files, retrieve images/documents/notes, work offline, and keep sensitive source material on the device.
This is close to the local RAG architecture described in the original article:
EmbeddingGemma 2
→ retrieve relevant local context
Gemma 4
→ reason over the retrieved context
The source article also collects several early developer experiments.
These are useful examples, but they should be treated as community-reported results rather than official Google benchmarks.
The Mac app Nativ reported tests using an 8-bit quantized model on an Apple M5 Max.
Its published figures included:
cosine similarity vs FP32:
0.9997
text embedding throughput at batch 32:
817 items/second

Those figures depend on the specific implementation, quantization, hardware, and batch size.
A Turkish developer cited in the source built a small test over personal notes.
The notes were mostly written in English, while the queries were in Turkish.
In his reported test:
keyword matching hit rate:
15%
EmbeddingGemma 2 hit rate:
97%
The setup checked whether the correct note appeared among the top 18 candidates.
He also tested typo-heavy real messages and reported that the semantic model found roughly twice as many relevant notes as keyword search. Each query reportedly took around 50 milliseconds locally.
This is not a standardized benchmark, but it illustrates one of the main reasons embedding search is useful: semantic similarity can work even when the query and stored text use different words or languages.
Developer Nick Lo also demonstrated a quantized EmbeddingGemma 2 build running through llama.cpp on a Nano development board.
He reported roughly:
~15ms
to convert a text query into an embedding for a small local image-search system.
After the matching image was found, an ESP32-S3 drew the image line by line.

Again, this is a developer experiment rather than an official reference latency.
The model is not only about media.
Code retrieval improved significantly over the previous EmbeddingGemma.
Google reports:
MTEB Code v1
EmbeddingGemma:
68.76
EmbeddingGemma 2:
78.68
That is a 9.92-point increase.
Google Gemma specifically highlights local codebase indexing, semantic code search, and coding-agent retrieval as target use cases.

Tools such as coding agents need to find the relevant portion of a repository before they can edit it.
A compact local embedder gives developers another architecture:
source repository
→ build embeddings locally
→ store index locally
→ natural-language query
→ retrieve relevant code
→ send only selected context to a larger coding model
The indexing and first-stage retrieval can remain on the developer's own machine.
Google had already launched Gemini Embedding 2 as a cloud API earlier in 2026.
EmbeddingGemma 2 takes a different approach.
Instead of requiring a hosted API for multimodal embedding generation, it is an open-weight model that can run locally.
| Approach | Main Advantage |
|---|---|
| Cloud embedding API | Managed infrastructure and easy scaling |
| On-device EmbeddingGemma 2 | Privacy, offline operation, low local latency |
For personal media, private files, enterprise notes, and local code repositories, the second option can be attractive because the raw source material does not necessarily have to leave the device.
Google explicitly promotes this as a privacy-first design.
That does not mean every application becomes private automatically. The rest of the application still has to be designed correctly, including secure local storage, access control, and careful handling of retrieved content.
EmbeddingGemma 2 is Google's open-weight multimodal embedding model built on Gemma 4 technology. It maps text, code, images, video, audio, and mixed inputs into a shared 768-dimensional vector space for search, RAG, similarity, classification, and clustering.
The full model has 740M parameters. The text component uses 270M parameters, while optional vision and audio encoders add 170M and 300M parameters respectively.
Yes. Google designed it for on-device use and provides deployment paths for phones, laptops, browsers, and edge hardware. Google's own AI Edge demos perform local search without requiring cloud embedding calls.
Google reports about 191MB active RAM for quantized text-only weights and about 567MB for the complete multimodal model on a Pixel 11 Pro. Actual memory use varies with runtime, hardware, precision, and enabled encoders.
Yes. All supported modalities are mapped into the same vector space, so text can retrieve images, audio, or video segments, and media can also be compared against other media.
Google releases the model weights under the commercially permissive Apache 2.0 license and describes it as an open model. Users must still follow the Gemma Prohibited Use Policy and applicable terms.
Yes. Google reports an MTEB Code score of 78.68, up from 68.76 for the previous EmbeddingGemma, and explicitly lists local repository indexing and coding-agent retrieval among the intended use cases.
EmbeddingGemma 2 is an embedding model: it converts content into vectors for retrieval and similarity. Gemma 4 is a generative model that can reason over retrieved information, so Google shows the two being combined for fully local RAG workflows.
EmbeddingGemma 2 takes multimodal semantic search out of the cloud and puts it within reach of ordinary consumer hardware. A single 740M model can embed text, code, images, video, and audio into one vector space while using about 567MB of active RAM in Google's full multimodal Pixel 11 Pro configuration.
Its strongest practical advantage is architectural simplicity: one compact model can support local media search, video-moment retrieval, private RAG, cross-language note search, and repository indexing without requiring every raw file to be uploaded to a server.
Google's official demos, model card, and deployment stack support the on-device story. The browser, Nativ, Turkish-note, and Nano-board figures in the source are useful early developer experiments, but they should be read as implementation-specific results rather than guaranteed performance.
EmbeddingGemma 2 makes local multimodal retrieval realistic enough that photos, recordings, videos, documents, and code can increasingly be searched where they already live—on the user's own device.
Start from one sentence and have a complete website in minutes.