
Applied AI · Multimodal RAG
With EmbeddingGemma 2, Google brings semantic search and local execution closer together. A comparison of RAG architectures, a Python tutorial, and the real-world limits of privacy on your devices.


Important information can be hidden in a PDF, a screenshot, a code file, or a meeting recording. What if you could find that content with a simple question, without sending the documents to a cloud service?
On October 6, 2026, Google DeepMind introduced EmbeddingGemma 2, an open embeddings model under the Apache 2.0 license, designed to search across multiple formats using a shared vector representation. Its appeal goes beyond its size: it makes a multimodal semantic search engine running locally more realistic.
This announcement highlights a specific stage of the RAG pipeline: indexing and semantically retrieving knowledge, before any answer is generated. Which content formats can be searched with a single model? What information can remain on the device? And when does on-device retrieval offer measurable advantages over a remote API?
EmbeddingGemma 2 turns text, code, images, video, and audio into comparable 768-dimensional vectors. Once installed, it can index and retrieve content offline. However, it does not write answers: to generate text, you need to pair it with an LLM, either local or remote. Privacy depends on the entire pipeline.
In a RAG (Retrieval-Augmented Generation) system, an application splits source materials into passages, creates their embeddings, retrieves items similar to the question, and may provide them to a generative model. This search can take place on remote infrastructure, on a private server, or directly on the device that holds the files.
When data is public and shared by thousands of users, a centralized embeddings API is often a pragmatic solution. The choice changes when dealing with a confidential contract, an internal software repository, meeting notes, or business screenshots: these documents should not move unnecessarily between multiple providers.
Cloud and local should not be treated as mutually exclusive. The best place to run embeddings depends on the sensitivity and volume of the source material, as well as the devices actually in use.
Based on Gemma 4, the model projects text, code, images, video frames, and audio into a 768-dimensional space. A text query can therefore be compared with an image or a sound clip: it does not necessarily need to go through transcription or a generated caption first. The different encoders produce vectors that are compatible with one another.
The architecture is modular. You do not need to load the audio components to build a search engine for a Git repository.
| Configuration | Parameters | Suitable use |
|---|---|---|
| Text and code | 270 M | Documentation, FAQs, source files |
| Text and vision | 440 M | Screenshots, diagrams, photos, and video frames |
| Text and audio | 570 M | Recordings, sounds, and speech-based search |
| Full multimodal | 740 M | Cross-media search across different formats |
The stated context window is 8,192 tokens, shared across modalities. For a long document or video, it is still essential to select coherent pages, passages, or segments. A single embedding for an entire corpus is not a substitute for indexing its contents.
The embeddings model turns a query and content into numerical representations. It can rank results by semantic similarity. A generative model is only involved if the application needs to summarize, explain, or answer in sentences. A local search paired with generation through a remote API is not a fully local RAG.
When choosing an architecture, it is more useful to compare three approaches than to look for “the best model” in the abstract.
The application sends content to a provider for vectorization. It is quick to set up and offloads computation. In return, you need to manage data flows, costs, and network dependency.
The text model and index reside on the device. This is often the first choice for notes, procedures, or a codebase. The work focuses on extraction, chunking, and retrieval quality.
Add the encoders needed to retrieve diagrams, screenshots, audio, or video sequences. This is useful when text alone is no longer enough, but media indexing and memory consumption need to be evaluated on the target hardware.
The three approaches can share building blocks: segmentation, business filters, metadata, and source display. Moving to multimodal does not require rewriting the entire application if the index has been designed properly.
An embeddings model does not replace document architecture. To retrieve the right information and show where it came from, you need to retain references to the files and manage permissions from the outset.
The last point matters: a high similarity score does not prove that the content is accurate. A professional interface should let users open the original document and, where appropriate, indicate that no sufficiently reliable answer was found.
You can start without a vector database or an LLM. This prototype takes two short passages, calculates their embeddings with the text configuration, and displays the one closest to a question. It illustrates content retrieval, not a production-ready RAG application.
Google documents using sentence-transformers from version 6.1.0 onward. The first run downloads the model weights; subsequent local offline execution is possible, provided the necessary files are cached.
python -m venv .venv
source .venv/bin/activate # Linux or macOS
# On Windows: .venv\Scripts\activate
pip install -U "sentence-transformers[image,audio,video]" transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)
documents = [
"The application enforces document access permissions.",
"The meeting notes describe a local retrieval architecture.",
]
question = "Where can I find the decisions about offline search?"
vectors = model.encode(
documents,
prompt_name="Document",
normalize_embeddings=True,
)
query_vector = model.encode(
question,
prompt_name="SearchQuery",
normalize_embeddings=True,
)
scores = model.similarity(query_vector, vectors)[0]
best_index = int(scores.argmax())
print(documents[best_index])
print("Score:", float(scores[best_index]))
The Document and SearchQuery prompts are the ones in Google's guide. The result is a ranked passage, not a written answer. A production-ready system also needs stable file IDs, permitted local paths or URLs, and metadata that can identify and cite the original excerpt.
For multimodal search, load the required encoders and pass the media to the model as an object that specifies its modality. The following example corresponds to the API presented in the developer guide:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
image_vector = model.encode({"image": "capture-interface.png"})
audio_vector = model.encode({"audio": "extrait-reunion.wav"})
query_vector = model.encode(
"discussion about the application's architecture",
prompt_name="SearchQuery",
)
print("Image:", model.similarity(query_vector, image_vector))
print("Audio:", model.similarity(query_vector, audio_vector))
The files must be available locally and compatible with the runtime. For video, the guide specifies MP4 support with frame sampling at one frame per second by default. Audio should be prepared in mono at 16 kHz.
A screenshot can be indexed using vision. A technical PDF often needs a hybrid approach: extractable text, diagrams, and possibly OCR for scanned pages.
To find a precise moment, index segments with their timestamps. A single representation for an entire video does not guarantee usable temporal localization.
Transcripts and captions are still useful. They can help with accessibility, auditing, exact-word filtering, and citing what someone actually said.
Google reports around 191 MB of active RAM for text only and 567 MB for full multimodal use on a Pixel 11 Pro. These figures come from optimized configurations on a specific device: they represent neither a universal memory peak nor the total cost of a Python or web pipeline. You also need to account for the runtime, media, activations, index, and possibly the LLM.
EmbeddingGemma 2 also uses Matryoshka Representation Learning (MRL), which allows 768-dimensional vectors to be truncated to 512, 256, or 128 dimensions. This choice mainly affects the space occupied by embeddings; its impact on recall needs to be measured using the project's data.
One million float32 vectors represent about 3.07 GB of raw data, excluding metadata and index structure.
The same million float32 vectors take up about 1.02 GB: three times less, with a quality trade-off to verify.
Google reports that at 256 dimensions, its evaluations retain around 95% of reference quality for image, video, and speech search. This is not a guaranteed result for every corpus. The call model.encode(..., truncate_dim=256, normalize_embeddings=True) lets you test this configuration; queries and documents must use the same dimension.
The model weights and integrations are available in several environments, including Transformers, sentence-transformers, MLX, Ollama, and LiteRT. Google also showcases two demos already announced in AI Edge Gallery: Instant Media Search and Video Moments Finder. The first uses a local SQLite index and ranking by cosine similarity, among other things.
For Mac, Google also presents AI Edge Foresight, an experimental meeting companion that combines notes, conversations, and private files with local processing. For a business project, these examples are architectural references, not a promise that browser integration is ready without adaptation.
Google says EmbeddingGemma 2 will arrive in ML Kit on Android in the weeks following October 6. The company also mentions cross-platform support via MediaPipe for iOS, macOS, Windows, Linux, and the web. As of October 7, 2026, check the released versions and CPU, GPU, or web constraints before promising identical execution everywhere.
For a web integration, the design needs to account for the initial download size, weight caching, memory limits in a tab, local index storage, and supported browsers. A prototype that works on a high-end computer is not automatically suitable for every smartphone.
Moving embeddings onto the device eliminates some transfers, but it is not enough to guarantee privacy. The vectors themselves can reveal information about their sources, and a remote generation component can reintroduce network traffic.
Local execution is an architectural property that must be verified: it applies to ingestion, embeddings, the index, search, and any generation—not just the model's name.
To compare these architectures, start with an authorized, representative test corpus combining technical documentation, source code, screenshots, and audio excerpts. The goal is not to prove that local execution is always better, but to identify which dataset sizes, devices, and privacy requirements make it a practical choice. The following protocol is a proposed evaluation method, not a report of tests already performed.
Prepare representative files, metadata, and 30 to 50 questions with known expected answers.
Evaluate a remote API, local text search, and multimodal search using the same data and hardware.
Track recall@k, source quality, latency, indexing time, memory, and index size at 768 and then 256 dimensions.
Add local generation only if the user experience needs a written answer, and test refusals when no sources are available.
One possible evaluation scenario is a sample documentation corpus containing code, diagrams, and meeting excerpts. The retrieval system should point to the correct file, passage, or timestamp. The resulting measurements can help determine whether a local multimodal index delivers meaningful benefits over other architectures. This is a suggested test scenario, not a description of a deployed product.
Find answers to common questions about local execution, models, and compatible environments.
Yes, for encoding and search, once the required weights and dependencies have been installed. This does not automatically make the rest of the application offline.
No. It ranks content by semantic similarity. Writing an answer requires an additional LLM, running locally if you want generation to stay on the device.
No. The encoders are modular. A text and code search can use the 270-million-parameter configuration.
Not necessarily. It offers convenient distribution, but memory, compatibility, and caching constraints may make a desktop or mobile application a better fit for some use cases.
Article written based on announcements published on October 6, 2026. Hardware figures come from Google's published measurements; no independent benchmarks were conducted for this article. The Python examples illustrate the documented APIs and should be adapted to the hardware, dependencies, and files being tested.
The key point: EmbeddingGemma 2 does not, by itself, solve questions of document quality, permissions, or security. However, it provides an interesting building block for text, image, and audio search running close to users, without requiring a remote API for every query. It is an option to evaluate, not a universal promise.
APPLIED AI
Computer vision, YOLO and RAG become useful when they connect to your real data, teams and constraints.
Enter at least 2 characters to start searching.