Loading
EmbeddingGemma-2-Local-Multimodal-RAG.webp
Applied AIBy SDX Development

Share this article

Important information can be hidden in a PDF, a screenshot, a code file, or a meeting recording. What if you could find that content with a simple question, without sending the documents to a cloud service?

On October 6, 2026, Google DeepMind introduced EmbeddingGemma 2, an open embeddings model under the Apache 2.0 license, designed to search across multiple formats using a shared vector representation. Its appeal goes beyond its size: it makes a multimodal semantic search engine running locally more realistic.

This announcement highlights a specific stage of the RAG pipeline: indexing and semantically retrieving knowledge, before any answer is generated. Which content formats can be searched with a single model? What information can remain on the device? And when does on-device retrieval offer measurable advantages over a remote API?

Key takeaway

Multimodal search on your device, not a new chatbot

EmbeddingGemma 2 turns text, code, images, video, and audio into comparable 768-dimensional vectors. Once installed, it can index and retrieve content offline. However, it does not write answers: to generate text, you need to pair it with an LLM, either local or remote. Privacy depends on the entire pipeline.

In this article

Why move semantic search onto the device?

In a RAG (Retrieval-Augmented Generation) system, an application splits source materials into passages, creates their embeddings, retrieves items similar to the question, and may provide them to a generative model. This search can take place on remote infrastructure, on a private server, or directly on the device that holds the files.

When data is public and shared by thousands of users, a centralized embeddings API is often a pragmatic solution. The choice changes when dealing with a confidential contract, an internal software repository, meeting notes, or business screenshots: these documents should not move unnecessarily between multiple providers.

  • Fewer transfers: encoding, indexing, and search can stay on the device.
  • Offline use: once the models have been downloaded and the index prepared, search can work without a network connection.
  • A different operating cost: fewer billable queries, but more CPU/GPU, storage, and maintenance on the device.

Cloud and local should not be treated as mutually exclusive. The best place to run embeddings depends on the sensitivity and volume of the source material, as well as the devices actually in use.

EmbeddingGemma 2: a shared space for five types of content

Based on Gemma 4, the model projects text, code, images, video frames, and audio into a 768-dimensional space. A text query can therefore be compared with an image or a sound clip: it does not necessarily need to go through transcription or a generated caption first. The different encoders produce vectors that are compatible with one another.

The architecture is modular. You do not need to load the audio components to build a search engine for a Git repository.

Configuration Parameters Suitable use
Text and code 270 M Documentation, FAQs, source files
Text and vision 440 M Screenshots, diagrams, photos, and video frames
Text and audio 570 M Recordings, sounds, and speech-based search
Full multimodal 740 M Cross-media search across different formats

The stated context window is 8,192 tokens, shared across modalities. For a long document or video, it is still essential to select coherent pages, passages, or segments. A single embedding for an entire corpus is not a substitute for indexing its contents.

Worth distinguishing

EmbeddingGemma retrieves; an LLM writes

The embeddings model turns a query and content into numerical representations. It can rank results by semantic similarity. A generative model is only involved if the application needs to summarize, explain, or answer in sentences. A local search paired with generation through a remote API is not a fully local RAG.

API, local text RAG, or local multimodal RAG: three approaches

When choosing an architecture, it is more useful to compare three approaches than to look for “the best model” in the abstract.

01

Embeddings via an API

The application sends content to a provider for vectorization. It is quick to set up and offloads computation. In return, you need to manage data flows, costs, and network dependency.

02

Local text RAG

The text model and index reside on the device. This is often the first choice for notes, procedures, or a codebase. The work focuses on extraction, chunking, and retrieval quality.

03

Local multimodal RAG

Add the encoders needed to retrieve diagrams, screenshots, audio, or video sequences. This is useful when text alone is no longer enough, but media indexing and memory consumption need to be evaluated on the target hardware.

The three approaches can share building blocks: segmentation, business filters, metadata, and source display. Moving to multimodal does not require rewriting the entire application if the index has been designed properly.

The pipeline for a reliable multimodal RAG

An embeddings model does not replace document architecture. To retrieve the right information and show where it came from, you need to retain references to the files and manage permissions from the outset.

From file to verifiable answer: six steps

  1. Select the sources. Define the permitted folders, file types, updates, and scope for each user.
  2. Segment the content. Split text, extract relevant pages, select images, and create timestamped audio or video windows.
  3. Calculate embeddings. Encode each segment with the appropriate modality and retain the model version and dimension used.
  4. Build an index. Associate each vector with the source file, page, passage, timestamp, and access rights.
  5. Retrieve results. Encode the question, compare similarities, and filter results by permissions and relevance.
  6. Display or generate. Show the sources directly, or provide the selected passages to an LLM for an explicitly cited answer.

The last point matters: a high similarity score does not prove that the content is accurate. A professional interface should let users open the original document and, where appropriate, indicate that no sufficiently reliable answer was found.

Tutorial: test local search with Python

You can start without a vector database or an LLM. This prototype takes two short passages, calculates their embeddings with the text configuration, and displays the one closest to a question. It illustrates content retrieval, not a production-ready RAG application.

1. Install the dependencies

Google documents using sentence-transformers from version 6.1.0 onward. The first run downloads the model weights; subsequent local offline execution is possible, provided the necessary files are cached.

terminal.shBash
python -m venv .venv
source .venv/bin/activate  # Linux or macOS
# On Windows: .venv\Scripts\activate
pip install -U "sentence-transformers[image,audio,video]" transformers

2. Encode the documents and the question

rag-text.pyPython
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "google/embeddinggemma-2",
    config_kwargs={"vision_config": None, "audio_config": None},
)

documents = [
    "The application enforces document access permissions.",
    "The meeting notes describe a local retrieval architecture.",
]
question = "Where can I find the decisions about offline search?"

vectors = model.encode(
    documents,
    prompt_name="Document",
    normalize_embeddings=True,
)
query_vector = model.encode(
    question,
    prompt_name="SearchQuery",
    normalize_embeddings=True,
)

scores = model.similarity(query_vector, vectors)[0]
best_index = int(scores.argmax())

print(documents[best_index])
print("Score:", float(scores[best_index]))

The Document and SearchQuery prompts are the ones in Google's guide. The result is a ranked passage, not a written answer. A production-ready system also needs stable file IDs, permitted local paths or URLs, and metadata that can identify and cite the original excerpt.

3. Prepare to scale

  • Split long documents into coherent passages without mixing unrelated topics.
  • Reindex only changed items and retain a stable ID for each segment.
  • Add a persistent local index (for example, SQLite for metadata and vector search suited to the volume).
  • Evaluate relevance using real questions before connecting a generative model.

Retrieve an image, sound, or moment in a video

For multimodal search, load the required encoders and pass the media to the model as an object that specifies its modality. The following example corresponds to the API presented in the developer guide:

rag-multimodal.pyPython
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

image_vector = model.encode({"image": "capture-interface.png"})
audio_vector = model.encode({"audio": "extrait-reunion.wav"})
query_vector = model.encode(
    "discussion about the application's architecture",
    prompt_name="SearchQuery",
)

print("Image:", model.similarity(query_vector, image_vector))
print("Audio:", model.similarity(query_vector, audio_vector))

The files must be available locally and compatible with the runtime. For video, the guide specifies MP4 support with frame sampling at one frame per second by default. Audio should be prepared in mono at 16 kHz.

Visual documents

Screenshots and PDFs

A screenshot can be indexed using vision. A technical PDF often needs a hybrid approach: extractable text, diagrams, and possibly OCR for scanned pages.

Time-based media

Meetings and videos

To find a precise moment, index segments with their timestamps. A single representation for an entire video does not guarantee usable temporal localization.

Transcripts and captions are still useful. They can help with accessibility, auditing, exact-word filtering, and citing what someone actually said.

How much memory and storage should you plan for?

Google reports around 191 MB of active RAM for text only and 567 MB for full multimodal use on a Pixel 11 Pro. These figures come from optimized configurations on a specific device: they represent neither a universal memory peak nor the total cost of a Python or web pipeline. You also need to account for the runtime, media, activations, index, and possibly the LLM.

EmbeddingGemma 2 also uses Matryoshka Representation Learning (MRL), which allows 768-dimensional vectors to be truncated to 512, 256, or 128 dimensions. This choice mainly affects the space occupied by embeddings; its impact on recall needs to be measured using the project's data.

Reference

768 dimensions

One million float32 vectors represent about 3.07 GB of raw data, excluding metadata and index structure.

More compact index

256 dimensions

The same million float32 vectors take up about 1.02 GB: three times less, with a quality trade-off to verify.

Google reports that at 256 dimensions, its evaluations retain around 95% of reference quality for image, video, and speech search. This is not a guaranteed result for every corpus. The call model.encode(..., truncate_dim=256, normalize_embeddings=True) lets you test this configuration; queries and documents must use the same dimension.

Desktop, mobile, browser: what is available and what is coming

The model weights and integrations are available in several environments, including Transformers, sentence-transformers, MLX, Ollama, and LiteRT. Google also showcases two demos already announced in AI Edge Gallery: Instant Media Search and Video Moments Finder. The first uses a local SQLite index and ranking by cosine similarity, among other things.

For Mac, Google also presents AI Edge Foresight, an experimental meeting companion that combines notes, conversations, and private files with local processing. For a business project, these examples are architectural references, not a promise that browser integration is ready without adaptation.

Availability

Don't confuse an announcement with immediate compatibility

Google says EmbeddingGemma 2 will arrive in ML Kit on Android in the weeks following October 6. The company also mentions cross-platform support via MediaPipe for iOS, macOS, Windows, Linux, and the web. As of October 7, 2026, check the released versions and CPU, GPU, or web constraints before promising identical execution everywhere.

For a web integration, the design needs to account for the initial download size, weight caching, memory limits in a tab, local index storage, and supported browsers. A prototype that works on a high-end computer is not automatically suitable for every smartphone.

“No cloud” does not automatically mean “secure”

Moving embeddings onto the device eliminates some transfers, but it is not enough to guarantee privacy. The vectors themselves can reveal information about their sources, and a remote generation component can reintroduce network traffic.

Five checks before a private deployment

  1. Network. Check downloads, telemetry, synchronization, logs, and any calls to generation APIs.
  2. Storage. Protect documents, embeddings, and metadata; define encryption, retention periods, and deletion.
  3. Permissions. Restrict access to each index and filter results according to user permissions.
  4. Untrusted content. Prevent an instruction hidden in a PDF from becoming a command executed by the assistant.
  5. Traceability. Display the source, page, or audio timestamp, and make it possible to say when the information is missing.

Local execution is an architectural property that must be verified: it applies to ingestion, embeddings, the index, search, and any generation—not just the model's name.

How to evaluate EmbeddingGemma 2: a reproducible approach

To compare these architectures, start with an authorized, representative test corpus combining technical documentation, source code, screenshots, and audio excerpts. The goal is not to prove that local execution is always better, but to identify which dataset sizes, devices, and privacy requirements make it a practical choice. The following protocol is a proposed evaluation method, not a report of tests already performed.

01

Build the corpus

Prepare representative files, metadata, and 30 to 50 questions with known expected answers.

02

Compare architectures

Evaluate a remote API, local text search, and multimodal search using the same data and hardware.

03

Measure instead of assuming

Track recall@k, source quality, latency, indexing time, memory, and index size at 768 and then 256 dimensions.

04

Add the LLM afterward

Add local generation only if the user experience needs a written answer, and test refusals when no sources are available.

One possible evaluation scenario is a sample documentation corpus containing code, diagrams, and meeting excerpts. The retrieval system should point to the correct file, passage, or timestamp. The resulting measurements can help determine whether a local multimodal index delivers meaningful benefits over other architectures. This is a suggested test scenario, not a description of a deployed product.

Frequently asked questions

Find answers to common questions about local execution, models, and compatible environments.

Does EmbeddingGemma 2 really work without an internet connection?

Yes, for encoding and search, once the required weights and dependencies have been installed. This does not automatically make the rest of the application offline.

Can it write an answer like ChatGPT?

No. It ranks content by semantic similarity. Writing an answer requires an additional LLM, running locally if you want generation to stay on the device.

Do you always need to load all 740 million parameters?

No. The encoders are modular. A text and code search can use the 270-million-parameter configuration.

Is the browser the best environment?

Not necessarily. It offers convenient distribution, but memory, compatibility, and caching constraints may make a desktop or mobile application a better fit for some use cases.

Official sources and methodology 3 references

Article written based on announcements published on October 6, 2026. Hardware figures come from Google's published measurements; no independent benchmarks were conducted for this article. The Python examples illustrate the documented APIs and should be adapted to the hardware, dependencies, and files being tested.

The key point: EmbeddingGemma 2 does not, by itself, solve questions of document quality, permissions, or security. However, it provides an interesting building block for text, image, and audio search running close to users, without requiring a remote API for every query. It is an option to evaluate, not a universal promise.

APPLIED AI

Turn an AI prototype into a production workflow.

Computer vision, YOLO and RAG become useful when they connect to your real data, teams and constraints.

  • Defined data and metrics
  • Deployment connected to your tools
  • Quality monitoring after launch