EmbeddingGemma 2 had reached 298 Hacker News points when MrKeyoor's feed captured it early on October 7. The number worth carrying into a deployment meeting, though, is 567MB. That is how little active RAM Google says the full quantized model used on a Pixel 11 Pro while embedding text, images, video, and audio. A private media search engine that once called for several models and a server could now fit on the device holding the media, provided Google's measurement survives testing on other hardware.
Google released EmbeddingGemma 2 on October 6 with 740 million parameters and an Apache 2.0 license. It maps text, code, images, sampled video, and audio into one 768-dimensional vector space. The weights are available through Google's Hugging Face model page, which also documents the exact prefixes, data types, and truncation settings needed to avoid plausible-looking bad results.
One index can cover four kinds of media
An embedding model does not answer a question in prose. It turns an input into a list of numbers whose position can be compared with other lists. If a written query lands near a video frame or an audio clip in that space, an application can retrieve the media without first asking a larger generative model to describe everything. EmbeddingGemma 2 extends that process across four modalities, so the same index can match a voice note to a clip or a text query to a photograph.
The model can also encode an interleaved item as one vector. Google's model card gives the example of a product listing containing copy, two photographs, and a demonstration video. That combined vector can then be compared with a text-only search such as waterproof shoes for trail running. For app developers, this removes the reconciliation step between separate text, image, and audio indexes. It does not remove the need to choose how each input is sampled.
Google split the 740 million parameters into modules: 270 million for text, 170 million for vision, and 300 million for audio. A text-only build can skip both media encoders. Text plus images uses 440 million parameters, while text plus audio uses 570 million. The launch post reports about 191MB of active RAM for quantized text-only weights on a Pixel 11 Pro and about 567MB for the full model. Those are measurements on one Google phone, not a promise for every Android device, browser, or laptop runtime.
That modularity is the first half of the local-search story. An app that only indexes documents and screenshots does not have to reserve memory for audio. A field recorder can make the opposite choice. The full checkpoint remains available when a mixed library needs all four inputs to meet in the same space.
The smaller index may save more than the smaller model
EmbeddingGemma 2 produces 768-number vectors by default, but it was trained with Matryoshka Representation Learning. Developers can keep the first 512, 256, or 128 dimensions and discard the rest. Google's published results show that 256-dimensional vectors cut storage to one third while the overall MMEB v2 score moves from 59.01 to 56.24. The MTEB Code score falls from 78.68 to 76.18.
The 128-dimensional option reaches the advertised sixfold storage reduction, with a much steeper trade. MMEB v2 drops to 45.65 and MTEB Code to 71.41. Google's own guidance says 128 dimensions are best suited to text-only work and tells teams to validate multimodal retrieval before using that size. On a collection with millions of items, 256 dimensions may be the more interesting setting: the vectors occupy one third of the space while Google's tables show a smaller loss than the 128-dimensional version.
There is a quiet failure mode here. After slicing a vector, an application must normalize it again before calculating cosine similarity. The model card warns that skipping this step can lower ranking quality without throwing an error. Queries and indexed items must also use the same dimension. A 768-dimensional query cannot search a 256-dimensional collection, which makes the selected dimension part of the index format rather than a harmless runtime toggle.
Code retrieval improved more than text retrieval
Google reports a 78.68 score for EmbeddingGemma 2 on MTEB Code, up from 68.76 for the first EmbeddingGemma. That 9.92-point increase is much larger than the movement on multilingual MTEB, where the score changes from 61.15 to 61.36. The release therefore looks most useful to developers who need code search or new media types, while teams satisfied with the first model's text retrieval have less reason to rebuild an index immediately.
The model card says EmbeddingGemma 2 understands more than 100 languages, and Google says its training collection included content in more than 140. It also warns that performance can vary by language. The launch tables report aggregate multilingual scores, so they cannot answer whether a particular language, codebase, accent, or noisy recording will retrieve well. That answer still has to come from a labeled sample drawn from the intended application.
An 8K window is shared across every input
The 8,192-token context window sounds roomy until media starts consuming it. At Google's default settings, each image uses 280 tokens, each sampled video frame uses 140, and audio uses 25 tokens per second. With no accompanying text, that works out to about 29 images, 58 video frames, or 327 seconds of audio. An interleaved record draws every component from the same budget.
Video is sampled at one frame per second by default, according to the Hugging Face model card, and the rate is configurable. That distinction matters for a clip containing a short title card, a quick defect on a production line, or any event shorter than the sampling interval. Developers can vary the vision token budget between 70 and 1,120 soft tokens per image or frame, trading detail for input length and latency. Google's maximum of 58 frames describes the default budget, not 58 seconds of complete motion analysis.
Audio should arrive as mono at 16kHz. Text, audio, pictures, and video can be mixed in one sequence with placeholder tokens marking their positions. This gives product search and personal archives a useful way to preserve relationships between a caption and its media, but it also means preprocessing choices become part of retrieval quality. A missed frame never reaches the embedder.
The easy install still has sharp edges
The shortest text-search setup uses Sentence Transformers, but a production configuration should choose its encoders, output dimension, and numeric precision explicitly. A text-only index with 256-dimensional normalized vectors can start like this:
import torch
from sentence_transformers import SentenceTransformer
dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
model_kwargs={"torch_dtype": dtype},
)
query = model.encode(
"find the authentication retry logic",
prompt_name="CodeRetrieval",
truncate_dim=256,
normalize_embeddings=True,
)
The data type in that snippet is more than an optimization. Google says the model's activation range exceeds float16; using it can produce NaNs or degraded embeddings without an exception. The recommended choices are bfloat16 on supported hardware and float32 elsewhere. For text retrieval, task prefixes also affect precision. Queries and documents use different prefixes, and titled documents need the format title: {title} | text: {content} rather than the default untitled-document shortcut.
At publication time, the Hugging Face page said no hosted inference provider had deployed the model. Local use is already supported through Transformers and Sentence Transformers, while Google's launch post names LiteRT, MediaPipe, MLX, llama.cpp, Ollama, vLLM, SGLang, and browser routes through Transformers.js or WebGPU. Support across that list does not imply identical quantization, memory use, or media preprocessing, so the Pixel figure should be treated as a reference configuration.
Apache 2.0 does not settle the production review
Both Google's launch post and the model repository label the weights Apache 2.0. The model card also tells deployments to follow Google's Gemma Prohibited Use Policy. Teams that distribute the weights or build a service around them should read both documents instead of inferring all operating terms from the repository badge.
Google says the pretraining data had a January 2025 cutoff and included web documents, code, images, video, audio, and paired examples across modalities. The company describes filters for certain personal information and child sexual abuse material. It does not disclose a dataset inventory in the model card. The same document says the embedder received no post-training alignment, safety tuning, or output moderation because it produces representations rather than user-facing text.
That design moves the safety check downstream. Google assigns developers responsibility for retrieval filtering and fairness testing, and it warns that language coverage is uneven and that learned representations can carry biases from training data. A local index can keep private files off a remote API, but local processing alone does not make ranking fair, accurate, or safe for automated decisions.
Independent tests now need to check the two numbers carrying this launch: 567MB for the full model and the retrieval scores retained at 256 dimensions. Results on older phones, CPUs, browsers, and multilingual collections will show how far the reference setup travels. If those figures hold outside Google's Pixel test, developers can shrink both sides of local search at once: the model that reads the media and the index that remembers it. If they do not, the model card already identifies the knobs that will explain why.