Skip to content

EmbeddingGemma 2 Maps Text, Images, Video and Audio On Device: 740M Apache 2.0 Multimodal Embedder With 8K Context

Google DeepMind launched EmbeddingGemma 2 on 6 Oct 2026: a 740M Apache 2.0 multimodal embedder for text, code, images, video and audio, with modular 270M text-only mode and vendor-reported MTEB Code 78.68.

EmbeddingGemma 2 multimodal embedding model banner from Google DeepMind

Google DeepMind launched EmbeddingGemma 2 on 6 Oct 2026: an open multimodal embedding model that maps text, code, images, video and audio into one 768-dimensional space.

The release is Apache 2.0 licensed, built on the Gemma 4 architecture, and sized for consumer hardware. Total parameters are 740 million when every encoder is loaded. Text-only work can run at 270 million.

That matters for local RAG, offline media search and coding agents that need retrieval without shipping private files to a cloud API. Weights are on Hugging Face and Kaggle. LiteRT builds are listed for on-device use.

What Google Claims EmbeddingGemma 2 Does

According to the Google Keyword blog and the model card, EmbeddingGemma 2 is natively multimodal. One model embeds mixed inputs and keeps them comparable in the same vector space.

Google says the first EmbeddingGemma release passed 20 million downloads for text embeddings. Version 2 adds vision and audio encoders on top of a stronger text and code backbone.

The same blog post says the model is built from the same technology line as Gemini Embedding models, while staying open weight and small enough for phones and laptops.

Confirmed Specs From The Model Card

These figures come from Google’s EmbeddingGemma 2 model card (updated 6 Oct 2026) and the developer guide. They are vendor-reported unless noted otherwise.

Item Confirmed Notes
License Apache 2.0 Commercial use allowed under Apache terms
Total parameters 740M full / 270M text-only Vision +170M, audio +300M optional
Output dimension 768 native MRL truncation to 512, 256 or 128
Context window 8,192 tokens 4x EmbeddingGemma 1; up to ~5.5 minutes of audio
Languages 100+ claimed Multilingual text retained vs v1 on MTEB
Price Free weights No Google hosting fee for self-run weights; infra cost is yours

Modular loading is a practical detail. With sentence-transformers, you can disable unused encoders so text-only stays at 270M, text plus vision at 440M, text plus audio at 570M, or full multimodal at 740M. Google says embeddings from those setups still share one vector space.

On a Google Pixel 11 Pro, Google reports about 191 MB active RAM for quantized text-only weights and about 567 MB for the full multimodal model. Those numbers are vendor measurements on that device, not independent lab results.

Benchmarks: Where Scores Moved

The model card publishes full-precision results at 768 dimensions. Multilingual text MTEB (v2) is essentially flat versus EmbeddingGemma 1: 61.36 versus 61.15 mean task score.

Code retrieval is the clear jump. MTEB Code (v1) NDCG@10 rises from 68.76 to 78.68, a 9.92-point gain. Google’s Keyword post and developer guide both call that roughly a 14% relative improvement on code.

New modality tables appear for the first time in this line:

  • MIEB (lite) image mean: 64.64
  • MMEB v2 image Hit@1: 57.28
  • MMEB v2 visual documents NDCG@5: 67.84
  • MMEB v2 video Hit@1: 50.67
  • MSEB audio retrieval MRR@10: 69.54
  • MAEB mean: 49.39

Matryoshka truncation keeps most quality at 256 dimensions and cuts storage about 3x. At 128 dimensions, Google warns multimodal quality drops harder (MMEB overall falls from 59.01 to 45.65 in the published table), so 128d is mainly for text-first indexes.

Treat every score as Google-reported until third-party leaderboards rerun the same protocols.

How Developers Are Meant To Use It

The Google Developers Blog guide shows sentence-transformers (v6.1.0 or later) as the quickest path: google/embeddinggemma-2 on Hugging Face.

Text tasks use short instruction prefixes via prompt_name. Retrieval is asymmetric: queries use SearchQuery, documents use Document. Code search has its own CodeRetrieval prompt. Images, video and audio go in without those text prefixes.

Interleaved inputs mix modalities in one embedding. Placeholders like <|image|>, <|video|> and <|audio|> mark where media sits inside the text sequence. That lets a product listing with photos and a clip become one vector you can query with plain text.

Default sampling notes from the model card: video at 1 frame per second through the vision encoder, audio as 16 kHz mono. Image and frame token budgets are configurable.

Google also lists MediaPipe, LiteRT, transformers.js, WebGPU, vLLM, MLX, llama.cpp, Ollama, LM Studio and Qdrant among supported paths. Fine-tuning guidance points to Unsloth.

One sharp footgun is called out: do not run inference in float16. The activation range can produce NaNs or silently bad embeddings. Prefer bfloat16 where supported, otherwise float32.

What It Means For Indian Developers

Self-hosted embeddings cut recurring USD API spend for search and RAG. That helps startups in India that already pay for GPU time or want offline apps for healthcare, legal, education and enterprise knowledge bases where data residency matters.

Text-only at 270M is realistic on mid-range laptops and many cloud CPU instances. Full multimodal at ~567 MB quantized RAM on Pixel-class phones is still a stretch for older Android devices, but the modular path lets teams ship text first and add vision later without reindexing into a different space.

Code retrieval gains matter for local coding agents and private repos. Teams that already experiment with Gemma 4 voice stacks (see our earlier LiveKit note) can keep tokenizer and audio encoder alignment when pairing EmbeddingGemma 2 with Gemma 4 for on-device RAG.

Hugging Face and Kaggle downloads avoid waiting on a paid Google embedding endpoint. You still need local storage for the index, and you should validate Hindi and other Indian language quality on your own corpus rather than relying only on the multilingual MTEB average.

Confirmed Versus Unconfirmed

Claim Status
Apache 2.0 weights on Hugging Face and Kaggle Confirmed in Google launch posts
740M / modular encoder sizes and 8K context Confirmed in model card
MTEB Code 78.68 vs 68.76 Vendor-reported; not independently audited here
Pixel 11 Pro RAM figures (~191 MB / ~567 MB) Vendor-reported on that device
Gemini Enterprise Agent Platform Model Garden availability Google says “coming soon”, not live at launch
Best-in-class among all sub-1B multimodal embedders Marketing claim; compare against peers on your data

Limits And Safety Notes

EmbeddingGemma 2 is a pre-trained embedding model. The model card states it does not get generative-style post-training alignment or output moderation. Safety work focused on pre-training filters.

Downstream risk sits in how you retrieve and rank content. Google puts application-level filtering and fairness testing on the deployer. The Gemma Prohibited Use Policy still applies.

Training data cutoff is January 2025 per the model card. Language performance is not claimed to be equal across all 100+ languages. Omitting task prefixes works but reduces precision on text tasks.

Related Reading On AI Magazine

For more open-weight and local-stack coverage, see Gemma 4 31B on LiveKit, Reflection AI Beam waitlist, AirLLM on low VRAM, and Google’s Gemini free-tier model limits.

FAQ

What is EmbeddingGemma 2?

It is Google DeepMind’s open multimodal embedding model, released 6 Oct 2026 under Apache 2.0. It turns text, code, images, video and audio into vectors in one 768-dimensional space for search, RAG and classification.

How large is EmbeddingGemma 2 and can it run on a phone?

Full size is 740 million parameters. Text-only is 270 million. Google reports quantized footprints around 191 MB RAM (text) and 567 MB (full) on a Pixel 11 Pro. Older phones may struggle with the full multimodal load.

Is EmbeddingGemma 2 free?

The weights are free to download under Apache 2.0. You pay for your own hardware, hosting and vector database. Google also said Model Garden availability on Gemini Enterprise Agent Platform is coming soon.

How much better is it at code search than EmbeddingGemma 1?

On Google’s MTEB Code (v1) numbers, the mean NDCG@10 score moves from 68.76 to 78.68. That is a 9.92-point gain, described by Google as about 14% relative improvement. Independent reruns were not available for this article.

Which libraries support EmbeddingGemma 2 today?

Google’s developer guide highlights sentence-transformers, Hugging Face Transformers, MediaPipe, LiteRT, vLLM, MLX, llama.cpp, Ollama, LM Studio and related tools. Use bfloat16 or float32, not float16.

Share this article

1 comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Loading the next article…

Continue reading