Google launches EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0
- EmbeddingGemma 2 maps text, images, audio, and video into one shared embedding space, built on the Gemma 4 architecture with 740M parameters and released under the commercially permissive Apache 2.0 license.
- The model is modular: 270M parameters cover text-only workloads, with optional vision (170M) and audio (300M) encoders, so developers enable only what they need.
- Matryoshka Representation Learning lets output vectors be truncated from 768 dimensions to 512, 256, or 128, cutting storage for local vector databases and memory use by up to 6x.
- On a Pixel 11 Pro the quantized model needs as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.
- Context grew 4x over EmbeddingGemma 1 to 8K tokens, enough for 5.5 minutes of audio, 29 images, or 58 video frames on local hardware, and code performance rose 9.92 points on MTEB Code from 68.76 to 78.6.
Hacker News opinions
Finally, a good moderate-size embedding model. 270M for text only, 440M for text plus vision, and it's multimodal on top of that. I had a local embedding tool calibrated for EmbeddingGemma that I sat on until something better showed up.
Audio too, not just vision with video. I want to know how it handles pairs where a word and the audio where it's spoken land near each other. I messed with CLIP and image embeddings carry both the clean text content and the stylistic tone, and the two are almost linearly separable.
Numbers from my M3 Pro: 78 text embeddings per second on short texts, 4 per second for images, 6 per second for 30-second audio chunks, and 0.2 per second per minute of video since the encoder runs at 1fps. Not bad for a laptop running a model this size, and an M5 Ultra should beat it 5x.
Apache 2.0 is the whole point for embeddings. You compute millions of vectors and store them, and a proprietary vendor will retire the model eventually, so you pay to re-embed everything. I'd still rather pay a provider, as long as I can fall back to the open weights.
Golden tests are the better answer. Even switching from CPU to GPU can give you different tokens.
Why is nobody comparing this to SigLIP 2, also from Google? Is it because that one isn't fully multimodal, or because it's a different team?
EmbeddingGemma 2 has no decoder, so you can't use it for OCR for one.
Hats off to Google for putting out open weights that are probably close to what they ship on Android phones anyway.
The parameter split is the interesting part: 270M text, 170M vision, 300M audio. Vision is smallest because it's just a still image, text has to deal with the entropy of human language, and audio is meaningless without time.
I'd like to see text comparisons against the Voyage AI embedding models. They beat the Qwen ones for me in the past.
Voyage was never a serious contender outside super-niche business domains. I'd bet this wins in 90%+ of use cases.
Unlike earlier on-device embedding models this one uses MRL, not MatFormers, so you don't get to shrink the weights along with the lower-dimensional embeddings. Probably because there's no good research on MatFormers for multimodal yet.
You can also use this for text and image classification like the MediaPipe examples. I ran the flight-refund case locally and it failed, said the request didn't involve a refund with p(true) of 0.22, while Laya got it right and was almost as fast.
For text the benchmarks are identical to the first EmbeddingGemma. The difference is you can enable or disable the parts you don't need and keep text only.
What are people actually using on-device multimodal embeddings for, and what's the hallucination rate like?
Semantic image search via text: encode the images and the query with the same model, then take nearest neighbors. That's the main one.
For law practice I'd search a case file for 'undamaged roof before Hurricane Katrina' versus 'damaged roof after' and pull both the deposition testimony and the photos.
JetBrains wrote about binary quantization instead of MRL the other day. Would that work with EmbeddingGemma 2 or is it incompatible?