EmbeddingGemma 2 scores 78.68 on the Massive Text Embedding Benchmark (Code), a jump of nearly 10 points over its predecessor (68.76). That puts it on par with much larger models. | Image: Google
Source: the-decoder.com
Small enough to run locally
EmbeddingGemma 2 needs about 191 MB of RAM. Google says it can reduce the storage needed for a local vector database by up to six times. For text-only tasks, a 270-million-parameter version is enough.
Paired with small open models such as Gemma 4, it can support offline applications with retrieval-augmented generation. The model weights are available on Hugging Face and Kaggle, alongside developer guidance and documentation.
The claim that needs a benchmark
The size comparison is the headline claim, but the announcement provides no benchmark results in the material here. I think the more important test is whether the model’s retrieval quality holds up on real workloads, especially when it is asked to search across images and video as well as text.
Local execution changes the trade-off: developers can keep data on-device and avoid an API dependency, but they still need to establish whether the model is accurate enough for their use case. A 191 MB memory requirement and 20–70 millisecond requests make that evaluation easier to attempt in a browser; they do not answer it.
Source: the-decoder.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X