Google launched EmbeddingGemma 2. This open model converts text, images, video, audio, and code into numerical vectors. These vectors help systems find and compare similar content easily.
The model contains 740 million parameters. Google states this makes it the most compact model of its kind. It outperforms competing models up to twice its size on multimodal embedding benchmarks.
The model runs locally without an API key
Each query takes about 20 to 70 milliseconds via WebGPU in the browser. The system requires only around 191 megabytes of RAM. It also cuts local vector database storage by up to six times.
Developers handling text-only tasks can use a smaller version. That version contains 270 million parameters.

Users can build offline applications
Developers can pair the software with small open models like Gemma 4. This setup allows users to run offline retrieval-augmented generation applications without sending data to external servers.
Google published the model weights on Hugging Face and Kaggle. The company also provided a developer guide and documentation.



