TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Google announced EmbeddingGemma 2 on Oct. 6, 2026, an Apache 2.0-licensed embedding model designed to represent text, code, images, audio and video in a shared vector space. The 740-million-parameter model is intended for on-device use; Google reports benchmark gains and low memory requirements, but independent verification and real-world performance details were not included in the announcement.
Google announced EmbeddingGemma 2 on Oct. 6, describing it as an open, 740-million-parameter multimodal embedding model for representing text, code, images, audio and video in a shared space. Released under the Apache 2.0 license, it is designed to run on consumer hardware, allowing developers to build local search and retrieval systems without sending every item to a cloud service.
The model extends the text-focused EmbeddingGemma, which Google says has been downloaded more than 20 million times, to support cross-modal retrieval. In Google’s examples, a developer could search video clips using a voice memo or query audio recordings with text. Those are described as possible applications; the company did not provide product-level deployments or independent user results in its announcement.
Google says EmbeddingGemma 2 is built on the Gemma 4 architecture and can be configured for different workloads. The company lists a text-only option with as few as 270 million parameters, plus optional vision and audio encoders identified as 170 million and 300 million parameters. Its output vectors can be reduced from 768 dimensions to 512, 256 or 128 using Matryoshka Representation Learning, which Google says can cut local vector storage and memory use by up to six times.
The model has an 8,000-token context window, four times the window of EmbeddingGemma 1, according to Google. The company says it can handle up to 5.5 minutes of audio, 29 images or 58 video frames, as well as combinations of those inputs. On a Google Pixel 11 Pro, Google reports active RAM use as low as about 191 MB for quantized text-only weights and about 567 MB for the full multimodal model. These figures are vendor-reported and depend on configuration.
Local Search Across Media Types
Embedding models convert content into numerical vectors so software can find items with related meaning, rather than relying only on exact words. Bringing several media types into a shared embedding space could let developers build search tools that connect a spoken request to a video, or retrieve audio and images alongside text. The model’s potential value is that these operations may run directly on a device, including when a network connection is unavailable.
Local processing can reduce the need to upload personal recordings, documents or code to a remote service. It may also reduce network delays and cloud-processing costs. Those benefits are not automatic: developers still need to implement storage, indexing and access controls, and actual speed and privacy depend on the complete application. Google’s announcement describes the model’s design goals, not a guarantee that every deployment will remain offline or keep all data local.
The model’s size and adjustable vector dimensions also matter for products with limited memory or storage. If Google’s reported performance holds in independent testing, developers could have more room to add semantic retrieval to phones, browsers and other edge devices. The key question for adopters is how quality, speed and memory use compare on their own hardware and data.
on-device multimodal search engine
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Text Embeddings to Multimodal Retrieval
Google introduced the original EmbeddingGemma as a lightweight option for text embeddings, which can support semantic search, recommendations and retrieval-augmented generation. The company says developers downloaded it more than 20 million times, though it did not specify the measurement period or provide a breakdown of usage in this announcement.
EmbeddingGemma 2 keeps that focus on compact, locally usable models while adding image, video and audio inputs. Google says its text tokenizer and audio encoder are shared with Gemma 4, which may allow the models to work together in an on-device retrieval-augmented generation pipeline with a lower combined memory footprint. The company points developers to its AI Edge tools, including MediaPipe and LiteRT, and says the model weights are available through Hugging Face and Kaggle. Availability in Gemini Enterprise Agent Platform Model Garden was described as coming soon.
Google reports a rise from 68.76 to 78.68 on MTEB Code compared with EmbeddingGemma, a 9.92-point gain. It also says the model scores strongly against other sub-billion-parameter systems across text, vision and audio benchmarks, including MTEB Code and MAEB. These are company-reported benchmark results; the announcement directs readers to the model card for evaluation details.
“EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings.”
As an affiliate, we earn on qualifying purchases.
Independent Results Still Pending
The announcement does not include independent benchmark testing, detailed comparisons across specific competing models, or results from production applications. Google’s claims about leading quality for models below one billion parameters and performance against larger or specialist systems should be read as vendor-reported findings until the model card and outside evaluations can be reviewed.
Some practical details also remain dependent on implementation. Google gives example input capacities and RAM figures, but does not specify in the announcement the full range of devices, processing speeds, power consumption or trade-offs associated with each configuration. It is also not clear how performance varies across languages, media formats and noisy or complex real-world inputs.
The company says the model is available on Hugging Face and Kaggle, but does not state whether all versions, encoders and optimized builds are available at once. Enterprise Model Garden availability is still pending. Developers will need to check the model card, license terms and deployment documentation for the precise weights and requirements they plan to use.
multimodal content retrieval device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Model Tests and Tool Support
Developers can access EmbeddingGemma 2 weights through Hugging Face and Kaggle, then test deployment with Google AI Edge MediaPipe or LiteRT. Google also lists support pathways through tools including Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, and references Unsloth guidance for fine-tuning.
The next useful milestone is independent evaluation on representative devices and datasets, including measurements of retrieval quality, latency, memory use and power consumption. Google says availability in Gemini Enterprise Agent Platform Model Garden is planned for later, but the announcement gives no date. Until those results and timelines are clearer, prospective users can assess the released weights and documentation directly.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is EmbeddingGemma 2?
It is Google’s 740-million-parameter embedding model, designed to represent text, code, images, audio and video in a shared space for search and retrieval tasks.
Is EmbeddingGemma 2 open source?
Google says it is released under the Apache 2.0 license, a commercially permissive license. The model weights are listed on Hugging Face and Kaggle; developers should check the accompanying license and model documentation for the specific release they use.
Can it run on a phone?
Google designed the model for on-device inference and reports that, when quantized on a Pixel 11 Pro, active RAM use can be about 191 MB for text-only weights or 567 MB for the full multimodal model. Those are company-provided figures, not independent device tests.
What is new compared with EmbeddingGemma?
The earlier model focused on text. EmbeddingGemma 2 adds image, video and audio support, an 8,000-token context window, and, according to Google, a 9.92-point improvement on MTEB Code.
Are its benchmark claims independently verified?
The announcement reports Google’s benchmark results and points to a model card, but it does not provide independent verification. External evaluations and tests on target devices can help establish how the model performs in specific applications.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
