Skip to content

Multimodal retrieval with embeddings in a shared space

In a post published on October 6, 2026, Google says EmbeddingGemma 2 represents text, code, images, video, and audio in a shared vector space and uses the Apache 2.0 license.

By Wendelmaques ·

Source: EmbeddingGemma 2 mira retrieval multimodal (blog.google). Text prepared with AI from this source.

What happened and what to do

In a post published on October 6, 2026, Google says EmbeddingGemma 2 maps text, code, images, video, and audio into a single shared vector space. The post also states that the model uses the Apache 2.0 license. This approach may simplify searches across different media types.

A company could assess a search use case spanning documents, images, recordings, and videos, then implement data collection and indexing and build a retrieval API for internal systems. The project should include relevance tests, access controls, and search-quality monitoring; the model and license should be assessed in the context of the application.

How the consultancy can help

Wendelmaques can assess data sources, search requirements, and necessary controls; scope an implementation of collection, indexing, an API, and monitoring; and support its operation on the company’s infrastructure.

Next step

Send a short description of your multimodal search case to receive a scoped proposal for diagnosis and implementation.

Consulting for your project

Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.

Quoted per project

Request a proposal