Google introduced Gemini Embedding 2 as its first natively multimodal embedding model. Embeddings represent content as vectors for tasks such as similarity search and retrieval; this is different from a chat model generating an answer.
The launch specifications described up to 8,192 text tokens, up to six PNG or JPEG images, MP4 or MOV video up to 120 seconds, native audio input, and PDFs up to six pages.

Putting different media types in a shared space can help search for related material across text, images, and recordings. Google reported strong evaluation results, but performance on a particular collection still needs testing.
See Google’s announcement for the launch details and current documentation links.
Adapted from the original Chinese article, published on March 11, 2026.

Comments NOTHING