Gemini Embedding 2: Text, Images, Audio, Video, and PDFs

XMLans Posted on 2026-03-10 18 Views


Google introduced Gemini Embedding 2 as its first natively multimodal embedding model. Embeddings represent content as vectors for tasks such as similarity search and retrieval; this is different from a chat model generating an answer.

The launch specifications described up to 8,192 text tokens, up to six PNG or JPEG images, MP4 or MOV video up to 120 seconds, native audio input, and PDFs up to six pages.

Gemini Embedding 2 multimodal capabilities

Putting different media types in a shared space can help search for related material across text, images, and recordings. Google reported strong evaluation results, but performance on a particular collection still needs testing.

See Google’s announcement for the launch details and current documentation links.

Adapted from the original Chinese article, published on March 11, 2026.

Hi! I frequently update with various articles about technology, practical tips, and cutting-edge news. I hope it will be helpful to you!
Last updated on 2026-09-30