Google has just launched an AI that understands text, video, images and audio at the same time: this is Gemini Embedding 2

Foto del autor

By Jack Ferson

Google has presented a new artificial intelligence model focused on multimodal information analysis. Named Gemini Embedding 2the system is currently available in public preview, marking an important step towards simultaneous processing of text, images, video and audio.

Unlike generative models, such as Gemini 3, embedding They do not focus on creating new content, but rather on understanding and representing information.

For this, convert different types of data into mathematical vectors, which machines can easily analyze. This capacity allows you to perform tasks such as semantic search, classification and grouping of information, offering more precise and contextual results than systems based only on keywords.

While Google’s first embedding model only worked with text, Gemini Embedding 2 expands the approach to integrate multiple types of content within a single representation space. The model processes text, images, video, audio, and documents, and can capture semantic intent in more than 100 languages.

According to Google, this system “simplifies complex processes and improves a wide variety of downstream multimodal tasks, from augmented generation by retrieval and semantic search to sentiment analysis and data clustering.” In addition, it allows analyzing relationships between different types of content, processing requests that simultaneously include text and imageswhich facilitates a combined analysis of the information.

Among the possible uses, Google highlights the legal field: during discovery processes, professionals could use Gemini Embedding 2 to locate critical information among millions of records more efficiently.

The model is available in public preview through the Gemini API and Vertex AI.

Deja un comentario