Google DeepMind Shares White Paper on Gemini Embedding 2, Its First Native Multimodal Embedding Model
Google DeepMind said it is sharing a white paper for Gemini Embedding 2, or GE 2, which it described as its first native multimodal embedding model. The model is intended to provide a unified representation for text, audio, video and image inputs.
The announcement focused on the research paper rather than broader product details. Beyond describing GE 2 as a single representation across four input types, Google DeepMind did not provide additional information on availability or performance benchmarks in the post.
From the sources (2 posts)
@googledeepmindRT @mseyed: Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini 🚀 Today, we’re sharing the @GoogleDeepMind white paper for…
@mseyedGemini Embedding 2: A Native Multimodal Embedding Model from Gemini 🚀 Today, we’re sharing the @GoogleDeepMind white paper for GE 2, our first native multimodal embedding model. Whether it’s text, audio, video, or image, GE 2 provides a un