Gemini Embedding 2 is now generally available in the Gemini API and Vertex AI! Start building with our first natively multimodal embedding model, now equipped with the stability and optimizations required for production apps. https://t.co/howDTPbG1Q
Google Launches Gemini Embedding 2 for Production Multimodal Search Applications
Google· Updated
Google's first natively multimodal embedding model, Gemini Embedding 2, is now generally available via the Gemini API and Vertex AI. The update enables developers to map text, images, audio, video, and documents into a single vector space for production-grade semantic search and clustering. This transition from preview to stable status introduces flexible vector dimensions and a new prompt-based task instruction system.
- Input token limit
- 8,192 tokens
- Output dimensions
- 128 to 3072
- Video duration limit
- 120 seconds
- Audio duration limit
- 180 seconds
- Image limit
- 6 images per request
- Document limit
- 6 PDF pages
- Batch API pricing
- 50% discount
This update simplifies Retrieval-Augmented Generation (RAG) systems (grounding AI responses with external data) that handle diverse media. By mapping modalities into one embedding space, you can perform cross-modal searches without separate, aligned pipelines. The model also uses Matryoshka Representation Learning to allow flexible vector sizes without quality loss.
Integrate these embeddings via the Gemini API or Vertex AI. The embedding space is incompatible with previous versions, requiring a full re-embedding of data. For high-volume tasks, the Batch API offers a 50% discount. You must now use specific task prefixes in prompts instead of the legacy task_type parameter.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →



