Google DeepMind's TIPSv2 Advances Multimodal AI with Enhanced Spatial Awareness

GoogleGoogle

Google DeepMind is presenting TIPSv2, a new foundational image-text encoder, at CVPR 2026. This model enhances spatial awareness and patch-text alignment, improving performance across vision and multimodal applications, including strong gains in zero-shot segmentation.

TIPSv2 is the next generation of Google DeepMind's foundational image-text encoders, designed to improve performance across multimodal and vision tasks. It introduces enhanced patch-text alignment—the ability for visual regions to precisely correspond with text descriptions—through an improved pretraining process. Key changes include iBOT++ for stronger dense alignment, Head-only EMA for reduced training cost, and Multi-Granularity Captions using PaliGemma and Gemini descriptions for richer text supervision.
Evaluated across
9 tasks, 20 datasets
Training parameter reduction (Head-only EMA)
42%
Zero-shot segmentation gain (iBOT++)
+14.1 mIoU on ADE150
Pretraining improvements
iBOT++, Head-only EMA, Multi-Granularity Captions
TIPSv2 (ViT-L) vs DINOv3 (ViT-L)
Wins 4 of 6 shared evaluations
TIPSv2-g vs PE-core G/14
Outperforms on 3 of 5 shared evals

This advancement matters as it addresses a surprising finding where smaller distilled models sometimes outperformed larger teachers in patch-text alignment, motivating TIPSv2's targeted improvements. The model demonstrates smoother feature maps with stronger semantic focus, delineating object boundaries more precisely than prior vision-language models like TIPS, SigLIP2, and DINOv3. TIPSv2 achieves performance generally on par with or better than recent vision encoders, with notable gains in zero-shot segmentation.

TIPSv2 is evaluated across 9 tasks and 20 datasets, with its ViT-L variant winning 4 of 6 shared evaluations against DINOv3 (ViT-L) despite DINOv3's teacher using significantly more parameters and images. You can explore TIPSv2's representations through a feature explorer for patch embeddings, zero-shot segmentation, and depth/normal prediction, with resources available on GitHub and HuggingFace.

Google Research
Google Research
@GoogleResearch
X

Check out TIPSv2 at our #CVPR2026 booth kiosk at 4pm! A foundational image-text encoder with spatial awareness, leading to strong results for vision & multimodal applications. Presented by Andre Araujo, Erik de Godoy & Gabriele Berton. https://t.co/OgVuQQKTeK @GoogleDeepMind https://t.co/iuuHByue4C

6retweets50likes
View on X

Still wondering? A few quick answers below.

TIPSv2 is Google DeepMind's next-generation foundational image-text encoder, designed to improve how AI systems understand and process visual information in conjunction with text. It focuses on enhancing spatial awareness and the alignment between image patches and their corresponding text descriptions.

TIPSv2 introduces three main pretraining improvements: iBOT++ for stronger dense patch-level alignment, Head-only EMA to reduce training costs while maintaining performance, and Multi-Granularity Captions that use richer descriptions from models like PaliGemma and Gemini for better text supervision.

TIPSv2 demonstrates strong performance across various vision and multimodal tasks, often matching or surpassing recent vision encoder models. It shows particularly strong gains in zero-shot segmentation and outperforms DINOv3 in several shared evaluations, even with DINOv3's larger teacher model.

Patch-text alignment refers to how accurately an AI model can connect specific visual regions (patches) within an image to corresponding words or phrases in a text description. Improved alignment means the model has a more precise understanding of what parts of an image relate to specific textual elements.

TIPSv2 is presented with a research paper, and its resources including GitHub, checkpoints, and a HuggingFace demo are available. These allow for exploring its patch embeddings, zero-shot segmentation, and other applications.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update