Check out TIPSv2 at our #CVPR2026 booth kiosk at 4pm! A foundational image-text encoder with spatial awareness, leading to strong results for vision & multimodal applications. Presented by Andre Araujo, Erik de Godoy & Gabriele Berton. https://t.co/OgVuQQKTeK @GoogleDeepMind https://t.co/iuuHByue4C
Google DeepMind's TIPSv2 Advances Multimodal AI with Enhanced Spatial Awareness
GoogleGoogle DeepMind is presenting TIPSv2, a new foundational image-text encoder, at CVPR 2026. This model enhances spatial awareness and patch-text alignment, improving performance across vision and multimodal applications, including strong gains in zero-shot segmentation.
- Evaluated across
- 9 tasks, 20 datasets
- Training parameter reduction (Head-only EMA)
- 42%
- Zero-shot segmentation gain (iBOT++)
- +14.1 mIoU on ADE150
- Pretraining improvements
- iBOT++, Head-only EMA, Multi-Granularity Captions
- TIPSv2 (ViT-L) vs DINOv3 (ViT-L)
- Wins 4 of 6 shared evaluations
- TIPSv2-g vs PE-core G/14
- Outperforms on 3 of 5 shared evals
This advancement matters as it addresses a surprising finding where smaller distilled models sometimes outperformed larger teachers in patch-text alignment, motivating TIPSv2's targeted improvements. The model demonstrates smoother feature maps with stronger semantic focus, delineating object boundaries more precisely than prior vision-language models like TIPS, SigLIP2, and DINOv3. TIPSv2 achieves performance generally on par with or better than recent vision encoders, with notable gains in zero-shot segmentation.
TIPSv2 is evaluated across 9 tasks and 20 datasets, with its ViT-L variant winning 4 of 6 shared evaluations against DINOv3 (ViT-L) despite DINOv3's teacher using significantly more parameters and images. You can explore TIPSv2's representations through a feature explorer for patch embeddings, zero-shot segmentation, and depth/normal prediction, with resources available on GitHub and HuggingFace.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →




