I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: "Do a where's Waldo style image but it's where is the raccoon holding a ham radio" https://t.co/KuPdFAEWUl
Simon Willison Benchmarks ChatGPT Images 2.0 Spatial Reasoning and Detail Gains
Simon Willison· Updated
Simon Willison tested OpenAI's new ChatGPT Images 2.0 using a custom Where's Waldo style benchmark to evaluate spatial reasoning and detail retention. His findings suggest the model significantly outperforms previous versions in rendering dense, complex scenes with precise object placement.
gpt-image-2 model successfully rendered specific, tiny details that previous models missed, demonstrating significantly improved spatial accuracy.- Model ID
- gpt-image-2
- Maximum resolution
- 3840x2160
- Cost per 4K image
- $0.40
- Output tokens per 4K image
- 13,342 tokens
- Benchmark
- Custom Where's Waldo test
- Quality setting
- high
These findings highlight a broader industry shift from "artistic vibes" to functional precision. Simon's results mirror the spatial accuracy gains recently seen in Google's models. His testing on the newly launched ChatGPT Images 2.0 confirms that high-resolution outputs can now handle dense infographics.
For developers, this benchmark suggests that gpt-image-2 is a strong candidate for technical workflows requiring high-resolution detail. While the model uses approximately 13,342 output tokens per 4K image—costing roughly $0.40—the trade-off provides significantly better accuracy. Access these capabilities via the OpenAI API using the high quality setting.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →



