Simon Willison Benchmarks ChatGPT Images 2.0 Spatial Reasoning and Detail Gains

Simon WillisonSimon Willison

· Updated

Simon Willison tested OpenAI's new ChatGPT Images 2.0 using a custom Where's Waldo style benchmark to evaluate spatial reasoning and detail retention. His findings suggest the model significantly outperforms previous versions in rendering dense, complex scenes with precise object placement.

Simon Willison, creator of Datasette, tested the new ChatGPT Images 2.0 model using a custom "Where's Waldo" style benchmark. By prompting for a raccoon holding a ham radio, he found that the gpt-image-2 model successfully rendered specific, tiny details that previous models missed, demonstrating significantly improved spatial accuracy.
Model ID
gpt-image-2
Maximum resolution
3840x2160
Cost per 4K image
$0.40
Output tokens per 4K image
13,342 tokens
Benchmark
Custom Where's Waldo test
Quality setting
high

These findings highlight a broader industry shift from "artistic vibes" to functional precision. Simon's results mirror the spatial accuracy gains recently seen in Google's models. His testing on the newly launched ChatGPT Images 2.0 confirms that high-resolution outputs can now handle dense infographics.

For developers, this benchmark suggests that gpt-image-2 is a strong candidate for technical workflows requiring high-resolution detail. While the model uses approximately 13,342 output tokens per 4K image—costing roughly $0.40—the trade-off provides significantly better accuracy. Access these capabilities via the OpenAI API using the high quality setting.

Simon Willison
Simon Willison
@simonw
X

I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: "Do a where's Waldo style image but it's where is the raccoon holding a ham radio" https://t.co/KuPdFAEWUl

2retweets33likes
View on X

Still wondering? A few quick answers below.

Simon Willison's custom benchmark tested the ability of ChatGPT Images 2.0 to render specific, small objects within a dense and complex scene. He found that the new model successfully included a raccoon holding a ham radio in a crowded illustration, a task that previous versions and some competitors failed to complete accurately.

According to Simon Willison's testing, ChatGPT Images 2.0 excels at spatial reasoning and maintaining detail in high-resolution 4K outputs. The model can accurately render complex scenes with many small elements and precise text, making it a significant improvement over earlier versions that often hallucinated or omitted specific details in crowded visual prompts.

Developers should consider ChatGPT Images 2.0 for technical illustrations and infographics that require high precision. While high-quality 4K images cost approximately forty cents each due to high token usage, the model provides the accuracy needed for functional design. Access is available via the API by specifying the high quality setting and the gpt-image-2 model identifier.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update