How do different model providers differ on cache hit rate and effective price? Now you can see real-time cache hit rate and historical traffic from the Pricing tab. Here's Opus 4.8: https://t.co/vkeUZLh7r9 https://t.co/z5WvzKcRK2
OpenRouter Reveals Real-Time Cache Hit Rates and Effective LLM Pricing by Provider
· Updated
OpenRouter now displays real-time cache hit rates and historical traffic data on its Pricing tab. This update provides transparency into how different model providers compare on effective pricing for LLMs like Anthropic's Claude Opus 4.8, enabling users to optimize costs.
- Weighted Avg Input Price
- $1.85/M tokens
- Weighted Avg Output Price
- $25.00/M tokens
- Highest Provider Cache Hit Rate
- 82.9% (Amazon Bedrock)
- Lowest Provider Cache Hit Rate
- 15.4% (Amazon Bedrock (US))
- Claude Platform on AWS Input Price
- $1.43/M tokens
- Claude Platform on AWS Cache Hit Rate
- 82.2%
This update addresses the challenge of optimizing inference costs, which can vary significantly based on provider-specific caching strategies. By exposing these metrics, OpenRouter moves beyond static pricing, offering dynamic insights into actual costs per million tokens. This builds on its existing Model Comparison Tool by adding more granular data for provider selection.
On the OpenRouter Pricing tab, you can now view provider-specific effective input and output prices, alongside cache hit rates and token share. For Claude Opus 4.8, the weighted average input price is $1.85/M tokens, with providers like Claude Platform on AWS showing an 82.2% cache hit rate.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →


