OpenRouter Reveals Real-Time Cache Hit Rates and Effective LLM Pricing by Provider

OpenRouterOpenRouter

· Updated

OpenRouter now displays real-time cache hit rates and historical traffic data on its Pricing tab. This update provides transparency into how different model providers compare on effective pricing for LLMs like Anthropic's Claude Opus 4.8, enabling users to optimize costs.

OpenRouter, a unified API for accessing large language models (LLMs), has added real-time cache hit rates and historical traffic data to its Pricing tab. This new visibility allows users to compare the effective cost and performance of models such as Anthropic's Claude Opus 4.8 across various providers. A cache hit rate indicates how often a repeated request is served from memory rather than being re-processed by the model.
Weighted Avg Input Price
$1.85/M tokens
Weighted Avg Output Price
$25.00/M tokens
Highest Provider Cache Hit Rate
82.9% (Amazon Bedrock)
Lowest Provider Cache Hit Rate
15.4% (Amazon Bedrock (US))
Claude Platform on AWS Input Price
$1.43/M tokens
Claude Platform on AWS Cache Hit Rate
82.2%

This update addresses the challenge of optimizing inference costs, which can vary significantly based on provider-specific caching strategies. By exposing these metrics, OpenRouter moves beyond static pricing, offering dynamic insights into actual costs per million tokens. This builds on its existing Model Comparison Tool by adding more granular data for provider selection.

On the OpenRouter Pricing tab, you can now view provider-specific effective input and output prices, alongside cache hit rates and token share. For Claude Opus 4.8, the weighted average input price is $1.85/M tokens, with providers like Claude Platform on AWS showing an 82.2% cache hit rate.

OpenRouter
OpenRouter
@OpenRouter
X

How do different model providers differ on cache hit rate and effective price? Now you can see real-time cache hit rate and historical traffic from the Pricing tab. Here's Opus 4.8: https://t.co/vkeUZLh7r9 https://t.co/z5WvzKcRK2

7retweets111likes
View on X

Still wondering? A few quick answers below.

OpenRouter's Pricing tab now displays real-time cache hit rates and historical traffic data, allowing users to see how different model providers perform on these metrics.

A higher cache hit rate means more requests are served from a cached response, reducing the need to re-process prompts through the LLM. This can significantly lower the effective cost per million tokens and improve response times.

The update specifically highlights the effective pricing and performance data for Anthropic's Claude Opus 4.8, showing how its costs and cache hit rates vary across different providers.

For Claude Opus 4.8, providers like Claude Platform on AWS, Anthropic, Amazon Bedrock (US), and Google Vertex show different cache hit rates, ranging from 15.4% to 82.9%.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update