Token Cost Calculator: Compare API Prices Across Providers
This page puts every major model's list price in one table and one calculator. Enter your input tokens, cached tokens, and output tokens, and the calculator returns the cost per request and per month for each model.
Current list prices, all providers
List prices verified 2026-08-10. All prices are USD per 1 million tokens, standard tier.
| Model | Input $/1M | Cached input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $1.00 | $50.00 | 1M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 1M |
| Claude Opus 4.8 | $5.00 | $0.50 | $25.00 | 1M |
| Claude Sonnet 5 | $3.00 | $0.30 | $15.00 | 1M |
| Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200K |
| Gemini 3.6 Flash | $1.50 | – | $7.50 | – |
| Gemini 3.5 Flash | $1.50 | – | $9.00 | – |
| Gemini 3.5 Flash-Lite | $0.30 | – | $2.50 | – |
| Gemini 3.1 Flash-Lite | $0.25 | – | $1.50 | – |
| Gemini 2.5 Pro (≤200K prompt) | $1.25 | – | $10.00 | 1,048,576 |
| Gemini 2.5 Pro (>200K prompt) | $2.50 | – | $15.00 | 1,048,576 |
| Gemini 2.5 Flash | $0.30 | – | $2.50 | 1,048,576 |
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | 1,050,000 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | – |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | – |
| GPT-5.5 | $5.00 | $0.50 | $30.00 | 1M |
| GPT-5 | $1.25 | $0.125 | $10.00 | 400K |
| GPT-5 mini | $0.25 | $0.025 | $2.00 | – |
| GPT-5 nano | $0.05 | $0.005 | $0.40 | – |
| GPT-4o | $2.50 | $1.25 | $10.00 | 128K |
| GPT-4o mini | $0.15 | $0.075 | $0.60 | 128K |
| GPT-4.1 | $2.00 | $0.50 | $8.00 | 1M |
| DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 | 1M |
| DeepSeek V4 Pro | $0.435 | $0.003625 | $0.87 | 1M |
| Kimi K3 | $3.00 | $0.30 | $15.00 | 1,048,576 |
| Kimi K2.6 | $0.95 | $0.16 | $4.00 | 262,144 |
Anthropic cached input is the prompt cache read price, 0.1x the input price. Claude Sonnet 5 has introductory pricing of $2.00 in / $10.00 out through 2026-08-31. Gemini 2.5 Pro switches to the higher price tier when the prompt exceeds 200K tokens; Gemini 3.x Flash models bill flat rates. DeepSeek's documentation states an overall price increase is planned in the near future.
How to compare these prices
Which column matters depends on your workload.
For chat workloads, the output price dominates. A chat response is long relative to its prompt, and output tokens cost 3x to 8x more than input tokens on most models in the table. Two models with similar input prices can produce very different bills once output volume grows.
For agent workloads, the cached input price dominates. An agent re-sends a large stable prefix, system prompt, tools, and history, on every step. When that prefix is cached, most of your input tokens bill at the cached rate. GPT-5 charges $0.125 per 1M cached input tokens against $1.25 uncached, and Claude cache reads bill at 0.1x the input price. The cached column is the one to compare for any loop that runs many steps.
Tokens are not comparable across providers
Each provider counts tokens with its own tokenizer, so identical text produces different token counts on different platforms. The same difference exists inside a provider. Claude models from Opus 4.7 onward use a newer tokenizer that produces roughly 1x to 1.35x the tokens of the older family on the same text. A price-per-token comparison is only exact after you count the tokens with each provider's own counter. Use the token counter before you trust a cross-provider estimate.
Per-provider pricing pages
All Claude models, cache and batch discounts, tokenizer differences by generation.
Gemini 2.5 Pro and Flash, and the 200K prompt price tier explained.
The full GPT lineup with cached input prices for every model.
Context limits also constrain what a model can do at any price. See context windows compared.