A fair LLM API token cost comparison starts with your own workload, not the rate card. Count the tokens your real prompts produce on each provider, then split the bill into fresh input, cached input, and output, and add the retries you see in production logs.
- The same text can bill as different token counts on different providers, so tokenize your own prompts instead of estimating from word count.
- Cached input is billed at roughly one tenth of fresh input on the models checked here, so a high cache hit rate can change which model is cheapest.
- Output tokens cost about five to six times input tokens on the flagship models listed, so average answer length drives most of the bill.
- Cache writes cost more than normal input, and retries plus reasoning tokens are billed without appearing in your feature request logs.
- Recheck the official pricing pages and record the date, because model names and rates change between comparisons.
Token Cost Comparison Starts With the Workload
A posted price is only one input. Your real cost is input tokens, cached input tokens, output tokens, and retries. A rate card tells you which model is cheaper per token, not which one is cheaper for your product.
Step 1: Count Tokens for the Same Text
Providers split text into tokens in different ways. The same support reply may bill as 700 tokens on one model and 900 on another. Use each provider's token counter instead of guessing from word count.
- Run the same three real prompts through every provider's counter and record the totals.
- Count the system prompt, tool schemas, and few-shot examples as well.
- Keep input and output totals separate, because they are billed at different rates.
Step 2: Compare Three Rates, Not One
Billable input splits into fresh input, cached input, and cache writes. Output usually costs several times more than input. Write down the date you checked each number.
| Provider and model | Input | Cached input | Output |
|---|---|---|---|
| OpenAI gpt-5.6-terra | $2.00 | $0.20 | $12.00 |
| OpenAI gpt-5.6-luna | $0.20 | $0.02 | $1.20 |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
Values are US dollars per one million tokens, taken from the official pricing pages checked on 2026-09-29.

Step 3: Measure Four Multipliers in Your Logs
Numbers from your own logs matter more than any benchmark.
- Cache hit rate: the share of input tokens served from cache.
- Output length: average output tokens per request, retries included.
- Retry rate: failed calls that are billed again on the next try.
- Reasoning tokens: hidden output tokens billed at the output rate.
A cheaper model can still cost more when it writes longer answers or needs more retries.
Step 4: Put the Numbers Into One Formula
cost = fresh_input/1e6 * input_rate
+ cached_input/1e6 * cached_rate
+ output/1e6 * output_rate
Run this on one week of logged usage. If you do not record token counts per request, start there, because no comparison is reliable without them.
Fix the Mistakes That Distort Results
- Comparing input-heavy and output-heavy jobs with one blended number.
- Ignoring cache writes, which cost more than normal input.
- Reusing last quarter's prices after a price change.
- Leaving out search and tool fees billed outside token usage.
Check the Official Pages Before Deciding
Model names and prices change often, so open the source page and note the date you checked it.
Sources
- Pricing | OpenAI API OpenAI · 2026-09-29
- Pricing – Claude Platform Docs Anthropic · 2026-09-29
- Claude Platform model pricing Anthropic · 2026-09-29