AI API token cost is not one rate; it is the sum of uncached input, cached input, cache writes, and output tokens, each billed separately. The most controllable reductions usually come from routing asynchronous jobs to batch pricing, keeping stable prompt prefixes cacheable, and measuring cost per completed task instead of cost per token. OpenAI and Anthropic both document these buckets and their batch and caching discounts on their official pricing pages, checked on 2026-09-29.
- Cached input reads are billed at roughly one tenth of the standard input rate on the OpenAI and Anthropic pricing pages, while cache writes carry a premium.
- Batch tiers are priced near half of standard rates for supported models, which makes asynchronous jobs such as classification, enrichment, and evaluation good candidates.
- Prompt caching only pays off while the cached prefix stays identical, so timestamps, reordered tool definitions, or user-specific text placed early in the prompt quietly cancel the discount.
- Tracking tokens per request and cost per completed task surfaces regressions that aggregate spend hides, such as growing conversation history or duplicated retrieved context.
- Per-key budgets, usage alerts, and fast key rotation are cost controls as well as security controls, because valid API keys can be abused to generate billable requests.
Where AI API Token Cost Comes From
AI API token cost is not a single number. Each request is billed across separate buckets: uncached input, cached input reads, cache writes, and output. OpenAI's pricing page lists a rate per million tokens for each bucket, and Anthropic's page lists the same four categories. Cache reads usually cost about one tenth of standard input. Output stays the most expensive bucket, so a service that writes long answers drains budget faster than one sending a large but stable cached context. Prices were checked on 2026-09-29, and both vendors update them often, so treat any figure here as a snapshot rather than a promise.
Choosing a Model Tier by Cost per Task
Before comparing model names, decide how much latency each job really needs. OpenAI separates standard, batch, flex, and fast tiers, and batch pricing is about half the standard rate. Anthropic lists a similar split: batch processing is discounted by 50 percent, and its prompt caching page describes caching as a way to cut repeated input cost by up to 90 percent. A person waiting for tokens on screen may justify a fast tier, while classification, enrichment, and nightly reporting usually belong on the batch tier. Compare options by cost per finished task, not by cost per token.

Instrument Token Use Before You Optimize
You cannot tune what you cannot attribute, so record the token buckets your provider reports per response. Log input tokens, cached tokens, cache write tokens, output tokens, the model ID, the service tier, and a feature or team label for every call. Dividing daily spend by completed tasks gives a cost per task that stays comparable when a model or prompt changes. Then watch two signals: tokens per request trending up, and cache hit ratio trending down. Rising tokens with flat traffic usually mean a prompt grew or retrieved context is attached twice.
Diagnosing a Sudden Cost Spike
Cache misses cause most cost surprises. Prefix caching keys on the exact token sequence, so a timestamp or a reordered tool list invalidates it on every call. Retry loops are the second cause, because a timeout that triggers immediate retries multiplies spend without producing usable output. Third, keys leak: security reporting in September 2026 described botnets that use valid API keys to send billable requests and exhaust credits, so per-key budgets, anomaly alerts, and quick rotation are cost controls as well. Fix one cause at a time, then re-measure cost per task.
Sources
- OpenAI API Pricing OpenAI · 2026-09-29
- Claude API Pricing Anthropic · 2026-09-29
- Claude Platform API: prompt caching and batch processing Anthropic · 2026-09-29
- Hackers sell AI token draining as a service Cybernews · 2026-09-24 · 2026-09-29