OpenAI restructures API pricing around cached and batch tiers
A simplified three-tier structure — realtime, cached, batch — replaces the growing menu of per-model discounts. Most workloads get cheaper, some get sharply more expensive. Demo content.
OpenAI has overhauled API pricing into three tiers — realtime, cached and batch — replacing the accumulated per-model discount matrix that had grown difficult to reason about.
The new structure
- ▸Realtime: standard interactive pricing, unchanged for most models
- ▸Cached: automatic discounts up to 90% for repeated prefixes
- ▸Batch: 50% off for asynchronous jobs with a 24-hour completion window
Winners and losers
High-volume automation pipelines with stable prompts benefit most. Chatbots with highly variable prompts and realtime requirements see effective prices rise on some models.
"This is pricing designed to reward exactly the workloads OpenAI wants: predictable, cacheable, schedulable." — API cost consultant
Competitors are expected to mirror the structure; Anthropic already offers similar cache economics, and Google's batch discounts are more aggressive still.
Source: OpenAI


