Prompt caching cost.
Agents resend the same system prompt and context on every turn. Those repeated tokens do not have to bill at the full rate, and on a long loop that difference is most of the invoice.
A 20-turn agent loop, with and without caching
Each turn carries a 100,000-token prefix, adds 1,500 new tokens, and returns 600 output tokens. Uncached, the prefix is billed at full input price20 times. Cached, it is written once at the provider's write price and read 19 times at the cached rate.
| Model | Provider | Write / 1M | Read / 1M | Uncached | Cached | Saved | Cut | Break-even |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek | $0.19* | $0.003 | $0.312 | $0.036 | $0.276 | 88% | 1 read |
| Claude Fable 5.1 | Anthropic | $12.50 | $0.250 | $20.90 | $2.63 | $18.27 | 87% | 1 read |
| GPT-6.1 Sol | OpenAI | $2.50 | $0.100 | $4.18 | $0.620 | $3.56 | 85% | 1 read |
| Claude Opus 5.5 | Anthropic | $5.00 | $0.200 | $8.36 | $1.24 | $7.12 | 85% | 1 read |
| GPT-6 Astra | OpenAI | $12.50 | $1.000 | $20.90 | $4.05 | $16.85 | 81% | 1 read |
| Gemini 3.8 Flash | $0.94* | $0.075 | $1.57 | $0.304 | $1.26 | 81% | 1 read | |
| Claude Sonnet 5.5 | Anthropic | $2.50 | $0.200 | $4.18 | $0.810 | $3.37 | 81% | 1 read |
| Qwen3.8 Max | Qwen | $2.50 | $0.250 | $4.13 | $0.857 | $3.27 | 79% | 1 read |
Rates are USD per 1,000,000 tokens from the maintained catalogue, using each provider's published short-lifetime write price. A star marks a model where no write price is published, where the table assumes 1.25x the input rate. A longer cache lifetime costs more to write and only pays off across many more reads.
The cached rate, not the headline rate, decides an agent bill
On this workload caching removes 88% of the bill on DeepSeek V4.1 Flash and 79% on Qwen3.8 Max. The spread comes from the discount each provider offers: DeepSeek V4.1 Flash reads cached tokens at 98% below its fresh rate, Qwen3.8 Max at 88%.
This is why two models with an identical headline price can produce very different invoices. A comparison that comes down to input and output rates alone will get an agent workload wrong, because the tokens that repeat are the ones that dominate.
The effect grows with prefix size and turn count. A short prompt asked once sees none of it.
Providers implement it differently
Anthropic charges to write: 1.25 times the base input rate for a five-minute lifetime, twice for an hour. Reads used to be a flat tenth of the input rate, but they are now set per model: a tenth on most, a twentieth on Claude Opus 5.5, and a fortieth on Claude Fable 5.1. The minimum cacheable prompt also varies by model, from 512 tokens on the newest models to 4,096 on Claude Haiku 4.5.
OpenAI used to cache for free. On GPT-5.6 and later a cache write now costs 1.25 times the input rate, reads cost a tenth, and GPT-6.1 Sol reads cost a twentieth. The cache lifetime on these models is a single 30-minute setting, where older models offered a short in-memory cache or 24 hours.
Google offers implicit caching at a smaller discount, and explicit caching that adds a storage fee charged per hour. That storage fee is unique among the three and makes explicit caching worth it mainly for large, long-lived content such as a code repository or a long video.
Anthropic's multipliers also stack with other modifiers: the batch discount halves them, and pinning inference to the United States multiplies every token category by 1.1.
Put the static part first
Caching matches on a prefix. Everything up to the first difference can be reused; everything after it cannot. So the order of a prompt decides whether caching works at all.
Static content at the top: system instructions, tool definitions, few-shot examples, retrieved context that does not change within the session. Variable content at the bottom: the user turn, timestamps, request identifiers.
The common failure is a single moving token near the start. A current date injected into the system prompt, or a user id in the first line, invalidates the cache on every request and produces the uncached column of the table above while looking like the cached one in the code.
Where caching does not help
One-shot requests over content that differs every time see no benefit, and the write premium makes them marginally more expensive if you cache anyway.
Traffic spread thinly across many different documents behaves the same way: each request writes a cache entry that expires before anything reads it. Caching rewards concentration, not volume.
If your workload is large single documents rather than repeated prefixes, the lever is architectural instead. Compare long context against retrieval →
Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
How much does prompt caching actually save?+
It depends on the cached-input discount and on how often the prefix is reused. In the twenty-turn agent loop priced on this page the saving ranges from about 79 to 88 percent of the total bill, depending on the model.
How many times must a prompt be reused before caching pays off?+
Usually once or twice. A cache write costs roughly 1.25 times the base input rate for a short lifetime, while a read costs a fraction of it, so the premium is repaid almost immediately and every later read is nearly free by comparison.
Do all providers charge for writing to the cache?+
Most now do. Anthropic charges 1.25x the input rate for a five-minute lifetime and 2x for one hour. OpenAI charges 1.25x on GPT-5.6 and later, where earlier models were free. Google offers implicit caching at a smaller discount, and explicit caching that adds a storage fee per hour.
How do I structure a prompt so it caches well?+
Put everything static at the top and everything variable at the bottom. Caching matches on a prefix, so a single changing token near the start, such as a timestamp or a user id, invalidates the whole cache for that request.
Does caching change the answer the model gives?+
No. It changes how the input is billed and how quickly the request is served. The tokens the model sees are identical.