Long context vs RAG cost.
Million-token windows made it possible to skip retrieval entirely. This page prices what that convenience costs per query, using current published rates.
The same answer, two input sizes
One question answered two ways. The long-context call pastes 200,000 tokens of source material into the prompt. The retrieval call sends 4,000 tokens, five retrieved passages plus the question. Both produce 800 output tokens.
| Model | Provider | Long context | Retrieval | Multiple | Saved per query |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $2.04 | $0.080 | 25.5x | $1.96 |
| Claude Fable 5.1 | Anthropic | $2.04 | $0.080 | 25.5x | $1.96 |
| GPT-5.6 Terra | OpenAI | $0.410 | $0.018 | 23.3x | $0.392 |
| Claude Sonnet 5 | Anthropic | $0.408 | $0.016 | 25.5x | $0.392 |
| Gemini 3.8 Flash | $0.153 | $0.0060 | 25.5x | $0.147 | |
| DeepSeek V4.1 Flash | DeepSeek | $0.030 | $0.0011 | 28.2x | $0.029 |
Rates are USD per 1,000,000 tokens from the maintained catalogue. Long-context tiers are applied where a provider charges one. Paste your own material into the calculator to replace these assumptions with measured token counts.
Retrieval sends 25 to 100 times fewer tokens
The mechanism is simple. Retrieval selects two to five relevant passages, roughly two thousand tokens. Stuffing sends the whole document set, commonly fifty thousand to two hundred thousand tokens. You pay for input tokens, so the bill follows that ratio almost exactly.
In the table above the multiple lands between 23x and 28x. It is not identical across models because output cost is a fixed part of each call, and output weighs more heavily where the model charges more for it.
Cheapest per long-context query here is DeepSeek V4.1 Flash at $0.030; the most expensive is Claude Fable 5.1 at $2.04. The gap between models is smaller than the gap between the two architectures.
Why long context still wins for most teams
Per-query cost is not the only cost. Retrieval means chunking, an embedding model, a vector store, a retrieval step to tune, and an evaluation harness to know whether it is picking the right passages. That is real engineering time and ongoing maintenance.
A rough framing: at the most expensive model in the table you save $1.96 per query. Against a nominal $600 developer-day, retrieval pays for that day after about 307 queries. If you serve a few hundred queries a week, stuffing the window is the cheaper decision once your own time is counted.
That crossover moves fast. Ten thousand queries a day changes the answer completely, which is why this is a volume question rather than a taste question.
Latency and recall, not just price
Published comparisons put a retrieval pipeline at roughly one second end to end, against thirty to sixty seconds for the equivalent long-context call. For an interactive product that difference decides the design before cost does.
Answer quality is the part most often assumed rather than measured. Needle-in-a-haystack scores are near perfect and widely quoted, but on realistic multi-fact questions spread across a long document, measured recall has been reported around sixty percent. A model that reads everything does not necessarily use everything.
The honest test is your own question set, scored by hand, on both architectures. Neither the token count nor the benchmark answers it for you.
Caching narrows the gap, when the context repeats
If many queries hit the same large context, the repeated portion can bill at the cached rate rather than the fresh one. On several providers that is roughly a tenth of the input price, which pulls long context much closer to retrieval on cost.
The condition matters. Caching helps when the prefix is genuinely identical and reused inside the cache lifetime. A corpus that differs per user, or a query pattern spread thinly across many documents, will not hit the cache and will pay full price every time.
Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Is a one-million-token context window cheaper than RAG?+
Almost never on a per-query basis. Retrieval typically sends 25 to 100 times fewer input tokens for the same answer, and input tokens are what you pay for. Long context wins on build cost and simplicity, not on unit cost.
When should I stuff the context window instead of building retrieval?+
When the corpus fits the window, query volume is low, and latency is not critical. For internal tools and document question-answering over a modest corpus, the engineering saving usually outweighs the token cost.
When does RAG become necessary?+
Three triggers: the corpus no longer fits the context window, per-request latency breaks your service level, or query volume pushes long-context spend past the cost of building and running retrieval.
Does long context give better answers than retrieval?+
Not reliably. Published needle-in-a-haystack scores are near perfect, but on realistic multi-fact retrieval across a long document, measured recall is far lower, around 60 percent in one widely cited evaluation. Test on your own questions.
Does prompt caching change the comparison?+
Substantially. If the same large context is reused across many queries, cached input bills at roughly a tenth of the fresh rate on several providers, which narrows the gap. It only helps when the context is genuinely repeated.