DeepSeek vs OpenAI reasoning model cost.
R1 and o1 are retired names. This page compares the reasoning models that replaced them, where output volume dominates the bill.
- Input gap
- 67×
- Output gap
- 83×
- Cached gap
- 333×
- Rates verified
- 2026-10-02
Current-model note: DeepSeek R1 and OpenAI o1 are legacy search terms. The current pair is DeepSeek V4.1 Flash, DeepSeek's newest reasoning model, and GPT-6 Astra, OpenAI's flagship. DeepSeek rates shown are off-peak.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| DeepSeek V4.1 FlashOfficial pricing ↗ | $0.15 | $0.003 | — | $0.60 | 1,000,000 |
| GPT-6 AstraOfficial pricing ↗ | $10.00 | $1.00 | $12.50 | $50.00 | 1,050,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Rates shown are DeepSeek's off-peak prices. From 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday (except Chinese public holidays), DeepSeek charges double.
DeepSeek V4.1 Flash: Provider-Calibrated UTF-8 Projection
GPT-6 Astra: exact local BPE tokenization
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
DeepSeek V4.1 Flash: unavailable
GPT-6 Astra: documented formula
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for DeepSeek V4.1 Flash and GPT-6 Astra, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | DeepSeek V4.1 Flash | GPT-6 Astra | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0004 | $0.035 | 77.78x cheaper on DeepSeek V4.1 Flash |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0023 | $0.160 | 70.18x cheaper on DeepSeek V4.1 Flash |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.010 | $0.700 | 68.63x cheaper on DeepSeek V4.1 Flash |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.047 | $6.30 (tier) | 132.91x cheaper on DeepSeek V4.1 Flash |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.137 | $18.30 (tier) | 133.19x cheaper on DeepSeek V4.1 Flash |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 333.3x: DeepSeek V4.1 Flash reads cached tokens at $0.0030 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.150 | $0.0030 | 98% | 1 read |
| GPT-6 Astra | $10.00 | $1.00 | 90% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Reasoning models bill differently from chat models
A reasoning model spends tokens thinking before it answers. Those tokens are billed as output, and output costs several times what input costs on every model in this comparison.
This inverts the usual budgeting instinct. For a chat assistant, prompt size drives the bill. For a reasoning model, the length of the reasoning does, and it varies with problem difficulty in a way that a fixed estimate will not capture.
Budget with a distribution, not an average. A task whose hard cases reason ten times longer than its easy ones will overrun a budget built on the mean.
The widest gap in this market
Off-peak, DeepSeek V4.1 Flash lists $0.15 per million input tokens and $0.60 output. GPT-6 Astra lists $10 and $50. That is about 67 times on input and 83 times on output. Even at DeepSeek's peak rates the gap is more than 40 times on output.
A gap that wide changes the question. It is no longer which model is cheaper, but whether the cheaper one is good enough, and on which tasks. DeepSeek's own table puts V4.1 Flash level with Claude Opus 5 on the DeepSWE coding benchmark, while independent indexes still place GPT-6 Astra near the top overall and V4.1 Flash well below it.
What the old names meant
DeepSeek R1 and OpenAI o1 were the first widely used reasoning models, and they established the pattern of billing visible thinking as output tokens. Both have been superseded, and neither appears in current published pricing.
Cost ratios from that generation do not transfer. Rates fell substantially and the models reason differently, so the token counts changed as well as the prices.
Model the workload, not the marketing price.
Reasoning workloads are output-heavy by construction. Set a realistic output-token budget before comparing daily or monthly spend, because output is where the difference lands.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Are DeepSeek R1 and OpenAI o1 still available?+
Neither appears in current published pricing. The closest current models are DeepSeek V4.1 Flash and GPT-6 Astra, which is the pair priced on this page.
Why do reasoning models cost more than expected?+
They emit reasoning tokens before answering, and those bill as output, which costs several times more than input. Reasoning length also varies with problem difficulty, so a budget built on an average output length will overrun on hard cases.
How wide is the price gap?+
Off-peak, DeepSeek V4.1 Flash is about 83 times cheaper than GPT-6 Astra on output, the side that dominates reasoning workloads. DeepSeek doubles its rates during peak hours, which still leaves a gap of more than 40 times.