DeepSeek V4.1 Flash vs V4 Pro cost.
DeepSeek's new Flash model is cheaper than its Pro model and beats it on most coding benchmarks. It does not beat it everywhere.
- Input gap
- 4.4×
- Output gap
- 3.3×
- Cached gap
- 7.3×
- Rates verified
- 2026-10-02
Current-model note: Both models are current. DeepSeek kept V4 Pro on its API after 14 September 2026 in response to user demand. Rates shown are off-peak, from DeepSeek's official pricing page.
| Current model | Input / 1M | Cached / 1M | Output / 1M | Context |
|---|---|---|---|---|
| DeepSeek V4.1 FlashOfficial pricing ↗ | $0.15 | $0.003 | $0.60 | 1,000,000 |
| DeepSeek V4 ProOfficial pricing ↗ | $0.66 | $0.022 | $1.98 | 1,000,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Rates shown are DeepSeek's off-peak prices. From 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday (except Chinese public holidays), DeepSeek charges double.
DeepSeek V4 Pro: the public catalogue lists an outdated rate, so this rate comes from the official pricing page.
DeepSeek V4.1 Flash: Provider-Calibrated UTF-8 Projection
DeepSeek V4 Pro: Provider-Calibrated UTF-8 Projection
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
DeepSeek V4.1 Flash: unavailable
DeepSeek V4 Pro: unavailable
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for DeepSeek V4.1 Flash and DeepSeek V4 Pro, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | DeepSeek V4.1 Flash | DeepSeek V4 Pro | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0004 | $0.0016 | 3.67x cheaper on DeepSeek V4.1 Flash |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0023 | $0.0095 | 4.17x cheaper on DeepSeek V4.1 Flash |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.010 | $0.044 | 4.27x cheaper on DeepSeek V4.1 Flash |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.047 | $0.206 | 4.34x cheaper on DeepSeek V4.1 Flash |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.137 | $0.602 | 4.38x cheaper on DeepSeek V4.1 Flash |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 7.3x: DeepSeek V4.1 Flash reads cached tokens at $0.0030 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.150 | $0.0030 | 98% | 1 read |
| DeepSeek V4 Pro | $0.660 | $0.022 | 97% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Flash is now both the cheaper model and the stronger coder
At off-peak rates DeepSeek V4.1 Flash lists $0.15 per million input tokens, $0.003 cached and $0.60 output. DeepSeek V4 Pro lists $0.66, $0.022 and $1.98. Flash is about four times cheaper on input, three times cheaper on output, and seven times cheaper on cached input.
DeepSeek's model card, run at maximum reasoning effort, puts V4.1 Flash ahead of V4 Pro on DeepSWE (74.2 against 62.7), Terminal-Bench 2.1 (90.6 against 87.9) and Codeforces (3471 against 3348). On DeepSWE it is level with Claude Opus 5, which scored 74.0 in the same table.
Where V4 Pro still leads
The same table shows V4 Pro ahead on knowledge-heavy tests: Humanity's Last Exam (42.7 against 36.8) and GPQA Diamond (92.4 against 90.9). If your workload is expert question answering rather than code, the older and more expensive model can still be the better choice.
DeepSeek first planned to route every V4 Pro request to V4.1 Flash from 14 September. It then decided to keep serving V4 Pro in response to user demand. DeepSeek has said a V4.1 Pro is coming, without giving a date.
Peak hours and verbosity change the real bill
DeepSeek bills double from 01:00 to 04:00 and from 06:00 to 10:00 UTC, Monday to Friday. The rates in the table are the off-peak ones. A workload that runs during Chinese business hours pays twice the figure shown, so schedule batch jobs outside those windows where you can.
V4.1 Flash is also verbose. Artificial Analysis recorded 250 million output tokens for it to run its Intelligence Index, against a median of 140 million for comparable models. A cheaper rate multiplied by more tokens is still cheaper here, but by less than the headline ratio suggests. Set a realistic output length in the calculator before comparing.
Model the workload, not the marketing price.
For coding and agent work, start with V4.1 Flash: it is cheaper and scores higher on DeepSeek's own coding benchmarks. Keep V4 Pro for knowledge-heavy questions, where it still scores higher, and measure output length, because V4.1 Flash tends to write more.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Is DeepSeek V4.1 Flash better than V4 Pro?+
On coding and agent benchmarks, yes: DeepSeek's own table puts V4.1 Flash ahead on DeepSWE, Terminal-Bench 2.1 and Codeforces. V4 Pro still scores higher on Humanity's Last Exam and GPQA Diamond, which test expert knowledge.
How much cheaper is DeepSeek V4.1 Flash than V4 Pro?+
At off-peak rates, about four times cheaper on input ($0.15 against $0.66 per million tokens), three times cheaper on output ($0.60 against $1.98), and seven times cheaper on cached input.
Is DeepSeek V4 Pro being retired?+
No. DeepSeek first planned to route V4 Pro requests to V4.1 Flash from 14 September 2026, then decided to keep serving V4 Pro in response to user demand. The old name deepseek-v4-flash, by contrast, is retired and now routes to V4.1 Flash.
Why is the real cost higher than the table at some times of day?+
DeepSeek charges double during peak hours: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, except Chinese public holidays. This page shows off-peak rates.