DeepSeek V4.1 Flash vs GPT-5.6 Luna pricing.
At the cheap end of the market, small rate differences decide budgets because the request count is enormous. DeepSeek adds a second variable: the time of day.
- Input gap
- 1.3×
- Output gap
- 2×
- Cached gap
- 6.7×
- Rates verified
- 2026-10-02
Current-model note: Both are current high-volume models. DeepSeek V4.1 Flash replaced V4 Flash on 10 September 2026, and this page replaces the V4 Flash comparison. DeepSeek rates shown are off-peak.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| DeepSeek V4.1 FlashOfficial pricing ↗ | $0.15 | $0.003 | — | $0.60 | 1,000,000 |
| GPT-5.6 LunaOfficial pricing ↗ | $0.20 | $0.02 | $0.25 | $1.20 | 1,050,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Rates shown are DeepSeek's off-peak prices. From 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday (except Chinese public holidays), DeepSeek charges double.
DeepSeek V4.1 Flash: Provider-Calibrated UTF-8 Projection
GPT-5.6 Luna: exact local BPE tokenization
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
DeepSeek V4.1 Flash: unavailable
GPT-5.6 Luna: documented formula
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for DeepSeek V4.1 Flash and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | DeepSeek V4.1 Flash | GPT-5.6 Luna | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0004 | $0.0008 | 1.78x cheaper on DeepSeek V4.1 Flash |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0023 | $0.0034 | 1.47x cheaper on DeepSeek V4.1 Flash |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.010 | $0.014 | 1.41x cheaper on DeepSeek V4.1 Flash |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.047 | $0.127 (tier) | 2.68x cheaper on DeepSeek V4.1 Flash |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.137 | $0.367 (tier) | 2.67x cheaper on DeepSeek V4.1 Flash |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 6.7x: DeepSeek V4.1 Flash reads cached tokens at $0.0030 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.150 | $0.0030 | 98% | 1 read |
| GPT-5.6 Luna | $0.200 | $0.020 | 90% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Half the output price off-peak, the same at peak
Off-peak, DeepSeek V4.1 Flash lists $0.15 per million input tokens and $0.60 output. GPT-5.6 Luna lists $0.20 and $1.20. Flash is 25 percent cheaper on input and half the price on output.
During DeepSeek's peak hours, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, Flash doubles to $0.30 and $1.20. Output then costs the same on both, and Luna becomes the cheaper model on input.
One request costs a fraction of a cent either way. Ten million requests a month does not, and that is where these differences become the whole line item.
What this tier is actually for
High-volume models are not small frontier models. They suit classification, routing, tagging, extraction from known formats, short summaries, and the first pass of a two-stage pipeline where a stronger model handles only the hard cases.
V4.1 Flash is stronger than that role usually needs: DeepSeek reports it ahead of its own V4 Pro on coding benchmarks. That makes it tempting to use for everything, but it writes long answers. Artificial Analysis recorded it among the more verbose models it tests, and output is the expensive side of the bill.
Cached input and long context
DeepSeek bills cached input at $0.003 per million off-peak, against $0.02 for GPT-5.6 Luna. A pipeline that resends a long system prompt on every call gains the most from that gap.
GPT-5.6 Luna applies a long-context rate above 272,000 input tokens: $0.40 input and $1.80 output for the whole request. DeepSeek bills its one-million-token window at one rate. For very large single requests, that rule matters more than the base price.
Model the workload, not the marketing price.
At this tier the per-request cost is trivial and the monthly total is not. Compare using your real requests per day and the hours they run, not the per-call price.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Which is cheaper for high-volume work?+
Off-peak, DeepSeek V4.1 Flash is cheaper on input, output and cached input. During DeepSeek's peak hours its rates double, output costs the same as GPT-5.6 Luna, and Luna is cheaper on input.
What happened to DeepSeek V4 Flash?+
DeepSeek retired it on 10 September 2026 and now routes the old model name to V4.1 Flash, billed at V4.1 Flash rates. This page replaces the earlier V4 Flash comparison.
When should I use a high-volume model instead of a balanced one?+
For classification, routing, tagging, extraction from known formats, and short summaries. A common pattern is a two-stage pipeline where the cheap model handles everything and only low-confidence cases escalate to a stronger model.