DeepSeek V4.1 Flash vs Qwen3.8 Flash pricing.
Two of the cheapest capable models on the market list the same input price. The bill still differs, and the time of day decides by how much.
- Input gap
- same
- Output gap
- 1.3×
- Cached gap
- 5.3×
- Rates verified
- 2026-10-02
Current-model note: Both models are current. DeepSeek rates shown are off-peak, from DeepSeek's official pricing page. Qwen3.8 Flash is served by Alibaba Cloud.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| DeepSeek V4.1 FlashOfficial pricing ↗ | $0.15 | $0.003 | — | $0.60 | 1,000,000 |
| Qwen3.8 FlashOfficial pricing ↗ | $0.15 | $0.016 | $0.20 | $0.47 | 1,000,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Rates shown are DeepSeek's off-peak prices. From 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday (except Chinese public holidays), DeepSeek charges double.
DeepSeek V4.1 Flash: Provider-Calibrated UTF-8 Projection
Qwen3.8 Flash: Provider-Calibrated UTF-8 Projection
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
DeepSeek V4.1 Flash: unavailable
Qwen3.8 Flash: unavailable
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for DeepSeek V4.1 Flash and Qwen3.8 Flash, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | DeepSeek V4.1 Flash | Qwen3.8 Flash | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0004 | $0.0004 | 1.17x cheaper on Qwen3.8 Flash |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0023 | $0.0022 | 1.05x cheaper on Qwen3.8 Flash |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.010 | $0.0099 | 1.03x cheaper on Qwen3.8 Flash |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.047 | $0.047 | 1.01x cheaper on Qwen3.8 Flash |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.137 | $0.137 | about equal |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 5.3x: DeepSeek V4.1 Flash reads cached tokens at $0.0030 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.150 | $0.0030 | 98% | 1 read |
| Qwen3.8 Flash | $0.150 | $0.016 | 89% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Same input price, different everything else
Off-peak, DeepSeek V4.1 Flash lists $0.15 per million input tokens, $0.003 cached and $0.60 output. Qwen3.8 Flash lists about $0.15 input, $0.016 cached and $0.47 output.
So Qwen3.8 Flash is about 20 percent cheaper on output, and DeepSeek V4.1 Flash is about five times cheaper on cached input. Which difference matters depends on the shape of your requests.
Peak hours double one side of the comparison
DeepSeek charges double from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays: $0.30 input and $1.20 output. At those hours Qwen3.8 Flash costs half as much on input and less than half on output.
A pipeline that runs around the clock pays a blend. About 35 of the week's 168 hours are peak, so a steady workload pays roughly 20 percent more than DeepSeek's off-peak rate. A workload concentrated in Asian business hours pays close to double.
Cached prefixes favour DeepSeek
Agent loops and retrieval pipelines resend the same prefix many times. DeepSeek bills cached input at $0.003 per million off-peak, about a fifth of Qwen's $0.016. On a workload dominated by a long cached prefix, that gap outweighs the output difference.
Both models have open weights. DeepSeek publishes V4.1 Flash under the MIT licence, and Alibaba published the model behind Qwen3.8 Flash as Qwen3.8-Flash-Next, with 6 billion active parameters. Either can be self-hosted if per-token pricing stops making sense at your volume.
Model the workload, not the marketing price.
Off-peak and with a large reused prompt, DeepSeek V4.1 Flash is usually cheaper. During DeepSeek's peak hours, or for output-heavy work, Qwen3.8 Flash usually is. Price your real traffic pattern, not the headline rate.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Which is cheaper, DeepSeek V4.1 Flash or Qwen3.8 Flash?+
It depends on timing and request shape. Input prices match off-peak. Qwen3.8 Flash is about 20 percent cheaper on output; DeepSeek V4.1 Flash is about five times cheaper on cached input. During DeepSeek's peak hours, Qwen3.8 Flash is cheaper on input and output.
When are DeepSeek peak hours?+
01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, except Chinese public holidays. DeepSeek charges double in those windows. The rates on this page are off-peak.
Are both models open-weight?+
Yes. DeepSeek publishes V4.1 Flash under the MIT licence, and Alibaba published the weights behind Qwen3.8 Flash as Qwen3.8-Flash-Next. The hosted API prices on this page do not apply if you self-host.