GPT-6 Luna vs GPT-5.6 Luna pricing.
OpenAI cut the price of its cheapest model by half on input and well over half on output. At high volume that is the whole decision.
- Input gap
- 2×
- Output gap
- 2.4×
- Cached gap
- 2×
- Rates verified
- 2026-10-02
Current-model note: Both models are current. GPT-6 Luna was released on 22 September 2026; GPT-5.6 Luna remains available.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| GPT-6 LunaOfficial pricing ↗ | $0.10 | $0.01 | $0.125 | $0.50 | 1,050,000 |
| GPT-5.6 LunaOfficial pricing ↗ | $0.20 | $0.02 | $0.25 | $1.20 | 1,050,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
GPT-6 Luna: exact local BPE tokenization
GPT-5.6 Luna: exact local BPE tokenization
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
GPT-6 Luna: documented formula
GPT-5.6 Luna: documented formula
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for GPT-6 Luna and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | GPT-6 Luna | GPT-5.6 Luna | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0003 | $0.0008 | 2.29x cheaper on GPT-6 Luna |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0016 | $0.0034 | 2.10x cheaper on GPT-6 Luna |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.0070 | $0.014 | 2.06x cheaper on GPT-6 Luna |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.063 (tier) | $0.127 (tier) | 2.02x cheaper on GPT-6 Luna |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.183 (tier) | $0.367 (tier) | 2.01x cheaper on GPT-6 Luna |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 2.0x: GPT-6 Luna reads cached tokens at $0.010 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| GPT-6 Luna | $0.100 | $0.010 | 90% | 1 read |
| GPT-5.6 Luna | $0.200 | $0.020 | 90% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Half the input price, 58 percent off output
GPT-6 Luna lists $0.10 per million input tokens, $0.01 cached and $0.50 output. GPT-5.6 Luna lists $0.20, $0.02 and $1.20. Input is exactly half, cached input is exactly half, and output is 58 percent lower.
At ten million requests a month with a 1,500-token prompt and a 100-token answer, that is the difference between about $3,000 and $1,350 a month, before any caching.
Both carry a long-context rate above 272,000 input tokens, where GPT-6 Luna charges $0.20 and $0.75 against $0.40 and $1.80.
A cheap model is not a small frontier model
This tier suits classification, routing, tagging, extraction from known formats and short summaries, and the first pass of a two-stage pipeline where only hard cases escalate.
The failure mode at this tier is not the rate, it is volume. A pipeline that quietly attaches a large system prompt to millions of calls spends more on the prefix than on the work. Cached input at $0.01 per million is the fix, and it needs the static part of the prompt to come first.
Where the floor is now
OpenAI is no longer alone at this price. DeepSeek V4.1 Flash lists $0.15 and $0.60 off-peak, Qwen3.8 Flash about $0.15 and $0.47, and Qwen3.7 Flash $0.03 and $0.13. GPT-6 Luna is competitive with all of them while keeping exact local token counting, which none of the others offer.
Compare the alternatives directly: <a href="/deepseek-v4-1-flash-vs-gpt-5-6-luna-pricing/">DeepSeek V4.1 Flash against GPT-5.6 Luna</a>, or <a href="/deepseek-v4-1-flash-vs-qwen3-8-flash-pricing/">DeepSeek against Qwen</a>.
Model the workload, not the marketing price.
For new high-volume work, GPT-6 Luna is cheaper on every rate. Re-run your monthly figure before assuming a workload is unaffordable: the cheap tier is now less than half what it was.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
How much cheaper is GPT-6 Luna than GPT-5.6 Luna?+
Half the price on input ($0.10 against $0.20 per million tokens) and 58 percent lower on output ($0.50 against $1.20). Cached input is also half, at $0.01 per million.
Is GPT-5.6 Luna being retired?+
It remains on the pricing page. OpenAI has announced shutdowns for older models such as gpt-5.1 and gpt-5.4-nano on 1 April 2027, with GPT-6 Sol and GPT-6 Luna named as replacements.
What is the cheapest model for high volume?+
It depends on timing and tokenizer. GPT-6 Luna lists $0.10 and $0.50 per million. Qwen3.7 Flash lists $0.03 and $0.13, and DeepSeek V4.1 Flash $0.15 and $0.60 off-peak, doubling at peak hours. Price your real monthly volume rather than comparing headline rates.