High-volume model pricing

GPT-6 Luna vs GPT-5.6 Luna pricing.

OpenAI cut the price of its cheapest model by half on input and well over half on output. At high volume that is the whole decision.

Input gap
2×
Output gap
2.4×
Cached gap
2×
Rates verified
2026-10-02

Current-model note: Both models are current. GPT-6 Luna was released on 22 September 2026; GPT-5.6 Luna remains available.

Current modelInput / 1MCached / 1MCache write / 1MOutput / 1MContext
GPT-6 LunaOfficial pricing ↗$0.10$0.01$0.125$0.501,050,000
GPT-5.6 LunaOfficial pricing ↗$0.20$0.02$0.25$1.201,050,000

Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.

Tokenization engine

GPT-6 Luna: exact local BPE tokenization
GPT-5.6 Luna: exact local BPE tokenization

Documents

PDF, DOCX, code, text, and data files are extracted locally and applied to both models.

Images

GPT-6 Luna: documented formula
GPT-5.6 Luna: documented formula

Privacy

Prompt text, document contents, and image pixels stay in the browser.

Everything stays in this browserLive

Text, code, PDF, DOCX, images, audio, or video · 12 MB per file

Computed from current rates

What five real workloads actually cost

Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for GPT-6 Luna and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.

WorkloadInputOutputGPT-6 LunaGPT-5.6 LunaDifference
Short chat turn
A typical assistant exchange.
1,000500$0.0003$0.00082.29x cheaper on GPT-6 Luna
RAG answer
Five retrieved chunks plus a question.
12,000800$0.0016$0.00342.10x cheaper on GPT-6 Luna
Code review
A medium pull request with surrounding files.
60,0002,000$0.0070$0.0142.06x cheaper on GPT-6 Luna
Whole-document analysis
A long report or contract read in one call.
300,0004,000$0.063 (tier)$0.127 (tier)2.02x cheaper on GPT-6 Luna
Full-context load
Filling most of a one-million-token window.
900,0004,000$0.183 (tier)$0.367 (tier)2.01x cheaper on GPT-6 Luna

"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.

Prompt caching

The cached-input rate is the number most comparisons miss

Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 2.0x: GPT-6 Luna reads cached tokens at $0.010 per million.

ModelFresh input / 1MCached input / 1MDiscountBreak-even
GPT-6 Luna$0.100$0.01090%1 read
GPT-5.6 Luna$0.200$0.02090%1 read

Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.

Half the input price, 58 percent off output

GPT-6 Luna lists $0.10 per million input tokens, $0.01 cached and $0.50 output. GPT-5.6 Luna lists $0.20, $0.02 and $1.20. Input is exactly half, cached input is exactly half, and output is 58 percent lower.

At ten million requests a month with a 1,500-token prompt and a 100-token answer, that is the difference between about $3,000 and $1,350 a month, before any caching.

Both carry a long-context rate above 272,000 input tokens, where GPT-6 Luna charges $0.20 and $0.75 against $0.40 and $1.80.

A cheap model is not a small frontier model

This tier suits classification, routing, tagging, extraction from known formats and short summaries, and the first pass of a two-stage pipeline where only hard cases escalate.

The failure mode at this tier is not the rate, it is volume. A pipeline that quietly attaches a large system prompt to millions of calls spends more on the prefix than on the work. Cached input at $0.01 per million is the fix, and it needs the static part of the prompt to come first.

Where the floor is now

OpenAI is no longer alone at this price. DeepSeek V4.1 Flash lists $0.15 and $0.60 off-peak, Qwen3.8 Flash about $0.15 and $0.47, and Qwen3.7 Flash $0.03 and $0.13. GPT-6 Luna is competitive with all of them while keeping exact local token counting, which none of the others offer.

Compare the alternatives directly: <a href="/deepseek-v4-1-flash-vs-gpt-5-6-luna-pricing/">DeepSeek V4.1 Flash against GPT-5.6 Luna</a>, or <a href="/deepseek-v4-1-flash-vs-qwen3-8-flash-pricing/">DeepSeek against Qwen</a>.

Decision rule

Model the workload, not the marketing price.

For new high-volume work, GPT-6 Luna is cheaper on every rate. Re-run your monthly figure before assuming a workload is unaffordable: the cheap tier is now less than half what it was.

Review calculation methodology →
Focused counters

Measure against the model you use

Current model comparisons

Other comparisons worth running

Cost guides

Compare approaches, not just models

Plain answers

Frequently asked questions

How much cheaper is GPT-6 Luna than GPT-5.6 Luna?+

Half the price on input ($0.10 against $0.20 per million tokens) and 58 percent lower on output ($0.50 against $1.20). Cached input is also half, at $0.01 per million.

Is GPT-5.6 Luna being retired?+

It remains on the pricing page. OpenAI has announced shutdowns for older models such as gpt-5.1 and gpt-5.4-nano on 1 April 2027, with GPT-6 Sol and GPT-6 Luna named as replacements.

What is the cheapest model for high volume?+

It depends on timing and tokenizer. GPT-6 Luna lists $0.10 and $0.50 per million. Qwen3.7 Flash lists $0.03 and $0.13, and DeepSeek V4.1 Flash $0.15 and $0.60 off-peak, doubling at peak hours. Price your real monthly volume rather than comparing headline rates.