High-volume model pricing

Gemini Flash vs GPT mini pricing.

GPT-4o mini is retired. This page compares today's Gemini Flash with the OpenAI model that took over the cheap tier.

Input gap
3.8×
Output gap
3.1×
Cached gap
3.8×
Rates verified
2026-10-02

Current-model note: GPT-4o mini is a legacy search term and no longer appears in current pricing. The live pair is Gemini 3.8 Flash, Google's current Flash model, and GPT-5.6 Luna, OpenAI's high-volume model.

Current modelInput / 1MCached / 1MCache write / 1MOutput / 1MContext
Gemini 3.8 FlashOfficial pricing ↗$0.75$0.075—$3.751,048,576
GPT-5.6 LunaOfficial pricing ↗$0.20$0.02$0.25$1.201,050,000

Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.

Tokenization engine

Gemini 3.8 Flash: Provider-Calibrated UTF-8 Projection
GPT-5.6 Luna: exact local BPE tokenization

Documents

PDF, DOCX, code, text, and data files are extracted locally and applied to both models.

Images

Gemini 3.8 Flash: documented formula
GPT-5.6 Luna: documented formula

Privacy

Prompt text, document contents, and image pixels stay in the browser.

Everything stays in this browserLive

Text, code, PDF, DOCX, images, audio, or video · 12 MB per file

Computed from current rates

What five real workloads actually cost

Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for Gemini 3.8 Flash and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.

WorkloadInputOutputGemini 3.8 FlashGPT-5.6 LunaDifference
Short chat turn
A typical assistant exchange.
1,000500$0.0026$0.00083.28x cheaper on GPT-5.6 Luna
RAG answer
Five retrieved chunks plus a question.
12,000800$0.012$0.00343.57x cheaper on GPT-5.6 Luna
Code review
A medium pull request with surrounding files.
60,0002,000$0.052$0.0143.65x cheaper on GPT-5.6 Luna
Whole-document analysis
A long report or contract read in one call.
300,0004,000$0.240$0.127 (tier)1.89x cheaper on GPT-5.6 Luna
Full-context load
Filling most of a one-million-token window.
900,0004,000$0.690$0.367 (tier)1.88x cheaper on GPT-5.6 Luna

"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.

Prompt caching

The cached-input rate is the number most comparisons miss

Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 3.8x: GPT-5.6 Luna reads cached tokens at $0.020 per million.

ModelFresh input / 1MCached input / 1MDiscountBreak-even
Gemini 3.8 Flash$0.750$0.07590%1 read
GPT-5.6 Luna$0.200$0.02090%1 read

Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.

The cheap tier got much cheaper

GPT-4o mini set the expectation for what a cheap model costs. The current tier is well below it: GPT-5.6 Luna lists $0.20 per million input and $1.20 output, following an OpenAI price reduction in mid-2026.

If your cost model still assumes GPT-4o mini rates, it is overstating the bill. Re-run it before deciding you cannot afford a workload.

Gemini Flash is no longer the cheap option

Gemini 3.8 Flash lists $0.75 per million input tokens and $3.75 output, almost four times GPT-5.6 Luna on input and about three times on output. Google now sells Flash as a capable mid-tier model; its low-cost line is Flash-Lite.

Google has also announced that Gemini 3.8 Flash rises to $1.50 input and $7.50 output on 1 January 2027. A budget built on today's rate will be half the real cost from January. For a workload that will still run next year, use the 2027 rate.

Per-call price is the wrong unit here

At these rates a single request costs a fraction of a cent, which makes per-call comparison useless. The meaningful figure is monthly total at your real request volume.

The two variables that actually move it are output length and any prefix you resend on every call. A pipeline that quietly attaches a large system prompt to millions of requests will spend more on the prefix than on the work.

Decision rule

Model the workload, not the marketing price.

Gemini Flash is no longer priced as a budget model. Compare it with GPT-5.6 Luna at your real request volume, and use Google's 2027 price if the workload will still run next year.

Review calculation methodology →
Focused counters

Measure against the model you use

Current model comparisons

Other comparisons worth running

Cost guides

Compare approaches, not just models

Plain answers

Frequently asked questions

Is GPT-4o mini still available?+

It no longer appears in current published pricing. The high-volume slot is now held by GPT-5.6 Luna at $0.20 per million input and $1.20 output, which is below the rate GPT-4o mini set.

Which is cheaper for high-volume work?+

GPT-5.6 Luna, at current published rates: $0.20 per million input and $1.20 output, against $0.75 and $3.75 for Gemini 3.8 Flash. The gap widens on 1 January 2027, when Gemini 3.8 Flash rises to $1.50 and $7.50.

Is Gemini 3.8 Flash getting more expensive?+

Yes. Google's pricing page lists $0.75 input and $3.75 output through 31 December 2026, and $1.50 input and $7.50 output from 1 January 2027.

How should I compare cheap models?+

By monthly total at your real request volume, not per-call price. Output length and any prefix resent on every call are the two variables that move the number most.