High-volume multimodal

Gemini 3.5 Flash-Lite vs GPT-5.6 Luna cost.

Two cheap models that both accept images, where a text-only price comparison will mislead you.

Input gap
1.5×
Output gap
2.1×
Cached gap
1.5×
Rates verified
2026-10-02

Current-model note: Both models are current. Google names Gemini 3.5 Flash-Lite as the replacement for Gemini 3.1 Flash-Lite, which shuts down on 7 May 2027. This page replaces the earlier 3.1 comparison.

Current modelInput / 1MCached / 1MCache write / 1MOutput / 1MContext
Gemini 3.5 Flash LiteOfficial pricing ↗$0.30$0.03—$2.501,048,576
GPT-5.6 LunaOfficial pricing ↗$0.20$0.02$0.25$1.201,050,000

Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.

Tokenization engine

Gemini 3.5 Flash Lite: Provider-Calibrated UTF-8 Projection
GPT-5.6 Luna: exact local BPE tokenization

Documents

PDF, DOCX, code, text, and data files are extracted locally and applied to both models.

Images

Gemini 3.5 Flash Lite: documented formula
GPT-5.6 Luna: documented formula

Privacy

Prompt text, document contents, and image pixels stay in the browser.

Everything stays in this browserLive

Text, code, PDF, DOCX, images, audio, or video · 12 MB per file

Computed from current rates

What five real workloads actually cost

Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for Gemini 3.5 Flash Lite and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.

WorkloadInputOutputGemini 3.5 Flash LiteGPT-5.6 LunaDifference
Short chat turn
A typical assistant exchange.
1,000500$0.0015$0.00081.94x cheaper on GPT-5.6 Luna
RAG answer
Five retrieved chunks plus a question.
12,000800$0.0056$0.00341.67x cheaper on GPT-5.6 Luna
Code review
A medium pull request with surrounding files.
60,0002,000$0.023$0.0141.60x cheaper on GPT-5.6 Luna
Whole-document analysis
A long report or contract read in one call.
300,0004,000$0.100$0.127 (tier)1.27x cheaper on Gemini 3.5 Flash Lite
Full-context load
Filling most of a one-million-token window.
900,0004,000$0.280$0.367 (tier)1.31x cheaper on Gemini 3.5 Flash Lite

"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.

Prompt caching

The cached-input rate is the number most comparisons miss

Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 1.5x: GPT-5.6 Luna reads cached tokens at $0.020 per million.

ModelFresh input / 1MCached input / 1MDiscountBreak-even
Gemini 3.5 Flash Lite$0.300$0.03090%1 read
GPT-5.6 Luna$0.200$0.02090%1 read

Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.

Images are priced as tokens, and the formulas differ

Neither provider bills images by the file. Both convert dimensions into a token count and charge the normal input rate on it, but they use different rules to do so.

OpenAI splits a resized image into tiles and charges a base allowance plus a per-tile cost, with a fixed low-detail option that ignores size entirely. Google allocates tokens by media size against its own published guidance. The same photograph therefore produces different token counts on the two models.

This is why a text-only price comparison is unreliable for a vision pipeline. Upload a representative image in the calculator below and compare the totals rather than the rates.

The low-detail option is the biggest lever on OpenAI

OpenAI's low-detail mode charges a flat allowance regardless of image dimensions. For thumbnails, screenshots where only layout matters, or a first-pass filter, it can cut image cost dramatically.

It is a quality trade, not a free saving: fine text and small features stop being legible to the model. Test it on your actual images before adopting it across a pipeline.

Text rates favour Luna; image volume decides

On text, GPT-5.6 Luna is the cheaper model. Gemini 3.5 Flash-Lite lists $0.30 per million input and $2.50 output; GPT-5.6 Luna lists $0.20 and $1.20, less than half on output.

Gemini 3.5 Flash-Lite also costs more than the 3.1 model it replaces, which listed $0.25 and $1.50. If you are migrating a Gemini pipeline, re-run the numbers rather than assuming the new generation is cheaper.

Once images enter the mix, the token counts diverge much more than the rates do. For an image-heavy workload, measure token counts on real files first and treat the published rates as a secondary factor.

Decision rule

Model the workload, not the marketing price.

For multimodal pipelines, include representative image sizes as well as text. Vision-token rules differ enough between providers that text-only tables hide the real request cost.

Review calculation methodology →
Focused counters

Measure against the model you use

Current model comparisons

Other comparisons worth running

Cost guides

Compare approaches, not just models

Plain answers

Frequently asked questions

How are image tokens calculated?+

Both providers convert image dimensions into a token count and bill it at the normal input rate, using different published formulas. OpenAI uses a tile-based method with a fixed low-detail option; Google allocates tokens by media size. The same image produces different counts on each.

Which is cheaper for a vision pipeline?+

It depends on image sizes as well as rates. On text, GPT-5.6 Luna is cheaper: $0.20 against $0.30 per million input, and $1.20 against $2.50 output. A given image produces different token counts on each, so upload a representative image to compare real totals.

Why not compare Gemini 3.1 Flash-Lite?+

Google has named Gemini 3.5 Flash-Lite as its replacement and will shut 3.1 Flash-Lite down on 7 May 2027. New projects should price the model they will still be able to use.

Does the calculator upload my images?+

No. Only the width and height are read locally in your browser and passed through each provider's published formula. The pixels never leave your device.