Gemini 3.5 Flash-Lite vs GPT-5.6 Luna cost.
Two cheap models that both accept images, where a text-only price comparison will mislead you.
- Input gap
- 1.5×
- Output gap
- 2.1×
- Cached gap
- 1.5×
- Rates verified
- 2026-10-02
Current-model note: Both models are current. Google names Gemini 3.5 Flash-Lite as the replacement for Gemini 3.1 Flash-Lite, which shuts down on 7 May 2027. This page replaces the earlier 3.1 comparison.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| Gemini 3.5 Flash LiteOfficial pricing ↗ | $0.30 | $0.03 | — | $2.50 | 1,048,576 |
| GPT-5.6 LunaOfficial pricing ↗ | $0.20 | $0.02 | $0.25 | $1.20 | 1,050,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Gemini 3.5 Flash Lite: Provider-Calibrated UTF-8 Projection
GPT-5.6 Luna: exact local BPE tokenization
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
Gemini 3.5 Flash Lite: documented formula
GPT-5.6 Luna: documented formula
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for Gemini 3.5 Flash Lite and GPT-5.6 Luna, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | Gemini 3.5 Flash Lite | GPT-5.6 Luna | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0015 | $0.0008 | 1.94x cheaper on GPT-5.6 Luna |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.0056 | $0.0034 | 1.67x cheaper on GPT-5.6 Luna |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.023 | $0.014 | 1.60x cheaper on GPT-5.6 Luna |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.100 | $0.127 (tier) | 1.27x cheaper on Gemini 3.5 Flash Lite |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $0.280 | $0.367 (tier) | 1.31x cheaper on Gemini 3.5 Flash Lite |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production. Here the gap is 1.5x: GPT-5.6 Luna reads cached tokens at $0.020 per million.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| Gemini 3.5 Flash Lite | $0.300 | $0.030 | 90% | 1 read |
| GPT-5.6 Luna | $0.200 | $0.020 | 90% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
Images are priced as tokens, and the formulas differ
Neither provider bills images by the file. Both convert dimensions into a token count and charge the normal input rate on it, but they use different rules to do so.
OpenAI splits a resized image into tiles and charges a base allowance plus a per-tile cost, with a fixed low-detail option that ignores size entirely. Google allocates tokens by media size against its own published guidance. The same photograph therefore produces different token counts on the two models.
This is why a text-only price comparison is unreliable for a vision pipeline. Upload a representative image in the calculator below and compare the totals rather than the rates.
The low-detail option is the biggest lever on OpenAI
OpenAI's low-detail mode charges a flat allowance regardless of image dimensions. For thumbnails, screenshots where only layout matters, or a first-pass filter, it can cut image cost dramatically.
It is a quality trade, not a free saving: fine text and small features stop being legible to the model. Test it on your actual images before adopting it across a pipeline.
Text rates favour Luna; image volume decides
On text, GPT-5.6 Luna is the cheaper model. Gemini 3.5 Flash-Lite lists $0.30 per million input and $2.50 output; GPT-5.6 Luna lists $0.20 and $1.20, less than half on output.
Gemini 3.5 Flash-Lite also costs more than the 3.1 model it replaces, which listed $0.25 and $1.50. If you are migrating a Gemini pipeline, re-run the numbers rather than assuming the new generation is cheaper.
Once images enter the mix, the token counts diverge much more than the rates do. For an image-heavy workload, measure token counts on real files first and treat the published rates as a secondary factor.
Model the workload, not the marketing price.
For multimodal pipelines, include representative image sizes as well as text. Vision-token rules differ enough between providers that text-only tables hide the real request cost.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
How are image tokens calculated?+
Both providers convert image dimensions into a token count and bill it at the normal input rate, using different published formulas. OpenAI uses a tile-based method with a fixed low-detail option; Google allocates tokens by media size. The same image produces different counts on each.
Which is cheaper for a vision pipeline?+
It depends on image sizes as well as rates. On text, GPT-5.6 Luna is cheaper: $0.20 against $0.30 per million input, and $1.20 against $2.50 output. A given image produces different token counts on each, so upload a representative image to compare real totals.
Why not compare Gemini 3.1 Flash-Lite?+
Google has named Gemini 3.5 Flash-Lite as its replacement and will shut 3.1 Flash-Lite down on 7 May 2027. New projects should price the model they will still be able to use.
Does the calculator upload my images?+
No. Only the width and height are read locally in your browser and passed through each provider's published formula. The pixels never leave your device.