Claude vs GPT token cost.
GPT-4o is retired from current pricing. This page compares the models that took its place in the balanced tier.
- Input gap
- same
- Output gap
- 1.2×
- Cached gap
- same
- Rates verified
- 2026-10-02
Current-model note: GPT-4o is a legacy search term and no longer appears in current pricing. The live pair below is Claude Sonnet 5 and GPT-5.6 Terra, the current balanced-tier models.
| Current model | Input / 1M | Cached / 1M | Cache write / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| Claude Sonnet 5Official pricing ↗ | $2.00 | $0.20 | $2.50 | $10.00 | 1,000,000 |
| GPT-5.6 TerraOfficial pricing ↗ | $2.00 | $0.20 | $2.50 | $12.00 | 1,050,000 |
Rates verified 2026-10-02. The browser may refresh them from the live catalogue. Provider pricing and tiers remain authoritative.
Claude Sonnet 5: Provider-Calibrated UTF-8 Projection
GPT-5.6 Terra: exact local BPE tokenization
PDF, DOCX, code, text, and data files are extracted locally and applied to both models.
Claude Sonnet 5: documented formula
GPT-5.6 Terra: documented formula
Prompt text, document contents, and image pixels stay in the browser.
Text, code, PDF, DOCX, images, audio, or video · 12 MB per file
Your input, measured across providers.
Choose a provider card to inspect every available model.
Workload & Scaling Planner +
What five real workloads actually cost
Headline per-token prices rarely decide a bill. Request shape does. These figures are calculated from the verified rates for Claude Sonnet 5 and GPT-5.6 Terra, including any long-context tier that applies once a request crosses its threshold.
| Workload | Input | Output | Claude Sonnet 5 | GPT-5.6 Terra | Difference |
|---|---|---|---|---|---|
| Short chat turn A typical assistant exchange. | 1,000 | 500 | $0.0070 | $0.0080 | 1.14x cheaper on Claude Sonnet 5 |
| RAG answer Five retrieved chunks plus a question. | 12,000 | 800 | $0.032 | $0.034 | 1.05x cheaper on Claude Sonnet 5 |
| Code review A medium pull request with surrounding files. | 60,000 | 2,000 | $0.140 | $0.144 | 1.03x cheaper on Claude Sonnet 5 |
| Whole-document analysis A long report or contract read in one call. | 300,000 | 4,000 | $0.640 | $1.27 (tier) | 1.99x cheaper on Claude Sonnet 5 |
| Full-context load Filling most of a one-million-token window. | 900,000 | 4,000 | $1.84 | $3.67 (tier) | 2.00x cheaper on Claude Sonnet 5 |
"(tier)" marks a request that crossed a long-context threshold, so it is billed above the headline rate. Output length is the assumption most worth changing for your own case: paste a real prompt into the calculator above to replace these with your numbers.
The cached-input rate is the number most comparisons miss
Agents, chat threads and RAG pipelines resend the same system prompt and context on every call. Those repeated tokens bill at the cached rate, not the headline rate, so a model can be cheaper on paper and more expensive in production.
| Model | Fresh input / 1M | Cached input / 1M | Discount | Break-even |
|---|---|---|---|---|
| Claude Sonnet 5 | $2.00 | $0.200 | 90% | 1 read |
| GPT-5.6 Terra | $2.00 | $0.200 | 90% | 1 read |
Break-even assumes a cache write costs about 1.25x the base input rate, which is what OpenAI and Anthropic currently document for a short time-to-live. It answers one question: how many times a cached prefix must be re-read before caching is cheaper than paying full price each call. Above that count, every further read saves the discount shown.
What replaced GPT-4o
GPT-4o occupied the balanced tier: capable enough for most production work, priced well below the frontier. That slot is now held by GPT-5.6 Terra on the OpenAI side and Claude Sonnet 5 on the Anthropic side.
If you are costing a migration from GPT-4o, compare against these rather than against a frontier model. Moving straight to the frontier tier is the most common way a migration budget is overestimated.
The rates that decide it
Both list $2 per million input tokens. Claude Sonnet 5 lists $10 per million output against $12 for GPT-5.6 Terra.
Because input price is identical, the entire comparison rests on output volume and on cached input. A workload that writes little will see almost no difference; an agent that writes long answers on every turn will.
Why old cost ratios do not transfer
Token counts are tokenizer-specific. A cost ratio you measured between GPT-4o and a Claude model in an earlier generation does not carry forward, because both the rates and the tokenizers changed.
Re-measure with your current prompt. The calculator below counts OpenAI text exactly in your browser and projects Claude counts deterministically, and labels which is which.
Model the workload, not the marketing price.
Compare input and output rates together. An output-heavy agent can reverse a decision that looks obvious when judged on prompt cost alone.
Review calculation methodology →Measure against the model you use
Other comparisons worth running
Compare approaches, not just models
Frequently asked questions
Is GPT-4o still available?+
GPT-4o no longer appears in current published pricing. Its role in the balanced tier is now filled by GPT-5.6 Terra. If you are costing a migration, compare against the balanced tier rather than a frontier model.
Which is cheaper, Claude or GPT, in the balanced tier?+
Input price is identical at $2 per million. Claude Sonnet 5 is cheaper on output at $10 per million against $12 for GPT-5.6 Terra, so output-heavy workloads favour Claude.
Can I reuse a cost ratio I measured last year?+
No. Both the published rates and the tokenizers have changed, so an old ratio will mislead you. Re-measure with your current prompt.