TokenCalculator.dev

Methodology

TokenCalculator.dev separates exact counts, provider formulas, and estimates. The accuracy badge beside every result describes which method was used.

Text

OpenAI content uses the matching local BPE encoding. Claude, Gemini, and DeepSeek use deterministic UTF-8 byte ratios calibrated by provider family. Estimates are useful for planning but the provider response remains the final billing record.

Documents

Text and source files are read directly. DOCX and text-based PDF files are converted to plain text locally. Scans, layout, embedded images, and OCR are excluded.

Images

Images are not OCR’d or uploaded. The browser reads their dimensions and applies the selected provider’s published patch, tile, detail, and cap rules. Unsupported model combinations return an error.

Prices

Pricing and context-window limits are refreshed from the open-source Models.dev catalog in the browser. A bundled, versioned registry provides an offline fallback. Prices are USD per million input tokens; provider invoices remain authoritative.

Calculation formula

Input cost equals input tokens divided by 1,000,000, multiplied by the selected model’s input price per million tokens. Applicable long-context tiers are selected by the pricing engine. Cached input and output projections are shown separately where the calculator supports them.

Known differences

Provider API responses may include message framing, system instructions, tool definitions, reasoning, or other server-side processing that is not part of raw pasted content. Code, JSON, punctuation, Unicode, and non-English text may also behave differently across tokenizers.

Review policy

Changing model facts are checked against Models.dev and linked provider sources. The bundled registry records a verification date for every model. Report reproducible discrepancies through the contact page.

Methodology reviewed: August 22, 2026.