Methodology
TokenCalculator.dev separates exact counts, provider formulas, and estimates. The accuracy badge beside every result describes which method was used.
Text
OpenAI content uses the matching local BPE encoding. Claude, Gemini, and DeepSeek use deterministic UTF-8 byte ratios calibrated by provider family. Estimates are useful for planning but the provider response remains the final billing record.
Documents
Text and source files are read directly. DOCX and text-based PDF files are converted to plain text locally. Scans, layout, embedded images, and OCR are excluded.
Images
Images are not OCR’d or uploaded. The browser reads their dimensions and applies the selected provider’s published patch, tile, detail, and cap rules. Unsupported model combinations return an error.
Prices
Pricing and context-window limits are refreshed from the open-source Models.dev catalog in the browser. A bundled, versioned registry provides an offline fallback. Prices are USD per million input tokens; provider invoices remain authoritative.
Calculation formula
Input cost equals input tokens divided by 1,000,000, multiplied by the selected model’s input price per million tokens. Applicable long-context tiers are selected by the pricing engine. Cached input and output projections are shown separately where the calculator supports them.
Known differences
Provider API responses may include message framing, system instructions, tool definitions, reasoning, or other server-side processing that is not part of raw pasted content. Code, JSON, punctuation, Unicode, and non-English text may also behave differently across tokenizers.
Review policy
Changing model facts are checked against Models.dev and linked provider sources. The bundled registry records a verification date for every model. Report reproducible discrepancies through the contact page.
Methodology reviewed: August 22, 2026.