Context-window planning · reversible estimate

Convert tokens to words—and words to tokens.

Translate an abstract context-window number into a readable workload range. Choose a content profile because prose, code, JSON, and multilingual text do not share one universal conversion ratio.

Conversion input

Natural-language paragraphs with ordinary punctuation and spacing.

Planning estimate

750,000

estimated words

Lower bound700,000
Upper bound800,000

1,000,000 tokens × 0.75 words/token

Have the actual content? Tokenize it instead of estimating.

Open the exact content calculator →
Conversion model

A planning interval, not fabricated precision.

The converter models words per token rather than claiming that every token equals a fixed fraction of a word. English prose starts from a 0.75 words-per-token center with a 0.70–0.80 interval. More syntactically dense profiles use lower centers and wider intervals.

In the reverse direction, the same profile is inverted: estimated tokens = words ÷ words per token. The calculation is deterministic, but its applicability depends on how closely the unknown content matches the chosen profile.

Review the token measurement methodology →
When accuracy matters

Use the content when you have the content.

Tokens only
Use this converter to estimate how much prose or code may fit inside a context window.
Actual text
Use the main calculator to apply exact local OpenAI BPE or a disclosed provider-family projection.
Files and images
Attach PDF, DOCX, source, data, and supported images so the complete request is measured together.
Plain answers

Frequently asked questions

How many words are 1 million tokens?+

For ordinary English prose, one million tokens is commonly about 750,000 words, with a planning interval of roughly 700,000 to 800,000 words. Code, structured data, technical notation, and multilingual text can differ substantially.

Is converting tokens to words exact?+

No. A token count alone does not contain the original words, language, whitespace, or punctuation. Token-to-word conversion is a workload estimate; exact tokenization requires the actual content and a model tokenizer.

Why does source code produce fewer word-equivalents per token?+

Identifiers, operators, punctuation, indentation, escaped strings, and compact syntax consume tokens without mapping cleanly to natural-language words.

Can I calculate exact tokens from my document?+

Yes. Use the main TokenCalculator.dev calculator with the actual text, PDF, DOCX, code, data file, or supported image. Processing runs locally in the browser.