Question 1

What is a token in an LLM?

Accepted Answer

A token is a chunk of text that a language model reads as a single unit. It is usually a common word, part of a longer word, or a piece of punctuation rather than a whole word or a single character. As a rough rule of thumb, one token is about four characters of English text, and 100 tokens is roughly 75 words.

Question 2

How accurate is this token counter?

Accepted Answer

For OpenAI and Llama models the count is exact, because it uses the same Byte Pair Encoding tokenizers those models ship (o200k_base for GPT-5 and GPT-4o, cl100k_base for GPT-4 and GPT-3.5, and the Llama 3 tokenizer for Llama 3 and 4). For Claude and Gemini the count is a labelled estimate, since those tokenizers are not publicly available to run in the browser. Estimates are typically within about 10% of the real value.

Question 3

Why do different models report different token counts?

Accepted Answer

Each model family is trained with its own tokenizer and vocabulary, so the same sentence can split into a different number of tokens depending on the model. Newer vocabularies like OpenAI's o200k_base are generally more efficient, packing more characters into each token, which lowers the count compared to older tokenizers.

Question 4

Is my text sent to a server?

Accepted Answer

No. All tokenization and counting happens locally in your browser using a tokenizer that loads on the page. Nothing you type or paste is uploaded, logged, or stored, which makes the tool safe to use for private prompts and confidential text.

Client-side token counter vs API-reported counts: which should you trust?

How this is calculated

Verdict

More Tokens scenarios

Frequently asked questions

LLM API Pricing Calculator

LLM VRAM Calculator

JSON Formatter