Skip to content
UniKit

LLM token counter

Count the tokens in a text with the gpt-tokenizer BPE vocabulary, together with characters, words, lines, UTF-8 bytes, a preview of the first N tokens and the estimated cost at a given price — empty input simply returns 0.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Text

Tokenisation uses the gpt-tokenizer default vocabulary (o200k_base — GPT-4o, 4.1, GPT-5 and the o-series), so Chinese and English are split by the same BPE rules.

Result

Tokens15
Characters (code points)55
Words18
Lines2
UTF-8 bytes85
Estimated cost$0.000038
within the limitToken limit check: 1000
Token preview
  • Uni
  • Kit
  • 是
  • 一个
  • 在线
  • 工具
  • 集合
  • ,
  • 永久免费
  • 。
  • Token
  • counting
  • is
  • deterministic
  • .

Tokens are separated by ·; spaces and newlines are shown as-is, exactly as the model sees them.

What this tool does

  • Measure a prompt before you ship it: context windows are hard limits, and checking the token count up front is cheaper than debugging truncation later.
  • Estimate what a call costs — pick a model or type your own price and see what that input costs before running a batch job.
  • See how much non-Latin text really costs: without spaces, guessing tokens from character counts underestimates badly, so read the real split here.
  • Understand how the model slices text: the preview decodes every token back into text, which explains why one English word can be two tokens or why a chunk of JSON is so expensive.

Example

Input

你好,世界

Output

Tokens 3 · characters (code points) 5 · words 4 · lines 1 · UTF-8 bytes 15 · preview: 你好 / , / 世界

Tokenisation uses the gpt-tokenizer default vocabulary (o200k_base): the common Chinese words 你好 and 世界 are one token each while the punctuation takes one of its own, so five characters cost three tokens.

Frequently asked questions

Which model’s tokenizer is this?

The gpt-tokenizer default vocabulary, o200k_base, which is what GPT-4o, GPT-4.1, GPT-5 and the o-series use — so the numbers are exact for those models. Claude, Gemini and older GPT-3.5 (cl100k_base) use different vocabularies where the same text can differ by 10–30%, so treat those counts as an order of magnitude.

Why are there fewer tokens than words in English but close to one per character in Chinese?

BPE merges frequent sequences: a common English word is a single token, so 100 words is usually around 130 tokens. Chinese words get merged too (你好 is one token) but at a finer granularity, so a Chinese character is typically 0.5–1 token. Estimating Chinese tokens from character counts is therefore optimistic — use the measured number instead.

How accurate is the cost estimate?

The formula is tokens ÷ 1,000,000 × price and covers input tokens only — no output tokens, cache discounts or batch pricing. The preset prices are reference values for a few common OpenAI models and they change over time, so check the vendor’s pricing page; the field is editable, so put in what you actually pay.

What happens with very long text?

Beyond 50000 characters only the first 50000 are measured and a warning is shown above the result, so you never get a silently wrong number for the whole document. For exact counts of long documents, paste it in chunks and add the results.

Is my text uploaded anywhere?

No. The vocabulary ships with the page and every calculation happens in your browser; no requests are made, and nothing you typed survives a refresh.

Keywords:token counterllm tokensgpt tokenizerbpe tokensprompt costToken 计数Token 计数器大模型 Token提示词长度费用估算

Related tools