Skip to content
UniKit

Keyword density analyzer

Analyse word frequency and keyword density for Chinese and English text (2-character sliding windows for Chinese), with custom stopwords, Top N and CSV export.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Options

Overview

Enter some text to see word frequency and density

Keyword density

0 rows

The text is too short or contains only stopwords, so there is nothing to count.

What this tool does

  • Check whether a piece of copy is keyword-stuffed: the Top N list shows the count and share of each term at a glance.
  • Chinese has no spaces, so text is split with a 2-character sliding window (bigrams) — no dictionary needed to surface frequent terms.
  • Mixed Chinese and English are handled separately: English and digits split on alphanumeric runs (keeping inner apostrophes like don't) while Chinese uses the sliding window.
  • Filter out meaningless function words with the built-in stopword list, add your own (brand or product names), then export the Top N as CSV for review.

Example

Input

UniKit runs locally. UniKit never uploads your data.

Output

Total words: 8
Unique words: 6

unikit   2  25.00%
data     1  12.50%
locally  1  12.50%
never    1  12.50%
runs     1  12.50%
uploads  1  12.50%

Case is ignored by default, so UniKit from both sentences merges into one term. The density denominator is every token produced (stopwords included), which is why it / is consume denominator space without appearing in the list.

Frequently asked questions

Why does Chinese produce fragments like "具站" that are not real words?

Chinese is split with a 2-character sliding window (bigrams) rather than dictionary segmentation, so 「工具站」 yields 「工具」 and 「具站」. That is the price of a dictionary-free approach, and the payoff is that any Chinese text can be analysed. Larger windows give longer fragments but the noise never fully disappears.

How is density calculated?

Density = occurrences ÷ total tokens × 100%, rounded to two decimals. The denominator includes tokens removed by stopwords, so the densities in the Top N list will not add up to 100%.

Which numbers do stopwords affect?

Stopwords are only removed from the keyword list, never from the total token count. Total words is always the raw token count; changing stopwords only moves the unique-word count and the Top N list.

What keyword density is acceptable?

This tool does not set a threshold. Search engines stopped ranking on density long ago, and the familiar "2–8%" figure is folklore. Use it to spot obvious stuffing — one term with an absurd share — and to confirm the top terms match your topic.

Is my text uploaded anywhere?

No. Tokenizing, counting and the CSV export all run locally in your browser with no network requests, and it works offline.

Keywords:keyword density关键词密度词频word frequencyseo停用词分词top n密度分析

Related tools