Skip to content
UniKit

Word cloud data

Turn text into word cloud data: tokenize, drop stopwords, count frequencies and export a word,count CSV, a ranked table and 1–100 weights for your visualisation tool. This tool produces data only — it never renders an image.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Text input

Text to analyse

Options

Remove stopwordsBuilt-in English and Chinese lists (the / and / 的 / 了 …)
Lowercase Latin words

1 – 500

Result

This tool only produces data: a CSV, a table and normalised weights. Render the cloud with your own visualisation or design tool.

Paste some text to see the word frequencies.

What this tool does

  • Feed a word cloud: paste reviews or an article, export the word,count CSV and drop it into ECharts, D3 or a design tool.
  • Spot the topic of a document fast — once the / and / 的 / 了 are filtered out, the top rows are usually the point.
  • Analyse Chinese text with 2-grams instead of single characters so words like 数据 and 统计 stay intact.
  • Scale font sizes from the 1–100 weights: the most frequent word is pinned to 100 and the rest follow proportionally.

Example

Input

The quick brown fox. The QUICK fox runs! (defaults: min length 2, stopwords on)

Output

CSV: word,count / fox,2 / quick,2 / brown,1 / runs,1 — top frequency 2, weights 100 for fox and quick, 50 for brown and runs

The built-in English stopword list removes "the", "QUICK" is folded into "quick" by lowercasing, and punctuation acts as a separator.

Frequently asked questions

Why does it only output data instead of drawing the cloud?

Rendering involves fonts, layout, colours and export formats, and dedicated chart or design tools do it better. This tool gets the error-prone part right — tokenizing, stopword removal, counts and weights — and exports a CSV plus a table.

How is Chinese text segmented?

Without any third-party segmenter: either per character, or with 2-grams (a sliding window over adjacent characters). 2-grams split 词云数据 into 词云 / 云数 / 数据, which usually gives more meaningful counts.

Can I change the stopword list?

The built-in English and Chinese lists can be switched off entirely, and you can add your own words under Extra stopwords (comma separated). Custom words are lowercased along with the Latin tokens when the lowercase option is on.

Are words escaped in the CSV?

Yes. Words containing a comma, quote or newline are wrapped in double quotes with inner quotes doubled, so importing the file into Excel or Google Sheets keeps every column intact.

Is my text uploaded anywhere?

No. Tokenizing and counting happen in the browser, the page makes no requests, and a refresh clears everything.

Keywords:word cloudword frequencyword counterword cloud datacsv exportstopwords词云词频统计词云数据分词停用词中文分词

Related tools