Skip to content
UniKit

Text similarity

Compare two texts item by item: Levenshtein distance and normalized similarity, Jaro-Winkler, Dice (bigram), cosine similarity (term frequency), longest common subsequence, containment — plus character-level difference highlighting, length statistics and a size guard for very large inputs.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Two texts

Similarity metrics

Levenshtein distance3
Normalized similarity0.5714
Jaro similarity0.7460
Jaro-Winkler0.7460
Dice coefficient (bigram)0.3636
Cosine similarity (term frequency)0.0000
Longest common subsequence4 (0.5714)
ContainmentNeither contains the other · 3

Single-character insertions, deletions and substitutions needed to turn A into B.

1 − distance / longer length; 1 means identical.

Uses a match window and transpositions; good for short strings.

Jaro with a bonus for a shared prefix; good for names and typos.

Overlap of adjacent character pairs, 0–1.

Cosine of the term-frequency vectors; CJK text is tokenized per character.

Length (and ratio) of the longest common subsequence, gaps allowed.

Whether one text is a substring of the other, plus the longest common substring.

Difference highlight

Characters in red are the differences (based on the longest common subsequence).

kitten
sitting

Length statistics

—Text AText B
Characters67
Characters without spaces67
Words11
Lines11
Bytes (UTF-8)67

What this tool does

  • Judge how alike two texts really are: a revised draft against the original, two title candidates, suspected duplicate reviews — all with several metrics at once.
  • Pick the metric that fits: Jaro-Winkler for short strings, Dice (bigram) or cosine for paragraphs, Levenshtein distance for character-level edits.
  • Locate the changes with the difference highlight — the characters in red are exactly where the two texts diverge.
  • Get length statistics for both sides (characters, characters without spaces, words, lines, UTF-8 bytes) without opening a separate counter.

Example

Input

A: kitten
B: sitting

Output

Levenshtein distance: 3
Normalized similarity: 0.5714
Jaro: 0.7460
Jaro-Winkler: 0.7460
Dice (bigram): 0.3636
Cosine (term frequency): 0
Longest common subsequence: 4 (0.5714)
Containment: neither contains the other · 3

Matching is case-insensitive by default; cosine is 0 because kitten and sitting share no identical word.

Frequently asked questions

Is a similarity of 0.57 high or low?

There is no universal threshold — it depends on the use case. For short strings 0.57 usually means a noticeable edit, while in a long paragraph the same number may only reflect a few reworded phrases. Read the Levenshtein distance and the highlighted diff alongside the percentage.

Why are Jaro and Jaro-Winkler identical here?

Because the prefix bonus only applies when Jaro is above 0.7. kitten and sitting have a Jaro of 0.746 but different first characters, so the shared-prefix bonus is 0 and both values match.

Why is cosine similarity 0?

Cosine compares term-frequency vectors, so only identical words overlap. kitten and sitting share no word at all (CJK text is tokenized per character), which makes the dot product 0. Use Dice (bigram) to compare character-level overlap instead.

What happens with very long texts?

Metrics accept up to 20000 characters per side and fail fast beyond that; difference highlighting is stricter — 3000 characters per side and at most 4 million DP cells. The limits exist so the O(n×m) dynamic programming never freezes the browser.

Is my text uploaded?

No. Every metric is computed locally in your browser with JavaScript and the page makes no network requests.

Keywords:文本相似度similaritylevenshtein编辑距离jaro-winklerdice余弦相似度cosinelcs差异diff

Related tools