Text similarity
Compare two texts item by item: Levenshtein distance and normalized similarity, Jaro-Winkler, Dice (bigram), cosine similarity (term frequency), longest common subsequence, containment — plus character-level difference highlighting, length statistics and a size guard for very large inputs.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Similarity metrics
30.57140.74600.74600.36360.00004 (0.5714)Neither contains the other · 3Single-character insertions, deletions and substitutions needed to turn A into B.
1 − distance / longer length; 1 means identical.
Uses a match window and transpositions; good for short strings.
Jaro with a bonus for a shared prefix; good for names and typos.
Overlap of adjacent character pairs, 0–1.
Cosine of the term-frequency vectors; CJK text is tokenized per character.
Length (and ratio) of the longest common subsequence, gaps allowed.
Whether one text is a substring of the other, plus the longest common substring.
Difference highlight
Characters in red are the differences (based on the longest common subsequence).
kitten
sitting
Length statistics
| — | Text A | Text B |
|---|---|---|
| Characters | 6 | 7 |
| Characters without spaces | 6 | 7 |
| Words | 1 | 1 |
| Lines | 1 | 1 |
| Bytes (UTF-8) | 6 | 7 |
What this tool does
- Judge how alike two texts really are: a revised draft against the original, two title candidates, suspected duplicate reviews — all with several metrics at once.
- Pick the metric that fits: Jaro-Winkler for short strings, Dice (bigram) or cosine for paragraphs, Levenshtein distance for character-level edits.
- Locate the changes with the difference highlight — the characters in red are exactly where the two texts diverge.
- Get length statistics for both sides (characters, characters without spaces, words, lines, UTF-8 bytes) without opening a separate counter.
Example
Input
A: kitten B: sitting
Output
Levenshtein distance: 3 Normalized similarity: 0.5714 Jaro: 0.7460 Jaro-Winkler: 0.7460 Dice (bigram): 0.3636 Cosine (term frequency): 0 Longest common subsequence: 4 (0.5714) Containment: neither contains the other · 3
Matching is case-insensitive by default; cosine is 0 because kitten and sitting share no identical word.
Frequently asked questions
Is a similarity of 0.57 high or low?
There is no universal threshold — it depends on the use case. For short strings 0.57 usually means a noticeable edit, while in a long paragraph the same number may only reflect a few reworded phrases. Read the Levenshtein distance and the highlighted diff alongside the percentage.
Why are Jaro and Jaro-Winkler identical here?
Because the prefix bonus only applies when Jaro is above 0.7. kitten and sitting have a Jaro of 0.746 but different first characters, so the shared-prefix bonus is 0 and both values match.
Why is cosine similarity 0?
Cosine compares term-frequency vectors, so only identical words overlap. kitten and sitting share no word at all (CJK text is tokenized per character), which makes the dot product 0. Use Dice (bigram) to compare character-level overlap instead.
What happens with very long texts?
Metrics accept up to 20000 characters per side and fail fast beyond that; difference highlighting is stricter — 3000 characters per side and at most 4 million DP cells. The limits exist so the O(n×m) dynamic programming never freezes the browser.
Is my text uploaded?
No. Every metric is computed locally in your browser with JavaScript and the page makes no network requests.
Keywords:文本相似度similaritylevenshtein编辑距离jaro-winklerdice余弦相似度cosinelcs差异diff