Skip to content
UniKit

Text deduplicator

Deduplicate text by line, by word or by a custom delimiter, with case/whitespace-insensitive matching, keep-first or keep-last, sorting, duplicate counts, plus CSV dedupe by column.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Source text

Split by
Keep

Deduplicated result

Paste text to deduplicate it automatically

CSV dedupe by column

What this tool does

  • Clean a column of values exported from a database or a log: dedupe by line and see exactly which line numbers each duplicate came from.
  • Tidy a keyword or tag list: split on a custom delimiter such as a comma, enumeration comma or pipe, dedupe, then sort the output.
  • Merge two lists with "Ignore case" on so Apple and apple count as the same entry instead of slipping through as duplicates.
  • Dedupe CSV by column: unique users by id, orders by order number, keeping either the first or the last record.

Example

Input

apple
banana
Apple
banana
cherry
apple (by line, case sensitive, keep first)

Output

apple
banana
Apple
cherry
Stats: 6 items, 4 unique, 2 removed, 2 duplicate groups
Removed: apple appeared 2× (positions 1, 6), banana appeared 2× (positions 2, 4)

With "Ignore case" off, Apple and apple are different keys, so only 2 of the 6 entries are removed. Turning it on drops the unique count to 3.

Frequently asked questions

What is the difference between keeping the first and the last occurrence?

Only the surviving text differs: for apple and Apple in the same list, keep-first yields apple and keep-last yields Apple. "Keep original order" decides whether the output is sorted by the position of the surviving entry or the first occurrence of its key, so the two settings also change ordering.

What does ignoring whitespace actually do?

It strips all whitespace (spaces, tabs, newlines) from the comparison key before matching, so "New York" and "NewYork" collapse into one entry. It only affects comparison — the output still shows the original spelling.

Why does CSV dedupe say the column is invalid?

Columns are 1-based and cannot exceed the real column count. Asking for column 2 in a file where some rows have only one field fails. It also fails when removing the header leaves no rows to deduplicate.

Will fields containing commas break the CSV parsing?

No. The CSV parser is character-by-character and understands quoted fields, commas and line breaks inside a field, and `""` escaping. Output only quotes when necessary, so pasting it back into a spreadsheet or program will not shift columns.

Is my text sent anywhere?

No. Deduplication, statistics and CSV parsing all run in the browser, with no network requests and no caching of your data.

Keywords:去重dedupededuplicateunique唯一重复duplicate文本textcsv按列去重

Related tools