Skip to content
UniKit

CSV validator

Check CSV or TSV files against RFC 4180 and get line numbers plus issue types: ragged rows, unclosed quotes, duplicated or empty headers, BOM and encoding problems, mixed line endings and control characters.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Validation settings

The parser follows RFC 4180; encoding checks rely on observable evidence: a BOM, replacement characters (U+FFFD) and ASCII control characters.

Validation result

Paste CSV to start validating

What this tool does

  • Check a file before importing it: database loaders usually report a single "error on line N", while this lists every offending line at once.
  • Audit exports from upstream systems: duplicated or empty headers silently drop columns when mapping to SQL or application fields, and they are flagged here.
  • Diagnose "Excel shows garbage" problems: see whether the file carries a UTF-8 BOM, replacement characters (U+FFFD) or invisible control characters.
  • Produce a quality report before a cleanup: the issue list is organised by line number so whoever writes the fix script knows exactly where to look.

Example

Input

id,id,,age
1,2,3,4
5,6
"unclosed,7,8

Output

Error line 4: the quote is never closed
Warning line 1, column 2: duplicated header "id"
Warning line 1, column 3: empty header
Error line 3: 2 columns instead of the expected 4
Error line 4: 1 column instead of the expected 4

Line numbers refer to where a record starts; newlines inside quoted fields shift the following numbers, which is exactly why "line 4" here is the record that begins on line 4.

Frequently asked questions

Which problems are checked?

Errors: rows whose column count differs from the header, unclosed quotes, and replacement characters (U+FFFD) that mean bytes were lost while decoding. Warnings: duplicated or empty headers, unescaped quotes inside a field, ASCII control characters, and mixed CRLF / LF line endings. Notes: a leading UTF-8 BOM, blank lines, and files with a header but no data rows.

How are line numbers calculated?

A line number is where the record starts, matching what your editor shows. Newlines inside a field belong to the same record, so the next record skips them — that is usually the missing piece when an import tool reports a line number you cannot find in the file.

My file passes validation but the import still fails. Why?

Validation only covers structural and encoding problems that can be judged offline. Business constraints — primary key clashes, oversized values, type mismatches such as abc in an INT column — are still reported by the database. The point here is to rule out format problems first so those errors are trustworthy.

Can it fix the problems automatically?

No, it only diagnoses. Auto-quoting or dropping columns can silently change your data, so the fix stays in your hands. Once repaired you can feed the file to the CSV to SQL or CSV viewer tools.

What if the delimiter is detected wrongly?

The default guess is the delimiter that appears most often in the first line. If it guesses wrong, pick comma, semicolon, tab or pipe in the dropdown and the issue list is recalculated immediately.

Keywords:csv validatorcsv checkerrfc 4180validate csvencoding checkdata qualityCSV格式验证CSV 校验列数不一致编码异常表头重复

Related tools