Charset & encoding converter
Convert the same text between UTF-8, UTF-16, GBK (a built-in 3,500 common-character subset), Unicode escapes, HTML entities, URL percent-encoding and octal/hex byte sequences, with a per-character byte table and byte counts.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Byte table
UTF-8 and GBK bytes per character; characters outside the built-in table show as —
What this tool does
- Debug mojibake: see exactly which bytes one piece of text produces in UTF-8, GBK and UTF-16 side by side instead of switching editors.
- Bridge a legacy GBK system: turn GBK bytes from a scraped page into UTF-8 text, or encode UTF-8 text as GBK bytes before posting it to an API that only speaks GBK.
- Check the escaping a spec asks for — Unicode escapes, HTML entities, URL percent-encoding, octal and hex byte sequences — and confirm how many bytes a CJK character really takes.
- Paste octal escapes like \344\275\240 from a log and get the original Chinese text back.
Example
Input
Hello, UniKit! 你好
Output
48656c6c6f2c20556e694b69742120e4bda0e5a5bd
The default pair is text → UTF-8 hex, so the Chinese characters are UTF-8 encoded first; switching the target to GBK yields c4e3bac3.
Frequently asked questions
Why does converting to GBK fail for some characters?
The built-in GBK table only covers GB2312 level-1 hanzi plus the symbol area — roughly 3,500 common characters. Level-2 hanzi, the GBK extensions and emoji are not in it, so the tool reports an error instead of emitting wrong bytes. Use a local iconv-based tool if you need the full GBK repertoire.
Why are plain ASCII letters not escaped in Unicode mode?
By default only non-ASCII characters and backslashes are escaped, so Chinese becomes \u4f60\u597d while UniKit! stays readable. Tick "escape ASCII as well" and every letter turns into \u00XX, which produces a strictly ASCII-safe string.
Is URL percent-encoding the same as encodeURIComponent?
No. This tool follows RFC 3986 and leaves -_.!~*'() unescaped, encoding CJK as UTF-8 bytes such as %E4%BD%A0. encodeURIComponent escapes a slightly different set, so the two outputs are not always identical.
What makes a hex byte string valid?
The number of hex digits must be even. Spaces, commas, colons, dashes and 0x prefixes are stripped before parsing. A five-digit string or a stray letter triggers a "not a valid hex byte sequence" error.
Is anything uploaded?
No. Every conversion runs in your browser with JavaScript, the page issues no network requests, and the GBK table ships with the bundle — it works offline.
Keywords:charsetencodinggbkgb2312utf-8utf-16unicode escapehtml entityurl encode字符集编码转换字节