Skip to content
UniKit

UTF-8 encoder

Convert between text and UTF-8 bytes: see the code point (U+XXXX) and hex, decimal and binary views of every character, or decode a byte sequence back to text with strict invalid-sequence detection.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Encoding result

Type something above to see the result

What this tool does

  • Debug mojibake: look at the UTF-8 bytes of a character and work out which encoding turned it into garbage.
  • Check bytes while writing hex or escape sequences by hand: Chinese characters take 3 bytes and emoji take 4, so a wrong length is obvious.
  • Decode a byte sequence back to text: paste E4 B8 AD or 228 184 173 to get 中, with per-byte hex, decimal and binary views.
  • Explain JavaScript string lengths: see how many code points, UTF-16 units and bytes a character takes, and why an emoji has length 2.

Example

Input

中文 A 🎉

Output

Code points: 6
Bytes: 13
Hex view: E4 B8 AD E6 96 87 20 41 20 F0 9F 8E 89
Per-character detail: 中 U+4E2D = E4 B8 AD; 文 U+6587 = E6 96 87; space U+0020 = 20; A U+0041 = 41; space U+0020 = 20; 🎉 U+1F389 = F0 9F 8E 89

This is the text → bytes direction: the two Chinese characters take 3 bytes each, the spaces and A take 1 each, and 🎉 takes 4, giving 13 bytes for only 6 code points. In the bytes → text direction, pasting E4 B8 AD gives 中 back.

Frequently asked questions

Why is the byte count so much larger than the code point count?

UTF-8 is variable length: 1 byte for ASCII, 2 for Latin extensions and common symbols, 3 for CJK and 4 for emoji and other astral characters. Chinese text therefore takes about three times as many bytes as characters, so length limits must be counted in bytes.

How can I enter the three number bases?

Hex accepts E4 B8 AD, e4b8ad, 0xE4,0xB8 or E4-B8-AD; decimal accepts 228 184 173 (each 0–255); binary accepts 11100100 10111000 (1–8 bits each). Spaces, commas, semicolons, colons, pipes and hyphens all work as separators.

Why does decoding report “not a valid UTF-8 byte sequence”?

Decoding runs in strict mode (TextDecoder with fatal), so overlong encodings, surrogate code points (U+D800–U+DFFF) and truncated multi-byte sequences are rejected. That usually means the bytes are not UTF-8 at all — often GBK or Latin-1 — and must be decoded with the matching charset.

What happens with a lone surrogate in the text?

For example a half-truncated emoji leaves an unpaired U+D83C. UTF-8 cannot represent such a code unit, so TextEncoder substitutes U+FFFD (); the page warns you and the byte count reflects the substituted output.

How is this different from the Unicode character table?

The table is for looking up a single code point’s properties (category, HTML entities, JS escapes, block) by clicking or searching. This tool works on whole strings, listing as many bytes as you type and converting bytes back to text. Neither one goes online — everything stays local.

Keywords:utf8utf-8encodedecodehexcode pointunicodeUTF-8 编码字节序列码点十六进制字符编码

Related tools