Unicode character table
A browsable table of 41 common Unicode blocks with character, U+XXXX, hex and decimal search, plus UTF-8 bytes, UTF-16 units, HTML entities, JS escapes and general category for every code point.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Range U+0000–U+007F · Total 128 code points · Shown 128
Click a character to copy it and see every encoding of it here.
What this tool does
- Find out exactly what a character is: type 中, ←, 🎉 or U+4E2D and get its code point, UTF-8 bytes, UTF-16 units and every escape notation at once.
- Debug mojibake and width issues by seeing which block a character belongs to, its general category, and whether it is a control or whitespace character.
- Get the right notation for HTML or JS: decimal and hex HTML entities plus \uXXXX or \u{XXXXX} escapes, ready to copy.
- Browse symbols by block — arrows, box drawing, mathematical operators, emoji and 37 more common blocks — and copy one on the spot.
Example
Input
U+4E2D
Output
Character: 中 Code point: U+4E2D Decimal: 20013 UTF-8 bytes: E4 B8 AD UTF-16 units: 4E2D HTML entity (decimal): 中 HTML entity (hex): 中 JS escape: \u4E2D General category: Lo Printable: Yes Control: No Whitespace: No
Typing 中, U+4E2D, 4e2d, 20013 or 中 all resolve to the same code point. Category Lo means “other letter” — letters without case, such as Chinese characters and kana.
Frequently asked questions
Why does an emoji take 4 UTF-8 bytes and 2 UTF-16 units?
Characters above U+FFFF need four bytes in UTF-8. JavaScript strings are UTF-16, so anything outside the basic multilingual plane must be encoded as a surrogate pair — 🎉 (U+1F389) is D83C DF89. That is also why str.length overcounts emoji by one.
What do the general categories Lu, Ll, Lo and So mean?
They are Unicode general category abbreviations: Lu uppercase letter, Ll lowercase letter, Lo other letter (CJK, kana), Nd decimal digit, Po other punctuation, So other symbol (most emoji) and Cc control. The tool derives them with built-in regular expressions, not an external data table.
Why does a block only show 256 characters?
CJK Unified Ideographs alone has more than 20,000 code points, and rendering them all at once visibly freezes the page. Blocks are capped at 256 characters and sparse blocks such as Miscellaneous Symbols and Pictographs are sampled with a step; searching for a specific character is more precise.
What if my code point returns no result?
Search only covers the 41 built-in blocks, so code points outside them (some minority scripts or private-use areas) report “no code point matched”. Use the UTF-8 encoder tool to inspect the bytes of such a character instead.
What happens when I click a character, and does it go online?
Clicking copies the character to your clipboard and shows all of its encodings below. Every value is generated locally with String.fromCodePoint and the page makes no network request.
Keywords:unicodecode point码点字符表utf-8utf-16html entityunicode table字符编码