Skip to content
UniKit

Text to Unicode escapes

Convert text to and from Unicode escapes: \uXXXX, \xXX, U+XXXX, 〹, HTML entities and URL percent-encoding, with correct handling of CJK text and emoji surrogate pairs.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Turn it off to keep ASCII as-is and escape only non-ASCII characters
Result

Decoding auto-detects every style above: backslash escapes, U+ code points, HTML entities and %XX bytes — even mixed together.

What this tool does

  • See what a string actually contains: paste \uXXXX, U+XXXX, HTML entities or %XX percent-encoding and get readable text back.
  • Turn Chinese or emoji into \uXXXX escapes before putting them into JSON, a regex or a source string, so encoding stops being a worry.
  • Debug mojibake: decode a server response such as %E4%BD%A0%E5%A5%BD and see which bytes were really sent.
  • Decoding auto-detects mixed styles — \u4f60, 你 and %E4%BD%A0 in the same line are all restored together, no format guessing needed.

Example

Input

你好 🎉

Output

\u4f60\u597d\u0020\ud83c\udf89

The default format is JS escapes with ASCII escaped too, so the space also becomes \u0020; U+1F389 is outside the basic plane, so it is written as the surrogate pair \ud83c\udf89.

Frequently asked questions

What is the difference between \uXXXX and U+XXXX?

U+XXXX is the notation for a Unicode code point — it documents which character you mean but cannot be placed inside a string literal. \uXXXX is a language escape that actually becomes that character. Code points above U+FFFF must be written as a surrogate pair in JavaScript, which is why 🎉 is \ud83c\udf89.

What changes if I turn off "escape ASCII too"?

Only non-ASCII characters are escaped, so English letters and digits stay as they are and the output is shorter and easier to read. Backslashes — and & < > in HTML format — are still escaped, because otherwise decoding would misread them.

Why does decoding say an escape sequence is broken?

A prefix such as \u, \x or &# appeared without the right number of following digits, for example \u12 or &#xZZ;. The tool reports this instead of guessing your intent — just complete the hexadecimal digits.

How do \xXX and %XX differ?

Both describe bytes rather than code points: \xe4\xbd\xa0 and %E4%BD%A0 both mean the three UTF-8 bytes of 你. If the resulting byte sequence is not valid UTF-8 the tool raises an error rather than silently emitting replacement characters.

Is my text uploaded?

No. Encoding and decoding run entirely in your browser with JavaScript and the page makes no network requests.

Keywords:unicode转义unescapeescapeu+code point码点emojihtml 实体url 编码

Related tools