Unicode Lookup

Drop in a mystery character, the weird whitespace your user pasted, a glyph from a font preview, an emoji that breaks your regex, and the lookup returns its Unicode code point (U+XXXX) and the raw UTF-8 bytes in hex, one row per character.

How to identify a Unicode character

  1. 1

    Paste the character

    You can paste one glyph or a whole string: each character is analysed in order.

  2. 2

    Read the code point

    Get the canonical U+ notation, uppercase and zero-padded to at least four hex digits.

  3. 3

    Inspect the UTF-8 bytes

    See the 1 to 4 byte sequence the character takes up on disk or on the wire.

  4. 4

    Copy the results

    The output is a CSV-style table (character, code point, UTF-8 bytes) you can select and paste into a spreadsheet or bug report.

Typical detective work you can do with a lookup

Most Unicode bugs look identical on screen. The only way to tell a regular space from a no-break space, or a curly apostrophe from a prime, is to inspect the code point.

Common gotchas the lookup solves

Looks like Actually is Code point
Space No-break space U+00A0
Space Narrow no-break space U+202F
(nothing) Zero-width space U+200B
(nothing) Zero-width joiner U+200D
Apostrophe Right single quotation mark U+2019
Apostrophe Prime (feet/minutes) U+2032
Hyphen En dash U+2013
Hyphen Non-breaking hyphen U+2011

UTF-8 byte layout at a glance

  • 1 byte, ASCII, U+0000 to U+007F (0xxxxxxx)
  • 2 bytes, Latin supplement, Arabic, Hebrew (110xxxxx 10xxxxxx)
  • 3 bytes, CJK, most symbols (1110xxxx 10xxxxxx 10xxxxxx)
  • 4 bytes, Emoji, rare scripts (11110xxx 10xxxxxx 10xxxxxx 10xxxxxx)

If your database column is a fixed byte length, a 4-byte emoji will eat four slots even though it renders as one glyph, a common source of truncation bugs.

Frequently Asked Questions

Usually BOM markers (U+FEFF) at the start of the file or stray zero-width joiners from a word processor. Paste one into the lookup and it will tell you immediately, then you can strip them with a simple find-and-replace.

Yes. The lookup iterates character by character, so a whole word produces one row per character with its code point and UTF-8 bytes.

A code point is the abstract identity of a character, U+00E9 is always “é” no matter where you store it. UTF-8 is one of several encodings that turn a code point into bytes; the same “é” is the two bytes C3 A9 in UTF-8 but the single byte E9 in Latin-1.

Yes. The lookup is computed on our server: the characters you paste are sent over, analysed, and their code points and UTF-8 bytes are returned right away. We do not store the text you submit.

Related Tools

Tool available in other languages