Unicode Normalizer

The same visible string can be stored as different byte sequences depending on whether an accent lives as a single code point or as a letter plus a combining mark. This normalizer converts text between the four Unicode normalization forms (NFC, NFD, NFKC, NFKD) so your strings compare cleanly in databases, search indexes, deduplication scripts and regex matches.

How to normalize Unicode text

  1. 1

    Paste your input

    Drop in the string: accented letters, CJK text, ligatures, anything that feels inconsistent.

  2. 2

    Pick a form

    Choose NFC (composed, the usual web form), NFD (decomposed), NFKC (compatibility composed) or NFKD (compatibility decomposed).

  3. 3

    Compare the counts

    The tool reports the character (code point) count and the byte length before and after, so you can see whether the form shortened or lengthened the string.

  4. 4

    Copy the normalized output

    Use the result in your database, API payload, slug generator or test fixture.

The four forms and when to use each

Unicode normalization is defined by the standard annex UAX #15. The four forms differ along two axes: canonical versus compatibility, and composed versus decomposed.

Side-by-side reference

Form Composition Mapping When to use
NFC Composed Canonical only Default for web content, databases, filenames
NFD Decomposed Canonical only Text processing that strips accents per character
NFKC Composed Includes compat. Search, identifiers, spam filters, display folding
NFKD Decomposed Includes compat. Aggressive normalization before removing diacritics

Canonical forms keep the meaning

Type é and switch forms. NFC reports 1 character and 2 bytes (the single code point U+00E9), while NFD reports 2 characters and 3 bytes (a plain e followed by a combining acute accent, U+0301). The two look identical on screen but are different byte sequences, and that mismatch is exactly what normalization fixes. Canonical forms only rearrange the bytes, so é in either form always means the letter e with an acute accent.

What “compatibility” actually changes

The K forms (NFKC, NFKD) also rewrite characters that look related but carry different formatting. This is lossy: you cannot get the original back. For example:

  • fi (U+FB01, Latin small ligature fi) becomes fi (two letters)
  • ① (U+2460, circled digit one) becomes 1
  • ㌀ (U+3300, the CJK square “apaato”) becomes アパート
  • Full-width A (U+FF21) becomes A

That is powerful for search and deduplication but destructive for typography, so reach for the K forms only when visual fidelity does not matter.

A practical rule

  • Store user content in NFC.
  • Compare identifiers in NFC so café written as one code point and café written as e + U+0301 match; step up to NFKC when you also need fi and fi, or full-width A and A, to compare equal, as identifier rules in UAX #31 do.
  • Build stripped-accent slugs by normalizing to NFD and deleting the combining-marks block U+0300 to U+036F.

Frequently Asked Questions

Usually it means your input was already in the target form. Paste text that mixes sources (a Mac filename, a copied PDF snippet, a legacy database row) and you are far more likely to see the character and byte counts change.

For display URLs, NFC. For slugs where you want plain ASCII, normalize to NFD, strip the combining marks with a regex on U+0300 to U+036F, then lowercase. NFKD is an alternative when you also want to fold ligatures and circled digits.

Canonical forms (NFC, NFD) never change meaning, they only rearrange bytes. Compatibility forms (NFKC, NFKD) do change it: ligatures split, full-width letters fold and superscripts flatten. Use them only when that is what you want.

Normalization runs on our server using PHP’s Normalizer class and the result comes back on the same request; your text is not stored. We log only an anonymous event noting which form was applied, never the content itself.

Related Tools

Tool available in other languages