How to normalize Unicode text
- Paste text and select a normalization form.
- Read the normalized text and compare the code point sequences.
- Copy the normalized result into your document or dataset.
Canonical versus compatibility forms
NFC composes canonical sequences where possible; NFD decomposes them. For example, an e followed by U+0301 becomes U+00E9 in NFC. NFKC and NFKD also apply compatibility mappings: full-width letters and ligatures can become ordinary letters. These mappings can remove meaningful formatting distinctions.
What normalization does not do
This tool does not translate, change case, remove accents, or detect look-alike characters. Code point counts are not visible character counts. Input whitespace is preserved. The inspector shows the first 200 code points; the normalized result includes all input, up to 100,000 UTF-16 units.
Local text processing
Normalization runs locally in JavaScript without uploading your text.