Invisible Unicode character list
The complete reference for what this site detects. Every row is a code point or range the engine classifies, with the action it takes and the reasoning behind it. This table is generated from the same data file the checker runs on, so it cannot drift out of date relative to the tool.
How to read the table
- Risk: high means the character can change how text is displayed or parsed, bidirectional controls, tag characters and control characters. These are removed and flagged in red.
- Risk: medium means the character is invisible and disruptive but does not alter display order. Removed or replaced, flagged in yellow.
- Kept or removed depending on context means the engine examines the neighbouring characters before deciding. The same code point is legitimate in one place and junk in another.
- Replaced with a normal space is used for every space-like character. They are never deleted, because deleting a space merges the words on either side of it.
Why context matters. U+200C is required by Persian orthography and meaningless in English. U+200D holds emoji together. U+2800 is a real braille cell. U+3164 is a Korean filler and the most popular way to make an "empty" username. A cleaner that ignores context does not clean text, it damages it.
57 individually named code points plus 7 ranges. This table is generated from the same data file the checker itself uses, so it always matches what the tool does.
| Code point | Name | Category | Risk | What the cleaner does | Why it matters and where it comes from |
|---|---|---|---|---|---|
| U+200B | ZERO WIDTH SPACE | Zero-width / invisible | Medium | Removed | Occupies no width but is a real character: it breaks string equality, splits words for search and regex, and survives copy-paste invisibly. Word-wrap hints inserted by CMS editors, PDF exporters and web typography scripts. |
| U+200C | ZERO WIDTH NON-JOINER | Zero-width / invisible | Medium | Kept or removed depending on context | Prevents two letters from forming a ligature. Meaningless in Latin text, but grammatically required in Persian, Urdu and several Indic scripts. Persian and Indic typing; also injected by some editors around emphasis markup. |
| U+200D | ZERO WIDTH JOINER | Joiner | Medium | Kept or removed depending on context | Glues code points into one glyph. It is what holds family and profession emoji together, and it is also required in Devanagari and Arabic-script typing. Emoji sequences, Indic and Arabic script text, occasionally editor artefacts. |
| U+2060 | WORD JOINER | Zero-width / invisible | Medium | Removed | A zero-width no-break point. Invisible, but it blocks line breaking and breaks word-boundary matching. Typographic tooling; the modern replacement for a leading U+FEFF. |
| U+2061 | FUNCTION APPLICATION | Zero-width / invisible | Medium | Removed | Invisible mathematical operator. Carries meaning only inside MathML; in prose it is silent noise. Copy-paste out of MathML, LaTeX-to-HTML converters and equation editors. |
| U+2062 | INVISIBLE TIMES | Zero-width / invisible | Medium | Removed | Invisible multiplication sign. Same story: meaningful in MathML, invisible corruption in plain text. Equation editors and scientific PDFs. |
| U+2063 | INVISIBLE SEPARATOR | Zero-width / invisible | Medium | Removed | Invisible comma used between mathematical indices. Splits tokens for anything that parses the text. Equation editors and scientific PDFs. |
| U+2064 | INVISIBLE PLUS | Zero-width / invisible | Medium | Removed | Invisible addition sign used in mixed numbers. Silent in prose, disruptive in parsers. Equation editors and scientific PDFs. |
| U+FEFF | ZERO WIDTH NO-BREAK SPACE (BOM) | Zero-width / invisible | Medium | Removed | The byte order mark. At the start of a file it is metadata; anywhere else it is an invisible character that corrupts JSON parsing, CSV headers and shell scripts. Files saved as "UTF-8 with BOM" by Windows editors and Excel exports. |
| U+180E | MONGOLIAN VOWEL SEPARATOR | Zero-width / invisible | Medium | Kept or removed depending on context | Reclassified as a zero-width format character in Unicode 6.3. Legitimate inside Mongolian text, invisible padding anywhere else. Mongolian typing; otherwise deliberate padding, because it long passed as a "space" in username filters. |
| U+034F | COMBINING GRAPHEME JOINER | Zero-width / invisible | Medium | Removed | Affects collation and combining-mark ordering without rendering anything. In ordinary text it only defeats string comparison. Linguistic and library data; occasionally used to defeat naive deduplication. |
| U+115F | HANGUL CHOSEONG FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | A placeholder for a missing Korean initial consonant. Renders as blank space and is widely used to fake empty nicknames. Korean text composition; also a classic "invisible username" trick. |
| U+1160 | HANGUL JUNGSEONG FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | A placeholder for a missing Korean vowel. Blank on screen, real in the byte stream. Korean text composition; also used for blank display names. |
| U+3164 | HANGUL FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | Renders as an empty box-width blank. The single most common character behind "invisible" chat messages and profile names. Korean input methods; deliberate use in games, chat apps and forums. |
| U+FFA0 | HALFWIDTH HANGUL FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | The halfwidth form of U+3164, with the same blank rendering and the same abuse pattern. Korean input methods and invisible-name tricks. |
| U+17B4 | KHMER VOWEL INHERENT AQ | Zero-width / invisible | Medium | Kept or removed depending on context | Normally invisible. It is a genuine part of Khmer orthography, so deleting it blindly damages Khmer text. Khmer typing; outside Khmer text it is padding. |
| U+17B5 | KHMER VOWEL INHERENT AA | Zero-width / invisible | Medium | Kept or removed depending on context | Same as U+17B4: invisible, legitimate in Khmer, junk elsewhere. Khmer typing; outside Khmer text it is padding. |
| U+2800 | BRAILLE PATTERN BLANK | Zero-width / invisible | Medium | Kept or removed depending on context | A braille cell with no raised dots. It is a real, meaningful character in braille content, and a wide invisible blank everywhere else. Braille text and ASCII-art; also used to pad messages past length filters. |
| U+00AD | SOFT HYPHEN | Zero-width / invisible | Medium | Removed | Shows a hyphen only if the word wraps. Invisible in the middle of a word, yet it defeats exact search, spellcheck and equality checks. Word processors, hyphenation tools and PDF-to-text conversion. |
| U+206A | INHIBIT SYMMETRIC SWAPPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control. No modern renderer needs it; its presence is a sign of legacy conversion or deliberate padding. Legacy document conversion. |
| U+206B | ACTIVATE SYMMETRIC SWAPPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control with no rendering effect in current software. Legacy document conversion. |
| U+206C | INHIBIT ARABIC FORM SHAPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion. |
| U+206D | ACTIVATE ARABIC FORM SHAPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion. |
| U+206E | NATIONAL DIGIT SHAPES | Zero-width / invisible | Medium | Removed | Deprecated control that switched digit rendering. Invisible, and ignored by modern text stacks. Legacy document conversion. |
| U+206F | NOMINAL DIGIT SHAPES | Zero-width / invisible | Medium | Removed | Deprecated counterpart to U+206E. Invisible and inert. Legacy document conversion. |
| U+061C | ARABIC LETTER MARK | Bidirectional control | High | Removed | An invisible character with strong right-to-left directionality. It changes how neighbouring text is ordered on screen without changing the stored string. Arabic and Persian layout fixes; also used to disguise reordering. |
| U+200E | LEFT-TO-RIGHT MARK | Bidirectional control | High | Removed | Forces left-to-right ordering for the text around it. Invisible, and it makes what you read differ from what is stored. Mixed-direction documents, spreadsheets and localisation files. |
| U+200F | RIGHT-TO-LEFT MARK | Bidirectional control | High | Removed | The mirror of U+200E. It can silently reverse the apparent order of adjacent characters. Mixed-direction documents; also the classic filename-spoofing character. |
| U+202A | LEFT-TO-RIGHT EMBEDDING | Bidirectional control | High | Removed | Opens a directional embedding that persists until it is popped. Unbalanced embeddings scramble the display of everything that follows. Legacy bidi markup; deprecated in favour of isolates. |
| U+202B | RIGHT-TO-LEFT EMBEDDING | Bidirectional control | High | Removed | Opens a right-to-left embedding. Same unbalanced-state risk as U+202A. Legacy bidi markup; deprecated in favour of isolates. |
| U+202C | POP DIRECTIONAL FORMATTING | Bidirectional control | High | Removed | Closes the most recent embedding or override. Harmless alone, but it is half of every bidi attack sequence. Legacy bidi markup. |
| U+202D | LEFT-TO-RIGHT OVERRIDE | Bidirectional control | High | Removed | Forces every following character to render left-to-right regardless of its own direction. One of the two characters that can make code display differently from how it runs. Almost never legitimate in modern text. Treat with suspicion in source code. |
| U+202E | RIGHT-TO-LEFT OVERRIDE | Bidirectional control | High | Removed | Reverses the rendered order of everything after it. This is the character behind reversed-extension filename tricks and misleading code display. Almost never legitimate. Its presence in code or a filename is a red flag. |
| U+2066 | LEFT-TO-RIGHT ISOLATE | Bidirectional control | High | Removed | Opens an isolated left-to-right run. The modern, safer bidi mechanism, but still invisible state that must be popped. Correct bidi markup in localisation files; also seen in code-display research demos. |
| U+2067 | RIGHT-TO-LEFT ISOLATE | Bidirectional control | High | Removed | Opens an isolated right-to-left run, invisibly changing display order for the text it wraps. Localisation files; also seen in code-display research demos. |
| U+2068 | FIRST STRONG ISOLATE | Bidirectional control | High | Removed | Opens an isolate whose direction is inferred from its first strong character, so its effect depends on the data inside it. Localisation frameworks handling user-supplied names. |
| U+2069 | POP DIRECTIONAL ISOLATE | Bidirectional control | High | Removed | Closes an isolate. Like U+202C, it is the closing half of a sequence that can rewrite what a reader sees. Localisation frameworks. |
| U+00A0 | NO-BREAK SPACE | Unusual space | Medium | Replaced with a normal space | Looks exactly like a normal space but is a different character, so exact matches, CSV columns and code indentation quietly fail. Word processors, web pages using , and PDF text extraction. The most common hidden character of all. |
| U+1680 | OGHAM SPACE MARK | Unusual space | Medium | Replaced with a normal space | A space character from the Ogham block that may render as a visible line in some fonts. Deliberate substitution; effectively never typed by accident. |
| U+2000 | EN QUAD | Unusual space | Medium | Replaced with a normal space | A fixed-width typographic space. Indistinguishable from a normal space in most UI fonts. Typesetting software and PDF extraction. |
| U+2001 | EM QUAD | Unusual space | Medium | Replaced with a normal space | A fixed-width typographic space one em wide. Typesetting software and PDF extraction. |
| U+2002 | EN SPACE | Unusual space | Medium | Replaced with a normal space | Half-em typographic space that passes for an ordinary space on screen. Typesetting software, InDesign and PDF extraction. |
| U+2003 | EM SPACE | Unusual space | Medium | Replaced with a normal space | Full-em typographic space, frequently mistaken for a tab or double space. Typesetting software and PDF extraction. |
| U+2004 | THREE-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Fixed-width space used in fine typesetting; breaks whitespace-sensitive parsing. Typesetting software and PDF extraction. |
| U+2005 | FOUR-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Fixed-width space used in fine typesetting. Typesetting software and PDF extraction. |
| U+2006 | SIX-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Very narrow fixed-width space, visually identical to a thin gap. Typesetting software and PDF extraction. |
| U+2007 | FIGURE SPACE | Unusual space | Medium | Replaced with a normal space | A non-breaking space as wide as a digit, used to align numbers in tables. It reliably breaks numeric parsing. Financial documents, spreadsheets and generated reports. |
| U+2008 | PUNCTUATION SPACE | Unusual space | Medium | Replaced with a normal space | A space as wide as a period, used for alignment. Typesetting software. |
| U+2009 | THIN SPACE | Unusual space | Medium | Replaced with a normal space | Narrow space used around dashes and units. Very common in text copied from well-typeset web pages. Web typography, Wikipedia, scientific writing. |
| U+200A | HAIR SPACE | Unusual space | Medium | Replaced with a normal space | The narrowest space in Unicode. Often invisible in practice, yet it still splits words for search. Web typography and typesetting software. |
| U+202F | NARROW NO-BREAK SPACE | Unusual space | Medium | Replaced with a normal space | Non-breaking and narrow. French typography and many number formats use it, so it arrives with copied prices and dates. French locale formatting, spreadsheets and currency output. |
| U+205F | MEDIUM MATHEMATICAL SPACE | Unusual space | Medium | Replaced with a normal space | Spacing used around mathematical operators. Invisible as anything but a space. MathML and equation editors. |
| U+3000 | IDEOGRAPHIC SPACE | Unusual space | Medium | Replaced with a normal space | The full-width space used in Chinese, Japanese and Korean typing. Legitimate in CJK text, and a silent mismatch in code and identifiers. CJK input methods, where it is the space produced in full-width mode. |
| U+2028 | LINE SEPARATOR | Line separator | Medium | Converted to a line break | A line break that is not a newline. It terminates lines in text engines but historically broke JSON and JavaScript string literals. Word processors, especially shift-enter line breaks exported to plain text. |
| U+2029 | PARAGRAPH SEPARATOR | Line separator | Medium | Converted to a line break | A paragraph break that most tools do not recognise as a newline, so paragraphs collapse into one line. Word processors and rich-text exports. |
| U+0085 | NEXT LINE (NEL) | Line separator | Medium | Converted to a line break | A legacy one-character line break combining CR and LF. Most parsers do not treat it as a newline, so lines silently merge. Mainframe (EBCDIC) exports and Windows-1252 data mis-decoded as UTF-8, where 0x85 is an ellipsis. |
| U+000D | CARRIAGE RETURN (WITHOUT LINE FEED) | Line separator | Medium | Converted to a line break | A lone CR with no LF after it. Old-Mac line endings and injection tricks use it; many tools show one line while parsers see two, or vice versa. Classic Mac OS files, HTTP header injection payloads and broken concatenation of mixed line endings. |
| U+0000–U+0008 | C0 CONTROL CHARACTERS (NUL to BACKSPACE) | Control character | High | Removed | Teletype-era control codes. A NUL byte truncates strings in C-based software, and the rest have no meaning in text. Binary data pasted as text, corrupted exports and malformed database dumps. |
| U+000B–U+000C | LINE TABULATION, FORM FEED | Control character | High | Removed | Legacy page and vertical-tab controls. Rendered inconsistently and treated as whitespace by some parsers and not others. Printer output, mainframe exports and old text files. |
| U+000E–U+001F | C0 CONTROL CHARACTERS (SHIFT OUT to UNIT SEPARATOR) | Control character | High | Removed | Device and record-separator controls. Invisible in most editors, and they corrupt CSV, TSV and log parsing. Legacy data interchange formats and terminal captures. |
| U+007F–U+009F | DELETE and C1 CONTROL CHARACTERS | Control character | High | Removed | DEL plus the C1 block. Common in text mis-decoded from Windows-1252, where they masquerade as punctuation. Encoding mismatches, especially Windows-1252 data read as UTF-8. |
| U+E0000–U+E007F | TAG CHARACTERS | Tag character | High | Kept or removed depending on context | An invisible mirror of ASCII. A full sentence can be encoded in tag characters and rendered as nothing at all, which makes this block the standard vehicle for prompt-injection smuggling. The one legitimate use is regional flag emoji. Deprecated language tagging, emoji flag sequences, and deliberate hidden payloads. |
| U+FE00–U+FE0F | VARIATION SELECTORS VS1 to VS16 | Variation selector | Medium | Kept or removed depending on context | Selects a glyph variant of the preceding character. VS16 (U+FE0F) is what turns a monochrome symbol into a colour emoji, so removing it blindly destroys emoji. Emoji presentation, keycap sequences, and CJK glyph variants. |
| U+E0100–U+E01EF | VARIATION SELECTORS SUPPLEMENT VS17 to VS256 | Variation selector | Medium | Kept or removed depending on context | Ideographic variation selectors that pick a specific glyph for a CJK character. Legitimate after an ideograph, and pure payload space after anything else. Japanese personal and place names; also used to hide data after non-CJK characters. |
Test it on your own text
Paste anything into the checker below to see which of these characters it contains and what the cleaner does with each one.
More: what was found, x-ray view and options
X-ray view
Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.
What was found
| Code point | Character | Count | Action |
|---|
What is deliberately not in this list
Three categories are out of scope, and each for a reason:
- Homoglyphs. Cyrillic а, Greek ο and their Latin lookalikes are visible characters and legitimate letters in their own scripts. Deciding that one is an impostor requires knowing the intended script of the surrounding text, which is a linter's job. See invisible characters in code for how they are misused.
- Mathematical alphanumeric symbols. The fake bold and italic letters common on social platforms are visible, so they fall outside a cleaner that removes invisible characters. They do break search and screen readers.
- Combining marks. Stacked diacritics can be used to distort layout, but they are the normal mechanism for writing accented text in many languages. Removing them would break far more than it fixed.
Notes on individual entries
U+00A0, the no-break space
The single most common find in real-world text. It looks exactly like a space, and it is produced by autocorrect in word processors, by in HTML, and by almost every PDF extractor. Most whitespace-trimming code does not trim it, so a value can look trimmed and still carry it.
U+FEFF, the byte order mark
Metadata at the start of a file, and an invisible character everywhere else. A BOM at the top of a JSON file makes parsing fail with a message that points at the wrong thing entirely; at the top of a shell script it breaks the shebang line.
U+200D, the zero-width joiner
Structural in two very different worlds. It builds emoji sequences, a family emoji is three people and two joiners, and it is required in Devanagari and Arabic-script typing. The engine keeps it when both neighbours are emoji, and when both neighbours belong to a script that uses it.
U+E0000 to U+E007F, the tag block
An invisible copy of ASCII. Any sentence can be encoded in it and rendered as absolutely nothing, which is why this block is the standard vehicle for hiding instructions inside text that looks harmless. The one legitimate use is regional flag emoji, where a tag sequence closed by U+E007F encodes the region code, and those are kept.
U+FE0F, variation selector 16
The character that turns a monochrome symbol into a colour emoji, and part of every keycap sequence. Removing it unconditionally is the most common way a naive cleaner mangles emoji. The engine keeps it after any emoji-capable character and before a keycap combining mark, and removes it only where there is nothing it could apply to.
Questions about the list
How many invisible Unicode characters are there in total?
There is no single official count, because "invisible" is a description of rendering rather than a Unicode property. This list covers the code points that appear in real text and cause real failures: the zero-width and format characters, the bidirectional controls, the space characters, the tag block, the variation selectors and the control characters.
Why are ordinary space characters in a list of invisible characters?
Because they are invisible in the sense that matters: a no-break space and an ordinary space are indistinguishable on screen and are different characters. That difference is what breaks lookups, numeric parsing and form validation.
Is this list the same as the one the tool uses?
It is literally the same data. The table on this page and the detection engine are generated from one file at build time, so a change to the engine's behaviour appears here automatically and the two cannot disagree.
How do I detect these characters in my own code?
Iterate over code points rather than UTF-16 units, and test each one against the ranges in this table. The important detail is the ordering: check for characters that must be kept in context before removing anything, or you will destroy emoji and non-Latin text on the way through.
Related pages
- Hidden characters in text - where each class comes from and what it breaks.
- Invisible characters in code - the practical consequences of the bidirectional controls above.
- Clean hidden characters out of text - the cleaner with its scope stated plainly.