AI Watermark Check

Invisible Unicode character list

The complete reference for what this site detects. Every row is a code point or range the engine classifies, with the action it takes and the reasoning behind it. This table is generated from the same data file the checker runs on, so it cannot drift out of date relative to the tool.

How to read the table

Why context matters. U+200C is required by Persian orthography and meaningless in English. U+200D holds emoji together. U+2800 is a real braille cell. U+3164 is a Korean filler and the most popular way to make an "empty" username. A cleaner that ignores context does not clean text, it damages it.

57 individually named code points plus 7 ranges. This table is generated from the same data file the checker itself uses, so it always matches what the tool does.

Code pointNameCategoryRiskWhat the cleaner doesWhy it matters and where it comes from
U+200BZERO WIDTH SPACEZero-width / invisibleMediumRemovedOccupies no width but is a real character: it breaks string equality, splits words for search and regex, and survives copy-paste invisibly. Word-wrap hints inserted by CMS editors, PDF exporters and web typography scripts.
U+200CZERO WIDTH NON-JOINERZero-width / invisibleMediumKept or removed depending on contextPrevents two letters from forming a ligature. Meaningless in Latin text, but grammatically required in Persian, Urdu and several Indic scripts. Persian and Indic typing; also injected by some editors around emphasis markup.
U+200DZERO WIDTH JOINERJoinerMediumKept or removed depending on contextGlues code points into one glyph. It is what holds family and profession emoji together, and it is also required in Devanagari and Arabic-script typing. Emoji sequences, Indic and Arabic script text, occasionally editor artefacts.
U+2060WORD JOINERZero-width / invisibleMediumRemovedA zero-width no-break point. Invisible, but it blocks line breaking and breaks word-boundary matching. Typographic tooling; the modern replacement for a leading U+FEFF.
U+2061FUNCTION APPLICATIONZero-width / invisibleMediumRemovedInvisible mathematical operator. Carries meaning only inside MathML; in prose it is silent noise. Copy-paste out of MathML, LaTeX-to-HTML converters and equation editors.
U+2062INVISIBLE TIMESZero-width / invisibleMediumRemovedInvisible multiplication sign. Same story: meaningful in MathML, invisible corruption in plain text. Equation editors and scientific PDFs.
U+2063INVISIBLE SEPARATORZero-width / invisibleMediumRemovedInvisible comma used between mathematical indices. Splits tokens for anything that parses the text. Equation editors and scientific PDFs.
U+2064INVISIBLE PLUSZero-width / invisibleMediumRemovedInvisible addition sign used in mixed numbers. Silent in prose, disruptive in parsers. Equation editors and scientific PDFs.
U+FEFFZERO WIDTH NO-BREAK SPACE (BOM)Zero-width / invisibleMediumRemovedThe byte order mark. At the start of a file it is metadata; anywhere else it is an invisible character that corrupts JSON parsing, CSV headers and shell scripts. Files saved as "UTF-8 with BOM" by Windows editors and Excel exports.
U+180EMONGOLIAN VOWEL SEPARATORZero-width / invisibleMediumKept or removed depending on contextReclassified as a zero-width format character in Unicode 6.3. Legitimate inside Mongolian text, invisible padding anywhere else. Mongolian typing; otherwise deliberate padding, because it long passed as a "space" in username filters.
U+034FCOMBINING GRAPHEME JOINERZero-width / invisibleMediumRemovedAffects collation and combining-mark ordering without rendering anything. In ordinary text it only defeats string comparison. Linguistic and library data; occasionally used to defeat naive deduplication.
U+115FHANGUL CHOSEONG FILLERZero-width / invisibleMediumKept or removed depending on contextA placeholder for a missing Korean initial consonant. Renders as blank space and is widely used to fake empty nicknames. Korean text composition; also a classic "invisible username" trick.
U+1160HANGUL JUNGSEONG FILLERZero-width / invisibleMediumKept or removed depending on contextA placeholder for a missing Korean vowel. Blank on screen, real in the byte stream. Korean text composition; also used for blank display names.
U+3164HANGUL FILLERZero-width / invisibleMediumKept or removed depending on contextRenders as an empty box-width blank. The single most common character behind "invisible" chat messages and profile names. Korean input methods; deliberate use in games, chat apps and forums.
U+FFA0HALFWIDTH HANGUL FILLERZero-width / invisibleMediumKept or removed depending on contextThe halfwidth form of U+3164, with the same blank rendering and the same abuse pattern. Korean input methods and invisible-name tricks.
U+17B4KHMER VOWEL INHERENT AQZero-width / invisibleMediumKept or removed depending on contextNormally invisible. It is a genuine part of Khmer orthography, so deleting it blindly damages Khmer text. Khmer typing; outside Khmer text it is padding.
U+17B5KHMER VOWEL INHERENT AAZero-width / invisibleMediumKept or removed depending on contextSame as U+17B4: invisible, legitimate in Khmer, junk elsewhere. Khmer typing; outside Khmer text it is padding.
U+2800BRAILLE PATTERN BLANKZero-width / invisibleMediumKept or removed depending on contextA braille cell with no raised dots. It is a real, meaningful character in braille content, and a wide invisible blank everywhere else. Braille text and ASCII-art; also used to pad messages past length filters.
U+00ADSOFT HYPHENZero-width / invisibleMediumRemovedShows a hyphen only if the word wraps. Invisible in the middle of a word, yet it defeats exact search, spellcheck and equality checks. Word processors, hyphenation tools and PDF-to-text conversion.
U+206AINHIBIT SYMMETRIC SWAPPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control. No modern renderer needs it; its presence is a sign of legacy conversion or deliberate padding. Legacy document conversion.
U+206BACTIVATE SYMMETRIC SWAPPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control with no rendering effect in current software. Legacy document conversion.
U+206CINHIBIT ARABIC FORM SHAPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion.
U+206DACTIVATE ARABIC FORM SHAPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion.
U+206ENATIONAL DIGIT SHAPESZero-width / invisibleMediumRemovedDeprecated control that switched digit rendering. Invisible, and ignored by modern text stacks. Legacy document conversion.
U+206FNOMINAL DIGIT SHAPESZero-width / invisibleMediumRemovedDeprecated counterpart to U+206E. Invisible and inert. Legacy document conversion.
U+061CARABIC LETTER MARKBidirectional controlHighRemovedAn invisible character with strong right-to-left directionality. It changes how neighbouring text is ordered on screen without changing the stored string. Arabic and Persian layout fixes; also used to disguise reordering.
U+200ELEFT-TO-RIGHT MARKBidirectional controlHighRemovedForces left-to-right ordering for the text around it. Invisible, and it makes what you read differ from what is stored. Mixed-direction documents, spreadsheets and localisation files.
U+200FRIGHT-TO-LEFT MARKBidirectional controlHighRemovedThe mirror of U+200E. It can silently reverse the apparent order of adjacent characters. Mixed-direction documents; also the classic filename-spoofing character.
U+202ALEFT-TO-RIGHT EMBEDDINGBidirectional controlHighRemovedOpens a directional embedding that persists until it is popped. Unbalanced embeddings scramble the display of everything that follows. Legacy bidi markup; deprecated in favour of isolates.
U+202BRIGHT-TO-LEFT EMBEDDINGBidirectional controlHighRemovedOpens a right-to-left embedding. Same unbalanced-state risk as U+202A. Legacy bidi markup; deprecated in favour of isolates.
U+202CPOP DIRECTIONAL FORMATTINGBidirectional controlHighRemovedCloses the most recent embedding or override. Harmless alone, but it is half of every bidi attack sequence. Legacy bidi markup.
U+202DLEFT-TO-RIGHT OVERRIDEBidirectional controlHighRemovedForces every following character to render left-to-right regardless of its own direction. One of the two characters that can make code display differently from how it runs. Almost never legitimate in modern text. Treat with suspicion in source code.
U+202ERIGHT-TO-LEFT OVERRIDEBidirectional controlHighRemovedReverses the rendered order of everything after it. This is the character behind reversed-extension filename tricks and misleading code display. Almost never legitimate. Its presence in code or a filename is a red flag.
U+2066LEFT-TO-RIGHT ISOLATEBidirectional controlHighRemovedOpens an isolated left-to-right run. The modern, safer bidi mechanism, but still invisible state that must be popped. Correct bidi markup in localisation files; also seen in code-display research demos.
U+2067RIGHT-TO-LEFT ISOLATEBidirectional controlHighRemovedOpens an isolated right-to-left run, invisibly changing display order for the text it wraps. Localisation files; also seen in code-display research demos.
U+2068FIRST STRONG ISOLATEBidirectional controlHighRemovedOpens an isolate whose direction is inferred from its first strong character, so its effect depends on the data inside it. Localisation frameworks handling user-supplied names.
U+2069POP DIRECTIONAL ISOLATEBidirectional controlHighRemovedCloses an isolate. Like U+202C, it is the closing half of a sequence that can rewrite what a reader sees. Localisation frameworks.
U+00A0NO-BREAK SPACEUnusual spaceMediumReplaced with a normal spaceLooks exactly like a normal space but is a different character, so exact matches, CSV columns and code indentation quietly fail. Word processors, web pages using  , and PDF text extraction. The most common hidden character of all.
U+1680OGHAM SPACE MARKUnusual spaceMediumReplaced with a normal spaceA space character from the Ogham block that may render as a visible line in some fonts. Deliberate substitution; effectively never typed by accident.
U+2000EN QUADUnusual spaceMediumReplaced with a normal spaceA fixed-width typographic space. Indistinguishable from a normal space in most UI fonts. Typesetting software and PDF extraction.
U+2001EM QUADUnusual spaceMediumReplaced with a normal spaceA fixed-width typographic space one em wide. Typesetting software and PDF extraction.
U+2002EN SPACEUnusual spaceMediumReplaced with a normal spaceHalf-em typographic space that passes for an ordinary space on screen. Typesetting software, InDesign and PDF extraction.
U+2003EM SPACEUnusual spaceMediumReplaced with a normal spaceFull-em typographic space, frequently mistaken for a tab or double space. Typesetting software and PDF extraction.
U+2004THREE-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceFixed-width space used in fine typesetting; breaks whitespace-sensitive parsing. Typesetting software and PDF extraction.
U+2005FOUR-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceFixed-width space used in fine typesetting. Typesetting software and PDF extraction.
U+2006SIX-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceVery narrow fixed-width space, visually identical to a thin gap. Typesetting software and PDF extraction.
U+2007FIGURE SPACEUnusual spaceMediumReplaced with a normal spaceA non-breaking space as wide as a digit, used to align numbers in tables. It reliably breaks numeric parsing. Financial documents, spreadsheets and generated reports.
U+2008PUNCTUATION SPACEUnusual spaceMediumReplaced with a normal spaceA space as wide as a period, used for alignment. Typesetting software.
U+2009THIN SPACEUnusual spaceMediumReplaced with a normal spaceNarrow space used around dashes and units. Very common in text copied from well-typeset web pages. Web typography, Wikipedia, scientific writing.
U+200AHAIR SPACEUnusual spaceMediumReplaced with a normal spaceThe narrowest space in Unicode. Often invisible in practice, yet it still splits words for search. Web typography and typesetting software.
U+202FNARROW NO-BREAK SPACEUnusual spaceMediumReplaced with a normal spaceNon-breaking and narrow. French typography and many number formats use it, so it arrives with copied prices and dates. French locale formatting, spreadsheets and currency output.
U+205FMEDIUM MATHEMATICAL SPACEUnusual spaceMediumReplaced with a normal spaceSpacing used around mathematical operators. Invisible as anything but a space. MathML and equation editors.
U+3000IDEOGRAPHIC SPACEUnusual spaceMediumReplaced with a normal spaceThe full-width space used in Chinese, Japanese and Korean typing. Legitimate in CJK text, and a silent mismatch in code and identifiers. CJK input methods, where it is the space produced in full-width mode.
U+2028LINE SEPARATORLine separatorMediumConverted to a line breakA line break that is not a newline. It terminates lines in text engines but historically broke JSON and JavaScript string literals. Word processors, especially shift-enter line breaks exported to plain text.
U+2029PARAGRAPH SEPARATORLine separatorMediumConverted to a line breakA paragraph break that most tools do not recognise as a newline, so paragraphs collapse into one line. Word processors and rich-text exports.
U+0085NEXT LINE (NEL)Line separatorMediumConverted to a line breakA legacy one-character line break combining CR and LF. Most parsers do not treat it as a newline, so lines silently merge. Mainframe (EBCDIC) exports and Windows-1252 data mis-decoded as UTF-8, where 0x85 is an ellipsis.
U+000DCARRIAGE RETURN (WITHOUT LINE FEED)Line separatorMediumConverted to a line breakA lone CR with no LF after it. Old-Mac line endings and injection tricks use it; many tools show one line while parsers see two, or vice versa. Classic Mac OS files, HTTP header injection payloads and broken concatenation of mixed line endings.
U+0000–U+0008C0 CONTROL CHARACTERS (NUL to BACKSPACE)Control characterHighRemovedTeletype-era control codes. A NUL byte truncates strings in C-based software, and the rest have no meaning in text. Binary data pasted as text, corrupted exports and malformed database dumps.
U+000B–U+000CLINE TABULATION, FORM FEEDControl characterHighRemovedLegacy page and vertical-tab controls. Rendered inconsistently and treated as whitespace by some parsers and not others. Printer output, mainframe exports and old text files.
U+000E–U+001FC0 CONTROL CHARACTERS (SHIFT OUT to UNIT SEPARATOR)Control characterHighRemovedDevice and record-separator controls. Invisible in most editors, and they corrupt CSV, TSV and log parsing. Legacy data interchange formats and terminal captures.
U+007F–U+009FDELETE and C1 CONTROL CHARACTERSControl characterHighRemovedDEL plus the C1 block. Common in text mis-decoded from Windows-1252, where they masquerade as punctuation. Encoding mismatches, especially Windows-1252 data read as UTF-8.
U+E0000–U+E007FTAG CHARACTERSTag characterHighKept or removed depending on contextAn invisible mirror of ASCII. A full sentence can be encoded in tag characters and rendered as nothing at all, which makes this block the standard vehicle for prompt-injection smuggling. The one legitimate use is regional flag emoji. Deprecated language tagging, emoji flag sequences, and deliberate hidden payloads.
U+FE00–U+FE0FVARIATION SELECTORS VS1 to VS16Variation selectorMediumKept or removed depending on contextSelects a glyph variant of the preceding character. VS16 (U+FE0F) is what turns a monochrome symbol into a colour emoji, so removing it blindly destroys emoji. Emoji presentation, keycap sequences, and CJK glyph variants.
U+E0100–U+E01EFVARIATION SELECTORS SUPPLEMENT VS17 to VS256Variation selectorMediumKept or removed depending on contextIdeographic variation selectors that pick a specific glyph for a CJK character. Legitimate after an ideograph, and pure payload space after anything else. Japanese personal and place names; also used to hide data after non-CJK characters.

Test it on your own text

Paste anything into the checker below to see which of these characters it contains and what the cleaner does with each one.

More: what was found, x-ray view and options

X-ray view

Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.

What is deliberately not in this list

Three categories are out of scope, and each for a reason:

Notes on individual entries

U+00A0, the no-break space

The single most common find in real-world text. It looks exactly like a space, and it is produced by autocorrect in word processors, by   in HTML, and by almost every PDF extractor. Most whitespace-trimming code does not trim it, so a value can look trimmed and still carry it.

U+FEFF, the byte order mark

Metadata at the start of a file, and an invisible character everywhere else. A BOM at the top of a JSON file makes parsing fail with a message that points at the wrong thing entirely; at the top of a shell script it breaks the shebang line.

U+200D, the zero-width joiner

Structural in two very different worlds. It builds emoji sequences, a family emoji is three people and two joiners, and it is required in Devanagari and Arabic-script typing. The engine keeps it when both neighbours are emoji, and when both neighbours belong to a script that uses it.

U+E0000 to U+E007F, the tag block

An invisible copy of ASCII. Any sentence can be encoded in it and rendered as absolutely nothing, which is why this block is the standard vehicle for hiding instructions inside text that looks harmless. The one legitimate use is regional flag emoji, where a tag sequence closed by U+E007F encodes the region code, and those are kept.

U+FE0F, variation selector 16

The character that turns a monochrome symbol into a colour emoji, and part of every keycap sequence. Removing it unconditionally is the most common way a naive cleaner mangles emoji. The engine keeps it after any emoji-capable character and before a keycap combining mark, and removes it only where there is nothing it could apply to.

Questions about the list

How many invisible Unicode characters are there in total?

There is no single official count, because "invisible" is a description of rendering rather than a Unicode property. This list covers the code points that appear in real text and cause real failures: the zero-width and format characters, the bidirectional controls, the space characters, the tag block, the variation selectors and the control characters.

Why are ordinary space characters in a list of invisible characters?

Because they are invisible in the sense that matters: a no-break space and an ordinary space are indistinguishable on screen and are different characters. That difference is what breaks lookups, numeric parsing and form validation.

Is this list the same as the one the tool uses?

It is literally the same data. The table on this page and the detection engine are generated from one file at build time, so a change to the engine's behaviour appears here automatically and the two cannot disagree.

How do I detect these characters in my own code?

Iterate over code points rather than UTF-16 units, and test each one against the ranges in this table. The important detail is the ordering: check for characters that must be kept in context before removing anything, or you will destroy emoji and non-Latin text on the way through.

Related pages