AI Watermark Check

Liste des caractères Unicode invisibles

Voici la liste de référence des caractères invisibles et ambigus que l’outil détecte : espaces de largeur nulle, contrôles bidirectionnels, caractères de balise, espaces inhabituelles, séparateurs de ligne et caractères de contrôle. Chaque ligne indique le point de code, la catégorie, le risque et ce que le nettoyeur en fait.

Plus : éléments trouvés, vue rayons X et options

Vue rayons X

Votre texte avec chaque caractère caché exposé sous forme d’étiquette. Survolez pour voir son nom Unicode.

Collez n’importe quel texte dans l’outil ci-dessus pour voir lesquels de ces caractères il contient et ce que le nettoyeur fait de chacun. Le tableau ci-dessous est généré à partir du même fichier de données que l’outil, il correspond donc toujours à son comportement. Les noms Unicode sont standard et affichés en anglais.

57 individually named code points plus 7 ranges. This table is generated from the same data file the checker itself uses, so it always matches what the tool does.

Code pointNameCategoryRiskWhat the cleaner doesWhy it matters and where it comes from
U+200BZERO WIDTH SPACEZero-width / invisibleMediumRemovedOccupies no width but is a real character: it breaks string equality, splits words for search and regex, and survives copy-paste invisibly. Word-wrap hints inserted by CMS editors, PDF exporters and web typography scripts.
U+200CZERO WIDTH NON-JOINERZero-width / invisibleMediumKept or removed depending on contextPrevents two letters from forming a ligature. Meaningless in Latin text, but grammatically required in Persian, Urdu and several Indic scripts. Persian and Indic typing; also injected by some editors around emphasis markup.
U+200DZERO WIDTH JOINERJoinerMediumKept or removed depending on contextGlues code points into one glyph. It is what holds family and profession emoji together, and it is also required in Devanagari and Arabic-script typing. Emoji sequences, Indic and Arabic script text, occasionally editor artefacts.
U+2060WORD JOINERZero-width / invisibleMediumRemovedA zero-width no-break point. Invisible, but it blocks line breaking and breaks word-boundary matching. Typographic tooling; the modern replacement for a leading U+FEFF.
U+2061FUNCTION APPLICATIONZero-width / invisibleMediumRemovedInvisible mathematical operator. Carries meaning only inside MathML; in prose it is silent noise. Copy-paste out of MathML, LaTeX-to-HTML converters and equation editors.
U+2062INVISIBLE TIMESZero-width / invisibleMediumRemovedInvisible multiplication sign. Same story: meaningful in MathML, invisible corruption in plain text. Equation editors and scientific PDFs.
U+2063INVISIBLE SEPARATORZero-width / invisibleMediumRemovedInvisible comma used between mathematical indices. Splits tokens for anything that parses the text. Equation editors and scientific PDFs.
U+2064INVISIBLE PLUSZero-width / invisibleMediumRemovedInvisible addition sign used in mixed numbers. Silent in prose, disruptive in parsers. Equation editors and scientific PDFs.
U+FEFFZERO WIDTH NO-BREAK SPACE (BOM)Zero-width / invisibleMediumRemovedThe byte order mark. At the start of a file it is metadata; anywhere else it is an invisible character that corrupts JSON parsing, CSV headers and shell scripts. Files saved as "UTF-8 with BOM" by Windows editors and Excel exports.
U+180EMONGOLIAN VOWEL SEPARATORZero-width / invisibleMediumKept or removed depending on contextReclassified as a zero-width format character in Unicode 6.3. Legitimate inside Mongolian text, invisible padding anywhere else. Mongolian typing; otherwise deliberate padding, because it long passed as a "space" in username filters.
U+034FCOMBINING GRAPHEME JOINERZero-width / invisibleMediumRemovedAffects collation and combining-mark ordering without rendering anything. In ordinary text it only defeats string comparison. Linguistic and library data; occasionally used to defeat naive deduplication.
U+115FHANGUL CHOSEONG FILLERZero-width / invisibleMediumKept or removed depending on contextA placeholder for a missing Korean initial consonant. Renders as blank space and is widely used to fake empty nicknames. Korean text composition; also a classic "invisible username" trick.
U+1160HANGUL JUNGSEONG FILLERZero-width / invisibleMediumKept or removed depending on contextA placeholder for a missing Korean vowel. Blank on screen, real in the byte stream. Korean text composition; also used for blank display names.
U+3164HANGUL FILLERZero-width / invisibleMediumKept or removed depending on contextRenders as an empty box-width blank. The single most common character behind "invisible" chat messages and profile names. Korean input methods; deliberate use in games, chat apps and forums.
U+FFA0HALFWIDTH HANGUL FILLERZero-width / invisibleMediumKept or removed depending on contextThe halfwidth form of U+3164, with the same blank rendering and the same abuse pattern. Korean input methods and invisible-name tricks.
U+17B4KHMER VOWEL INHERENT AQZero-width / invisibleMediumKept or removed depending on contextNormally invisible. It is a genuine part of Khmer orthography, so deleting it blindly damages Khmer text. Khmer typing; outside Khmer text it is padding.
U+17B5KHMER VOWEL INHERENT AAZero-width / invisibleMediumKept or removed depending on contextSame as U+17B4: invisible, legitimate in Khmer, junk elsewhere. Khmer typing; outside Khmer text it is padding.
U+2800BRAILLE PATTERN BLANKZero-width / invisibleMediumKept or removed depending on contextA braille cell with no raised dots. It is a real, meaningful character in braille content, and a wide invisible blank everywhere else. Braille text and ASCII-art; also used to pad messages past length filters.
U+00ADSOFT HYPHENZero-width / invisibleMediumRemovedShows a hyphen only if the word wraps. Invisible in the middle of a word, yet it defeats exact search, spellcheck and equality checks. Word processors, hyphenation tools and PDF-to-text conversion.
U+206AINHIBIT SYMMETRIC SWAPPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control. No modern renderer needs it; its presence is a sign of legacy conversion or deliberate padding. Legacy document conversion.
U+206BACTIVATE SYMMETRIC SWAPPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control with no rendering effect in current software. Legacy document conversion.
U+206CINHIBIT ARABIC FORM SHAPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion.
U+206DACTIVATE ARABIC FORM SHAPINGZero-width / invisibleMediumRemovedDeprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion.
U+206ENATIONAL DIGIT SHAPESZero-width / invisibleMediumRemovedDeprecated control that switched digit rendering. Invisible, and ignored by modern text stacks. Legacy document conversion.
U+206FNOMINAL DIGIT SHAPESZero-width / invisibleMediumRemovedDeprecated counterpart to U+206E. Invisible and inert. Legacy document conversion.
U+061CARABIC LETTER MARKBidirectional controlHighRemovedAn invisible character with strong right-to-left directionality. It changes how neighbouring text is ordered on screen without changing the stored string. Arabic and Persian layout fixes; also used to disguise reordering.
U+200ELEFT-TO-RIGHT MARKBidirectional controlHighRemovedForces left-to-right ordering for the text around it. Invisible, and it makes what you read differ from what is stored. Mixed-direction documents, spreadsheets and localisation files.
U+200FRIGHT-TO-LEFT MARKBidirectional controlHighRemovedThe mirror of U+200E. It can silently reverse the apparent order of adjacent characters. Mixed-direction documents; also the classic filename-spoofing character.
U+202ALEFT-TO-RIGHT EMBEDDINGBidirectional controlHighRemovedOpens a directional embedding that persists until it is popped. Unbalanced embeddings scramble the display of everything that follows. Legacy bidi markup; deprecated in favour of isolates.
U+202BRIGHT-TO-LEFT EMBEDDINGBidirectional controlHighRemovedOpens a right-to-left embedding. Same unbalanced-state risk as U+202A. Legacy bidi markup; deprecated in favour of isolates.
U+202CPOP DIRECTIONAL FORMATTINGBidirectional controlHighRemovedCloses the most recent embedding or override. Harmless alone, but it is half of every bidi attack sequence. Legacy bidi markup.
U+202DLEFT-TO-RIGHT OVERRIDEBidirectional controlHighRemovedForces every following character to render left-to-right regardless of its own direction. One of the two characters that can make code display differently from how it runs. Almost never legitimate in modern text. Treat with suspicion in source code.
U+202ERIGHT-TO-LEFT OVERRIDEBidirectional controlHighRemovedReverses the rendered order of everything after it. This is the character behind reversed-extension filename tricks and misleading code display. Almost never legitimate. Its presence in code or a filename is a red flag.
U+2066LEFT-TO-RIGHT ISOLATEBidirectional controlHighRemovedOpens an isolated left-to-right run. The modern, safer bidi mechanism, but still invisible state that must be popped. Correct bidi markup in localisation files; also seen in code-display research demos.
U+2067RIGHT-TO-LEFT ISOLATEBidirectional controlHighRemovedOpens an isolated right-to-left run, invisibly changing display order for the text it wraps. Localisation files; also seen in code-display research demos.
U+2068FIRST STRONG ISOLATEBidirectional controlHighRemovedOpens an isolate whose direction is inferred from its first strong character, so its effect depends on the data inside it. Localisation frameworks handling user-supplied names.
U+2069POP DIRECTIONAL ISOLATEBidirectional controlHighRemovedCloses an isolate. Like U+202C, it is the closing half of a sequence that can rewrite what a reader sees. Localisation frameworks.
U+00A0NO-BREAK SPACEUnusual spaceMediumReplaced with a normal spaceLooks exactly like a normal space but is a different character, so exact matches, CSV columns and code indentation quietly fail. Word processors, web pages using  , and PDF text extraction. The most common hidden character of all.
U+1680OGHAM SPACE MARKUnusual spaceMediumReplaced with a normal spaceA space character from the Ogham block that may render as a visible line in some fonts. Deliberate substitution; effectively never typed by accident.
U+2000EN QUADUnusual spaceMediumReplaced with a normal spaceA fixed-width typographic space. Indistinguishable from a normal space in most UI fonts. Typesetting software and PDF extraction.
U+2001EM QUADUnusual spaceMediumReplaced with a normal spaceA fixed-width typographic space one em wide. Typesetting software and PDF extraction.
U+2002EN SPACEUnusual spaceMediumReplaced with a normal spaceHalf-em typographic space that passes for an ordinary space on screen. Typesetting software, InDesign and PDF extraction.
U+2003EM SPACEUnusual spaceMediumReplaced with a normal spaceFull-em typographic space, frequently mistaken for a tab or double space. Typesetting software and PDF extraction.
U+2004THREE-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceFixed-width space used in fine typesetting; breaks whitespace-sensitive parsing. Typesetting software and PDF extraction.
U+2005FOUR-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceFixed-width space used in fine typesetting. Typesetting software and PDF extraction.
U+2006SIX-PER-EM SPACEUnusual spaceMediumReplaced with a normal spaceVery narrow fixed-width space, visually identical to a thin gap. Typesetting software and PDF extraction.
U+2007FIGURE SPACEUnusual spaceMediumReplaced with a normal spaceA non-breaking space as wide as a digit, used to align numbers in tables. It reliably breaks numeric parsing. Financial documents, spreadsheets and generated reports.
U+2008PUNCTUATION SPACEUnusual spaceMediumReplaced with a normal spaceA space as wide as a period, used for alignment. Typesetting software.
U+2009THIN SPACEUnusual spaceMediumReplaced with a normal spaceNarrow space used around dashes and units. Very common in text copied from well-typeset web pages. Web typography, Wikipedia, scientific writing.
U+200AHAIR SPACEUnusual spaceMediumReplaced with a normal spaceThe narrowest space in Unicode. Often invisible in practice, yet it still splits words for search. Web typography and typesetting software.
U+202FNARROW NO-BREAK SPACEUnusual spaceMediumReplaced with a normal spaceNon-breaking and narrow. French typography and many number formats use it, so it arrives with copied prices and dates. French locale formatting, spreadsheets and currency output.
U+205FMEDIUM MATHEMATICAL SPACEUnusual spaceMediumReplaced with a normal spaceSpacing used around mathematical operators. Invisible as anything but a space. MathML and equation editors.
U+3000IDEOGRAPHIC SPACEUnusual spaceMediumReplaced with a normal spaceThe full-width space used in Chinese, Japanese and Korean typing. Legitimate in CJK text, and a silent mismatch in code and identifiers. CJK input methods, where it is the space produced in full-width mode.
U+2028LINE SEPARATORLine separatorMediumConverted to a line breakA line break that is not a newline. It terminates lines in text engines but historically broke JSON and JavaScript string literals. Word processors, especially shift-enter line breaks exported to plain text.
U+2029PARAGRAPH SEPARATORLine separatorMediumConverted to a line breakA paragraph break that most tools do not recognise as a newline, so paragraphs collapse into one line. Word processors and rich-text exports.
U+0085NEXT LINE (NEL)Line separatorMediumConverted to a line breakA legacy one-character line break combining CR and LF. Most parsers do not treat it as a newline, so lines silently merge. Mainframe (EBCDIC) exports and Windows-1252 data mis-decoded as UTF-8, where 0x85 is an ellipsis.
U+000DCARRIAGE RETURN (WITHOUT LINE FEED)Line separatorMediumConverted to a line breakA lone CR with no LF after it. Old-Mac line endings and injection tricks use it; many tools show one line while parsers see two, or vice versa. Classic Mac OS files, HTTP header injection payloads and broken concatenation of mixed line endings.
U+0000–U+0008C0 CONTROL CHARACTERS (NUL to BACKSPACE)Control characterHighRemovedTeletype-era control codes. A NUL byte truncates strings in C-based software, and the rest have no meaning in text. Binary data pasted as text, corrupted exports and malformed database dumps.
U+000B–U+000CLINE TABULATION, FORM FEEDControl characterHighRemovedLegacy page and vertical-tab controls. Rendered inconsistently and treated as whitespace by some parsers and not others. Printer output, mainframe exports and old text files.
U+000E–U+001FC0 CONTROL CHARACTERS (SHIFT OUT to UNIT SEPARATOR)Control characterHighRemovedDevice and record-separator controls. Invisible in most editors, and they corrupt CSV, TSV and log parsing. Legacy data interchange formats and terminal captures.
U+007F–U+009FDELETE and C1 CONTROL CHARACTERSControl characterHighRemovedDEL plus the C1 block. Common in text mis-decoded from Windows-1252, where they masquerade as punctuation. Encoding mismatches, especially Windows-1252 data read as UTF-8.
U+E0000–U+E007FTAG CHARACTERSTag characterHighKept or removed depending on contextAn invisible mirror of ASCII. A full sentence can be encoded in tag characters and rendered as nothing at all, which makes this block the standard vehicle for prompt-injection smuggling. The one legitimate use is regional flag emoji. Deprecated language tagging, emoji flag sequences, and deliberate hidden payloads.
U+FE00–U+FE0FVARIATION SELECTORS VS1 to VS16Variation selectorMediumKept or removed depending on contextSelects a glyph variant of the preceding character. VS16 (U+FE0F) is what turns a monochrome symbol into a colour emoji, so removing it blindly destroys emoji. Emoji presentation, keycap sequences, and CJK glyph variants.
U+E0100–U+E01EFVARIATION SELECTORS SUPPLEMENT VS17 to VS256Variation selectorMediumKept or removed depending on contextIdeographic variation selectors that pick a specific glyph for a CJK character. Legitimate after an ideograph, and pure payload space after anything else. Japanese personal and place names; also used to hide data after non-CJK characters.

Pour aller plus loin