Elenco dei caratteri Unicode invisibili
Questo è l'elenco di riferimento dei caratteri invisibili e ambigui che lo strumento rileva: spazi a larghezza zero, controlli bidirezionali, caratteri tag, spazi insoliti, separatori di riga e caratteri di controllo. Ogni riga indica il punto di codice, la categoria, il rischio e cosa ne fa il pulitore.
Altro: cosa è stato trovato, vista a raggi X e opzioni
Vista a raggi X
Il tuo testo con ogni carattere nascosto esposto come etichetta. Passa il mouse per il nome Unicode.
Cosa è stato trovato
| Punto di codice | Carattere | Conteggio | Azione |
|---|
Incolla qualsiasi testo nello strumento sopra per vedere quali di questi caratteri contiene e cosa ne fa il pulitore. La tabella seguente è generata dallo stesso file di dati usato dallo strumento, quindi corrisponde sempre al suo comportamento. I nomi Unicode sono standard e mostrati in inglese.
57 individually named code points plus 7 ranges. This table is generated from the same data file the checker itself uses, so it always matches what the tool does.
| Code point | Name | Category | Risk | What the cleaner does | Why it matters and where it comes from |
|---|---|---|---|---|---|
| U+200B | ZERO WIDTH SPACE | Zero-width / invisible | Medium | Removed | Occupies no width but is a real character: it breaks string equality, splits words for search and regex, and survives copy-paste invisibly. Word-wrap hints inserted by CMS editors, PDF exporters and web typography scripts. |
| U+200C | ZERO WIDTH NON-JOINER | Zero-width / invisible | Medium | Kept or removed depending on context | Prevents two letters from forming a ligature. Meaningless in Latin text, but grammatically required in Persian, Urdu and several Indic scripts. Persian and Indic typing; also injected by some editors around emphasis markup. |
| U+200D | ZERO WIDTH JOINER | Joiner | Medium | Kept or removed depending on context | Glues code points into one glyph. It is what holds family and profession emoji together, and it is also required in Devanagari and Arabic-script typing. Emoji sequences, Indic and Arabic script text, occasionally editor artefacts. |
| U+2060 | WORD JOINER | Zero-width / invisible | Medium | Removed | A zero-width no-break point. Invisible, but it blocks line breaking and breaks word-boundary matching. Typographic tooling; the modern replacement for a leading U+FEFF. |
| U+2061 | FUNCTION APPLICATION | Zero-width / invisible | Medium | Removed | Invisible mathematical operator. Carries meaning only inside MathML; in prose it is silent noise. Copy-paste out of MathML, LaTeX-to-HTML converters and equation editors. |
| U+2062 | INVISIBLE TIMES | Zero-width / invisible | Medium | Removed | Invisible multiplication sign. Same story: meaningful in MathML, invisible corruption in plain text. Equation editors and scientific PDFs. |
| U+2063 | INVISIBLE SEPARATOR | Zero-width / invisible | Medium | Removed | Invisible comma used between mathematical indices. Splits tokens for anything that parses the text. Equation editors and scientific PDFs. |
| U+2064 | INVISIBLE PLUS | Zero-width / invisible | Medium | Removed | Invisible addition sign used in mixed numbers. Silent in prose, disruptive in parsers. Equation editors and scientific PDFs. |
| U+FEFF | ZERO WIDTH NO-BREAK SPACE (BOM) | Zero-width / invisible | Medium | Removed | The byte order mark. At the start of a file it is metadata; anywhere else it is an invisible character that corrupts JSON parsing, CSV headers and shell scripts. Files saved as "UTF-8 with BOM" by Windows editors and Excel exports. |
| U+180E | MONGOLIAN VOWEL SEPARATOR | Zero-width / invisible | Medium | Kept or removed depending on context | Reclassified as a zero-width format character in Unicode 6.3. Legitimate inside Mongolian text, invisible padding anywhere else. Mongolian typing; otherwise deliberate padding, because it long passed as a "space" in username filters. |
| U+034F | COMBINING GRAPHEME JOINER | Zero-width / invisible | Medium | Removed | Affects collation and combining-mark ordering without rendering anything. In ordinary text it only defeats string comparison. Linguistic and library data; occasionally used to defeat naive deduplication. |
| U+115F | HANGUL CHOSEONG FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | A placeholder for a missing Korean initial consonant. Renders as blank space and is widely used to fake empty nicknames. Korean text composition; also a classic "invisible username" trick. |
| U+1160 | HANGUL JUNGSEONG FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | A placeholder for a missing Korean vowel. Blank on screen, real in the byte stream. Korean text composition; also used for blank display names. |
| U+3164 | HANGUL FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | Renders as an empty box-width blank. The single most common character behind "invisible" chat messages and profile names. Korean input methods; deliberate use in games, chat apps and forums. |
| U+FFA0 | HALFWIDTH HANGUL FILLER | Zero-width / invisible | Medium | Kept or removed depending on context | The halfwidth form of U+3164, with the same blank rendering and the same abuse pattern. Korean input methods and invisible-name tricks. |
| U+17B4 | KHMER VOWEL INHERENT AQ | Zero-width / invisible | Medium | Kept or removed depending on context | Normally invisible. It is a genuine part of Khmer orthography, so deleting it blindly damages Khmer text. Khmer typing; outside Khmer text it is padding. |
| U+17B5 | KHMER VOWEL INHERENT AA | Zero-width / invisible | Medium | Kept or removed depending on context | Same as U+17B4: invisible, legitimate in Khmer, junk elsewhere. Khmer typing; outside Khmer text it is padding. |
| U+2800 | BRAILLE PATTERN BLANK | Zero-width / invisible | Medium | Kept or removed depending on context | A braille cell with no raised dots. It is a real, meaningful character in braille content, and a wide invisible blank everywhere else. Braille text and ASCII-art; also used to pad messages past length filters. |
| U+00AD | SOFT HYPHEN | Zero-width / invisible | Medium | Removed | Shows a hyphen only if the word wraps. Invisible in the middle of a word, yet it defeats exact search, spellcheck and equality checks. Word processors, hyphenation tools and PDF-to-text conversion. |
| U+206A | INHIBIT SYMMETRIC SWAPPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control. No modern renderer needs it; its presence is a sign of legacy conversion or deliberate padding. Legacy document conversion. |
| U+206B | ACTIVATE SYMMETRIC SWAPPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control with no rendering effect in current software. Legacy document conversion. |
| U+206C | INHIBIT ARABIC FORM SHAPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion. |
| U+206D | ACTIVATE ARABIC FORM SHAPING | Zero-width / invisible | Medium | Removed | Deprecated Unicode format control, superseded by ZWJ and ZWNJ. Legacy Arabic document conversion. |
| U+206E | NATIONAL DIGIT SHAPES | Zero-width / invisible | Medium | Removed | Deprecated control that switched digit rendering. Invisible, and ignored by modern text stacks. Legacy document conversion. |
| U+206F | NOMINAL DIGIT SHAPES | Zero-width / invisible | Medium | Removed | Deprecated counterpart to U+206E. Invisible and inert. Legacy document conversion. |
| U+061C | ARABIC LETTER MARK | Bidirectional control | High | Removed | An invisible character with strong right-to-left directionality. It changes how neighbouring text is ordered on screen without changing the stored string. Arabic and Persian layout fixes; also used to disguise reordering. |
| U+200E | LEFT-TO-RIGHT MARK | Bidirectional control | High | Removed | Forces left-to-right ordering for the text around it. Invisible, and it makes what you read differ from what is stored. Mixed-direction documents, spreadsheets and localisation files. |
| U+200F | RIGHT-TO-LEFT MARK | Bidirectional control | High | Removed | The mirror of U+200E. It can silently reverse the apparent order of adjacent characters. Mixed-direction documents; also the classic filename-spoofing character. |
| U+202A | LEFT-TO-RIGHT EMBEDDING | Bidirectional control | High | Removed | Opens a directional embedding that persists until it is popped. Unbalanced embeddings scramble the display of everything that follows. Legacy bidi markup; deprecated in favour of isolates. |
| U+202B | RIGHT-TO-LEFT EMBEDDING | Bidirectional control | High | Removed | Opens a right-to-left embedding. Same unbalanced-state risk as U+202A. Legacy bidi markup; deprecated in favour of isolates. |
| U+202C | POP DIRECTIONAL FORMATTING | Bidirectional control | High | Removed | Closes the most recent embedding or override. Harmless alone, but it is half of every bidi attack sequence. Legacy bidi markup. |
| U+202D | LEFT-TO-RIGHT OVERRIDE | Bidirectional control | High | Removed | Forces every following character to render left-to-right regardless of its own direction. One of the two characters that can make code display differently from how it runs. Almost never legitimate in modern text. Treat with suspicion in source code. |
| U+202E | RIGHT-TO-LEFT OVERRIDE | Bidirectional control | High | Removed | Reverses the rendered order of everything after it. This is the character behind reversed-extension filename tricks and misleading code display. Almost never legitimate. Its presence in code or a filename is a red flag. |
| U+2066 | LEFT-TO-RIGHT ISOLATE | Bidirectional control | High | Removed | Opens an isolated left-to-right run. The modern, safer bidi mechanism, but still invisible state that must be popped. Correct bidi markup in localisation files; also seen in code-display research demos. |
| U+2067 | RIGHT-TO-LEFT ISOLATE | Bidirectional control | High | Removed | Opens an isolated right-to-left run, invisibly changing display order for the text it wraps. Localisation files; also seen in code-display research demos. |
| U+2068 | FIRST STRONG ISOLATE | Bidirectional control | High | Removed | Opens an isolate whose direction is inferred from its first strong character, so its effect depends on the data inside it. Localisation frameworks handling user-supplied names. |
| U+2069 | POP DIRECTIONAL ISOLATE | Bidirectional control | High | Removed | Closes an isolate. Like U+202C, it is the closing half of a sequence that can rewrite what a reader sees. Localisation frameworks. |
| U+00A0 | NO-BREAK SPACE | Unusual space | Medium | Replaced with a normal space | Looks exactly like a normal space but is a different character, so exact matches, CSV columns and code indentation quietly fail. Word processors, web pages using , and PDF text extraction. The most common hidden character of all. |
| U+1680 | OGHAM SPACE MARK | Unusual space | Medium | Replaced with a normal space | A space character from the Ogham block that may render as a visible line in some fonts. Deliberate substitution; effectively never typed by accident. |
| U+2000 | EN QUAD | Unusual space | Medium | Replaced with a normal space | A fixed-width typographic space. Indistinguishable from a normal space in most UI fonts. Typesetting software and PDF extraction. |
| U+2001 | EM QUAD | Unusual space | Medium | Replaced with a normal space | A fixed-width typographic space one em wide. Typesetting software and PDF extraction. |
| U+2002 | EN SPACE | Unusual space | Medium | Replaced with a normal space | Half-em typographic space that passes for an ordinary space on screen. Typesetting software, InDesign and PDF extraction. |
| U+2003 | EM SPACE | Unusual space | Medium | Replaced with a normal space | Full-em typographic space, frequently mistaken for a tab or double space. Typesetting software and PDF extraction. |
| U+2004 | THREE-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Fixed-width space used in fine typesetting; breaks whitespace-sensitive parsing. Typesetting software and PDF extraction. |
| U+2005 | FOUR-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Fixed-width space used in fine typesetting. Typesetting software and PDF extraction. |
| U+2006 | SIX-PER-EM SPACE | Unusual space | Medium | Replaced with a normal space | Very narrow fixed-width space, visually identical to a thin gap. Typesetting software and PDF extraction. |
| U+2007 | FIGURE SPACE | Unusual space | Medium | Replaced with a normal space | A non-breaking space as wide as a digit, used to align numbers in tables. It reliably breaks numeric parsing. Financial documents, spreadsheets and generated reports. |
| U+2008 | PUNCTUATION SPACE | Unusual space | Medium | Replaced with a normal space | A space as wide as a period, used for alignment. Typesetting software. |
| U+2009 | THIN SPACE | Unusual space | Medium | Replaced with a normal space | Narrow space used around dashes and units. Very common in text copied from well-typeset web pages. Web typography, Wikipedia, scientific writing. |
| U+200A | HAIR SPACE | Unusual space | Medium | Replaced with a normal space | The narrowest space in Unicode. Often invisible in practice, yet it still splits words for search. Web typography and typesetting software. |
| U+202F | NARROW NO-BREAK SPACE | Unusual space | Medium | Replaced with a normal space | Non-breaking and narrow. French typography and many number formats use it, so it arrives with copied prices and dates. French locale formatting, spreadsheets and currency output. |
| U+205F | MEDIUM MATHEMATICAL SPACE | Unusual space | Medium | Replaced with a normal space | Spacing used around mathematical operators. Invisible as anything but a space. MathML and equation editors. |
| U+3000 | IDEOGRAPHIC SPACE | Unusual space | Medium | Replaced with a normal space | The full-width space used in Chinese, Japanese and Korean typing. Legitimate in CJK text, and a silent mismatch in code and identifiers. CJK input methods, where it is the space produced in full-width mode. |
| U+2028 | LINE SEPARATOR | Line separator | Medium | Converted to a line break | A line break that is not a newline. It terminates lines in text engines but historically broke JSON and JavaScript string literals. Word processors, especially shift-enter line breaks exported to plain text. |
| U+2029 | PARAGRAPH SEPARATOR | Line separator | Medium | Converted to a line break | A paragraph break that most tools do not recognise as a newline, so paragraphs collapse into one line. Word processors and rich-text exports. |
| U+0085 | NEXT LINE (NEL) | Line separator | Medium | Converted to a line break | A legacy one-character line break combining CR and LF. Most parsers do not treat it as a newline, so lines silently merge. Mainframe (EBCDIC) exports and Windows-1252 data mis-decoded as UTF-8, where 0x85 is an ellipsis. |
| U+000D | CARRIAGE RETURN (WITHOUT LINE FEED) | Line separator | Medium | Converted to a line break | A lone CR with no LF after it. Old-Mac line endings and injection tricks use it; many tools show one line while parsers see two, or vice versa. Classic Mac OS files, HTTP header injection payloads and broken concatenation of mixed line endings. |
| U+0000–U+0008 | C0 CONTROL CHARACTERS (NUL to BACKSPACE) | Control character | High | Removed | Teletype-era control codes. A NUL byte truncates strings in C-based software, and the rest have no meaning in text. Binary data pasted as text, corrupted exports and malformed database dumps. |
| U+000B–U+000C | LINE TABULATION, FORM FEED | Control character | High | Removed | Legacy page and vertical-tab controls. Rendered inconsistently and treated as whitespace by some parsers and not others. Printer output, mainframe exports and old text files. |
| U+000E–U+001F | C0 CONTROL CHARACTERS (SHIFT OUT to UNIT SEPARATOR) | Control character | High | Removed | Device and record-separator controls. Invisible in most editors, and they corrupt CSV, TSV and log parsing. Legacy data interchange formats and terminal captures. |
| U+007F–U+009F | DELETE and C1 CONTROL CHARACTERS | Control character | High | Removed | DEL plus the C1 block. Common in text mis-decoded from Windows-1252, where they masquerade as punctuation. Encoding mismatches, especially Windows-1252 data read as UTF-8. |
| U+E0000–U+E007F | TAG CHARACTERS | Tag character | High | Kept or removed depending on context | An invisible mirror of ASCII. A full sentence can be encoded in tag characters and rendered as nothing at all, which makes this block the standard vehicle for prompt-injection smuggling. The one legitimate use is regional flag emoji. Deprecated language tagging, emoji flag sequences, and deliberate hidden payloads. |
| U+FE00–U+FE0F | VARIATION SELECTORS VS1 to VS16 | Variation selector | Medium | Kept or removed depending on context | Selects a glyph variant of the preceding character. VS16 (U+FE0F) is what turns a monochrome symbol into a colour emoji, so removing it blindly destroys emoji. Emoji presentation, keycap sequences, and CJK glyph variants. |
| U+E0100–U+E01EF | VARIATION SELECTORS SUPPLEMENT VS17 to VS256 | Variation selector | Medium | Kept or removed depending on context | Ideographic variation selectors that pick a specific glyph for a CJK character. Legitimate after an ideograph, and pure payload space after anything else. Japanese personal and place names; also used to hide data after non-CJK characters. |