Hidden characters in text
Invisible Unicode characters are the most common cause of the bug where two identical-looking strings refuse to match. This page covers where they come from, what each class breaks, and how to get rid of them without collateral damage.
More: what was found, x-ray view and options
X-ray view
Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.
What was found
| Code point | Character | Count | Action |
|---|
What counts as a hidden character
Four groups, with different origins and different levels of danger. The invisible Unicode list documents every individual code point; this is the map.
1. Zero-width and invisible formatting
Characters with no visual width at all: the zero-width space U+200B, the zero-width no-break space U+FEFF, the word joiner U+2060, the soft hyphen U+00AD, and the invisible mathematical operators. They occupy a position in the string and contribute nothing to what you see.
2. Bidirectional controls
Marks, embeddings, overrides and isolates that tell the text renderer how to order characters for display. They are the only class where what you read can differ from what is stored, which is why they are treated as high risk here. See invisible characters in code for what that enables.
3. Unusual spaces
Roughly twenty characters that look like a space and are not one: the no-break space U+00A0 above all, plus thin, hair, figure, punctuation and ideographic spaces. These are the most frequent finds in real documents, by a wide margin.
4. Tag characters and control characters
The tag block U+E0000 to U+E007F is an invisible mirror of ASCII: an entire paragraph can be encoded in it and render as nothing whatsoever. C0 and C1 control characters are teletype-era codes that survive in text through binary pastes and encoding mismatches.
Where they actually come from
| Source | Signature you will see |
|---|---|
| PDF text extraction | Soft hyphens inside words, thin and hair spaces between them, line separators instead of newlines |
| Word processors | No-break spaces from autocorrect, soft hyphens from hyphenation, curly quotes, U+2028 from shift-enter breaks |
| Spreadsheets | No-break and narrow no-break spaces inside numbers, figure spaces used for column alignment |
| Web pages | Non-breaking spaces from , thin spaces around dashes, zero-width spaces inserted by CMS line-break logic |
| Chat and AI interfaces | Whatever typography the page renders, carried along by the copy |
| Localisation files | Bidirectional marks and isolates wrapping user-supplied names |
| Encoding mistakes | C1 control characters, from Windows-1252 bytes read as UTF-8 |
| Windows text editors | A byte order mark at the start of a file saved as "UTF-8 with BOM" |
Notice what is not on that list: language models. Text generated by a model is text; the invisible characters arrive when it is rendered into a page and copied out again. That is why finding a zero-width space tells you a document has been through a browser, not who wrote it.
What they break
String equality
The defining failure. "invoice" and "invoice" with a zero-width space between the n and the v are different strings, so a database lookup returns nothing, a deduplication pass creates a second record, and a cache key never hits. Nobody sees the difference until they check byte lengths.
Regular expressions
Word boundaries stop being where you think they are. \bword\b does not match across an invisible separator, character classes do not include the exotic spaces, and \s in some engines does not cover U+00A0 or U+3000. Anchored patterns fail when a BOM sits at the start of the input.
CSV, TSV and data imports
A no-break space inside a numeric field turns a number into text. A figure space used for alignment survives the export. A control character terminates a field early. All three produce imports that succeed loudly and are wrong quietly.
Authentication and forms
A password copied from a document with a trailing no-break space is not the password. Validation that trims \s often does not trim U+00A0, so the value looks trimmed and is not. Email fields reject perfectly valid addresses that carry an invisible character.
Source code
Bidirectional controls make code render differently from how it compiles, covered in full on the invisible characters in code page. Zero-width characters inside identifiers create two variables that look identical. A BOM at the top of a shell script breaks the shebang line.
How to fix them properly
The naive fix is a regular expression that deletes everything invisible. It is worse than the problem, in three specific ways:
- Deleting spaces glues words together. Every unusual space must be replaced with U+0020, never removed. Delete a no-break space and
1 200becomes1200; do that inside a sentence and two words merge. - Deleting zero-width joiners destroys emoji. A family emoji is several people held together by invisible joiners. Remove them and the glyph falls apart into its components.
- Deleting non-joiners damages other scripts. The zero-width non-joiner is grammatically required in Persian and Urdu, and used throughout Indic scripts. It is not noise there; it is orthography.
A correct cleaner works code point by code point and decides with context:
- Iterate over code points, not UTF-16 units, so characters outside the Basic Multilingual Plane are handled as single units rather than surrogate halves.
- Replace every unusual space with U+0020. Convert U+2028 and U+2029 to a newline.
- Remove bidirectional controls, control characters and stray tag characters.
- For joiners, non-joiners, variation selectors and script-specific invisibles, test the neighbouring characters and keep the ones that are doing real work.
- Keep the tag characters that form a valid flag emoji sequence, closed by U+E007F.
That is exactly what the checker on this page does, and the x-ray view exists so you can watch it make each decision rather than trusting the output blindly.
How to stop them coming back
- Normalise at the boundary. Clean text as it enters your system, on form submission, on import, on paste, not later when a bug report arrives.
- Save files as UTF-8 without a BOM. Every modern editor offers the choice explicitly.
- Configure your editor to render invisible characters. Most support it; almost nobody turns it on.
- In code review, treat any diff touching a line with no visible change as a red flag.
- Compare byte lengths, not appearances, when two strings that should match do not.
Common questions
How do I find a hidden character without a tool?
Compare the visible length of the string with its actual length in code points, or open the text in an editor with invisible-character rendering enabled. In a terminal, piping the text through a hex dump shows every byte, which makes a stray U+00A0 or a BOM obvious immediately.
Is a non-breaking space a hidden character?
Yes, in the sense that matters: it is visually identical to an ordinary space and is a completely different character. It is by far the most common one found in real documents, and it is the usual cause of a spreadsheet number that refuses to behave like a number.
Do hidden characters mean my text was written by AI?
No. They indicate that the text has been copied between applications, through a PDF, a word processor, a spreadsheet or a web page. They carry no information about authorship whatsoever.
Which invisible characters are actually dangerous?
Bidirectional controls, because they change display order and can make code or a filename read differently from what it is. Tag characters come next, since an entire hidden message can be encoded in them. The rest are disruptive rather than dangerous.
Can I just strip everything that is not ASCII?
Only if your text will never contain a non-English name, a currency symbol, an emoji or an accented word. For anything user-facing it is destructive. Targeted, context-aware cleaning is not much more work and does not break your users' names.
Go deeper
- The invisible Unicode list - every code point this checker detects, with its rule.
- Invisible characters in code - how bidirectional overrides mislead code review, with a safe demo.
- Clean hidden characters out of AI text - the cleaner, with its scope stated plainly.
- Statistical versus character watermarks - why these characters are not watermarks.