AI Watermark Check

Hidden characters in text

Invisible Unicode characters are the most common cause of the bug where two identical-looking strings refuse to match. This page covers where they come from, what each class breaks, and how to get rid of them without collateral damage.

More: what was found, x-ray view and options

X-ray view

Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.

What counts as a hidden character

Four groups, with different origins and different levels of danger. The invisible Unicode list documents every individual code point; this is the map.

1. Zero-width and invisible formatting

Characters with no visual width at all: the zero-width space U+200B, the zero-width no-break space U+FEFF, the word joiner U+2060, the soft hyphen U+00AD, and the invisible mathematical operators. They occupy a position in the string and contribute nothing to what you see.

2. Bidirectional controls

Marks, embeddings, overrides and isolates that tell the text renderer how to order characters for display. They are the only class where what you read can differ from what is stored, which is why they are treated as high risk here. See invisible characters in code for what that enables.

3. Unusual spaces

Roughly twenty characters that look like a space and are not one: the no-break space U+00A0 above all, plus thin, hair, figure, punctuation and ideographic spaces. These are the most frequent finds in real documents, by a wide margin.

4. Tag characters and control characters

The tag block U+E0000 to U+E007F is an invisible mirror of ASCII: an entire paragraph can be encoded in it and render as nothing whatsoever. C0 and C1 control characters are teletype-era codes that survive in text through binary pastes and encoding mismatches.

Where they actually come from

SourceSignature you will see
PDF text extractionSoft hyphens inside words, thin and hair spaces between them, line separators instead of newlines
Word processorsNo-break spaces from autocorrect, soft hyphens from hyphenation, curly quotes, U+2028 from shift-enter breaks
SpreadsheetsNo-break and narrow no-break spaces inside numbers, figure spaces used for column alignment
Web pagesNon-breaking spaces from  , thin spaces around dashes, zero-width spaces inserted by CMS line-break logic
Chat and AI interfacesWhatever typography the page renders, carried along by the copy
Localisation filesBidirectional marks and isolates wrapping user-supplied names
Encoding mistakesC1 control characters, from Windows-1252 bytes read as UTF-8
Windows text editorsA byte order mark at the start of a file saved as "UTF-8 with BOM"

Notice what is not on that list: language models. Text generated by a model is text; the invisible characters arrive when it is rendered into a page and copied out again. That is why finding a zero-width space tells you a document has been through a browser, not who wrote it.

What they break

String equality

The defining failure. "invoice" and "invoice" with a zero-width space between the n and the v are different strings, so a database lookup returns nothing, a deduplication pass creates a second record, and a cache key never hits. Nobody sees the difference until they check byte lengths.

Regular expressions

Word boundaries stop being where you think they are. \bword\b does not match across an invisible separator, character classes do not include the exotic spaces, and \s in some engines does not cover U+00A0 or U+3000. Anchored patterns fail when a BOM sits at the start of the input.

CSV, TSV and data imports

A no-break space inside a numeric field turns a number into text. A figure space used for alignment survives the export. A control character terminates a field early. All three produce imports that succeed loudly and are wrong quietly.

Authentication and forms

A password copied from a document with a trailing no-break space is not the password. Validation that trims \s often does not trim U+00A0, so the value looks trimmed and is not. Email fields reject perfectly valid addresses that carry an invisible character.

Source code

Bidirectional controls make code render differently from how it compiles, covered in full on the invisible characters in code page. Zero-width characters inside identifiers create two variables that look identical. A BOM at the top of a shell script breaks the shebang line.

How to fix them properly

The naive fix is a regular expression that deletes everything invisible. It is worse than the problem, in three specific ways:

A correct cleaner works code point by code point and decides with context:

  1. Iterate over code points, not UTF-16 units, so characters outside the Basic Multilingual Plane are handled as single units rather than surrogate halves.
  2. Replace every unusual space with U+0020. Convert U+2028 and U+2029 to a newline.
  3. Remove bidirectional controls, control characters and stray tag characters.
  4. For joiners, non-joiners, variation selectors and script-specific invisibles, test the neighbouring characters and keep the ones that are doing real work.
  5. Keep the tag characters that form a valid flag emoji sequence, closed by U+E007F.

That is exactly what the checker on this page does, and the x-ray view exists so you can watch it make each decision rather than trusting the output blindly.

How to stop them coming back

Common questions

How do I find a hidden character without a tool?

Compare the visible length of the string with its actual length in code points, or open the text in an editor with invisible-character rendering enabled. In a terminal, piping the text through a hex dump shows every byte, which makes a stray U+00A0 or a BOM obvious immediately.

Is a non-breaking space a hidden character?

Yes, in the sense that matters: it is visually identical to an ordinary space and is a completely different character. It is by far the most common one found in real documents, and it is the usual cause of a spreadsheet number that refuses to behave like a number.

Do hidden characters mean my text was written by AI?

No. They indicate that the text has been copied between applications, through a PDF, a word processor, a spreadsheet or a web page. They carry no information about authorship whatsoever.

Which invisible characters are actually dangerous?

Bidirectional controls, because they change display order and can make code or a filename read differently from what it is. Tag characters come next, since an entire hidden message can be encoded in them. The rest are disruptive rather than dangerous.

Can I just strip everything that is not ASCII?

Only if your text will never contain a non-English name, a currency symbol, an emoji or an accented word. For anything user-facing it is destructive. Targeted, context-aware cleaning is not much more work and does not break your users' names.

Go deeper