Why AI-generated code breaks, and how to fix the real cause
A linter throws an error on a line that reads perfectly. A JSON payload fails validation on a field that matches the schema exactly, same length, same characters, same everything. The bug usually isn't in your logic. It's sitting between two letters, invisible, and it got there because you pasted code out of a chat window.
Text copied from an AI chat interface often carries invisible Unicode characters, zero-width spaces, joiners, non-breaking spaces, left over from how the interface renders its answer as typeset HTML. They aren't a watermark and nobody put them there on purpose. They're also completely invisible in a normal editor, which is why they cause syntax errors, corrupted string lengths and failed API calls that look impossible to explain. Paste the code into the checker below and every one of them shows up.
More: what was found, x-ray view and options
X-ray view
Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.
What was found
| Code point | Character | Count | Action |
|---|
What these characters actually are
Real, standard Unicode code points, not malware and not a deliberate signature. A zero-width joiner fuses two emoji into one combined glyph. A zero-width non-joiner keeps letters from merging in Persian and several Indic scripts. A word joiner tells a renderer not to break a line at that point. Legitimate jobs, every one, just not jobs that belong inside a variable name or a SQL string.
| Character | Code point | What it's for | What it does to code |
|---|---|---|---|
| Zero-width space | U+200B | Marks a soft line-break point in dense text | Splits identifiers, breaks \bword\b regex matches |
| Zero-width joiner | U+200D | Glues emoji sequences together | Silently changes string length; .length lies to you |
| Word joiner | U+2060 | Prevents an unwanted line break | Corrupts tokens that get diffed or hashed |
| Non-breaking space | U+00A0 | Keeps a number glued to its unit | Looks like a space, isn't one; form and schema validators reject it |
None of these show up in a normal editor. That is the entire problem: your code review, your diff, your own eyes, nothing catches them, because there is nothing visible to catch.
Why they end up in AI output specifically
Not fingerprinting, and not a deliberate signal from the model. A chat interface renders its answer as typeset HTML, the same way any well-built web page does: non-breaking spaces keep units attached to numbers, soft break points sit in long unbroken strings, an occasional joiner keeps text wrapping cleanly inside a chat bubble. Selecting and copying the rendered answer copies that typography along with the words. It is a side effect of how the text reached your screen, not a mark the model chose to leave. The hidden characters guide covers where else this debris comes from, PDFs, word processors, spreadsheet exports, since chat interfaces are one source among several.
It is worth being precise here, because a real statistical watermark, the kind Anthropic and Google have started shipping this year, works nothing like this. That signal lives in which words the model chose, not in any character you could select. See statistical versus character watermarks for the full mechanism. What breaks your code is a separate, far more mundane problem, and unlike a statistical watermark, it is fully fixable.
Where this actually bites you
A linter fails on a perfect line
Usually a zero-width space wedged inside an identifier, splitting one token into two as far as the parser is concerned.
A CSV or database import corrupts one row
A stray character turns a clean numeric field into text, or pushes a value past a length constraint nobody set on purpose.
An API call returns 400 with no visible cause
Strict JSON validators aren't forgiving about unexpected code points inside a string value, and the request body looks completely normal in every tool you check it with.
Code review misses something it shouldn't
Bidirectional control characters can make a line display in an order it doesn't execute in. That's a distinct, more serious class covered in full in invisible characters in code.
Hunting them yourself
If you'd rather see it before trusting a tool, the short version in JavaScript:
const suspects = /[]/;
if (suspects.test(yourString)) {
console.log('Found something invisible in there.');
}
That catches the common ones. It also flags a zero-width joiner sitting correctly inside an emoji sequence, a false positive you would need to sort out by hand. Fine for a one-off check, less fine as something you paste into production without thinking about what it might mangle, which is the whole reason the checker above classifies by context instead of deleting on sight.
Questions about hidden characters in AI-generated code
Are these characters an AI watermark?
No. No major provider uses invisible characters as a watermark, it would be trivially removed by any text-cleaning step, which makes it useless for that purpose. They're typographic debris from how a chat interface renders its answer, not a deliberate signal.
Will cleaning break my code's formatting?
No. Tabs, line feeds and indentation are preserved exactly. Only invisible code points that carry no visible meaning are touched, and unusual spaces are replaced with an ordinary space rather than deleted.
Does this happen with every AI tool, or just one?
Any chat interface that renders typeset HTML can leave this behind, it isn't specific to one vendor. Whether the vendor also applies a real statistical watermark is a separate question with a different answer per provider.
Why not just delete every invisible character automatically?
Because some of them are doing real work: zero-width joiners inside emoji, zero-width non-joiners required in Persian and Indic scripts, blank braille cells that are meaningful content. Deleting on sight breaks those. The checker above checks context first.
Related reading
- The checker - paste code and see this in action.
- Invisible characters in code - the bidirectional-control class in full, with a safe live demo.
- How to remove AI watermarks from your text - the broader guide, including which vendors have a confirmed statistical watermark today.
- Hidden characters in text - every source of this debris, not just AI chat interfaces.