Em dash remover
Text from a chat model arrives with em dashes where most people would have typed a comma or a hyphen. They look correct on screen, then break a CSV import, an identifier or an exact match. Paste the text below to see every one, and copy a version with ordinary hyphens instead.
More: what was found, x-ray view and options
X-ray view
Your text with every hidden character exposed as a labelled chip. Hover a chip for its Unicode name.
What was found
| Code point | Character | Count | Action |
|---|
Why AI text is full of em dashes
Chat models were trained on edited prose: books, journalism, magazine copy, documentation. In that material the em dash is ordinary punctuation, and a model that has absorbed it reaches for one wherever a sentence needs a break stronger than a comma. People typing at a keyboard mostly do not, because on most layouts the em dash is not on a key. The result is that AI-assisted text carries em dashes at a rate that reads as unusual, not because anything was inserted on purpose, but because the model writes like the edited prose it learned from and you write like someone using a keyboard.
Chat interfaces make it more visible. They render answers as formatted web text, so a dash that a model produced as U+2014 arrives in your clipboard as U+2014, not as the two hyphens a word processor would have turned into a dash under your control.
An em dash is not a watermark
This needs saying plainly, because a lot of pages in this niche imply otherwise. An em dash is a visible, standard punctuation mark documented in every style guide in use. It carries no hidden information, no identifier, and no signal that a model put it there. Counting em dashes tells you about writing style, in the same way that counting semicolons does, and nothing more.
Removing them does not change what AI-detection software concludes about your text. Style-based detectors work on word choice and sentence structure across a whole passage; punctuation frequency is at most one weak input among hundreds, and none of them decide anything on a dash. Statistical watermarks such as SynthID live in word choice itself and are untouched by punctuation of any kind. If what you are actually looking for is a way to hide that a model helped write something, this page will not do it, and neither will the rest of this site.
What replacing em dashes does do is fix concrete, boring, technical breakage, which is the rest of this page.
What an em dash actually breaks
Spreadsheets and CSV imports
The dash itself is valid UTF-8 and imports cleanly when the encoding is declared. The failure comes when it is not: a CSV written as UTF-8 and opened by a spreadsheet that assumes Windows-1252 turns every U+2014 into â€", three visible characters where one was intended. A file that looked correct in your editor arrives at a colleague full of debris, and the column that was supposed to be a clean product name no longer matches anything in the database.
Identifiers, keys and code
Inside a comment or a string literal an em dash is harmless. Anywhere a parser reads, it is a syntax error, and at normal editor font sizes it is very difficult to tell from a hyphen by eye. The characteristic version of this bug is a config key or a command copied out of a chat answer, where --verbose was rendered as an en dash and the program reports an unrecognised option that looks, on screen, exactly like the one in the documentation.
Search and exact matching
Someone searching your page for state-of-the-art does not match state—of—the—art, and a database WHERE clause does not either. Product names, part numbers and version strings are the usual casualties: the record exists, the query is correct, and nothing comes back.
URLs and slugs
Slug generators handle the hyphen because that is what they are built for. An em dash is either dropped, leaving two words fused, or percent-encoded, leaving %E2%80%94 sitting in the middle of a URL that then has to live that way permanently because changing it would break the link.
Terminals, PDFs and older fonts
A fixed-width context that has no glyph for the character shows a box or a question mark. Text extracted back out of a PDF frequently keeps the em dash while losing the spaces around it, which is why a copied paragraph so often comes out with words welded to punctuation.
Em dash, en dash, hyphen and minus
Four characters, routinely confused, and only one of them is on your keyboard.
| Character | Code point | Name | What it is for | This tool |
|---|---|---|---|---|
| — | U+2014 | Em dash | A break in a sentence stronger than a comma | Replaced with - between word characters |
| – | U+2013 | En dash | Ranges: 2010–2020, pp. 40–52 | Replaced with - between word characters |
| - | U+002D | Hyphen-minus | Compound words, the key on your keyboard | Kept, it is the replacement |
| − | U+2212 | Minus sign | Mathematics, typeset subtraction | Left alone, it is not a dash |
How this tool replaces them
The rule is deliberately narrow, and worth knowing before you rely on it. A dash is replaced only when it sits directly between two letters or digits with no space on either side.
word—wordbecomesword-word. This is the form chat models produce most often, and the form that breaks matching.2010–2020becomes2010-2020, which is what a spreadsheet or a parser expects from a range.word — word, with spaces, is left exactly as it is. A spaced dash is doing a punctuation job that a hyphen cannot do, and swapping it would produce something no style guide permits.- A dash at the start or the end of a line is left alone, because there it is a list marker or a dialogue mark, not a join.
The word test is Unicode-aware, not ASCII, so it behaves the same in Cyrillic, Greek, Devanagari and every other script: минуле—майбутнє is joined by a dash between two letters exactly as word—word is, and is treated the same way.
Off by default, and on purpose. Em dashes are visible characters, not hidden ones. The cleaner never touches them unless you tick Em dashes (—) to hyphens in the options, because for most text the correct number of em dashes to remove is zero. The hidden-character cleaning that runs by default is a separate pass and does not change any punctuation you can see.
Smart quotes have the same problem
The second option, Straighten smart quotes, handles the other half of the same story. A chat model writes don't with U+2019, the right single quotation mark, not with the apostrophe on your keyboard. Visually there is almost nothing in it. Functionally it is a different character, and it fails in exactly the same places: string comparison, JSON that was hand-edited, a SQL literal, a filename, a search box.
Ticking the option maps the curly forms back to their straight equivalents: U+2018 and U+2019 and the low-9 and reversed variants all become ', the four double-quote forms become ", and the prime and double prime used in measurements become ' and ". Nothing else about the text changes.
The same warning applies in reverse: if what you are writing is prose for publication, straight quotes are the wrong choice typographically, and you should leave this off. It exists for text going into a parser, not text going into a magazine.
Removing em dashes without this tool
You do not need a website for this, and it is worth knowing the manual route for the cases where pasting is not practical.
Microsoft Word
Open Find and Replace, and in the Find field type ^+ for an em dash or ^= for an en dash. Replace with a hyphen. To stop Word creating them in the first place, turn off AutoFormat As You Type and clear the Hyphens (--) with dash (—) option, which is what converts your typing behind your back.
Google Docs
Find and replace does not have escape codes, so paste the actual character into the search field. To stop it happening again, remove the substitution under Tools, Preferences, Substitutions.
VS Code, Sublime and any regex-capable editor
Switch the search to regular expressions and use [–—], which catches both dashes in one pass. If you want to match the narrow case this tool implements, (?<=[\p{L}\p{N}])[–—](?=[\p{L}\p{N}]) with the Unicode flag enabled does it.
Command line
A single sed call handles a whole directory: sed -i 's/[\xe2\x80\x93\xe2\x80\x94]/-/g' works on UTF-8 input in the GNU version. In Python, text.replace('—', '-').replace('–', '-') is the direct equivalent and does not need the shell to agree about encoding.
Common questions
Does removing em dashes hide that AI helped write my text?
No. Style-based detectors evaluate word choice and sentence construction across a whole passage, not punctuation counts, and the statistical watermarks some providers embed live in word choice itself where punctuation cannot reach them. Replacing dashes changes how text parses, not what a detector concludes.
Why does ChatGPT use so many em dashes?
Because the edited prose it learned from uses them and most people typing do not. It is a property of the training material showing through, not a deliberate mark, and every major chat model does it to some degree.
Is it safe to replace every em dash with a hyphen?
Not blindly. In running prose a spaced em dash and a hyphen are not interchangeable, and a global replacement produces sentences that read as though the punctuation was done by a machine, which is the opposite of what most people want. That is why this tool only replaces dashes that sit between two word characters, where the hyphen is genuinely the correct character.
Will this fix the â€" characters I am seeing?
No, and it is worth knowing why. Those are not dashes, they are a correctly-encoded em dash being read with the wrong encoding, and the fix is to open the file as UTF-8 rather than to edit the text. If you clean the text while it is displayed wrongly you will save the mistake permanently.
Does this change my wording?
No. Every replacement is one punctuation character for another. Words, word order, line breaks and paragraph structure come out exactly as they went in, and nothing is deleted.
Are em dashes wrong?
Not at all. They are correct punctuation and good writing uses them. This page exists because a correct character in prose is still the wrong character in a spreadsheet column, a URL slug or a database key, and the fix is to change it where it is causing damage rather than to avoid it everywhere.
Go deeper
- Hidden characters in text - the invisible ones, which unlike em dashes you cannot see at all.
- The invisible Unicode list - every code point the checker knows about, with what each one does.
- Why AI-assisted text will not format right - the wider set of paste and upload failures.
- Statistical versus character watermarks - why punctuation is not, and cannot be, a watermark.