This text cleaner fixes messy text instantly and entirely in your browser — nothing is uploaded and nothing is stored. Paste text from Word, PDF, email or a web page and strip out the formatting junk that comes with it: extra spaces, stray line breaks, smart (curly) quotes, non-breaking spaces, invisible Unicode, HTML tags and accents. Every change happens live as you type, with a running change log so you can see exactly what was cleaned.
Normalize smart quotes, em/en dashes and ellipses to plain ASCII, and collapse the double spaces and stray line breaks that copying from documents leaves behind.
Remove HTML tags, accents/diacritics, emoji, URLs, email addresses, punctuation or numbers — keeping only the text you actually want.
Change case, wrap or align columns, add or remove line numbers, sort, dedupe, reverse, slugify, and convert tabs ↔ spaces.
Plain or full regular-expression search and replace with case sensitivity, plus a live match counter.
.txt file straight onto it.There is a reason “how do I get rid of the em dashes” became a question in 2025. Language models learned from edited prose, where the em dash is ordinary punctuation, and they reproduce it far more often than people type it. The same goes for curly quotes, the single-character ellipsis, and the narrow no-break space. None of that is a watermark — it is a house style — but it is a tell, and it is also genuinely inconvenient the moment the text has to go into code, a spreadsheet or a CMS.
The invisible characters are a separate story. Zero-width spaces, joiners, byte-order marks and soft hyphens arrive from web pages, PDFs and editors far more often than from any model. Two Unicode ranges — variation selectors and tag characters — genuinely can carry hidden data, which is where the watermark rumours come from. Rather than guess, this page lists what is actually in your text, one row per character, with a count. If there is nothing hidden in it, you will see that too.
Every one of these is invisible or near-invisible on screen, and every one of them breaks something. The inspector on this page names them as it finds them.
| Code point | Name | Where it comes from | What it breaks |
|---|---|---|---|
U+00A0 | No-break space | Word, web pages, French typography | Does not wrap, does not split, often survives trim() |
U+202F | Narrow no-break space | Model output, French locales | Same, and narrower — invisible in most fonts |
U+200B | Zero-width space | Web pages, copy-paste | Breaks search and exact matching |
U+200C / U+200D | Zero-width non-joiner / joiner | Persian, Indic scripts, emoji sequences | Meaningful in those scripts; noise elsewhere; removing the joiner splits emoji |
U+FEFF | Byte-order mark | Files saved as UTF-8 with BOM | First CSV column never matches |
U+00AD | Soft hyphen | PDFs, justified documents | Invisible until the word wraps |
U+2060 | Word joiner | Typesetting | Invisible, prevents a break |
U+FE00–U+FE0F | Variation selectors | Emoji, and hidden-data tricks | Invisible; a long run of them is carrying something |
U+E0000–U+E007F | Tag characters | Steganography | Invisible; can encode a whole hidden message |
U+202A–U+202E, U+2066–U+2069 | Bidi controls | Right-to-left text, Trojan Source attacks | Text can display in a different order than it is stored |
U+2014 / U+2013 | Em dash / en dash | Edited prose, AI output | Not an ASCII hyphen; breaks code and matching |
U+2018 U+2019 U+201C U+201D | Curly quotes | Word, autocorrect | Break SQL, JSON, CSV and code |
U+2026 | Ellipsis | Word autocorrect | One character where three dots were meant |
U+FB01 / U+FB02 | fi / fl ligatures | PDF extraction | Searching for “file” misses “file” |
Cyrillic а е о р с | Look-alike letters | Transliteration, OCR, deliberate evasion | Identical on screen, different in every comparison |
| Feature | This tool | convertcase.net | textcleaner.net | typical “AI text cleaner” |
|---|---|---|---|---|
| Shows you what is in the text before deleting it | ✓ full character inspector | ✗ no | ✗ no | ~ counts only |
| Names each character with its Unicode code point | ✓ yes | ✗ no | ✗ no | ✗ no |
| Em dash handling with a choice of replacement | ✓ hyphen, comma, space, remove | ✓ yes | ~ remove only | ~ remove only |
| Invisible characters (zero-width, BOM, soft hyphen, variation selectors) | ✓ detected and listed | ~ removed silently | ~ removed silently | ~ removed silently |
| Bidi / Trojan-Source controls | ✓ flagged | ✗ no | ✗ no | ✗ no |
| Look-alike (homoglyph) letters | ✓ flagged and foldable | ✗ no | ✗ no | ✗ no |
| Unwrap PDF line breaks, keeping paragraphs | ✓ yes, plus de-hyphenation | ~ line breaks only | ~ line breaks only | ✓ yes |
| Strip Markdown from AI output | ✓ headings, bold, links, tables | ~ partial | ✗ no | ~ partial |
| Extract e-mails, URLs, numbers, IPs | ✓ 9 extractors | ~ e-mails, URLs | ✗ no | ~ some |
| Runs entirely in the browser | ✓ no upload, no backend | ✓ yes | ~ unclear | ✓ yes |
| Ads | ✓ none | ✗ ad-supported | ✗ ad-supported | ✗ ad-supported |
Checked on 18 September 2026 against convertcase.net, textcleaner.net and the current crop of “AI text cleaner” pages. They are competent at deleting; the difference here is being told what was there.
Think about what people paste into a text cleaner: a contract with names in it, an exported customer list, an incident report, a draft nobody has seen yet. A tool that posts that to a server for processing has just moved your document somewhere you cannot see, for a job your own browser can do in a millisecond. This page has no backend at all. The HTML is static, the JavaScript is in the page, and the network stays silent while you work — open the developer tools and watch, or pull the plug and keep cleaning.
Set Em/en dashes to the replacement you want — a hyphen, a comma, a plain space, or nothing at all — and every — and – in the text is converted as you type. The tool handles the spacing around them too, so “word — word” becomes “word - word” rather than “word - word”. The em dash became a talking point because language models produce it far more often than most people type it: they learned from edited prose, where it is standard punctuation.
The ones that actually turn up are zero-width spaces (U+200B), zero-width joiners and non-joiners, word joiners, byte-order marks, soft hyphens, narrow no-break spaces (U+202F) and variation selectors. Most arrive by accident: they come from web pages, PDFs, editors and tokenisation, not from a deliberate mark. A run of variation selectors or Unicode tag characters (U+E0000–U+E007F) can carry hidden data, which is why they are worth seeing rather than guessing about. This page lists every one it finds, with its code point and count, and lets you decide.
Use the AI output preset. It removes invisible and direction-control characters, converts the special spaces to ordinary ones, straightens curly quotes, turns em dashes into hyphens, strips the Markdown (**bold**, ## headings, link syntax, code fences and tables) and normalises the bullet glyphs to -. What you get is the same words in plain text that will paste anywhere without surprises.
Because it is not the characters you think it is. Word gives you curly quotes instead of straight ones, non-breaking spaces instead of spaces, en dashes instead of hyphens, and sometimes a byte-order mark at the very start of the text that makes the first column fail to match anything. The Pasted from Word preset converts all of them; the inspector shows you which ones were there in the first place.
That is two problems at once, and the Pasted from a PDF preset fixes both. Words split across lines with a hyphen are re-joined (“exam-\nple” → “example”), and the hard line breaks inside paragraphs are removed while the blank line between paragraphs is kept. Lists and headings are left on their own lines, because joining those is what makes other “remove line breaks” tools produce a wall of text.
It reads your input and lists every character that is not plain ASCII: the character itself, its code point, its Unicode name, how many there are, and one line on what it breaks. Invisible and direction-control characters are shown with a placeholder, because by definition you cannot see them. Each row has a button that switches on the option which removes that kind of character, so you can go from “something is wrong with this text” to a fix without knowing any of the terminology.
U+200B is a character with no width: it is there, it takes up a position, and it shows nothing. Web pages use it to allow a line break inside a long word; copy-paste then carries it along. The damage is quiet — a search for “invoice” will not match “invoice”, two identical-looking strings compare as different, and a primary key that contains one is a bug nobody can see.
They tell the renderer to display text in a different order than it is stored. That is essential for Arabic and Hebrew, and it is also the basis of the “Trojan Source” trick, where source code reads one way to a human reviewer and compiles to something else, and of file names like “gnp.exe” that display as “exe.png”. If you find one in a pasted snippet that has nothing to do with right-to-left languages, it deserves a look.
A homoglyph is a character that looks like another one. Cyrillic а, Greek ο and full-width a are different letters from Latin a, o and a, and no amount of staring at them will tell you which is which. They are used in lookalike domains and in text meant to defeat copy-detection, and they arrive innocently from transliteration and bad OCR. The inspector names them and the cleaner can fold them back to Latin.
U+00AD is an invisible hint that says “you may break the word here”. Justified documents and PDFs are full of them. They are invisible until the word wraps, which means a word can contain one and still look perfectly normal while failing every search and comparison you throw at it.
A non-breaking space (U+00A0) looks identical and behaves differently: text never wraps at it, split(' ') does not split on it, and trim() in some languages does not trim it. It is the most common invisible troublemaker in pasted text, along with the narrow no-break space (U+202F) that turns up in French typography and in model output.
Set Unwrap lines to paragraphs. Lines inside a paragraph are joined into one, and the blank line that separates paragraphs is kept. Set it to everything instead if you genuinely want one long line — useful for a CSV cell or a single-line prompt, and rarely what you want for prose.
There are two switches, because there are two questions. Remove duplicate lines is exact. Ignore case and spacing treats “Jane@Example.com ” and “jane@example.com” as the same line, which is what you want for a list of addresses. Keep only duplicates inverts the whole thing and shows you what repeats — the fastest way to find the rows a spreadsheet exported twice.
Use the E-mail list preset, or set Extract to e-mail addresses. Every address in the text is pulled out, duplicates dropped, one per line (or comma-separated, if that is what the next tool wants). The same extractor handles URLs, domains, numbers, IP addresses, hashtags, mentions, dates and quoted strings.
The Slug preset strips accents, lowercases, and replaces everything that is not a letter or a number with a hyphen — so “Café — Crème Brûlée!” becomes cafe-creme-brulee. Check the result before publishing: transliteration of non-Latin scripts is a judgement call no tool makes perfectly.
Yes, both ways. Join lines with puts them on one line with a comma, semicolon, pipe, tab or space between them; Split on does the reverse. Quote every line wraps each one in quotes, with a CSV option that adds the trailing comma, which is how most of this ends up being used: turning a column from a spreadsheet into something you can paste into an IN (…) clause.
No. The page is static HTML and every transformation happens in JavaScript in your tab. There is no backend, no request when you type, no analytics on the content, and nothing is stored. You can open the developer tools and watch the network stay silent, or disconnect entirely — it keeps working. That matters for the text people actually paste into cleaners: contracts, customer lists, incident reports.
Only the ones you ask it to. Everything under whitespace, characters and typography touches formatting, not content. The transforms that do change words — remove punctuation, remove numbers, case conversion, slugify, extraction — are separate switches, off by default, and the change log lists every step that fired.
Tens of thousands of lines are fine; the work is done as you type, so very large files make typing feel sluggish long before anything fails. For a multi-megabyte file, paste it in sections. There is no upload limit because there is no upload.
Because U+200D is doing a job there. A family emoji is several people glued together with joiners; a profession emoji is a person joined to an object. Strip the joiners and you get the pieces back. That is why the inspector names the character and counts it instead of quietly deleting it — in ordinary prose it is noise, inside an emoji it is the glue.
Cleaning never edits your input: the left pane stays exactly as you pasted it, and the right pane is the result. Switch an option off and the output goes back. Use as input is the one deliberate exception — it moves the cleaned text across so you can run a second pass on it.
Free, no account, no ads, no upload limit. It is one of the free browser tools at jasperbernaers.com.
Part of 122 free, no-signup browser tools. A few that pair well with the text cleaner: