PasteDetox — Free Text Sanitizer for Pasted Content

Free in-browser tool that cleans messy pasted text from Word, Google Docs, Slack, Jira, and PDFs. Fixes smart quotes, em-dashes, zero-width characters, soft hyphens, and weird Unicode that breaks code, JSON, and LLM prompts.

Why use PasteDetox?

Fix Smart Quotes

Modern word processors like Microsoft Word, Google Docs, and Slack automatically convert standard straight quotes into "smart" curly quotes. While these look nice in print, they cause devastating syntax errors in programming languages like Python, JavaScript, JSON, and HTML. PasteDetox instantly reverts these characters back to code-safe straight quotes.

Remove Invisible Characters

Copying text from PDFs or websites often introduces hidden characters that are invisible to the naked eye but fatal to compilers. Common culprits include the Zero Width Space (ZWSP / u200B) and Left-to-Right marks. Our tool scans the byte-level data of your clipboard and purges these phantom characters, ensuring your strings are truly empty when they look empty.

Sanitize Non-Breaking Spaces

The Non-Breaking Space is the most common copy-paste bug in web development. It looks exactly like a space, but it is treated as a different character by interpreters. This frequently breaks CSS selectors, SQL queries, and variable names. PasteDetox identifies every variation of whitespace and normalizes them into standard ASCII spaces.

PasteDetox is a free developer utility provided by id8. We believe in privacy-first tools. This entire application runs client-side in your browser using JavaScript. No text you paste is ever sent to our servers.

PasteDetox Technical FAQ and Unicode Sanitization Guide

What PasteDetox does for developers

PasteDetox cleans text that looks normal to humans but breaks parsers, compilers, query engines, and config readers. It normalizes smart punctuation to plain ASCII where appropriate, removes invisible control and formatting code points, and standardizes whitespace so text behaves predictably across tooling. This matters in real engineering workflows where content is copied from PDFs, chat tools, office documents, issue trackers, and CMS editors before being pasted into code, JSON payloads, SQL, or prompt templates.

The key value is deterministic sanitation. Manual cleanup is unreliable because problematic characters are often not visible in editors, and different fonts hide problems differently. A sanitizer can scan byte-level or code-point-level data and enforce a strict output policy every time. That policy can be conservative for prose, or strict for code contexts where only a narrow character set is valid. PasteDetox is designed for this repeatable cleanup path so copy-paste errors stop reaching production systems.

History of the copy-paste corruption problem

Rich text systems introduced typographic substitutions long before modern web development workflows existed. Word processors replaced straight quotes with directional quotes and inserted non-breaking spaces for layout control. Later, web editors and chat applications inherited similar behavior. At the same time, Unicode added many format and directionality characters to support global scripts. These features are useful in publishing and international text, but they create failure modes when the same text is treated as code or strict data.

Invisible Unicode marks such as zero-width space or bidirectional control characters are especially dangerous because they can change token boundaries or visual ordering without obvious visual cues. This caused a long tail of confusing bugs: broken selectors, failed imports, invalid JSON keys, and mismatched identifiers that looked identical on screen. Security research also highlighted how confusable characters and bidi controls can mask malicious code intent. As a result, text sanitation moved from convenience feature to essential hygiene in developer tooling.

How to sanitize text programmatically without this tool

Build a multi-stage sanitation pipeline. Start with Unicode normalization, usually NFKC for code-adjacent text, to collapse compatibility variants. Then remove disallowed code points by class, including zero-width format characters, unwanted control characters, and optionally bidi overrides outside approved locales. Next, map punctuation substitutions such as curly quotes and long dashes to ASCII equivalents when your target grammar requires plain tokens. After that, normalize whitespace by converting non-breaking spaces and unusual separators to standard spaces or newlines.

Implement allowlist and denylist modes. In allowlist mode, keep only characters permitted by a target grammar, such as JSON string-safe text or identifier-safe subsets. In denylist mode, remove only known problematic ranges and preserve broad Unicode for multilingual content. Add diagnostics to your sanitizer output: report which character classes were changed and how many replacements were made. This audit data helps teams detect upstream sources that repeatedly inject toxic formatting and prioritize fixes at the source.

Engineering tutorial: integrating sanitation into CI and apps

Treat sanitation as a boundary control, not an afterthought. Apply it at ingestion points: clipboard handlers in web apps, file upload preprocessors, and API endpoints that receive user text. In repositories, add pre-commit hooks that scan staged files for hidden characters and fail fast with line-level diagnostics. In backend services, store both raw and sanitized versions when compliance permits, so you can trace user input while executing only safe normalized text in downstream pipelines.

For language tooling, combine sanitation with parser validation. After cleaning, immediately parse the result using the same parser used in production. This catches cases where character cleanup was not enough because structural syntax is still broken. For prompt pipelines, sanitize before token counting and before caching so duplicate detection remains stable. For SQL or command templates, sanitation is not a replacement for parameterization, but it removes accidental character-level corruption that makes debugging and observability unnecessarily difficult.

Edge cases and safe handling strategy

Not every non-ASCII character is bad. Teams working with multilingual data, mathematical notation, or legal names must preserve valid Unicode. The correct strategy is context-aware profiles: strict ASCII for code snippets, broad Unicode for user-facing prose, and mixed policies for AI prompts where structure is strict but content can be multilingual. Keep profile selection explicit and visible in your codebase. That prevents accidental over-sanitization while still eliminating hidden characters that routinely break developer workflows.

You might also like

🧵 Tweet Splitter 🩺 JSON Surgeon