PasteDetox — Free Text Sanitizer for Pasted Content
Free in-browser tool that cleans messy pasted text from Word, Google Docs, Slack, Jira, and PDFs. Fixes smart quotes, em-dashes, zero-width characters, soft hyphens, and weird Unicode that breaks code, JSON, and LLM prompts.
PasteDetox Technical FAQ and Unicode Sanitization Guide
What PasteDetox does for developers
PasteDetox cleans text that looks normal to humans but breaks parsers, compilers, query engines, and
config readers. It normalizes smart punctuation to plain ASCII where appropriate, removes invisible
control and formatting code points, and standardizes whitespace so text behaves predictably across
tooling. This matters in real engineering workflows where content is copied from PDFs, chat tools,
office documents, issue trackers, and CMS editors before being pasted into code, JSON payloads, SQL, or
prompt templates.
The key value is deterministic sanitation. Manual cleanup is unreliable because problematic characters
are often not visible in editors, and different fonts hide problems differently. A sanitizer can scan
byte-level or code-point-level data and enforce a strict output policy every time. That policy can be
conservative for prose, or strict for code contexts where only a narrow character set is valid.
PasteDetox is designed for this repeatable cleanup path so copy-paste errors stop reaching production
systems.
History of the copy-paste corruption problem
Rich text systems introduced typographic substitutions long before modern web development workflows
existed. Word processors replaced straight quotes with directional quotes and inserted non-breaking
spaces for layout control. Later, web editors and chat applications inherited similar behavior. At the
same time, Unicode added many format and directionality characters to support global scripts. These
features are useful in publishing and international text, but they create failure modes when the same
text is treated as code or strict data.
Invisible Unicode marks such as zero-width space or bidirectional control characters are especially
dangerous because they can change token boundaries or visual ordering without obvious visual cues. This
caused a long tail of confusing bugs: broken selectors, failed imports, invalid JSON keys, and
mismatched identifiers that looked identical on screen. Security research also highlighted how
confusable characters and bidi controls can mask malicious code intent. As a result, text sanitation
moved from convenience feature to essential hygiene in developer tooling.
How to sanitize text programmatically without this
tool
Build a multi-stage sanitation pipeline. Start with Unicode normalization, usually NFKC for
code-adjacent text, to collapse compatibility variants. Then remove disallowed code points by class,
including zero-width format characters, unwanted control characters, and optionally bidi overrides
outside approved locales. Next, map punctuation substitutions such as curly quotes and long dashes to
ASCII equivalents when your target grammar requires plain tokens. After that, normalize whitespace by
converting non-breaking spaces and unusual separators to standard spaces or newlines.
Implement allowlist and denylist modes. In allowlist mode, keep only characters permitted by a target
grammar, such as JSON string-safe text or identifier-safe subsets. In denylist mode, remove only known
problematic ranges and preserve broad Unicode for multilingual content. Add diagnostics to your
sanitizer output: report which character classes were changed and how many replacements were made. This
audit data helps teams detect upstream sources that repeatedly inject toxic formatting and prioritize
fixes at the source.
Engineering tutorial: integrating sanitation into CI
and apps
Treat sanitation as a boundary control, not an afterthought. Apply it at ingestion points: clipboard
handlers in web apps, file upload preprocessors, and API endpoints that receive user text. In
repositories, add pre-commit hooks that scan staged files for hidden characters and fail fast with
line-level diagnostics. In backend services, store both raw and sanitized versions when compliance
permits, so you can trace user input while executing only safe normalized text in downstream pipelines.
For language tooling, combine sanitation with parser validation. After cleaning, immediately parse the
result using the same parser used in production. This catches cases where character cleanup was not
enough because structural syntax is still broken. For prompt pipelines, sanitize before token counting
and before caching so duplicate detection remains stable. For SQL or command templates, sanitation is
not a replacement for parameterization, but it removes accidental character-level corruption that makes
debugging and observability unnecessarily difficult.
Edge cases and safe handling strategy
Not every non-ASCII character is bad. Teams working with multilingual data, mathematical notation, or
legal names must preserve valid Unicode. The correct strategy is context-aware profiles: strict ASCII
for code snippets, broad Unicode for user-facing prose, and mixed policies for AI prompts where
structure is strict but content can be multilingual. Keep profile selection explicit and visible in your
codebase. That prevents accidental over-sanitization while still eliminating hidden characters that
routinely break developer workflows.