Escaping vs Sanitizing: The Distinction That Keeps XSS Out
· 6 min read
Escaping shows markup as text, sanitizing removes unsafe markup and keeps the rest. When to use each, and why a converter needs the second.

Two techniques get used interchangeably in conversation and they are not interchangeable. Escaping makes dangerous characters visible, so a script tag becomes text on the page. Sanitizing decides what is allowed, so a script tag is removed while the paragraph around it survives. One changes how something is displayed, the other changes whether it exists.
Getting this wrong produces one of two bad outcomes. A system that escapes when it should sanitize shows users the raw markup, which breaks the feature. A system that sanitizes when it should escape strips legitimate content, which is quieter and arguably worse because nothing looks broken. The rule that separates them is simple: escaping is for text you did not author, sanitizing is for markup somebody deliberately wrote.
Key Takeaways
- Escaping converts markup into visible text. Sanitizing removes unsafe markup and keeps the rest.
- Escaping is right for values interpolated into a page. Sanitizing is right when users supply HTML on purpose.
- A Markdown converter is the sanitizing case, because rendering means honoring the author's markup.
- They are not alternatives. Systems that accept user HTML commonly need both, in different places.
What Is the Difference?
Escaping and sanitizing both take a string that might contain dangerous characters and make it safe to render. They differ in what they do to the string.
Escaping replaces the significant characters with entity references, as the W3C HTML specification defines them, so < becomes < and & becomes &. The browser receives text, not markup, and displays exactly what you encoded. Nothing is removed, because nothing was unsafe once encoded.
Sanitizing parses the string as HTML, walks the result, and rebuilds it from the nodes that pass a policy. A paragraph stays a paragraph. A heading stays a heading. A script element and its contents are dropped. Attributes that can execute code are stripped from the elements that remain.
| Escaping | Sanitizing | |
|---|---|---|
| Result is | Text | Markup |
<b>bold</b> |
Shows as literal characters | Renders bold |
<script> |
Shows as literal text | Removed entirely |
| Use when | Displaying a value | Accepting authored HTML |
The distinction is the difference between showing someone a tag and letting them have one. OWASP's XSS Prevention Cheat Sheet draws it explicitly: use encoding to display untrusted values as text, and use sanitization when users are intentionally providing HTML.
When escaping is the right answer
Escaping is correct when the data is a value rather than markup. A username, a search term, a filename, a comment body: these are strings a person supplied, and the correct rendering is to show them.
commentEl.textContent = userInput; // safe, no escaping needed
The MDN documentation on textContent notes that assigning to it treats the value as text, which is why it is the recommended alternative to innerHTML for untrusted data. Where a template language does the work for you, the same principle applies, and OWASP's guidance on templating treats auto-escaping as the default posture for exactly this reason.
The characteristic of the escaping case is that the user did not intend to write markup. If someone types <b> into a comment field, they want to see those characters, not bold text.
When sanitizing is the right answer
Sanitizing is correct when the input is markup by design. A rich-text editor, a wiki page, an issue body, a Markdown document: in each case the user is writing formatted content and expects it to render.
Here escaping is wrong twice over. It shows the user their own markup as literal text, and it does not actually make the system meaningfully safer, because the developer will eventually need to allow some formatting and will turn escaping off.
DOMPurify is the sanitizer OWASP names for this case. It parses input into a real DOM, applies an allowlist, and returns the rebuilt markup, covering HTML, MathML, and SVG with protections against the specific mutation techniques that bypass naive filters. It runs in the browser, which is what makes it usable in a client-side tool with no server in the loop.
Why a converter needs the sanitizing one
A Markdown converter is unambiguously a sanitizing case, and the reason is that rendering is the entire point. The author writes **bold** expecting a <strong> element, and headings, lists, links, tables, and code blocks are all markup they deliberately produced. Escaping all of that would produce a page full of visible asterisks and hash signs, which is not a converter at all.
The complication is that Markdown also permits raw HTML in the source, which means the output contains elements the Markdown syntax never described. That is the path an attacker uses, and it is why parsing alone is not enough.
const html = await marked.parse(markdown);
const clean = DOMPurify.sanitize(html); // the necessary second step
The OWASP Cross Site Scripting Prevention Cheat Sheet and the Python-Markdown documentation state the general rule plainly: the library does not sanitize its output, and cleaning untrusted input is the developer's responsibility. This is not a criticism of one project. It is the contract of the category, and treating it as anything else is the actual bug.
The case for using both
The techniques are not competitors. In a real application, both appear, usually in the same codebase, in different roles.
A comment field uses escaping, because the input is a value. A wiki page uses sanitizing, because the input is authored markup. A page title might use escaping when it comes from a form field and sanitizing when it comes from a rich-text block. A template that auto-escapes everything still needs a sanitizer anywhere a user can deliberately supply HTML.
The mistake is picking one technique and applying it uniformly. A system that escapes everything loses its formatting. A system that sanitizes everything strips values it should have displayed as text, and a username containing an apostrophe or a comparison symbol comes back subtly mangled.
Testing which one you have
The OWASP cheat sheet lists the encoding rules per context, and a real system applies the one matching where the value lands. There is a quick empirical check. Paste a raw script tag into the system and observe the result. If you see the characters displayed, you are looking at escaping. If the tag is gone and the surrounding text remains, you are looking at sanitizing. If the tag is gone and the text around it is also gone, something is over-aggressive.
A second check catches the more common bug. Enter a legitimate bold tag. If it renders bold, you are sanitizing and user formatting works. If you see literal characters, you are escaping where you should be sanitizing, and the feature is broken in a way that is easy to mistake for a styling issue.
Related tools and further reading
The Markdown to HTML converter is a sanitizing case end to end, running marked and DOMPurify in your browser on every parse. For the sanitizing step alone, the HTML Sanitizer / XSS Filter does that on existing markup. The HTML Entity Encoder/Decoder is the escaping tool, useful when you need to display markup as text. For the security reasoning behind the second step, sanitizing Markdown output against XSS covers the pipeline in detail.
Written by
Abhay Khant
Abhay Khant is the founder of ToolSura, a privacy-first developer tools platform. Writes about client-side architecture, AI tooling, and the open web.