Why Markdown output needs sanitizing before it reaches a page
· 6 min read
Why Markdown output is not safe HTML, when to sanitize rather than escape, and how to parse-then-sanitize with DOMPurify against injected scripts.

Markdown parsers do not make your HTML safe. Most of them pass raw HTML in the source straight through to the output, which means a document containing <script>alert(1)</script> produces a page containing that script. The Python-Markdown documentation states plainly that the library does not sanitize its output and that cleaning untrusted input is the developer's job. This is not an oversight in one library; it is the documented contract of the category.
The fix is a second step. Parse the Markdown, then run the resulting HTML through a sanitizer before it is injected into the page. OWASP's XSS Prevention Cheat Sheet calls this HTML sanitization, and it is the correct tool specifically when you intend to let users supply HTML, which is exactly what a Markdown converter does.
Key Takeaways
- Markdown is not safe HTML. Most parsers forward raw tags in the source without filtering them.
- Encoding and sanitizing solve different problems: encoding shows a tag as text, sanitizing decides which tags are allowed to render.
- Parse first, sanitize second. Sanitizing the Markdown source misses HTML the parser itself generates.
- DOMPurify is the sanitizer OWASP recommends for untrusted HTML, and it runs in the browser.
What Is HTML Sanitization?
HTML sanitization is the process of parsing untrusted HTML and returning a version with anything dangerous removed, while keeping the safe structure intact. Where encoding converts a tag into visible text, sanitization keeps the tag and drops the attributes or elements that can execute code.
The distinction matters because the two techniques are not interchangeable. If you encode everything, a user's legitimate bold text and headings come back as literal <strong> characters on screen. If you sanitize, they render properly and the dangerous parts are gone. A converter needs the second behavior, because rendering Markdown means honoring the markup the author wrote.
OWASP's guidance draws the line this way: use encoding to display untrusted values as text, and use sanitization when users are intentionally allowed to provide HTML, such as through a rich-text editor. A Markdown converter is the second case. The MDN's glossary entry on cross-site scripting describes the same class of risk, where injected script runs with the full privileges of the page it lands on.
The vulnerability, concretely
Markdown allows raw HTML in two ways. Inline, a tag can sit inside a paragraph:
Hello <img src=x onerror=alert(1)> world
And as a block, on its own lines:
<div onclick="steal()">
click me
</div>
A parser that does not filter these will hand them straight to the browser. The onerror attribute in the first example fires because the image source fails to load. This is a stored XSS vector: anyone who can submit Markdown to your system can run script in the browser of anyone who later views the rendered output.
Raw HTML passthrough is documented behavior, not a bug. GitHub Flavored Markdown's specification lists disallowed raw HTML as a documented feature, and the same is true of CommonMark. Disabling raw HTML entirely is one option, but it breaks the legitimate case of a technical writer embedding a <br> or a video embed, which is why sanitizing the output is the more common approach. The CommonMark specification defines this passthrough as the standard behavior, and the format dates from John Gruber's original 2004 release whose goal was readable plain text rather than a security model.
Why the order matters
There are two places you could sanitize, and picking the wrong one leaves a gap.
Sanitizing the Markdown source first is a mistake. At that point you are looking at text, not markup. The parser has not run, so you cannot tell whether <b> in the source came from the user or will be produced by a Markdown construct like _emphasis_. A sanitizer inspecting the source either strips text a legitimate author needed or misses constructs that only become tags after parsing.
Parse, then sanitize, closes the gap. You run marked.parse() to get HTML, then pass that HTML through the sanitizer. At that point every tag is explicit, including the ones the parser generated from Markdown syntax, and the sanitizer sees the complete document.
The order is the whole point. It is the difference between filtering an ambiguous text format and filtering a known markup language.
Encoding versus sanitizing, side by side
| Encoding | Sanitizing | |
|---|---|---|
| Purpose | Display untrusted data as text | Allow some HTML, remove the rest |
Effect on <b> |
Shows as literal <b> |
Renders as bold |
Effect on <script> |
Shows as literal text | Removed |
| Right for | Values interpolated into a page | User-supplied HTML |
| OWASP name | HTML Entity Encoding | HTML sanitization |
OWASP lists entity, attribute, URL, JavaScript, and CSS encoding as separate rules because each context parses differently. The takeaway for a converter is narrower: you are building a system that deliberately accepts HTML from users, so sanitization is the mechanism, and encoding alone would defeat the tool's purpose.
What DOMPurify actually does
DOMPurify is the sanitizer OWASP names in its cheat sheet, described as a DOM-only sanitizer covering HTML, MathML, and SVG. It parses the input into a real DOM, walks it, and rebuilds the document from nodes that pass its allowlist.
Two properties matter for a converter. First, it runs in the browser, so the same code path works offline and in an in-page tool with no server in the loop. Second, its defaults remove the constructs that execute code, including event-handler attributes, javascript: URLs, and elements like <script> and <iframe> that have no business in rendered Markdown.
DOMPurify also integrates with Trusted Types, the browser mechanism that makes assigning untrusted strings to DOM sinks a compile-time-visible operation rather than a silent one. There is one rule that catches people out. DOMPurify's own documentation states that changes made after sanitization can invalidate the protection, so sanitized content should not be transformed into a different context afterward. If you sanitize and then run the result through a second renderer, you have an unsanitized gap again. Sanitize last, or sanitize once and do not touch the output.
Applying it in a converter
The pattern is short. Parse, sanitize, and only then inject:
const result = await marked.parse(markdown, { gfm: true, breaks: true });
const clean = DOMPurify.sanitize(result);
container.innerHTML = clean;
Two properties of this shape are worth keeping. Sanitizing happens once, at the point the HTML is produced, rather than at each place it gets rendered, so a preview pane and an export path cannot drift apart. And the sanitized value is the only thing ever assigned to innerHTML, so both call sites inherit the same guarantee instead of each needing their own check.
If you want to test whether a converter is doing this, paste a raw script tag into the editor and look at the HTML output. A converter that sanitizes shows the script stripped or escaped, with the surrounding text intact. One that does not will show your tag reflected back, which is the signal to stop using it on untrusted input.
Related tools and further reading
The format itself is described in the Markdown syntax cheat sheet, which covers where raw HTML is allowed in the source. The Markdown to HTML converter applies this pattern on every keystroke, running marked and DOMPurify in your browser so nothing is uploaded. If you already have HTML and need to strip it, the HTML Sanitizer / XSS Filter does the sanitizing step on its own. For the gap between the two ideas, HTML escaping vs sanitizing works through when each applies. The Markdown syntax cheat sheet covers the raw HTML passthrough rule that makes this necessary in the first place.
Written by
Abhay Khant
Abhay Khant is the founder of ToolSura, a privacy-first developer tools platform. Writes about client-side architecture, AI tooling, and the open web.