Your Markdown Parser Is Not Your XSS Boundary
A Markdown parser can produce exactly the right HTML and still leave your application exposed to XSS. Parsing answers what the input means. Sanitization decides which parts of that meaning are allowed to reach an HTML sink. I tested that boundary with Node.js 25.3.0, Marked 18.0.7, DOMPurify 3.4.12, and jsdom 30.0.1. The important comparison is not a screenshot. It is the HTML before and after…
A Markdown parser can create safe HTML, yet fail to protect against Cross-Site Scripting (XSS) attacks. Parsing identifies the input's meaning, while sanitization determines which aspects of that meaning are allowed to become part of the HTML output. Testing this boundary in Node.js 25.3.0, using Marked 18.0.7, DOMPurify 3.4.12, and jsdom 30.0.1 revealed several critical insights.
The key distinction lies in comparing the HTML before and after sanitization, rather than just looking at screenshots. By deliberately keeping parsing and sanitization separate, we can better understand the security implications.
Five scenarios illustrate the boundary:
1. Normal content, when parsed from Markdown, usually survives both stages. For example, a heading and HTTPS link remain intact. A sanitizer should preserve the intended structure, not convert it to plain text.
2. Raw HTML within Markdown is valid syntax but doesn't guarantee safe HTML. Marked 18.0.7 returns the element and its event attribute without modification. DOMPurify, however, removes the potentially dangerous onerror attribute, returning only the img tag.
3. URL schemes require their own policy. An image tag with a javascript:alert(1) link is rendered with the link intact, but sanitized HTML removes the onerror attribute, mitigating the XSS risk.
4. Code examples within Markdown must not be sanitized as malicious attacks. The parser escapes the payload in a preformatted code block, preserving the code's integrity. However, removing the attack-looking source before parsing could damage legitimate security documentation.
5. XSS defenses extend beyond script tags. For forms with user-controlled attributes, applying the SANITIZE_NAMED_PROPS option prevents DOM clobbering by ensuring user-controlled names do not interfere with expected application properties.
The recommended pipeline is: untrusted Markdown → parser → Markdown/HTML AST transforms → sanitizer → serializer → matching HTML sink. SANITIZE_NAMED_PROPS should be applied after the last unsafe operation, as later plugins might reintroduce unsafe properties.
DOMPurify's current threat model advises against sanitizing and then freely modifying the result. The policy and sink must remain aligned. While turning off raw HTML reduces attack surface, it is not a universal solution. Additional considerations include plugins, link protocols, generated IDs, and later transforms.
When verifying a Markdown-to-HTML conversion workflow, three questions should be addressed: Did normal Markdown preserve the intended structure? Did fenced code examples remain inert? Is the untrusted output safe for the application's sink and policy?
To ensure safety, treat external Markdown as untrusted by default, disable raw HTML where unnecessary, sanitize the final HTML tree with a maintained allow-list sanitizer, explicitly model URL schemes, id, name, and styling capabilities, avoid running arbitrary HTML-mutating plugins after sanitization, and thoroughly test the AST shape, rendered HTML, and sanitized HTML separately. Keep sanitizer versions pinned and up-to-date, as security fixes are an integral part of the boundary.
The ultimate question remains: should Markdown renderers disable raw HTML by default, or should they only expose it through an explicit, host-supplied security policy?
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.