Join our Newsletter — 33% off our NHI Course
Home› FAQ› Cyber Security› Why do lexical HTML sanitizers still fail to…
Cyber Security

Why do lexical HTML sanitizers still fail to stop some XSS payloads?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 30, 2026 Domain: Cyber Security

They fail when the browser and the sanitizer do not interpret the same input in the same way. An attacker can place markup inside a context that the sanitizer treats as text, while the browser later reparses it as active HTML or JavaScript. That mismatch lets dangerous nodes survive sanitization and execute in the DOM.

How lexical sanitizers and browser parsers get out of sync

Lexical sanitizers usually inspect a string before the browser has turned it into a DOM. That is the weak point: they decide whether input “looks safe” based on tokens, delimiters, and context assumptions, but the browser may later interpret the same bytes differently after decoding, entity expansion, or reparsing. XSS appears when those two interpretation models diverge.

A sanitizer can miss payloads that are inert in one parsing context but become executable when the browser normalises them into HTML, SVG, MathML, URL, or script-bearing nodes. The problem is not only bad filtering rules, it is that string-level validation cannot fully model all browser parsing states and mutation steps.

This is why payloads that survive “as text” can still become active markup after insertion into the page. Once the browser builds the DOM, later operations such as innerHTML assignment, template insertion, attribute repair, or namespace switching can surface executable content that never looked dangerous to the sanitizer in its original form.

Why parser differentials create bypasses

Many bypasses are parser differentials: the sanitizer and the browser do not agree on where tags start and end, how quotes are balanced, or whether a fragment belongs to HTML at all. That mismatch lets an attacker hide an event handler, script URL, or malformed element inside input the sanitizer classed as harmless text.

The browser is also more forgiving than most lexical filters. It will often recover from broken markup, ignore certain syntax errors, and reinterpret fragments after decoding. A sanitizer that blocks obvious strings like “<script>” may still fail when the payload uses entity encoding, obscure tag forms, nested contexts, or namespace-sensitive content that the browser later promotes into executable structure.

The key lesson is that lexical sanitizers are trying to infer final browser behaviour from incomplete text evidence. When the target sink is DOM-based, template-based, or mutation-driven, the gap between “safe string” and “safe DOM” is exactly where XSS survives.

What defenders should verify before trusting a sanitizer

Sanitization only holds when the output is tested against the exact browser sink and context it will reach. A policy that is safe for plain HTML text may fail for attribute values, inline SVG, rich-text editors, Markdown renderers, or frameworks that later reserialize and reparse content.

Defenders should verify three things: the sanitizer’s parsing model, the browser context where the output lands, and whether any subsequent DOM mutation can reintroduce active markup. If any step after sanitization can reinterpret the content, the sanitizer is only a partial control, not a guarantee.

For this reason, output encoding at the final sink remains more reliable than “sanitize once, use everywhere.” In practice, robust defences combine context-aware encoding, strict allowlists, and avoidance of dangerous DOM sinks, rather than treating the sanitizer as the only barrier.

Risk and Threat Considerations

The security risk is not simply “bad input gets through,” but that a seemingly inert fragment can become executable after browser reparsing. That makes bypasses hard to spot in review and easy to miss in testing, especially when the payload only activates inside a specific browser context or after a downstream DOM transformation.

Failure mechanism: The sanitizer and the browser apply different parsing rules, so input judged harmless as a string later becomes active HTML, scriptable attributes, or executable DOM content when the browser normalises it.

Impact: An attacker can achieve XSS despite filtering, leading to session theft, action forgery, content injection, or stored compromise in any application that reuses the same sanitised content across contexts.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP ASVS provides the primary governance reference for this topic.

FrameworkControl / ReferenceRelevance
OWASP ASVSV1 — Encoding and SanitizationCovers context-aware encoding and sanitization for untrusted HTML inputs.
V3 — Web Frontend SecurityApplies because browser parsing and DOM mutation drive the XSS bypass.
V16 — Security Logging and Error HandlingSupports monitoring and test evidence for failed sanitization and injection attempts.
Recommendation — Use context-aware encoding and sanitizer validation before any HTML reaches the browser sink. Review frontend sinks and DOM mutation paths that can reintroduce executable content. Log rejected payloads and test cases to detect sanitizer gaps and parser differentials.

Practitioner Guidance

What to verify: Test the exact browser sink, not just the sanitizer library. If sanitised content is ever inserted with innerHTML, template rendering, rich-text editing, or a framework reparse step, treat the output as untrusted until you have context-specific proof.

Common mistake: Teams often validate against the library’s documented examples and stop there. That misses the real control boundary, which is the browser’s final interpretation, not the sanitizer’s intermediate string result.

Decision rule: If content must remain user-editable or HTML-like, use a well-maintained allowlist sanitizer plus strict output encoding at the final sink. If you cannot guarantee stable context, prefer plain text rendering over partial HTML preservation.

Practitioner takeaway: The safest mental model is that sanitization reduces risk, but only the browser’s final parsing context determines whether the payload is actually dead.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 30, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org