A sanitization approach that tries to classify input as text or code before deciding what to allow, block, or rewrite. In rich text workflows it can preserve safe HTML while removing dangerous constructs, but it depends on matching the browser’s parsing behavior closely enough to avoid inconsistent interpretation.
What Lexical Parsing Does
Lexical parsing is a sanitization strategy that classifies incoming markup or text before deciding what to keep, block, or rewrite. It aims to preserve legitimate rich-text content while stripping out unsafe constructs that could be interpreted differently by a browser.
Why It Matters in Rich Text Sanitization
The core value of lexical parsing is that it reasons about the input as data, then applies policy before the browser gets a chance to reinterpret it. That makes it useful in editors, comment systems, CMS workflows, and any application that accepts user-supplied HTML, Markdown, or mixed-format content.
The approach is especially important when the application wants to allow some formatting, because blanket escaping can be too restrictive and naive allowlists can be too permissive. A lexical parser tries to keep the allowed surface area small while still supporting practical content authoring.
Where Parsing Mismatches Create Security Problems
Lexical parsing only works if its interpretation closely tracks the browser’s own parsing behavior. If the sanitizer and the browser disagree about token boundaries, attribute handling, tag nesting, or malformed markup recovery, an attacker can smuggle active content through an apparently safe input path.
That is why this technique is not just about string filtering. It is about correctly understanding how markup will actually be reconstructed and rendered after sanitization.
How It Differs From Simpler Filtering
Simple blacklists, regex-based stripping, and ad hoc replacements often fail when markup is malformed or intentionally ambiguous. Lexical parsing is more structured: it tokenizes input, applies rules to the parsed representation, and then emits a cleaned version that should remain safe under browser interpretation.
In practice, this can preserve useful formatting such as basic headings, links, or emphasis while rejecting scripts, event handlers, dangerous URL schemes, and other executable constructs. The more expressive the allowed HTML, the more carefully the parser must mirror real browser behavior.
When Practitioners Should Prefer It
Common misunderstanding: lexical parsing is not a guarantee of safety by itself. It is a control technique whose security depends on completeness, parser fidelity, and continuous maintenance as browser behavior and HTML features evolve.
Governance implication: teams should treat the sanitizer as security-sensitive parsing logic, not as a formatting convenience. Changes to allowed tags, attributes, and rewrite rules can alter exposure in ways that deserve review and test coverage.
Risk and Threat Considerations
Lexical parsing reduces the chance that malicious markup survives sanitization, but it also creates a trust boundary between the sanitizer and the browser. If that boundary is imperfect, attackers can use malformed HTML, nested constructs, or encoding tricks to produce a different runtime interpretation than the sanitizer expected.
Failure mechanism: the sanitizer classifies the input one way, while the browser reparses the output into a more dangerous structure, which can lead to cross-site scripting, content injection, or unsafe link behavior.
Impact: a single parsing mismatch can turn a “safe” rich-text field into an execution path, affecting user sessions, stored content, and any downstream page that renders the sanitized output.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP ASVS and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP ASVS | V1 — Encoding and Sanitization | Lexical parsing is a sanitization approach for untrusted markup and text. |
| V15 — Secure Coding and Architecture | This control family supports safe handling of rich-text parsing and browser-facing transformations. | |
| V2 — Validation and Business Logic | Parsing and classification rules are part of safe handling for structured user input. | |
| Recommendation — Use V1 to validate that untrusted input is encoded or sanitized before rendering. Design parsing and rewrite logic so sanitizer output matches the browser's interpretation. Use V2 to enforce strict validation rules for accepted text and markup forms. | ||
| NIST SP 800-53 Rev 5 | SI-10 — Information Input Validation | Lexical parsing is an input-handling control that validates and constrains untrusted content. |
| SC-18 — Mobile Code | Markup that can execute or behave dynamically must be constrained to prevent unsafe code-like content. | |
| Recommendation — Apply SI-10 to validate and normalize untrusted markup before it reaches rendering logic. Use SC-18 to restrict or neutralize executable content embedded in user input. | ||
Practitioner Guidance
What to watch for: test cases should include malformed tags, broken nesting, mixed encodings, unusual attribute quoting, and browser-specific edge cases. These inputs are the ones most likely to expose disagreements between a sanitizer and the rendering engine.
Practitioner takeaway: treat lexical parsing as a security control that must be validated against the browser model it is trying to constrain, not as a one-time text-cleaning utility.