Skip to content

🧪 test(fuzz): add sanitizer wrong-output oracles - #971

Merged
gaborbernat merged 3 commits into
tox-dev:mainfrom
gaborbernat:fuzz/sanitizer-oracles
Oct 2, 2026
Merged

gaborbernat merged 3 commits into
tox-dev:mainfrom
gaborbernat:fuzz/sanitizer-oracles

Conversation

@gaborbernat

@gaborbernat gaborbernat commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

HTML sanitizers rarely crash when they fail. They return markup a browser parses into a different tree from the one the sanitizer judged, or keep a URL whose scheme the browser reads differently from the allowlist check. The published bypasses follow that pattern: bleach let schemes through when code points above U+00A0 split them, loofah missed javascript: split by character references, html-sanitizer normalized lookalike characters into markup after it had sanitized, HtmlSanitizer passed template content that Chromium's 512-depth flattening pops out, and Chromium's Sanitizer API let a fast-path javascript: check disagree with its own URL parser. Most of these need a non-default configuration. turbohtml's fuzz loop reports crashes only, and its re-parse idempotence check runs turbohtml's own parser under one policy, so a quirk shared by sanitize and re-parse cancels out.

🔬 fuzz.py --mode oracle (tox -e fuzz-oracle) adds three wrong-output oracles. The first serializes the sanitized tree and re-parses it with the vendored html5lib-python as a div fragment with scripting on, then requires the same tags, namespaces, attributes and text; parse5 re-parses each minimized divergence and labels it either a turbohtml mutation or a stale html5lib rule. Neither parser applies the 512 open-element cap turbohtml shares with Blink and Gecko, so inputs nested past it are skipped. The second oracle draws a random Policy and requires every element and attribute of that re-parse to be allowed by it, with no event handler, baseline element or disallowed URL. The third compares each URL keep-or-drop decision with the scheme turbohtml's public WHATWG normalize_url reads from the decoded attribute value.

The re-parse oracle follows OWASP java-html-sanitizer's validator.nu round trip and the tree comparison from Klein and Johns' MutaGen study, which also recommends parsing with scripting on; its inside-out generator is re-implemented here because the MutaGen repository carries no license. Policies come from the hazard pool behind the advisories above: foreign content, raw-text elements, SVG animate and set, URL attributes and custom schemes. The URL obfuscations replay the bleach, loofah, html-sanitizer and Chromium bypass shapes. Before the run starts, each oracle must fire on a planted negative control in the style of DOMPurify's fast-check suite, so an oracle gone blind fails the run.

Seeds come from the DOMPurify fixtures, cure53's H5SC vectors, the WPT sanitizer-api tree cases with their #config sections mapped onto Policy, and payloads generated from Dharma's xss.dg grammar through the published dharma package. Splicing draws from google/fuzzing's HTML, SVG, MathML and CSS dictionaries, concatenated the way jsoup's OSS-Fuzz build does. H5SC, the dictionaries and WPT arrive as pinned shallow submodules under tools/fuzz-data; WPT is marked update = none and checked out sparsely to sanitizer-api/, since a full clone is about a gigabyte.

Pull requests replay the deterministic seed pass only. The scheduled run adds generation from a secret per-run seed and writes minimized findings to a JSON report; its public log shows only their hashes, and sanitizer reports from a generating run go to the encrypted crash directory, the disclosure rule #970 set for the crash fuzzer.

@gaborbernat gaborbernat added the enhancement New feature or request label Oct 1, 2026
@codspeed

codspeed Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 6.3%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

❌ 1 regressed benchmark
✅ 580 untouched benchmarks
⏩ 32 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_feature[select-relative-sibling] 42.7 µs 45.6 µs -6.3%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing gaborbernat:fuzz/sanitizer-oracles (10b38e9) with main (99d0327)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@gaborbernat
gaborbernat force-pushed the fuzz/sanitizer-oracles branch 3 times, most recently from 69c0c2a to 8de17c9 Compare October 2, 2026 01:36
The sanitizer oracles need hostile inputs nobody here wrote: H5SC
vectors, the google/fuzzing html, svg, mathml and css dictionaries, and
the WPT sanitizer-api tree cases with their #config. All three come in as
pinned submodules under tools/fuzz-data, the way the bench corpora live
under tools/bench-data.

A depth-1 WPT clone is about 1 GB while sanitizer-api is a few files, so
the WPT submodule is marked update = none. Plain and recursive submodule
updates, including the conformance job's checkout, skip it, and the fuzz
driver checks out only that directory. The directory stays out of ruff,
ty and the sdist like the other vendored suites.
Sanitizer bugs show up as unsafe or altered markup, so the crash-only
loop misses them, and the existing re-parse checks use turbohtml's own
parser, where a bug shared by sanitize and re-parse cancels out.

fuzz.py --mode oracle (tox -e fuzz-oracle) re-parses each sanitized tree
with the vendored html5lib-python and requires the same tree, the way
OWASP java-html-sanitizer re-parses with validator.nu. On that re-parse
every element and attribute must be allowed by a random Policy drawn
from the mXSS hazard pool, with no on* handler and no disallowed URL
scheme. A third oracle compares each URL keep or drop with the scheme
turbohtml.extract.normalize_url reads from the decoded value under the
WHATWG rules. Trees past the 512 nesting cap are skipped, and trees
holding a select use turbohtml's own re-parse, since html5lib and parse5
predate the customizable select rules.

Inputs come from DOMPurify, H5SC, WPT sanitizer-api, Dharma's xss.dg and
a MutaGen-style generator re-implemented from its unlicensed gen.ml.
Negative controls run first and fail the run when an oracle goes blind.
Findings are ddmin-reduced, re-checked with parse5, and written to a
JSON report while the log shows only their hashes. Pull requests replay
the seed pass; scheduled runs add generation.
@gaborbernat
gaborbernat force-pushed the fuzz/sanitizer-oracles branch from 8de17c9 to 6ff5d22 Compare October 2, 2026 05:44
tox-dev#977 kept xmlns:p="" and minified a script body holding a parsed empty
comment, with a length guard on each NULL copy. tox-dev#989 now rejects that
declaration before the copy, and tox-dev#994 hands a body with a comment child
to the verbatim path, so main asserted both old and new outcomes and the
guards could never take their false branch. Drop the two stale tests and
both guards; an empty Text child under a minified script or style keeps
a test.
@gaborbernat
gaborbernat merged commit f0bcb8f into tox-dev:main Oct 2, 2026
52 of 53 checks passed
@gaborbernat
gaborbernat deleted the fuzz/sanitizer-oracles branch October 2, 2026 13:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant