Repository navigation
🔒 fix(url): attribute hosts the WHATWG way - #960
Merged
Merged
Conversation
Resolve the URL/host layer to the WHATWG URL host parser so the host turbohtml attributes matches the one a browser connects to. - End a special-scheme authority at a backslash in url_split, parse_ref, the relative join, and the sanitizer media-host scan, so evil.example\@good.example is read as evil.example. - Parse an IPv4 host in all four notations and emit dotted-decimal, and percent-decode the host before domain-to-ASCII. - Canonicalize an IPv6 literal with zero-run compression. - resolve_links joins through the C _url_join instead of stdlib urljoin. This fixes a media_hosts allowlist bypass and host-confusion that defeated external_only scoping and blocklist/SSRF checks keyed on the extraction helpers' output.
gaborbernat
force-pushed
the
fix/url-host-attribution
branch
from
October 1, 2026 16:47
ba5b714 to
ccb02da
Compare
Merging this PR will regress 1 benchmark
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ❌ | test_feature[urls-ascii-host-long] |
1.7 ms | 1.9 ms | -8.44% |
| ⚡ | test_feature[links-absolutize] |
5.6 ms | 1.7 ms | ×3.3 |
| ⚡ | test_feature[select-relative-sibling] |
45.6 µs | 42.7 µs | +6.85% |
| ⚡ | test_feature[transform-number-count-last] |
585.3 µs | 552 µs | +6.04% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing gaborbernat:fix/url-host-attribution (ccb02da) with main (2af1136)
Footnotes
-
32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
turbohtml attributed a URL's host differently from a browser, so the sanitizer's
media_hostsallowlist, theexternal_onlycrawl boundary, and the host inclean_url/normalize_urloutput could name a host the browser never connects to (CWE-918, CWE-20). 🔒 The fix resolves the whole URL-host layer through the WHATWG URL host parser.A special-scheme authority now ends at a backslash, as the WHATWG authority state requires, so
https://evil.example\@good.example/x.mp4is read with hostevil.example, not thegood.examplethat turbohtml took from after the last@. A mediasrctherefore no longer passes amedia_hosts={good.example}allowlist while the browser fetches from the attacker's host, andextract_links(external_only=True)no longer classifies an off-site link as internal. The same rule drivesclean_url,normalize_url,extract_linksand the relative join, andresolve_linksnow joins through the C_url_joinrather than stdliburllib.parse.urljoin, which kept the backslash misattribution.A host that ends in a number is now parsed as IPv4 in all four notations and re-emitted dotted-decimal, the host is percent-decoded before domain-to-ASCII, and an IPv6 literal is zero-compressed, so a link-safety or SSRF check keyed on the output sees the address a browser fetches.
Ending a special-scheme authority at a backslash and canonicalising the host is what the WHATWG-conformant parsers already do, and turbohtml now joins them. The C++ ada-url and Node's whatwg-url both terminate the authority at
\, read IPv4 in decimal, hex, octal and short forms and re-emit it dotted-decimal, compress IPv6 and percent-decode the host, exactly the behaviournormalize_url's docstring already promised and did not deliver.Two widely used parsers keep the raw spelling, and turbohtml deliberately does not follow them, because neither claims WHATWG host semantics while turbohtml's sanitizer and extraction helpers are documented to match what a browser fetches. CPython's
urlsplitand Go'snet/urlboth keep a backslash in the authority, keep an IPv4 host verbatim and do not percent-decode, which is correct for a generic RFC 3986 parser but wrong for a host-attribution check a browser will second-guess. Non-special schemes stayurllib-compatible and unchanged.\urlsplit, Gonet/url\An attacker who controls a media
srcor a page's links could load media from a host outside amedia_hostsallowlist, sending the viewer's IP and Referer to an unapproved origin, or make a host blocklist, SSRF filter or crawl-scope check built onnormalize_url/clean_url/extract_linksdecide on a host the browser never uses. There is no integrity or availability impact beyond that misattribution.This fixes the browser-divergent URL host attribution across the sanitizer and extraction helpers, tracked privately in GHSA-c5m2-cm7w-7v9c.