Skip to content

🔒 fix(url): attribute hosts the WHATWG way - #960

Merged
gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/url-host-attribution
Oct 1, 2026
Merged

gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/url-host-attribution

Conversation

@gaborbernat

@gaborbernat gaborbernat commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

turbohtml attributed a URL's host differently from a browser, so the sanitizer's media_hosts allowlist, the external_only crawl boundary, and the host in clean_url/normalize_url output could name a host the browser never connects to (CWE-918, CWE-20). 🔒 The fix resolves the whole URL-host layer through the WHATWG URL host parser.

A special-scheme authority now ends at a backslash, as the WHATWG authority state requires, so https://evil.example\@good.example/x.mp4 is read with host evil.example, not the good.example that turbohtml took from after the last @. A media src therefore no longer passes a media_hosts={good.example} allowlist while the browser fetches from the attacker's host, and extract_links(external_only=True) no longer classifies an off-site link as internal. The same rule drives clean_url, normalize_url, extract_links and the relative join, and resolve_links now joins through the C _url_join rather than stdlib urllib.parse.urljoin, which kept the backslash misattribution.

A host that ends in a number is now parsed as IPv4 in all four notations and re-emitted dotted-decimal, the host is percent-decoded before domain-to-ASCII, and an IPv6 literal is zero-compressed, so a link-safety or SSRF check keyed on the output sees the address a browser fetches.

from turbohtml.extract import extract_links, normalize_url

normalize_url("http://2130706433/")
# before: http://2130706433/ ; after: http://127.0.0.1/

extract_links('<a href="http://0x7f.0.0.1/a">x</a>', "http://127.0.0.1/", external_only=True)
# before: keeps the link as external; after: recognised as the same host

Ending a special-scheme authority at a backslash and canonicalising the host is what the WHATWG-conformant parsers already do, and turbohtml now joins them. The C++ ada-url and Node's whatwg-url both terminate the authority at \, read IPv4 in decimal, hex, octal and short forms and re-emit it dotted-decimal, compress IPv6 and percent-decode the host, exactly the behaviour normalize_url's docstring already promised and did not deliver.

Two widely used parsers keep the raw spelling, and turbohtml deliberately does not follow them, because neither claims WHATWG host semantics while turbohtml's sanitizer and extraction helpers are documented to match what a browser fetches. CPython's urlsplit and Go's net/url both keep a backslash in the authority, keep an IPv4 host verbatim and do not percent-decode, which is correct for a generic RFC 3986 parser but wrong for a host-attribution check a browser will second-guess. Non-special schemes stay urllib-compatible and unchanged.

Library / version Backslash authority IPv4 (dec/oct/hex/short) IPv6 compression Percent-decode host
ada-url 4.0.0 / Node whatwg-url ends authority at \ parsed → dotted-decimal compressed yes
CPython urlsplit, Go net/url kept (no WHATWG claim) kept verbatim kept no
turbohtml ≤ 1.13.1 kept (host misattributed) kept verbatim kept no
turbohtml (fixed) ends authority at \ parsed → dotted-decimal compressed yes

An attacker who controls a media src or a page's links could load media from a host outside a media_hosts allowlist, sending the viewer's IP and Referer to an unapproved origin, or make a host blocklist, SSRF filter or crawl-scope check built on normalize_url/clean_url/extract_links decide on a host the browser never uses. There is no integrity or availability impact beyond that misattribution.

This fixes the browser-divergent URL host attribution across the sanitizer and extraction helpers, tracked privately in GHSA-c5m2-cm7w-7v9c.

@gaborbernat gaborbernat added the bug Something isn't working label Oct 1, 2026
Resolve the URL/host layer to the WHATWG URL host parser so the host
turbohtml attributes matches the one a browser connects to.

- End a special-scheme authority at a backslash in url_split, parse_ref,
  the relative join, and the sanitizer media-host scan, so
  evil.example\@good.example is read as evil.example.
- Parse an IPv4 host in all four notations and emit dotted-decimal, and
  percent-decode the host before domain-to-ASCII.
- Canonicalize an IPv6 literal with zero-run compression.
- resolve_links joins through the C _url_join instead of stdlib urljoin.

This fixes a media_hosts allowlist bypass and host-confusion that
defeated external_only scoping and blocklist/SSRF checks keyed on the
extraction helpers' output.
@gaborbernat
gaborbernat force-pushed the fix/url-host-attribution branch from ba5b714 to ccb02da Compare October 1, 2026 16:47
@codspeed

codspeed Bot commented Oct 1, 2026

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 3 improved benchmarks
❌ 1 regressed benchmark
✅ 576 untouched benchmarks
⏩ 32 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_feature[urls-ascii-host-long] 1.7 ms 1.9 ms -8.44%
⚡ test_feature[links-absolutize] 5.6 ms 1.7 ms ×3.3
⚡ test_feature[select-relative-sibling] 45.6 µs 42.7 µs +6.85%
⚡ test_feature[transform-number-count-last] 585.3 µs 552 µs +6.04%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing gaborbernat:fix/url-host-attribution (ccb02da) with main (2af1136)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@gaborbernat
gaborbernat merged commit 9dbe8c0 into tox-dev:main Oct 1, 2026
51 of 52 checks passed
@gaborbernat
gaborbernat deleted the fix/url-host-attribution branch October 1, 2026 21:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant