Skip to content

🐛 fix(url): reject mapped host delimiters - #1193

Merged
gaborbernat merged 5 commits into
tox-dev:mainfrom
gaborbernat:feat/url-reparse-1010
Oct 7, 2026
Merged

gaborbernat merged 5 commits into
tox-dev:mainfrom
gaborbernat:feat/url-reparse-1010

Conversation

@gaborbernat

@gaborbernat gaborbernat commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Percent-decoded or IDNA-mapped host characters can become URL component delimiters on a second normalization. For example, http://o%23 produced http://o#/ and then http://o/#/. Rejecting forbidden domain characters after host mapping follows the WHATWG domain-to-ASCII rule. Node 26.10.0 with Ada 4.0.0 rejects these inputs, matching the pinned Ada host parser.

The split/reparse oracle composes the public normalization result through urllib.parse.urlsplit, preserving authority syntax and empty delimiter presence. It detects changed composition, changed second normalization and rejection of the serialized output. This completes the remaining URL requirement alongside the merged IDNA, NFC and encoding invariant checks.

normalize_url raises ValueError for a forbidden mapped host; clean_url returns None. Empty file authorities, IPv6 and the documented Unicode fallback retain their behavior. Validated nonnumeric hosts skip the IPv4 buffer allocation. Sixteen of the 17 measured workloads use fewer TOTAL instructions, including a 4.0% reduction for long ASCII hosts; deep-link extraction retains a measured 2.0% increase.

Closes #1010

Decoded or IDNA-mapped host delimiters changed URL components during a
second normalization. Reject forbidden domain characters after mapping
and check serialization through a reference split and public reparse.

Refs tox-dev#1010
Rejecting mapped host delimiters requires checking each character. The
character search added 18.6% instructions to deep link extraction. Keep
the same forbidden set with a constant lookup in the existing copy loop.

Refs tox-dev#1010
@gaborbernat gaborbernat added bug Something isn't working enhancement New feature or request labels Oct 7, 2026
@gaborbernat
gaborbernat marked this pull request as ready for review October 7, 2026 00:52
@codspeed

codspeed Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 5.14%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

❌ 1 regressed benchmark
✅ 580 untouched benchmarks
⏩ 32 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_feature[prune-shared-single] 90.8 µs 95.8 µs -5.14%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing gaborbernat:feat/url-reparse-1010 (451ff89) with main (c389009)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Mapped domain validation must run before IPv4 fallback. Nonnumeric
first characters cannot begin an IPv4 number, so retain the validated
host without allocating or copying the IPv4 buffer.
@gaborbernat
gaborbernat merged commit ad6aa23 into tox-dev:main Oct 7, 2026
54 of 55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fuzz URL, IDNA and encoding detection against their invariants

1 participant