Repository navigation
🐛 fix(url): reject mapped host delimiters - #1193
Merged
Merged
Conversation
Decoded or IDNA-mapped host delimiters changed URL components during a second normalization. Reject forbidden domain characters after mapping and check serialization through a reference split and public reparse. Refs tox-dev#1010
Rejecting mapped host delimiters requires checking each character. The character search added 18.6% instructions to deep link extraction. Keep the same forbidden set with a constant lookup in the existing copy loop. Refs tox-dev#1010
gaborbernat
marked this pull request as ready for review
October 7, 2026 00:52
Merging this PR will degrade performance by 5.14%
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ❌ | test_feature[prune-shared-single] |
90.8 µs | 95.8 µs | -5.14% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing gaborbernat:feat/url-reparse-1010 (451ff89) with main (c389009)
Footnotes
-
32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Mapped domain validation must run before IPv4 fallback. Nonnumeric first characters cannot begin an IPv4 number, so retain the validated host without allocating or copying the IPv4 buffer.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Percent-decoded or IDNA-mapped host characters can become URL component delimiters on a second normalization. For example,
http://o%23producedhttp://o#/and thenhttp://o/#/. Rejecting forbidden domain characters after host mapping follows the WHATWG domain-to-ASCII rule. Node 26.10.0 with Ada 4.0.0 rejects these inputs, matching the pinned Ada host parser.The split/reparse oracle composes the public normalization result through
urllib.parse.urlsplit, preserving authority syntax and empty delimiter presence. It detects changed composition, changed second normalization and rejection of the serialized output. This completes the remaining URL requirement alongside the merged IDNA, NFC and encoding invariant checks.normalize_urlraisesValueErrorfor a forbidden mapped host;clean_urlreturnsNone. Empty file authorities, IPv6 and the documented Unicode fallback retain their behavior. Validated nonnumeric hosts skip the IPv4 buffer allocation. Sixteen of the 17 measured workloads use fewer TOTAL instructions, including a 4.0% reduction for long ASCII hosts; deep-link extraction retains a measured 2.0% increase.Closes #1010