Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/changelog/960.bugfix.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
Attribute URL hosts the way a browser resolves them in ``normalize_url``, ``clean_url``, ``extract_links``,
``resolve_links``, and the sanitizer's ``media_hosts`` allowlist: a special-scheme authority ends at a backslash, an
IPv4 address is read in all four notations and emitted dotted-decimal, the host is percent-decoded, and an IPv6 literal
is zero-compressed.
7 changes: 7 additions & 0 deletions docs/how-to/links.rst
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,13 @@ with :func:`turbohtml.extract.normalize_url`. It applies the WHATWG URL standard

https://example.org/page?a=1&b=2

The host is attributed the way a browser resolves it, so a link-safety or SSRF check keyed on the result sees the host a
browser fetches: a special-scheme authority ends at a backslash (``http://a\@b/`` has host ``a``), an IPv4 address is
read in decimal, octal, hexadecimal, and short forms and re-emitted dotted-decimal (``http://127.1/`` becomes
``http://127.0.0.1/``), the host is percent-decoded, and an IPv6 literal is zero-compressed (``[0:0:0:0:0:0:0:1]``
becomes ``[::1]``). An IPv6 literal with an embedded-IPv4 tail (``[::ffff:1.2.3.4]``) keeps its given spelling. The same
host parsing backs ``extract_links(external_only=True)`` and :meth:`~turbohtml.Node.resolve_links`.

For URLs scraped out of markup, :func:`turbohtml.extract.clean_url` first scrubs HTML damage (stray whitespace,
``&``, a truncating quote) and answers ``None`` for anything that is not a fetchable web URL, so a scraping pipeline
can filter and normalize in one call. :class:`turbohtml.extract.UrlCleaning` carries the knobs: a strict query-parameter
Expand Down
7 changes: 4 additions & 3 deletions docs/reference/clean.rst
Original file line number Diff line number Diff line change
Expand Up @@ -66,9 +66,10 @@ drops the declaration holding it.

``Policy.attribute_filter`` replacements and ``Policy.set_attributes`` additions pass through the mandatory safety
checks before serialization. These checks remove event handlers, disallowed URL, ``srcset`` and meta refresh schemes,
unsafe CSS, media hosts outside ``media_hosts``, and values outside ``attribute_values``. Template stripping and
named-property isolation run on the final values. The sanitizer ASCII-lowercases HTML names created by these rules
before the checks.
unsafe CSS, media hosts outside ``media_hosts`` (matched against the host a browser resolves, so an authority ending at
a backslash cannot smuggle an off-allowlist host past the check), and values outside ``attribute_values``. Template
stripping and named-property isolation run on the final values. The sanitizer ASCII-lowercases HTML names created by
these rules before the checks.

``Policy.transform_tags`` renames elements during the same walk, sanitize-html's ``transformTags``. Key it by source
tag: map to a bare string to rename, or to a :class:`Transform` to rename and add attributes. The rename runs *before*
Expand Down
28 changes: 19 additions & 9 deletions src/turbohtml/_c/clean/sanitize.c
Original file line number Diff line number Diff line change
Expand Up @@ -552,31 +552,41 @@ static int is_media_host_tag(uint16_t atom) {
}
}

/* The authority marker bytes that end a URL host: a path, query, or fragment. */
/* The bytes that end a URL authority: a path, query, or fragment delimiter, plus the '\' a browser treats as a
separator for a special scheme (WHATWG authority state, https://url.spec.whatwg.org/#authority-state). A real host
never contains '\', so ending the authority at one only tightens the host the allowlist sees, closing the
`evil.example\@good.example` userinfo trick a browser resolves to evil.example. */
static int ends_authority(Py_UCS4 c) {
switch (c) {
case '/':
case '?':
case '#':
case '\\':
return 1;
default:
return 0;
}
}

/* Locate the authority host of a URL value: the host after "scheme://" or a protocol-relative "//". The authority is
bounded here without preprocessing the value (the WHATWG tab/newline stripping a browser applies is intentionally not
done, so an obfuscated host never masquerades as an allowlisted one), then th_url_authority -- the same decomposition
url_split runs -- splits off any "userinfo@" and ":port" and reports the host span and its literal kind. Sets
*start,*end to the host span and *kind to the host literal, returns 0, or returns -1 when the URL carries no
authority (a relative or opaque src, which has no host to match). */
/* A URL authority opener slash: '/', or the '\' a browser treats alike for a special scheme. */
static int authority_slash(Py_UCS4 c) {
return c == '/' || c == '\\';
}

/* Locate the authority host of a URL value: the host after "scheme://" or a protocol-relative "//" (with '\' accepted
like '/', as a browser does for a special scheme). The authority is bounded here without preprocessing the value (the
WHATWG tab/newline stripping a browser applies is intentionally not done, so an obfuscated host never masquerades as
an allowlisted one), then th_url_authority -- the same decomposition url_split runs -- splits off any "userinfo@" and
":port" and reports the host span and its literal kind. Sets *start,*end to the host span and *kind to the host
literal, returns 0, or returns -1 when the URL carries no authority (a relative or opaque src, which has no host to
match). */
static int url_host_span(const Py_UCS4 *value, Py_ssize_t len, Py_ssize_t *start, Py_ssize_t *end, int *kind) {
Py_ssize_t authority = -1;
if (len >= 2 && value[0] == '/' && value[1] == '/') {
if (len >= 2 && authority_slash(value[0]) && authority_slash(value[1])) {
authority = 2; /* protocol-relative //host/path */
}
for (Py_ssize_t index = 0; authority < 0 && index + 2 < len; index++) {
if (value[index] == ':' && value[index + 1] == '/' && value[index + 2] == '/') {
if (value[index] == ':' && authority_slash(value[index + 1]) && authority_slash(value[index + 2])) {
authority = index + 3; /* scheme://host */
}
}
Expand Down
9 changes: 9 additions & 0 deletions src/turbohtml/_c/core/common.h
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,15 @@ PyObject *turbohtml_url_language_matches(PyObject *module, PyObject *args);
PyObject *th_url_to_ascii(PyObject *host);
PyObject *turbohtml_url_to_ascii(PyObject *module, PyObject *arg);

/* Implemented in url/url.c. th_url_host_canonical runs the rest of the WHATWG host parser over the bracket-stripped
host span url_split reports (https://url.spec.whatwg.org/#concept-host-parser): a bracketed IPv6 literal (kind
TH_HOST_IPV6) is parsed and serialized with zero-run compression; any other host is percent-decoded, run through
domain-to-ASCII, and -- when it ends in a number -- parsed as IPv4 and re-serialized in dotted-decimal, so two
spellings of one address compare equal. The host is a borrowed str, the result a new str (brackets dropped for IPv6);
NULL with an error only on allocation failure. A host that is not valid for its form falls back to its lowercased
spelling, the advisory behavior normalize_url already takes for an unencodable label. */
PyObject *th_url_host_canonical(PyObject *host, int kind);

/* Implemented in url/registrable.c. _registrable_domain(host) returns a lowercased
host's registrable domain (eTLD+1) from the shipped IANA and Public Suffix List
tables, the site boundary behind extract_links(external_only=True). METH_O. */
Expand Down
19 changes: 10 additions & 9 deletions src/turbohtml/_c/extract/links.c
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
This file is the walk that locates every link-bearing location and, for the rewrite path, splices a replacement back
in place. The genuinely new capability over iterating <a href> by hand is the URLs embedded in CSS url()/@import (in
a style attribute and in <style> text), in a <meta http-equiv=refresh> content value, and in the srcset/ping/archive
list attributes. URL resolution itself (resolve_links) is stdlib urllib.parse.urljoin, bound through
functools.partial here, so RFC 3986 is not reinvented. */
list attributes. URL resolution itself (resolve_links) is the C _url_join, bound through functools.partial here, so
RFC 3986 resolution is not reinvented and a reference's host matches the one a browser resolves. */

#include "core/ascii.h"
#include "core/common.h"
Expand Down Expand Up @@ -724,7 +724,8 @@ static Py_ssize_t base_scheme_of(PyObject *base_url, char *out, Py_ssize_t cap)
}

/* Node.resolve_links(base_url) -> None. Rewrites every link absolute against base_url with functools.partial bound over
stdlib urllib.parse.urljoin, so RFC 3986 resolution is not reinvented. */
the C _url_join, so a reference's host matches the one a browser resolves -- including the WHATWG backslash authority
forms stdlib urllib.parse.urljoin misattributes -- without reinventing RFC 3986 resolution. */
PyObject *turbohtml_node_resolve_links(PyObject *owner, th_tree *tree, th_node *root, PyObject *base_url) {
if (!PyUnicode_Check(base_url)) {
PyErr_SetString(PyExc_TypeError, "resolve_links expected a base URL string");
Expand All @@ -734,13 +735,13 @@ PyObject *turbohtml_node_resolve_links(PyObject *owner, th_tree *tree, th_node *
Py_ssize_t base_scheme_len = base_scheme_of(base_url, base_scheme, (Py_ssize_t)sizeof(base_scheme));
int base_netloc_scheme = (base_scheme_len == 4 && memcmp(base_scheme, "http", 4) == 0) ||
(base_scheme_len == 5 && memcmp(base_scheme, "https", 5) == 0);
PyObject *parse_module = PyImport_ImportModule("urllib.parse");
if (parse_module == NULL) { /* GCOVR_EXCL_BR_LINE: a stdlib import cannot be forced to fail from a test */
return NULL; /* GCOVR_EXCL_LINE: import-failure path */
PyObject *html_module = PyImport_ImportModule("turbohtml._html");
if (html_module == NULL) { /* GCOVR_EXCL_BR_LINE: the extension's own module is already imported */
return NULL; /* GCOVR_EXCL_LINE: import-failure path */
}
PyObject *urljoin = PyObject_GetAttrString(parse_module, "urljoin");
Py_DECREF(parse_module);
if (urljoin == NULL) { /* GCOVR_EXCL_BR_LINE: urllib.parse.urljoin always exists */
PyObject *urljoin = PyObject_GetAttrString(html_module, "_url_join");
Py_DECREF(html_module);
if (urljoin == NULL) { /* GCOVR_EXCL_BR_LINE: _url_join is always registered on the module */
return NULL; /* GCOVR_EXCL_LINE: attribute-failure path */
}
PyObject *functools_module = PyImport_ImportModule("functools");
Expand Down
48 changes: 10 additions & 38 deletions src/turbohtml/_c/url/clean.c
Original file line number Diff line number Diff line change
Expand Up @@ -115,30 +115,6 @@ static int str_holds(PyObject *text, Py_UCS4 needle) {
return PyUnicode_FindChar(text, needle, 0, PyUnicode_GET_LENGTH(text), 1) >= 0;
}

/* The ASCII (punycode) form of a registered name the way the URL standard's host parser produces it (spec 3.5): the
lowercased host when it is already ASCII, else UTS #46 ToASCII in C; a label punycode cannot encode (an unpaired
surrogate) leaves the lowercased host as it is, which the later encode step then rejects. */
static PyObject *ascii_host(PyObject *host) {
PyObject *lowered = PyObject_CallMethod(host, "lower", NULL);
if (lowered == NULL) { /* GCOVR_EXCL_BR_LINE: str.lower cannot fail on a host */
return NULL; /* GCOVR_EXCL_LINE: allocation-failure path */
}
if (PyUnicode_IS_ASCII(lowered)) {
return lowered;
}
PyObject *encoded = th_url_to_ascii(lowered);
if (encoded != NULL) {
Py_DECREF(lowered);
return encoded;
}
if (!PyErr_ExceptionMatches(PyExc_ValueError)) { /* GCOVR_EXCL_BR_LINE: ToASCII raises nothing else */
Py_DECREF(lowered); /* GCOVR_EXCL_LINE: allocation-failure path */
return NULL; /* GCOVR_EXCL_LINE */
}
PyErr_Clear();
return lowered;
}

/* The ":port" suffix, or "" for an absent, empty, or scheme-default port (port state, URL standard 4.4). A port of
digits is read as the integer it spells, so leading zeros fall away and "0080" is the http default. */
static PyObject *port_suffix(const th_url_parts *parts) {
Expand Down Expand Up @@ -176,20 +152,16 @@ static PyObject *port_suffix(const th_url_parts *parts) {
return suffix;
}

/* The authority rebuilt from its normalized host and port, keeping userinfo verbatim: a registered name goes through
domain-to-ASCII, an IPv4/IPv6 literal is already ASCII and only lowercases (IPv6 keeping its brackets). */
/* The authority rebuilt from its WHATWG-canonical host and port, keeping userinfo verbatim: the host is
percent-decoded, domain-to-ASCII'd, and IPv4/IPv6-canonicalized by th_url_host_canonical, then a bracketed IPv6
literal is re-wrapped. */
static PyObject *normalize_netloc(const th_url_parts *parts) {
PyObject *host;
if (parts->kind == TH_HOST_REGNAME) {
host = ascii_host(parts->part[TH_URL_HOST]);
} else {
PyObject *lowered = PyObject_CallMethod(parts->part[TH_URL_HOST], "lower", NULL);
if (lowered == NULL) { /* GCOVR_EXCL_BR_LINE: str.lower cannot fail on a host */
return NULL; /* GCOVR_EXCL_LINE: allocation-failure path */
}
host = parts->kind == TH_HOST_IPV6 ? th_str_format("[%U]", lowered) : Py_NewRef(lowered);
Py_DECREF(lowered);
PyObject *canonical = th_url_host_canonical(parts->part[TH_URL_HOST], parts->kind);
if (canonical == NULL) { /* GCOVR_EXCL_BR_LINE: the host parse only fails on allocation failure */
return NULL; /* GCOVR_EXCL_LINE: allocation-failure path */
}
PyObject *host = parts->kind == TH_HOST_IPV6 ? th_str_format("[%U]", canonical) : Py_NewRef(canonical);
Py_DECREF(canonical);
if (host == NULL) { /* GCOVR_EXCL_BR_LINE: the host fold only fails on allocation failure */
return NULL; /* GCOVR_EXCL_LINE: allocation-failure path */
}
Expand Down Expand Up @@ -465,9 +437,9 @@ static PyObject *site_of(PyObject *url) {
if (th_url_split(url, &parts) < 0) {
return NULL;
}
PyObject *host = ascii_host(parts.part[TH_URL_HOST]);
PyObject *host = th_url_host_canonical(parts.part[TH_URL_HOST], parts.kind);
th_url_parts_clear(&parts);
if (host == NULL) { /* GCOVR_EXCL_BR_LINE: the host fold only fails on allocation failure */
if (host == NULL) { /* GCOVR_EXCL_BR_LINE: the host parse only fails on allocation failure */
return NULL; /* GCOVR_EXCL_LINE: allocation-failure path */
}
PyObject *site = turbohtml_registrable_domain(NULL, host);
Expand Down
Loading
Loading