Skip to content

🐛 fix(detect): BOM codec decodes with replace - #961

Merged
gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/detect-bom-codec
Oct 1, 2026
Merged

gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/detect-bom-codec

Conversation

@gaborbernat

@gaborbernat gaborbernat commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

detect()'s EncodingMatch.codec for a byte-order mark named a strict CPython codec, so the documented data.decode(match.codec) recipe raised UnicodeDecodeError on truncated or ill-formed input even though turbohtml.parse decodes the same bytes to U+FFFD (CWE-20, CWE-248). 🔒 The docstring and the encoding how-to promise match.codec reproduces the parser's text with no strict mode to trip over.

For input starting with a mark, detect (and detect_all, EncodingDetector.close) reports a codec of whatwg-utf-8-sig, whatwg-utf-16le, whatwg-utf-16be, whatwg-utf-32le or whatwg-utf-32be. Those five names resolved through a delegate that wraps codecs.lookup(label).decode, CPython's default strict decoder, while every other whatwg-* codec replaces, so detection and decoding disagreed only on these five.

The fix forces the delegated decoder to "replace", matching the parser and the other whatwg-* codecs while keeping utf-8-sig's mark stripping and UTF-32 support. The change is confined to the Python codec-registration shim's error mode; no decoding of valid input, no detection result and no C path changes.

from turbohtml.detect import detect

data = b"\xef\xbb\xbfhi\xc3"        # UTF-8 BOM + a truncated 2-byte sequence
match = detect(data)               # match.codec == "whatwg-utf-8-sig"
data.decode(match.codec)
# before: UnicodeDecodeError; after: "hi�", the text turbohtml.parse already produced

Replacement is the only conformant behaviour here, so matching it aligns turbohtml with the standard rather than with any one competitor's quirk. The WHATWG Encoding "decode" algorithm never fails and maps every malformed sequence to U+FFFD, and the Rust encoding_rs that browsers build on has no strict mode on its decode-to-string path. turbohtml's own native whatwg-* decoders already replace, so the five byte-order-mark names were the lone exception. CPython's codecs default to errors="strict" and raise, which is correct for bytes.decode but wrong for a codec documented to reproduce a WHATWG parse, so the shim pins the error mode rather than inheriting that default.

Library / version Malformed bytes after detection Mechanism
WHATWG Encoding "decode" U+FFFD, never fails replacement error mode
encoding_rs U+FFFD; no strict decode-to-string Decoder replacement
CPython 3.14 utf-16-le / utf-8-sig raise errors="strict" default
turbohtml ≤ 1.13.1 raised for the 5 BOM codecs, replaced for the rest strict delegate

A service that followed the documented decode path on untrusted bytes got an uncaught UnicodeDecodeError on byte-order-mark input the parser accepts, a denial of service on that path with no confidentiality or integrity impact.

This fixes the strict-decode crash on byte-order-mark input through EncodingMatch.codec, tracked privately in GHSA-p5gq-fcfg-8fjj.

@gaborbernat gaborbernat added the bug Something isn't working label Oct 1, 2026
For byte-order-mark input, detect() returned a codec name that delegated to
CPython's strict utf-8-sig/utf-16/utf-32 decoder, so data.decode(match.codec)
raised UnicodeDecodeError on truncated or ill-formed bytes while
turbohtml.parse handled the same input. That broke the documented recipe for
untrusted bytes.

Force the delegated decoder to "replace", matching the parser and the other
whatwg-* codecs, which already map malformed sequences to U+FFFD.
@gaborbernat
gaborbernat force-pushed the fix/detect-bom-codec branch from aebb6e8 to f953781 Compare October 1, 2026 17:17
@codspeed

codspeed Bot commented Oct 1, 2026

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 580 untouched benchmarks
⏩ 32 skipped benchmarks1


Comparing gaborbernat:fix/detect-bom-codec (f953781) with main (2af1136)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@gaborbernat
gaborbernat merged commit d95dc7b into tox-dev:main Oct 1, 2026
45 of 52 checks passed
@gaborbernat
gaborbernat deleted the fix/detect-bom-codec branch October 1, 2026 21:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant