Repository navigation
🐛 fix(detect): BOM codec decodes with replace - #961
Merged
Merged
Conversation
For byte-order-mark input, detect() returned a codec name that delegated to CPython's strict utf-8-sig/utf-16/utf-32 decoder, so data.decode(match.codec) raised UnicodeDecodeError on truncated or ill-formed bytes while turbohtml.parse handled the same input. That broke the documented recipe for untrusted bytes. Force the delegated decoder to "replace", matching the parser and the other whatwg-* codecs, which already map malformed sequences to U+FFFD.
gaborbernat
force-pushed
the
fix/detect-bom-codec
branch
from
October 1, 2026 17:17
aebb6e8 to
f953781
Compare
Merging this PR will not alter performance
Comparing Footnotes
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
detect()'sEncodingMatch.codecfor a byte-order mark named a strict CPython codec, so the documenteddata.decode(match.codec)recipe raisedUnicodeDecodeErroron truncated or ill-formed input even thoughturbohtml.parsedecodes the same bytes to U+FFFD (CWE-20, CWE-248). 🔒 The docstring and the encoding how-to promisematch.codecreproduces the parser's text with no strict mode to trip over.For input starting with a mark,
detect(anddetect_all,EncodingDetector.close) reports a codec ofwhatwg-utf-8-sig,whatwg-utf-16le,whatwg-utf-16be,whatwg-utf-32leorwhatwg-utf-32be. Those five names resolved through a delegate that wrapscodecs.lookup(label).decode, CPython's default strict decoder, while every otherwhatwg-*codec replaces, so detection and decoding disagreed only on these five.The fix forces the delegated decoder to
"replace", matching the parser and the otherwhatwg-*codecs while keepingutf-8-sig's mark stripping and UTF-32 support. The change is confined to the Python codec-registration shim's error mode; no decoding of valid input, no detection result and no C path changes.Replacement is the only conformant behaviour here, so matching it aligns turbohtml with the standard rather than with any one competitor's quirk. The WHATWG Encoding "decode" algorithm never fails and maps every malformed sequence to U+FFFD, and the Rust encoding_rs that browsers build on has no strict mode on its decode-to-string path. turbohtml's own native
whatwg-*decoders already replace, so the five byte-order-mark names were the lone exception. CPython'scodecsdefault toerrors="strict"and raise, which is correct forbytes.decodebut wrong for a codec documented to reproduce a WHATWG parse, so the shim pins the error mode rather than inheriting that default.Decoderreplacementutf-16-le/utf-8-sigerrors="strict"defaultA service that followed the documented decode path on untrusted bytes got an uncaught
UnicodeDecodeErroron byte-order-mark input the parser accepts, a denial of service on that path with no confidentiality or integrity impact.This fixes the strict-decode crash on byte-order-mark input through
EncodingMatch.codec, tracked privately in GHSA-p5gq-fcfg-8fjj.