Repository navigation
🐛 fix(fuzz): size harness buffers to their input - #1235
Merged
Merged
Conversation
The IDNA and phone harnesses decoded UTF-8 into one slot per input byte. Multibyte input decodes to fewer code points than bytes, so the engines read from a block with spare slots past the last code point, and ASan missed any read into them. All three standalone harnesses also gave an empty input one slot. Count the code points first and allocate that many, and let an empty input or file get a zero-size block. libFuzzer copies each input into a block of its exact size for the same reason.
Merging this PR will regress 1 benchmark
Warning Please fix the performance issues or acknowledge them on CodSpeed. Performance Changes
Tip Investigate this regression by commenting Comparing Footnotes
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The IDNA and phone harnesses decoded UTF-8 into a block with one slot per input byte. A multibyte input decodes to fewer code points than bytes, so the engines got a buffer with spare slots past the last code point, and ASan missed a read into them. All three standalone harnesses, the JS one included, also gave an empty input one slot.
Each harness now counts the code points first and allocates that many, none for an empty input. The file readers allocate the file size, none for an empty file. libFuzzer copies each input into a heap block of
Sizebytes for the same reason, so ASan catches an overflow of it (FuzzerLoop.cpp).With a harness patched to read one code point past the end, the old allocation let a multibyte and an empty input pass clean, and the new one stops both with a heap-buffer-overflow report. The fixed harnesses run the seed corpora and
tox -e asan-jsclean under ASan and UBSan.