Skip to content

🐛 fix(markdown): drop whitespace in link destinations - #1237

Merged
gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/markdown-url-whitespace
Oct 8, 2026
Merged

gaborbernat merged 1 commit into
tox-dev:mainfrom
gaborbernat:fix/markdown-url-whitespace

Conversation

@gaborbernat

Copy link
Copy Markdown
Member

to_markdown copied a tab or line break in a link or image destination into the Markdown. <a href="a\nb">x</a> became [x](a\nb), which reads back as literal text, since no destination form holds a line ending (CommonMark 6.3).

The URL parser removes every tab and newline from a URL (URL Standard, basic URL parser), so a\nb and ab name the same URL. The destination writer now leaves them out before it picks the bare or the <...> form, giving [x](ab).

markdownify 1.2.3 and html2text 2025.4.15 both write [x](a\nb).

@gaborbernat gaborbernat added the bug Something isn't working label Oct 8, 2026
@gaborbernat
gaborbernat force-pushed the fix/markdown-url-whitespace branch from 9234df8 to 0c37ec5 Compare October 8, 2026 13:47
@codspeed

codspeed Bot commented Oct 8, 2026

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 579 untouched benchmarks
⏩ 32 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_feature[prune-shared-single] 89.1 µs 96.1 µs -7.31%
⚡ test_feature[shadow-slot-comments] 137.9 µs 83.2 µs +65.74%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing gaborbernat:fix/markdown-url-whitespace (0c37ec5) with main (af31f20)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

A destination cannot hold a line ending (CommonMark 6.3), so an href with a
newline came out as literal text. The URL parser removes every tab and
newline anyway (URL Standard 4.4), so the destination writer leaves them out
and the URL stays the same.
@gaborbernat
gaborbernat force-pushed the fix/markdown-url-whitespace branch from 0c37ec5 to 32e53ca Compare October 8, 2026 15:59
@gaborbernat
gaborbernat merged commit 14f3415 into tox-dev:main Oct 8, 2026
7 of 61 checks passed
gaborbernat added a commit to gaborbernat/turbohtml that referenced this pull request Oct 8, 2026
On tox-dev#1237's run the summary listed one input and its stack trace as two
crashes, both under the fuzzer name "undefined". CIFuzz saves each
input as out/artifacts/<target>/<sanitizer>/<input> with an
<input>.summary beside it (fuzz_target.py _target_artifact_path and
_save_crash). cflite.py read the parent directory, the sanitizer, and
counted the summary too.

cflite.py now takes the target from the first directory under
out/artifacts and skips summary files. A maintainer needs that name to
reproduce the crash, and it is not secret.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant