Repository navigation
🐛 fix(markdown): preserve Unicode space emphasis - #1291
Merged
gaborbernat merged 3 commits intoOct 10, 2026
Merged
Conversation
CommonMark treats Unicode space separators as whitespace at emphasis boundaries. The exporter treated them as ordinary characters and emitted literal delimiter runs instead of emphasis. Classify those spaces so the existing raw HTML fallback preserves the element and its content. A direct lookup leaves the ASCII path unchanged.
Merging this PR will degrade performance by 18.45%
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ❌ | test_feature[shadow-slot-comments] |
84.3 µs | 138.8 µs | -39.3% |
| ❌ | test_feature[serialize-inner] |
5.1 ms | 5.4 ms | -5.64% |
| ❌ | test_feature[serialize-inner-indent] |
5.5 ms | 5.9 ms | -5.32% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing gaborbernat:fix/markdown-nbsp-emphasis (7ce4636) with main (d4701ec)
Footnotes
-
32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
to_markdown()turned<strong> </strong>into** **. CommonMark's emphasis rules treat U+00A0 and the other Unicode space separators as whitespace, so the stars remain literal. Both md4c and markdown-it-py parse that output without strong emphasis. For the same input, markdownify 1.2.3 drops the content, while html2text 2025.4.15 replaces the nonbreaking space with an ordinary space.Fix
The flanking classifier now recognizes every Unicode
Zsspace. Runs that cannot open or close use the serializer's existing raw HTML fallback, preserving the element and its space. Asterisks cannot delimit strong emphasis around this whitespace under CommonMark, so another Markdown delimiter would still lose the element. A direct indexed table keeps the ordinary ASCII path to one bounds check and one load; the extension grows by 12,288 bytes (0.26%).Turbohtml: strong element and nonbreaking space preserved. The before and after files below contain the exact
to_markdown()output for the same input. This hunk comes fromgit diff --no-index --unified=0:Competitors
markdownify 1.2.3 and html2text 2025.4.15: strong element around NBSP lost
markdownify 1.2.3 output
html2text 2025.4.15 output
For
<p><b> </b></p>, markdownify emits an empty string, dropping the element and U+00A0. html2text emits** **\n\n, replacing U+00A0 with U+0020; markdown-it-py parses its stars as a thematic break. Turbohtml uses raw HTML to retain both the strong element and its nonbreaking space. The diffs contain the actual U+00A0 character.Performance
Fixed-seed Callgrind runs on warmed CPython 3.15 compared the same GCC
-O3builds. For 1,000 conversions, the existing article corpus rose from 88,451,703 to 88,475,569 instructions (+0.027%), while the 100-block emphasis corpus went from 401,304,703 to 401,304,569. The affected two-emphasis case rose from 127,280,856 to 139,600,722 instructions over 20,000 conversions (+9.68%) because it now emits longer raw HTML in place of broken delimiters.CodSpeed flags
shadow-slot-commentsas a 39.3% regression and notes that its comparison uses different runtime environments. That benchmark assigns shadow slots to 1,000 HTML comments without calling the Markdown serializer. With its exact loader and 1,000 warmed invocations under fixed-seed Callgrind, the base uses 9,291,132 instructions and this branch uses 9,291,125 (-7).CodSpeed also flags
serialize-innerandserialize-inner-indentat 5.1 to 5.4 ms and 5.5 to 5.9 ms. On their exact 234,774-byte WHATWG spec corpus, fresh CPython 3.15 release builds at based4701ec6and head7ce46364used 200 warmed calls after one parse. Two subsequent fixed-seed paired Callgrind runs counted 842,636,402 instructions on both revisions forserialize-innerand 926,219,073 on both forserialize-inner-indent. The first head runs counted 849,313,319 and 933,033,128 against the same base counts; the extra instructions were concentrated inmemcpy, whileserialize_compact_stepself instructions were unchanged. Repeated paired counts do not support a persistent serializer regression.