Skip to content

🐛 fix(markdown): preserve Unicode space emphasis - #1291

Merged
gaborbernat merged 3 commits into
tox-dev:mainfrom
gaborbernat:fix/markdown-nbsp-emphasis
Oct 10, 2026
Merged

gaborbernat merged 3 commits into
tox-dev:mainfrom
gaborbernat:fix/markdown-nbsp-emphasis

Conversation

@gaborbernat

@gaborbernat gaborbernat commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

to_markdown() turned <strong>&nbsp;</strong> into ** **. CommonMark's emphasis rules treat U+00A0 and the other Unicode space separators as whitespace, so the stars remain literal. Both md4c and markdown-it-py parse that output without strong emphasis. For the same input, markdownify 1.2.3 drops the content, while html2text 2025.4.15 replaces the nonbreaking space with an ordinary space.

Fix

The flanking classifier now recognizes every Unicode Zs space. Runs that cannot open or close use the serializer's existing raw HTML fallback, preserving the element and its space. Asterisks cannot delimit strong emphasis around this whitespace under CommonMark, so another Markdown delimiter would still lose the element. A direct indexed table keeps the ordinary ASCII path to one bounds check and one load; the extension grows by 12,288 bytes (0.26%).

Turbohtml: strong element and nonbreaking space preserved. The before and after files below contain the exact to_markdown() output for the same input. This hunk comes from git diff --no-index --unified=0:

diff --git a/before.md b/after.md
index 3ba7d4f..558ddfd 100644
--- a/before.md
+++ b/after.md
@@ -1 +1 @@
-** **
\ No newline at end of file
+<strong> </strong>
\ No newline at end of file

Competitors

markdownify 1.2.3 and html2text 2025.4.15: strong element around NBSP lost

markdownify 1.2.3 output

diff --git a/after.md b/markdownify.md
index 558ddfd..e69de29 100644
--- a/after.md
+++ b/markdownify.md
@@ -1 +0,0 @@
-<strong> </strong>
\ No newline at end of file

html2text 2025.4.15 output

diff --git a/after.md b/html2text.md
index 558ddfd..1a1cdf8 100644
--- a/after.md
+++ b/html2text.md
@@ -1 +1,2 @@
-<strong> </strong>
\ No newline at end of file
+** **
+

For <p><b>&nbsp;</b></p>, markdownify emits an empty string, dropping the element and U+00A0. html2text emits ** **\n\n, replacing U+00A0 with U+0020; markdown-it-py parses its stars as a thematic break. Turbohtml uses raw HTML to retain both the strong element and its nonbreaking space. The diffs contain the actual U+00A0 character.

Performance

Fixed-seed Callgrind runs on warmed CPython 3.15 compared the same GCC -O3 builds. For 1,000 conversions, the existing article corpus rose from 88,451,703 to 88,475,569 instructions (+0.027%), while the 100-block emphasis corpus went from 401,304,703 to 401,304,569. The affected two-emphasis case rose from 127,280,856 to 139,600,722 instructions over 20,000 conversions (+9.68%) because it now emits longer raw HTML in place of broken delimiters.

CodSpeed flags shadow-slot-comments as a 39.3% regression and notes that its comparison uses different runtime environments. That benchmark assigns shadow slots to 1,000 HTML comments without calling the Markdown serializer. With its exact loader and 1,000 warmed invocations under fixed-seed Callgrind, the base uses 9,291,132 instructions and this branch uses 9,291,125 (-7).

CodSpeed also flags serialize-inner and serialize-inner-indent at 5.1 to 5.4 ms and 5.5 to 5.9 ms. On their exact 234,774-byte WHATWG spec corpus, fresh CPython 3.15 release builds at base d4701ec6 and head 7ce46364 used 200 warmed calls after one parse. Two subsequent fixed-seed paired Callgrind runs counted 842,636,402 instructions on both revisions for serialize-inner and 926,219,073 on both for serialize-inner-indent. The first head runs counted 849,313,319 and 933,033,128 against the same base counts; the extra instructions were concentrated in memcpy, while serialize_compact_step self instructions were unchanged. Repeated paired counts do not support a persistent serializer regression.

CommonMark treats Unicode space separators as whitespace at emphasis
boundaries. The exporter treated them as ordinary characters and emitted
literal delimiter runs instead of emphasis.

Classify those spaces so the existing raw HTML fallback preserves the
element and its content. A direct lookup leaves the ASCII path unchanged.
@gaborbernat gaborbernat added the bug Something isn't working label Oct 10, 2026
@codspeed

codspeed Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 18.45%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

❌ 3 regressed benchmarks
✅ 579 untouched benchmarks
⏩ 32 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_feature[shadow-slot-comments] 84.3 µs 138.8 µs -39.3%
❌ test_feature[serialize-inner] 5.1 ms 5.4 ms -5.64%
❌ test_feature[serialize-inner-indent] 5.5 ms 5.9 ms -5.32%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing gaborbernat:fix/markdown-nbsp-emphasis (7ce4636) with main (d4701ec)

Open in CodSpeed

Footnotes

  1. 32 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@gaborbernat
gaborbernat merged commit c05d8d4 into tox-dev:main Oct 10, 2026
57 of 58 checks passed
@gaborbernat
gaborbernat deleted the fix/markdown-nbsp-emphasis branch October 10, 2026 21:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant