Repository navigation
TextDecoder is wrong and very slow #61041
Description
Activity
cc @nodejs/performance perhaps
- changed the title
[-]TextDecoder is wrong and slow[/-][+]TextDecoder is wrong and very slow[/+]on Dec 13, 2025 Status update: I'm now at a point "Chrome decodes fetch responses wrong in
await res.text()" for utf-8
https://issues.chromium.org/issues/468458744Reacted by René, Benjamin Gruenbaum and Sukka- addedperformanceIssues and PRs related to the performance of Node.js.Issues and PRs related to the performance of Node.js.
on Dec 14, 2025 I published my own benchmark for UTF-8 decoding.
Here are my results.
Node 24 with Apple M4
Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 19.69 GiB/s 0.004 ms Arabic lipsum 79.771 KiB 0.40 GiB/s 0.193 ms Chinese lipsum 68.203 KiB 0.45 GiB/s 0.144 ms Bun 1.3.4 with Apple M4
Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 55.83 GiB/s 0.002 ms Arabic lipsum 79.771 KiB 2.51 GiB/s 0.031 ms Chinese lipsum 68.203 KiB 5.69 GiB/s 0.012 ms Doing profiling on the Node code, I get the following results...
49.96% MainThread node [.] v8::internal::Utf8DecoderBase<v8::internal::Utf8Decoder>::Utf8DecoderBase(v8::base::Vector<unsigned char const>) 30.33% MainThread node [.] void v8::internal::Utf8DecoderBase<v8::internal::Utf8Decoder>::Decode<unsigned short>(unsigned short*, v8::base::Vector<unsigned char const>) 9.21% MainThread libc.so.6 [.] __memmove_evex_unaligned_erms 1.17% MainThread node [.] v8::internal::(anonymous namespace)::IterateObjectCache(v8::internal::Isolate*, std::vector<v8::internal::Tagged<v8::internal::Object>, std::allocator<v8::i 0.58% MainThread node [.] v8::internal::RootScavengeVisitor::VisitRootPointer(v8::internal::Root, char const*, v8::internal::FullObjectSlot)If I trust this output, then I have to conclude that Node is bottlenecked by v8.
Reacted by Nikita Skovoroda, Mert Can Altin and Sukka@lemire Thanks for the benchmark!
I added one line:
import { TextDecoder, TextEncoder } from '@exodus/bytes/encoding.js'
Then ran your benchmark as-is per instructions
On Node.js v25.2.1, without it:
Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 17.35 GiB/s 0.006 ms Arabic lipsum 79.771 KiB 0.26 GiB/s 0.305 ms Chinese lipsum 68.203 KiB 0.32 GiB/s 0.207 ms On Node.js v25.2.1, with it:
Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 33.83 GiB/s 0.003 ms Arabic lipsum 79.771 KiB 0.27 GiB/s 0.281 ms Chinese lipsum 68.203 KiB 0.33 GiB/s 0.198 ms I see a 2x improvement on ASCII with 0 native code involved
Reacted by Sukka@srl295 ... just fyi
The impl for utf8 encoder/decoder for Node.js this uses is here: https://github.com/ExodusOSS/bytes/blob/master/utf8.node.js
See comment at: https://github.com/ExodusOSS/bytes/blob/4b758ba6aa7171efec77053e9170bb2ddbdabba2/utf8.node.js#L40-L42
Moreover, this could be made even better for worst-case and win even in those too with minor changes in native side by replacing
isAsciicall with a method that returns the position of the first non-ASCII char (or the length of the ASCII prefix) instead of true/false, see the logic hereThe difference on ASCII is even more significant in
{ fatal: true }mode:Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 15.03 GiB/s 0.006 ms Arabic lipsum 79.771 KiB 0.26 GiB/s 0.290 ms Chinese lipsum 68.203 KiB 0.32 GiB/s 0.207 ms Test Size Throughput Mean Time Latin lipsum (ASCII) 84.902 KiB 33.80 GiB/s 0.003 ms Arabic lipsum 79.771 KiB 0.26 GiB/s 0.291 ms Chinese lipsum 68.203 KiB 0.32 GiB/s 0.205 ms There is no need for a slowdown in fatal mode when we are in ASCII fast path, but Node.js looses an additional 15% there (on top of the ~2x for non-fatal mode)
Reacted by SukkaAlso
utf16-ledecoder returns invalid results in Node.js without ICU, updated the list of issues: #61041 (comment)I think that #60893 caused a very significant perf degradation in main (which is not yet accounted for in the table above)
Why did we chose to keep it instead of a revert? It just hurts performance
WPT PR: web-platform-tests/wpt#56892
While it’s great that this gets fixed upstream, it would also be great if a package with that implementation could be created, similar to what was done with the
streammodule. Otherwise, there will be inconsistencies for packages, for example,iconv-litewill expose those inconsistencies once the version withTextDecoderis released.@bjohansebas that package was already created.
@exodus/bytes/encoding.jsimplements all encodings per spec and provides zero-dep TextDecoder / TextEncoder APIs even on barebone engines.
That is also faster than iconv-lite 😉I'm upstreaming fixes in Node.js and was filing issues in browsers after making a stand-alone impl to compare to.
Reacted by Carlos Fuentesgithub-actions commented
on Jul 20, 2026 on Jul 20, 2026 – with GitHub ActionsContributorMore actionsThis issue has been marked as stale due to 90 days of inactivity.
It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.- addedstaleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.Issues and PRs marked stale due to inactivity and scheduled for automatic closure.
on Jul 20, 2026 Unstale. This is still replicable on Node.js 26, and if IIUC this is still ongoing by the awesome team.
- removedstaleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.Issues and PRs marked stale due to inactivity and scheduled for automatic closure.
on Jul 29, 2026 - addedconfirmed-bugIssues and PRs for confirmed bugs.Issues and PRs for confirmed bugs.
on Jul 31, 2026 - added 2 commits that reference this issue
on Sep 4, 2026


Correctness
Encodings that return invalid results:
ibm866(fails at even ascii input)koi8-uwindows-874windows-1252windows-1253windows-1255gb18030):gbk(should be identical togb18030but it is instead broken)big5euc-jpiso-2022-jpshift_jis(fails at even ascii input)euc-krUnimplemented encodings that throw:
iso-8859-16x-user-definedIf built without
icu,utf-16leencoding also returns invalid results:Performance
utf-8(aka default)TextDecoderis much slower on ascii input than it can and should be1.3xon 4096 bytes,~3xon 1 MiB inputbuffer.toString()tooIt's much slower on ASCII input than a checked js impl (same
1.3x-3x)windows-1252akanew TextDecoder('ascii')akanew TextDecoder('latin1')is ~
2x-4xslower than an optimized impl on ascii inputwindows-1252akanew TextDecoder('latin1')is ~
6x-12xslower than an optimized impl on latin1 inputwindows-1252is ~7x-12xslower than an optimized js impliso-8859-3,iso-8859-6,iso-8859-7,iso-8859-8,iso-8859-8-i,windows-1253,windows-1255,windows-1257windows-1252are>=10xslower than the js impl on ascii input(
windows-1252is only ~2-4xslower)References
Nothing of the above requires any changes on the native side, I compared to a somewhat optimized JS implementation
See https://docs.google.com/spreadsheets/d/1pdEefRG6r9fZy61WHGz0TKSt8cO4ISWqlpBN5KntIvQ/edit
See tests in https://github.com/ExodusOSS/bytes/blob/master/tests/encoding/mistakes.test.js (comment out the import and it can be run on Node.js without deps with only that file)
Suggestions
buffer.toString()src: improve StringBytes::Encode perf on ASCII #61119
new TextDecoder().decode(arg)src: improve StringBytes::Encode perf on ASCII #61119
Or at least replace the slow, unsupported, or invalid ones.
lib: implement all 1-byte encodings in js #61093
gbkdecoder path and make it do the same asgb18030as the spec sayslib: gbk decoder is gb18030 decoder per spec #61099
lib: add utf16 fast path for TextDecoder #61559
lib: unify ICU and no-ICU TextDecoder #61409
lib: use utf8 fast path for streaming TextDecoder #61549
lib: add utf16 fast path for TextDecoder #61559
big5,euc-jp,iso-2022-jp,shift_jis,euc-krTo fix legacy multi-byte decoders, attempt to re-use what Chromium has or import js code from
@exodus/bytes