Skip to content

TextDecoder is wrong and very slow #61041

Description

@ChALkeR

Correctness

Encodings that return invalid results:

  • Single-byte:
    • ibm866 (fails at even ascii input)
    • koi8-u
    • windows-874
    • windows-1252
    • windows-1253
    • windows-1255
  • Multi-byte (all except gb18030):
    • gbk (should be identical to gb18030 but it is instead broken)
    • big5
    • euc-jp
    • iso-2022-jp
    • shift_jis (fails at even ascii input)
    • euc-kr

Unimplemented encodings that throw:

  • iso-8859-16
  • x-user-defined

If built without icu, utf-16le encoding also returns invalid results:

> new TextDecoder('utf-16le').decode(Uint16Array.of(0xd800))
'�' // correct
'\ud800' // no ICU

Performance

  • utf-8 (aka default) TextDecoder is much slower on ascii input than it can and should be
    1.3x on 4096 bytes, ~3x on 1 MiB input
  • The above applies to buffer.toString() too
    It's much slower on ASCII input than a checked js impl (same 1.3x-3x)
  • windows-1252 aka new TextDecoder('ascii') aka new TextDecoder('latin1')
    is ~2x-4x slower than an optimized impl on ascii input
  • windows-1252 aka new TextDecoder('latin1')
    is ~6x-12x slower than an optimized impl on latin1 input
  • windows-1252 is ~7x-12x slower than an optimized js impl
  • Other single-byte encodings that are significantly slower than js impl even on non-ascii input:
    iso-8859-3, iso-8859-6, iso-8859-7, iso-8859-8, iso-8859-8-i, windows-1253, windows-1255, windows-1257
  • None of the single-byte encodings are faster than the js impl even on non-ascii input
  • All of the single-byte encodings except windows-1252 are >=10x slower than the js impl on ascii input
    (windows-1252 is only ~2-4x slower)

References

Nothing of the above requires any changes on the native side, I compared to a somewhat optimized JS implementation

See https://docs.google.com/spreadsheets/d/1pdEefRG6r9fZy61WHGz0TKSt8cO4ISWqlpBN5KntIvQ/edit

See tests in https://github.com/ExodusOSS/bytes/blob/master/tests/encoding/mistakes.test.js (comment out the import and it can be run on Node.js without deps with only that file)

Suggestions

  1. Add a proper ASCII fast path to buffer.toString()
    src: improve StringBytes::Encode perf on ASCII #61119
  2. Add a proper ASCII fast path to new TextDecoder().decode(arg)
    src: improve StringBytes::Encode perf on ASCII #61119
  3. Perhaps replace single-byte decoders with a js impl, remove native paths and lib usage. They all are just mappers, the implementation for all of them is identical
    Or at least replace the slow, unsupported, or invalid ones.
    lib: implement all 1-byte encodings in js #61093
  4. Remove gbk decoder path and make it do the same as gb18030 as the spec says
    lib: gbk decoder is gb18030 decoder per spec #61099
  5. For utf16 decode optimistically using existing fast apis, then check the string for validity
    lib: add utf16 fast path for TextDecoder #61559
  6. Fix bugs in the non-ICU codepath
    lib: unify ICU and no-ICU TextDecoder #61409
    lib: use utf8 fast path for streaming TextDecoder #61549
    lib: add utf16 fast path for TextDecoder #61559
  7. Fix or replace implementations for big5, euc-jp, iso-2022-jp, shift_jis, euc-kr
    To fix legacy multi-byte decoders, attempt to re-use what Chromium has or import js code from @exodus/bytes

Activity

  1. ChALkeR commented on Dec 13, 2025

    @ChALkeR
    MemberAuthor

    Here is a comparison to an implementation that works and passes WPT and extra tests:

    Image

    Upd: results updated.

    Old ones Image
  2. ChALkeR commented on Dec 13, 2025

    @ChALkeR
    MemberAuthor

    cc @nodejs/performance perhaps

  3. changed the title [-]TextDecoder is wrong and slow[/-] [+]TextDecoder is wrong and very slow[/+] on Dec 13, 2025
  4. ChALkeR commented on Dec 14, 2025

    @ChALkeR
    MemberAuthor

    Status update: I'm now at a point "Chrome decodes fetch responses wrong in await res.text()" for utf-8
    https://issues.chromium.org/issues/468458744

  5. added
    performanceIssues and PRs related to the performance of Node.js.
    on Dec 14, 2025
  6. lemire commented on Dec 15, 2025

    @lemire
    Member

    I published my own benchmark for UTF-8 decoding.

    Here are my results.

    Node 24 with Apple M4

    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 19.69 GiB/s 0.004 ms
    Arabic lipsum 79.771 KiB 0.40 GiB/s 0.193 ms
    Chinese lipsum 68.203 KiB 0.45 GiB/s 0.144 ms

    Bun 1.3.4 with Apple M4

    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 55.83 GiB/s 0.002 ms
    Arabic lipsum 79.771 KiB 2.51 GiB/s 0.031 ms
    Chinese lipsum 68.203 KiB 5.69 GiB/s 0.012 ms

    Doing profiling on the Node code, I get the following results...

      49.96%  MainThread       node                  [.] v8::internal::Utf8DecoderBase<v8::internal::Utf8Decoder>::Utf8DecoderBase(v8::base::Vector<unsigned char const>)
      30.33%  MainThread       node                  [.] void v8::internal::Utf8DecoderBase<v8::internal::Utf8Decoder>::Decode<unsigned short>(unsigned short*, v8::base::Vector<unsigned char const>)
       9.21%  MainThread       libc.so.6             [.] __memmove_evex_unaligned_erms
       1.17%  MainThread       node                  [.] v8::internal::(anonymous namespace)::IterateObjectCache(v8::internal::Isolate*, std::vector<v8::internal::Tagged<v8::internal::Object>, std::allocator<v8::i
       0.58%  MainThread       node                  [.] v8::internal::RootScavengeVisitor::VisitRootPointer(v8::internal::Root, char const*, v8::internal::FullObjectSlot)
    

    If I trust this output, then I have to conclude that Node is bottlenecked by v8.

  7. ChALkeR commented on Dec 16, 2025

    @ChALkeR
    MemberAuthor

    @lemire Thanks for the benchmark!

    I added one line:

    import { TextDecoder, TextEncoder } from '@exodus/bytes/encoding.js'

    Then ran your benchmark as-is per instructions

    On Node.js v25.2.1, without it:

    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 17.35 GiB/s 0.006 ms
    Arabic lipsum 79.771 KiB 0.26 GiB/s 0.305 ms
    Chinese lipsum 68.203 KiB 0.32 GiB/s 0.207 ms

    On Node.js v25.2.1, with it:

    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 33.83 GiB/s 0.003 ms
    Arabic lipsum 79.771 KiB 0.27 GiB/s 0.281 ms
    Chinese lipsum 68.203 KiB 0.33 GiB/s 0.198 ms

    I see a 2x improvement on ASCII with 0 native code involved

  8. jasnell commented on Dec 16, 2025

    @jasnell
    Member

    @srl295 ... just fyi

  9. ChALkeR commented on Dec 16, 2025

    @ChALkeR
    MemberAuthor

    The impl for utf8 encoder/decoder for Node.js this uses is here: https://github.com/ExodusOSS/bytes/blob/master/utf8.node.js

    See comment at: https://github.com/ExodusOSS/bytes/blob/4b758ba6aa7171efec77053e9170bb2ddbdabba2/utf8.node.js#L40-L42

    Moreover, this could be made even better for worst-case and win even in those too with minor changes in native side by replacing isAscii call with a method that returns the position of the first non-ASCII char (or the length of the ASCII prefix) instead of true/false, see the logic here

    The difference on ASCII is even more significant in { fatal: true } mode:

    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 15.03 GiB/s 0.006 ms
    Arabic lipsum 79.771 KiB 0.26 GiB/s 0.290 ms
    Chinese lipsum 68.203 KiB 0.32 GiB/s 0.207 ms
    Test Size Throughput Mean Time
    Latin lipsum (ASCII) 84.902 KiB 33.80 GiB/s 0.003 ms
    Arabic lipsum 79.771 KiB 0.26 GiB/s 0.291 ms
    Chinese lipsum 68.203 KiB 0.32 GiB/s 0.205 ms

    There is no need for a slowdown in fatal mode when we are in ASCII fast path, but Node.js looses an additional 15% there (on top of the ~2x for non-fatal mode)

  10. ChALkeR commented on Dec 17, 2025

    @ChALkeR
    MemberAuthor

    Also utf16-le decoder returns invalid results in Node.js without ICU, updated the list of issues: #61041 (comment)

  11. ChALkeR commented on Dec 17, 2025

    @ChALkeR
    MemberAuthor

    I think that #60893 caused a very significant perf degradation in main (which is not yet accounted for in the table above)

    Why did we chose to keep it instead of a revert? It just hurts performance

  12. ChALkeR commented on Dec 20, 2025

    @ChALkeR
    MemberAuthor
  13. bjohansebas commented on Dec 28, 2025

    @bjohansebas
    Member

    While it’s great that this gets fixed upstream, it would also be great if a package with that implementation could be created, similar to what was done with the stream module. Otherwise, there will be inconsistencies for packages, for example, iconv-lite will expose those inconsistencies once the version with TextDecoder is released.

  14. ChALkeR commented on Dec 28, 2025

    @ChALkeR
    MemberAuthor

    @bjohansebas that package was already created. @exodus/bytes/encoding.js implements all encodings per spec and provides zero-dep TextDecoder / TextEncoder APIs even on barebone engines.
    That is also faster than iconv-lite 😉

    I'm upstreaming fixes in Node.js and was filing issues in browsers after making a stand-alone impl to compare to.

  15. github-actions commented on Jul 20, 2026

    @github-actions
    Contributor

    This issue has been marked as stale due to 90 days of inactivity.
    It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.

  16. added
    staleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.
    on Jul 20, 2026
  17. SukkaW commented on Jul 28, 2026

    @SukkaW

    Unstale. This is still replicable on Node.js 26, and if IIUC this is still ongoing by the awesome team.

  18. removed
    staleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.
    on Jul 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    confirmed-bugIssues and PRs for confirmed bugs.performanceIssues and PRs related to the performance of Node.js.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions