Skip to content

perf(core): *Async functions copy their input on the main thread #548

Description

@derodero24

Problem

All 23 *Async functions copy their input on the JS thread before they queue the task. Examples:

The copy runs synchronously inside the call. It blocks the event loop and doubles the input's memory footprint. For large inputs, which are what the async API is recommended for, a significant share of the work therefore stays on the main thread.

#339 investigated this and closed it on the basis that napi-rs Buffer is not Send. That no longer holds for the napi version this repo pins (=3.9.1):

  • Buffer and the typed arrays implement Send / Sync: unsafe impl Send for Buffer {} (napi-3.9.1/src/bindgen_runtime/js_values/buffer.rs, line 364), and the same for typed arrays in arraybuffer.rs, line 470.
  • Buffer holds a napi_ref to the JS object. When it is dropped on a worker thread, Drop hands the reference back to the JS thread through a thread-safe function. A task can therefore keep the caller's buffer alive without copying it.

There is a catch, which is why this needs a decision rather than a straight fix. napi-rs's own comment on that impl says that this is undefined behaviour if JS touches the buffer concurrently. Node-API offers no way to pin an ArrayBuffer against detachment.

Reproduction

Linux x64, 4 vCPU, Node 22.22.0, comprs 2.0.2 release build, random input.

const c = require('@derodero24/comprs');
const data = require('crypto').randomBytes(100 << 20);
(async () => {
  await c.zstdCompressAsync(data); // warm-up
  const t0 = performance.now();
  const promise = c.zstdCompressAsync(data);
  console.log('call returned after', (performance.now() - t0).toFixed(0), 'ms'); // JS thread blocked
  await promise;
})();
Call (100 MB) Time until the call returns (median of 5) Sync variant, for reference
zstdCompressAsync 67 ms 126 ms
lz4CompressAsync 65 ms 88 ms
gzipCompressAsync 59 ms 2,252 ms

Prototype: a stand-alone napi 3.9.1 addon calling comprs_core::lz4::compress, comparing the current pattern with a task that stores the Either<Buffer, Uint8Array> itself. 128 MB random input, three runs each:

Prototype tasks
pub struct CopyTask { data: Vec<u8> }            // current pattern
pub struct RefTask { data: Either<Buffer, Uint8Array> } // holds the JS buffer

#[napi]
pub fn copy_async(data: Either<Buffer, Uint8Array>) -> AsyncTask<CopyTask> {
    AsyncTask::new(CopyTask { data: as_bytes(&data).to_vec() })
}

#[napi]
pub fn ref_async(data: Either<Buffer, Uint8Array>) -> AsyncTask<RefTask> {
    AsyncTask::new(RefTask { data })
}
// Both `compute()`s call `comprs_core::lz4::compress(...)`; both `resolve()`s return `o.into()`.
Task JS thread blocked per call Time per call Peak RSS
copy (current) 82-107 ms 306-445 ms 591-597 MB
hold reference 0.0 ms 126-155 ms 464-465 MB

What happens if the caller interferes while the task runs:

Caller behaviour while the task runs Copy (current) Hold reference
Drops its last reference and GC runs Correct output Correct output; the napi_ref keeps the buffer alive
Mutates the buffer (buf.fill(2)) Correct output Output is a timing-dependent mix of old and new contents (2-34% old bytes in three runs)
Transfers the ArrayBuffer (ab.transfer()), drops the new owner and GC runs Correct output Segmentation fault: the backing store is freed while being read

For comparison, zlib.deflate() in Node 22.22.0 shows the same segmentation fault when its input ArrayBuffer is transferred mid-operation, so node:zlib has the same exposure.

Proposed fix

This needs a maintainer decision. Options:

  1. Keep the copy. Document the main-thread cost and point users with very large inputs to worker threads. Add a note to perf: investigate async task input data copy overhead #339 with the corrected reasoning.
  2. Hold a reference by default. Document that the input must not be mutated, transferred or detached until the promise settles, matching what node:zlib implicitly requires today. Violating that contract can crash the process.
  3. Make zero-copy opt-in. Keep copying by default. Add an opt-in for callers who own the buffer, e.g. an options object { copyInput: false } (this fits the unified v3 options API RFC, tracked in Tracking: 2026-10 audit follow-ups #535) or a separate function family.
  4. Hybrid. Copy below a size threshold, where the copy is cheap, and reference above it. This still carries the risks of option 2 for large inputs.

Whichever option is chosen:

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance benchmarks and optimization

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions