You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf(streams): stream transforms compress and decompress synchronously on the main thread #554
The stream APIs do all compression work synchronously on the JS thread:
Every ctx.transform() / flush() / finish() is a synchronous napi call, made inline from the Web TransformStream callbacks in streams.js (e.g. createBrotliCompressStream) and from the Node Transform callbacks in node.js (e.g. createZstdCompressTransform).
Each chunk blocks the event loop for its full compression time.
When the source is in memory, the transform callbacks complete synchronously, so the stream machinery moves on to the next chunk through microtasks and process.nextTick. Timers and I/O do not run until the whole stream has been processed.
node:zlib streams run each chunk on the libuv thread pool, so the event loop keeps turning.
The README does not mention this. Its "Choosing an API mode" table (README.md#L200-L207) recommends Async "when event loop must stay free (servers, UIs)" and Streaming for "unknown/unbounded data size". Users of the streaming APIs in servers, including @derodero24/comprs-middleware, can reasonably assume that streaming does not block either.
Reproduction
Linux x64, 4 vCPU, Node 22.22.0, comprs 2.0.2 release build. The input is 16 MB of random words from a 4,096-word vocabulary. A 1 ms setInterval counts how often the event loop got to run timers:
comprs Node transform, 256 × 64 KiB from fs.createReadStream
5,766 ms
103
134 ms
node:zlib, 256 × 64 KiB from fs.createReadStream
5,305 ms
4,622
12 ms
comprs Web stream (createBrotliCompressStream(9)), 256 × 64 KiB from memory
5,558 ms
0
5,558 ms
The LZ4 decompression stream is the extreme case. It buffers all input and decodes everything in flush() (tracked separately in #535). For 100 MB of output, that flush() alone blocked for 124-142 ms in three runs.
Proposed fix
Async methods. Add transformAsync(chunk) / flushAsync() / finishAsync() to the context classes, implemented as AsyncTasks:
Keep the codec state in an Arc<Mutex<...>> (or move it into the task and back), so compute() runs on the thread pool.
The JS wrappers already serialize calls per stream, so the lock is uncontended.
Reject calls made while another call is still in flight, instead of blocking on the lock.
Wrappers. Use the async methods from the Node transform / flush callbacks (call callback when the promise settles) and from the Web transform / flush callbacks (return the promise).
Small chunks. Thread-pool dispatch costs tens of microseconds per call. The wrappers could keep the synchronous path for chunks below a threshold, e.g. 16-64 KiB, as long as the total synchronous work per tick stays bounded.
Problem
The stream APIs do all compression work synchronously on the JS thread:
ctx.transform()/flush()/finish()is a synchronous napi call, made inline from the WebTransformStreamcallbacks instreams.js(e.g.createBrotliCompressStream) and from the NodeTransformcallbacks innode.js(e.g.createZstdCompressTransform).process.nextTick. Timers and I/O do not run until the whole stream has been processed.node:zlibstreams run each chunk on the libuv thread pool, so the event loop keeps turning.The README does not mention this. Its "Choosing an API mode" table (README.md#L200-L207) recommends Async "when event loop must stay free (servers, UIs)" and Streaming for "unknown/unbounded data size". Users of the streaming APIs in servers, including
@derodero24/comprs-middleware, can reasonably assume that streaming does not block either.Reproduction
Linux x64, 4 vCPU, Node 22.22.0, comprs 2.0.2 release build. The input is 16 MB of random words from a 4,096-word vocabulary. A 1 ms
setIntervalcounts how often the event loop got to run timers:node:zlib, 1 chunk of 16 MBnode:zlib, 256 × 64 KiB from memoryfs.createReadStreamnode:zlib, 256 × 64 KiB fromfs.createReadStreamcreateBrotliCompressStream(9)), 256 × 64 KiB from memoryThe LZ4 decompression stream is the extreme case. It buffers all input and decodes everything in
flush()(tracked separately in #535). For 100 MB of output, thatflush()alone blocked for 124-142 ms in three runs.Proposed fix
transformAsync(chunk)/flushAsync()/finishAsync()to the context classes, implemented asAsyncTasks:Arc<Mutex<...>>(or move it into the task and back), socompute()runs on the thread pool.transform/flushcallbacks (callcallbackwhen the promise settles) and from the Webtransform/flushcallbacks (return the promise).*Asyncfunctions (perf(core): *Async functions copy their input on the main thread #548).*Asyncor a worker thread for CPU-heavy settings such as brotli quality ≥ 9.Related
*Asyncinput copy.close()/ cancellation in mind, e.g. close only after an in-flight task finishes.