Repository navigation
stdout/stderr buffering considerations #6379
Description
Activity
- addeddiscussIssues opened for discussion and feedback.Issues opened for discussion and feedback.
on Apr 25, 2016 process.stdout._handle.setBlocking(true);as proposed in #1741 (comment) and #6297 (comment) would considerably hurt the performance. Exposing that isn't a good solution to any of this.- changed the title
[-]stdio buffering considerations[/-][+]stdout/stderr buffering considerations[/+]on Apr 25, 2016 - addednetIssues and PRs related to the net subsystem.Issues and PRs related to the net subsystem.processIssues and PRs related to the process subsystem.Issues and PRs related to the process subsystem.libuvIssues and PRs related to the libuv dependency or the uv binding.Issues and PRs related to the libuv dependency or the uv binding.
on Apr 25, 2016 process.stdout._handle.setBlocking(true);
would considerably hurt the performanceDevelopers are a smart bunch. With the proposed
setBlockingAPI they can decide for themselves if they want stdout/stderr to block on writes. The default could still remain non-blocking. There's no harm in giving people the option. Having CLI applications block on writes for stdout/stderr is quite desirable. This is not unlike all the Sync calls in thefsmodule. In a high-performance server scenario you don't have to use them, but they are there if you need them.Reacted by Peter Kelly, Tushar Pokle, Andrei Botezatu, Arel Cordero, Filip Jeretina, Arthur Wolf and Lars KappertReacted by Paul Irish and Harold HuntThere's no harm in giving people the option.
There is, every option comes at a cost from both sides. That option would need to be supported, and it would actually just toggle the behaviour between two states that are far from being ideal in many cases.
As for the option here, the said cyclic buffer maximum size could perhaps be configured using a command-line argument, at the buffer size 0 turning in into a fully blocking API.
toggle the behaviour between two states that are far from being ideal in many cases
The proposed API
process.stdout._handle.setBlocking(Boolean)could be documented to only be guaranteed to work before any output is sent to the stream.On UNIX it is a trivial system call on a fd. It's pretty much the same thing on Windows as far as I recall.
Edit: See #6297 (comment)
Reacted by Josh Junon and GPThe proposed API process.stdout._handle.setBlocking(Boolean) could be documented to only be guaranteed to work before any output is sent to the stream.
We should not be documenting underscore-prefixed properties. This would need to be a separate, true public API.
Reacted by kzc, Nikita Skovoroda, Jiewei Qian, GP and Claudia MeadowsIn CLI tools, I don't want to have stdout/stderr be blocking though, rather I want to be able to explicitly exit the process and have stdout/stderr flushed before exiting. Has there been any discussion about introducing a way to explicitly exit the process gracefully? Meaning flushing the stdout/stderr and then exit. Maybe something like a
process.gracefulExit()method? In ES2015 modules, we can no longerreturnin the top-scope, so short-circuiting will then be effectively impossible, without a nesting mess.Reacted by Sam Verschueren, Kevin Mårtensson, James Talmage, Scott Santucci, kuzyn, GP, Jiri Spac, Ruben Bridgewater, Jordan Wallet, bl-ue and 1 moreThe exit module is a drop in replacement for
process.exit()that effectively waits for stdout/stderr to be drained. Other streams can optionally be drained as well. Something like it really ought to rolled into node core with a different name.Edit: upon further analysis, the
exitmodule is no longer recommended.I think there's a slight discrepancy in interpretation here.
Are we talking about console blocking, or console buffering? Albeit similar, these are two different considerations.
Further, I'm not entirely sure what the problem is you're trying to fix, @ChALkeR. Could you elaborate on a problem that is currently being faced? I'm not groking what it is that's broken here.
One thing to note is that on most systems (usually UNIX systems)
stdoutis buffered to begin with, whereasstderris usually not (or sometimes just line buffered).However, an explicit
fflush(stdout)orfflush(stderr)is not necessary upon a process closing because these buffers are handled by the system and are flushed whenmain()returns. If the buffer grows too large during execution, the system automatically flushes it.This is why the following program still outputs while the loop is running, and on most terminal emulators the same reason why it looks like it's 'jittery'.
#include <stdio.h> int main(void) { for (int i = 10000000; i > 0; i--) { printf("This is a relatively not-short line of text. :) %d\n", i); } return 0; }
Blocking, on the other hand, doesn't make much sense. What do you want to block? Are you suggesting we
fflush(stdout)every time weconsole.log()? Because if so, that is a terrible idea.Developers are a smart bunch. With the proposed
setBlockingAPI they can decide for themselves if they want stdout/stderr to block on writes.I respectfully disagree. There is absolutely no point in blocking on stdout/stderr. It's very rarely done in the native world and when it is it's usually for concurrency problems (two
FILE*handles being written to by two threads that cannot otherwise be locked; i.e. TTY descriptors). Let the system handle it.The above program without
fflush(stdout)on every iteration produced the following average:real 6.32s sys 0.54 405 involuntary context switchesand with
fflush(stdout)on each iteration:real 7.48s sys 0.56s 1607 involuntary context switchesThats 18.3% of a time increase on this microbenchmark alone. I'm in the 'time is relative and isn't a great indicator of actual performance' camp, so I included the involuntary context switch count (voluntary context switches would be the same between the two runs, obviously).
The results show that
fflush(stdout)upon every message encouraged the system to switch out of the process 1202 more times - that's 296% higher than ye olde system buffering.Further, the feature that was requested by @sindresorhus is part of how executables are formed - at least on UNIX systems, all file descriptors that are buffered are flushed prior to/during
close()-ing orfclose()-ing, flushing prior to a process exiting is a needless step and exposing a new function for it would be completely unnecessary.
I'm still lost as to what is trying to be fixed here... I wouldn't mind seeing a
process.stdout.flush()method, but it'd be abused in upstream dependencies and I'd find myself writing some pretty annoyed issues in that case. I've never run into a problem with Node's use of standard I/O...Also, just to reiterate - please don't even start with a
.setBlocking()method. That's asking for trouble. The second a dependency makes standard I/O blocking then everything else is affected and then I have to answer to my technical lead about why my simple little program is switching out of context a bazillion times when it didn't before.Reacted by Sindre Sorhus, Evan Carroll, Evan Crane and Peter PistoriusFurther, I'm not entirely sure what the problem is you're trying to fix, @ChALkeR. Could you elaborate on a problem that is currently being faced? I'm not groking what it is that's broken here.
See the issues I linked above.
Blocking, on the other hand, doesn't make much sense. What do you want to block? Are you suggesting we fflush(stdout) every time we console.log()? Because if so, that is a terrible idea.
No, of course I don't.
console.logis not the same asprintfin your example.
It's an async operation. I don't proposefflush, I propose to somehow make sure that we don't fill up the async queue withconsole.logs.@ChALkeR are we actually seeing this happen?
var str = new Array(8000).join('.'); while (true) { console.log(str); }
works just fine. In what realistic scenario are we seeing the async queue fill up? The three linked issues are toy programs that are meant to break Node, and would break in most languages, in just about all situations, on all systems (resources are finite, obviously).
The way I personally see it, there are tradeoffs to having coroutine-like execution paradigms and having slightly higher memory usage is a perfectly acceptable tradeoff for the asynchronicity Node provides. It doesn't handle extreme cases because those cases are well outside any normal or even above normal scenario.
The first case, which mentions a break in execution to flush the async buffer, is perfectly reasonable and expected.
Side note: thanks for clarifying :) that makes more sense. I was hoping it wasn't seriously proposed we
fflush()all the time. Just had to make sure to kill that idea before anyone ran with it.44 remaining items
@gireeshpunathil Yes, that's the basic idea. The real implementation would be more involved: take a list of handles, a timeout, possible some flags.
Reacted by Gireesh Punathil, Jiri Spac, Ehsan Khakifirooz and bl-ueSeems like this ran its course and can be closed. If I'm wrong about that, don't hesitate to re-open.
Reacted by Julio Bastida@Trott, really? Nothing has been done, this is not at all resolved.
Reacted by Andrei Kazakou, Cole Carter and Alert Debug- removeddiscussIssues opened for discussion and feedback.Issues opened for discussion and feedback.
on Nov 10, 2018 - addedstdioIssues and PRs related to standard input, output, and error streams.Issues and PRs related to standard input, output, and error streams.
on Jun 25, 2020 The idea that this is "expected" for non blocking streams is fatally flawed: nobody expects logging or stdio to be non blocking, and if you've changed the file descriptor flags, that's an implementation detail that should be hidden from the outside world.
Moreover buffering constitutes an implicit promise that it will eventually be written, so at least at exit all buffers should be fully written out.
Likewise OOM should trigger memory recovery, including writing out buffers.
Delays are less important than broken promises.
Reacted by Justin M. Keyes, Claudia Meadows, Mikael Wistoft-Ibsen, Slava, silverwind, Alexis Tyler and Mike LangJust posting workaround for this issue for future reference:
for (const stream of [process.stdout, process.stderr]) { stream?._handle?.setBlocking?.(true); }
Before:
$ node -p 'process.stdout.write("x".repeat(5e6)); process.exit()' | wc -c 65536
After:
$ node -p 'for (const stream of [process.stdout, process.stderr]) stream?._handle?.setBlocking?.(true); process.stdout.write("x".repeat(5e6)); process.exit()' | wc -c 5000000
Reacted by Claudia Meadows, James Brock, rafalglowacz and Maja CieślakReacted by Karl Horky, Maja Cieślak, Jason Kraus and Brady HoltIt seems that the phrase "non blocking" has been used in this discussion to mean something different from the POSIX definition, and that may account for some of the confusion.
In POSIX terms, buffering is what happens within a process, while non blocking and asynchronous are two different ways of interacting with the kernel.
A non blocking (
O_NONBLOCKorO_NDELAY) kernel call will fail rather than block; this is certainly not what is envisioned by most participants in this discussion. (Moreover, it will only do so for streams that in principle can block for an arbitrary period, such as terminals, pipes, and network sockets; it does not affect writing to most ordinary files and block devices. )In contrast, an asynchronous kernel call will neither succeed nor fail, but rather it returns to the running process immediately and subsequently uses a signal to notify the process of the success or failure of the requested operation. That's also not what is being discussed here, but could be useful as an implementation detail for buffering in userspace.
What is wanted here is the equivalent of the C language's
setvbuffunction: the ability to change the buffer size from "infinite" down to something smaller.I would strongly urge avoiding using the term "blocking" when defining a new public method to control this, as it unnecessarily muddles the meaning of related terms cf
O_NONBLOCK.@kurahaupo - Irrespective of whether the usage of the term
blockingis consistent with POSIX definition or not (I don't know if that consistency is relevant here or not), its usage in this project's documentation is consistent: a call that effectively blocks the Node.js event loop's forward progress (by virtue of the current thread yielding to the kernel for actively processing I/O). And I guess our users had not much issues with that definition.Reacted by Alexis Tyler and Noah AllenIf I understand this issue correctly this is the place to post my feedback. I wrote a Node.js Native Messaging host that streams output from
child.stdout.on('data' (data) => {})to the browser withprocess.stdout.write(data).I observed that RSS constantly increases during usage. Perhaps someone in the Node.js realm can explain why this occurs and how to fix this memory leak https://thankful-butter-handball.glitch.me/.
This issue has been open for nearly 7 years. In the mean time mitigating work has been done and lots of discussion took place in other issues. I'm going to put this one out to pasture.
Reacted by Andrzej Wódkiewicz, Justin M. Keyes, Alexis Tyler and Jonas ZeigerFor people that are looking for a workarounds, many are given here: https://unix.stackexchange.com/questions/25372/turn-off-buffering-in-pipe.
The one I choose was https://unix.stackexchange.com/questions/25372/turn-off-buffering-in-pipe/25378#25378, which is to wrap the command you want to execute with the 'script' command, which has a way to unbuffer the stdout.
script -q /dev/null long_running_command | print_progress # (FreeBSD, Mac OS X) script -q -c "long_running_command" /dev/null | print_progress # (Linux)All the other solutions in that linked thread required installing a new program, but the
scriptcommand is available on most POSIX-ish systems, albeit with different cmd line arg styles. This works for me - hopefully it can work for you.Reacted by Justin M. Keyes and Chirayu Krishnappa
I tried to discuss this some time ago at IRC, but postponed it for quite a long time. Also I started the discussion of this in #1741, but I would like to extract the more specific discussion to a separate issue.
I could miss some details, but will try to give a quick overview here.
Several issues here:
console.log(e.g. calling it in a loop) could chew up all the memory and die — Why does node/io.js hang when printing to the console inside a loop? #1741, Silly program or memory leak? #2970, Strange memory leaks #3171, memory leaks of console.log #18013.console.loghas different behavior while printing to a terminal and being redirected to a file. — Why does node/io.js hang when printing to the console inside a loop? #1741 (comment).As I understand it — the output has an implicit write buffer (as it's non-blocking) of unlimited size.
One approach to fixing this would be to:
For almost all cases, except for the ones that are currently broken, this would behave as a non-blocking buffer (because writes to the buffer are considerably faster than writes from the buffer to file/terminal).
For cases when the data is being piped to the output too quickly and when the output file/terminal does not manage to output it at the same rate — the write would turn into a blocking operation. It would also be blocking at the exit until all the data is written.
Another approach would be to monitor (and limit) the size of data that is contained in the implicit buffer coming from the async queue, and make the operations block when that limit is reached.