Skip to content

Commit 729cce1

Browse files
committed
fix(cli): make a failed stderr write non-fatal on the published entry
`bin/run.js` writes to `process.stderr` with no `error` listener, so an `error` event there is an uncaught exception. #15564 was filed NOT REPRODUCED, and its fence was explicit: symmetry with `bin/run-dev.js` is not evidence, establish reachability first. Both of the card's probes were re-run before anything was written here and both still read clean — exit 2, no `uncaughtException`, 3/3 each. They were not a guard, they were the wrong lifecycle: everything a failing invocation puts on stderr is written after `run()` has settled, by `handle()`, which exits on top of its own report, and a failing write reports through libuv's completion callback that a synchronous exit is never told about. Widening to a lifecycle that outlives its first failed write reaches it. `os serve` on `examples/app-todo`, read end destroyed: uncaughtException code=EPIPE msg=write EPIPE at afterWriteDispatched (node:internal/stream_base_commons:159:15) exit code=1 3 of 3 runs, 3049-3433 ms in — the same frame and status #14858 traced on the dev shim. The same child read by a draining parent boots, serves, and exits 0 after 7926 bytes over 16.6 s. Pinned with the hazard manufactured inside the published binary's own process, alongside a live positive control that removes the listener in the child and re-crashes it, so the guarded arm's silence is a reading. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFY46JydE1gMxQG1TqBcMZ
1 parent d5c4022 commit 729cce1

4 files changed

Lines changed: 445 additions & 0 deletions

File tree

‎.changeset/tidy-cars-repeat.md‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
---
2+
'@objectstack/cli': patch
3+
---
4+
5+
Stop the published CLI from dying of an uncaught `write EPIPE` when its caller's stderr read end is gone.
6+
7+
`bin/run.js` — the file `bin.objectstack` / `bin.os` point at, and the only thing under `bin/` npm packs — now attaches the same no-op `error` listener to `process.stderr` that the in-repo dev shim has carried since the original finding. `process.stderr` is an `EventEmitter`, so an `error` event with nothing listening is an uncaught exception.
8+
9+
Measured on the published entry with the read end destroyed (`stdio: ['ignore','ignore','pipe']`, then `child.stderr.destroy()`), traced with an observer that installs no listener and wraps no write:
10+
11+
```
12+
uncaughtException code=EPIPE msg=write EPIPE
13+
at afterWriteDispatched (node:internal/stream_base_commons:159:15)
14+
exit code=1
15+
```
16+
17+
3 of 3 runs, 3049-3433 ms in, on `os serve` over `examples/app-todo`. Read by a draining parent the same child boots and serves and exits 0, having written 7926 bytes over 16.6 s — so the crash was costing the run at its first diagnostic line and 20 of its 21 stderr writes. Failing invocations do not reach it: everything they put on stderr is written after `run()` has settled, by a handler that exits on top of its own report.
18+
19+
Behaviour change worth knowing about: a long-running command (`os serve`, `os dev`, `os start`) whose reader has gone now keeps running and reports its own exit status, instead of dying on its first diagnostic write. A supervisor that destroyed the read end and relied on that crash to end the child needs to end it itself.

‎packages/cli/bin/run.js‎

Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -70,6 +70,77 @@ try {
7070
// Unbuilt or half-built tree — nothing to install and nothing to say.
7171
}
7272

73+
/**
74+
* Make a FAILED stderr write non-fatal, so a caller whose read end is gone
75+
* still gets this CLI's own exit status instead of a crash. #14858, reached on
76+
* THIS entry point by the #15564 measurement.
77+
*
78+
* `process.stderr` is an `EventEmitter`, and an `error` event with nothing
79+
* listening IS an uncaught exception. `bin/run-dev.js` has carried this
80+
* listener since #14858; the published entry did not, and #15564 was filed
81+
* NOT REPRODUCED because the two probes that had been run against it — a
82+
* bad command id, and `OBJECTSTACK_DEBUG=1` over an unbuilt `@objectstack/spec`
83+
* — both answered exit 2 with no `uncaughtException`. Re-run here, they still
84+
* do (3/3 each, 57 and 35528 bytes drained). ⭐ They were not a guard; they
85+
* were the wrong lifecycle, and the difference is measurable rather than
86+
* arguable:
87+
*
88+
* leg (bin/run.js, read end destroyed) stderr writes exit
89+
* --------------------------------------- ------------- ----------------
90+
* `definitely-not-a-command` 1 @ 3231 ms 2, no crash
91+
* OBJECTSTACK_DEBUG=1 + unbuilt spec 60 @ 932-960 ms 2, no crash
92+
* `serve objectstack.config.ts` 21 @ 3180 ms on 1, `write EPIPE`
93+
* 3/3
94+
*
95+
* Two things separate the last row, and BOTH are needed:
96+
*
97+
* • an event-loop TURN between the failing write and `process.exit`. A
98+
* failing write reports through libuv's completion callback, so a write
99+
* followed by a synchronous exit is never told. Both probe legs are that
100+
* shape: everything they put on stderr is written after `run()` has already
101+
* settled, by `handle()`, which exits on top of its own report — measured
102+
* at one write 1 ms before exit, and at 59 warning blocks whose EPIPE
103+
* arrives synchronously inside the write.
104+
* • a RAW `process.stderr.write`. Node's `console.error` carries
105+
* `ignoreErrors`, which parks a temporary `error` listener across the write
106+
* — so oclif's warning blocks cannot crash this process at any size
107+
* (measured: 1 MiB through `console.error` does not, one line through
108+
* `process.stderr.write` does, 3/3 each).
109+
*
110+
* `os serve` is both: `printDiagnostic` in `src/commands/serve.ts` writes
111+
* straight to stderr (#7915) and the boot around it is asynchronous, so the
112+
* process is alive across the whole sequence. Measured on `examples/app-todo`
113+
* through this file, read end destroyed (`stdio: ['ignore','ignore','pipe']`,
114+
* then `child.stderr.destroy()`), traced with a `--import` observer that
115+
* installs NO listener here and wraps no write:
116+
*
117+
* uncaughtException code=EPIPE msg=write EPIPE
118+
* at afterWriteDispatched (node:internal/stream_base_commons:159:15)
119+
* exit code=1
120+
*
121+
* 3 of 3 runs, 3049-3433 ms in — the same frame and the same status #14858
122+
* traced on the dev shim. The same child read by a draining parent boots and
123+
* serves, exit 0 at a 20 s SIGTERM, having written 7926 bytes over 16.6 s. So
124+
* the crash costs the run at its FIRST diagnostic line and 20 of its 21 stderr
125+
* writes, on the entry point a customer's install actually runs (`files` names
126+
* only `dist`, but npm packs a `bin` target regardless — #14874).
127+
*
128+
* ⛔ Deliberately NOT narrowed to `error.code === 'EPIPE'`, for the reason
129+
* `bin/run-dev.js` records: the reason to tolerate is not WHICH error it is.
130+
* Every event here means one thing — a write to stderr failed — the only
131+
* channel it could be reported on is the stream that just failed, and there is
132+
* no other action to take.
133+
*
134+
* ⚠️ What it costs: a long-running command whose reader has gone now keeps
135+
* running instead of dying on its first diagnostic. That is the point (the
136+
* server is still serving, and its caller still gets the CLI's own status), but
137+
* it is a real behaviour change for a supervisor that destroyed the read end
138+
* and relied on the crash to end the child.
139+
*/
140+
process.stderr.on('error', () => {
141+
// Nothing to report, and nowhere left to report it.
142+
});
143+
73144
await run(process.argv.slice(2), import.meta.url)
74145
.then(async (result) => {
75146
flush();
Lines changed: 121 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,121 @@
1+
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.
2+
3+
/**
4+
* The #14858 crash class, manufactured INSIDE the published entry point's own
5+
* process — driven by `published-entry-stderr-error-listener.e2e.test.ts`.
6+
*
7+
* Loaded with `node --import <this> bin/run.js …` against a read end the parent
8+
* has destroyed, so everything below runs in the same process as the shipped
9+
* CLI, on the same open file description, after `bin/run.js` has had its chance
10+
* to attach the `error` listener.
11+
*
12+
* ## Why the failing write is manufactured rather than taken from a command
13+
*
14+
* The field reproduction is `os serve`: `printDiagnostic` writes straight to
15+
* stderr (#7915), the boot around it is asynchronous, and #15564 measured the
16+
* published entry dying there — `write EPIPE` at `afterWriteDispatched`, exit
17+
* 1, 3 of 3 runs, 3049-3433 ms in. Reproducing THAT needs a fixture app, a
18+
* database, a bound port and four seconds per leg, and it pins the crash to one
19+
* command that could stop writing raw tomorrow. What the entry point owes is
20+
* narrower and does not move: **a failed stderr write in this process must not
21+
* be fatal.** One raw write to a destroyed pipe is the whole of that hazard,
22+
* and it costs milliseconds.
23+
*
24+
* ⛔ The write is deliberately `process.stderr.write` and NOT `console.error`.
25+
* Node's `console.error` carries `ignoreErrors`, which parks a temporary
26+
* `error` listener across the write, so it cannot crash a process at any
27+
* payload size — measured at 1 MiB, 0 of 3, against one line through
28+
* `process.stderr.write` at 3 of 3. A probe written with `console.error` would
29+
* be green with the listener REMOVED, which is the one thing it must not be.
30+
*
31+
* ## The two arms, and why the unguarded one is not an ablation
32+
*
33+
* `OS_PUBLISHED_ENTRY_ERROR_PROBE_ARM=unguarded` makes this probe remove the
34+
* entry's listener in its own process before writing. That is the harness's
35+
* LIVE POSITIVE CONTROL: it shows, in the same run and against the same tree,
36+
* that this instrument can still see the crash — so the guarded arm's silence
37+
* is a reading rather than a zero. Nothing on disk is touched, so it costs no
38+
* restore and cannot leave a mutated tree behind.
39+
*
40+
* Markers go to a file: stderr is the thing under test and, on the arm that is
41+
* supposed to fail, the thing that is already broken.
42+
*
43+
* env: `OS_PUBLISHED_ENTRY_ERROR_PROBE_MARKS` — the marker file.
44+
* env: `OS_PUBLISHED_ENTRY_ERROR_PROBE_ARM` — `guarded` (default) | `unguarded`.
45+
*/
46+
47+
import { appendFileSync } from 'node:fs';
48+
49+
const MARKS = process.env.OS_PUBLISHED_ENTRY_ERROR_PROBE_MARKS;
50+
const UNGUARDED = process.env.OS_PUBLISHED_ENTRY_ERROR_PROBE_ARM === 'unguarded';
51+
52+
const mark = (line) => appendFileSync(MARKS, `${line}\n`);
53+
54+
/**
55+
* How long to wait for `bin/run.js` to attach its listener before proceeding
56+
* anyway.
57+
*
58+
* A CONSTANT, and far above anything the attach legitimately needs: it happens
59+
* at the top of `bin/run.js`, after one dynamic `import()` of a dependency-free
60+
* module, and every `@oclif/core` byte is written later, inside `run()`. The
61+
* bound exists only so an absent listener is REPORTED rather than waited on
62+
* forever — it is not an oracle over how fast the attach is, and the harness
63+
* asserts the mark this produces rather than the number in it.
64+
*/
65+
const ATTACH_WAIT_MS = 15_000;
66+
67+
/** Comfortably finer than anything being timed. */
68+
const POLL_MS = 10;
69+
70+
/**
71+
* ⛔ This probe installs NO `error` listener of its own on `process.stderr`.
72+
* `uncaughtExceptionMonitor` observes the default action without preventing it,
73+
* so an unguarded run still dies exactly as it would unobserved — an
74+
* `uncaughtException` handler would have changed the very thing being read.
75+
*/
76+
process.on('uncaughtExceptionMonitor', (error) => {
77+
mark(`UNCAUGHT code=${error?.code} msg=${error?.message}`);
78+
});
79+
80+
process.on('exit', (code) => mark(`EXIT code=${code}`));
81+
82+
function writeAndOutliveIt() {
83+
// ONE raw write. The read end is already gone, so this fails; whether that
84+
// failure is fatal is the entire subject.
85+
process.stderr.write('published-entry-stderr-error-probe: one line to a read end that is gone\n');
86+
mark('WROTE');
87+
88+
// ⚠️ The turn is the point, not the delay. A failing write reports through
89+
// libuv's completion callback, so a write followed by a SYNCHRONOUS exit is
90+
// never told at all — which is exactly why the two probes on #15564's card
91+
// read clean, and why a probe that exited here would reproduce their zero
92+
// reading instead of testing anything.
93+
//
94+
// ⛔ NOT unref'd: this timer is what keeps the process alive across that
95+
// turn, and an unref'd one would let the CLI's own exit race it away.
96+
setTimeout(() => {
97+
mark('SURVIVED');
98+
// A distinctive status, so "ended on its own past the write" is evidence
99+
// about THIS probe rather than about any process that happens to exit 0.
100+
process.exit(7);
101+
}, 250);
102+
}
103+
104+
let waited = 0;
105+
const poll = setInterval(() => {
106+
const attached = process.stderr.listenerCount('error') > 0;
107+
if (!attached && waited < ATTACH_WAIT_MS) {
108+
waited += POLL_MS;
109+
return;
110+
}
111+
clearInterval(poll);
112+
mark(attached ? `LISTENER ATTACHED after ${waited} ms` : `LISTENER ABSENT after ${waited} ms`);
113+
if (UNGUARDED) {
114+
// The live positive control — in this process only, never on disk.
115+
process.stderr.removeAllListeners('error');
116+
mark(`ARM unguarded listeners=${process.stderr.listenerCount('error')}`);
117+
} else {
118+
mark(`ARM guarded listeners=${process.stderr.listenerCount('error')}`);
119+
}
120+
writeAndOutliveIt();
121+
}, POLL_MS);

0 commit comments

Comments
 (0)