Repository navigation
[Bug]: headless t3 serve crashes on an unhandled 'read ECONNRESET' from one provider child's pipe; all 44 sessions stopped #16794
Description
Activity
Note
Grok responding on behalf of Julius.
Thanks for the stack trace. It's the same crash fixed by the open #16555.
Root cause.
@effect/platform-node-shared@4.0.1'sNodeChildProcessSpawneronly listens for"error"on a child's input pipes (stdin andtype: "input"extra fds) while itsfromWritablesink is writing. When a child exits with bytes still unread, the pipe fails later with no listener attached. Node then throwsUnhandled 'error' eventand the whole server goes down, along with every session. Output pipes are fine: stdout and stderr get a permanent listener (NodeChildProcessSpawner.ts:331,342), and so do output fds (:266).Which pipe this was. Your trace says
read ECONNRESETwithsyscall: 'read'. Node never reads a child's stdin (it's write-only, so a dead child shows up there aswrite EPIPE). And stdout/stderr already have listeners. That leaves an extra fd that is both written and read. In the server, the one that fits ishtml_render/html_preview's headless Chrome CDP pipe (apps/server/src/htmlRender/headlessChrome.ts:219-221,fd3: { type: "input" }). Tearing down a capture interrupts the fd3 writer and then kills Chrome, so this crash can happen on any thread that renders HTML. #16555 reports the same trace, and with ~40 agent threads using html_render it's the likely trigger here, not a provider's own stdio.Does #16555 cover provider sessions? Yes, as far as this can fail:
- Codex (
CodexAdapterV2.ts:1345) and ACP drivers (provider/acp/AcpSessionRuntime.ts:1584) spawn through the same Effect spawner with piped stdin. fix(server): a child killed with unread input no longer crashes the server on html_render #16555 adds a lasting listener on that stdin too, so the provider-side variant (write EPIPE) is covered as well. - Claude Code spawns through
@anthropic-ai/claude-agent-sdk, which already guards its child's stdin and stderr. It isn't exposed. - fix(server): a child killed with unread input no longer crashes the server on html_render #16555 does not cover direct
node:child_processspawns (serviceLauncher.ts:429,cli/triage.ts,cli/uninstall.ts). Those use inherited stdio or aren't long-lived, so they shouldn't produce this crash.
So
Fixes #16794on #16555 looks right.Proposed fix. Land #16555. A follow-up worth considering: report this upstream to Effect, so the patch can be dropped once the spawner keeps an error listener on its input pipes itself.
Workaround until then. Keep the systemd
Restart=you already have. If you can live withouthtml_render/html_previewon the headless box, avoiding them should stop these crashes. If a crash still happens with no HTML render in flight, please post the log lines from just before it. That would point to a different pipe.History: #5621 had the same symptom on a different socket (the TCP server socket, an Effect beta.103 regression fixed in beta.104). It's unrelated to this one.
- Codex (
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 7, 2026 Second occurrence, same machine, now on
t3@0.0.46-nightly.20261007.2761(upgraded 2026-10-07 05:33 Panama after the first crash):2026-10-07T08:37:51-05:00 t3[1231909]: throw er; // Unhandled 'error' event 2026-10-07T08:37:51-05:00 t3[1231909]: Error: read ECONNRESET 2026-10-07T08:37:51-05:00 t3[1231909]: errno: -104, 2026-10-07T08:37:51-05:00 t3[1231909]: code: 'ECONNRESET', 2026-10-07T08:37:51-05:00 t3[1231909]: syscall: 'read' systemd: t3-server.service: Main process exited, code=exited, status=1/FAILUREServer uptime before the crash: 3 h 04 m (05:33 → 08:37). The first crash was after 2 h 21 m. ~32 stopped sessions and ~12 running provider sessions at the time; the preceding log line is a provider session id for an
opencodethread, as in the first report.Memory: RSS of the
t3 serveprocess reached ~18 GB after 3 h this time (14 GB after 2 h 21 m in the first crash), on a 61 GB box with ~45 threads. Growth looks linear with uptime, so a leak on top of the unhandled socket error. Happy to run a heap snapshot or--inspectbuild if that helps.One more thing seen every ~80 s before both crashes:
GitCommandError ... GitManager.branchPullRequest.remotes (<project root>): not_a_repositoryfor a registered project whose root is deliberately not a git repository, and the same call failing with ENOENT for a worktree that was removed while its thread still exists. Both are logged as warnings and keep retrying.More detail from the journal around both crashes (Node.js v26.8.2 embedded in the t3 binary, Linux, systemd user service):
Crash 2 (08:37:51 local), the 17 seconds before it:
[08:37:34.804] WARN (#1159200): orchestration-v2.provider-session-scope-close-timeout { providerSessionId: 'provider-session:provider-instance:opencode:thread:86e2bfcf-…:e7e584d7-…', reason: 'idle_timeout', timeoutMs: 30000 } 08:37:51 node:events:505 throw er; // Unhandled 'error' event Error: read ECONNRESET at Pipe.onStreamRead (node:internal/stream_base_commons:216:20) Emitted 'error' event on Socket instance at: at emitErrorNT (node:internal/streams/destroy:170:8) at emitErrorCloseNT (node:internal/streams/destroy:129:3)So the socket is a
Pipe(child-process stdio), and the last orchestration event before the crash is the server closing an opencode provider session scope onidle_timeout. All 4provider-session-scope-close-timeoutwarnings in the last two days areprovider-instance:opencode; no other provider has produced one. The opencode thread in question had already ended in an error state from the UI's point of view. Hypothesis: when the scope close times out, the child's stdio pipe is torn down while a reader is still attached and the resultingread ECONNRESEThas noerrorlistener, so it is unhandled and takes the whole server down.Crash 1 (05:03:17 local): same stack. Preceding warnings were
orchestration V2 provider event ingestion failed(04:58:00) and twothread pull request update failed(05:00:57, 05:02:18); no scope-close warning that time, so the scope-close may be one trigger among several for the same unguarded pipe.systemd accounting at exit:
crash 1: Consumed 2h 55min CPU over 2h 21min wall, 14G memory peak crash 2: Consumed 5h 24min CPU over 3h 04min wall, 18G memory peakRoughly 1.3 to 1.8 cores busy on average with ~45 threads registered, and peak RSS growing with uptime.
Workaround on our side until this is fixed: a planned restart when RSS passes 12 GB plus a nightly restart, and
Restart=alwaysso the crash costs ~30 s. If aprocess.on('uncaughtException')/socketerrorguard around the provider child pipes is the fix, we are happy to test a nightly build.Thanks, that reading fits our logs better than my opencode guess. Checking against it:
- Crash 2 (08:37:51): 71 s earlier the server logged
[08:36:40.416] WARN (#1369819): html_preview failed. Consistent with a torn-down capture on the fd3 CDP pipe. - Crash 1 (05:03:17): nothing HTML-related in the 13 minutes before it, but only failed renders are logged at WARN, and this box had used
html_render/html_previewfrom agent threads during that server's life, so I cannot rule it out. Theopencodescope-close warning before crash 2 is probably coincidence; I withdraw that hypothesis.
Until #16555 ships we are stopping all
html_render/html_previewuse on this box and keepingRestart=alwaysplus a planned restart when RSS passes 12 GB. If a crash still happens with no HTML render in flight I will post the preceding log lines here. The memory growth (14 GB after 2 h 21 m, 18 GB after 3 h 04 m) looks separate from this crash; happy to open a second issue with a heap snapshot if useful.- Crash 2 (08:37:51): 71 s earlier the server logged
Steps to reproduce
Run
t3 serveheadlessly on Linux as a long-lived systemd user service with ~40 concurrent threads (Claude Code, Codex and ACP provider sessions), each driven as a child process over stdio pipes. Leave it running for days. At some point one child's pipe is reset.Expected behavior
A
read ECONNRESETon a single provider child's pipe should fail that one session (or be retried by the adapter), not crash the whole server.Actual behavior
The server process exits with an unhandled
'error'event on aSocket(pipe) and systemd restarts it. Every session is stopped: the restart reconciliation loggedstoppedSessions: 44. The orchestrator came back after ~33 s, but all 44 threads had to be re-woken externally.Impact
All sessions on the box interrupted once; the first crash of this kind in 3 days on this install. Earlier builds had a similar pattern (#5621, closed).
Version or commit
0.0.46-nightly.20261005.2702 (cfa4f76). Upgrading to 0.0.46-nightly.20261007.2761 today; no commit between the two mentions ECONNRESET or the pipe error path, so reporting.
Environment
Linux (Ubuntu, kernel 7.0.0-31-generic), Node.js v26.8.2,
t3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3undersystemd --user, Tailscale serve in front, ~14 GB RSS peak, ~40-45 live threads (claudeAgent, codex, ACP drivers).Logs or stack traces