Skip to content

[Bug]: headless t3 serve crashes on an unhandled 'read ECONNRESET' from one provider child's pipe; all 44 sessions stopped #16794

Description

@MattRiddell

Steps to reproduce

Run t3 serve headlessly on Linux as a long-lived systemd user service with ~40 concurrent threads (Claude Code, Codex and ACP provider sessions), each driven as a child process over stdio pipes. Leave it running for days. At some point one child's pipe is reset.

Expected behavior

A read ECONNRESET on a single provider child's pipe should fail that one session (or be retried by the adapter), not crash the whole server.

Actual behavior

The server process exits with an unhandled 'error' event on a Socket (pipe) and systemd restarts it. Every session is stopped: the restart reconciliation logged stoppedSessions: 44. The orchestrator came back after ~33 s, but all 44 threads had to be re-woken externally.

Impact

All sessions on the box interrupted once; the first crash of this kind in 3 days on this install. Earlier builds had a similar pattern (#5621, closed).

Version or commit

0.0.46-nightly.20261005.2702 (cfa4f76). Upgrading to 0.0.46-nightly.20261007.2761 today; no commit between the two mentions ECONNRESET or the pipe error path, so reporting.

Environment

Linux (Ubuntu, kernel 7.0.0-31-generic), Node.js v26.8.2, t3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3 under systemd --user, Tailscale serve in front, ~14 GB RSS peak, ~40-45 live threads (claudeAgent, codex, ACP drivers).

Logs or stack traces

2026-10-07T05:03:17-05:00 t3[302398]: node:events:505
    throw er; // Unhandled 'error' event
    ^
Error: read ECONNRESET
    at Pipe.onStreamRead (node:internal/stream_base_commons:216:20)
Emitted 'error' event on Socket instance at:
    at emitErrorNT (node:internal/streams/destroy:170:8)
    at emitErrorCloseNT (node:internal/streams/destroy:129:3)
    at process.processTicksAndRejections (node:internal/process/task_queues:90:21) {
  errno: -104,
  code: 'ECONNRESET',
  syscall: 'read'
}
Node.js v26.8.2
systemd[1816]: t3-server.service: Main process exited, code=exited, status=1/FAILURE
systemd[1816]: t3-server.service: Scheduled restart job, restart counter is at 1.
... 33 s later ...
t3[46070]: V2 orchestration shutdown reconciliation completed { stoppedSessions: 44, ... }

Activity

  1. juliusmarminge commented on Oct 7, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the stack trace. It's the same crash fixed by the open #16555.

    Root cause. @effect/platform-node-shared@4.0.1's NodeChildProcessSpawner only listens for "error" on a child's input pipes (stdin and type: "input" extra fds) while its fromWritable sink is writing. When a child exits with bytes still unread, the pipe fails later with no listener attached. Node then throws Unhandled 'error' event and the whole server goes down, along with every session. Output pipes are fine: stdout and stderr get a permanent listener (NodeChildProcessSpawner.ts:331,342), and so do output fds (:266).

    Which pipe this was. Your trace says read ECONNRESET with syscall: 'read'. Node never reads a child's stdin (it's write-only, so a dead child shows up there as write EPIPE). And stdout/stderr already have listeners. That leaves an extra fd that is both written and read. In the server, the one that fits is html_render/html_preview's headless Chrome CDP pipe (apps/server/src/htmlRender/headlessChrome.ts:219-221, fd3: { type: "input" }). Tearing down a capture interrupts the fd3 writer and then kills Chrome, so this crash can happen on any thread that renders HTML. #16555 reports the same trace, and with ~40 agent threads using html_render it's the likely trigger here, not a provider's own stdio.

    Does #16555 cover provider sessions? Yes, as far as this can fail:

    So Fixes #16794 on #16555 looks right.

    Proposed fix. Land #16555. A follow-up worth considering: report this upstream to Effect, so the patch can be dropped once the spawner keeps an error listener on its input pipes itself.

    Workaround until then. Keep the systemd Restart= you already have. If you can live without html_render/html_preview on the headless box, avoiding them should stop these crashes. If a crash still happens with no HTML render in flight, please post the log lines from just before it. That would point to a different pipe.

    History: #5621 had the same symptom on a different socket (the TCP server socket, an Effect beta.103 regression fixed in beta.104). It's unrelated to this one.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 7, 2026
  3. MattRiddell commented on Oct 7, 2026

    @MattRiddell
    Author

    Second occurrence, same machine, now on t3@0.0.46-nightly.20261007.2761 (upgraded 2026-10-07 05:33 Panama after the first crash):

    2026-10-07T08:37:51-05:00 t3[1231909]:     throw er; // Unhandled 'error' event
    2026-10-07T08:37:51-05:00 t3[1231909]: Error: read ECONNRESET
    2026-10-07T08:37:51-05:00 t3[1231909]:   errno: -104,
    2026-10-07T08:37:51-05:00 t3[1231909]:   code: 'ECONNRESET',
    2026-10-07T08:37:51-05:00 t3[1231909]:   syscall: 'read'
    systemd: t3-server.service: Main process exited, code=exited, status=1/FAILURE
    

    Server uptime before the crash: 3 h 04 m (05:33 → 08:37). The first crash was after 2 h 21 m. ~32 stopped sessions and ~12 running provider sessions at the time; the preceding log line is a provider session id for an opencode thread, as in the first report.

    Memory: RSS of the t3 serve process reached ~18 GB after 3 h this time (14 GB after 2 h 21 m in the first crash), on a 61 GB box with ~45 threads. Growth looks linear with uptime, so a leak on top of the unhandled socket error. Happy to run a heap snapshot or --inspect build if that helps.

    One more thing seen every ~80 s before both crashes: GitCommandError ... GitManager.branchPullRequest.remotes (<project root>): not_a_repository for a registered project whose root is deliberately not a git repository, and the same call failing with ENOENT for a worktree that was removed while its thread still exists. Both are logged as warnings and keep retrying.

  4. MattRiddell commented on Oct 7, 2026

    @MattRiddell
    Author

    More detail from the journal around both crashes (Node.js v26.8.2 embedded in the t3 binary, Linux, systemd user service):

    Crash 2 (08:37:51 local), the 17 seconds before it:

    [08:37:34.804] WARN (#1159200): orchestration-v2.provider-session-scope-close-timeout
      {
        providerSessionId: 'provider-session:provider-instance:opencode:thread:86e2bfcf-…:e7e584d7-…',
        reason: 'idle_timeout',
        timeoutMs: 30000
      }
    08:37:51  node:events:505  throw er; // Unhandled 'error' event
              Error: read ECONNRESET
                  at Pipe.onStreamRead (node:internal/stream_base_commons:216:20)
              Emitted 'error' event on Socket instance at:
                  at emitErrorNT (node:internal/streams/destroy:170:8)
                  at emitErrorCloseNT (node:internal/streams/destroy:129:3)
    

    So the socket is a Pipe (child-process stdio), and the last orchestration event before the crash is the server closing an opencode provider session scope on idle_timeout. All 4 provider-session-scope-close-timeout warnings in the last two days are provider-instance:opencode; no other provider has produced one. The opencode thread in question had already ended in an error state from the UI's point of view. Hypothesis: when the scope close times out, the child's stdio pipe is torn down while a reader is still attached and the resulting read ECONNRESET has no error listener, so it is unhandled and takes the whole server down.

    Crash 1 (05:03:17 local): same stack. Preceding warnings were orchestration V2 provider event ingestion failed (04:58:00) and two thread pull request update failed (05:00:57, 05:02:18); no scope-close warning that time, so the scope-close may be one trigger among several for the same unguarded pipe.

    systemd accounting at exit:

    crash 1: Consumed 2h 55min CPU over 2h 21min wall, 14G memory peak
    crash 2: Consumed 5h 24min CPU over 3h 04min wall, 18G memory peak
    

    Roughly 1.3 to 1.8 cores busy on average with ~45 threads registered, and peak RSS growing with uptime.

    Workaround on our side until this is fixed: a planned restart when RSS passes 12 GB plus a nightly restart, and Restart=always so the crash costs ~30 s. If a process.on('uncaughtException')/socket error guard around the provider child pipes is the fix, we are happy to test a nightly build.

  5. MattRiddell commented on Oct 7, 2026

    @MattRiddell
    Author

    Thanks, that reading fits our logs better than my opencode guess. Checking against it:

    • Crash 2 (08:37:51): 71 s earlier the server logged [08:36:40.416] WARN (#1369819): html_preview failed. Consistent with a torn-down capture on the fd3 CDP pipe.
    • Crash 1 (05:03:17): nothing HTML-related in the 13 minutes before it, but only failed renders are logged at WARN, and this box had used html_render/html_preview from agent threads during that server's life, so I cannot rule it out. The opencode scope-close warning before crash 2 is probably coincidence; I withdraw that hypothesis.

    Until #16555 ships we are stopping all html_render/html_preview use on this box and keeping Restart=always plus a planned restart when RSS passes 12 GB. If a crash still happens with no HTML render in flight I will post the preceding log lines here. The memory growth (14 GB after 2 h 21 m, 18 GB after 3 h 04 m) looks separate from this crash; happy to open a second issue with a heap snapshot if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions