Repository navigation
Server self-update leaks the old server's SSH device-host tunnel #17438
Description
Activity
Note
Grok responding on behalf of Julius.
Report
On Linux under
t3 __service-launcher, a server self-update appears to leave the old process’s SSH device-host tunnel (ssh … -N -L …) alive after the oldt3 serveis gone. The orphan keeps its loopback listeners and SSH session, reparented undersystemd --user, insidet3code.service’s cgroup.Findings on main (
43f8a8de17a7)-
Launcher kill window targets the child pid only.
terminateChildsendsSIGTERMviachild.kill, thenSIGKILLafterTERMINATE_GRACE_MS = 5_000. Self-update waitsHANDOFF_DELAY_MS = 2_000then calls it at#beginTrial; rollback uses the same helper at#returnToPrevious. There is no process-group or process-tree kill after the child exits. -
Device-host tunnels are detached session leaders; cleanup goes through
stop.SshDeviceHost.connectOncebuilds a standaloneScope.make()and spawnsssh -N -Lwithout settingdetached. Effect’s Node spawner defaultsdetachedto true on non-Windows (nodeChildProcessSpawner.ts), and scope release kills the process group (NodeChildProcessSpawner.ts,#L551-L559). HoststopclosesconnectionScopefirst, then runs remotebootstrap(…, "stop")with a 45s timeout. Thatstopis the host finalizer (:428);DeviceServicecloses host scopes on shutdown (:1141-1145). Orchestration shutdown reconciliation runs in its own finalizer (serverRuntimeStartup.ts:435-453) and can consume a large share of the 5s window before device-host teardown finishes. IfSIGKILLlands mid-stop(or whilestopwaits on the host lock), Effect’s group kill may not complete, and-Nwith ignored stdin means ssh does not notice the parent exit on its own. -
Environment SSH tunnels share the detached default, with different cleanup.
packages/sshtunnel spawn also uses Effect’s spawner withoutdetached: false, so it is detached on Linux too. Its finalizer explicitlykills withforceKillAfter: 2_000before remote stop, and it setsControlMaster=no(:1147-1151). Device-host tunnels do neither. -
systemd unit / launcher lifetime. The rendered user unit sets
KillMode=mixedand runs__service-launcherasExecStart. Self-update terminates and respawns the server child; the launcher stays as MainPID, so systemd’s mixed cgroup sweep does not run. Comments inbootService.ts:148-151describe that intentional trade-off (agent children surviving updates). -
Secondary path (code-only).
ensureReadyclosesconnectionScopeintapError(not on interruption), andconnectOnceassigns a new scope without closing a prior one. An interrupted connect could orphan a tunnel without a hard kill; not required to explain this update incident.
Related
- #17437 — same reporter / same device-host tunnel, ControlMaster port leak (sibling, not a duplicate).
- #14268 / open #14276 — desktop SSH forwards surviving desktop update (same class, different process).
- #15749 — shutdown work that does not finish before a hard stop (related timing class).
- #12507 / open #12513 — helpers outliving the backend; that PR’s
detached: false+forceKillAfterpattern is a plausible direction here. - Local hub already reaps stale pids on boot (
LocalDeviceHost.ts:317-370).
No open PR looks aimed at this server device-host tunnel leak under the service launcher.
Possible directions (any one helps)
- Close device-host tunnel scopes early in server shutdown (before slow orchestration / remote
bootstrap stop), and/or boundstopduring shutdown. - Spawn the tunnel with
detached: false, or kill the process group from the launcher after the child exits (and/or lengthen grace past worst-case shutdown). - Startup pid ledger for
ssh -N -Lorphans (hub reap is a precedent); tying ssh lifetime to the server (piped stdin, no-N) is another option.
Severity / next step
Likely a real bug on the self-update / hard-kill path: leaked loopback forwards and SSH sessions until manual kill or a full unit restart. Worth fixing; labels below.
-
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 9, 2026
What happened
On Linux, T3 Code runs as a systemd user service through
t3 __service-launcher. A server self-update from 0.0.46-nightly.20261008.2849 to 0.0.46-nightly.20261009.2861 at 16:23 JST left the old server's SSH device-host tunnel running. That process wasssh … -N -L 127.0.0.1:<port>:127.0.0.1:<hubPort> -L 127.0.0.1:<port>:127.0.0.1:<daemonPort> <host>, started at 16:20:39.It has been reparented to
systemd --user. It still holds both loopback listeners and an ESTABLISHED SSH session to the device host. Nothing will ever stop it:Every update or rollback that interrupts shutdown can leak another one.
Diagnosis
There are two parts: a short kill window, and children that no one else reaps.
1. The launcher SIGKILLs the old server 5 s after SIGTERM.
request-update, the launcher waitsHANDOFF_DELAY_MS = 2000, then callsterminateChild(apps/server/src/serviceLauncher.ts:524,:534; rollback at:613).terminateChild(:261-273) sends SIGTERM to the server pid only, then SIGKILL afterTERMINATE_GRACE_MS = 5000(:35).2. The old server's shutdown took longer than 5 s.
t3 serveruns its scope finalizers on SIGTERM. In this run, orchestration shutdown reconciliation alone took about 2.5 s.stopstarted a remotebootstrap(…, "stop")ssh, which has a 45 s timeout (SshDeviceHost.ts:70,:399). That ssh never finished.3. Nothing else reaps the device-host tunnel.
Scope.make()(SshDeviceHost.ts:234). Onlystopcloses it, andstopis registered as the host's finalizer (:393-401,:428; host scopes are closed byDeviceService.ts:1142-1144).stopnever closed it before the SIGKILL. The trace doesn't show whether it was waiting on the host lock or simply hadn't been reached.SshDeviceHostdoesn't setdetached, but the Effect Node spawner detaches by default on non-Windows, and its finalizer kills the process group. So killing the server does not kill the tunnel.-Nwith stdin ignored means ssh never notices that its parent is gone.KillMode=mixedsweep never runs. The orphan stays insidet3code.service's cgroup until the whole unit is restarted.The same leak follows from any hard kill of
t3 serve, such as an OOM kill, a crash, orkill -9. A graceful foreground Ctrl-C / SIGTERM with no deadline does clean up, becauserunMainwaits for the finalizers.A possibly related path, found by reading the code only and not observed:
ensureReadyclosesconnectionScopeinEffect.tapError(:385-389). That doesn't run on interruption, and the nextconnectOnceoverwritesconnectionScope(:235) without closing it. An interruptedensureReadycould therefore orphan a tunnel even without a restart.Suggested directions (any one helps; together they make this robust):
bootstrap stop. Bound thestopfinalizer with a short timeout during shutdown.-N, so it exits when the server dies.Steps to reproduce
t3 __service-launcher→t3 serve). Connect one SSH device host.ps -o pid,pgid,sid,args -C ssh. The tunnel ssh has pgid = sid = pid.kill -9 <t3 serve pid>. The launcher starts a new server.bootstrap stopblocks), then trigger a server update from the UI.systemd --user. It keeps its-Llisteners (ss -ltnp) and stays in the unit's cgroup (systemd-cgls --user-unit t3code.service). The new server opens its own tunnel when the host reconnects.Version
The old server was 0.0.46-nightly.20261008.2849, updated to 0.0.46-nightly.20261009.2861. The service launcher is 0.0.46-nightly.20261008.2801; its launcher code is identical to 2861's. The relevant code is unchanged on
main43f8a8d.Environment
CachyOS Linux 7.2 x64, Node 26.8.2, OpenSSH 10.5p1. systemd user unit
t3code.service:KillMode=mixed,TimeoutStopUSec=10s,Restart=always. Two macOS SSH device hosts over Tailscale.Evidence
Related issues
#14268 (desktop: an SSH local forward survives desktop update and relaunch) and its PR #14276. Those are the desktop app's own forwards; this issue is the server's device-host tunnel under the service launcher. The general class: #12507 (child processes left running after the app exits) and #15357 (a backend crash leaves Claude agents running). The ControlMaster port leak on the same tunnel is #17437.
Fix applied or workaround
None automatic yet. Kill the orphan by hand (
kill <pid>; it exits on SIGTERM), or runsystemctl --user restart t3code, which sweeps the cgroup but also stops running agent sessions. To find orphans:pgrep -af -P "$(pgrep -xo -u "$USER" systemd)" 'ssh .* -N -L 127\.0\.0\.1:'.Filed by
Claude Code (claude-opus-5-5) via
t3 triage