Skip to content

Server self-update leaks the old server's SSH device-host tunnel #17438

Description

@kcazin

What happened

On Linux, T3 Code runs as a systemd user service through t3 __service-launcher. A server self-update from 0.0.46-nightly.20261008.2849 to 0.0.46-nightly.20261009.2861 at 16:23 JST left the old server's SSH device-host tunnel running. That process was ssh … -N -L 127.0.0.1:<port>:127.0.0.1:<hubPort> -L 127.0.0.1:<port>:127.0.0.1:<daemonPort> <host>, started at 16:20:39.

It has been reparented to systemd --user. It still holds both loopback listeners and an ESTABLISHED SSH session to the device host. Nothing will ever stop it:

  • The new server doesn't know about it.
  • The launcher, which is the unit's MainPID, never exits, so systemd never sweeps the unit's cgroup.

Every update or rollback that interrupts shutdown can leak another one.

Diagnosis

There are two parts: a short kill window, and children that no one else reaps.

1. The launcher SIGKILLs the old server 5 s after SIGTERM.

  • On request-update, the launcher waits HANDOFF_DELAY_MS = 2000, then calls terminateChild (apps/server/src/serviceLauncher.ts:524, :534; rollback at :613).
  • terminateChild (:261-273) sends SIGTERM to the server pid only, then SIGKILL after TERMINATE_GRACE_MS = 5000 (:35).

2. The old server's shutdown took longer than 5 s.

  • t3 serve runs its scope finalizers on SIGTERM. In this run, orchestration shutdown reconciliation alone took about 2.5 s.
  • After that a device host's stop started a remote bootstrap(…, "stop") ssh, which has a 45 s timeout (SshDeviceHost.ts:70, :399). That ssh never finished.
  • The old server's last log line came 4 s after SIGTERM, about 1 s before the SIGKILL deadline.

3. Nothing else reaps the device-host tunnel.

  • The tunnel lives in a standalone Scope.make() (SshDeviceHost.ts:234). Only stop closes it, and stop is registered as the host's finalizer (:393-401, :428; host scopes are closed by DeviceService.ts:1142-1144).
  • For the leaked host, stop never closed it before the SIGKILL. The trace doesn't show whether it was waiting on the host lock or simply hadn't been reached.
  • The ssh child runs in its own session: on the live orphan, pid = pgid = sid. SshDeviceHost doesn't set detached, but the Effect Node spawner detaches by default on non-Windows, and its finalizer kills the process group. So killing the server does not kill the tunnel.
  • -N with stdin ignored means ssh never notices that its parent is gone.
  • Because the launcher survives the update, systemd's KillMode=mixed sweep never runs. The orphan stays inside t3code.service's cgroup until the whole unit is restarted.

The same leak follows from any hard kill of t3 serve, such as an OOM kill, a crash, or kill -9. A graceful foreground Ctrl-C / SIGTERM with no deadline does clean up, because runMain waits for the finalizers.

A possibly related path, found by reading the code only and not observed: ensureReady closes connectionScope in Effect.tapError (:385-389). That doesn't run on interruption, and the next connectOnce overwrites connectionScope (:235) without closing it. An interrupted ensureReady could therefore orphan a tunnel even without a restart.

Suggested directions (any one helps; together they make this robust):

  • Launcher: give the old child a grace period at least as long as worst-case shutdown, or wait for an explicit "shutdown done" message. After the child exits, kill whatever is left of its process tree or session.
  • SshDeviceHost: close all tunnel scopes first during shutdown, before taking the host lock or running any remote bootstrap stop. Bound the stop finalizer with a short timeout during shutdown.
  • Startup: keep a pid ledger of spawned tunnels and reap stale ones on boot (cloudflared recovery is a precedent). Alternatively, tie ssh to the server's lifetime, for example with a piped stdin and no -N, so it exits when the server dies.

Steps to reproduce

  1. Use Linux, with T3 Code installed as a systemd user service (t3 __service-launcher → t3 serve). Connect one SSH device host.
  2. Run ps -o pid,pgid,sid,args -C ssh. The tunnel ssh has pgid = sid = pid.
  3. Do any of the following:
    • Simplest: kill -9 <t3 serve pid>. The launcher starts a new server.
    • Make shutdown take more than 5 s (for example, a second device host that is slow or unreachable, so its bootstrap stop blocks), then trigger a server update from the UI.
  4. The old tunnel survives with PPID = systemd --user. It keeps its -L listeners (ss -ltnp) and stays in the unit's cgroup (systemd-cgls --user-unit t3code.service). The new server opens its own tunnel when the host reconnects.

Version

The old server was 0.0.46-nightly.20261008.2849, updated to 0.0.46-nightly.20261009.2861. The service launcher is 0.0.46-nightly.20261008.2801; its launcher code is identical to 2861's. The relevant code is unchanged on main 43f8a8d.

Environment

CachyOS Linux 7.2 x64, Node 26.8.2, OpenSSH 10.5p1. systemd user unit t3code.service: KillMode=mixed, TimeoutStopUSec=10s, Restart=always. Two macOS SSH device hosts over Tailscale.

Evidence

# server.trace.ndjson / boot-service.log, 2026-10-09 JST
16:20:39.0   last SshDeviceHost.connectOnce of the old server (2849); this spawned the tunnel that later leaked
16:23:34.076 "Server update prepared; handing off to the service launcher" (target 2861)
16:23:36.08  SIGTERM received: releaseManagedTunnelOnShutdown, ws subscriptions torn down (incl. subscribeDeviceState)
16:23:38.547 "V2 orchestration shutdown reconciliation completed" {terminalizedRuns:1, stoppedSessions:1}
16:23:38.556 ssh/auth.buildSshChildEnvironment: an ssh command (device host stop) starts and never finishes
16:23:40.117 last old-server log line                 (SIGKILL due ~16:23:41.1)
16:23:45.92  first new-server spans; 16:23:46.334 "Listening on http://127.0.0.1:3773"

# live, 18 minutes after the update
$ ps -o pid,ppid,pgid,sid,stat,lstart -p <pid>
    PID    PPID    PGID     SID STAT  STARTED
<pid>    <systemd --user>  <pid>  <pid>  Ss    Fri Oct  9 16:20:39 2026
$ cat /proc/<pid>/cgroup
0::/user.slice/user-1000.slice/user@1000.service/app.slice/t3code.service
$ ss -H -ltnp | grep <pid>     -> 127.0.0.1:<port> and 127.0.0.1:<port> still listening
$ ss -H -tnp  | grep <pid>     -> ESTABLISHED to <device host>:22
# The new server (2861) has no ssh children; the launcher (MainPID) has run since Oct 8.

Related issues

#14268 (desktop: an SSH local forward survives desktop update and relaunch) and its PR #14276. Those are the desktop app's own forwards; this issue is the server's device-host tunnel under the service launcher. The general class: #12507 (child processes left running after the app exits) and #15357 (a backend crash leaves Claude agents running). The ControlMaster port leak on the same tunnel is #17437.

Fix applied or workaround

None automatic yet. Kill the orphan by hand (kill <pid>; it exits on SIGTERM), or run systemctl --user restart t3code, which sweeps the cgroup but also stops running agent sessions. To find orphans: pgrep -af -P "$(pgrep -xo -u "$USER" systemd)" 'ssh .* -N -L 127\.0\.0\.1:'.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

  1. juliusmarminge commented on Oct 9, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Report

    On Linux under t3 __service-launcher, a server self-update appears to leave the old process’s SSH device-host tunnel (ssh … -N -L …) alive after the old t3 serve is gone. The orphan keeps its loopback listeners and SSH session, reparented under systemd --user, inside t3code.service’s cgroup.

    Findings on main (43f8a8de17a7)

    1. Launcher kill window targets the child pid only. terminateChild sends SIGTERM via child.kill, then SIGKILL after TERMINATE_GRACE_MS = 5_000. Self-update waits HANDOFF_DELAY_MS = 2_000 then calls it at #beginTrial; rollback uses the same helper at #returnToPrevious. There is no process-group or process-tree kill after the child exits.

    2. Device-host tunnels are detached session leaders; cleanup goes through stop. SshDeviceHost.connectOnce builds a standalone Scope.make() and spawns ssh -N -L without setting detached. Effect’s Node spawner defaults detached to true on non-Windows (nodeChildProcessSpawner.ts), and scope release kills the process group (NodeChildProcessSpawner.ts, #L551-L559). Host stop closes connectionScope first, then runs remote bootstrap(…, "stop") with a 45s timeout. That stop is the host finalizer (:428); DeviceService closes host scopes on shutdown (:1141-1145). Orchestration shutdown reconciliation runs in its own finalizer (serverRuntimeStartup.ts:435-453) and can consume a large share of the 5s window before device-host teardown finishes. If SIGKILL lands mid-stop (or while stop waits on the host lock), Effect’s group kill may not complete, and -N with ignored stdin means ssh does not notice the parent exit on its own.

    3. Environment SSH tunnels share the detached default, with different cleanup. packages/ssh tunnel spawn also uses Effect’s spawner without detached: false, so it is detached on Linux too. Its finalizer explicitly kills with forceKillAfter: 2_000 before remote stop, and it sets ControlMaster=no (:1147-1151). Device-host tunnels do neither.

    4. systemd unit / launcher lifetime. The rendered user unit sets KillMode=mixed and runs __service-launcher as ExecStart. Self-update terminates and respawns the server child; the launcher stays as MainPID, so systemd’s mixed cgroup sweep does not run. Comments in bootService.ts:148-151 describe that intentional trade-off (agent children surviving updates).

    5. Secondary path (code-only). ensureReady closes connectionScope in tapError (not on interruption), and connectOnce assigns a new scope without closing a prior one. An interrupted connect could orphan a tunnel without a hard kill; not required to explain this update incident.

    Related

    • #17437 — same reporter / same device-host tunnel, ControlMaster port leak (sibling, not a duplicate).
    • #14268 / open #14276 — desktop SSH forwards surviving desktop update (same class, different process).
    • #15749 — shutdown work that does not finish before a hard stop (related timing class).
    • #12507 / open #12513 — helpers outliving the backend; that PR’s detached: false + forceKillAfter pattern is a plausible direction here.
    • Local hub already reaps stale pids on boot (LocalDeviceHost.ts:317-370).

    No open PR looks aimed at this server device-host tunnel leak under the service launcher.

    Possible directions (any one helps)

    • Close device-host tunnel scopes early in server shutdown (before slow orchestration / remote bootstrap stop), and/or bound stop during shutdown.
    • Spawn the tunnel with detached: false, or kill the process group from the launcher after the child exits (and/or lengthen grace past worst-case shutdown).
    • Startup pid ledger for ssh -N -L orphans (hub reap is a precedent); tying ssh lifetime to the server (piped stdin, no -N) is another option.

    Severity / next step

    Likely a real bug on the self-update / hard-kill path: leaked loopback forwards and SSH sessions until manual kill or a full unit restart. Worth fixing; labels below.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions