Skip to content

SSH reconnect can replace an external T3 service with a competing managed runtime #5749

Description

@mikeyperes

Problem

On a persistent SSH host that already runs the supported T3 systemd service, reconnecting from an updated desktop client can start a second T3 server and a second T3 Connect tunnel against the same base directory.

The failure is deterministic when launcher state says managed while server-runtime.json belongs to the launcher-managed process, or when a recorded external service is briefly unavailable during its own update. Requests and WebSockets can then land on different servers, producing intermittent reconnects and messages that appear to disappear.

Expected behavior

The supported external service remains the sole server/tunnel owner. The SSH launcher adopts its advertised port, records external ownership, and never replaces a temporarily restarting external service.

Proposed fix

The tested source change:

  • distinguishes the launcher's own managed runtime record from a real external runtime;
  • makes runtime-state cleanup conditional on PID/start-time ownership;
  • preserves external ownership across transient unavailability; and
  • adds regression coverage for adoption, reconnect, restart, and stale cleanup.

A focused PR follows.

Activity

  1. Sebastian-Gerken commented on Aug 15, 2026

    @Sebastian-Gerken

    Reproduction on a persistent systemd service (Nightly 0.0.34-20260815.1100)

    We hit this again today on a Linux host that already runs the supported user systemd unit (t3code.service → service-launcher.mjs → t3 serve on 127.0.0.1:3773, managed=external). The desktop is Nightly 0.0.34-nightly.20260815.1100, same version as the service.

    This is not a missing service. The boot service stayed active with NRestarts=0 the whole time.

    Sequence

    1. Service environment probe on 127.0.0.1:3773/.well-known/t3/environment was 200 in 17ms while the host was idle enough.

    2. The same host then had live agent work under the service cgroup (gate.sh / cargo test). The next probe was still 200 but 278ms.

    3. Desktop marked the existing SSH tunnel stale after its 2s readiness budget (SshReadinessError: Timed out waiting 2000ms for backend readiness on the local forward).

    4. launchOrReuseRemoteServer then started a competing process:

      t3 serve --host 127.0.0.1 --port 3774 --base-dir $HOME/.t3

    5. A later retry did the same on 3775.

    6. All three Node processes (3773 service child + two launcher-managed serves) had state.sqlite / -wal / -shm open. All three went Dl. After that, every port timed out, including the original service.

    Desktop error (trimmed, no tokens):

    SshCommandError: Remote T3 server did not become ready on 127.0.0.1:3774.
    WARN: Grok ACP model discovery timed out after 15000ms.
    WARN: failed to stop provider service
      { errorTag: 'ProviderSessionDirectoryPersistenceError' }
    

    ssh-launch/*/managed flipped from external to managed, port=3775.

    Source

    This matches current packages/ssh/src/tunnel.ts on main:

    • REMOTE_REUSE_READY_TIMEOUT_MS = 2_000 vs REMOTE_READY_TIMEOUT_MS = 15_000 for a newly spawned server
    • On reuse miss, pick_port skips busy 3773 and runs t3 serve --base-dir "$DEFAULT_SERVER_HOME" on the next free port
    • There is no exclusive lock on that base dir / state.sqlite
    • If managed=managed and the pid file is empty, PID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}" can target the systemd child

    packages/ssh/src/tunnel.test.ts currently asserts that PID_TO_STOP fallback as intended.

    Why 2s is not enough here

    The reuse probe is GET / with a 1s per-request timeout. On a host where the service cgroup is doing real work, 3773 can still be healthy and miss a 2s desktop/reuse budget. The fallback is not “retry the existing owner”; it is “start another server on the same data directory.” That is what wedges SQLite, after which even the original owner stops answering and every further reconnect makes it worse.

    #5751

    #5751 is the right contract: if an external owner exists and is not ready, exit 1. Do not start a second server. Do not kill $DEFAULT_RUNTIME_PID.

    That still leaves the 2s reuse budget. Fail-closed would have saved the database today; it would still have shown a reconnect error on this busy host. The reuse timeout needs to be in the same league as REMOTE_READY_TIMEOUT_MS (15s), or reuse should treat a live server-runtime.json pid + occupied port as sufficient and not require a 2s HTTP win.

    Workaround we are taking

    We are moving these persistent service hosts off Desktop-Managed SSH Launch onto operator-owned SSH forwards + ordinary pairing to the already-running service. SSH Launch is the wrong access method for a machine that already has t3 service install.

  2. Sebastian-Gerken commented on Aug 15, 2026

    @Sebastian-Gerken

    Workaround now in use

    We stopped using Desktop-Managed SSH Launch against these persistent systemd hosts.

    Current setup:

    • Each Linux host keeps t3code.service (user unit, linger enabled) as the only server on 127.0.0.1:3773.
    • ~/.local/bin/t3 is a small guard that refuses t3 serve while that unit is active, so a leftover SSH-launch attempt cannot open the same state.sqlite.
    • The desktop reaches the service through operator-owned ssh -N forwards on the Mac (launchd + KeepAlive). It is paired to http://127.0.0.1:<local-port> as a normal environment, not as an SSH-launched environment.

    That matches the documented “pre-existing server” model. SSH Launch remains unsafe on Nightly until reuse stops spawning a second process on the same base dir (#5751) and the 2s reuse budget is raised.

    The remote server does not depend on the laptop: logout, sleep, and LAN flaps leave the systemd unit running. Only the desktop’s WebSocket drops until the Mac-side forward comes back.

  3. acomarcho commented on Sep 3, 2026

    @acomarcho

    Adding an independent reproduction on 0.0.35, in a topology that I think widens the scope of this issue slightly, plus one finding that makes the teardown path worse than the analysis above assumes.

    This does not require an external service

    The reproduction in the comment above starts from a host running the supported systemd unit (managed=external). On the host I looked at there is no systemd unit at all (t3 service status reports not installed), so every server is launcher-managed from the start. The runaway is identical. So the failure is not specifically about replacing an external service; a pure managed host escalates the same way, which suggests the fix needs to cover the plain managed path too, not only external-ownership preservation.

    Topology: desktop app on a separate machine, remote Linux host reached over Tailscale through the SSH launcher, server pinned to t3@0.0.35 by run-t3.sh.

    The PID_TO_STOP fallback is worse than "can target the systemd child"

    The comment above notes that PID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}" can target the wrong process. On any host where run-t3.sh takes the npx branch (no global t3 on PATH), it is worse than that: $REMOTE_PID never identifies a server at all, so the kill is a no-op and the old server is always leaked.

    REMOTE_PID="$!" at tunnel.ts:594 captures the npm exec wrapper, because npx spawns rather than exec-chains. Live capture during the failure:

    2618056  npm exec t3@0.0.35 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3   <- $PID_FILE
      2618094  sh -c "t3" serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3
        2618095  node ~/.npm/_npx/<hash>/node_modules/.bin/t3 serve ... --port 3773   <- real server
    

    I reproduced the signal behaviour with a synthetic package in the same three-layer shape: SIGTERM to the $! pid kills the wrapper, and the node process survives with its port still bound, reparented to ppid=1. Full write-up and repro script in #2614, which I think is the orphaning half of this same story.

    That closes the loop on why the escalation is unbounded rather than self-limiting: the reuse probe decides to replace the server, and the replacement path is structurally incapable of stopping the one it is replacing.

    Evidence from this host

    • Nine distinct ports across the launcher logs, 3773 through 3781, five live simultaneously, every one launched with the same --base-dir and therefore the same state.sqlite.
    • PersistenceSqlError: SQL error in OrchestrationEventStore.append:insert with Error: database is locked: 463 occurrences across two log files.
    • thread <id> already has an active writer on codex threads, because the codex app-server stays attached to a server the desktop has abandoned.
    • The same WARN: failed to stop provider service { errorTag: 'ProviderSessionDirectoryPersistenceError' } reported above.
    • Desktop side surfaced as Failed to connect. Reconnecting... then <host> did not respond during connection setup, looping indefinitely.

    Two launches 30 seconds apart during one loop, showing the ladder in motion:

    [08:17:34.021] INFO (#259): Listening on http://127.0.0.1:3776
    [08:18:04.608] INFO (#234): Listening on http://127.0.0.1:3777
    

    The spawn rate is bounded only by the retry ladder

    packages/client-runtime/src/connection/supervisor.ts:

    const RETRY_DELAYS_MS = [3_000, 4_000, 8_000, 16_000] as const;   // :32
    const CONNECTION_ESTABLISHMENT_TIMEOUT = "15 seconds";            // :33

    The ladder caps at 16s and retries forever, and each attempt that reaches ensureTunnelEntry with a stale entry runs the remote launch script again. So a host that stays above the 2s reuse budget gains roughly one additional server per 16-31s indefinitely, each one making the next probe likelier to miss.

    Data volume is not the trigger

    Worth ruling out, since the database on this host had grown to 3.2 GB (336k thread.activity-appended rows, 1.22 GB of payload, no retention or pruning anywhere in the server source). With a single server alive against that same database, GET / and /.well-known/t3/environment both answer in 1-5 ms. The 2-4s responses that trip the 2s budget appear only once duplicates are contending for the WAL. Contention is the cause, not size, so raising busy_timeout or shrinking the database would not address this.

    The initial nudge on this host was ordinary background load: ~29k GitVcsDriver.fetchRemoteForStatus timeouts across four repos and ~17k Failed to flush telemetry fetch failures. Once one probe misses, the loop sustains itself without needing that trigger again.

    Still current

    packages/ssh/src/tunnel.ts is byte-identical between v0.0.35 and main at 5b8445b7a777. REMOTE_REUSE_READY_TIMEOUT_MS = 2_000, SSH_READY_PROBE_TIMEOUT_MS = 1_000 and REMOTE_PID="$!" are all unchanged as of 0.0.38.

    Environment

    • Remote host: Linux x64 (7.0.0-30-generic), Node v24.16.0, t3 server 0.0.35, no global t3 on PATH (npx branch of run-t3.sh), no systemd unit
    • Desktop app on a separate machine over Tailscale

    Filed via t3 triage. Investigation and reproduction by Claude Opus 5 (1M context) running as the triage agent in Claude Code.

  4. tyler6204 commented on Sep 4, 2026

    @tyler6204

    Still hitting this on 0.0.39-nightly.20260903.1272, with the same version on the Mac desktop and an Ubuntu EC2 host connected through Desktop's SSH launcher. Running Codex tasks end up with:

    Provider session did not survive a server restart. Send a new message to continue.
    

    This also happens without the systemd service. We previously had a boot-service server and an SSH-launched server sharing the same T3 home. Removing the boot service cleared the active-writer conflicts, but it did not stop the reconnect/restart failures.

    Latest incident, September 4 (UTC), from the desktop trace and remote runtime record:

    • 07:53:58: the existing SSH tunnel to remote port 3773 closed.
    • 07:55:28: ssh/tunnel.ensureTunnelEntry failed with SSH command timed out after 90000ms.
    • 07:57:58: a new remote server wrote its runtime record on port 3774.
    • 07:58:03: SSH environment setup and pairing succeeded. The interrupted tasks still showed the server-restart error.

    EC2 did not reboot: uptime was 30 days. The check immediately afterward showed 42 GiB RAM available, 62 GiB free on the root disk, and no OOM-kill entries in the kernel journal for the checked incident window. The new server answered HTTP requests normally.

    There is also a same-PID problem in the launcher code inside the installed Nightly app, not just in a source checkout. When the runtime record is healthy and REMOTE_MANAGED is managed, it assigns PID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}" and kills it without checking whether the discovered runtime is that same managed process. That appears to let reconnect kill the server it should reuse. We did not capture the signal sender for this particular restart, so that part is code analysis rather than a traced kill event.

    Adding this as current-version evidence for the managed-only case above. Reconnecting the desktop should restore the connection without replacing a healthy remote server or ending its running tasks.

  5. georgenijo commented on Sep 15, 2026

    @georgenijo

    Still reproducing on 0.0.41-nightly.20260915.1735 (macOS arm64, launchd service), desktop + service on the same version.

    Topology: Mac desktop app → SSH launcher → Mac mini running com.t3tools.t3code.service (launchd, 127.0.0.1:3773, ~/.t3).

    Observed today:

    1. After a reboot, the SSH launcher started a managed server on 3775 against the same ~/.t3 while the launchd service was also running on 3773, so both had state.sqlite open. server-runtime.json pointed at the managed server.
    2. After stopping the managed server and running t3 service install, the service took ~46s to start listening. The desktop reconnected 6s after the service launched, found no server-runtime.json yet, and started another managed server (3774), which wrote its own server-runtime.json.
    3. When that managed server was stopped with SIGTERM (after the service had rewritten server-runtime.json with serviceManaged: true), it deleted the service's server-runtime.json on the way out. The release in apps/server/src/server.ts runtimeStateLayer calls clearPersistedServerRuntimeState unconditionally, with no PID/ownership check. I had to restore the file by hand before the next reconnect would adopt the service as external.

    So there are two windows on this version: (a) service startup, before activation writes server-runtime.json, and (b) any non-owner server exiting and clearing the file. The ownership-conditional cleanup proposed above would fix (b). (a) probably needs the launcher to wait for or probe a known service unit (e.g. service-state.json / launchd label) before falling back to a managed launch.

    Filed via t3 triage by Claude Code (Claude Opus 5).

  6. marcellocurto commented on Sep 16, 2026

    @marcellocurto

    I hit the competing-server behavior on stable 0.0.42 after an in-app remote update. Investigation identified another trigger: SSH bootstrap replaces the running service's npm installation with a standalone archive before checking whether it can reuse that service.

    Observed sequence: I updated the Mac desktop through its update button, which succeeded, then accepted its prompt to update the remote server from 0.0.40 to 0.0.42. The remote update committed, but the desktop stayed at “Finishing an update,” then “did not respond during connection setup.” Restarting the desktop and reconnecting did not recover it. I ran no manual installation commands during this sequence.

    SSH configuration: My server permits SSH forwarding only to port 3773. T3's fallback server started on 3774, so that restriction blocked reconnection. Separately, the original service's runtime files had been deleted and its endpoint on 3773 returned 404. I did not test whether allowing 3774 would restore desktop connectivity. Reinstalling the managed service restored the connection on 3773 without changing the SSH policy.

    Environment: macOS 26.6.2 arm64 → built-in SSH connection → Debian 13 x86-64, systemd user service on 127.0.0.1:3773. The original service used node service-launcher.mjs; its only drop-in set PATH, with no ExecStart override.

    Evidence and mechanism

    Before repair:

    service-state.json: protocol=2, activeVersion=0.0.42, update.status=committed
    systemd service: active/running
    /proc/<service-child>/exe:
      ~/.t3/runtime/versions/0.0.42/node_modules/@t3code/t3-linux-x64/t3 (deleted)
    GET / on 3773: 404
    GET / on duplicate server at 3774: 200
    

    Source inspection explains the sequence:

    1. The running 0.0.40 updater installs the target through npm. The 0.0.42 npm package retains a compatibility entrypoint, allowing the old launcher to start it and commit the update.
    2. On reconnect, the 0.0.42 SSH launcher initializes its archive runner before discovering the existing service. Its installation check requires <versionDir>/t3, which the complete npm installation lacks.
    3. It then executes rm -rf "$T3_RUNTIME_DIR" and replaces that same version directory with the archive. The running process survives, but its executable and original static assets are gone.
    4. The resulting 404 fails the HTTP readiness check. The launcher starts another server on 3774 against the same T3 home, triggering the competing-server behavior discussed here.

    Isolated reproduction: I reproduced the replacement step using the captured SSH script and a disposable fixture with the npm directory layout. The old executable and assets were deleted, the process survived with /proc/<pid>/exe marked (deleted), and HTTP changed from 200 to 404. This used executable and HTTP stand-ins; I did not replay the full desktop update. A control run preserved a complete standalone installation.

    The legacy-launcher compatibility check in #11940, addressing #11934, merged after v0.0.42. In this incident, the update committed before SSH bootstrap removed its runtime. The SSH bootstrap file is unchanged between v0.0.42 and the inspected main commit, b900fc94.

    Recovery

    After backing up configuration and launch state, I disabled desktop retries, stopped the service and duplicate, cleared the duplicate's ownership files, and ran:

    ~/.t3/runtime/versions/0.0.42/t3 service install --base-dir "$HOME/.t3"

    The desktop then showed Connected, with only the service on 3773 remaining. Data was preserved.

    Suggested regression coverage: SSH bootstrap preserves a complete npm installation used by a live service, and a failed readiness check does not start a second server against the same T3 home.

  7. kyrregjerstad commented on Sep 18, 2026

    @kyrregjerstad

    Reproduced on stable 0.0.42: macOS Desktop connected over SSH/Tailscale to a Debian host running t3code.service on port 3773. During a temporary disk I/O stall, the service remained active but missed the 2-second readiness check, so Desktop launched another server on 3774 against the same ~/.t3; both then contended for SQLite and stopped responding. Restarting with only the systemd server restored service, and the ownership lock proposed in #9652 would have prevented the duplicate.

  8. taylormadearmy commented on Sep 24, 2026

    @taylormadearmy

    Reproduced on stable 0.0.42. Windows desktop clients connect over SSH to an Ubuntu host running t3code.service (t3 serve on 0.0.0.0:3773, ~/.t3).

    What happened: each time a desktop reconnected over SSH, the launcher started a managed server on 127.0.0.1:3774 against the same ~/.t3, while the service kept running. ssh-launch/<id>/server.log shows this happening repeatedly since July. Two different client machines did it.

    Symptoms:

    1. Codex active-writer conflict. Each server starts its own codex app-server. The managed server's Codex process kept a thread's rollout open, so a turn sent through the service failed with:
      ProviderAdapterProcessError: Provider adapter process error (codex) ...: thread 01a0c88a-… already has an active writer
        at startSession → ensureSessionForThread → processTurnStartRequested
      
      /proc/<pid>/fd confirmed that only the managed server's Codex process had the rollout open.
    2. The managed server deletes the service's server-runtime.json when it exits. The managed server overwrites the file on startup. On shutdown (killed, or when its SSH session closed) it removes the file, even though the service is still running. After that, t3 pair fails with NoRunningServerError until the file is rewritten by hand or the service restarts. While the managed server is running, t3 pair pairs with the localhost-only 3774 server instead of the service. This matches the "make runtime-state cleanup conditional on PID/start-time ownership" item in the proposed fix. [Bug]: t3 project CLI treats any 1s live-server probe failure as "no server": deletes server-runtime.json and writes offline behind a running server #7504 describes a similar deletion from the t3 project side.

    Workaround: remove the SSH environments from every desktop, pair them to the service as remote environments (http://<host>:3773), stop the managed server, and rewrite server-runtime.json with the service's PID.

  9. Rasalas commented on Sep 26, 2026

    @Rasalas

    Reproduced on 0.0.43-nightly.20260926.2282 on September 26: macOS desktop → built-in SSH connection over Tailscale → Ubuntu host.

    The original incident involved a background service and an SSH-launched server. The failure recurred after the background service had been disabled, with two servers running the same Nightly version against the same T3 home.

    Observed sequence, timestamps in UTC:

    • 08:33:40: Desktop logged ssh.environment.tunnel.existing.stale, with a 2,000 ms readiness budget and a 1,000 ms individual probe timeout.
    • 08:33:41: Desktop closed the SSH tunnel to remote port 3774.
    • 08:34:05: Desktop reported a newly launched managed server on 3773.
    • The previous server on 3774 was still alive. Both used the same --base-dir.
    • Codex processes remained children of the previous server and held actual thread writer locks. Attempts through the replacement server failed with already has an active writer; other threads showed Provider session did not survive a server restart.

    A previous workaround had marked the surviving server as externally managed. That prevented its termination but did not prevent a competing server from being launched.

    After gracefully stopping the older server:

    • Its Codex writer locks were released.
    • The shared server-runtime.json disappeared even though the newer server remained alive.
    • After restoring the surviving server’s runtime metadata, all eleven affected Codex histories successfully passed thread/resume. Eight interrupted project threads were subsequently resumed through T3.

    The host was under substantial load, but we have not established the exact reason for the slow readiness probe.

    Our temporary mitigation is external ownership plus a local startup guard that rejects another server targeting the same data directory. This protects the current installation; it does not fix the desktop timeout.

    This corroborates both the reconnect lifecycle problem addressed by #13521 and the need for the single-owner protection in #9652.

  10. PPPartners commented on Sep 26, 2026

    @PPPartners

    Local repro (no SSH): desktop's built-in backend + launchd background service on the same Mac, both driving the same Claude threads

    Posted by Claude (Claude Code, running as an agent inside T3 Code) on behalf of @StefanPernek. I was one of the affected sessions: my own thread got a second agent process too. Everything below comes from read-only inspection of processes, ~/.t3/userdata/logs and a read-only open of state.sqlite.

    Setup: T3 Code (Alpha) desktop 0.0.42 on macOS 26.6.2 arm64, Claude provider (Claude Code 2.1.282), serverExposureMode: network-accessible, Tailscale Serve enabled. The owner uses T3 Connect across two machines and one phone.

    What was running: two servers on the same ~/.t3, both with state.sqlite open:

    1. Desktop built-in backend: PID 4697, a child of the Electron app, listening on *:3773, recorded in server-runtime.json.
    2. Background service: ~/Library/LaunchAgents/com.t3tools.t3code.service.plist (RunAtLoad, KeepAlive), running ~/.t3/runtime/versions/0.0.42/t3 __service-launcher → t3 serve (PID 1573) on 127.0.0.1:59440, plus a cloudflared tunnel. The plist was created 2026-09-25 08:21 UTC through the npx package @t3code/t3-darwin-arm64, one minute before the desktop app was last launched. We believe T3 Connect setup installed it, which matches the trigger described in fix(server): prevent duplicate servers for one state directory #9652.

    What happened: at 17:01:51 UTC on 2026-09-26 the service started a new claude process for every open thread at once (8 threads), while the desktop backend's agents for those same threads were still running.

    • orchestration_events shows, for each of the 8 threads within ~180 ms, two identical thread.session-set events (status: ready, activeTurnId: null) with no causation event.
    • Afterwards each thread had two claude processes in the same worktree, one parented by 4697 (desktop) and one by 1573 (service).
    • The duplicates worked on the same task at the same time and saw each other's edits. One even "handed over" the ticket to "another session in this worktree".
    • About 6 s earlier (17:01:45 UTC), desktop.trace.ndjson shows the desktop re-running its environment bootstrap (getLocalEnvironmentBootstraps, backendConfiguration.resolvePrimaryLabel). That's the likely trigger, but we can't confirm it because the connection catalog is encrypted.
    • The same pattern had already happened twice earlier that day. boot-service.log has provider command reactor restarting provider session / claude.session.replacing at 13:47 and 13:57 UTC, which matches second agents on two other worktrees.

    Stopped by: launchctl bootout gui/$UID/com.t3tools.t3code.service. The service's orphaned agents exited within seconds, and each thread is back to one agent. The plist still has RunAtLoad, so this comes back at the next login.

    Why it matters beyond the SSH cases here: there's no SSH launcher involved. On a single Mac, "desktop app + T3 Connect background service" is enough to get two servers on one T3 home, and with the Claude provider that becomes duplicate agents running with full tool access in the same worktree. An ownership lock like #9652 (one server per state directory, service setup refusing a takeover) would prevent it. Ideally the desktop app would also use a running local service instead of starting its own backend. Happy to share more trace excerpts.

  11. elliason commented on Oct 7, 2026

    @elliason

    Another reproduction of the load-triggered case described above.

    Setup: desktop app, built-in SSH environment, Linux host that already runs T3 as a background service on the same base directory.

    The running server stayed healthy but answered slowly. On reconnect, the launcher's reuse check (REMOTE_REUSE_READY_TIMEOUT_MS = 2_000 in packages/ssh/src/tunnel.ts, repeated 1,000 ms probes) got no answer within that budget, treated the server as gone, and started a second managed server against the same base directory. Both servers then ran against the same base directory and state, producing the symptoms described above. The reuse budget is unchanged on current main.

    The single-owner lock proposed in #9652 and #14694 (both closed unmerged) would have stopped the second server from starting, though reconnect would still fail while the probe misses. Treating a live recorded server process as sufficient for reuse, as suggested above, would avoid both. +1 to raising the reuse budget toward the fresh-launch readiness timeout, as also suggested above.

  12. matheustimbo commented on Oct 8, 2026

    @matheustimbo
    Contributor

    Reproduced on 0.0.46-nightly.20261008.2801, ending in a malformed statev2.sqlite.

    Topology: macOS desktop → built-in SSH environment over Tailscale → Fedora 44 host running the t3code.service user unit (__service-launcher, Restart=always, RestartSec=5) on 127.0.0.1:3773, base dir ~/.t3.

    Timeline (host-local times):

    1. 21:24: added an SSH route to the existing environment. The launcher adopted the service correctly (managed file = external, port 3773).
    2. ~22:51: the service restarted onto a new nightly. During the restart window, the desktop reconnected, the reuse probe (REMOTE_REUSE_READY_TIMEOUT_MS = 2_000) failed, and pick_port returned 3773 itself, since the default port was briefly free. The launcher started its own t3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3.
    3. From then on, every __service-launcher child died with listen EADDRINUSE: address already in use 127.0.0.1:3773 and was restarted every 5 s. Each attempt boots against the same base dir before failing to bind.
    4. The two runtimes were also on different nightlies (service child 20261007.2787, launcher-managed 20261008.2801) sharing one database.
    5. ~22:55: the client showed database disk image is malformed and stayed on "Syncing messages…". The failing reads were ws.orchestrationV2.subscribeThread (event count over orchestration_events), ScheduledTaskService.listDueTasks, and ProjectionStore.getLimitRecoveryCandidates.

    DB damage: PRAGMA quick_check reported ~45 contiguous unreadable pages (2493913–2493955) at the very end of the file, past the main file's last page (the file had 2,493,909 pages while page_count reported 2,493,978). In other words, the lost pages were the ones that lived only in the WAL. .recover produced a clean DB: 34 orchestration_events, 2 orchestration_v2_projection_nodes and 5 orchestration_v2_projection_turn_items rows were lost.

    Second trigger during recovery: after I stopped everything and started t3code.service again, the desktop reconnected within ~5 s and the launcher started another managed server on 3774 with the same --base-dir, while the service was coming up on 3773. So any restart of the external service (update, crash, manual restart) is a window for a competing runtime.

    I can't prove that the overlap alone caused the WAL loss, since SQLite WAL normally tolerates multiple processes, and #11084 tracks the corruption side. But mixed versions plus a crash-looping process opening the same base dir every 5 s is the setup in which it happened.

    Suggestions on top of the proposed fix: have pick_port never return the port recorded in server-runtime.json or the default 3773 while an external owner is recorded, and take an exclusive lock on the base dir in t3 serve so a second runtime fails fast instead of sharing the DB.

    Workaround: removed the SSH route and kept only the Tailscale route to the service.

  13. g1331 commented on Oct 9, 2026

    @g1331

    Independent field confirmation on 0.0.45: Codex writer conflicts with two servers sharing one home

    Adding evidence from an incident on 2026-10-09. This matches the competing-service failure described here, including the runtime-file cleanup problem. This is a captured incident and verified recovery, not a fresh deterministic reproduction of the initial trigger.

    Environment and impact

    • Windows desktop, T3 Code 0.0.45; the executable was named T3 Code (Alpha), while Settings showed version 0.0.45 and the Stable update track.
    • Linux x86_64 SSH host, kernel 6.8.0-138-generic, T3 server 0.0.45.
    • Supported systemd user service: t3code.service → t3 __service-launcher → t3 serve.
    • Node v24.20.0; Codex CLI 0.162.0.
    • Built-in desktop SSH connection and T3 Connect for mobile access, using the default $HOME/.t3.
    • Two existing Codex conversations could not accept follow-ups. Both showed Failed; one showed idle/resumable subagents while follow-ups failed with already has an active writer.

    Observed state before recovery

    Two live T3 servers used the same data home:

    <service-launcher-pid> -> <service-server-pid>
      ~/.t3/runtime/versions/0.0.45/t3 serve
      LISTEN 127.0.0.1:46415
    
    <ssh-server-pid>, PPID 1
      ~/.t3/runtime/versions/0.0.45/t3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3
      LISTEN 127.0.0.1:3773
    
    desktop SSH setup:
      remotePort=3773
      remoteServerKind=managed
    

    Each server had a cloudflared child. The logs showed both registering the same T3 Connect tunnel ID and name.

    The service-owned Codex processes remained alive. Inspection of /proc/<codex-pid>/fd identified the actual writer owners: the service-owned Codex processes held the affected conversations' thread-writer-locks/<thread-id>.lock files and rollout files. The replacement server's trace recorded:

    CodexSessionRuntime.start
      CodexAppServerRequestError:
      thread <codex-thread-A> already has an active writer
    
    startSession -> ensureSessionForThread -> processTurnStartRequested
      provider.instance_id=codex
      provider.resume_cursor.source=persisted
      provider.resume_cursor.present=true
    

    The second conversation's saved error referred to a different Codex thread, whose lock was also held by a service-owned Codex process. This was not inferred solely from the UI's Failed labels.

    Runtime-file ownership and recovery

    While both servers were alive, userdata/server-runtime.json advertised the SSH-managed server on 3773, rather than the still-running service on 46415.

    1. Switched off the desktop SSH environment. The SSH-managed server and its tunnel exited; the systemd server and its Codex children remained alive.
    2. Checked immediately afterward: server-runtime.json no longer existed, although the systemd server was still healthy. The SSH launch ownership files were also gone.
    3. Backed up the conversation database and normally restarted only t3code.service. No saved session had an active turn at that point. The service published a new runtime record with serviceManaged: true, on 3773.
    4. Re-enabled the same desktop SSH environment. Desktop setup now reported remoteServerKind=external; ssh-launch/<state-key>/managed contained external. Only one T3 server and one cloudflared child remained.
    5. Sent a minimal connectivity-only prompt in each original conversation, explicitly prohibiting tool calls, file changes, or continuation of project work. Both replied successfully. Read-back showed status=ready, active_turn_id=NULL, and last_error=NULL; both Failed labels disappeared.
    6. T3 Connect logged successful tunnel registrations again. The phone itself was not operated during verification.

    No session histories, writer lock files, or project files were deleted. No application source was patched.

    What is established, and what is still unknown

    The immediate failure is established: two T3 runtimes shared one home, the original service's Codex processes retained the writers, and the desktop was sending resumes through the competing runtime.

    The first reason SSH failed to adopt the service is not established by the retained incident evidence. A readiness timeout, startup race, or another discovery failure may explain it, but I cannot claim which occurred. The later overwrite/removal of the shared runtime record was directly observed.

    The installed bundle agrees with the v0.0.45 source:

    • SSH discovery and launch: discovers the default service through the shared runtime record and can launch against that same home.
    • Runtime-state finalizer: clears that runtime path on release, without a caller-side PID/ownership check.

    This supports the existing requests for one live server per data home, ownership-aware runtime cleanup, and preserving external-service ownership through failed readiness checks. The recovery above is a workaround; it does not demonstrate that reconnect/update races are fixed.

    Redaction

    Host/IP addresses, usernames, project names and paths, repository/PR details, real process and conversation IDs, SSH state keys, tunnel identifiers, pairing codes, tokens, and raw screenshots/databases/logs have been omitted or replaced with placeholders. Only relevant diagnostic excerpts are included.

    Investigation and report preparation used Codex on the affected user's behalf.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions