Repository navigation
SSH reconnect can replace an external T3 service with a competing managed runtime #5749
Description
Activity
Reproduction on a persistent systemd service (Nightly
0.0.34-20260815.1100)We hit this again today on a Linux host that already runs the supported user systemd unit (
t3code.service→service-launcher.mjs→t3 serveon127.0.0.1:3773,managed=external). The desktop is Nightly0.0.34-nightly.20260815.1100, same version as the service.This is not a missing service. The boot service stayed
activewithNRestarts=0the whole time.Sequence
-
Service environment probe on
127.0.0.1:3773/.well-known/t3/environmentwas200in 17ms while the host was idle enough. -
The same host then had live agent work under the service cgroup (
gate.sh/cargo test). The next probe was still200but 278ms. -
Desktop marked the existing SSH tunnel stale after its 2s readiness budget (
SshReadinessError: Timed out waiting 2000ms for backend readinesson the local forward). -
launchOrReuseRemoteServerthen started a competing process:t3 serve --host 127.0.0.1 --port 3774 --base-dir $HOME/.t3 -
A later retry did the same on 3775.
-
All three Node processes (
3773service child + two launcher-managed serves) hadstate.sqlite/-wal/-shmopen. All three wentDl. After that, every port timed out, including the original service.
Desktop error (trimmed, no tokens):
SshCommandError: Remote T3 server did not become ready on 127.0.0.1:3774. WARN: Grok ACP model discovery timed out after 15000ms. WARN: failed to stop provider service { errorTag: 'ProviderSessionDirectoryPersistenceError' }ssh-launch/*/managedflipped fromexternaltomanaged,port=3775.Source
This matches current
packages/ssh/src/tunnel.tsonmain:REMOTE_REUSE_READY_TIMEOUT_MS = 2_000vsREMOTE_READY_TIMEOUT_MS = 15_000for a newly spawned server- On reuse miss,
pick_portskips busy3773and runst3 serve --base-dir "$DEFAULT_SERVER_HOME"on the next free port - There is no exclusive lock on that base dir /
state.sqlite - If
managed=managedand the pid file is empty,PID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}"can target the systemd child
packages/ssh/src/tunnel.test.tscurrently asserts thatPID_TO_STOPfallback as intended.Why 2s is not enough here
The reuse probe is
GET /with a 1s per-request timeout. On a host where the service cgroup is doing real work, 3773 can still be healthy and miss a 2s desktop/reuse budget. The fallback is not “retry the existing owner”; it is “start another server on the same data directory.” That is what wedges SQLite, after which even the original owner stops answering and every further reconnect makes it worse.#5751
#5751 is the right contract: if an external owner exists and is not ready, exit 1. Do not start a second server. Do not kill
$DEFAULT_RUNTIME_PID.That still leaves the 2s reuse budget. Fail-closed would have saved the database today; it would still have shown a reconnect error on this busy host. The reuse timeout needs to be in the same league as
REMOTE_READY_TIMEOUT_MS(15s), or reuse should treat a liveserver-runtime.jsonpid + occupied port as sufficient and not require a 2s HTTP win.Workaround we are taking
We are moving these persistent service hosts off Desktop-Managed SSH Launch onto operator-owned SSH forwards + ordinary pairing to the already-running service. SSH Launch is the wrong access method for a machine that already has
t3 service install.-
Workaround now in use
We stopped using Desktop-Managed SSH Launch against these persistent systemd hosts.
Current setup:
- Each Linux host keeps
t3code.service(user unit, linger enabled) as the only server on127.0.0.1:3773. ~/.local/bin/t3is a small guard that refusest3 servewhile that unit is active, so a leftover SSH-launch attempt cannot open the samestate.sqlite.- The desktop reaches the service through operator-owned
ssh -Nforwards on the Mac (launchd+KeepAlive). It is paired tohttp://127.0.0.1:<local-port>as a normal environment, not as an SSH-launched environment.
That matches the documented “pre-existing server” model. SSH Launch remains unsafe on Nightly until reuse stops spawning a second process on the same base dir (#5751) and the 2s reuse budget is raised.
The remote server does not depend on the laptop: logout, sleep, and LAN flaps leave the systemd unit running. Only the desktop’s WebSocket drops until the Mac-side forward comes back.
- Each Linux host keeps
Adding an independent reproduction on 0.0.35, in a topology that I think widens the scope of this issue slightly, plus one finding that makes the teardown path worse than the analysis above assumes.
This does not require an external service
The reproduction in the comment above starts from a host running the supported systemd unit (
managed=external). On the host I looked at there is no systemd unit at all (t3 service statusreports not installed), so every server is launcher-managedfrom the start. The runaway is identical. So the failure is not specifically about replacing an external service; a puremanagedhost escalates the same way, which suggests the fix needs to cover the plain managed path too, not only external-ownership preservation.Topology: desktop app on a separate machine, remote Linux host reached over Tailscale through the SSH launcher, server pinned to
t3@0.0.35byrun-t3.sh.The
PID_TO_STOPfallback is worse than "can target the systemd child"The comment above notes that
PID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}"can target the wrong process. On any host whererun-t3.shtakes the npx branch (no globalt3on PATH), it is worse than that:$REMOTE_PIDnever identifies a server at all, so the kill is a no-op and the old server is always leaked.REMOTE_PID="$!"attunnel.ts:594captures thenpm execwrapper, becausenpxspawns rather than exec-chains. Live capture during the failure:2618056 npm exec t3@0.0.35 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3 <- $PID_FILE 2618094 sh -c "t3" serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3 2618095 node ~/.npm/_npx/<hash>/node_modules/.bin/t3 serve ... --port 3773 <- real serverI reproduced the signal behaviour with a synthetic package in the same three-layer shape: SIGTERM to the
$!pid kills the wrapper, and the node process survives with its port still bound, reparented toppid=1. Full write-up and repro script in #2614, which I think is the orphaning half of this same story.That closes the loop on why the escalation is unbounded rather than self-limiting: the reuse probe decides to replace the server, and the replacement path is structurally incapable of stopping the one it is replacing.
Evidence from this host
- Nine distinct ports across the launcher logs, 3773 through 3781, five live simultaneously, every one launched with the same
--base-dirand therefore the samestate.sqlite. PersistenceSqlError: SQL error in OrchestrationEventStore.append:insertwithError: database is locked: 463 occurrences across two log files.thread <id> already has an active writeron codex threads, because the codex app-server stays attached to a server the desktop has abandoned.- The same
WARN: failed to stop provider service { errorTag: 'ProviderSessionDirectoryPersistenceError' }reported above. - Desktop side surfaced as
Failed to connect. Reconnecting...then<host> did not respond during connection setup, looping indefinitely.
Two launches 30 seconds apart during one loop, showing the ladder in motion:
[08:17:34.021] INFO (#259): Listening on http://127.0.0.1:3776 [08:18:04.608] INFO (#234): Listening on http://127.0.0.1:3777The spawn rate is bounded only by the retry ladder
packages/client-runtime/src/connection/supervisor.ts:const RETRY_DELAYS_MS = [3_000, 4_000, 8_000, 16_000] as const; // :32 const CONNECTION_ESTABLISHMENT_TIMEOUT = "15 seconds"; // :33
The ladder caps at 16s and retries forever, and each attempt that reaches
ensureTunnelEntrywith a stale entry runs the remote launch script again. So a host that stays above the 2s reuse budget gains roughly one additional server per 16-31s indefinitely, each one making the next probe likelier to miss.Data volume is not the trigger
Worth ruling out, since the database on this host had grown to 3.2 GB (336k
thread.activity-appendedrows, 1.22 GB of payload, no retention or pruning anywhere in the server source). With a single server alive against that same database,GET /and/.well-known/t3/environmentboth answer in 1-5 ms. The 2-4s responses that trip the 2s budget appear only once duplicates are contending for the WAL. Contention is the cause, not size, so raisingbusy_timeoutor shrinking the database would not address this.The initial nudge on this host was ordinary background load: ~29k
GitVcsDriver.fetchRemoteForStatustimeouts across four repos and ~17kFailed to flush telemetryfetch failures. Once one probe misses, the loop sustains itself without needing that trigger again.Still current
packages/ssh/src/tunnel.tsis byte-identical betweenv0.0.35andmainat5b8445b7a777.REMOTE_REUSE_READY_TIMEOUT_MS = 2_000,SSH_READY_PROBE_TIMEOUT_MS = 1_000andREMOTE_PID="$!"are all unchanged as of 0.0.38.Environment
- Remote host: Linux x64 (7.0.0-30-generic), Node v24.16.0, t3 server 0.0.35, no global
t3on PATH (npx branch ofrun-t3.sh), no systemd unit - Desktop app on a separate machine over Tailscale
Filed via
t3 triage. Investigation and reproduction by Claude Opus 5 (1M context) running as the triage agent in Claude Code.- Nine distinct ports across the launcher logs, 3773 through 3781, five live simultaneously, every one launched with the same
Still hitting this on
0.0.39-nightly.20260903.1272, with the same version on the Mac desktop and an Ubuntu EC2 host connected through Desktop's SSH launcher. Running Codex tasks end up with:Provider session did not survive a server restart. Send a new message to continue.This also happens without the systemd service. We previously had a boot-service server and an SSH-launched server sharing the same T3 home. Removing the boot service cleared the active-writer conflicts, but it did not stop the reconnect/restart failures.
Latest incident, September 4 (UTC), from the desktop trace and remote runtime record:
07:53:58: the existing SSH tunnel to remote port3773closed.07:55:28:ssh/tunnel.ensureTunnelEntryfailed withSSH command timed out after 90000ms.07:57:58: a new remote server wrote its runtime record on port3774.07:58:03: SSH environment setup and pairing succeeded. The interrupted tasks still showed the server-restart error.
EC2 did not reboot: uptime was 30 days. The check immediately afterward showed 42 GiB RAM available, 62 GiB free on the root disk, and no OOM-kill entries in the kernel journal for the checked incident window. The new server answered HTTP requests normally.
There is also a same-PID problem in the launcher code inside the installed Nightly app, not just in a source checkout. When the runtime record is healthy and
REMOTE_MANAGEDismanaged, it assignsPID_TO_STOP="${REMOTE_PID:-$DEFAULT_RUNTIME_PID}"and kills it without checking whether the discovered runtime is that same managed process. That appears to let reconnect kill the server it should reuse. We did not capture the signal sender for this particular restart, so that part is code analysis rather than a traced kill event.Adding this as current-version evidence for the managed-only case above. Reconnecting the desktop should restore the connection without replacing a healthy remote server or ending its running tasks.
Still reproducing on
0.0.41-nightly.20260915.1735(macOS arm64, launchd service), desktop + service on the same version.Topology: Mac desktop app → SSH launcher → Mac mini running
com.t3tools.t3code.service(launchd,127.0.0.1:3773,~/.t3).Observed today:
- After a reboot, the SSH launcher started a
managedserver on 3775 against the same~/.t3while the launchd service was also running on 3773, so both hadstate.sqliteopen.server-runtime.jsonpointed at the managed server. - After stopping the managed server and running
t3 service install, the service took ~46s to start listening. The desktop reconnected 6s after the service launched, found noserver-runtime.jsonyet, and started anothermanagedserver (3774), which wrote its ownserver-runtime.json. - When that managed server was stopped with SIGTERM (after the service had rewritten
server-runtime.jsonwithserviceManaged: true), it deleted the service'sserver-runtime.jsonon the way out. The release inapps/server/src/server.tsruntimeStateLayercallsclearPersistedServerRuntimeStateunconditionally, with no PID/ownership check. I had to restore the file by hand before the next reconnect would adopt the service asexternal.
So there are two windows on this version: (a) service startup, before activation writes
server-runtime.json, and (b) any non-owner server exiting and clearing the file. The ownership-conditional cleanup proposed above would fix (b). (a) probably needs the launcher to wait for or probe a known service unit (e.g.service-state.json/ launchd label) before falling back to a managed launch.Filed via
t3 triageby Claude Code (Claude Opus 5).- After a reboot, the SSH launcher started a
I hit the competing-server behavior on stable 0.0.42 after an in-app remote update. Investigation identified another trigger: SSH bootstrap replaces the running service's npm installation with a standalone archive before checking whether it can reuse that service.
Observed sequence: I updated the Mac desktop through its update button, which succeeded, then accepted its prompt to update the remote server from 0.0.40 to 0.0.42. The remote update committed, but the desktop stayed at “Finishing an update,” then “did not respond during connection setup.” Restarting the desktop and reconnecting did not recover it. I ran no manual installation commands during this sequence.
SSH configuration: My server permits SSH forwarding only to port 3773. T3's fallback server started on 3774, so that restriction blocked reconnection. Separately, the original service's runtime files had been deleted and its endpoint on 3773 returned 404. I did not test whether allowing 3774 would restore desktop connectivity. Reinstalling the managed service restored the connection on 3773 without changing the SSH policy.
Environment: macOS 26.6.2 arm64 → built-in SSH connection → Debian 13 x86-64, systemd user service on
127.0.0.1:3773. The original service usednode service-launcher.mjs; its only drop-in setPATH, with noExecStartoverride.Evidence and mechanism
Before repair:
service-state.json: protocol=2, activeVersion=0.0.42, update.status=committed systemd service: active/running /proc/<service-child>/exe: ~/.t3/runtime/versions/0.0.42/node_modules/@t3code/t3-linux-x64/t3 (deleted) GET / on 3773: 404 GET / on duplicate server at 3774: 200Source inspection explains the sequence:
- The running 0.0.40 updater installs the target through npm. The 0.0.42 npm package retains a compatibility entrypoint, allowing the old launcher to start it and commit the update.
- On reconnect, the 0.0.42 SSH launcher initializes its archive runner before discovering the existing service. Its installation check requires
<versionDir>/t3, which the complete npm installation lacks. - It then executes
rm -rf "$T3_RUNTIME_DIR"and replaces that same version directory with the archive. The running process survives, but its executable and original static assets are gone. - The resulting 404 fails the HTTP readiness check. The launcher starts another server on 3774 against the same T3 home, triggering the competing-server behavior discussed here.
Isolated reproduction: I reproduced the replacement step using the captured SSH script and a disposable fixture with the npm directory layout. The old executable and assets were deleted, the process survived with
/proc/<pid>/exemarked(deleted), and HTTP changed from 200 to 404. This used executable and HTTP stand-ins; I did not replay the full desktop update. A control run preserved a complete standalone installation.The legacy-launcher compatibility check in #11940, addressing #11934, merged after v0.0.42. In this incident, the update committed before SSH bootstrap removed its runtime. The SSH bootstrap file is unchanged between v0.0.42 and the inspected main commit,
b900fc94.Recovery
After backing up configuration and launch state, I disabled desktop retries, stopped the service and duplicate, cleared the duplicate's ownership files, and ran:
~/.t3/runtime/versions/0.0.42/t3 service install --base-dir "$HOME/.t3"
The desktop then showed Connected, with only the service on 3773 remaining. Data was preserved.
Suggested regression coverage: SSH bootstrap preserves a complete npm installation used by a live service, and a failed readiness check does not start a second server against the same T3 home.
Reproduced on stable
0.0.42: macOS Desktop connected over SSH/Tailscale to a Debian host runningt3code.serviceon port 3773. During a temporary disk I/O stall, the service remained active but missed the 2-second readiness check, so Desktop launched another server on 3774 against the same~/.t3; both then contended for SQLite and stopped responding. Restarting with only the systemd server restored service, and the ownership lock proposed in #9652 would have prevented the duplicate.Reproduced on stable 0.0.42. Windows desktop clients connect over SSH to an Ubuntu host running
t3code.service(t3 serveon0.0.0.0:3773,~/.t3).What happened: each time a desktop reconnected over SSH, the launcher started a
managedserver on127.0.0.1:3774against the same~/.t3, while the service kept running.ssh-launch/<id>/server.logshows this happening repeatedly since July. Two different client machines did it.Symptoms:
- Codex active-writer conflict. Each server starts its own
codex app-server. The managed server's Codex process kept a thread's rollout open, so a turn sent through the service failed with:ProviderAdapterProcessError: Provider adapter process error (codex) ...: thread 01a0c88a-… already has an active writer at startSession → ensureSessionForThread → processTurnStartRequested/proc/<pid>/fdconfirmed that only the managed server's Codex process had the rollout open. - The managed server deletes the service's
server-runtime.jsonwhen it exits. The managed server overwrites the file on startup. On shutdown (killed, or when its SSH session closed) it removes the file, even though the service is still running. After that,t3 pairfails withNoRunningServerErroruntil the file is rewritten by hand or the service restarts. While the managed server is running,t3 pairpairs with the localhost-only 3774 server instead of the service. This matches the "make runtime-state cleanup conditional on PID/start-time ownership" item in the proposed fix. [Bug]:t3 projectCLI treats any 1s live-server probe failure as "no server": deletes server-runtime.json and writes offline behind a running server #7504 describes a similar deletion from thet3 projectside.
Workaround: remove the SSH environments from every desktop, pair them to the service as remote environments (
http://<host>:3773), stop the managed server, and rewriteserver-runtime.jsonwith the service's PID.- Codex active-writer conflict. Each server starts its own
Reproduced on
0.0.43-nightly.20260926.2282on September 26: macOS desktop → built-in SSH connection over Tailscale → Ubuntu host.The original incident involved a background service and an SSH-launched server. The failure recurred after the background service had been disabled, with two servers running the same Nightly version against the same T3 home.
Observed sequence, timestamps in UTC:
- 08:33:40: Desktop logged
ssh.environment.tunnel.existing.stale, with a 2,000 ms readiness budget and a 1,000 ms individual probe timeout. - 08:33:41: Desktop closed the SSH tunnel to remote port 3774.
- 08:34:05: Desktop reported a newly launched managed server on 3773.
- The previous server on 3774 was still alive. Both used the same
--base-dir. - Codex processes remained children of the previous server and held actual thread writer locks. Attempts through the replacement server failed with
already has an active writer; other threads showedProvider session did not survive a server restart.
A previous workaround had marked the surviving server as externally managed. That prevented its termination but did not prevent a competing server from being launched.
After gracefully stopping the older server:
- Its Codex writer locks were released.
- The shared
server-runtime.jsondisappeared even though the newer server remained alive. - After restoring the surviving server’s runtime metadata, all eleven affected Codex histories successfully passed
thread/resume. Eight interrupted project threads were subsequently resumed through T3.
The host was under substantial load, but we have not established the exact reason for the slow readiness probe.
Our temporary mitigation is external ownership plus a local startup guard that rejects another server targeting the same data directory. This protects the current installation; it does not fix the desktop timeout.
This corroborates both the reconnect lifecycle problem addressed by #13521 and the need for the single-owner protection in #9652.
- 08:33:40: Desktop logged
Local repro (no SSH): desktop's built-in backend + launchd background service on the same Mac, both driving the same Claude threads
Posted by Claude (Claude Code, running as an agent inside T3 Code) on behalf of @StefanPernek. I was one of the affected sessions: my own thread got a second agent process too. Everything below comes from read-only inspection of processes,
~/.t3/userdata/logsand a read-only open ofstate.sqlite.Setup: T3 Code (Alpha) desktop 0.0.42 on macOS 26.6.2 arm64, Claude provider (Claude Code 2.1.282),
serverExposureMode: network-accessible, Tailscale Serve enabled. The owner uses T3 Connect across two machines and one phone.What was running: two servers on the same
~/.t3, both withstate.sqliteopen:- Desktop built-in backend: PID 4697, a child of the Electron app, listening on
*:3773, recorded inserver-runtime.json. - Background service:
~/Library/LaunchAgents/com.t3tools.t3code.service.plist(RunAtLoad,KeepAlive), running~/.t3/runtime/versions/0.0.42/t3 __service-launcher→t3 serve(PID 1573) on127.0.0.1:59440, plus a cloudflared tunnel. The plist was created 2026-09-25 08:21 UTC through the npx package@t3code/t3-darwin-arm64, one minute before the desktop app was last launched. We believe T3 Connect setup installed it, which matches the trigger described in fix(server): prevent duplicate servers for one state directory #9652.
What happened: at 17:01:51 UTC on 2026-09-26 the service started a new
claudeprocess for every open thread at once (8 threads), while the desktop backend's agents for those same threads were still running.orchestration_eventsshows, for each of the 8 threads within ~180 ms, two identicalthread.session-setevents (status: ready,activeTurnId: null) with no causation event.- Afterwards each thread had two
claudeprocesses in the same worktree, one parented by 4697 (desktop) and one by 1573 (service). - The duplicates worked on the same task at the same time and saw each other's edits. One even "handed over" the ticket to "another session in this worktree".
- About 6 s earlier (17:01:45 UTC),
desktop.trace.ndjsonshows the desktop re-running its environment bootstrap (getLocalEnvironmentBootstraps,backendConfiguration.resolvePrimaryLabel). That's the likely trigger, but we can't confirm it because the connection catalog is encrypted. - The same pattern had already happened twice earlier that day.
boot-service.loghasprovider command reactor restarting provider session/claude.session.replacingat 13:47 and 13:57 UTC, which matches second agents on two other worktrees.
Stopped by:
launchctl bootout gui/$UID/com.t3tools.t3code.service. The service's orphaned agents exited within seconds, and each thread is back to one agent. The plist still hasRunAtLoad, so this comes back at the next login.Why it matters beyond the SSH cases here: there's no SSH launcher involved. On a single Mac, "desktop app + T3 Connect background service" is enough to get two servers on one T3 home, and with the Claude provider that becomes duplicate agents running with full tool access in the same worktree. An ownership lock like #9652 (one server per state directory, service setup refusing a takeover) would prevent it. Ideally the desktop app would also use a running local service instead of starting its own backend. Happy to share more trace excerpts.
- Desktop built-in backend: PID 4697, a child of the Electron app, listening on
Another reproduction of the load-triggered case described above.
Setup: desktop app, built-in SSH environment, Linux host that already runs T3 as a background service on the same base directory.
The running server stayed healthy but answered slowly. On reconnect, the launcher's reuse check (
REMOTE_REUSE_READY_TIMEOUT_MS = 2_000inpackages/ssh/src/tunnel.ts, repeated 1,000 ms probes) got no answer within that budget, treated the server as gone, and started a secondmanagedserver against the same base directory. Both servers then ran against the same base directory and state, producing the symptoms described above. The reuse budget is unchanged on currentmain.The single-owner lock proposed in #9652 and #14694 (both closed unmerged) would have stopped the second server from starting, though reconnect would still fail while the probe misses. Treating a live recorded server process as sufficient for reuse, as suggested above, would avoid both. +1 to raising the reuse budget toward the fresh-launch readiness timeout, as also suggested above.
Reproduced on
0.0.46-nightly.20261008.2801, ending in a malformedstatev2.sqlite.Topology: macOS desktop → built-in SSH environment over Tailscale → Fedora 44 host running the
t3code.serviceuser unit (__service-launcher,Restart=always,RestartSec=5) on127.0.0.1:3773, base dir~/.t3.Timeline (host-local times):
- 21:24: added an SSH route to the existing environment. The launcher adopted the service correctly (
managedfile =external, port3773). - ~22:51: the service restarted onto a new nightly. During the restart window, the desktop reconnected, the reuse probe (
REMOTE_REUSE_READY_TIMEOUT_MS = 2_000) failed, andpick_portreturned 3773 itself, since the default port was briefly free. The launcher started its ownt3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3. - From then on, every
__service-launcherchild died withlisten EADDRINUSE: address already in use 127.0.0.1:3773and was restarted every 5 s. Each attempt boots against the same base dir before failing to bind. - The two runtimes were also on different nightlies (service child
20261007.2787, launcher-managed20261008.2801) sharing one database. - ~22:55: the client showed
database disk image is malformedand stayed on "Syncing messages…". The failing reads werews.orchestrationV2.subscribeThread(event count overorchestration_events),ScheduledTaskService.listDueTasks, andProjectionStore.getLimitRecoveryCandidates.
DB damage:
PRAGMA quick_checkreported ~45 contiguous unreadable pages (2493913–2493955) at the very end of the file, past the main file's last page (the file had 2,493,909 pages whilepage_countreported 2,493,978). In other words, the lost pages were the ones that lived only in the WAL..recoverproduced a clean DB: 34orchestration_events, 2orchestration_v2_projection_nodesand 5orchestration_v2_projection_turn_itemsrows were lost.Second trigger during recovery: after I stopped everything and started
t3code.serviceagain, the desktop reconnected within ~5 s and the launcher started another managed server on3774with the same--base-dir, while the service was coming up on3773. So any restart of the external service (update, crash, manual restart) is a window for a competing runtime.I can't prove that the overlap alone caused the WAL loss, since SQLite WAL normally tolerates multiple processes, and #11084 tracks the corruption side. But mixed versions plus a crash-looping process opening the same base dir every 5 s is the setup in which it happened.
Suggestions on top of the proposed fix: have
pick_portnever return the port recorded inserver-runtime.jsonor the default3773while an external owner is recorded, and take an exclusive lock on the base dir int3 serveso a second runtime fails fast instead of sharing the DB.Workaround: removed the SSH route and kept only the Tailscale route to the service.
- 21:24: added an SSH route to the existing environment. The launcher adopted the service correctly (
Independent field confirmation on 0.0.45: Codex writer conflicts with two servers sharing one home
Adding evidence from an incident on 2026-10-09. This matches the competing-service failure described here, including the runtime-file cleanup problem. This is a captured incident and verified recovery, not a fresh deterministic reproduction of the initial trigger.
Environment and impact
- Windows desktop, T3 Code 0.0.45; the executable was named T3 Code (Alpha), while Settings showed version 0.0.45 and the Stable update track.
- Linux x86_64 SSH host, kernel 6.8.0-138-generic, T3 server 0.0.45.
- Supported systemd user service:
t3code.service→t3 __service-launcher→t3 serve. - Node v24.20.0; Codex CLI 0.162.0.
- Built-in desktop SSH connection and T3 Connect for mobile access, using the default
$HOME/.t3. - Two existing Codex conversations could not accept follow-ups. Both showed Failed; one showed idle/resumable subagents while follow-ups failed with
already has an active writer.
Observed state before recovery
Two live T3 servers used the same data home:
<service-launcher-pid> -> <service-server-pid> ~/.t3/runtime/versions/0.0.45/t3 serve LISTEN 127.0.0.1:46415 <ssh-server-pid>, PPID 1 ~/.t3/runtime/versions/0.0.45/t3 serve --host 127.0.0.1 --port 3773 --base-dir ~/.t3 LISTEN 127.0.0.1:3773 desktop SSH setup: remotePort=3773 remoteServerKind=managedEach server had a cloudflared child. The logs showed both registering the same T3 Connect tunnel ID and name.
The service-owned Codex processes remained alive. Inspection of
/proc/<codex-pid>/fdidentified the actual writer owners: the service-owned Codex processes held the affected conversations'thread-writer-locks/<thread-id>.lockfiles and rollout files. The replacement server's trace recorded:CodexSessionRuntime.start CodexAppServerRequestError: thread <codex-thread-A> already has an active writer startSession -> ensureSessionForThread -> processTurnStartRequested provider.instance_id=codex provider.resume_cursor.source=persisted provider.resume_cursor.present=trueThe second conversation's saved error referred to a different Codex thread, whose lock was also held by a service-owned Codex process. This was not inferred solely from the UI's Failed labels.
Runtime-file ownership and recovery
While both servers were alive,
userdata/server-runtime.jsonadvertised the SSH-managed server on 3773, rather than the still-running service on 46415.- Switched off the desktop SSH environment. The SSH-managed server and its tunnel exited; the systemd server and its Codex children remained alive.
- Checked immediately afterward:
server-runtime.jsonno longer existed, although the systemd server was still healthy. The SSH launch ownership files were also gone. - Backed up the conversation database and normally restarted only
t3code.service. No saved session had an active turn at that point. The service published a new runtime record withserviceManaged: true, on 3773. - Re-enabled the same desktop SSH environment. Desktop setup now reported
remoteServerKind=external;ssh-launch/<state-key>/managedcontainedexternal. Only one T3 server and one cloudflared child remained. - Sent a minimal connectivity-only prompt in each original conversation, explicitly prohibiting tool calls, file changes, or continuation of project work. Both replied successfully. Read-back showed
status=ready,active_turn_id=NULL, andlast_error=NULL; both Failed labels disappeared. - T3 Connect logged successful tunnel registrations again. The phone itself was not operated during verification.
No session histories, writer lock files, or project files were deleted. No application source was patched.
What is established, and what is still unknown
The immediate failure is established: two T3 runtimes shared one home, the original service's Codex processes retained the writers, and the desktop was sending resumes through the competing runtime.
The first reason SSH failed to adopt the service is not established by the retained incident evidence. A readiness timeout, startup race, or another discovery failure may explain it, but I cannot claim which occurred. The later overwrite/removal of the shared runtime record was directly observed.
The installed bundle agrees with the v0.0.45 source:
- SSH discovery and launch: discovers the default service through the shared runtime record and can launch against that same home.
- Runtime-state finalizer: clears that runtime path on release, without a caller-side PID/ownership check.
This supports the existing requests for one live server per data home, ownership-aware runtime cleanup, and preserving external-service ownership through failed readiness checks. The recovery above is a workaround; it does not demonstrate that reconnect/update races are fixed.
Redaction
Host/IP addresses, usernames, project names and paths, repository/PR details, real process and conversation IDs, SSH state keys, tunnel identifiers, pairing codes, tokens, and raw screenshots/databases/logs have been omitted or replaced with placeholders. Only relevant diagnostic excerpts are included.
Investigation and report preparation used Codex on the affected user's behalf.
- added a commit that references this issue
on Oct 9, 2026
Problem
On a persistent SSH host that already runs the supported T3 systemd service, reconnecting from an updated desktop client can start a second T3 server and a second T3 Connect tunnel against the same base directory.
The failure is deterministic when launcher state says
managedwhileserver-runtime.jsonbelongs to the launcher-managed process, or when a recorded external service is briefly unavailable during its own update. Requests and WebSockets can then land on different servers, producing intermittent reconnects and messages that appear to disappear.Expected behavior
The supported external service remains the sole server/tunnel owner. The SSH launcher adopts its advertised port, records external ownership, and never replaces a temporarily restarting external service.
Proposed fix
The tested source change:
A focused PR follows.