Repository navigation
[Bug]: Restarting the desktop client breaks running remote Codex threads over SSH #10934
Description
Activity
Triage
Confirmed on current
main(6c583620). This is a real managed-SSH lifecycle regression, not a Codex or persistence bug.Restarting the desktop client tears down
SshEnvironmentManager. That scope close walks every tunnel and the tunnel finalizer always callsstopRemoteServer()(packages/ssh/src/tunnel.ts). The generated stop script SIGTERMs the recorded PID unless the host is markedexternal.#9843 is the trigger, and it should stay. It made the runner
execthe resolved CLI so the recorded PID is the actualt3 serve. Before that, the same finalizer only killed the npm/npx wrapper and the server (and the in-flight turn) survived. That accidental persistence was the remote workflow people were using.Reconnect is not the killer. The launch script already reuses a live managed PID. By the time the client comes back, the server is already gone, so startup reconciliation marks running sessions with
Provider session did not survive a server restart….Continue threads after restartsdefaults to off and would only start a new continuation turn anyway — it would not keep the live provider session.Distinct from
- #2614: unbounded orphans because stop missed the real PID. Opposite failure.
- #5749: reconnect replacing an external systemd service.
- #10928: continue-after-restart drops Codex reasoning effort.
No open PR already fixes the quit-time stop (9687 / 10588 / 9652 are adjacent, not this).
Product contract
User docs say removing the connection stops a server T3 launched; a server that was already running is left alone. Internals say the same: launcher-owned vs discovered/external. The finalizer currently treats any tunnel/app teardown — including Quit — as a managed-server stop. That is stricter than the documented “remove connection” rule and is what broke long-running remote work.
#2614 wanted Disconnect or Quit to stop the server so orphans could not accumulate. After #9843, stop-when-invoked is reliable. Leaving one correctly tracked server alive across a client quit is not that leak.
Suggested fix
Keep #9843. Change ownership of the stop:
- Client quit / window close / manager teardown: close the local SSH forward only. Do not stop a healthy managed remote server.
- Explicit Disconnect / Remove connection: keep
stopRemoteServer()so managed hosts can still be torn down on purpose.
npx t3 service install(host recordedexternal) remains the supported workaround and the right setup for a persistent remote that must outlive the laptop.Needs a maintainer stamp on that quit-vs-disconnect contract before implementation.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 9, 2026
Restarting my T3 Code desktop client breaks running Codex threads on my remote machine connected through T3's managed SSH connection. When I reopen the client and reconnect, the threads show "Provider session did not survive a server restart. Send a new message to continue." Long-running agent work is interrupted and needs another message to continue. This previously worked for me: closing and reopening the client did not interrupt the remote work.
Observations
Read-only inspection of the remote server logs, persisted orchestration events, provider logs, and service state found the following on September 9, 2026. Times below are UTC:
The remote host had been up seven days. The kernel journal contained no OOM kills or segfaults that day.
The logs establish that server restarts interrupted the turns. They do not independently identify the initiator of each shutdown; the association with client close/reopen comes from my repeated experience and the matching SSH lifecycle code. This report concerns interrupted active turns, not confirmed loss of stored conversation history.
Probable cause
#9843, merged September 4 and included in stable v0.0.39 and v0.0.40, changed the SSH runner to execute the resolved T3 CLI directly instead of an npm/npx wrapper. The generated runner on the affected host contains this change.
That PR fixes an orphaned-process problem: previously the recorded PID could identify the npm wrapper, leaving the actual server alive after cleanup. The existing SSH finalizer now reaches the actual server, so closing the client can terminate the remote server and interrupt active provider turns. Fixing process ownership appears to have exposed a regression in the remote workflow.
At upstream
6c583620ff7ad3235b135af7107c0543467eecfa, SSH scope cleanup still callsstopRemoteServer(). This file is unchanged from v0.0.40. I do not have the prior working installation's exact version/logs, so #9843 is the probable trigger rather than a conclusively established version bisect.Workaround
Installing T3 as an independent background service on the remote host worked:
I interrupted the two active turns before switching, installed the service, reconnected the desktop client to that existing server, and sent new messages to continue both threads. SSH now records the server as
external, which exempts it from managed-server shutdown.Verification after these steps on September 9, 2026 (UTC):
t3code.servicestarts.Systemd reported zero service restarts, and neither affected thread emitted a provider error event between the service switch and the end of the inspection.
Before submitting
Area
apps/desktop (managed SSH lifecycle in
packages/ssh, affecting the remote server and provider sessions)Steps to reproduce
Expected behavior
Remote agent work continues while the client is closed. Reconnecting shows the same running turn and current activity without requiring another prompt. Closing the client should not implicitly stop the remote environment's active work.
Actual behavior
The remote server restarts and active turns are marked failed. The client displays the restart error and requires another message to continue.
Impact
Major degradation or frequent failure. Long-running remote work is repeatedly interrupted when restarting the client.
Version or commit
Desktop and remote server: v0.0.40. Relevant behavior is still present on upstream main at
6c583620ff7ad3235b135af7107c0543467eecfa(checked September 9, 2026).Environment
T3 Code desktop client using built-in SSH to a Linux remote; remote Node v24.16.0; Codex app-server, model
gpt-6-astra. Client OS and exact Codex CLI version were not captured during this inspection.Logs or stack traces
Selected remote
~/.t3/ssh-launch/<target-key>/server.logentries (timestamps converted to UTC):Corresponding persisted
thread.session-setevents (UTC, thread identifiers omitted):Related: #2614 describes the orphaned-process problem addressed by #9843. This report is about the resulting interruption of active remote work when the client exits.