Before submitting
Area
apps/server
Steps to reproduce
Uses the same HTTP orchestration API the web client uses (POST /api/orchestration/dispatch), so this can be hit by any client that archives right after stopping (or tears down several threads at once).
Deterministic (archive before stop):
- Start a thread on any provider (seen with Claude Code and Codex) and let the first turn finish, so the session is
ready and a claude / codex app-server process is alive.
- Dispatch
{"type":"thread.archive","threadId":…}, then immediately {"type":"thread.session.stop","threadId":…}.
- Wait 30s and check the provider process.
Realistic (bulk teardown):
- Start 4 threads (any providers) and let them reach
ready.
- For each thread in turn, without waiting in between, dispatch
thread.turn.interrupt, then thread.session.stop, then thread.archive.
- Wait 30s and check the provider processes.
Expected behavior
thread.session.stop stops the provider session and its process whether or not the thread is archived, or the archive stops the session itself. An archived thread should never keep a live provider process.
Actual behavior
- Deterministic case: the provider process keeps running indefinitely (1/1).
- Bulk case: only the first thread's session is stopped. The other 3 of 4 provider processes keep running with no visible thread in the UI.
In load testing, a 10-thread teardown leaked 9 of 10 provider processes (Claude Code + Codex, about 2.4 GB RSS on a 6.6 GB machine), which stayed until I killed them by hand. In a 24-thread teardown, only 1 of 24 stop commands actually stopped a session.
The server trace shows the stop commands are received but not acted on. The first stopSessionInternal takes ~494 ms; the remaining 23 processSessionStopRequested spans are then drained in the same millisecond (~0.2 ms each) with no corresponding stopSessionInternal. By then those threads' thread.archive commands had already been dispatched. It looks like the stop reactor skips (or can't resolve) the session once the thread is archived, so any stop that is still queued behind a slow stop is silently dropped.
Control: stop, wait ~3s, then archive → process exits every time.
Impact
Major degradation or frequent failure
Version or commit
t3 v0.0.42
Environment
Ubuntu, Linux 7.0.0, headless t3 serve (background service). Providers: Claude Code 2.1.283 (claudeAgent), codex-cli 0.157.1 (codex app-server).
Logs or stack traces
# server.trace.ndjson, bulk teardown of 24 threads (interrupt → stop → archive per thread)
1790647304.85 stopSessionInternal 493.6ms # thread 1: actually stopped
1790647305.34 processSessionStopRequested 0.5ms # threads 2..24: all drained here,
1790647305.34 processSessionStopRequested 0.2ms # no stopSessionInternal follows
... (22 more, all ~0.2ms)
1790647310.71 stopSessionInternal 0.1ms # these only appear after I SIGTERM'd
... (10 more) # the leaked processes by hand
# counts in the window: 24 thread.session.stop commands, 24 processSessionStopRequested, 1 real stopSessionInternal
Workaround
Wait for the session to reach stopped before dispatching thread.archive, and tear threads down one at a time. Leaked processes have to be killed manually (they're children of the t3 serve process, with cwd = the thread's worktree/project).
Before submitting
Area
apps/server
Steps to reproduce
Uses the same HTTP orchestration API the web client uses (
POST /api/orchestration/dispatch), so this can be hit by any client that archives right after stopping (or tears down several threads at once).Deterministic (archive before stop):
readyand aclaude/codex app-serverprocess is alive.{"type":"thread.archive","threadId":…}, then immediately{"type":"thread.session.stop","threadId":…}.Realistic (bulk teardown):
ready.thread.turn.interrupt, thenthread.session.stop, thenthread.archive.Expected behavior
thread.session.stopstops the provider session and its process whether or not the thread is archived, or the archive stops the session itself. An archived thread should never keep a live provider process.Actual behavior
In load testing, a 10-thread teardown leaked 9 of 10 provider processes (Claude Code + Codex, about 2.4 GB RSS on a 6.6 GB machine), which stayed until I killed them by hand. In a 24-thread teardown, only 1 of 24 stop commands actually stopped a session.
The server trace shows the stop commands are received but not acted on. The first
stopSessionInternaltakes ~494 ms; the remaining 23processSessionStopRequestedspans are then drained in the same millisecond (~0.2 ms each) with no correspondingstopSessionInternal. By then those threads'thread.archivecommands had already been dispatched. It looks like the stop reactor skips (or can't resolve) the session once the thread is archived, so any stop that is still queued behind a slow stop is silently dropped.Control: stop, wait ~3s, then archive → process exits every time.
Impact
Major degradation or frequent failure
Version or commit
t3 v0.0.42
Environment
Ubuntu, Linux 7.0.0, headless
t3 serve(background service). Providers: Claude Code 2.1.283 (claudeAgent), codex-cli 0.157.1 (codex app-server).Logs or stack traces
Workaround
Wait for the session to reach
stoppedbefore dispatchingthread.archive, and tear threads down one at a time. Leaked processes have to be killed manually (they're children of thet3 serveprocess, with cwd = the thread's worktree/project).