Repository navigation
[Bug]: Idle session reaper kills in-flight background agent work (dynamic workflows / subagents) #4198
Description
Activity
Same root cause here, different entry point: a background shell command (deploy watch) instead of subagents. The agent ends its turn saying it will report back, the reaper tears the session down (
provider.session.reaped, reason: 'inactivity_threshold'), and on the next user message the session resumes with "No completion record was found for this background shell command from the previous session" — the promised follow-up is silently lost. Filed #4265 before finding this issue; closing it as a duplicate. One addition to the proposal: besides exempting sessions with in-flight background work, delivering the task result by respawning the session on completion would also repair the contract when a reap does happen. Observed on 0.0.28, self-hosted Ubuntu 24.04.Hitting this too on 0.0.28 (Ubuntu 26.04, Bun 1.3.14, Claude provider). Adding my logs as another data point.
Thread
e79655a5, all timestamps from the provider log:Time (UTC) Event 05:49:37.867 turn.started—5cb4c19d06:08:37.190 turn.started—3cea60d706:17:21.017 turn.started—ccdd04f906:19:19.081 claude/system/task_started(background task 1)06:19:20.177 claude/system/task_started(background task 2)06:19:32.442 turn.completedforccdd04f9→activeTurnIdcleared06:20:20.443 session.exited—reason: "inactivity_threshold",idleDurationMs: 1837486The two background tasks were killed 61 and 60 seconds after starting. I didn't notice until the session resumed 40 minutes later and reported them as orphaned, which made it look like a much longer hang than it was.
One detail that might be useful for scoping: back-computing from the reap,
1837486 msbefore06:20:20.443is05:49:42.957. That matches the first of those three turns — solast_seen_atsat unchanged while two later foreground turns ran and completed. In my case the stale timestamp wasn't only about background work; ordinary provider-started turns didn't refresh it either.For what it's worth, #4199 looks like it would have caught my specific incident, since
task_startedfired ~60s before the sweep landed.Confirmed independently on t3code 0.0.31 / current upstream
mainwith two separate Claude background workflows.Sanitized incident evidence
In both incidents:
- the foreground turn ended normally with a successful provider result and
stop_reason: end_turn; - exactly one background
task.startedremained without a matching terminaltask.completedevent; - the T3 server stayed healthy throughout (no restart, crash, OOM, or provider/API error);
- the reaper then emitted
provider.session.reapedwithreason: "inactivity_threshold", followed by a gracefulsession.exited; - the workflow resumed only after a new user message started/recovered the provider session.
The key timing discrepancy reproduced twice:
Incident Reaper-reported idle duration Time since the latest foreground turn completed A 31m 18s 1m 27s B 36m 48s 1m 41s Back-calculating
lastSeenAtfrom the reaper log placed it near initial session startup/early activity, despite several later foreground and automatically generated turns. This confirms the stale liveness timestamp is broader than just quiet background work: subsequent provider activity was not consistently refreshing the binding.Clean-fix criteria
#4199 addresses an important part of this by touching
last_seen_atfrom runtime events, with throttling and without resurrecting stopped rows. I think the robust lifecycle fix should also include a hard busy guard:- Refresh liveness for provider runtime activity, including automatically generated follow-up turns and background task lifecycle/progress events.
- Do not reap while the thread has known outstanding background tasks, even if a task is temporarily quiet and emits no progress event for longer than the inactivity threshold.
- Preserve the existing
activeTurnIdguard and ensure stopped sessions cannot be revived by late events. - Add a fake-clock regression test: stale binding + no active foreground turn + outstanding task must not call
stopSession; after the task reaches a terminal state and the session remains idle past the threshold, it may be reaped. - As defense in depth, if a reap still races with task completion, preserve or deliver the completion rather than silently orphaning it.
This report contains only synthetic/sanitized details; no private thread IDs, paths, hostnames, usernames, or credentials are included.
- the foreground turn ended normally with a successful provider result and
This is still happening on 0.0.32-nightly.20260803.986
- added a commit that references this issue
on Aug 6, 2026 We hit this on 0.0.32 nightlies (a session reaped 6 times in ~12 h, three reaps landing 105 s–5 min after a turn had completed) and traced the same two roots documented here. Two PRs are up:
- fix(server): refresh provider session lastSeenAt on turn and task activity #5524 fixes the liveness staleness:
last_seen_atis currently written only bystartSession/sendTurn/stopSession, so provider-delivered turns and background-task events never refresh it. The PR touches the binding (throttled, 60 s per thread) onturn.started/completed/abortedandtask.started/progress/updated/completed. Back-computing theidleDurationMsvalues in this thread against the turn timestamps, every logged incident here (@NoorChasib's stale-across-two-turns table, @Mooned8's deploy-watch, @Zeus-Deus's two workflows) is an event-emitting case that PR covers. - feat(server): make provider session reaper timing configurable #5525 + Provider session reaper inactivity threshold and sweep interval are not configurable #5523 make the threshold and sweep interval server settings (currently hardcoded 30 min/5 min), which is the practical mitigation for long-quiet work in the meantime.
Remaining, not covered by either PR: @Zeus-Deus's criterion 2 (a hard busy-guard for tasks that emit nothing for longer than the threshold). The thread shell already carries
backgroundLivenessfed from task lifecycle events and cleared onsession.exited, so a small follow-up can make the reaper skip sessions with live background state, plus a hard max-idle cap so a wedged liveness entry cannot pin a session forever. Planned next, unless a maintainer would rather see it folded elsewhere. Criterion 5 (re-delivering results after a reap, #4265) is a larger lifecycle change and intentionally deferred.Written by Claude Fable 5 via Claude Code.
- fix(server): refresh provider session lastSeenAt on turn and task activity #5524 fixes the liveness staleness:
Closing as fixed by #5677: the reaper no longer kills active background work.
Before submitting
Area
apps/server
Steps to reproduce
sendTurn) for longer than the reaper's inactivity threshold (default 30 min; swept every 5 min).ProviderSessionReapercallsstopSessionfor the thread and tears down the provider session while the background agents are still in flight.Expected behavior
A session with live background agent work should not be reaped for inactivity. The reaper already skips sessions with an active foreground turn; the same protection should extend to in-flight background work.
Actual behavior
ProviderSessionReaper(apps/server/src/provider/Layers/ProviderSessionReaper.ts) stops any provider session after 30 min oflast_seen_atinactivity. Its only busy-guard is the projection snapshot'sactiveTurnId. When the foreground turn completes, the adapter session goesreadywith no active turn, and nothing refresheslast_seen_atfrom ongoing background activity. So a long-running background workflow with no foreground turn is torn down mid-run. The next interaction resumes the conversation from a transcript cursor as a fresh process, but the in-flight background orchestration (subagent work, queued steps) is gone, leaving the run half-completed.Impact
Major degradation or frequent failure
Version or commit
mainat5d34f9ff2Environment
T3 Code server, running background dynamic workflows / subagents.
Logs or stack traces
No exception is thrown. The reaper logs
provider.session.reapedwithreason: "inactivity_threshold"for a thread whose adapter session is still running background work.Screenshots, recordings, or supporting files
A deterministic regression test is included in the linked fix.
Workaround
Send a message in the thread periodically (every <30 min) to refresh
last_seen_atand keep the session alive while the background work runs.Suggested fix
Refresh
last_seen_atfrom runtime activity: bump the binding's timestamp when the adapter emits runtime events (e.g.task.progressfrom background subagents), throttled per thread so an event burst is at most one lightweight write per window. The adapter keeps emitting runtime events during background work even with no active foreground turn, so this keeps the session warm exactly when the reaper would otherwise reap it.