Skip to content

[Bug]: Idle session reaper kills in-flight background agent work (dynamic workflows / subagents) #4198

Description

@ChamaruAmasara

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Start a long-running background agent task in a thread — a dynamic workflow or a task that fans out subagents that keep working after the foreground turn settles.
  2. Let the foreground turn complete so the thread has no active user turn, while the background work keeps running.
  3. Leave the thread otherwise idle (no new sendTurn) for longer than the reaper's inactivity threshold (default 30 min; swept every 5 min).
  4. Observe that ProviderSessionReaper calls stopSession for the thread and tears down the provider session while the background agents are still in flight.

Expected behavior

A session with live background agent work should not be reaped for inactivity. The reaper already skips sessions with an active foreground turn; the same protection should extend to in-flight background work.

Actual behavior

ProviderSessionReaper (apps/server/src/provider/Layers/ProviderSessionReaper.ts) stops any provider session after 30 min of last_seen_at inactivity. Its only busy-guard is the projection snapshot's activeTurnId. When the foreground turn completes, the adapter session goes ready with no active turn, and nothing refreshes last_seen_at from ongoing background activity. So a long-running background workflow with no foreground turn is torn down mid-run. The next interaction resumes the conversation from a transcript cursor as a fresh process, but the in-flight background orchestration (subagent work, queued steps) is gone, leaving the run half-completed.

Impact

Major degradation or frequent failure

Version or commit

main at 5d34f9ff2

Environment

T3 Code server, running background dynamic workflows / subagents.

Logs or stack traces

No exception is thrown. The reaper logs provider.session.reaped with reason: "inactivity_threshold" for a thread whose adapter session is still running background work.

Screenshots, recordings, or supporting files

A deterministic regression test is included in the linked fix.

Workaround

Send a message in the thread periodically (every <30 min) to refresh last_seen_at and keep the session alive while the background work runs.

Suggested fix

Refresh last_seen_at from runtime activity: bump the binding's timestamp when the adapter emits runtime events (e.g. task.progress from background subagents), throttled per thread so an event burst is at most one lightweight write per window. The adapter keeps emitting runtime events during background work even with no active foreground turn, so this keeps the session warm exactly when the reaper would otherwise reap it.

Activity

  1. Mooned8 commented on Jul 22, 2026

    @Mooned8

    Same root cause here, different entry point: a background shell command (deploy watch) instead of subagents. The agent ends its turn saying it will report back, the reaper tears the session down (provider.session.reaped, reason: 'inactivity_threshold'), and on the next user message the session resumes with "No completion record was found for this background shell command from the previous session" — the promised follow-up is silently lost. Filed #4265 before finding this issue; closing it as a duplicate. One addition to the proposal: besides exempting sessions with in-flight background work, delivering the task result by respawning the session on completion would also repair the contract when a reap does happen. Observed on 0.0.28, self-hosted Ubuntu 24.04.

  2. NoorChasib commented on Jul 27, 2026

    @NoorChasib

    Hitting this too on 0.0.28 (Ubuntu 26.04, Bun 1.3.14, Claude provider). Adding my logs as another data point.

    Thread e79655a5, all timestamps from the provider log:

    Time (UTC) Event
    05:49:37.867 turn.started — 5cb4c19d
    06:08:37.190 turn.started — 3cea60d7
    06:17:21.017 turn.started — ccdd04f9
    06:19:19.081 claude/system/task_started (background task 1)
    06:19:20.177 claude/system/task_started (background task 2)
    06:19:32.442 turn.completed for ccdd04f9 → activeTurnId cleared
    06:20:20.443 session.exited — reason: "inactivity_threshold", idleDurationMs: 1837486

    The two background tasks were killed 61 and 60 seconds after starting. I didn't notice until the session resumed 40 minutes later and reported them as orphaned, which made it look like a much longer hang than it was.

    One detail that might be useful for scoping: back-computing from the reap, 1837486 ms before 06:20:20.443 is 05:49:42.957. That matches the first of those three turns — so last_seen_at sat unchanged while two later foreground turns ran and completed. In my case the stale timestamp wasn't only about background work; ordinary provider-started turns didn't refresh it either.

    For what it's worth, #4199 looks like it would have caught my specific incident, since task_started fired ~60s before the sweep landed.

  3. Zeus-Deus commented on Jul 30, 2026

    @Zeus-Deus

    Confirmed independently on t3code 0.0.31 / current upstream main with two separate Claude background workflows.

    Sanitized incident evidence

    In both incidents:

    • the foreground turn ended normally with a successful provider result and stop_reason: end_turn;
    • exactly one background task.started remained without a matching terminal task.completed event;
    • the T3 server stayed healthy throughout (no restart, crash, OOM, or provider/API error);
    • the reaper then emitted provider.session.reaped with reason: "inactivity_threshold", followed by a graceful session.exited;
    • the workflow resumed only after a new user message started/recovered the provider session.

    The key timing discrepancy reproduced twice:

    Incident Reaper-reported idle duration Time since the latest foreground turn completed
    A 31m 18s 1m 27s
    B 36m 48s 1m 41s

    Back-calculating lastSeenAt from the reaper log placed it near initial session startup/early activity, despite several later foreground and automatically generated turns. This confirms the stale liveness timestamp is broader than just quiet background work: subsequent provider activity was not consistently refreshing the binding.

    Clean-fix criteria

    #4199 addresses an important part of this by touching last_seen_at from runtime events, with throttling and without resurrecting stopped rows. I think the robust lifecycle fix should also include a hard busy guard:

    1. Refresh liveness for provider runtime activity, including automatically generated follow-up turns and background task lifecycle/progress events.
    2. Do not reap while the thread has known outstanding background tasks, even if a task is temporarily quiet and emits no progress event for longer than the inactivity threshold.
    3. Preserve the existing activeTurnId guard and ensure stopped sessions cannot be revived by late events.
    4. Add a fake-clock regression test: stale binding + no active foreground turn + outstanding task must not call stopSession; after the task reaches a terminal state and the session remains idle past the threshold, it may be reaped.
    5. As defense in depth, if a reap still races with task completion, preserve or deliver the completion rather than silently orphaning it.

    This report contains only synthetic/sanitized details; no private thread IDs, paths, hostnames, usernames, or credentials are included.

  4. diziebol commented on Aug 3, 2026

    @diziebol

    This is still happening on 0.0.32-nightly.20260803.986

  5. kraptor23 commented on Aug 6, 2026

    @kraptor23

    We hit this on 0.0.32 nightlies (a session reaped 6 times in ~12 h, three reaps landing 105 s–5 min after a turn had completed) and traced the same two roots documented here. Two PRs are up:

    Remaining, not covered by either PR: @Zeus-Deus's criterion 2 (a hard busy-guard for tasks that emit nothing for longer than the threshold). The thread shell already carries backgroundLiveness fed from task lifecycle events and cleared on session.exited, so a small follow-up can make the reaper skip sessions with live background state, plus a hard max-idle cap so a wedged liveness entry cannot pin a session forever. Planned next, unless a maintainer would rather see it folded elsewhere. Criterion 5 (re-delivering results after a reap, #4265) is a larger lifecycle change and intentionally deferred.

    Written by Claude Fable 5 via Claude Code.

  6. t3-code commented on Aug 10, 2026

    @t3-code
    Contributor

    Closing as fixed by #5677: the reaper no longer kills active background work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions