Skip to content

[BUG] V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls #15124

Description

@NK-Works

What happened

A V2 thread whose provider turn succeeded but whose checkpoint capture has not finished yet keeps showing "Agent is working" and keeps offering Stop targets for work that already finished. waiting is the new post-success, pre-checkpoint transient (completed is persisted as waiting until capture flips it to completed). While a run sits in waiting, the shared awareness rule reports running and the pending-background roster treats the run as settled, so the user sees working + stoppable background work for a turn that is already done. If capture stalls or fails and retries, this state persists.

Diagnosis

Grounded in source at main@fed41fa88:

  • apps/server/src/orchestration-v2/RunExecutionService.ts:633-634 persists a successful terminal as waiting:
    persistedStatus = input.terminal.status === "completed" ? "waiting" : input.terminal.status.
  • packages/shared/src/agentAwareness.ts:99-105 maps both running and waiting to phase "running" ("Agent is working").
  • packages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43 includes waiting in SETTLED_FOR_BACKGROUND_WAIT_RUN_STATUSES, and isLatestRunSettledForBackgroundWait (:109-116) plus derivePendingBackgroundWork (:224-235) therefore surface the provider-thread roster and turn-item Stop targets while the latest run is waiting. hasActiveRun only covers preparing/starting/running, so waiting does not suppress the roster.
  • apps/server/src/orchestration-v2/CheckpointCaptureService.ts:90-104 leaves a run in waiting when the capture target is incomplete (missing run/rootNode/scope, scope mismatch, or no provider thread) by returning a CheckpointCaptureExecutionError without changing status. A failed capture that waits to retry therefore leaves the thread in the working + stoppable state above.

V1 had no waiting state, so this classification is introduced by the orchestrator V2 merge (de3439142).

Steps to reproduce

Deterministic logic repro (no provider CLI needed):

  1. backgroundWorkHoldsCompletion-adjacent check: derivePendingBackgroundWork with latestRun.status = "waiting" returns the roster instead of [].
  2. projectThreadAwarenessV2 / phaseFor("waiting") returns "running" ("Agent is working") even though the provider turn already succeeded.
  3. Minimal script replicating the two predicates:
function phaseFor(status) {
  switch (status) {
    case "running":
    case "waiting": return "running";
    case "completed": return "completed";
  }
}
console.log(phaseFor("waiting")); // "running" — UI shows working for a finished turn
node /tmp/opencode/repro-v2.mjs
# waiting run -> running => UI shows 'Agent is working' even though provider turn succeeded, checkpoint pending

Observable product repro (when capture is slow to flip waiting -> completed):

  1. Start a V2 thread, send a message, let the provider turn succeed (run enters waiting before checkpoint capture).
  2. While the run is waiting, observe the thread row/composer: "Agent is working" plus Stop/background targets for already-finished work.

Version

main@fed41fa88bb27cb4325cb208d571393850bc63c2 (post orchestrator V2 merge de3439142); checked 2026-10-03.

Environment

Linux x64 (Debian 6.12), Node v20.19.2, repo checkout at main@fed41fa88. No provider CLI needed for the logic repro; product symptom applies to any provider with checkpoint capture enabled.

Evidence

Source refs (all at main@fed41fa88):

  • apps/server/src/orchestration-v2/RunExecutionService.ts:633-634
  • packages/shared/src/agentAwareness.ts:99-110 (waiting -> "running")
  • packages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43,109-116,224-235
  • apps/server/src/orchestration-v2/CheckpointCaptureService.ts:90-104

Focused checks run:

vp test run packages/shared/src/agentAwareness.test.ts packages/shared/src/orchestrationV2PendingBackgroundWork.test.ts
# Test Files 2 passed, Tests 41 passed

vp test run apps/server/src/orchestration-v2/Orchestrator.migration.test.ts apps/server/src/orchestration-v2/ProviderTurnStartService.test.ts
# Test Files 2 passed, Tests 19 passed

Repro output:

repro waiting: waiting run -> running => UI shows 'Agent is working' even though provider turn succeeded, checkpoint pending

No live trace available (logic-level defect, no crash); willing to add a unit test asserting waiting renders distinctly (e.g. "Finishing…") or times out back to completed/failed.

Related issues

Fix applied or workaround

None. Possible directions (not tested): render waiting distinctly instead of running, or exclude waiting from the Stop-target roster, or time waiting back to completed/failed if capture does not land.

Filed by

opencode (muse-spark-1.3-contributor-free) via triage session

Activity

  1. NK-Works commented on Oct 3, 2026

    @NK-Works
    Author

    Evidence addendum (logic repro, no provider CLI needed):

    phaseFor(waiting) returns running per packages/shared/src/agentAwareness.ts:99-105, and derivePendingBackgroundWork treats waiting as settled per packages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43,109-116. Successful turns are persisted as waiting at apps/server/src/orchestration-v2/RunExecutionService.ts:633-634 until CheckpointCaptureService.ts:90-104 flips them. Repro script output:

    repro waiting: waiting run -> running => UI shows 'Agent is working' even though provider turn succeeded, checkpoint pending
    

    Visual diagram prepared locally at /tmp/opencode/shot.svg and /tmp/opencode/v2-waiting-evidence.html for the reporter to screenshot and drag into this issue (no desktop browser connected in this session, so no PNG captured here).

  2. changed the title [-]V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls[/-] [+][BUG] V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls[/+] on Oct 3, 2026
  3. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful source walk-through, @NK-Works! What the logic repro shows is the designed post-success state, not a misclassification.

    How waiting is supposed to behave

    A successful provider turn is saved as waiting until checkpoint capture commits completed (persistedStatus in RunExecutionService). While a run is in that state:

    • Awareness and the composer treat waiting as still working, and Stop stays available. projectThreadAwarenessV2 maps waiting to the running phase ("Agent is working"), and derivePhase and deriveCanInterruptRunningThread do the same. apps/web/src/session-logic.test.ts checks this.
    • waiting is in SETTLED_FOR_BACKGROUND_WAIT_RUN_STATUSES on purpose, so background tasks that are still running can show up before the checkpoint lands. derivePendingBackgroundWork only includes items that are still active (pending, running, or waiting), not finished work. The test "returns pending work when the latest run is waiting (post-success, pre-checkpoint)" covers this.
    • If that list includes work that will wake the agent, like a subagent or a monitor, the thread is shown as idle rather than stuck on the checkpoint waiting state. The background stop strip appears, and the thread doesn't read as "Agent is working" (presentThreadShell, see the test "parks runtime idle over stale checkpoint waiting when the roster is nonempty").

    So the tests you ran pass because they lock this behavior in. A short "working" window while capture is in flight is expected. active_run_id also excludes waiting, so this isn't an interruptible provider turn.

    What would be a real bug

    This report doesn't show a run that stays stuck after capture gives up. checkpoint.capture retries, and after the worker's maximum attempts (5 by default) the outbox row fails without moving the run off waiting.

    Because waiting still blocks the thread, it would keep looking busy, and queued turns wouldn't start. Stop would target that run, but dispatchRunInterrupt rejects it as "not interruptible" when there's no running provider turn and no active background work. That failure to finish the run is a separate problem from showing the short in-flight state as "working".

    If you've seen a thread stay on waiting after capture stopped retrying, please comment with:

    • The run status
    • The checkpoint.capture effect row (status, attempt count, and error)
    • Roughly how long it stayed that way

    A maintainer will decide on the fix direction once that's clear.

  4. NK-Works commented on Oct 3, 2026

    @NK-Works
    Author

    Thanks for the detailed triage — verified your refs against main@fed41fa88 and you're right that the short in-flight window is by design. Withdrawing the broad "waiting should not show working" claim:

    • apps/web/src/session-logic.ts:1013 maps waiting to running, locked by session-logic.test.ts:103-106.
    • packages/client-runtime/src/state/models.ts:168-171 parks runtime at idle when the background roster is nonempty, locked by entities.test.ts:302 ("parks runtime idle over stale checkpoint waiting"). So with background work present the thread does not read as "Agent is working" — my report overstated that case.
    • EffectWorker.ts:508 defaults maxAttempts to 5 for the checkpoint.capture outbox row.

    Narrowing to the stuck case you describe, which I have not reproduced live in this session: a run remaining in waiting after checkpoint.capture exhausts retries, where the outbox row fails without moving the run off waiting, the thread keeps looking busy, queued turns do not start, and Stop targets the run but dispatchRunInterrupt rejects it as not interruptible (Orchestrator.ts:7779 path).

    I don't have a stuck run here to pull the three items from (run status, checkpoint.capture effect row status/attempts/error, duration). If it helps, I can gather them from a reporter machine with read-only queries against statev2.sqlite (runs + effect outbox rows for checkpoint.capture:<runId>) plus the surrounding server.trace.ndjson — or, if you prefer, I can re-title this issue to that stuck-after-retries shape and drop the in-flight presentation part. Let me know which evidence format you want and I'll supply it.

  5. NK-Works commented on Oct 3, 2026

    @NK-Works
    Author

    Evidence gathered while you review (code-construction + local snapshot, no live stuck run here):

    1. Stuck mechanism exists by construction (verified on main@fed41fa88)

    • Failure path emits no run transition: CheckpointCaptureService.ts:90-104 returns CheckpointCaptureExecutionError when the capture target is incomplete (missing run/rootNode/scope, scope mismatch, no provider thread). Only the success path (:136-174) commits run.updated (waiting -> completed) plus checkpoint.captured. The existing test exercises this: CheckpointCaptureService.test.ts:316-321 executes with a missing scope and asserts the error is returned before attempting the real capture.
    • Retry exhaustion also leaves the run untouched: EffectWorker.ts:688-691 calls outbox.fail once attemptCount >= maxAttempts (default 5, :508) without any run-status write. So a persistently failing checkpoint.capture:<runId> effect ends as failed in the outbox while the run row stays waiting — the exact stuck shape you describe (busy-looking thread, queued turns blocked, Stop targeting a non-interruptible run per Orchestrator.ts:7779).

    2. Local snapshot has no stuck rows (queries included for reporters)

    Read-only copy of this machine's statev2.sqlite (copied to /tmp, never opened read-write):

    runs by status: cancelled 3, completed 35, failed 5, running 2, waiting 0
    checkpoint.capture effects: 35 succeeded (max attempt_count 1), 0 failed
    

    Queries a reporter can rerun against their statev2.sqlite (read-only open recommended):

    SELECT status, COUNT(*) FROM orchestration_v2_projection_runs GROUP BY status;
    SELECT run_id, thread_id, ordinal, status
    FROM orchestration_v2_projection_runs WHERE status = 'waiting';
    SELECT status, COUNT(*), MAX(attempt_count)
    FROM orchestration_v2_effect_outbox
    WHERE effect_type = 'checkpoint.capture' GROUP BY status;
    SELECT effect_id, thread_id, status, attempt_count, substr(last_error,1,300)
    FROM orchestration_v2_effect_outbox
    WHERE effect_type = 'checkpoint.capture' AND status = 'failed';

    Plus the surrounding server.trace.ndjson window for the stuck runId (the checkpoint.capture attempts and the final fail).

    3. What I still lack

    A live thread actually stuck in waiting with its failed checkpoint.capture row (effect status, attempt count, last_error) and how long it stayed that way. If you can point me at a repro that forces capture to fail persistently (e.g. a scope-mismatch fixture or a failing CheckpointStore.capture), I'll drive it here and attach the run + outbox rows. Otherwise, happy for this issue to be re-titled to the stuck-after-retries shape.

  6. NK-Works commented on Oct 3, 2026

    @NK-Works
    Author

    Deterministic repro now in hand (temporary test, since removed — not committed):

    Temp test: failed capture leaves the run in waiting with zero events committed

    Setup mirrors CheckpointCaptureService.test.ts: thread + run in waiting applied to an in-memory ProjectionStore, then CheckpointCaptureService.execute against a missing scope (incomplete target, same class of failure as a persistently failing capture):

    • execute returns CheckpointCaptureExecutionError (Failure exit).
    • The mocked EventSink.commitCommand is never called (0 commits — no run.updated, no checkpoint.captured).
    • getThreadProjection still shows the run as waiting.

    Result under Node 24 (repo tests need node:sqlite):

    RUN  v5.0.1
    Test Files  1 passed (1)
    Tests  1 passed (1)
    

    Combined with EffectWorker.ts:688-691 (outbox.fail after maxAttempts, default 5 at :508, with no run write), a checkpoint.capture effect that keeps failing ends as failed in orchestration_v2_effect_outbox while the run row stays waiting indefinitely — thread looks busy, queued turns stay blocked, Stop targets a run dispatchRunInterrupt rejects (Orchestrator.ts:7779).

    Live snapshot (this machine, read-only copy): no stuck rows currently — runs: 3 cancelled / 35 completed / 5 failed / 2 running / 0 waiting; checkpoint.capture effects: 35 succeeded, 0 failed. Reporter queries from my previous comment still apply for any machine that does hit it (run status + effect row status/attempts/last_error + duration).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs more infoInitial triage showed no bug. Awaiting more infovia-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions