Repository navigation
[BUG] V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls #15124
Description
Activity
Evidence addendum (logic repro, no provider CLI needed):
phaseFor(waiting)returnsrunningperpackages/shared/src/agentAwareness.ts:99-105, andderivePendingBackgroundWorktreatswaitingas settled perpackages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43,109-116. Successful turns are persisted aswaitingatapps/server/src/orchestration-v2/RunExecutionService.ts:633-634untilCheckpointCaptureService.ts:90-104flips them. Repro script output:repro waiting: waiting run -> running => UI shows 'Agent is working' even though provider turn succeeded, checkpoint pendingVisual diagram prepared locally at
/tmp/opencode/shot.svgand/tmp/opencode/v2-waiting-evidence.htmlfor the reporter to screenshot and drag into this issue (no desktop browser connected in this session, so no PNG captured here).- changed the title
[-]V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls[/-][+][BUG] V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls[/+]on Oct 3, 2026 Note
Grok responding on behalf of Julius.
Triage
Thanks for the careful source walk-through, @NK-Works! What the logic repro shows is the designed post-success state, not a misclassification.
How
waitingis supposed to behaveA successful provider turn is saved as
waitinguntil checkpoint capture commitscompleted(persistedStatusinRunExecutionService). While a run is in that state:- Awareness and the composer treat
waitingas still working, and Stop stays available.projectThreadAwarenessV2mapswaitingto therunningphase ("Agent is working"), andderivePhaseandderiveCanInterruptRunningThreaddo the same.apps/web/src/session-logic.test.tschecks this. waitingis inSETTLED_FOR_BACKGROUND_WAIT_RUN_STATUSESon purpose, so background tasks that are still running can show up before the checkpoint lands.derivePendingBackgroundWorkonly includes items that are still active (pending,running, orwaiting), not finished work. The test "returns pending work when the latest run is waiting (post-success, pre-checkpoint)" covers this.- If that list includes work that will wake the agent, like a subagent or a monitor, the thread is shown as idle rather than stuck on the checkpoint
waitingstate. The background stop strip appears, and the thread doesn't read as "Agent is working" (presentThreadShell, see the test "parks runtime idle over stale checkpoint waiting when the roster is nonempty").
So the tests you ran pass because they lock this behavior in. A short "working" window while capture is in flight is expected.
active_run_idalso excludeswaiting, so this isn't an interruptible provider turn.What would be a real bug
This report doesn't show a run that stays stuck after capture gives up.
checkpoint.captureretries, and after the worker's maximum attempts (5 by default) the outbox row fails without moving the run offwaiting.Because
waitingstill blocks the thread, it would keep looking busy, and queued turns wouldn't start. Stop would target that run, butdispatchRunInterruptrejects it as "not interruptible" when there's no running provider turn and no active background work. That failure to finish the run is a separate problem from showing the short in-flight state as "working".If you've seen a thread stay on
waitingafter capture stopped retrying, please comment with:- The run status
- The
checkpoint.captureeffect row (status, attempt count, and error) - Roughly how long it stayed that way
A maintainer will decide on the fix direction once that's clear.
- Awareness and the composer treat
- addedneeds more infoInitial triage showed no bug. Awaiting more infoInitial triage showed no bug. Awaiting more infovia-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 3, 2026 Thanks for the detailed triage — verified your refs against
main@fed41fa88and you're right that the short in-flight window is by design. Withdrawing the broad "waiting should not show working" claim:apps/web/src/session-logic.ts:1013mapswaitingtorunning, locked bysession-logic.test.ts:103-106.packages/client-runtime/src/state/models.ts:168-171parks runtime atidlewhen the background roster is nonempty, locked byentities.test.ts:302("parks runtime idle over stale checkpoint waiting"). So with background work present the thread does not read as "Agent is working" — my report overstated that case.EffectWorker.ts:508defaultsmaxAttemptsto 5 for thecheckpoint.captureoutbox row.
Narrowing to the stuck case you describe, which I have not reproduced live in this session: a run remaining in
waitingaftercheckpoint.captureexhausts retries, where the outbox row fails without moving the run offwaiting, the thread keeps looking busy, queued turns do not start, and Stop targets the run butdispatchRunInterruptrejects it as not interruptible (Orchestrator.ts:7779path).I don't have a stuck run here to pull the three items from (run status,
checkpoint.captureeffect row status/attempts/error, duration). If it helps, I can gather them from a reporter machine with read-only queries againststatev2.sqlite(runs + effect outbox rows forcheckpoint.capture:<runId>) plus the surroundingserver.trace.ndjson— or, if you prefer, I can re-title this issue to that stuck-after-retries shape and drop the in-flight presentation part. Let me know which evidence format you want and I'll supply it.Evidence gathered while you review (code-construction + local snapshot, no live stuck run here):
1. Stuck mechanism exists by construction (verified on
main@fed41fa88)- Failure path emits no run transition:
CheckpointCaptureService.ts:90-104returnsCheckpointCaptureExecutionErrorwhen the capture target is incomplete (missing run/rootNode/scope, scope mismatch, no provider thread). Only the success path (:136-174) commitsrun.updated(waiting->completed) pluscheckpoint.captured. The existing test exercises this:CheckpointCaptureService.test.ts:316-321executes with a missing scope and asserts the error is returned before attempting the real capture. - Retry exhaustion also leaves the run untouched:
EffectWorker.ts:688-691callsoutbox.failonceattemptCount >= maxAttempts(default 5,:508) without any run-status write. So a persistently failingcheckpoint.capture:<runId>effect ends asfailedin the outbox while the run row stayswaiting— the exact stuck shape you describe (busy-looking thread, queued turns blocked, Stop targeting a non-interruptible run perOrchestrator.ts:7779).
2. Local snapshot has no stuck rows (queries included for reporters)
Read-only copy of this machine's
statev2.sqlite(copied to/tmp, never opened read-write):runs by status: cancelled 3, completed 35, failed 5, running 2, waiting 0 checkpoint.capture effects: 35 succeeded (max attempt_count 1), 0 failedQueries a reporter can rerun against their
statev2.sqlite(read-only open recommended):SELECT status, COUNT(*) FROM orchestration_v2_projection_runs GROUP BY status; SELECT run_id, thread_id, ordinal, status FROM orchestration_v2_projection_runs WHERE status = 'waiting'; SELECT status, COUNT(*), MAX(attempt_count) FROM orchestration_v2_effect_outbox WHERE effect_type = 'checkpoint.capture' GROUP BY status; SELECT effect_id, thread_id, status, attempt_count, substr(last_error,1,300) FROM orchestration_v2_effect_outbox WHERE effect_type = 'checkpoint.capture' AND status = 'failed';
Plus the surrounding
server.trace.ndjsonwindow for the stuckrunId(thecheckpoint.captureattempts and the finalfail).3. What I still lack
A live thread actually stuck in
waitingwith its failedcheckpoint.capturerow (effect status, attempt count,last_error) and how long it stayed that way. If you can point me at a repro that forces capture to fail persistently (e.g. a scope-mismatch fixture or a failingCheckpointStore.capture), I'll drive it here and attach the run + outbox rows. Otherwise, happy for this issue to be re-titled to the stuck-after-retries shape.- Failure path emits no run transition:
Deterministic repro now in hand (temporary test, since removed — not committed):
Temp test: failed capture leaves the run in
waitingwith zero events committedSetup mirrors
CheckpointCaptureService.test.ts: thread + run inwaitingapplied to an in-memoryProjectionStore, thenCheckpointCaptureService.executeagainst a missing scope (incomplete target, same class of failure as a persistently failing capture):executereturnsCheckpointCaptureExecutionError(Failure exit).- The mocked
EventSink.commitCommandis never called (0 commits — norun.updated, nocheckpoint.captured). getThreadProjectionstill shows the run aswaiting.
Result under Node 24 (repo tests need
node:sqlite):RUN v5.0.1 Test Files 1 passed (1) Tests 1 passed (1)Combined with
EffectWorker.ts:688-691(outbox.failaftermaxAttempts, default 5 at:508, with no run write), acheckpoint.captureeffect that keeps failing ends asfailedinorchestration_v2_effect_outboxwhile the run row stayswaitingindefinitely — thread looks busy, queued turns stay blocked, Stop targets a rundispatchRunInterruptrejects (Orchestrator.ts:7779).Live snapshot (this machine, read-only copy): no stuck rows currently — runs: 3 cancelled / 35 completed / 5 failed / 2 running / 0 waiting;
checkpoint.captureeffects: 35 succeeded, 0 failed. Reporter queries from my previous comment still apply for any machine that does hit it (run status + effect row status/attempts/last_error+ duration).
What happened
A V2 thread whose provider turn succeeded but whose checkpoint capture has not finished yet keeps showing "Agent is working" and keeps offering Stop targets for work that already finished.
waitingis the new post-success, pre-checkpoint transient (completedis persisted aswaitinguntil capture flips it tocompleted). While a run sits inwaiting, the shared awareness rule reportsrunningand the pending-background roster treats the run as settled, so the user sees working + stoppable background work for a turn that is already done. If capture stalls or fails and retries, this state persists.Diagnosis
Grounded in source at
main@fed41fa88:apps/server/src/orchestration-v2/RunExecutionService.ts:633-634persists a successful terminal aswaiting:persistedStatus = input.terminal.status === "completed" ? "waiting" : input.terminal.status.packages/shared/src/agentAwareness.ts:99-105maps bothrunningandwaitingto phase"running"("Agent is working").packages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43includeswaitinginSETTLED_FOR_BACKGROUND_WAIT_RUN_STATUSES, andisLatestRunSettledForBackgroundWait(:109-116) plusderivePendingBackgroundWork(:224-235) therefore surface the provider-thread roster and turn-item Stop targets while the latest run iswaiting.hasActiveRunonly coverspreparing/starting/running, sowaitingdoes not suppress the roster.apps/server/src/orchestration-v2/CheckpointCaptureService.ts:90-104leaves a run inwaitingwhen the capture target is incomplete (missing run/rootNode/scope, scope mismatch, or no provider thread) by returning aCheckpointCaptureExecutionErrorwithout changing status. A failed capture that waits to retry therefore leaves the thread in the working + stoppable state above.V1 had no
waitingstate, so this classification is introduced by the orchestrator V2 merge (de3439142).Steps to reproduce
Deterministic logic repro (no provider CLI needed):
backgroundWorkHoldsCompletion-adjacent check:derivePendingBackgroundWorkwithlatestRun.status = "waiting"returns the roster instead of[].projectThreadAwarenessV2/phaseFor("waiting")returns"running"("Agent is working") even though the provider turn already succeeded.node /tmp/opencode/repro-v2.mjs # waiting run -> running => UI shows 'Agent is working' even though provider turn succeeded, checkpoint pendingObservable product repro (when capture is slow to flip
waiting->completed):waitingbefore checkpoint capture).waiting, observe the thread row/composer: "Agent is working" plus Stop/background targets for already-finished work.Version
main@fed41fa88bb27cb4325cb208d571393850bc63c2(post orchestrator V2 mergede3439142); checked 2026-10-03.Environment
Linux x64 (Debian 6.12), Node v20.19.2, repo checkout at
main@fed41fa88. No provider CLI needed for the logic repro; product symptom applies to any provider with checkpoint capture enabled.Evidence
Source refs (all at
main@fed41fa88):apps/server/src/orchestration-v2/RunExecutionService.ts:633-634packages/shared/src/agentAwareness.ts:99-110(waiting->"running")packages/shared/src/orchestrationV2PendingBackgroundWork.ts:37-43,109-116,224-235apps/server/src/orchestration-v2/CheckpointCaptureService.ts:90-104Focused checks run:
Repro output:
No live trace available (logic-level defect, no crash); willing to add a unit test asserting
waitingrenders distinctly (e.g. "Finishing…") or times out back tocompleted/failed.Related issues
commandbackground work) — same files, opposite edge (completed vs waiting). Not a duplicate.waiting. Not a duplicate.waitingawareness item.derivePendingBackgroundWork/activityRunStatus: no match.Fix applied or workaround
None. Possible directions (not tested): render
waitingdistinctly instead ofrunning, or excludewaitingfrom the Stop-target roster, or timewaitingback tocompleted/failedif capture does not land.Filed by
opencode (muse-spark-1.3-contributor-free) via triage session