Repository navigation
[Bug]: An environment's threads silently stop updating until the client is restarted #4589
Description
Activity
- changed the title
[-][Bug]: Client holds stale running session state after server reports ready, blocking settle client-side with no command sent[/-][+][Bug]: Clicking Settle silently does nothing on an idle thread until the client is restarted[/+]on Jul 26, 2026 - changed the title
[-][Bug]: Clicking Settle silently does nothing on an idle thread until the client is restarted[/-][+][Bug]: Thread acts as if it's already settled — Settle silently does nothing until the client is restarted[/+]on Jul 26, 2026 I'm seeing all sorts of issues with the recent nightlies along the same lines (doesn't matter if I have new sidebar on or off).
Often can't settle a thread, new threads aren't always added to the sidebar and the only way to get back to them is hitting the new button - but that means I also can't create new threads.
Woudlnt' be surprised if ti was somehow related to the above one (I'm assumign I have a thread in a broken state that's preventing other actions).
- changed the title
[-][Bug]: Thread acts as if it's already settled — Settle silently does nothing until the client is restarted[/-][+][Bug]: Thread already settled on the server shows as unsettled and can't be settled (observed multiple times)[/+]on Jul 26, 2026 Confirmed instance with server-side evidence, from a case where the server runs on a different machine than the client — so the two sides could be inspected independently.
Thread
e22c69b9-f1c5-4b14-8a6f-8f0084752556("Investigate Host-Specific Favicon Issue"). Client on host A, execution environment on host B.The server has it settled. From host B's
state.sqlite:settled_override = settled settled_at = 2026-07-27T01:06:03.253Z archived_at = NULL pending_approval_count = 0 pending_user_input_count = 0 projection_thread_sessions.status = ready, active_turn_id = NULLTen settle commands, every one accepted. The user clicked repeatedly because the UI never changed — two bursts, ~20 minutes apart:
sequence event_type occurred_at settledAt (payload) updatedAt (payload) 20752 thread.settled 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20753 thread.settled 2026-07-27T01:06:04.793Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20754 thread.settled 2026-07-27T01:06:06.283Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20755 thread.settled 2026-07-27T01:06:06.795Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20756 thread.settled 2026-07-27T01:06:12.521Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20757 thread.settled 2026-07-27T01:08:22.410Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20758 thread.settled 2026-07-27T01:08:24.456Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20759 thread.settled 2026-07-27T01:08:25.919Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20760 thread.settled 2026-07-27T01:08:26.985Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253Z 20761 thread.settled 2026-07-27T01:08:28.057Z 2026-07-27T01:06:03.253Z 2026-07-27T01:06:03.253ZEvery corresponding row in
orchestration_command_receiptsisaccepted, with norejectedrows. The first settle did the work; the other nine re-emitted with the originalsettledAtandupdatedAt, exactly asdecider.ts:489-505intends.This settles the earlier ambiguity in this issue: it is not a client-side pre-send block. The commands reach the server, the server accepts them, the thread is settled — and the client never reflects it.
The server did emit the shell events.
threadUpsertOrRemove(ws.ts:725-750) refetches the current projected row and emitsthread-upsertedafter each of these. Its one silent-drop path is a projection refetch that fails twice ("If both attempts fail, log and drop the stream item",ws.ts:663-684) — and that path logs a warning.journalctlfor the server has zero occurrences ofshell projection refetch failedover its whole uptime, which covers both bursts. So the events were emitted and the client discarded or never received them.Also ruled out while narrowing this:
- Client shell cache cross-contamination. Sequence spaces are per-environment and differ by an order of magnitude here (host B head is
20761; the client's other, local environment is at195842). A snapshot sequence leaking across environments would makeitem.sequence > snapshot.snapshotSequence(shell.ts:149-158) discard every event from the lower-numbered environment permanently. But the shell cache is keyed byenvironmentIdand re-verifiesstored.environmentId === environmentIdon load (apps/web/src/connection/storage.ts:464-476), so this is not it. - Server-side snapshot sequencing.
computeSnapshotSequencetakes the min across required projectors in the same transaction as the projection reads.
What remains is a client whose shell subscription has stopped applying events for one environment while the socket stays healthy enough to round-trip commands on that same connection. The connection is not "down" in any way the client can currently notice — which is why nothing recovers it and why only a full restart (fresh HTTP snapshot + resubscribe) clears it.
Two properties make this unrecoverable rather than merely wrong:
- The user's only tool, clicking Settle, is a guaranteed no-op once the server already agrees — by design.
- Reconciliation is reachable only on initial load, foreground wake, or reconnect. A stalled-but-connected subscription hits none of them.
Worth noting the client had been running ~45 minutes when the first burst happened, and the server was under sustained load at the time (
VcsStatusBroadcastergit fetch timeouts firing continuously), if that helps anyone reproduce it.- Client shell cache cross-contamination. Sequence spaces are per-environment and differ by an order of magnitude here (host B head is
Follow-up: the same client is stale for the whole environment, not just the settled thread — and other machines connected to the same server show correct state. That isolates it to one client's subscriptions and points at a specific mechanism.
The tell: RPC works while subscriptions do not. The ten settle commands were accepted, and the client received their responses —
SidebarV2.attemptSettletoasts on failure and nothing was toasted, so the mutations resolved as successes. Commands and subscription pushes share one WebSocket, so this is not transport-level. The RPC channel is healthy; the subscriptions on it are not delivering.Where that can happen. In
subscribeDynamic(packages/client-runtime/src/rpc/client.ts:211-222), a transport failure on a subscription stream logs and drains:if (isTransportFailure) { return Stream.fromEffect( Effect.logWarning("Durable RPC subscription lost its transport; waiting for the next session.", ...) ).pipe(Stream.drain); }
Draining ends the inner stream. Because it sits inside
Stream.switchMapoversessions, nothing more is produced untilsessionsemits again. Andsupervisor.sessionis only written when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243).So if a subscription stream fails with an RPC client error while the underlying session stays alive — a blip that does not tear down the lease — every affected subscription for that environment is permanently dead, with no retry and no error surfaced. New RPC calls on that same session continue to work, which is exactly why the app looks connected and commands still round-trip.
That matches every observation here:
- whole environment frozen in this client, other clients fine
- commands accepted and answered on the same socket
- no reconnect, because nothing was disconnected
- only a full client restart recovers it — a restart makes a new session, which is the one thing that resubscribes
- a thread displaying as Working long after it finished, per its stale last-known shell row
Note the contrast with the branch immediately below it: an expected failure with
retryExpectedFailureAfterset does retry (shell.tspasses"250 millis"). Only the transport branch drains unconditionally. If the intent is "the supervisor will bring us back", that holds only when the failure also tears down the lease.Two things would help independently of which layer gets the fix:
- Something must be able to ask for reconciliation without a restart. Today reconciliation is reachable only on initial load, foreground wake, or reconnect, and a drained-but-connected subscription hits none of them.
- A subscription that has stopped delivering is invisible. The shell state machine still reports
live, so no indicator, no toast, nothing distinguishes it from a quiet environment.
I have a PR up for the first (#4593), which adds a
resync-requestedwakeup routed into the sameresubscribestream thatsubscribeDynamicmerges intosessions— so requesting a resync does resubscribe a drained subscription. That was written as a mitigation for the settle symptom before this mechanism was understood; it happens to recover this case, but it only fires from the settle path, so it is not a general fix for the drain.One more observation that pins the mechanism down, because it distinguishes a dead subscription from a dead connection.
While the client was in this state, a new thread was created successfully and never appeared in the sidebar. Server-side it exists and is running normally:
thread b7d3f304-eb83-4184-89ed-73feb2af74d6 "Start coding conversation" created_at 2026-07-27T01:17:07.394Z thread.activity-appended / thread.session-set events continuing through 01:18:05So on one connection, at the same moment:
- RPC works — thread creation round-tripped, as did ten settle commands earlier
- The pre-existing shell subscription is dead — the new thread produces a
thread-upsertedthat never lands, so the sidebar does not list it - A newly created thread-detail subscription works — the thread opens and streams
That asymmetry is the signature of the drain in
subscribeDynamic(packages/client-runtime/src/rpc/client.ts:211-222). The drain is per subscription stream: a stream that hit the transport branch is over and waits forsessionsto emit, whilesubscribeToSession()runs fresh for any subscription started afterward and works fine against the same still-healthy session. A connection-level failure could not produce this — it would take out the new subscription too.Practical blast radius: every subscription established before the blip is dead for that environment, and everything opened afterward works. The app therefore reads as selectively broken rather than disconnected — the sidebar frozen (stale rows, threads stuck showing Working long after they finished, new threads missing) while any thread opened now behaves normally. Other clients against the same server are unaffected, which is what first ruled out the server here.
Nothing surfaces any of this. The shell state machine still reports
live, so there is no indicator, no toast, and no way for the user to tell a frozen environment from a quiet one — which is why the reachable-for-humans symptom was "Settle does nothing".The narrow fix is at the drain itself: retry the transport branch with backoff rather than waiting for a session change that never comes when the lease survives. Note the branch directly below it already does this for expected failures when
retryExpectedFailureAfteris set —shell.tspasses"250 millis". Only the transport branch drains unconditionally.A liveness signal seems worth having regardless. A subscription that has stopped delivering is currently indistinguishable from an idle one, at every layer above it.
- changed the title
[-][Bug]: Thread already settled on the server shows as unsettled and can't be settled (observed multiple times)[/-][+][Bug]: An environment's threads silently stop updating until the client is restarted[/+]on Jul 27, 2026 Root-cause fix opened as #4602 — a durable subscription that loses its transport now asks the supervisor to re-establish, instead of draining and waiting for a session that nothing produces. Reconnect requests are coalesced per session, since a dead transport takes every subscription for an environment at once and each
retryNowaborts an in-flight establishment attempt.#4593 (reconcile on settle) remains useful as a targeted recovery for the settle path, but #4602 is the actual bug.
Two adjacent gaps found while narrowing this, neither addressed by either PR:
- Unary
request(rpc/client.ts:107) has no timeout, so a call dispatched into a hung session never settles. In the web app that can latchattemptSettle's in-flight guard (SidebarV2.tsx:1694) permanently, since the key is cleared in afinallythat never runs — which would make every later click return before any check, send, or toast. - The only liveness probe runs on an
application-activewakeup, so a desktop window that stays visible never health-checks its connection.
- Unary
Following #4602 and #4405 from a production incident on a long-lived fork deployment (desktop client, two providers): sharing three additional failure modes we hit in the same subsystem, since neither PR covers them.
-
Terminal misses retry forever. Deleting a thread whose subscription was mid-flight leaves
orchestration.subscribeThreadfailing withOrchestrationGetSnapshotError: Thread … was not found— classified as an expected failure and retried at the fixed delay for the rest of the session. We counted 12,000+ consecutive misses (~4/s for 80+ minutes) against a healthy connection. fix(client): self-heal empty thread details and back off failed thread subscriptions #4405's backoff shrinks the storm but keeps it eternal; a terminal classification (all-Fail causes whose errors all identify a permanent miss → stop, mark the thread deleted via the existing deleted-thread path) ends it. -
Input construction failures kill the subscription fiber silently. In
subscribeDynamic,makeInputruns outside thecatchCauseenvelope. Its real work includes the HTTP snapshot load plus reducer application, so a transient failure there dies out ofStream.runForEachwith no log, no retry, no error state — RPC works, streams never come back until restart. Matches the "RPC works while subscriptions do not" signature described upthread. Moving the input step inside the same classifier makes it recoverable without touching transport semantics. -
A dead stream renders as healthy work. With fibers dead, the sidebar keeps "Working" (with a ticking timer) and chat shows optimistic echoes only. We added pure liveness derivations (platform-neutral, ready for both web and mobile) and a distinct Reconnecting presentation whenever the stream that would deliver the completion isn't alive. Related to [Bug]: Mobile thread list projection goes stale over a churning remote connection — stuck "Working", missing new threads, fixed only by app restart #5742's stuck-Working symptoms.
Happy to send these as small focused PRs (terminal handling; input-envelope; Reconnecting surfaces) and/or stack on #4602/#4405 whatever shape you prefer. Fork branch with working code + tests available on request.
-
Before submitting
Area
apps/web
Steps to reproduce
Observed multiple times, no deterministic repro yet. The pattern:
Expected behavior
A settled thread displays as settled. Failing that, clicking Settle converges the view — or says why it can't.
Actual behavior
The thread displays as unsettled while the server considers it settled, and clicking Settle does nothing at all — no state change, no toast, no error, no feedback of any kind. There is no way to move the thread into the settled group from the UI. Repeated clicks are equally silent.
Restarting the desktop client clears it.
Technical detail
The command round-trip is itself a state sync:
threadUpsertOrRemove(apps/server/src/ws.ts:725-750) re-reads the current projected thread row and emits athread-upsertedcarrying it after any thread event — including the idempotent settle re-emit below. So a settle sent against a client that disagrees with the server should return the authoritative row and converge the client.Two candidate explanations, in order of how well they fit "nothing visibly happens":
1. The settle succeeds and the client never applies the resulting update.
thread.settleon an already-settled thread re-emitsthread.settledwith the originalsettledAtand the existingupdatedAt(apps/server/src/orchestration/decider.ts:489-505), deliberately, so bulk-settle and double-click stay quiet no-ops. Success is not toasted (it is a high-frequency lifecycle action). So if the client's shell stream is stalled or its snapshot is stale, every click succeeds server-side, emits an authoritativethread-upserted, and the user sees absolutely nothing change — indefinitely. A client restart reloads the HTTP shell snapshot and resubscribes, which is consistent with restart being the only known fix.2. A client-side pre-send block.
settleThread(apps/web/src/hooks/useThreadActions.ts:415) returns without sending whenreadEnvironmentSupportsSettlement(...)is false (line 419 — fires if the environment's server config is missing or stale, sincethreadSettlementdefaults tofalseon decode) or whencanSettle(...)is false (line 433 — fires while the client's copy ofsession.statusisstarting | running,packages/client-runtime/src/state/threadSettled.ts:86). Note this path is not silent in the web app: every settle call site routes throughSidebarV2.attemptSettle(:1690), which surfaces a "Failed to settle thread" toast on failure. So this explains a settle that visibly fails, not one that does nothing.Worth noting:
canSettleis used in exactly two places, both action handlers, and never to disable or hide the control — so the control is always live regardless of state.One instance captured in detail, from
state.sqlite, stream9cf5c431-…. Caveat: this instance shows the server having never settled the thread, so it is likely path 2 rather than the already-settled case, and may be a distinct bug presenting similarly. Including it for the forensics:orchestration_command_receipts: nothread.settlereceipt for the aggregate, and norejectedrows — the command never reached the serverorchestration_events: nothread.settled/thread.unsettledrows for the stream, everprojection_threads.settled_override:NULLthroughoutthread.session-set→readyat20:27:44.228Z(turn completed), nextstartingnot until20:28:26.074ZFor what it's worth, the server-side sync design looks sound on inspection —
computeSnapshotSequencetakes the min across required projectors inside the same transaction as the projection reads (ProjectionSnapshotQuery.ts:160-180), and a missing or too-largeafterSequencefalls back to a full snapshot (ws.ts:1286-1327). If this is a stale-view bug, the interesting question is what leaves a client subscribed but not applying updates.Suggested mitigation regardless of root cause: have the settle action reconcile the thread's authoritative state when it is pressed, rather than acting on a local view that may be stale, and give the user feedback when a settle cannot change anything. Today a no-op success and a healthy settle are indistinguishable from the UI.
Impact
Major degradation or frequent failure
Version or commit
0.0.28-flakedesktop build from a local integration stack on top of upstreammain@5719e8ac4Environment
NixOS, Electron 41.9.1 desktop app, Node 24.18.0,
claudeAgentprovider (Opus)Logs or stack traces
Workaround
Restart the desktop client. Settle behaves normally afterwards.
Possibly related but distinct: #4561 and #4584 land in the same user-visible family (thread won't settle / stuck showing Working), but both are server-side persisted-state bugs where
projection_thread_sessionsitself is wrong after a stop-then-quit or a mid-turn app death. Here the disagreement is between the server's view and the client's copy of it.