Skip to content

[Bug]: An environment's threads silently stop updating until the client is restarted #4589

Description

@colonelpanic8

Update — root cause identified. This is not specific to settling. A client's subscriptions to an environment can die permanently after a transport blip while the connection stays healthy, leaving that environment's threads frozen until the client restarts. "Settle does nothing" was just the first reachable symptom. See the confirmed server-side evidence, the mechanism in subscribeDynamic, and the dead-subscription vs dead-connection asymmetry. The original report follows unchanged.


Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/web

Steps to reproduce

Observed multiple times, no deterministic repro yet. The pattern:

  1. Desktop app, thread that the server already considers settled.
  2. The thread shows in the unsettled/active group anyway.
  3. Click Settle on it.
  4. Nothing happens — it stays showing as unsettled.

Expected behavior

A settled thread displays as settled. Failing that, clicking Settle converges the view — or says why it can't.

Actual behavior

The thread displays as unsettled while the server considers it settled, and clicking Settle does nothing at all — no state change, no toast, no error, no feedback of any kind. There is no way to move the thread into the settled group from the UI. Repeated clicks are equally silent.

Restarting the desktop client clears it.

Technical detail

The command round-trip is itself a state sync: threadUpsertOrRemove (apps/server/src/ws.ts:725-750) re-reads the current projected thread row and emits a thread-upserted carrying it after any thread event — including the idempotent settle re-emit below. So a settle sent against a client that disagrees with the server should return the authoritative row and converge the client.

Two candidate explanations, in order of how well they fit "nothing visibly happens":

1. The settle succeeds and the client never applies the resulting update. thread.settle on an already-settled thread re-emits thread.settled with the original settledAt and the existing updatedAt (apps/server/src/orchestration/decider.ts:489-505), deliberately, so bulk-settle and double-click stay quiet no-ops. Success is not toasted (it is a high-frequency lifecycle action). So if the client's shell stream is stalled or its snapshot is stale, every click succeeds server-side, emits an authoritative thread-upserted, and the user sees absolutely nothing change — indefinitely. A client restart reloads the HTTP shell snapshot and resubscribes, which is consistent with restart being the only known fix.

2. A client-side pre-send block. settleThread (apps/web/src/hooks/useThreadActions.ts:415) returns without sending when readEnvironmentSupportsSettlement(...) is false (line 419 — fires if the environment's server config is missing or stale, since threadSettlement defaults to false on decode) or when canSettle(...) is false (line 433 — fires while the client's copy of session.status is starting | running, packages/client-runtime/src/state/threadSettled.ts:86). Note this path is not silent in the web app: every settle call site routes through SidebarV2.attemptSettle (:1690), which surfaces a "Failed to settle thread" toast on failure. So this explains a settle that visibly fails, not one that does nothing.

Worth noting: canSettle is used in exactly two places, both action handlers, and never to disable or hide the control — so the control is always live regardless of state.

One instance captured in detail, from state.sqlite, stream 9cf5c431-…. Caveat: this instance shows the server having never settled the thread, so it is likely path 2 rather than the already-settled case, and may be a distinct bug presenting similarly. Including it for the forensics:

  • orchestration_command_receipts: no thread.settle receipt for the aggregate, and no rejected rows — the command never reached the server
  • orchestration_events: no thread.settled / thread.unsettled rows for the stream, ever
  • projection_threads.settled_override: NULL throughout
  • thread.session-set → ready at 20:27:44.228Z (turn completed), next starting not until 20:28:26.074Z

For what it's worth, the server-side sync design looks sound on inspection — computeSnapshotSequence takes the min across required projectors inside the same transaction as the projection reads (ProjectionSnapshotQuery.ts:160-180), and a missing or too-large afterSequence falls back to a full snapshot (ws.ts:1286-1327). If this is a stale-view bug, the interesting question is what leaves a client subscribed but not applying updates.

Suggested mitigation regardless of root cause: have the settle action reconcile the thread's authoritative state when it is pressed, rather than acting on a local view that may be stale, and give the user feedback when a settle cannot change anything. Today a no-op success and a healthy settle are indistinguishable from the UI.

Impact

Major degradation or frequent failure

Version or commit

0.0.28-flake desktop build from a local integration stack on top of upstream main @ 5719e8ac4

Environment

NixOS, Electron 41.9.1 desktop app, Node 24.18.0, claudeAgent provider (Opus)

Logs or stack traces

# captured instance — session status transitions (orchestration_events, stream 9cf5c431-…)
187909  2026-07-26T20:26:21.975Z  running
188015  2026-07-26T20:27:18.405Z  running
188051  2026-07-26T20:27:44.228Z  ready      <-- settle attempted in the window after this
188071  2026-07-26T20:28:26.074Z  starting
188072  2026-07-26T20:28:26.147Z  running

# settle events for this stream
select sequence,event_type from orchestration_events
  where stream_id='9cf5c431-…' and event_type in ('thread.settled','thread.unsettled');
-- (0 rows)

# receipts for this aggregate: no thread.settle, nothing rejected
select status,count(*) from orchestration_command_receipts where aggregate_id='9cf5c431-…' group by status;
accepted|…

# projection state
projection_threads.settled_override = NULL
projection_threads.archived_at      = NULL

Workaround

Restart the desktop client. Settle behaves normally afterwards.


Possibly related but distinct: #4561 and #4584 land in the same user-visible family (thread won't settle / stuck showing Working), but both are server-side persisted-state bugs where projection_thread_sessions itself is wrong after a stop-then-quit or a mid-turn app death. Here the disagreement is between the server's view and the client's copy of it.

Activity

  1. changed the title [-][Bug]: Client holds stale running session state after server reports ready, blocking settle client-side with no command sent[/-] [+][Bug]: Clicking Settle silently does nothing on an idle thread until the client is restarted[/+] on Jul 26, 2026
  2. changed the title [-][Bug]: Clicking Settle silently does nothing on an idle thread until the client is restarted[/-] [+][Bug]: Thread acts as if it's already settled — Settle silently does nothing until the client is restarted[/+] on Jul 26, 2026
  3. bzbetty commented on Jul 26, 2026

    @bzbetty

    I'm seeing all sorts of issues with the recent nightlies along the same lines (doesn't matter if I have new sidebar on or off).

    Often can't settle a thread, new threads aren't always added to the sidebar and the only way to get back to them is hitting the new button - but that means I also can't create new threads.

    Woudlnt' be surprised if ti was somehow related to the above one (I'm assumign I have a thread in a broken state that's preventing other actions).

  4. changed the title [-][Bug]: Thread acts as if it's already settled — Settle silently does nothing until the client is restarted[/-] [+][Bug]: Thread already settled on the server shows as unsettled and can't be settled (observed multiple times)[/+] on Jul 26, 2026
  5. colonelpanic8 commented on Jul 27, 2026

    @colonelpanic8
    ContributorAuthor

    Confirmed instance with server-side evidence, from a case where the server runs on a different machine than the client — so the two sides could be inspected independently.

    Thread e22c69b9-f1c5-4b14-8a6f-8f0084752556 ("Investigate Host-Specific Favicon Issue"). Client on host A, execution environment on host B.

    The server has it settled. From host B's state.sqlite:

    settled_override = settled
    settled_at       = 2026-07-27T01:06:03.253Z
    archived_at      = NULL
    pending_approval_count = 0
    pending_user_input_count = 0
    projection_thread_sessions.status = ready, active_turn_id = NULL
    

    Ten settle commands, every one accepted. The user clicked repeatedly because the UI never changed — two bursts, ~20 minutes apart:

    sequence  event_type      occurred_at               settledAt (payload)       updatedAt (payload)
    20752     thread.settled  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20753     thread.settled  2026-07-27T01:06:04.793Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20754     thread.settled  2026-07-27T01:06:06.283Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20755     thread.settled  2026-07-27T01:06:06.795Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20756     thread.settled  2026-07-27T01:06:12.521Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20757     thread.settled  2026-07-27T01:08:22.410Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20758     thread.settled  2026-07-27T01:08:24.456Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20759     thread.settled  2026-07-27T01:08:25.919Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20760     thread.settled  2026-07-27T01:08:26.985Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    20761     thread.settled  2026-07-27T01:08:28.057Z  2026-07-27T01:06:03.253Z  2026-07-27T01:06:03.253Z
    

    Every corresponding row in orchestration_command_receipts is accepted, with no rejected rows. The first settle did the work; the other nine re-emitted with the original settledAt and updatedAt, exactly as decider.ts:489-505 intends.

    This settles the earlier ambiguity in this issue: it is not a client-side pre-send block. The commands reach the server, the server accepts them, the thread is settled — and the client never reflects it.

    The server did emit the shell events. threadUpsertOrRemove (ws.ts:725-750) refetches the current projected row and emits thread-upserted after each of these. Its one silent-drop path is a projection refetch that fails twice ("If both attempts fail, log and drop the stream item", ws.ts:663-684) — and that path logs a warning. journalctl for the server has zero occurrences of shell projection refetch failed over its whole uptime, which covers both bursts. So the events were emitted and the client discarded or never received them.

    Also ruled out while narrowing this:

    • Client shell cache cross-contamination. Sequence spaces are per-environment and differ by an order of magnitude here (host B head is 20761; the client's other, local environment is at 195842). A snapshot sequence leaking across environments would make item.sequence > snapshot.snapshotSequence (shell.ts:149-158) discard every event from the lower-numbered environment permanently. But the shell cache is keyed by environmentId and re-verifies stored.environmentId === environmentId on load (apps/web/src/connection/storage.ts:464-476), so this is not it.
    • Server-side snapshot sequencing. computeSnapshotSequence takes the min across required projectors in the same transaction as the projection reads.

    What remains is a client whose shell subscription has stopped applying events for one environment while the socket stays healthy enough to round-trip commands on that same connection. The connection is not "down" in any way the client can currently notice — which is why nothing recovers it and why only a full restart (fresh HTTP snapshot + resubscribe) clears it.

    Two properties make this unrecoverable rather than merely wrong:

    1. The user's only tool, clicking Settle, is a guaranteed no-op once the server already agrees — by design.
    2. Reconciliation is reachable only on initial load, foreground wake, or reconnect. A stalled-but-connected subscription hits none of them.

    Worth noting the client had been running ~45 minutes when the first burst happened, and the server was under sustained load at the time (VcsStatusBroadcaster git fetch timeouts firing continuously), if that helps anyone reproduce it.

  6. colonelpanic8 commented on Jul 27, 2026

    @colonelpanic8
    ContributorAuthor

    Follow-up: the same client is stale for the whole environment, not just the settled thread — and other machines connected to the same server show correct state. That isolates it to one client's subscriptions and points at a specific mechanism.

    The tell: RPC works while subscriptions do not. The ten settle commands were accepted, and the client received their responses — SidebarV2.attemptSettle toasts on failure and nothing was toasted, so the mutations resolved as successes. Commands and subscription pushes share one WebSocket, so this is not transport-level. The RPC channel is healthy; the subscriptions on it are not delivering.

    Where that can happen. In subscribeDynamic (packages/client-runtime/src/rpc/client.ts:211-222), a transport failure on a subscription stream logs and drains:

    if (isTransportFailure) {
      return Stream.fromEffect(
        Effect.logWarning("Durable RPC subscription lost its transport; waiting for the next session.", ...)
      ).pipe(Stream.drain);
    }

    Draining ends the inner stream. Because it sits inside Stream.switchMap over sessions, nothing more is produced until sessions emits again. And supervisor.session is only written when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243).

    So if a subscription stream fails with an RPC client error while the underlying session stays alive — a blip that does not tear down the lease — every affected subscription for that environment is permanently dead, with no retry and no error surfaced. New RPC calls on that same session continue to work, which is exactly why the app looks connected and commands still round-trip.

    That matches every observation here:

    • whole environment frozen in this client, other clients fine
    • commands accepted and answered on the same socket
    • no reconnect, because nothing was disconnected
    • only a full client restart recovers it — a restart makes a new session, which is the one thing that resubscribes
    • a thread displaying as Working long after it finished, per its stale last-known shell row

    Note the contrast with the branch immediately below it: an expected failure with retryExpectedFailureAfter set does retry (shell.ts passes "250 millis"). Only the transport branch drains unconditionally. If the intent is "the supervisor will bring us back", that holds only when the failure also tears down the lease.

    Two things would help independently of which layer gets the fix:

    1. Something must be able to ask for reconciliation without a restart. Today reconciliation is reachable only on initial load, foreground wake, or reconnect, and a drained-but-connected subscription hits none of them.
    2. A subscription that has stopped delivering is invisible. The shell state machine still reports live, so no indicator, no toast, nothing distinguishes it from a quiet environment.

    I have a PR up for the first (#4593), which adds a resync-requested wakeup routed into the same resubscribe stream that subscribeDynamic merges into sessions — so requesting a resync does resubscribe a drained subscription. That was written as a mitigation for the settle symptom before this mechanism was understood; it happens to recover this case, but it only fires from the settle path, so it is not a general fix for the drain.

  7. colonelpanic8 commented on Jul 27, 2026

    @colonelpanic8
    ContributorAuthor

    One more observation that pins the mechanism down, because it distinguishes a dead subscription from a dead connection.

    While the client was in this state, a new thread was created successfully and never appeared in the sidebar. Server-side it exists and is running normally:

    thread b7d3f304-eb83-4184-89ed-73feb2af74d6  "Start coding conversation"
    created_at 2026-07-27T01:17:07.394Z
    thread.activity-appended / thread.session-set events continuing through 01:18:05
    

    So on one connection, at the same moment:

    • RPC works — thread creation round-tripped, as did ten settle commands earlier
    • The pre-existing shell subscription is dead — the new thread produces a thread-upserted that never lands, so the sidebar does not list it
    • A newly created thread-detail subscription works — the thread opens and streams

    That asymmetry is the signature of the drain in subscribeDynamic (packages/client-runtime/src/rpc/client.ts:211-222). The drain is per subscription stream: a stream that hit the transport branch is over and waits for sessions to emit, while subscribeToSession() runs fresh for any subscription started afterward and works fine against the same still-healthy session. A connection-level failure could not produce this — it would take out the new subscription too.

    Practical blast radius: every subscription established before the blip is dead for that environment, and everything opened afterward works. The app therefore reads as selectively broken rather than disconnected — the sidebar frozen (stale rows, threads stuck showing Working long after they finished, new threads missing) while any thread opened now behaves normally. Other clients against the same server are unaffected, which is what first ruled out the server here.

    Nothing surfaces any of this. The shell state machine still reports live, so there is no indicator, no toast, and no way for the user to tell a frozen environment from a quiet one — which is why the reachable-for-humans symptom was "Settle does nothing".

    The narrow fix is at the drain itself: retry the transport branch with backoff rather than waiting for a session change that never comes when the lease survives. Note the branch directly below it already does this for expected failures when retryExpectedFailureAfter is set — shell.ts passes "250 millis". Only the transport branch drains unconditionally.

    A liveness signal seems worth having regardless. A subscription that has stopped delivering is currently indistinguishable from an idle one, at every layer above it.

  8. changed the title [-][Bug]: Thread already settled on the server shows as unsettled and can't be settled (observed multiple times)[/-] [+][Bug]: An environment's threads silently stop updating until the client is restarted[/+] on Jul 27, 2026
  9. colonelpanic8 commented on Jul 27, 2026

    @colonelpanic8
    ContributorAuthor

    Root-cause fix opened as #4602 — a durable subscription that loses its transport now asks the supervisor to re-establish, instead of draining and waiting for a session that nothing produces. Reconnect requests are coalesced per session, since a dead transport takes every subscription for an environment at once and each retryNow aborts an in-flight establishment attempt.

    #4593 (reconcile on settle) remains useful as a targeted recovery for the settle path, but #4602 is the actual bug.

    Two adjacent gaps found while narrowing this, neither addressed by either PR:

    • Unary request (rpc/client.ts:107) has no timeout, so a call dispatched into a hung session never settles. In the web app that can latch attemptSettle's in-flight guard (SidebarV2.tsx:1694) permanently, since the key is cleared in a finally that never runs — which would make every later click return before any check, send, or toast.
    • The only liveness probe runs on an application-active wakeup, so a desktop window that stays visible never health-checks its connection.
  10. Coriou commented on Aug 24, 2026

    @Coriou

    Following #4602 and #4405 from a production incident on a long-lived fork deployment (desktop client, two providers): sharing three additional failure modes we hit in the same subsystem, since neither PR covers them.

    1. Terminal misses retry forever. Deleting a thread whose subscription was mid-flight leaves orchestration.subscribeThread failing with OrchestrationGetSnapshotError: Thread … was not found — classified as an expected failure and retried at the fixed delay for the rest of the session. We counted 12,000+ consecutive misses (~4/s for 80+ minutes) against a healthy connection. fix(client): self-heal empty thread details and back off failed thread subscriptions #4405's backoff shrinks the storm but keeps it eternal; a terminal classification (all-Fail causes whose errors all identify a permanent miss → stop, mark the thread deleted via the existing deleted-thread path) ends it.

    2. Input construction failures kill the subscription fiber silently. In subscribeDynamic, makeInput runs outside the catchCause envelope. Its real work includes the HTTP snapshot load plus reducer application, so a transient failure there dies out of Stream.runForEach with no log, no retry, no error state — RPC works, streams never come back until restart. Matches the "RPC works while subscriptions do not" signature described upthread. Moving the input step inside the same classifier makes it recoverable without touching transport semantics.

    3. A dead stream renders as healthy work. With fibers dead, the sidebar keeps "Working" (with a ticking timer) and chat shows optimistic echoes only. We added pure liveness derivations (platform-neutral, ready for both web and mobile) and a distinct Reconnecting presentation whenever the stream that would deliver the completion isn't alive. Related to [Bug]: Mobile thread list projection goes stale over a churning remote connection — stuck "Working", missing new threads, fixed only by app restart #5742's stuck-Working symptoms.

    Happy to send these as small focused PRs (terminal handling; input-envelope; Reconnecting surfaces) and/or stack on #4602/#4405 whatever shape you prefer. Fork branch with working code + tests available on request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions