Skip to content

[Bug]: New thread sometimes stays on "Thinking" with a spinning send button after a quick first turn until reload #14046

Description

@Jardo-51

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/web

Steps to reproduce

  1. In the web app, create a new thread.
  2. Send a simple prompt that the agent finishes in a couple of seconds.
  3. Watch the composer and the thread timeline while the turn runs and after it ends.

Expected behavior

  • The send button turns red with a Stop icon while the turn runs.
  • The thread streams the agent's answer.
  • When the agent finishes, the send button returns to normal and the thread shows the complete answer.

Actual behavior

Sometimes the thread gets stuck:

  • The send button shows a spinning icon.
  • The thread shows "Thinking" but never renders any of the agent's response.
  • Nothing changes after the agent has finished replying. The thread stays in this state indefinitely.

Only a full app reload recovers it. After the reload the complete answer is shown and the composer is ready for another prompt. That means the server completed and stored the turn, and the client never applied the updates for it.

Impact

Major degradation or frequent failure

Version or commit

Observed on 0.0.43-nightly.20260926.2282 (6530de0339). It was not seen on 0.0.41-nightly.20260915.1766, but that is probably a coincidence rather than a regression. The trigger is browser storage eviction under low disk space, which started around the same time as the upgrade. See #14046 (comment).

Environment

Web app, Linux host, Claude provider

Logs or stack traces

No response

Screenshots, recordings, or supporting files

No response

Workaround

Reload the app.

Related issues and PRs

Activity

  1. juliusmarminge commented on Sep 28, 2026

    @juliusmarminge
    Member

    Triage

    This is a real client sync bug, and it is still present on current main. It is not a duplicate. The turn is finishing on the server; the sending view never applies it.

    The stuck controls are the local send state, not a running turn. The spinning icon is the send button’s busy spinner (isSendBusy). The red Stop icon is a different state, and it only appears once the session phase is running. “Thinking” is the timeline working row, and that row is shown whenever a send is still busy, even if no assistant text has arrived. A send stays busy until the server thread acknowledges it. Acknowledgement is suppressed the whole time the session is starting (derivePhase reports connecting).

    On a new thread that acknowledgement never gets a chance to arrive:

    • The draft stays on screen until the client has seen a persisted user message, a started turn, or a failed startup (resolveDraftPromotionNavigationTarget).
    • Until then, thread detail is not subscribed. useThread waits for the shell index while the draft still exists.
    • Leaving the draft is also what clears the local send (draftId / threadId change). A reload works because it reads the completed snapshot. That matches this report: the server stored the turn, and this client never did.

    #9521 is already in 6530de0339 and on main. It only keeps events that arrive while an existing subscribeThread snapshot is loading. This report is the step before that subscription exists.

    A short first turn makes that window smaller than the shell coalesce (one refetch per burst, latest event per thread). The client will not open the thread until that refetch lands. If the refetch fails, the failure is logged and the item is dropped (retryShellProjectionRead in subscribeShell). Nothing later asks for the thread, so the draft waits forever. After a detail cursor does exist, applyItemLocked still advances lastSequence before it has thread data and then discards the event. A later snapshot does not replay what was skipped.

    Related, but not the same bug:

    A server trace around the send would show which gap it was (shell refetch dropped, or detail cursor already past the turn). The report is enough to treat this as a bug without that.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 28, 2026
  3. Jardo-51 commented on Sep 28, 2026

    @Jardo-51
    ContributorAuthor

    Here is a server trace for one occurrence: thread 82fb15c6-c2bf-4b51-a213-0d4d80819809, prompt sent at 04:52:30 UTC, app reloaded at 04:53:37 UTC. The trace slice trace-82fb15c6-0443-0454.ndjson.gz is attached. It contains every browser span that was exported to the server between 04:43:57 and the reload, plus every server span for this thread. The host name is redacted. It contains no message text.

    trace-82fb15c6-0443-0454.ndjson.gz

    Findings

    The server handled the turn correctly. Every command span succeeded between thread.create at 04:52:30.320 and thread.session.set at 04:53:05.346. Reasoning and assistant deltas were dispatched throughout, and the assistant message completed at 04:53:05.

    The browser's IndexedDB connection was already dead before the prompt was sent.

    • No browser spans exist between about 22:54 the previous day and 04:44:00. The tab was asleep, most likely while the machine was suspended.
    • At 04:44:00 the client woke up and resubscribed (environment.initialSync, interrupted subscribeShell / subscribeThread).
    • Starting at 04:44:01, every IndexedDB access in the client failed with InvalidStateError: Failed to execute 'transaction' on 'IDBDatabase': The database connection is closing. The failing spans include EnvironmentRegistry.run, VcsRefsState.invalidateCached, EnvironmentThreadState.persist, EnvironmentServerConfigState.persist and web.connectionStorage.*.
    • This continued until the reload, and there was no recovery attempt in between.

    The new thread's detail state failed to build, so the client never subscribed to it. At 04:52:30.659 EnvironmentThreadState.make failed:

    InvalidStateError: Failed to execute 'transaction' on 'IDBDatabase': The database connection is closing.
        at web.connectionStorage.readDatabaseValue
        at EnvironmentThreadState.make
    

    Between that failure and the reload, the client made no fetchEnvironmentThreadSnapshot, no orchestration.subscribeThread and no EnvironmentThreadState.applyItem for the thread. Only shell updates arrived, which is why the sidebar showed activity while the chat view stayed on "Thinking". After the reload, EnvironmentThreadState.make succeeded at 04:53:37.180, the snapshot was fetched and the subscription started.

    Likely mechanism

    1. readDatabaseValue calls database.transaction(...) synchronously inside Effect.callback. On a closed connection that call throws. The throw becomes a defect, not a ConnectionTransientError.
    2. Because it is a defect, the fallback in makeEnvironmentThreadState, which is meant to treat a failed cache load as "no cached thread", never runs. The whole thread-state construction dies instead.
    3. The layer opens the IDBDatabase once and never listens for the close event or reopens after the browser force-closes the connection. So after a long sleep, every later cache access fails for the lifetime of the page.

    The cache is only an optimization. Making readDatabaseValue and the other IndexedDB helpers turn synchronous throws into typed failures would likely fix the stuck thread by itself. Reopening the connection after a close event would also restore the cache.

    Open questions

    • Which change caused the regression. The synchronous transaction() call was already present in 0.0.41-nightly.20260915.1766. The only in-range change to storage.ts (chore: clear Effect language service suggestions #13536) is cosmetic. The regression is therefore more likely in what triggers the connection closing, or in which code paths now read the cache when a new thread is created, than in storage.ts itself.
    • The later failures. The same IndexedDB error appears again after the reload, at 05:09 and 05:27, including another failed EnvironmentThreadState.make at 05:27:37, with no sleep gap before them. Client spans carry no tab or page-load identifier, so the trace cannot show whether they come from a second tab that still holds a dead connection.
  4. Jardo-51 commented on Sep 28, 2026

    @Jardo-51
    ContributorAuthor

    Update: the trigger is browser storage eviction under low disk space

    It happened again in a second thread, this time with no sleep beforehand. The trace and the browser's quota state point to the same client failure, with a different trigger than I first guessed.

    Second occurrence

    • The server handled both turns without errors.
    • On the client, EnvironmentThreadState.make failed again with InvalidStateError: ... The database connection is closing. 0.4 s after the thread was created. The thread got no snapshot fetch and no subscribeThread until a reload.
    • After a reload at 04:53:37 UTC, the page's IndexedDB worked for about 14 minutes. The last successful write was at 05:07:38.697 and the first failure at 05:09:30.607. The client did nothing in between that could close the connection: no page load, no storage-layer teardown, no gap in client spans. It was only sending periodic server.reportClientActivity calls. So the browser closed the connection.

    Why the browser closed it

    The machine's disk was nearly full, with 1.41 GB available of 399 GB. The browser is Chromium-based. Its chrome://quota-internals page showed:

    • Eviction statistics: 14 evicted buckets across 18 eviction rounds since the browser started. Total site storage was under 90 MB, so evicting could never relieve the pressure, and rounds kept happening.
    • The http://localhost:3773/ bucket: a use count of 2, and a "last accessed" time equal to the reload time. The trace shows the page accessing IndexedDB hundreds of times before that. The bucket had been evicted and recreated by the reload.

    Under storage pressure, Chromium evicts best-effort buckets, and deleting a bucket force-closes its open IndexedDB connections. Every later transaction() call then throws. The first occurrence fits the same pattern: the connection was already dead when the tab woke up.

    What this changes

    • The regression between 0.0.41-nightly.20260915.1766 and 0.0.43-nightly.20260926.2282 is probably a coincidence. The disk filling up likely overlapped with the upgrade.
    • The app bug still stands. Browsers can evict storage or close IndexedDB connections at any time, and the cache is optional. One eviction currently breaks every new thread until the page is reloaded. The fix outlined above still applies:
      1. Turn synchronous throws in the IndexedDB helpers (readDatabaseValue, writeDatabaseValue, removeDatabaseValuesInRange) into typed failures, so the existing "no cached thread" fallback runs.
      2. Listen for the IDBDatabase close event and reopen the connection instead of keeping the dead handle for the rest of the page's life.

    Likely deterministic repro

    Not yet verified: with the web app open, clear the site's storage (DevTools → Application → Storage → "Clear site data"), then start a new thread and send a short prompt. Clearing the storage deletes the IndexedDB database while the page holds it open, which should force-close the connection the same way eviction does.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions