Repository navigation
[Bug]: New thread sometimes stays on "Thinking" with a spinning send button after a quick first turn until reload #14046
Description
Activity
Triage
This is a real client sync bug, and it is still present on current
main. It is not a duplicate. The turn is finishing on the server; the sending view never applies it.The stuck controls are the local send state, not a running turn. The spinning icon is the send button’s busy spinner (
isSendBusy). The red Stop icon is a different state, and it only appears once the session phase isrunning. “Thinking” is the timeline working row, and that row is shown whenever a send is still busy, even if no assistant text has arrived. A send stays busy until the server thread acknowledges it. Acknowledgement is suppressed the whole time the session isstarting(derivePhasereportsconnecting).On a new thread that acknowledgement never gets a chance to arrive:
- The draft stays on screen until the client has seen a persisted user message, a started turn, or a failed startup (
resolveDraftPromotionNavigationTarget). - Until then, thread detail is not subscribed.
useThreadwaits for the shell index while the draft still exists. - Leaving the draft is also what clears the local send (
draftId/threadIdchange). A reload works because it reads the completed snapshot. That matches this report: the server stored the turn, and this client never did.
#9521 is already in
6530de0339and onmain. It only keeps events that arrive while an existingsubscribeThreadsnapshot is loading. This report is the step before that subscription exists.A short first turn makes that window smaller than the shell coalesce (one refetch per burst, latest event per thread). The client will not open the thread until that refetch lands. If the refetch fails, the failure is logged and the item is dropped (
retryShellProjectionReadinsubscribeShell). Nothing later asks for the thread, so the draft waits forever. After a detail cursor does exist,applyItemLockedstill advanceslastSequencebefore it has thread data and then discards the event. A later snapshot does not replay what was skipped.Related, but not the same bug:
- [Bug]: New thread in a freshly-created project gets stuck on "Loading messages..." forever, even though the backend completes the turn #11681 stays on “Loading messages…” and a reload does not recover. It also involves a second project row.
- [Bug]: An environment's threads silently stop updating until the client is restarted #4589 is a subscription that dies after the thread is already open.
- fix(threads): stop missing-thread subscription retries #12711 stops missing-thread retry loops and makes the draft shell gate explicit. It does not deliver a turn the shell never announced.
A server trace around the send would show which gap it was (shell refetch dropped, or detail cursor already past the turn). The report is enough to treat this as a bug without that.
- The draft stays on screen until the client has seen a persisted user message, a started turn, or a failed startup (
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 28, 2026 Here is a server trace for one occurrence: thread
82fb15c6-c2bf-4b51-a213-0d4d80819809, prompt sent at 04:52:30 UTC, app reloaded at 04:53:37 UTC. The trace slicetrace-82fb15c6-0443-0454.ndjson.gzis attached. It contains every browser span that was exported to the server between 04:43:57 and the reload, plus every server span for this thread. The host name is redacted. It contains no message text.trace-82fb15c6-0443-0454.ndjson.gz
Findings
The server handled the turn correctly. Every command span succeeded between
thread.createat 04:52:30.320 andthread.session.setat 04:53:05.346. Reasoning and assistant deltas were dispatched throughout, and the assistant message completed at 04:53:05.The browser's IndexedDB connection was already dead before the prompt was sent.
- No browser spans exist between about 22:54 the previous day and 04:44:00. The tab was asleep, most likely while the machine was suspended.
- At 04:44:00 the client woke up and resubscribed (
environment.initialSync, interruptedsubscribeShell/subscribeThread). - Starting at 04:44:01, every IndexedDB access in the client failed with
InvalidStateError: Failed to execute 'transaction' on 'IDBDatabase': The database connection is closing.The failing spans includeEnvironmentRegistry.run,VcsRefsState.invalidateCached,EnvironmentThreadState.persist,EnvironmentServerConfigState.persistandweb.connectionStorage.*. - This continued until the reload, and there was no recovery attempt in between.
The new thread's detail state failed to build, so the client never subscribed to it. At 04:52:30.659
EnvironmentThreadState.makefailed:InvalidStateError: Failed to execute 'transaction' on 'IDBDatabase': The database connection is closing. at web.connectionStorage.readDatabaseValue at EnvironmentThreadState.makeBetween that failure and the reload, the client made no
fetchEnvironmentThreadSnapshot, noorchestration.subscribeThreadand noEnvironmentThreadState.applyItemfor the thread. Only shell updates arrived, which is why the sidebar showed activity while the chat view stayed on "Thinking". After the reload,EnvironmentThreadState.makesucceeded at 04:53:37.180, the snapshot was fetched and the subscription started.Likely mechanism
readDatabaseValuecallsdatabase.transaction(...)synchronously insideEffect.callback. On a closed connection that call throws. The throw becomes a defect, not aConnectionTransientError.- Because it is a defect, the fallback in
makeEnvironmentThreadState, which is meant to treat a failed cache load as "no cached thread", never runs. The whole thread-state construction dies instead. - The layer opens the
IDBDatabaseonce and never listens for thecloseevent or reopens after the browser force-closes the connection. So after a long sleep, every later cache access fails for the lifetime of the page.
The cache is only an optimization. Making
readDatabaseValueand the other IndexedDB helpers turn synchronous throws into typed failures would likely fix the stuck thread by itself. Reopening the connection after acloseevent would also restore the cache.Open questions
- Which change caused the regression. The synchronous
transaction()call was already present in0.0.41-nightly.20260915.1766. The only in-range change tostorage.ts(chore: clear Effect language service suggestions #13536) is cosmetic. The regression is therefore more likely in what triggers the connection closing, or in which code paths now read the cache when a new thread is created, than instorage.tsitself. - The later failures. The same IndexedDB error appears again after the reload, at 05:09 and 05:27, including another failed
EnvironmentThreadState.makeat 05:27:37, with no sleep gap before them. Client spans carry no tab or page-load identifier, so the trace cannot show whether they come from a second tab that still holds a dead connection.
Update: the trigger is browser storage eviction under low disk space
It happened again in a second thread, this time with no sleep beforehand. The trace and the browser's quota state point to the same client failure, with a different trigger than I first guessed.
Second occurrence
- The server handled both turns without errors.
- On the client,
EnvironmentThreadState.makefailed again withInvalidStateError: ... The database connection is closing.0.4 s after the thread was created. The thread got no snapshot fetch and nosubscribeThreaduntil a reload. - After a reload at 04:53:37 UTC, the page's IndexedDB worked for about 14 minutes. The last successful write was at 05:07:38.697 and the first failure at 05:09:30.607. The client did nothing in between that could close the connection: no page load, no storage-layer teardown, no gap in client spans. It was only sending periodic
server.reportClientActivitycalls. So the browser closed the connection.
Why the browser closed it
The machine's disk was nearly full, with 1.41 GB available of 399 GB. The browser is Chromium-based. Its
chrome://quota-internalspage showed:- Eviction statistics: 14 evicted buckets across 18 eviction rounds since the browser started. Total site storage was under 90 MB, so evicting could never relieve the pressure, and rounds kept happening.
- The
http://localhost:3773/bucket: a use count of 2, and a "last accessed" time equal to the reload time. The trace shows the page accessing IndexedDB hundreds of times before that. The bucket had been evicted and recreated by the reload.
Under storage pressure, Chromium evicts best-effort buckets, and deleting a bucket force-closes its open IndexedDB connections. Every later
transaction()call then throws. The first occurrence fits the same pattern: the connection was already dead when the tab woke up.What this changes
- The regression between
0.0.41-nightly.20260915.1766and0.0.43-nightly.20260926.2282is probably a coincidence. The disk filling up likely overlapped with the upgrade. - The app bug still stands. Browsers can evict storage or close IndexedDB connections at any time, and the cache is optional. One eviction currently breaks every new thread until the page is reloaded. The fix outlined above still applies:
- Turn synchronous throws in the IndexedDB helpers (
readDatabaseValue,writeDatabaseValue,removeDatabaseValuesInRange) into typed failures, so the existing "no cached thread" fallback runs. - Listen for the
IDBDatabasecloseevent and reopen the connection instead of keeping the dead handle for the rest of the page's life.
- Turn synchronous throws in the IndexedDB helpers (
Likely deterministic repro
Not yet verified: with the web app open, clear the site's storage (DevTools → Application → Storage → "Clear site data"), then start a new thread and send a short prompt. Clearing the storage deletes the IndexedDB database while the page holds it open, which should force-close the connection the same way eviction does.
Before submitting
Area
apps/web
Steps to reproduce
Expected behavior
Actual behavior
Sometimes the thread gets stuck:
Only a full app reload recovers it. After the reload the complete answer is shown and the composer is ready for another prompt. That means the server completed and stored the turn, and the client never applied the updates for it.
Impact
Major degradation or frequent failure
Version or commit
Observed on
0.0.43-nightly.20260926.2282(6530de0339). It was not seen on0.0.41-nightly.20260915.1766, but that is probably a coincidence rather than a regression. The trigger is browser storage eviction under low disk space, which started around the same time as the upgrade. See #14046 (comment).Environment
Web app, Linux host, Claude provider
Logs or stack traces
No response
Screenshots, recordings, or supporting files
No response
Workaround
Reload the app.
Related issues and PRs