Skip to content

Threads remain stuck on “Loading messages…” after disk space is freed #10418

Description

@dyc3

What happened

When the disk fills up, some threads in the desktop app stop loading and show “Loading messages…” indefinitely. Other threads still work.

Freeing disk space does not recover the affected threads. Restarting T3 Code has restored access in previous occurrences, but restarting interrupts ongoing work. The current occurrence was investigated after freeing space and before restarting.

Expected behavior: show an actionable storage error and recover after space becomes available.

Diagnosis

Root cause remains unconfirmed.

Read-only investigation found:

  • The existing server process remained alive and answered its root HTTP endpoint.
  • Both an affected thread and a working thread had readable SQLite records, ready Codex sessions, and no recorded session error.
  • The affected thread contained 61 messages and 1,066 activity records. Its message attachment and activity JSON parsed successfully.
  • Retained server and desktop traces contained no explicit ENOSPC or SQLITE_FULL errors.
  • In v0.0.38, apps/web/src/threadSync.ts displays this loading banner when a thread shell exists but its detail data has not loaded.

A direct snapshot request without authentication returned HTTP 401, so it did not test the affected snapshot. Main-app DevTools could not be opened during triage; client console and network evidence are unavailable. These checks do not rule out runtime database or client state failures.

Steps to reproduce

User-observed sequence; no deterministic reproduction established:

  1. Use T3 Code desktop while the filesystem runs out of space.
  2. Open threads. Some remain on “Loading messages…”, while others work.
  3. Free disk space.
  4. Affected threads remain stuck.
  5. Restart T3 Code to restore access, as observed in previous occurrences.

Version

0.0.38, source inspected at tag v0.0.38.

Environment

Linux x64, kernel 6.8.0-138-generic; desktop app with local server. Triage runtime: Node v22.22.0.

Evidence

Screenshot shows an empty conversation with the persistent loading banner. After cleanup, the home filesystem had approximately 26 GB available.

Related issues

PR #5104 describes SQLite I/O failures requiring restart, but was not merged. Issue #10206 describes the same loading symptom from client decoding failures. Neither is a confirmed duplicate.

Fix applied or workaround

No changes, database writes, or restarts were performed during triage. Freeing space alone did not help. Restarting has helped previously.

Filed by

Codex, GPT-6, via t3 triage.

Screenshot

Image

Activity

  1. juliusmarminge commented on Sep 6, 2026

    @juliusmarminge
    Member

    Triage

    Valid desktop/web bug. Not a duplicate of #10206, #9414, or #5104.

    What we believe happened

    The “Loading messages…” pill is apps/web/src/threadSync.ts: a thread shell exists and detail has not loaded. That is independent of disk-full messaging — there is no ENOSPC/SQLITE_FULL UI on this path.

    Reporter notes line up with a stuck client subscribe, not a dead server or unreadable DB:

    • Server stayed up; other threads still loaded.
    • Affected thread had readable SQLite rows, a ready Codex session, and parseable attachments/activity (61 messages / 1,066 activities).
    • Freeing space did not recover; restart did (same as prior occurrences).
    • No ENOSPC / SQLITE_FULL in retained traces, and the authenticated snapshot was never exercised (unauthenticated GET → 401).

    On v0.0.38 (reported; #10216 is not in that tag), subscribeDynamic treats many RPC failures as transport loss and waits for the next session. A local desktop session stays healthy, so the atom can sit in synchronizing with no detail and no error until process restart.

    On current main, #10216 records state.error (Could not synchronize the thread.) and can retry on application-active. Web/desktop still ignore that error: threadSync does not read it, and ThreadErrorBanner uses session.lastError, not the thread atom (errorAtom is unused on web). #10216 already documented that web still shows this pill after a terminated load.

    That also explains partial impact: the shell list is a different subscription. Threads that already had detail, or whose subscribe succeeded after space was freed, keep working. Sidebar detail prewarm can keep the failed atom mounted, so leaving the thread may not retry.

    #5104 is the SQLite durability / “I/O until restart” theme only (closed, unmerged). A poisoned shared connection would usually break every thread; that does not match this report. #1231/#961 are startup-corruption recovery, not this live stuck-detail case.

    Suggested direction

    1. Web/desktop: treat EnvironmentThreadState.error as a failed load (hide the loading pill; show the existing error banner or equivalent). Mirror what mobile does after fix(client-runtime): report terminated thread loads #10216.
    2. Retry a failed snapshot while the local session is still up (visibility wakeup, remount, or an explicit retry). Do not require a full restart after space is available.
    3. Optional later: map ENOSPC / SQLITE_FULL to an actionable storage error. Not required to stop the infinite-loading lie.

    Workaround

    Restart T3 Code (known). Before the next restart, worth trying hide/show the window or leaving the thread long enough to unmount — sidebar prewarm may still hold the failed subscribe.

    Next occurrence (if it happens again)

    Before restart: DevTools on GET /api/orchestration/threads/:id and console lines Could not load the thread snapshot over HTTP / Durable RPC subscription lost its transport. No thread contents needed.

    Related: #9414 (iOS; web presentation called out as unfixed), #10206 (mobile decode; closed duplicate of #9414), #10216 (runtime error reporting; not in v0.0.38; does not fix web), #5104 (SQLite durability only).

  2. added
    via-triageFiled through npx t3 triage
    bugSomething is broken or behaving incorrectly.
    on Sep 6, 2026
  3. dyc3 commented on Sep 9, 2026

    @dyc3
    Author

    Recurred on v0.0.40, Linux desktop with a local server, while the filesystem had approximately 38 GB free.

    An affected thread remained on “Loading messages…”. Read-only database checks found 5 messages, 82 activities, valid attachment/activity JSON, and a ready Codex session with no recorded session error. Its server-side updates succeeded.

    DevTools showed repeated environment-data:vcs:refresh-status defects:

    InvalidStateError: Failed to execute 'transaction' on 'IDBDatabase': The database connection is closing.
    

    It also showed “Could not persist environment shell cache.” with ConnectionPersistenceError.

    Source inspection at v0.0.40 found that thread caching and Git cache operations share an IndexedDB connection, with no reopening mechanism in that storage layer. Synchronous transaction exceptions can bypass the thread loader’s ordinary cache-error fallback. This is a plausible explanation for the stuck loading state, but the original connection-close trigger remains unknown.

    Confirmed workaround: pressing Ctrl+R in the main app restored thread loading without restarting the backend.

    This recurrence provides evidence of a client IndexedDB failure with disk space available.

    Investigated with Codex, GPT-6, via t3 triage.

    Image
  4. MaximilianMauroner commented on Sep 16, 2026

    @MaximilianMauroner

    Confirmed a related desktop occurrence on macOS with 0.0.43-nightly.20260916.1811.

    The initial thread snapshot returned HTTP 400:

    SchemaError: Expected JSON value at ["thread"]["activities"][134]["payload"]
    

    The offending historical activity was a failed question tool call whose stored input contained a second question without its required question field. After repairing the projected activity and canonical event, the authenticated snapshot returned HTTP 200. The desktop still showed “Syncing messages…” until only that thread’s record was removed from IndexedDB:

    Database: t3code:connection-runtime
    Store: thread
    Key: <environmentId>:<threadId>
    

    This restored the thread history, but it was not a complete recovery. The fresh snapshot exposed an unresolved user-input.requested activity whose provider session had stopped. Answering it then triggered #5454.

    This provides evidence for two separate persistent states:

    1. A failed snapshot can leave the per-thread IndexedDB state unable to recover after the server data becomes valid.
    2. Reloading from a clean snapshot can reveal an orphaned provider callback that still blocks the composer.

    Per-thread cache eviction is therefore only a recovery for the snapshot path. It does not make the thread usable when a pending provider request outlives its session.

    Investigated with GPT-5.6 Sol via Codex in T3 Code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions