Skip to content

Persisted state corruption can leave T3 Code unusable with no clear recovery path #961

Description

@copypasteitworks

Summary

Persisted state corruption can leave T3 Code unusable with no clear recovery path.

This should be treated as a separate reliability track, distinct from the already-open diff and version fixes.

Symptom

  • the app can become broken or unusable on startup
  • current bootstrap behavior does not clearly explain that persisted state load failed
  • one malformed persisted runtime row can poison list-based recovery paths
  • some Codex resume failures caused by state corruption are still treated as fatal even though a fresh thread start would be an acceptable fallback

Known workaround

In at least one repro, renaming state.sqlite allowed the app to recover.

That suggests persisted corruption can poison startup, but the mental model here should be broader than a single DB file.

Important

For this track, state should include all persisted surfaces that can poison startup or recovery:

  • T3CODE_STATE_DIR
  • state.sqlite
  • files under the state dir such as keybindings, logs, attachments, and runtime metadata
  • CODEX_HOME
  • persisted browser/Electron client state

Activity

  1. eggfriedrice24 commented on Mar 20, 2026

    @eggfriedrice24
    Contributor

    Working on a server side fix for the SQLite corruption crash loop.

    Root cause: NodeSqliteClient.ts calls openDatabase() synchronously with no error handling. When state.sqlite is corrupted, new DatabaseSync() throws, becoming an unhandled defect that crashes the process. The desktop app then restarts the backend in a loop, never recovering.

    Repro (manual):

    # write garbage into the database file
    echo "not a db" > ~/.t3/userdata/state.sqlite
    # start the app, observe crash loop:
    # "Error: file is not a database" repeated every ~500ms
    t3code

    Fix will:

    • Validate the SQLite file header before opening (valid files start with SQLite format 3\0, or the file doesn't exist yet)
    • If corrupted, back up the bad file (.corrupted.<timestamp>) and let a fresh database be created
    • Wrap openDatabase() in Effect.try so open failures produce typed errors instead of process crashes
    • Log a clear warning when corruption is detected and auto recovered

    Note: PR #964 (draft) addresses the web/UI recovery view for bootstrap snapshot failures. This fix is complementary, targeting the server side crash loop so the backend never gets stuck.

  2. israelvf commented on Mar 27, 2026

    @israelvf

    Hey folks! I was experiencing the exact same infinite backend crash loop on startup (code=1 signal=null), but after doing a deep dive, I found that my issue was not related to SQLite corruption (a clean wipe of ~/.t3 and ~/.config/t3code didn't help).

    It turns out it's an unhandled ENOTDIR exception caused by the Claude CLI path resolution.

    The Root Cause

    Using strace, I caught the backend dying because it attempts to execute claude inside the Go compiler binary:

    [pid 1036996] execve("/usr/bin/go/claude", ["claude", "--version"], 0x11b00010ce00) = -1 ENOTDIR (Not a directory)
    

    Because /usr/bin/go is a binary file (not a directory) on my system (Arch Linux/CachyOS), the OS throws an ENOTDIR error. The backend doesn't seem to catch this specific spawn error, which immediately crashes the node process:

    ERROR (#10): Error: spawn ENOTDIR
        at ChildProcess.spawn (node:internal/child_process:421:11)
        at Module.spawn (node:child_process:810:9)
        at MixedScheduler.<anonymous> (file:///opt/t3code-bin/resources/app.asar/node_modules/@effect/platform-node-shared/dist/NodeChildProcessSpawner.js:228:37)
    

    The Workaround

    I managed to bypass the crash loop completely by installing the Claude CLI. Since the app now successfully finds the valid binary earlier in its search, it never reaches the buggy /usr/bin/go/claude check, and the backend survives.

    Suggested Fix

    It looks like the path resolution logic (perhaps iterating over $PATH) is blindly appending /claude to paths that are actually files instead of directories. Wrapping the spawn or path check in a try/catch that handles ENOTDIR should prevent the entire backend from going down when hunting for the agent. Hope this helps!

  3. Zedyas commented on Sep 25, 2026

    @Zedyas

    Another occurrence on 0.0.43-nightly.20260924.2213 (macOS). The corruption came from a hard shutdown, which is filed separately as #13544. This comment covers two things: why one small corrupted region made the whole app unusable, and a recovery procedure that worked and lost only the damaged rows.

    How a small corrupted region blocked startup

    The damage was 377 zero-filled pages at the end of an 835 MB file. They held 23 events (seq 85017–85039, a 30 s window of one thread), 1 message row, and 16 activity rows. Everything else was readable.

    projection_state at the time:

    projection.projects, .thread-messages, .thread-activities, .threads, ... (9 projectors)   85055
    projection.attachment-cleanup                                                             80614
    
    1. The 9 regular projectors were already at the last event, so their bootstrap read nothing.
    2. projection.attachment-cleanup was ~4,440 events behind. Its cursor is only written at the end of bootstrap() (ProjectionPipeline.ts L2135-L2186; the projector name appears nowhere else in the file). So it always points at the last event that existed at the previous launch.
    3. Every start therefore replays all events since the previous launch through readFromSequence(cleanupStart, Number.MAX_SAFE_INTEGER), only to collect thread.reverted and thread.deleted events.
    4. That scan reached the zeroed pages, failed with PersistenceSqlError ... SQLITE(11), and the server exited with code 1.
    5. The desktop app restarted it 124 times over 92 minutes, ~40 s apart. No window was created, because window creation waits for backend readiness ([Bug]: Backend startup crashes leave desktop blank with no visible error #10517; fix(desktop): show startup failures instead of a blank window #10523 would show a dialog after 5 failures).

    After salvaging the file with .recover, startup failed a second way:

    PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows: Composite(Pointer(Composite(Pointer(Encoding(InvalidValue)))))
      [cause]: SchemaError: Expected a valid JSON string
        at [402]["payload"]
    

    One row at the edge of the damaged region came back with payload_json = '' and metadata_json = ''. So a single undecodable row anywhere between the cleanup cursor and the end of the log is also fatal to startup.

    Suggestions

    1. Make attachment cleanup non-fatal in bootstrap(). It is best-effort file cleanup, and none of the other projectors needed the damaged rows. On a SQL or decode error: log it, leave the cursor where it is, and continue startup.
    2. Advance the cleanup cursor while events are applied at runtime, or bound the replay. Then a start does not re-read everything written since the previous launch. That also removes the startup cost of replaying thousands of events.
    3. On SQLITE_CORRUPT during startup, run PRAGMA quick_check and put the result and the recovery steps below into the fix(desktop): show startup failures instead of a blank window #10523 dialog, instead of a generic crash.
    4. Stop tracing each readiness probe attempt. Every http.client GET /.well-known/t3/environment attempt (~every 100 ms) is its own span. During the restart loop, desktop.trace.ndjson filled a 10 MB file every ~9 minutes, so the 10-file rotation holds only ~90 minutes of a loop. By the time I recovered, one file of pre-crash desktop history was left. One span per readiness wait would keep that history.

    Recovery procedure that worked

    This worked on the database above (835 MB, 85k events). It needs the sqlite3 CLI, which macOS ships (3.51). Paths are for the default state dir.

    # 1. Quit T3 Code completely (the backend restarts itself while the app runs)
    osascript -e 'quit app "T3 Code (Nightly)"'      # or "T3 Code"
    pgrep -fl 'T3 Code' || echo stopped
    
    cd ~/.t3/userdata
    
    # 2. Back up the database together with its -wal/-shm files, keeping their names
    #    so SQLite still pairs them when reading the backup
    mkdir -p corrupt-backup
    cp -p state.sqlite corrupt-backup/
    for f in state.sqlite-wal state.sqlite-shm; do [ -f "$f" ] && cp -p "$f" corrupt-backup/; done
    
    # 3. Copy every readable row into a new file
    sqlite3 corrupt-backup/state.sqlite ".recover" | sqlite3 recovered.sqlite
    sqlite3 recovered.sqlite "PRAGMA integrity_check;"          # expect: ok
    
    # 4. Find rows that .recover only partly restored (invalid or empty JSON)
    sqlite3 recovered.sqlite "SELECT m.name, p.name FROM sqlite_schema m
      JOIN pragma_table_info(m.name) p WHERE m.type = 'table' AND p.name LIKE '%json%';" |
    while IFS='|' read -r t c; do
      n=$(sqlite3 recovered.sqlite "SELECT count(*) FROM \"$t\" WHERE \"$c\" IS NOT NULL AND json_valid(\"$c\") = 0;")
      [ "$n" != 0 ] && echo "$t.$c: $n"
    done
    #    Inspect, then delete each reported row, e.g.:
    #    sqlite3 recovered.sqlite "DELETE FROM orchestration_events WHERE json_valid(payload_json) = 0;"
    
    # 5. Restore WAL mode and swap the file in. Remove the old -wal/-shm first:
    #    SQLite would otherwise apply the old WAL to the new file.
    sqlite3 recovered.sqlite "PRAGMA journal_mode = WAL;"
    rm -f state.sqlite-wal state.sqlite-shm
    mv recovered.sqlite state.sqlite

    Here, step 4 found 2 rows: orchestration_events seq 85039 and one projection_thread_activities row, both with empty JSON. After deleting them, the app started in 3 s and appended new events immediately.

    Why the event log still works with gaps after recovery:

    • .recover copies sqlite_sequence, so new events continue from 85056 and never reuse lost sequence numbers.
    • readFromSequence pages by sequence > cursor, with the cursor set to the last row read (OrchestrationEventStore.ts L287-L324), so missing sequence numbers are skipped.
    • append computes stream_version as the stream's current max + 1 (L149), so a gap in one stream's versions does not block new appends.

    The server has been running and appending events since the recovery. I did not audit every consumer of stream_version for assumptions that versions are contiguous.


    A note from me: I am not a developer at all, but it felt worth helping out when I ran into an unusual issue and Claude helped me debug it. So apologies if anything here is just flat out wrong. I genuinely don't understand most of the technical detail, so I'm really sorry if I'm wasting your time.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions