Repository navigation
Persisted state corruption can leave T3 Code unusable with no clear recovery path #961
Description
Activity
Working on a server side fix for the SQLite corruption crash loop.
Root cause:
NodeSqliteClient.tscallsopenDatabase()synchronously with no error handling. Whenstate.sqliteis corrupted,new DatabaseSync()throws, becoming an unhandled defect that crashes the process. The desktop app then restarts the backend in a loop, never recovering.Repro (manual):
# write garbage into the database file echo "not a db" > ~/.t3/userdata/state.sqlite # start the app, observe crash loop: # "Error: file is not a database" repeated every ~500ms t3code
Fix will:
- Validate the SQLite file header before opening (valid files start with
SQLite format 3\0, or the file doesn't exist yet) - If corrupted, back up the bad file (
.corrupted.<timestamp>) and let a fresh database be created - Wrap
openDatabase()inEffect.tryso open failures produce typed errors instead of process crashes - Log a clear warning when corruption is detected and auto recovered
Note: PR #964 (draft) addresses the web/UI recovery view for bootstrap snapshot failures. This fix is complementary, targeting the server side crash loop so the backend never gets stuck.
Reacted by shivam and skrow- Validate the SQLite file header before opening (valid files start with
Hey folks! I was experiencing the exact same infinite backend crash loop on startup (
code=1 signal=null), but after doing a deep dive, I found that my issue was not related to SQLite corruption (a clean wipe of~/.t3and~/.config/t3codedidn't help).It turns out it's an unhandled
ENOTDIRexception caused by the Claude CLI path resolution.The Root Cause
Using
strace, I caught the backend dying because it attempts to executeclaudeinside the Go compiler binary:[pid 1036996] execve("/usr/bin/go/claude", ["claude", "--version"], 0x11b00010ce00) = -1 ENOTDIR (Not a directory)Because
/usr/bin/gois a binary file (not a directory) on my system (Arch Linux/CachyOS), the OS throws anENOTDIRerror. The backend doesn't seem to catch this specificspawnerror, which immediately crashes the node process:ERROR (#10): Error: spawn ENOTDIR at ChildProcess.spawn (node:internal/child_process:421:11) at Module.spawn (node:child_process:810:9) at MixedScheduler.<anonymous> (file:///opt/t3code-bin/resources/app.asar/node_modules/@effect/platform-node-shared/dist/NodeChildProcessSpawner.js:228:37)The Workaround
I managed to bypass the crash loop completely by installing the Claude CLI. Since the app now successfully finds the valid binary earlier in its search, it never reaches the buggy
/usr/bin/go/claudecheck, and the backend survives.Suggested Fix
It looks like the path resolution logic (perhaps iterating over
$PATH) is blindly appending/claudeto paths that are actually files instead of directories. Wrapping thespawnor path check in atry/catchthat handlesENOTDIRshould prevent the entire backend from going down when hunting for the agent. Hope this helps!Another occurrence on
0.0.43-nightly.20260924.2213(macOS). The corruption came from a hard shutdown, which is filed separately as #13544. This comment covers two things: why one small corrupted region made the whole app unusable, and a recovery procedure that worked and lost only the damaged rows.How a small corrupted region blocked startup
The damage was 377 zero-filled pages at the end of an 835 MB file. They held 23 events (seq 85017–85039, a 30 s window of one thread), 1 message row, and 16 activity rows. Everything else was readable.
projection_stateat the time:projection.projects, .thread-messages, .thread-activities, .threads, ... (9 projectors) 85055 projection.attachment-cleanup 80614- The 9 regular projectors were already at the last event, so their bootstrap read nothing.
projection.attachment-cleanupwas ~4,440 events behind. Its cursor is only written at the end ofbootstrap()(ProjectionPipeline.tsL2135-L2186; the projector name appears nowhere else in the file). So it always points at the last event that existed at the previous launch.- Every start therefore replays all events since the previous launch through
readFromSequence(cleanupStart, Number.MAX_SAFE_INTEGER), only to collectthread.revertedandthread.deletedevents. - That scan reached the zeroed pages, failed with
PersistenceSqlError ... SQLITE(11), and the server exited with code 1. - The desktop app restarted it 124 times over 92 minutes, ~40 s apart. No window was created, because window creation waits for backend readiness ([Bug]: Backend startup crashes leave desktop blank with no visible error #10517; fix(desktop): show startup failures instead of a blank window #10523 would show a dialog after 5 failures).
After salvaging the file with
.recover, startup failed a second way:PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows: Composite(Pointer(Composite(Pointer(Encoding(InvalidValue))))) [cause]: SchemaError: Expected a valid JSON string at [402]["payload"]One row at the edge of the damaged region came back with
payload_json = ''andmetadata_json = ''. So a single undecodable row anywhere between the cleanup cursor and the end of the log is also fatal to startup.Suggestions
- Make attachment cleanup non-fatal in
bootstrap(). It is best-effort file cleanup, and none of the other projectors needed the damaged rows. On a SQL or decode error: log it, leave the cursor where it is, and continue startup. - Advance the cleanup cursor while events are applied at runtime, or bound the replay. Then a start does not re-read everything written since the previous launch. That also removes the startup cost of replaying thousands of events.
- On
SQLITE_CORRUPTduring startup, runPRAGMA quick_checkand put the result and the recovery steps below into the fix(desktop): show startup failures instead of a blank window #10523 dialog, instead of a generic crash. - Stop tracing each readiness probe attempt. Every
http.client GET /.well-known/t3/environmentattempt (~every 100 ms) is its own span. During the restart loop,desktop.trace.ndjsonfilled a 10 MB file every ~9 minutes, so the 10-file rotation holds only ~90 minutes of a loop. By the time I recovered, one file of pre-crash desktop history was left. One span per readiness wait would keep that history.
Recovery procedure that worked
This worked on the database above (835 MB, 85k events). It needs the
sqlite3CLI, which macOS ships (3.51). Paths are for the default state dir.# 1. Quit T3 Code completely (the backend restarts itself while the app runs) osascript -e 'quit app "T3 Code (Nightly)"' # or "T3 Code" pgrep -fl 'T3 Code' || echo stopped cd ~/.t3/userdata # 2. Back up the database together with its -wal/-shm files, keeping their names # so SQLite still pairs them when reading the backup mkdir -p corrupt-backup cp -p state.sqlite corrupt-backup/ for f in state.sqlite-wal state.sqlite-shm; do [ -f "$f" ] && cp -p "$f" corrupt-backup/; done # 3. Copy every readable row into a new file sqlite3 corrupt-backup/state.sqlite ".recover" | sqlite3 recovered.sqlite sqlite3 recovered.sqlite "PRAGMA integrity_check;" # expect: ok # 4. Find rows that .recover only partly restored (invalid or empty JSON) sqlite3 recovered.sqlite "SELECT m.name, p.name FROM sqlite_schema m JOIN pragma_table_info(m.name) p WHERE m.type = 'table' AND p.name LIKE '%json%';" | while IFS='|' read -r t c; do n=$(sqlite3 recovered.sqlite "SELECT count(*) FROM \"$t\" WHERE \"$c\" IS NOT NULL AND json_valid(\"$c\") = 0;") [ "$n" != 0 ] && echo "$t.$c: $n" done # Inspect, then delete each reported row, e.g.: # sqlite3 recovered.sqlite "DELETE FROM orchestration_events WHERE json_valid(payload_json) = 0;" # 5. Restore WAL mode and swap the file in. Remove the old -wal/-shm first: # SQLite would otherwise apply the old WAL to the new file. sqlite3 recovered.sqlite "PRAGMA journal_mode = WAL;" rm -f state.sqlite-wal state.sqlite-shm mv recovered.sqlite state.sqlite
Here, step 4 found 2 rows:
orchestration_eventsseq 85039 and oneprojection_thread_activitiesrow, both with empty JSON. After deleting them, the app started in 3 s and appended new events immediately.Why the event log still works with gaps after recovery:
.recovercopiessqlite_sequence, so new events continue from 85056 and never reuse lost sequence numbers.readFromSequencepages bysequence > cursor, with the cursor set to the last row read (OrchestrationEventStore.tsL287-L324), so missing sequence numbers are skipped.appendcomputesstream_versionas the stream's current max + 1 (L149), so a gap in one stream's versions does not block new appends.
The server has been running and appending events since the recovery. I did not audit every consumer of
stream_versionfor assumptions that versions are contiguous.
A note from me: I am not a developer at all, but it felt worth helping out when I ran into an unusual issue and Claude helped me debug it. So apologies if anything here is just flat out wrong. I genuinely don't understand most of the technical detail, so I'm really sorry if I'm wasting your time.
Summary
Persisted state corruption can leave T3 Code unusable with no clear recovery path.
This should be treated as a separate reliability track, distinct from the already-open diff and version fixes.
Symptom
Known workaround
In at least one repro, renaming
state.sqliteallowed the app to recover.That suggests persisted corruption can poison startup, but the mental model here should be broader than a single DB file.
Important
For this track,
stateshould include all persisted surfaces that can poison startup or recovery:T3CODE_STATE_DIRstate.sqliteCODEX_HOME