Repository navigation
[Bug]: Hard shutdown on macOS corrupts state.sqlite: SQLite never uses F_FULLFSYNC #13544
Description
Activity
Decision
Valid bug. Fix on
main. Not a duplicate. The macOS desktop server can corruptstate.sqliteon an unclean power loss because WAL checkpoints sync withfsync(), which does not flush the drive cache on macOS.PRAGMA checkpoint_fullfsyncis off. Currentmainstill has that configuration.Severity: high for the macOS desktop app. A corrupt
state.sqliteexits the server on every boot, and the shell restarts it indefinitely, so the window never appears. This needs a power loss or hard reset during a checkpoint. A process kill does not reproduce it.Confidence: high on the missing barrier and on the page-level signature. Medium on this specific shutdown, because the damaged file and WAL are not attached and the cause of the hard stop is unknown.
What matches the code
apps/server/src/persistence/Layers/Sqlite.tsis the only production setup forstate.sqlite. It runs on the singlenode:sqliteconnection used by the desktop server and the CLI:const setup = Layer.effectDiscard( Effect.gen(function* () { const sql = yield* SqlClient.SqlClient; // CLI and server write from separate processes; wait rather than fail with SQLITE_BUSY. yield* sql`PRAGMA busy_timeout = 5000;`; yield* sql`PRAGMA foreign_keys = ON;`; yield* sql`PRAGMA journal_mode = WAL;`; yield* runMigrations(); }), );
There is no
fullfsyncorcheckpoint_fullfsyncanywhere in the repo. Those pragmas are per connection and are not stored in the file, so every writer must set them on open.On this machine,
node:sqlite(SQLite 3.47.2) already reportssynchronous = 2(FULL) before any T3 pragma, and it stays 2 afterjournal_mode = WAL.fullfsyncandcheckpoint_fullfsyncboth read back as 0, andcheckpoint_fullfsync = ONreads back as 1. That matches the reporter's measurement on the nightly binary (SQLite 3.53.4, Electron 44.4.2): FULL is already the effective setting, and neither fullfsync flag is on. Closed PR #5104 (synchronous = FULL) would not have changed this runtime.SQLite's own corruption notes say that in WAL mode a lying sync corrupts the database only during a checkpoint. A commit-time sync failure loses recent transactions. The reported damage is that checkpoint shape: the header page count matches the file size, interior pages point at overflow pages, and those pages are 4096 zero bytes. The ~92-page zero runs match one ~310 KB activity payload. WAL replay would have repaired that if a valid WAL copy had survived the checkpoint. macOS documents that
fsync()does not provide that ordering and that databases needF_FULLFSYNC.PRAGMA checkpoint_fullfsyncis the SQLite switch that usesF_FULLFSYNCfor checkpoint syncs. Upstreamnode:sqliteis not Apple's libsqlite, so this pragma is the realF_FULLFSYNCpath, not Apple'sF_BARRIERFSYNCsubstitute.The startup failure matches the source.
OrchestrationEventStore.readFromSequenceis what the log names (OrchestrationEventStore.readFromSequence:query), and a malformed database fails that read. The desktop shell then restarts the child while it is supposed to be running. Backoff caps at 10 seconds (DesktopBackendManager.ts); the ~40 second spacing in the trace is failed startup plus that cap. Nothing stops the loop forSQLITE_CORRUPT.Version
0.0.43-nightly.20260924.2213and currentmainuse the same three pragmas. No newer commit fixes this.What the evidence does not prove
The hard stop itself is unexplained: logging ends at 00:23:52 UTC, boot is 01:00:56 UTC, with no panic report and no shutdown record. That is consistent with the drive cache being discarded, and it is not a captured kernel crash. The zero-page signature is the evidence for write reordering. A clean
kill -9would not do this, and the report is right about that.checkpoint_fullfsync = ONmakes the checkpoint barrier real. Commits since the previous checkpoint can still disappear, because ordinary commitfsync()still does not flush the drive. SQLite checksums turn that into a clean rollback. That is the outcome the report asks for.fullfsync = ONwould also flush every commit, at the reported ~3.1 ms per commit on the server's single synchronous connection. The integrity fix does not require that.The benchmark was not rerun here (Linux VM, no Electron nightly, no APFS SSD). The reported shape fits the mechanism: ~0.04 ms p50 with FULL and no fullfsync means the drive cache was not flushed;
checkpoint_fullfsyncstays in the thousands of commits per second on that internal SSD, against a reported peak of about 2.9 events per second. Treat those numbers as one-machine measurements. The cost lands on checkpoint, and one large image payload is on the order of 90 pages of the default 1000-page autocheckpoint.Related issues
- [Bug]: Overlapping server lifecycles during a restart leave state.sqlite malformed and make the app unusable #11084 is a different trigger. These traces show one server lifecycle and no overlapping bootstrap.
- fix(server): harden SQLite durability under streaming writes #5104 was closed asking for a disk-backed cost. This report supplies that cost. The missing setting is
checkpoint_fullfsync, andsynchronousis already FULL. - Persisted state corruption can leave T3 Code unusable with no clear recovery path #961, [Bug]: Backend startup crashes leave desktop blank with no visible error #10517, and fix(desktop): show startup failures instead of a blank window #10523 are the reason a corrupt file becomes a blank window. This pragma does not add a recovery path. Already-damaged databases still need the manual salvage described on Persisted state corruption can leave T3 Code unusable with no clear recovery path #961.
- fix(server): snapshot service-update databases with VACUUM INTO #8431 is backups, not this failure.
SQLite 3.53.4 is past the WAL-reset race fixed in 3.51.3, and the single-writer trace does not look like that race.
Fix
Add this to the shared setup in
Sqlite.ts, immediately after WAL mode, with no platform check. The pragma is a no-op whereF_FULLFSYNCdoes not exist:yield* sql`PRAGMA journal_mode = WAL;`; // macOS fsync() does not flush the drive cache. Checkpoint syncs must use // F_FULLFSYNC so a power loss cannot publish page references without the pages. yield* sql`PRAGMA checkpoint_fullfsync = ON;`;
Extend
Sqlite.test.tssoPRAGMA checkpoint_fullfsyncreads back as 1. That locks the setting. It cannot simulate power loss in CI.Also set the same pragma on the direct
NodeSqliteClientconnection inapps/server/scripts/t3-sqlite-state.tsbeforeVACUUM INTOor other writes. That script bypassesSqlite.tsand can checkpointstate.sqlite.migrate-dev-db.tsis dev-only; same gap if it is considered a writer.Leave
fullfsyncoff unless per-commit durability is an explicit goal. The reported workload does not require the ~28× commit slowdown.apps/mobile/src/persistence/mobile-database.tsis a separate Expo database that also enables WAL without this pragma. Worth a follow-up on iOS. It is not the desktopstate.sqlitefailure.- addedvia-triageFiled through npx t3 triageFiled through npx t3 triagebugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request accepted
on Sep 25, 2026
Before submitting
Area
apps/server
Summary
A hard shutdown on macOS (no clean shutdown, no kernel panic report) left
state.sqlitewith zero-filled pages insideorchestration_events,projection_thread_messages, andprojection_thread_activities. Every server start afterwards failed withSQLITE(11) database disk image is malformed, and the desktop app never opened a window.SQLite in WAL mode is designed to stay consistent through power loss, but only if its syncs reach stable storage in order. On macOS,
fsync()does not flush the drive's write cache. Onlyfcntl(F_FULLFSYNC)does, and SQLite uses it only whenPRAGMA fullfsyncorPRAGMA checkpoint_fullfsyncis on. T3 Code sets neither (Sqlite.tsL15-L17), and both are0in the shipped runtime.This is not #11084. The server traces show a single server lifecycle up to the shutdown (details below).
Steps to reproduce
Not deterministic, because it depends on losing writes that are still in the drive's cache:
thread.activity-appendedevents).kill -9of the server process does not reproduce this. The kernel page cache survives process death, so every completedwrite()still reaches the drive. Only losing the drive's volatile cache does. SQLite's own crash testing simulates power loss at the VFS layer for this reason.Expected behavior
After power loss the database is consistent. At most the last few commits are missing.
Actual behavior
B-tree pages in the file point at pages that contain only zeros. The server exits with code 1 on every start. The desktop app restarted it 124 times over 92 minutes and never created a window. Startup-path details and a working recovery procedure are in my comment on #961.
Evidence
Timeline (UTC, 2026-09-25)
desktop.trace.ndjsonorchestration_eventskern.boottimelasthas no shutdown record. No.panicreport.pmset -g loghas no sleep or shutdown entry; battery was at 93%.server-child.logSQLITE(11).desktop.trace.ndjsondesktop.backendInstance.startspans, ~40 s apart, until manual recovery.Server traces covering 23:48:54 – 00:23:46 contain no startup, bootstrap, or migration spans, so only one server lifecycle was writing the database before the shutdown.
PRAGMA integrity_check(complete result, summarized)Tree 3 =
orchestration_events, 27 =projection_thread_messages, 29 =projection_thread_activities.Page contents of the damaged file
btreeInitPage()are 4096 zero bytes.This is a write-ordering failure at the storage level: some writes to the file reached stable storage and earlier or neighboring writes did not. On macOS,
fsync()does not prevent this. Fromman 2 fsync:Effective settings in the shipped runtime
Measured by running
node:sqlitefrom the app binary (ELECTRON_RUN_AS_NODE=1) and applying the same pragmas asSqlite.ts:Context for #5104:
synchronousis already FULL undernode:sqlite, so that PR would not have changed the effective setting. The missing piece isF_FULLFSYNC.Cost of the fix (disk-backed benchmark)
#5104 was closed with a request to "show the streaming-write cost of the chosen setting". Setup:
node:sqlite, SQLite 3.53.4, Electron 44.4.2), database file on the internal APFS SSD.Sqlite.ts, plus the setting under test. Defaultwal_autocheckpoint(1000 pages).orchestration_events-shaped table. 10% are 310 KB payloads and 90% are 500 B. For comparison, in the real database 0.8% of events are over 100 KB, but they hold 60% of payload bytes.synchronous=FULL, no fullfsync)checkpoint_fullfsync = ONfullfsync = ONsynchronous = NORMAL+fullfsync = ONsynchronous=FULLin WAL mode callsfsync()on every commit. A p50 of 0.04 ms for that commit is only possible because the drive cache is not flushed. WithF_FULLFSYNCthe same commit takes ~3.1 ms.DatabaseSyncis synchronous, so the cost is event-loop time. At the busiest minute,fullfsync = ONadds about 171 × 3.1 ms ≈ 0.53 s per minute (<1%), multiplied by the number of commits per event.checkpoint_fullfsync = ONonly adds cost at checkpoints.Benchmark script
Run with the app's runtime:
ELECTRON_RUN_AS_NODE=1 "/Applications/T3 Code (Nightly).app/Contents/MacOS/T3 Code (Nightly)" fsync-bench.cjs <dir>Proposed fix
checkpoint_fullfsync = ONis the minimum for integrity. During a checkpoint, SQLite syncs the WAL before copying pages into the database file, and syncs the database file before the WAL can be reset. Both syncs become real flushes. A power loss can still drop the last few commits since the previous checkpoint. WAL frame checksums detect that cleanly, so it does not corrupt the file.fullfsync = ONalso makes every commit durable, at ~3 ms per commit on the event loop.state.sqlite.F_FULLFSYNC, so no platform check is needed.What I could not verify
checkpoint_fullfsync = ON. The evidence is the write-ordering signature above plus the documented macOSfsync()behavior.Impact
Blocks work completely
Version or commit
0.0.43-nightly.20260924.2213.main@ 568c9bc sets the same three pragmas.Environment
macOS 26.6.2 (25G83), MacBook Pro M3 (Mac15,3), 16 GB, internal SSD (APFS). Electron 44.4.2, Node 24.21.0, SQLite 3.53.4 via
node:sqlite.state.sqlitewas 835 MB.Logs or stack traces
Workaround
Salvage the database with the
sqlite3CLI. Steps are in my comment on #961. Loss in this case: 23 events (seq 85017–85039, a 30 s window of one thread), 1 message row, and 16 activity rows. Every other table was recovered with identical row counts.Additional context
tool.completedactivity payloads (activity.payload.data.result.content[0].source.data, 309 KB average). The zero-filled runs are these payloads' overflow chains. Larger checkpoints mean more data waiting in the drive cache at any moment.synchronous = FULL; already the effective value), fix(server): snapshot service-update databases with VACUUM INTO #8431 (VACUUM INTObackups).A note from me: I am not a developer at all, but it felt worth helping out when I ran into an unusual issue and Claude helped me debug it. So apologies if anything here is just flat out wrong. I genuinely don't understand most of the technical detail, so I'm really sorry if I'm wasting your time.