Repository navigation
[Bug]: Server: service self-update stops the server before checking the database backup fits #14515
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Confirmed on
mainat0cf482b08bc. When the background service updates itself, it stops the server before it knows whether the database backup will fit. If the backup then fails, you get an outage plus a cold restart. I didn't find an open duplicate.#handleUpdateRequestaccepts the update without checking free space, waits 2 seconds, and then#beginTrialstops the running child. Only after that does#startTrialcallbackupDatabaseOnce, which copiesdbPath,-wal, and-shmwithfs.promises.copyFileand no clone flag. If that throws, the update is recorded asfailed/db-backup-failedand#returnToPreviousstarts the previous version again, which matches yourservice-state.json. The copy goes into a staging directory and is only renamed into place at the end, so a failed copy leaves the live database untouched. The damage is the stop plus the cold start.Stopping first is deliberate under the current rollback rule. The snapshot is taken after the old child exits so the three SQLite files aren't changing during the copy (see
docs/internals/server-updates.md). Rejections inside#handleUpdateRequestgo out asupdate-rejectedwhile the server is still up, and the client shows that reason. The backup runs afterupdate-accepted, though, so the client only learns it failed by reconnecting to a server that reports the failure. A cold start on a database this large can also take longer than the client's 4-minute update resume window, which fits the long reconnect loop you saw.fs.copyFilewithCOPYFILE_FICLONEdoesn't clone on macOS. libuv implements those flags with the LinuxFICLONEioctl. On Darwin, the force flag returnsENOSYSand the plain flag falls back to a fullsendfilecopy, which matches your measurement.cp -cusesclonefile, and Node never calls that. libuv used to clone on APFS but removed it after macOS kernel bugs (libuv#3987), so a clone-based backup would need its own Darwin code path with a fallback.t3 updatedoes skip this snapshot. With a restart, it stops the service, rewrites the service state to the new version, and starts it throughBootService.install, which never callsbackupDatabaseOnce. That avoids this outage, but it also skips the rollback copy, so a bad migration on the new version can't be restored the way a failed trial is.Closed PR #8431 (not merged) proposed
VACUUM INTOfor this snapshot and was closed for missing verification. It would still write a full extra database file, so it wouldn't remove the free-space requirement. It would also pull SQLite into the launcher, which only uses Node built-ins.A fix could reject the update in
#handleUpdateRequestwhen there isn't enough free space for a full copy, and clone on filesystems that support it so the backup doesn't need a second full copy's worth of space. Whatever changes, the backup still needs to capture all three SQLite files while the server is stopped, or replace that rule with something equally safe. Nothing more is needed from you on this report.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 1, 2026 An option to investigate, not a proven replacement for the stopped-server copy. On a live 22 GB V2 database (build 0.0.46-nightly.20261010.2922, Linux, NVMe, btrfs), with the server up, we copied from a read-only connection in one read transaction (
BEGIN, one read, then SQLite 3.51.2's.backup --asyncinto a staging file). The copy took 20.8 s with noserver.eventLoop.stallduring it or 120 s after, andintegrity_checkreturned ok. The 20.8 s measures copying with destination journaling and fsync disabled; durable publication happened later.
Without a held read snapshot, stepped backups can repeatedly restart under an external writer. With a writer committing every 20 ms, the CLI's.backupand Python'sbackup(pages=100)were still running at 30 s; one-stepbackup(pages=-1)andVACUUM INTOfinished in 0.1-0.2 s on the same approximately 234 MiB synthetic test database. Node 24'snode:sqlite.backup()defaults to 100-page batches and documents restarts on writes from other connections, so an unpinned backup has the same restart risk; we did not benchmark Node here.
This demonstrates a consistent online point-in-time backup. Using it for update rollback still needs a way to preserve every commit through old-server shutdown; otherwise rollback drops writes made after the snapshot. SQLite integration and #14520's free-space preflight also remain necessary.Drafted with Claude Opus 5.5 in Claude Code (T3 Code harness).
🤖 Generated with Claude Code
Before submitting
Area
apps/server
Steps to reproduce
t3 service install) on a machine whose free disk space is smaller than the server database. In my case:statev2.sqlitewas 32.8 GB and the volume had 24 GB free.server.updateServerWithProgress,serverSelfUpdate: "boot-service").Expected behavior
The update is refused up front with a reason the client can show ("not enough free disk space to back up the database"), and the running server is left alone. Or the backup takes no extra space where the filesystem supports clones (APFS, btrfs, XFS).
Actual behavior
The launcher stops the server first and only then runs
backupDatabaseOnce, which copies the whole database withfs.copyFile. The copy fails for lack of space, the update is recorded as failed and rolled back, and the old server restarts cold. Every attempt is an outage plus a cold restart. With a large database the cold server then took long enough to answer that clients reconnect-looped for about 10 minutes.Even with enough space, the full copy of a tens-of-GB database is slow and briefly needs double the space. On macOS,
fs.copyFilewithCOPYFILE_FICLONEdoes not clone: it falls back to a full copy (1 GiB took 805 ms and consumed 1 GiB), andCOPYFILE_FICLONE_FORCEfails with ENOSYS.cp -cclones on APFS (a 32 MiB copy changed free space by 176 KiB).Impact
Major degradation or frequent failure
Version or commit
Seen on a 0.0.42 nightly build; the launcher code is unchanged on upstream/main @ 0cf482b (
apps/server/src/serviceLauncher.ts,backupDatabaseOnce,#handleUpdateRequest).Environment
macOS 26 (Darwin 25.5), APFS, background service via launchd, desktop client connected to the service.
Logs or stack traces
~/.t3/runtime/service-state.jsonafter the attempt:Workaround
Free disk space larger than the database before updating, or update with
t3 update <version>from the CLI, which does not take the launcher's backup step.