Skip to content

[Bug]: Server: service self-update stops the server before checking the database backup fits #14515

Description

@saphid

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Run T3 Code as the background service (t3 service install) on a machine whose free disk space is smaller than the server database. In my case: statev2.sqlite was 32.8 GB and the volume had 24 GB free.
  2. From a client, accept the prompt to update the server (server.updateServerWithProgress, serverSelfUpdate: "boot-service").

Expected behavior

The update is refused up front with a reason the client can show ("not enough free disk space to back up the database"), and the running server is left alone. Or the backup takes no extra space where the filesystem supports clones (APFS, btrfs, XFS).

Actual behavior

The launcher stops the server first and only then runs backupDatabaseOnce, which copies the whole database with fs.copyFile. The copy fails for lack of space, the update is recorded as failed and rolled back, and the old server restarts cold. Every attempt is an outage plus a cold restart. With a large database the cold server then took long enough to answer that clients reconnect-looped for about 10 minutes.

Even with enough space, the full copy of a tens-of-GB database is slow and briefly needs double the space. On macOS, fs.copyFile with COPYFILE_FICLONE does not clone: it falls back to a full copy (1 GiB took 805 ms and consumed 1 GiB), and COPYFILE_FICLONE_FORCE fails with ENOSYS. cp -c clones on APFS (a 32 MiB copy changed free space by 176 KiB).

Impact

Major degradation or frequent failure

Version or commit

Seen on a 0.0.42 nightly build; the launcher code is unchanged on upstream/main @ 0cf482b (apps/server/src/serviceLauncher.ts, backupDatabaseOnce, #handleUpdateRequest).

Environment

macOS 26 (Darwin 25.5), APFS, background service via launchd, desktop client connected to the service.

Logs or stack traces

~/.t3/runtime/service-state.json after the attempt:

"status": "failed", "reason": "db-backup-failed"

Workaround

Free disk space larger than the database before updating, or update with t3 update <version> from the CLI, which does not take the launcher's backup step.

Activity

  1. juliusmarminge commented on Oct 1, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Confirmed on main at 0cf482b08bc. When the background service updates itself, it stops the server before it knows whether the database backup will fit. If the backup then fails, you get an outage plus a cold restart. I didn't find an open duplicate.

    #handleUpdateRequest accepts the update without checking free space, waits 2 seconds, and then #beginTrial stops the running child. Only after that does #startTrial call backupDatabaseOnce, which copies dbPath, -wal, and -shm with fs.promises.copyFile and no clone flag. If that throws, the update is recorded as failed / db-backup-failed and #returnToPrevious starts the previous version again, which matches your service-state.json. The copy goes into a staging directory and is only renamed into place at the end, so a failed copy leaves the live database untouched. The damage is the stop plus the cold start.

    Stopping first is deliberate under the current rollback rule. The snapshot is taken after the old child exits so the three SQLite files aren't changing during the copy (see docs/internals/server-updates.md). Rejections inside #handleUpdateRequest go out as update-rejected while the server is still up, and the client shows that reason. The backup runs after update-accepted, though, so the client only learns it failed by reconnecting to a server that reports the failure. A cold start on a database this large can also take longer than the client's 4-minute update resume window, which fits the long reconnect loop you saw.

    fs.copyFile with COPYFILE_FICLONE doesn't clone on macOS. libuv implements those flags with the Linux FICLONE ioctl. On Darwin, the force flag returns ENOSYS and the plain flag falls back to a full sendfile copy, which matches your measurement. cp -c uses clonefile, and Node never calls that. libuv used to clone on APFS but removed it after macOS kernel bugs (libuv#3987), so a clone-based backup would need its own Darwin code path with a fallback.

    t3 update does skip this snapshot. With a restart, it stops the service, rewrites the service state to the new version, and starts it through BootService.install, which never calls backupDatabaseOnce. That avoids this outage, but it also skips the rollback copy, so a bad migration on the new version can't be restored the way a failed trial is.

    Closed PR #8431 (not merged) proposed VACUUM INTO for this snapshot and was closed for missing verification. It would still write a full extra database file, so it wouldn't remove the free-space requirement. It would also pull SQLite into the launcher, which only uses Node built-ins.

    A fix could reject the update in #handleUpdateRequest when there isn't enough free space for a full copy, and clone on filesystems that support it so the backup doesn't need a second full copy's worth of space. Whatever changes, the backup still needs to capture all three SQLite files while the server is stopped, or replace that rule with something equally safe. Nothing more is needed from you on this report.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 1, 2026
  3. only21mil commented on Oct 10, 2026

    @only21mil
    Contributor

    An option to investigate, not a proven replacement for the stopped-server copy. On a live 22 GB V2 database (build 0.0.46-nightly.20261010.2922, Linux, NVMe, btrfs), with the server up, we copied from a read-only connection in one read transaction (BEGIN, one read, then SQLite 3.51.2's .backup --async into a staging file). The copy took 20.8 s with no server.eventLoop.stall during it or 120 s after, and integrity_check returned ok. The 20.8 s measures copying with destination journaling and fsync disabled; durable publication happened later.
    Without a held read snapshot, stepped backups can repeatedly restart under an external writer. With a writer committing every 20 ms, the CLI's .backup and Python's backup(pages=100) were still running at 30 s; one-step backup(pages=-1) and VACUUM INTO finished in 0.1-0.2 s on the same approximately 234 MiB synthetic test database. Node 24's node:sqlite.backup() defaults to 100-page batches and documents restarts on writes from other connections, so an unpinned backup has the same restart risk; we did not benchmark Node here.
    This demonstrates a consistent online point-in-time backup. Using it for update rollback still needs a way to preserve every commit through old-server shutdown; otherwise rollback drops writes made after the snapshot. SQLite integration and #14520's free-space preflight also remain necessary.

    Drafted with Claude Opus 5.5 in Claude Code (T3 Code harness).

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions