Skip to content

Interrupted-drop recovery blocks the event loop: rmSync on the schema-rescan and on-demand-open paths #2598

Description

@kriszyp

recoverInterruptedDrop finishes a crashed drop_database by deleting the database directory and every blob root with a synchronous, recursive rmSync. It is reached from two places, and neither can await:

  • databasesBlockedByLifecycle (resources/databases.ts), inside getDatabases() — which runs on every thread, on every schema event, and at worker boot.
  • throwIfBlockedByRestore (resources/databases.ts), on an on-demand open — i.e. inside a request.

So after a crash mid-drop of a large database, the first thread to win the lifecycle lock blocks its event loop for the whole removal: on an HTTP worker that stalls every in-flight request, at boot it stalls startup, and nothing is logged until it finishes.

The same module already has the async, one-entry-at-a-time removeSteadily for exactly this reason — the online drop path uses it. The recovery path cannot, because getDatabases() is synchronous all the way down.

Why this is a design task, not a narrow fix

Making the recovery asynchronous means changing getDatabases()'s synchronous contract across its callers, which is a real API-surface change rather than swapping one call. The shape the review converged on:

  1. Mark the database blocked synchronously (the lifecycle marker already does this — every rescan skips it and every on-demand open refuses it).
  2. Enqueue the deletion to a single owner rather than running it on whichever thread scanned first.
  3. Signal reload only once that completes.

The marker keeps the database unavailable throughout, so nothing observes a half-deleted database while the asynchronous cleanup runs — which is what makes the split safe.

Provenance

Raised independently by three reviewers across rounds 13, 39, 40, 42 and 43 of the pre-push review on Stop a recycled Windows PID from wedging deploy_component and release dropped databases on every thread, and named as alternative (b) by that PR's --mode plan framing recheck. @kriszyp ruled it out of scope for #2470 and into its own issue; alternative (a), recording the drop's deletion targets in its marker, landed there.

The synchronous recovery is not a regression — it is how the drop protocol is written in #2470 — but it is the one part of it that puts a destructive filesystem operation on the request and rescan paths.

Activity

  1. added theissue type on Sep 14, 2026
  2. kriszyp commented on Oct 10, 2026

    @kriszyp
    MemberAuthor

    Closing: recoverInterruptedDrop and the synchronous rmSync recovery path never reached main. That design lived only in #2470, closed unmerged, and the drop lifecycle was since reworked (#2604, #2962, #3119). Reopen if the current drop path has an equivalent sync delete.

    🤖 Claude (Anthropic, Opus 5.5), posted via @kriszyp during an issue grouping pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions