recoverInterruptedDrop finishes a crashed drop_database by deleting the database directory and every blob root with a synchronous, recursive rmSync. It is reached from two places, and neither can await:
databasesBlockedByLifecycle (resources/databases.ts), inside getDatabases() — which runs on every thread, on every schema event, and at worker boot.
throwIfBlockedByRestore (resources/databases.ts), on an on-demand open — i.e. inside a request.
So after a crash mid-drop of a large database, the first thread to win the lifecycle lock blocks its event loop for the whole removal: on an HTTP worker that stalls every in-flight request, at boot it stalls startup, and nothing is logged until it finishes.
The same module already has the async, one-entry-at-a-time removeSteadily for exactly this reason — the online drop path uses it. The recovery path cannot, because getDatabases() is synchronous all the way down.
Why this is a design task, not a narrow fix
Making the recovery asynchronous means changing getDatabases()'s synchronous contract across its callers, which is a real API-surface change rather than swapping one call. The shape the review converged on:
- Mark the database blocked synchronously (the lifecycle marker already does this — every rescan skips it and every on-demand open refuses it).
- Enqueue the deletion to a single owner rather than running it on whichever thread scanned first.
- Signal reload only once that completes.
The marker keeps the database unavailable throughout, so nothing observes a half-deleted database while the asynchronous cleanup runs — which is what makes the split safe.
Provenance
Raised independently by three reviewers across rounds 13, 39, 40, 42 and 43 of the pre-push review on Stop a recycled Windows PID from wedging deploy_component and release dropped databases on every thread, and named as alternative (b) by that PR's --mode plan framing recheck. @kriszyp ruled it out of scope for #2470 and into its own issue; alternative (a), recording the drop's deletion targets in its marker, landed there.
The synchronous recovery is not a regression — it is how the drop protocol is written in #2470 — but it is the one part of it that puts a destructive filesystem operation on the request and rescan paths.
recoverInterruptedDropfinishes a crasheddrop_databaseby deleting the database directory and every blob root with a synchronous, recursivermSync. It is reached from two places, and neither can await:databasesBlockedByLifecycle(resources/databases.ts), insidegetDatabases()— which runs on every thread, on every schema event, and at worker boot.throwIfBlockedByRestore(resources/databases.ts), on an on-demand open — i.e. inside a request.So after a crash mid-drop of a large database, the first thread to win the lifecycle lock blocks its event loop for the whole removal: on an HTTP worker that stalls every in-flight request, at boot it stalls startup, and nothing is logged until it finishes.
The same module already has the async, one-entry-at-a-time
removeSteadilyfor exactly this reason — the online drop path uses it. The recovery path cannot, becausegetDatabases()is synchronous all the way down.Why this is a design task, not a narrow fix
Making the recovery asynchronous means changing
getDatabases()'s synchronous contract across its callers, which is a real API-surface change rather than swapping one call. The shape the review converged on:The marker keeps the database unavailable throughout, so nothing observes a half-deleted database while the asynchronous cleanup runs — which is what makes the split safe.
Provenance
Raised independently by three reviewers across rounds 13, 39, 40, 42 and 43 of the pre-push review on Stop a recycled Windows PID from wedging deploy_component and release dropped databases on every thread, and named as alternative (b) by that PR's
--mode planframing recheck. @kriszyp ruled it out of scope for #2470 and into its own issue; alternative (a), recording the drop's deletion targets in its marker, landed there.The synchronous recovery is not a regression — it is how the drop protocol is written in #2470 — but it is the one part of it that puts a destructive filesystem operation on the request and rescan paths.