Summary
When a write fails at the filesystem level, RocksDB latches a background error and refuses all further writes. There is currently no way for a consumer to recover in-process — no exposed DB::Resume(), and no way to observe the latched state. The only remedy is restarting the process, so a transient, self-correcting condition becomes an indefinite outage.
Observed
An instance exhausted its filesystem quota. RocksDB recorded exactly one background error and stopped accepting writes:
rocksdb.error.handler.bg.error.count COUNT : 1
rocksdb.error.handler.bg.io.error.count COUNT : 1
rocksdb.error.handler.bg.retryable.io.error.count COUNT : 0
rocksdb.error.handler.autoresume.count COUNT : 0
rocksdb.error.handler.autoresume.success.count COUNT : 0
Note bg.retryable.io.error.count: 0 — the error was classified non-retryable, so RocksDB's own auto-resume path was never eligible. That is arguably correct classification; the problem is what follows.
Every subsequent write failed with a stored, stale message naming a WAL file that no longer existed on disk (the newest WAL present was several generations older). This repeated 19,149 times over 7h45m. Freeing space did not help — the quota was more than doubled and writes kept failing with the identical message. Only reopening the database recovered it.
Two consequences worth separating:
- No in-process recovery. Once latched, the state is unreachable from JS. The caller cannot retry, cannot call
Resume(), and cannot even detect that it is latched rather than merely failing.
- The latch blocks the engine's own cleanup. While in the error state, RocksDB could not delete obsolete files. Reopening the database immediately freed ~644 MB of files it had been holding. So a database that fails because storage is full then keeps itself full — the failure is self-sustaining rather than self-correcting.
Why this hits self-hosted consumers hardest
A full disk is routine outside managed environments. The current behavior converts "disk filled up briefly, then I freed space" into "the database is read-only until someone notices and restarts it," with no diagnostic that says so — the error message actively misleads, since it names a file that does not exist.
Suggested direction
- Expose
DB::Resume() so a consumer can attempt recovery after the underlying condition clears.
- Expose the latched background-error state (and ideally its retryable/non-retryable classification) so a caller can distinguish "this write failed" from "this database is now read-only."
- Consider whether the error message surfaced to callers should be regenerated rather than replayed from the stored error, since the stored text becomes stale and misdirects diagnosis.
- Worth considering whether obsolete-file cleanup can proceed while in a write-latched state, given that it is the action most likely to clear the cause.
🤖 Generated with Claude Code
Summary
When a write fails at the filesystem level, RocksDB latches a background error and refuses all further writes. There is currently no way for a consumer to recover in-process — no exposed
DB::Resume(), and no way to observe the latched state. The only remedy is restarting the process, so a transient, self-correcting condition becomes an indefinite outage.Observed
An instance exhausted its filesystem quota. RocksDB recorded exactly one background error and stopped accepting writes:
Note
bg.retryable.io.error.count: 0— the error was classified non-retryable, so RocksDB's own auto-resume path was never eligible. That is arguably correct classification; the problem is what follows.Every subsequent write failed with a stored, stale message naming a WAL file that no longer existed on disk (the newest WAL present was several generations older). This repeated 19,149 times over 7h45m. Freeing space did not help — the quota was more than doubled and writes kept failing with the identical message. Only reopening the database recovered it.
Two consequences worth separating:
Resume(), and cannot even detect that it is latched rather than merely failing.Why this hits self-hosted consumers hardest
A full disk is routine outside managed environments. The current behavior converts "disk filled up briefly, then I freed space" into "the database is read-only until someone notices and restarts it," with no diagnostic that says so — the error message actively misleads, since it names a file that does not exist.
Suggested direction
DB::Resume()so a consumer can attempt recovery after the underlying condition clears.🤖 Generated with Claude Code