Skip to content

No in-process recovery from a latched background error; only a restart clears it #730

Description

@heskew

Summary

When a write fails at the filesystem level, RocksDB latches a background error and refuses all further writes. There is currently no way for a consumer to recover in-process — no exposed DB::Resume(), and no way to observe the latched state. The only remedy is restarting the process, so a transient, self-correcting condition becomes an indefinite outage.

Observed

An instance exhausted its filesystem quota. RocksDB recorded exactly one background error and stopped accepting writes:

rocksdb.error.handler.bg.error.count                COUNT : 1
rocksdb.error.handler.bg.io.error.count             COUNT : 1
rocksdb.error.handler.bg.retryable.io.error.count   COUNT : 0
rocksdb.error.handler.autoresume.count              COUNT : 0
rocksdb.error.handler.autoresume.success.count      COUNT : 0

Note bg.retryable.io.error.count: 0 — the error was classified non-retryable, so RocksDB's own auto-resume path was never eligible. That is arguably correct classification; the problem is what follows.

Every subsequent write failed with a stored, stale message naming a WAL file that no longer existed on disk (the newest WAL present was several generations older). This repeated 19,149 times over 7h45m. Freeing space did not help — the quota was more than doubled and writes kept failing with the identical message. Only reopening the database recovered it.

Two consequences worth separating:

  1. No in-process recovery. Once latched, the state is unreachable from JS. The caller cannot retry, cannot call Resume(), and cannot even detect that it is latched rather than merely failing.
  2. The latch blocks the engine's own cleanup. While in the error state, RocksDB could not delete obsolete files. Reopening the database immediately freed ~644 MB of files it had been holding. So a database that fails because storage is full then keeps itself full — the failure is self-sustaining rather than self-correcting.

Why this hits self-hosted consumers hardest

A full disk is routine outside managed environments. The current behavior converts "disk filled up briefly, then I freed space" into "the database is read-only until someone notices and restarts it," with no diagnostic that says so — the error message actively misleads, since it names a file that does not exist.

Suggested direction

  • Expose DB::Resume() so a consumer can attempt recovery after the underlying condition clears.
  • Expose the latched background-error state (and ideally its retryable/non-retryable classification) so a caller can distinguish "this write failed" from "this database is now read-only."
  • Consider whether the error message surfaced to callers should be regenerated rather than replayed from the stored error, since the stored text becomes stale and misdirects diagnosis.
  • Worth considering whether obsolete-file cleanup can proceed while in a write-latched state, given that it is the action most likely to clear the cause.

🤖 Generated with Claude Code

Activity

  1. added 2 commits that reference this issue on Aug 14, 2026
    7370c28
    ff92423
  2. added a commit that references this issue on Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions