Skip to content

Disk-space safety thresholds: reclamation / write-stop / hard-shutdown #594

Description

@kriszyp

Add a configurable disk-space safety model so Harper can deliberately throttle and then halt writes before disk exhaustion, instead of failing catastrophically when the disk fills (which is exceptionally difficult and time-consuming to recover from).

Model (per Jeff Darnton's writeup)

Conceptually three thresholds:

  1. Reclamation threshold (X% free) — reclamation runs on an interval; gets progressively more aggressive as disk pressure rises.
  2. Backpressure / warning — clear warnings (logs at minimum) before unsafe levels are reached; "user" operations start receiving backpressure.
  3. Hard refusal threshold (Y% free) — Harper refuses incoming user writes to protect itself. Internal/system writes (status flips, log rotation, replication acknowledgement, etc.) continue.
  4. Nuclear option — if the disk would otherwise fill anyway, Harper shuts itself down before the filesystem hits 100% (recovery is far easier from a clean shutdown than from a full disk).

Important constraint (per Kris)

Blocking all writes — including the hdb_status system table — would prevent us from setting the instance to unavailable. The throttle MUST be partitioned: "user" writes off, "system" writes on. Open question: does hdb_analytics count as system or user? (It's by far the most heavily written system table; consider moving it to its own database).

Asks

  • Config knobs: storage.reclamationThreshold, storage.writeStopThreshold, storage.hardShutdownThreshold (or equivalent).
  • Apply the threshold against quota when one is configured (see CORE-3011 / HarperFast/harper#593), not against raw filesystem.
  • Partition writes: system writes always allowed; user writes throttled then refused.
  • Auto-flip hdb_status to unavailable when the write-stop threshold trips.
  • Clear logging of the state transitions.

Acceptance criteria

  • Operator can configure all three thresholds (or sensible defaults are documented).
  • Reaching the write-stop threshold:
    • User writes return a clear, actionable error.
    • System writes (including hdb_status flip to unavailable) succeed.
    • Logged at WARN or ERROR.
  • Reaching the hard-shutdown threshold:
    • Harper shuts down cleanly before disk hits 100%.
    • Recovery path: operator clears space, restarts, instance comes back.

Related


Tracked in Jira: CORE-3006

🤖 Filed by Claude on behalf of Kris.

Activity

  1. added
    enhancementNew feature or request
    from-jiraMigrated or originated from a Jira ticket
    area:storageStorage engine, LMDB/RocksDB, compaction
    on May 19, 2026
  2. added this to the v5.1 milestone on May 21, 2026
  3. modified the milestones: v5.1, v5.3 on Jul 23, 2026
  4. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:storageStorage engine, LMDB/RocksDB, compactionenhancementNew feature or requestfrom-jiraMigrated or originated from a Jira ticket

Type

No type

Fields

Priority

P0

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions