Skip to content

remove_node leaves the departed peer's per-origin transaction-log store orphaned forever #686

Description

@heskew

Summary

Every replication peer gets a per-origin transaction-log store on each database (ensureLogExists → rootStore.useLog(name)). Nothing ever deletes one: remove_node (replication/setNode.ts) does no transaction-log cleanup, removeLog only splices a log out of an in-flight iterator (in-memory), and core's RocksTransactionLogStore.remove() is an explicit no-op stub (// TODO: this function can likely be removed once the call to purgeLogs() is added in resources/Table.ts — v5.2.1). destroy: true appears nowhere in either repo.

Impact

A removed or renamed peer leaves its log store (and sequence files) on disk on every node in the cluster, permanently. Store-wide purgeLogs prunes aged entries inside it (when pruning runs at all — see HarperFast/harper#2140) but the store itself and its floor never go away. Clusters that churn membership — migrations, node replacement, bridge nodes — accumulate dead per-origin stores per database.

Suggested fix

remove_node (and add_node_back's reciprocal-removal path) should destroy the departed peer's per-origin log stores across databases after the removal settles — or at minimum mark them for deletion at the next purge pass.


🤖 Investigated and filed by Claude (Fable) on Nathan's behalf

Activity

  1. added theissue type on Aug 11, 2026
  2. maurice-harper commented on Sep 24, 2026

    @maurice-harper

    Production impact 2026-09-24: this is what turned a deploy into a 3-hour outage on J.Jill prod (prod.jjl, harper-pro 5.1.12, 10 nodes)

    Six nodes removed from this cluster before its June 25 rebuild (us-central-1, us-central-2, us-east-1, us-west-1, us-sea-1, us-southeast-3) still had per-origin transaction-log stores on every node in all four databases. remove_node (replication/setNode.ts, hdbNodes.delete at L91 / remove_node_back at L97 on main) deleted only the hdb_nodes row. Also left behind: the REMOTE_NODE_IDS name-to-id entries, the __dbis__ seq/copyCursor keys, and the hdb_certificate rows.

    Chain: a component deploy at 01:24Z restarted all 10 nodes' http workers within 8 s. Every node re-subscribed from its saved cursor and hit two records about the dead nodes (hdb_user clone-temp-admin from us-west-1, hdb_nodes row for 08x-us-east-1 from us-central-2). Core routes each write to the origin's store (RocksTransactionLogStore.put: logById(nodeId) ?? logById(viaNodeId) ?? local), the dead-node store still existed so the write bound there, the next record in the batch wanted a different store, rocksdb-js threw Transaction N is already bound to the log store "<dead node>" (see HarperFast/harper#1162), the cursor never advanced, and 10 nodes re-delivered the same two records at ~25/s each for 3 hours: 106k error lines in 40 min, 52 worker OOMs, 5 process crashes, ord-1 swap-thrashed at 28 GB, www.jjill.com TTFB 19 to 28 s.

    Fix applied live (approved by @kriszyp): per node, out of GTM, stop container, mv database/{system,jjillCache,harperfast_nextjs,data}/transaction_logs/<dead-node>/ aside, start. Errors went to zero on the cleaned node immediately while peers still stormed; rolled to all 10 in 65 min. With no store for the dead origin the write falls back to the sender's store and cannot conflict. Residual: a bounded startup burst (mia-1: 588 lines in 90 s, then 0) when a batch mixed a dead-origin record with a live one.

    Fix for this issue is small and shippable on the 5.1 line: rocksdb-js has had purgeLogs({ name, destroy: true }) since 2.2.0 (the version prod runs). remove_node and remove_node_back should call it per database for the departed peer after the row delete, and drop the id-map entry, cursors and certs. Core's RocksTransactionLogStore.remove() stub (L634) is the natural place for the store part.

    Also observed: a store directory is recreated if a stale hdb_nodes row for the dead peer is transiently replayed (ensureLogExists at replicationConnection.ts L7724, seen at 01:28:47 for 77o-us-central-1), so the cleanup must be paired with #1162's cursor-advance fix.

    Related: #793 (removed peer re-authorizes), HarperFast/rocksdb-js#808 (purge once per process). Needed for J.Jill's 5.1.29 ask (they upgrade Sept 30). Full timeline available from Maurice.

    🤖 Investigated and filed by Claude (Fable) on Maurice's behalf

  3. DavidCockerill commented on Sep 24, 2026

    @DavidCockerill
    Member

    Second occurrence in 12 h: preprod.jjl, 2026-09-24, on 5.1.28 — plus the recovery shape that actually converges

    Same mechanism as the prod incident described above, this time on 5.1.28 (confirming the version does not change it): a component deploy at 17:34Z restarted http workers on all nodes at once, and within 75 minutes every node's orphan store for a node removed months ago (database/system/transaction_logs/<removed-node>/) had grown from empty to 0.3–2 GB of re-logged hdb_user entries (~300k entries per 16 MB file, almost all under a single source version). Workers OOM'd cluster-wide, and once a process died its boot-time replay of that store heap-OOM'd the main thread → 5-minute boot loop on 4 of 10 nodes (details and measured version distribution in harper#2161).

    Two findings that change the runbook:

    1. A rolling clean re-poisons the cleaned node. After quarantining the stores on one node and starting it, a live peer immediately streamed its own copy of the junk: 93k entries landed in the cleaned node's store for that peer within 10 minutes (origin id unknown → logById(viaNodeId)), RSS 3 → 13 GB. Every node streams its logs to every other node, so the junk circulates as a ring for as long as any node still holds it — this is the mechanism behind the ord-1 relapse on prod. With no drain constraint, stop all → quarantine on all → start all together converged in ~10 minutes: all 10 nodes booted in ~20 s, RSS ~3 GB (from 22–31 GB), zero re-log errors, stores not recreated, 9/9 legs × 4 databases connected.

    2. The "already bound" log line is not a liveness signal. Live nodes showed 0 of those lines for 15+ minutes while their orphan stores kept growing ~2 MB/min. Watch the store's newest-file mtime / size instead.

    Quarantined (moved, not deleted) on every node, in all four databases: the removed-node stores (six of them here) and the empty numeric stores from #2778. Nothing else touched.

    Reinforces the fix direction: remove_node should destroy the departed peer's per-origin stores (purgeLogs({ name, destroy: true })) and drop its id-map / cursor / cert entries; and boot replay should not walk stores whose origin is absent from hdb_nodes (proposed on harper#2161).

    🤖 Investigated and written by Claude (Fable 5.1) on David's behalf

  4. kriszyp commented on Sep 24, 2026

    @kriszyp
    Member

    Reinforces the fix direction: remove_node should destroy the departed peer's per-origin stores

    That would be reckless/dangerous deletion of live production data. Membership removal does permit the deletion of data that that member previously wrote (and was accepted by other nodes). There may be some residual empty-store/metadata cleanup we can look into here, but considering the premise of this ticket is fundamentally wrong, and the priority is P2, will be deferring this for a while.

  5. maurice-harper commented on Sep 25, 2026

    @maurice-harper

    Two follow-ups from a production cleanup of this problem (10-node cluster on 5.1.12, 2026-09-24).

    Evidence that departed-peer stores stay live. Two nodes had been removed from the cluster via remove_node three months earlier. Their system logs were still present on the remaining nodes, held 9.1M to 11.7M entries (first removed node) and 22k to 137k entries (second removed node) on the two nodes checked, all stamped with the removed nodes' origin ids, and were still being appended to until 2026-09-24 13:41Z. Once a peer restarted and replayed, every node with such a store threw already bound to the log store and looped; nodes without one could not (put() falls back to the via-node log). Moving the stores aside on all 10 nodes cleared the loop; leaving them on one node re-poisoned it as soon as a cleaned peer restarted. Frame-level detail with origin ids is on HarperFast/harper#2793.

    Nothing was deleted. Each departed-peer store was moved to a quarantine directory on the same host filesystem, reversible and still there. The ask here can be worded as: after remove_node, stop opening and replaying a log whose node is no longer in hdb_nodes, and let retention purge it. It does not need to be a delete.

  6. kriszyp commented on Oct 6, 2026

    @kriszyp
    Member

    The restart-triggered replay of dead-node records described above has a cursor root cause, now filed as #989. The sender starts every log that is neither excluded nor listed in the request at position 0, and a removed node's log is exactly that case. Its full retained history is therefore replayed to every peer on every resubscribe, and the replayed entries are written into its log again. #989 fixes this with per-origin cursors (W4's startByLog vector) rather than by deleting stores.

    🤖 Claude (Opus 5.5) on Kris's behalf

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions