Repository navigation
remove_node leaves the departed peer's per-origin transaction-log store orphaned forever #686
Description
Activity
Production impact 2026-09-24: this is what turned a deploy into a 3-hour outage on J.Jill prod (prod.jjl, harper-pro 5.1.12, 10 nodes)
Six nodes removed from this cluster before its June 25 rebuild (us-central-1, us-central-2, us-east-1, us-west-1, us-sea-1, us-southeast-3) still had per-origin transaction-log stores on every node in all four databases.
remove_node(replication/setNode.ts,hdbNodes.deleteat L91 /remove_node_backat L97 on main) deleted only the hdb_nodes row. Also left behind: theREMOTE_NODE_IDSname-to-id entries, the__dbis__seq/copyCursorkeys, and the hdb_certificate rows.Chain: a component deploy at 01:24Z restarted all 10 nodes' http workers within 8 s. Every node re-subscribed from its saved cursor and hit two records about the dead nodes (hdb_user
clone-temp-adminfrom us-west-1, hdb_nodes row for08x-us-east-1from us-central-2). Core routes each write to the origin's store (RocksTransactionLogStore.put:logById(nodeId) ?? logById(viaNodeId) ?? local), the dead-node store still existed so the write bound there, the next record in the batch wanted a different store, rocksdb-js threwTransaction N is already bound to the log store "<dead node>"(see HarperFast/harper#1162), the cursor never advanced, and 10 nodes re-delivered the same two records at ~25/s each for 3 hours: 106k error lines in 40 min, 52 worker OOMs, 5 process crashes, ord-1 swap-thrashed at 28 GB, www.jjill.com TTFB 19 to 28 s.Fix applied live (approved by @kriszyp): per node, out of GTM, stop container,
mv database/{system,jjillCache,harperfast_nextjs,data}/transaction_logs/<dead-node>/aside, start. Errors went to zero on the cleaned node immediately while peers still stormed; rolled to all 10 in 65 min. With no store for the dead origin the write falls back to the sender's store and cannot conflict. Residual: a bounded startup burst (mia-1: 588 lines in 90 s, then 0) when a batch mixed a dead-origin record with a live one.Fix for this issue is small and shippable on the 5.1 line: rocksdb-js has had
purgeLogs({ name, destroy: true })since 2.2.0 (the version prod runs).remove_nodeandremove_node_backshould call it per database for the departed peer after the row delete, and drop the id-map entry, cursors and certs. Core'sRocksTransactionLogStore.remove()stub (L634) is the natural place for the store part.Also observed: a store directory is recreated if a stale hdb_nodes row for the dead peer is transiently replayed (
ensureLogExistsat replicationConnection.ts L7724, seen at 01:28:47 for 77o-us-central-1), so the cleanup must be paired with #1162's cursor-advance fix.Related: #793 (removed peer re-authorizes), HarperFast/rocksdb-js#808 (purge once per process). Needed for J.Jill's 5.1.29 ask (they upgrade Sept 30). Full timeline available from Maurice.
🤖 Investigated and filed by Claude (Fable) on Maurice's behalf
Second occurrence in 12 h: preprod.jjl, 2026-09-24, on 5.1.28 — plus the recovery shape that actually converges
Same mechanism as the prod incident described above, this time on 5.1.28 (confirming the version does not change it): a component deploy at 17:34Z restarted http workers on all nodes at once, and within 75 minutes every node's orphan store for a node removed months ago (
database/system/transaction_logs/<removed-node>/) had grown from empty to 0.3–2 GB of re-loggedhdb_userentries (~300k entries per 16 MB file, almost all under a single source version). Workers OOM'd cluster-wide, and once a process died its boot-time replay of that store heap-OOM'd the main thread → 5-minute boot loop on 4 of 10 nodes (details and measured version distribution in harper#2161).Two findings that change the runbook:
-
A rolling clean re-poisons the cleaned node. After quarantining the stores on one node and starting it, a live peer immediately streamed its own copy of the junk: 93k entries landed in the cleaned node's store for that peer within 10 minutes (origin id unknown →
logById(viaNodeId)), RSS 3 → 13 GB. Every node streams its logs to every other node, so the junk circulates as a ring for as long as any node still holds it — this is the mechanism behind the ord-1 relapse on prod. With no drain constraint, stop all → quarantine on all → start all together converged in ~10 minutes: all 10 nodes booted in ~20 s, RSS ~3 GB (from 22–31 GB), zero re-log errors, stores not recreated, 9/9 legs × 4 databases connected. -
The "already bound" log line is not a liveness signal. Live nodes showed 0 of those lines for 15+ minutes while their orphan stores kept growing ~2 MB/min. Watch the store's newest-file mtime / size instead.
Quarantined (moved, not deleted) on every node, in all four databases: the removed-node stores (six of them here) and the empty numeric stores from #2778. Nothing else touched.
Reinforces the fix direction:
remove_nodeshould destroy the departed peer's per-origin stores (purgeLogs({ name, destroy: true })) and drop its id-map / cursor / cert entries; and boot replay should not walk stores whose origin is absent fromhdb_nodes(proposed on harper#2161).🤖 Investigated and written by Claude (Fable 5.1) on David's behalf
-
Reinforces the fix direction: remove_node should destroy the departed peer's per-origin stores
That would be reckless/dangerous deletion of live production data. Membership removal does permit the deletion of data that that member previously wrote (and was accepted by other nodes). There may be some residual empty-store/metadata cleanup we can look into here, but considering the premise of this ticket is fundamentally wrong, and the priority is P2, will be deferring this for a while.
Two follow-ups from a production cleanup of this problem (10-node cluster on 5.1.12, 2026-09-24).
Evidence that departed-peer stores stay live. Two nodes had been removed from the cluster via remove_node three months earlier. Their
systemlogs were still present on the remaining nodes, held 9.1M to 11.7M entries (first removed node) and 22k to 137k entries (second removed node) on the two nodes checked, all stamped with the removed nodes' origin ids, and were still being appended to until 2026-09-24 13:41Z. Once a peer restarted and replayed, every node with such a store threwalready bound to the log storeand looped; nodes without one could not (put() falls back to the via-node log). Moving the stores aside on all 10 nodes cleared the loop; leaving them on one node re-poisoned it as soon as a cleaned peer restarted. Frame-level detail with origin ids is on HarperFast/harper#2793.Nothing was deleted. Each departed-peer store was moved to a quarantine directory on the same host filesystem, reversible and still there. The ask here can be worded as: after remove_node, stop opening and replaying a log whose node is no longer in hdb_nodes, and let retention purge it. It does not need to be a delete.
The restart-triggered replay of dead-node records described above has a cursor root cause, now filed as #989. The sender starts every log that is neither excluded nor listed in the request at position 0, and a removed node's log is exactly that case. Its full retained history is therefore replayed to every peer on every resubscribe, and the replayed entries are written into its log again. #989 fixes this with per-origin cursors (W4's
startByLogvector) rather than by deleting stores.🤖 Claude (Opus 5.5) on Kris's behalf
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
Every replication peer gets a per-origin transaction-log store on each database (
ensureLogExists→rootStore.useLog(name)). Nothing ever deletes one:remove_node(replication/setNode.ts) does no transaction-log cleanup,removeLogonly splices a log out of an in-flight iterator (in-memory), and core'sRocksTransactionLogStore.remove()is an explicit no-op stub (// TODO: this function can likely be removed once the call to purgeLogs() is added in resources/Table.ts— v5.2.1).destroy: trueappears nowhere in either repo.Impact
A removed or renamed peer leaves its log store (and sequence files) on disk on every node in the cluster, permanently. Store-wide
purgeLogsprunes aged entries inside it (when pruning runs at all — see HarperFast/harper#2140) but the store itself and its floor never go away. Clusters that churn membership — migrations, node replacement, bridge nodes — accumulate dead per-origin stores per database.Suggested fix
remove_node(andadd_node_back's reciprocal-removal path) should destroy the departed peer's per-origin log stores across databases after the removal settles — or at minimum mark them for deletion at the next purge pass.🤖 Investigated and filed by Claude (Fable) on Nathan's behalf