Epic: Replication reliability & scale — foundations, adaptive topology, and locking
This epic organizes a multi-release plan for Harper replication into tracked workstreams. It comes out of a deep architecture review of replication/ and the core substrate it depends on (full review notes available on request). The goal: turn a long tail of hard-won reliability fixes into a sound foundation that scales across topologies (full-mesh, hub-and-spoke/transitive, sharded, WAN, many-node) and supports adaptive behavior and distributed locking as additive features rather than new complexity on cracks.
Storage assumption: RocksDB is the canonical (and going-forward only) replicated engine. LMDB is being deprecated — this plan does not design around LMDB . That simplifies the keystone work below, because the RocksDB transaction log is already partitioned per origin.
The core finding
Reliability has been hard-won because a few foundational design choices each radiate a whole family of bugs, and the monolithic protocol engine (replicationConnection.ts, ~4,400 lines and growing — up ~800 lines since this epic was filed; replicateOverWS is a single ~3,250-line closure) makes every fix high-risk. Five of our strategic concerns are blocked on, or dramatically de-risked by, two foundations. Land these first and adaptive routing, dedicated threads, robust sharding, and locking become straightforward.
Foundation 1 — A single source of truth for connection & health state. Today the main-thread orchestrator keeps an edge-triggered, inferred mirror of connection state; the real sockets live on worker threads; live metrics (latency, back-pressure) live in shared-memory buffers the main thread never reads. That split is the direct cause of the connection-truth bug class (#289 , #349 , #357 , #233 ) and is exactly the missing bridge that adaptive replication (#218 ) needs. Fix it once → retire a bug class and unlock adaptive routing + adaptive threading.
Foundation 2 — Per-origin transaction-log convergence. Provably-correct transitive replication (#399 ), per-originating-node connections (#193 ), and per-origin observability (#192 ) all require resume cursors that track each origin's own position . The RocksDB transaction log already partitions per origin (RocksTransactionLogStore nodeLogs[]/logById); the work is to make replication's cursor layer consume per-origin positions as the primary mechanism and carry them transitively across hops. This is the keystone for transitive correctness and topology scale.
Root-cause themes (each generates multiple bugs)
A. No single source of truth for connection/health — the orchestrator's connected bit desyncs whenever a terminal/idle state is reached without the expected transition event; back-pressure/latency sit in shared buffers the main thread never reads. → Replication: ~9/11 outgoing peers show connected:false in connectionReplicationMap after restart despite being reachable #289 , Outbound replication WebSocket can silently die without firing 'close', leaving (peer, db) stuck with no retry until process restart #233 , subscriptionManager registers a worker.on('exit') per database subscription → MaxListenersExceededWarning with >10 databases #357 , Replication: 5.0.31 wedge-reconcile spins at 100% CPU re-subscribing an unreachable peer (forceResubscribe leaks into persistent listeners) #349 , cluster_status on a removed node should clearly indicate removal and close sockets #217
B. The resume cursor is trusted , not checked — narrow error classification over-holds (a permanently bad source pins the cursor forever) or silently gaps (a [T, head] hole is never detected because live tail traffic keeps connected:true). No receive-side sequence-gap detection. → rapid-reconnect stress: follower permanently loses replication backlog after restart churn (data loss; fast-skip ruled out) #426 , Incomplete (truncated) source blob is misclassified as transient → permanently wedges replication cursor (ENOENT is handled; 'Blob is incomplete' is not) #429 , Blob repair sweep has no automatic trigger — a materialized corrupt stub needs an operator to run repair_blob_data #385 , Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB) #386 , Proactive blob backfill: repair already-committed records whose blobs are missing/corrupt (no recovery once the resume cursor advances past them) #388
C. Inconsistent backoff; one-shot intents leak into persistent listeners — subscription-setup has no backoff (spins the main thread); forceResubscribe binds to a persistent listener; expensive per-reconnect work (TLS context creation) amplifies into OOM. → replication: no-backoff subscription-setup retry storm on transient boot-time DNS failure ends in OOM #327 , Replication: 5.0.31 wedge-reconcile spins at 100% CPU re-subscribing an unreachable peer (forceResubscribe leaks into persistent listeners) #349 , Investigate: heap corruption / OOM from rapid TLS context churn in monitorNodeCAs during reconnect cycles #288
D. Per-origin vs. interleaved cursor — transitive flooding happens because a relay re-streams an already-applied tail (the relayed cursor is the proxy's position, not the origin's ). Fixed by Foundation 2. → Replication: transitive/proxied re-delivery floods peers with already-applied out-of-order writes (reduce volume; complements harper#1310) #399 , Use separate WS connections per originating-node subscription (avoid full audit scan resets on fail-over) #193 , Per-originating-node transaction log partitioning (dependent on new transaction log format) #192 , [sharding] FATAL ERROR: NewSpace:: EnsureCurrentCapacity Allocation failed - JavaScript heap out of memory (verify on current build) #197
E. Residency is overloaded — replicateTo (transient per-request routing) is laundered into the same persistent residencyId as durable shard policy, triggering invalidation broadcasts to never-target nodes. → X-Replicate-From: none not respected for blob GETs — still fetches from remote nodes #208 , replicateTo update leaks to non-target node after SQL SELECT triggers cache load #211 , replicateTo update deletes record on non-target node #212 , plus the sharding crash family [sharding] FATAL ERROR: NewSpace:: EnsureCurrentCapacity Allocation failed - JavaScript heap out of memory (verify on current build) #197 –201, Records written to all nodes despite sharding configuration #257
F. Monolithic code blocks safe change — a ~3,250-line closure with ~50 interdependent locals, no protocol version negotiation, and an inbound decode loop that swallows errors (silent partial-message loss).
G. Wall-clock LWW makes clock skew a correctness property — concurrent full-record writes resolve by node name and silently discard the loser. This is the gap exclusive locking (harper#483) closes for explicitly-serialized keys.
H. Sync consumers of MaybePromise storage reads (added 2026-07-01) — RocksDB store.get() returns a Promise on a block-cache miss, so sync consumers of system-table point reads (handshake resume-cursor reads, hdb_nodes lookups) work on warm caches and small tables, then silently break as tables grow or caches go cold — the definitive root cause of the 5.1.x post-upgrade stall family. Fixed pointwise by getSync conversions (fix(replication): use getSync for hdb_nodes point reads (RocksDB MaybePromise) #476 , fix(replication): use getSync for seq/copyCursor resume-cursor reads (RocksDB MaybePromise) #484 ) and by typing the __dbis__ store so sync get() misuse is a compile error (feat(types): type the __dbis__/seq store so sync get() misuse is a compile error (#484 follow-up) #485 ). The residual — typed Store<V> for the remaining system stores + a lint rule — is tracked in W11 (Replication W11: Code organization, protocol versioning & decode-loop safety #440 ). → fix(replication): use getSync for hdb_nodes point reads (RocksDB MaybePromise) #476 , fix(replication): use getSync for seq/copyCursor resume-cursor reads (RocksDB MaybePromise) #484 , feat(types): type the __dbis__/seq store so sync get() misuse is a compile error (#484 follow-up) #485 , plus the superseded scan-fallback bandaids fix(replication): decode-resilient outbound subscriptions + idempotent node-update watcher on deploy reload (#460) #461 /harper#1463
Workstreams
Foundations (do first):
Features & cross-cutting:
Proposed sequencing
Phase 1 — Foundations & safety: W1 (Replication W1: Connection & health — single source of truth #431 ), W2 (Replication W2: Cursor correctness & divergence detection #432 ), W3 (Replication W3: Reconnect/storm control & resource bounds #433 ), W11 early steps (Replication W11: Code organization, protocol versioning & decode-loop safety #440 — decode-loop fix, protocol version negotiation, FrameWriter), W8 Tier 1 (Replication W8: Observability & metrics pipeline #437 — cluster_status quick wins), W13 correctness items (Replication W13: Base-copy & catch-up path (consolidation & correctness) #510 — MQTT SUBACK granted before cross-node delivery path is live after cluster bootstrap #495 /Replication: txn-log "table-reload" marker to extend copyApply to the system DB (#480 follow-up) #489 are already in flight).
Phase 2 — Keystone & observability backbone: W4 (Replication W4: Per-origin transaction-log convergence (keystone) #434 ), W8 Tier 2 (Replication W8: Observability & metrics pipeline #437 — metrics + divergence detection), W11 continued (Replication W11: Code organization, protocol versioning & decode-loop safety #440 ), W13 watchdog consolidation (Replication W13: Base-copy & catch-up path (consolidation & correctness) #510 , once W1 lands).
Phase 3 — Adaptive behavior & robust sharding: W5 (Adaptive replication routing: dynamic fan-out under egress back-pressure #218 ), W6 static pool (Replication W6: Adaptive / dedicated replication threads #435 ), W7 (Replication W7: Robust sharding (residency-vs-routing split) #436 ), W10 incremental (Replication W10: Performance #439 ), W13 copy-transport perf (Replication W13: Base-copy & catch-up path (consolidation & correctness) #510 — direct binary relay).
Phase 4 — Locking & advanced scale: W9 (Replication W9: Exclusive distributed locking (replication integration) #438 ), W6 elastic (Replication W6: Adaptive / dedicated replication threads #435 , if justified), W10 structural (Replication W10: Performance #439 ), large-cluster validation (Large cluster scalability testing (many-node horizontal scale) #263 , Stress: extend replicationLoad to 500K records + WAN latency + high-freq small-record replication pattern #300 ). W12 (Replication W12: Hierarchical aggregation / re-origin relay mode (large fan-in trees) #444 ) re-origin relay is design-spike-gated and opportunity-driven — may pull earlier than this phase.
W9 (locking) depends only on the core metadata/audit wiring, not on W1/W4, so it can start in parallel once the metadata bit lands.
Shipped since filing (ledger, updated 2026-07-01)
Point fixes landed since this epic was filed, mapped to the theme they belong to. They confirm the theme taxonomy — every one falls into a predicted family — but none retires a theme; the workstreams do that.
Theme A (connection truth): Outbound replication WebSocket can silently die without firing 'close', leaving (peer, db) stuck with no retry until process restart #233 closed; Replication wedges permanently after simultaneous cluster restart (reconciler skips open-but-idle sockets) → blocks replicated deploys #420 open-but-idle watchdog (fix(replication): recover open-but-idle wedged subscriptions via watchdog-driven reconnect #424 ); fix(replication): reconcile-level fallback for connected:true/Receiving copy-stall wedge #463 reconcile-level fallback for connected:true copy stalls; Replication: outbound subscription to a restarted peer wedges connected:false with no reconnect attempt (no SYN), reconcile never re-drives — base-copy + post-restart TLS window #466 pause-stall watchdog + wedge re-drive + never-connected backstop; replication: empty-subscription delayed close finishes a STILL-DESIRED peer as intentional after base-copy resync — permanent connected:false wedge, no reconnect #471 empty-subscription delayed-close fix (fix(replication): don't finish a still-desired peer on empty-subscription delayed close (#471) #475 ). Note the shape of these fixes: they are now ~6 layered edge-triggered watchdog/fallback recovery mechanisms (~97 watchdog references in replicationConnection.ts), and they have already produced their first interaction bug (the Replication: outbound subscription to a restarted peer wedges connected:false with no reconnect attempt (no SYN), reconcile never re-drives — base-copy + post-restart TLS window #466 false-positive force-reconnect). The risk has shifted from "wedges with no recovery" to "overlapping recovery layers with subtle interactions" — which sharpens, not weakens, the case for W1: its acceptance criterion now includes demoting these watchdogs to telemetry/assertions (see Replication W1: Connection & health — single source of truth #431 ).
Theme B (cursor trust): blob error classification broadened well past ENOENT — Incomplete (truncated) source blob is misclassified as transient → permanently wedges replication cursor (ENOENT is handled; 'Blob is incomplete' is not) #429 /fix(replication): classify gone/corrupt source blobs as permanent via forwarded statusCode (#429) #443 statusCode-forwarded permanent classification, Blob replication wedges permanently when source blob is gone (ENOENT) on an expiration cache table — held resume cursor never recovers #403 /fix(replication): advance resume cursor past source-missing (ENOENT) blobs instead of wedging (#403) #405 ENOENT advance-past, the core blob-write taxonomy (harper#1480: PENDING_TYPE stamp, 503-transient vs 500-permanent), and the Proactive blob backfill: repair already-committed records whose blobs are missing/corrupt (no recovery once the resume cursor advances past them) #388 proactive blob-repair sweep (blobRepair.ts). rapid-reconnect stress: follower permanently loses replication backlog after restart churn (data loss; fast-skip ruled out) #426 closed via fix(replication): full-copy when a source has no resume cursor (#426) #428 (cursorless-start full copy) + the fast-skip; its send-side stale-cursor residual is explicitly owned by W2/W4. Still missing: the bounded-retry escalation budget (a deliberately-transient 503 can still pin the cursor forever — observed live as the v4→v5 circular-503 wedge) and receive-side gap detection. See the W2 status update (Replication W2: Cursor correctness & divergence detection #432 ).
Theme C (storm control): subscription-setup retries now have fixed delays in places, but replication: no-backoff subscription-setup retry storm on transient boot-time DNS failure ends in OOM #327 /Replication: 5.0.31 wedge-reconcile spins at 100% CPU re-subscribing an unreachable peer (forceResubscribe leaks into persistent listeners) #349 /Investigate: heap corruption / OOM from rapid TLS context churn in monitorNodeCAs during reconnect cycles #288 all remain open; the uniform backoff discipline is still the gap.
Theme F (monolith): got worse — the file grew ~800 lines absorbing all of the above; the decode-loop error swallow and the absence of protocol versioning are unchanged. A DESIGN.md navigation guide and a proven low-risk decomposition pattern (extracted pure decision helpers + unit tests) emerged — W11 codifies it (Replication W11: Code organization, protocol versioning & decode-loop safety #440 ).
Theme H: discovered and mostly retired post-filing (see above); residual in W11.
Copy path: the Replicated hdb_analytics floods system transaction logs and spins a worker in native code; system base-copy wedges until worker recycle #480 family (copyApply durable snapshots, copy-flush pacer fix(replication): pace bulk full-copy flushes under the receive watchdog #483 , copy-progress watchdog, control-plane-first ordering Order replication base copy: control-plane tables before bulk tables (#421) #422 ) plus new follow-ups (Replication: txn-log "table-reload" marker to extend copyApply to the system DB (#480 follow-up) #489 , MQTT SUBACK granted before cross-node delivery path is live after cluster bootstrap #495 ) — consolidated into the new W13 (Replication W13: Base-copy & catch-up path (consolidation & correctness) #510 ) .
Also shipped: Enforce directional controlled-flow replication config on live connections #506 directional controlled-flow enforcement on live connections (closes the Controlled-flow replication: directional fields in replication.routes[] config are not enforced on live connections #498 gap — directional route config was stored but not honored).
Notes
🤖 Filed by Claude on behalf of Kris.
Epic: Replication reliability & scale — foundations, adaptive topology, and locking
This epic organizes a multi-release plan for Harper replication into tracked workstreams. It comes out of a deep architecture review of
replication/and the core substrate it depends on (full review notes available on request). The goal: turn a long tail of hard-won reliability fixes into a sound foundation that scales across topologies (full-mesh, hub-and-spoke/transitive, sharded, WAN, many-node) and supports adaptive behavior and distributed locking as additive features rather than new complexity on cracks.Storage assumption: RocksDB is the canonical (and going-forward only) replicated engine. LMDB is being deprecated — this plan does not design around LMDB. That simplifies the keystone work below, because the RocksDB transaction log is already partitioned per origin.
The core finding
Reliability has been hard-won because a few foundational design choices each radiate a whole family of bugs, and the monolithic protocol engine (
replicationConnection.ts, ~4,400 lines and growing — up ~800 lines since this epic was filed;replicateOverWSis a single ~3,250-line closure) makes every fix high-risk. Five of our strategic concerns are blocked on, or dramatically de-risked by, two foundations. Land these first and adaptive routing, dedicated threads, robust sharding, and locking become straightforward.Foundation 1 — A single source of truth for connection & health state. Today the main-thread orchestrator keeps an edge-triggered, inferred mirror of connection state; the real sockets live on worker threads; live metrics (latency, back-pressure) live in shared-memory buffers the main thread never reads. That split is the direct cause of the connection-truth bug class (#289, #349, #357, #233) and is exactly the missing bridge that adaptive replication (#218) needs. Fix it once → retire a bug class and unlock adaptive routing + adaptive threading.
Foundation 2 — Per-origin transaction-log convergence. Provably-correct transitive replication (#399), per-originating-node connections (#193), and per-origin observability (#192) all require resume cursors that track each origin's own position. The RocksDB transaction log already partitions per origin (
RocksTransactionLogStorenodeLogs[]/logById); the work is to make replication's cursor layer consume per-origin positions as the primary mechanism and carry them transitively across hops. This is the keystone for transitive correctness and topology scale.Root-cause themes (each generates multiple bugs)
connectedbit desyncs whenever a terminal/idle state is reached without the expected transition event; back-pressure/latency sit in shared buffers the main thread never reads. → Replication: ~9/11 outgoing peers show connected:false in connectionReplicationMap after restart despite being reachable #289, Outbound replication WebSocket can silently die without firing 'close', leaving (peer, db) stuck with no retry until process restart #233, subscriptionManager registers a worker.on('exit') per database subscription → MaxListenersExceededWarning with >10 databases #357, Replication: 5.0.31 wedge-reconcile spins at 100% CPU re-subscribing an unreachable peer (forceResubscribe leaks into persistent listeners) #349, cluster_status on a removed node should clearly indicate removal and close sockets #217[T, head]hole is never detected because live tail traffic keepsconnected:true). No receive-side sequence-gap detection. → rapid-reconnect stress: follower permanently loses replication backlog after restart churn (data loss; fast-skip ruled out) #426, Incomplete (truncated) source blob is misclassified as transient → permanently wedges replication cursor (ENOENT is handled; 'Blob is incomplete' is not) #429, Blob repair sweep has no automatic trigger — a materialized corrupt stub needs an operator to run repair_blob_data #385, Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB) #386, Proactive blob backfill: repair already-committed records whose blobs are missing/corrupt (no recovery once the resume cursor advances past them) #388forceResubscribebinds to a persistent listener; expensive per-reconnect work (TLS context creation) amplifies into OOM. → replication: no-backoff subscription-setup retry storm on transient boot-time DNS failure ends in OOM #327, Replication: 5.0.31 wedge-reconcile spins at 100% CPU re-subscribing an unreachable peer (forceResubscribe leaks into persistent listeners) #349, Investigate: heap corruption / OOM from rapid TLS context churn in monitorNodeCAs during reconnect cycles #288replicateTo(transient per-request routing) is laundered into the same persistentresidencyIdas durable shard policy, triggering invalidation broadcasts to never-target nodes. → X-Replicate-From: none not respected for blob GETs — still fetches from remote nodes #208, replicateTo update leaks to non-target node after SQL SELECT triggers cache load #211, replicateTo update deletes record on non-target node #212, plus the sharding crash family [sharding] FATAL ERROR: NewSpace:: EnsureCurrentCapacity Allocation failed - JavaScript heap out of memory (verify on current build) #197–201, Records written to all nodes despite sharding configuration #257store.get()returns a Promise on a block-cache miss, so sync consumers of system-table point reads (handshake resume-cursor reads,hdb_nodeslookups) work on warm caches and small tables, then silently break as tables grow or caches go cold — the definitive root cause of the 5.1.x post-upgrade stall family. Fixed pointwise bygetSyncconversions (fix(replication): use getSync for hdb_nodes point reads (RocksDB MaybePromise) #476, fix(replication): use getSync for seq/copyCursor resume-cursor reads (RocksDB MaybePromise) #484) and by typing the__dbis__store so syncget()misuse is a compile error (feat(types): type the __dbis__/seq store so sync get() misuse is a compile error (#484 follow-up) #485). The residual — typedStore<V>for the remaining system stores + a lint rule — is tracked in W11 (Replication W11: Code organization, protocol versioning & decode-loop safety #440). → fix(replication): use getSync for hdb_nodes point reads (RocksDB MaybePromise) #476, fix(replication): use getSync for seq/copyCursor resume-cursor reads (RocksDB MaybePromise) #484, feat(types): type the __dbis__/seq store so sync get() misuse is a compile error (#484 follow-up) #485, plus the superseded scan-fallback bandaids fix(replication): decode-resilient outbound subscriptions + idempotent node-update watcher on deploy reload (#460) #461/harper#1463Workstreams
Foundations (do first):
Features & cross-cutting:
Proposed sequencing
W9 (locking) depends only on the core metadata/audit wiring, not on W1/W4, so it can start in parallel once the metadata bit lands.
Shipped since filing (ledger, updated 2026-07-01)
Point fixes landed since this epic was filed, mapped to the theme they belong to. They confirm the theme taxonomy — every one falls into a predicted family — but none retires a theme; the workstreams do that.
replicationConnection.ts), and they have already produced their first interaction bug (the Replication: outbound subscription to a restarted peer wedges connected:false with no reconnect attempt (no SYN), reconcile never re-drives — base-copy + post-restart TLS window #466 false-positive force-reconnect). The risk has shifted from "wedges with no recovery" to "overlapping recovery layers with subtle interactions" — which sharpens, not weakens, the case for W1: its acceptance criterion now includes demoting these watchdogs to telemetry/assertions (see Replication W1: Connection & health — single source of truth #431).blobRepair.ts). rapid-reconnect stress: follower permanently loses replication backlog after restart churn (data loss; fast-skip ruled out) #426 closed via fix(replication): full-copy when a source has no resume cursor (#426) #428 (cursorless-start full copy) + the fast-skip; its send-side stale-cursor residual is explicitly owned by W2/W4. Still missing: the bounded-retry escalation budget (a deliberately-transient 503 can still pin the cursor forever — observed live as the v4→v5 circular-503 wedge) and receive-side gap detection. See the W2 status update (Replication W2: Cursor correctness & divergence detection #432).DESIGN.mdnavigation guide and a proven low-risk decomposition pattern (extracted pure decision helpers + unit tests) emerged — W11 codifies it (Replication W11: Code organization, protocol versioning & decode-loop safety #440).Notes
from-jiraissues reference the abandoned NATS transport (Add NATS message stream backlog counts to cluster_status #261 backlog counts, NATS MAX_PAYLOAD_EXCEEDED error when cluster cloning large data #266 large-clone payload). The concerns map onto the live WebSocket engine (audit backlog depth; oversized-message chunking throughMAX_PAYLOAD); the NATS-specific issues should be reframed or closed.🤖 Filed by Claude on behalf of Kris.