Skip to content

[Epic] Replication reliability & scale — foundations, adaptive topology, and locking #430

Description

@kriszyp

Epic: Replication reliability & scale — foundations, adaptive topology, and locking

This epic organizes a multi-release plan for Harper replication into tracked workstreams. It comes out of a deep architecture review of replication/ and the core substrate it depends on (full review notes available on request). The goal: turn a long tail of hard-won reliability fixes into a sound foundation that scales across topologies (full-mesh, hub-and-spoke/transitive, sharded, WAN, many-node) and supports adaptive behavior and distributed locking as additive features rather than new complexity on cracks.

Storage assumption: RocksDB is the canonical (and going-forward only) replicated engine. LMDB is being deprecated — this plan does not design around LMDB. That simplifies the keystone work below, because the RocksDB transaction log is already partitioned per origin.

The core finding

Reliability has been hard-won because a few foundational design choices each radiate a whole family of bugs, and the monolithic protocol engine (replicationConnection.ts, ~4,400 lines and growing — up ~800 lines since this epic was filed; replicateOverWS is a single ~3,250-line closure) makes every fix high-risk. Five of our strategic concerns are blocked on, or dramatically de-risked by, two foundations. Land these first and adaptive routing, dedicated threads, robust sharding, and locking become straightforward.

Foundation 1 — A single source of truth for connection & health state. Today the main-thread orchestrator keeps an edge-triggered, inferred mirror of connection state; the real sockets live on worker threads; live metrics (latency, back-pressure) live in shared-memory buffers the main thread never reads. That split is the direct cause of the connection-truth bug class (#289, #349, #357, #233) and is exactly the missing bridge that adaptive replication (#218) needs. Fix it once → retire a bug class and unlock adaptive routing + adaptive threading.

Foundation 2 — Per-origin transaction-log convergence. Provably-correct transitive replication (#399), per-originating-node connections (#193), and per-origin observability (#192) all require resume cursors that track each origin's own position. The RocksDB transaction log already partitions per origin (RocksTransactionLogStore nodeLogs[]/logById); the work is to make replication's cursor layer consume per-origin positions as the primary mechanism and carry them transitively across hops. This is the keystone for transitive correctness and topology scale.

Root-cause themes (each generates multiple bugs)

Workstreams

Foundations (do first):

Features & cross-cutting:

Proposed sequencing

W9 (locking) depends only on the core metadata/audit wiring, not on W1/W4, so it can start in parallel once the metadata bit lands.

Shipped since filing (ledger, updated 2026-07-01)

Point fixes landed since this epic was filed, mapped to the theme they belong to. They confirm the theme taxonomy — every one falls into a predicted family — but none retires a theme; the workstreams do that.

Notes

🤖 Filed by Claude on behalf of Kris.

Activity

  1. self-assigned this
    on Jun 20, 2026
  2. 12 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:replicationReplication, cluster sync, peer connectionsenhancementNew feature or request

Type

No type

Fields

Priority

P1

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions