Repository navigation
Adaptive replication routing: dynamic fan-out under egress back-pressure #218
Description
Activity
- addedenhancementNew feature or requestNew feature or requestarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connectionsfrom-jiraMigrated or originated from a Jira ticketMigrated or originated from a Jira ticket
on May 22, 2026 Adopted as W5 of the replication epic #430. Capturing the concrete build breakdown from the architecture review, in rough dependency order:
The good news: the measurement, the latency ranking, and the re-route machinery mostly already exist. Back-pressure ratio is computed today (sender writes it at
replicationConnection.ts:1012); per-upstream latency is already tracked and already used byReplicator.loadfor cache-miss routing; and the disconnect-failover path already re-points a subscription through another node. The missing core is the bridge (sender-thread back-pressure → sender main → wire → subscriber main → re-subscribe-indirect), the cross-node "shed" directive, and cycle safety.The single biggest de-risker is W1 (#431) — its shared-memory connection/health registry is exactly the channel this needs to get back-pressure to the main thread. Build W5 on top of W1 and pieces 1–2 below become nearly free.
- Egress back-pressure → main-thread channel. Today back-pressure is written to the shared buffer and read only by
cluster_status— it never reaches the orchestrator. Add a worker→main report (edge-triggered on threshold cross, mirroring the existingconnected/disconnected-from-nodeIPC), or have main read it from the shared-memory registry once W1 lands. - Per-subscriber inventory on the sending worker → main. The orchestrator tracks upstream nodes only; to name "which connection is back-pressured" it needs the sender's downstream
(db, subscriberNode)list. The data exists inreplicateOverWS(nodeSubscriptions); it just isn't reported. - Threshold + hysteresis (enter/exit) on the back-pressure ratio so topology decisions don't flap. Today there's only a decaying ratio with no decision boundary.
- Cross-node "shed" directive (the genuinely new protocol surface). A frame for the back-pressured sender to tell its subscriber: "suspend your direct subscription to node H (your highest-latency upstream), re-subscribe indirectly via node X." Can potentially reuse the
excluded-nodes machinery already inSUBSCRIPTION_REQUEST. - Subscriber-side re-route execution. Pick the highest-latency upstream (slot 4, already populated), suspend that
(db,node)subscription, issue the indirect subscription — generalizing the existing disconnect-failover re-pointing into a directive-driven path. - Resume-direct on decay. When the ratio falls below the exit threshold, tear down the indirect hop and re-subscribe direct (reuses the failover-restore logic).
- Loop/convergence safety. Indirect fan-out must not create subscription cycles — the existing
excluded-nodes list and origin-loop prevention (SKIPPED_MESSAGE_SEQUENCE_UPDATE_DELAY) are building blocks, but adaptive re-routing adds new cycle risk that needs explicit guarding. - Observability + tests. Surface the active topology (direct vs indirect-via-X) in
cluster_status; add cluster integration tests (the existingreplicationLoad/replicationTopologysuites are the natural homes).
Note: landing protocol version negotiation (W11 / #440) before piece 4 is strongly advised, so the new directive frame rolls out safely across mixed-version clusters.
🤖 Filed by Claude on behalf of Kris.
- Egress back-pressure → main-thread channel. Today back-pressure is written to the shared buffer and read only by
- added a commit that references this issue
on Sep 4, 2026 - added a commit that references this issue
on Sep 8, 2026 - added a commit that references this issue
on Sep 9, 2026 - added a commit that references this issue
on Sep 9, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Overview
Dynamically adjust replication routing topology based on back-pressure from egress network limits, switching to fan-out only when needed.
Design
Under normal conditions, each node replicates directly to all subscribers (lowest latency). When egress limits are hit:
Key constraint
Harper doesn't proactively push replication — nodes subscribe to receive it. Back-pressure detection must bridge the gap between the sending thread (which sees network limits) and the main thread (which manages topology).
🤖 Filed by Claude on behalf of Kris.