Skip to content

Replication: transitive/proxied re-delivery floods peers with already-applied out-of-order writes (reduce volume; complements harper#1310) #399

Description

@kriszyp

Summary

Transitive/proxied replication re-delivers already-applied writes to peers out of order, generating a high volume of redundant out-of-order applies. On records with deep audit histories this drove the core out-of-order resequencing walk to its depth cap repeatedly, starving the event loop (→ replication ping-timeouts, wedged subscriptions). Observed in production on [customer-cluster].

This is the "reduce the volume" counterpart to harper#1310 (which absorbs these cheaply in core via an up-front keyed dedup). #1310 stops them from being catastrophic, but they're still generated and still cost work every cycle — this issue tracks stopping them at the source.

Evidence ([customer-cluster], [node] on 5.1.1)

A non-pausing logpoint at the core depth-cap site captured the capping writes:

  • type=patch, fullUpdate=false, every one carrying viaNodeId (relayed via a proxy node, not direct).
  • Multiple source nodes (srcNode 4/8/9), all older than the local record head (out-of-order).
  • The post-cap keyed lookup matched (dupVerMatch=true, dupNode===srcNode) — i.e. they were exact (version, nodeId) duplicates already applied, redundantly re-delivered via the relay path, each forcing an ~858–1000-step walk before being discarded.

So the re-deliveries are real, redundant, and proxied.

Suspected mechanism

The proxy-resume start-time derivation in replication/replicationConnection.ts (~the indirect-connection block that reads a proxy node's seqId, around the Using sequence id from proxy node path) appears to re-stream a longer already-applied tail than necessary — when the proxied seq cursor isn't relayed/persisted as tightly as the direct cursor, the leader streams from too-early a point and re-delivers writes the follower already has. (High confidence on the symptom — proxied out-of-order duplicate re-delivery — from the logpoint; the exact start-time-relay gap as the generator is the leading hypothesis from reading the resume code, not yet confirmed end-to-end.)

Relationship to existing work

Proposed directions

  1. Tighten the proxy seq-cursor relay so a proxied resume starts from an accurate point and doesn't re-stream an already-applied tail (attacks the cause).
  2. Arm perf(replication): fast-skip leading duplicates on resume #370's leading-dup-skip for proxied subscriptions (per-origin-node cursor from the proxy seqId) — earlier skip for the in-order subset.
  3. Consider whether redundant transitive delivery is needed at all when a follower already receives the writes directly (reduce proxied fan-in).

Filed from the [customer-cluster] incident investigation. Investigation + writeup by Claude (Opus 4.8).

Activity

  1. self-assigned this
    on Jun 16, 2026
  2. kriszyp commented on Jun 16, 2026

    @kriszyp
    MemberAuthor

    Direction 2 implemented in #401 (arms the leading-duplicate fast-skip for proxied subscriptions; drops the re-streamed already-applied tail cheaply at the receive layer).

    Root-cause finding while investigating direction 1 — the proxied resume override is dead code. The indirect-connection block in replicationConnection.ts matched the relayed per-source cursor by seqNode.name === node.name, but core's updateRecordedSequenceId only ever persists { id, seqId, lastTxnTime } (never name). So seqNode.name is always undefined and the override never fired — every proxied resume has silently fallen back to the lastTxnTimes overlap / Date.now() - 60000 start, which is exactly the "stream from too-early a point" behavior described here.

    Direction 1 (tighten the proxied resume start) carries a real data-loss risk and is intentionally not in #401. The relayed seqId is the proxy connection's local time (Math.max(existingSeq.seqId, event.localTime)), which for an out-of-order source can sit ahead of what the follower has actually applied for that source. Raising startTime to it as the leader's resume lower bound could skip un-applied writes. The smallest safe shape is: fix the match to by-id and keep the lastTxnTimes overlap as the data-loss floor (startTime = max(proxiedSeqId, lastTxnTime)), gated behind a 3-node-relay repro that also asserts a write committed-but-not-yet-applied during the resume is still delivered.

    #401 deliberately leaves startTime untouched (so no behavior change vs main); it only adds the receive-layer skip. This issue stays open to track direction 1.

    Posted by Claude (Opus 4.8) on behalf of Kris.

  3. added this to the v5.3 milestone on Aug 7, 2026
  4. added theissue type on Aug 7, 2026
  5. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P2

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions