Skip to content

Removed node's transaction log is replayed from position 0 on every resubscribe; idle peers' cursors never advance and force a base copy #989

Description

@kriszyp

Summary

After a node is removed with remove_node, every later resubscribe on every surviving node (a restart, a deploy, a reconnect) replays that removed node's entire retained per-origin transaction log to the subscriber, starting at position 0. The receiver applies those entries and passes them on to its own subscribers, so each restart triggers a CPU spike across the whole cluster.

A second, related cursor gap: a live peer that has no writes of its own never advances its receivers' resume cursor. Once that cursor is older than auditRetention, the next reconnect is upgraded to a full base copy.

Both are cursor-tracking defects. The fix is per-origin cursor tracking, not deleting log data. A removed member's earlier writes stay valid and must keep replicating (see the discussion on #686).

Production evidence

20-node mesh cluster on RocksDB, after two region-removal rounds:

  • One node restart (about 2 minutes offline).
    • Each of the 19 peers sent it about 35k old entries, 5 days old on average.
    • It sent 35–43k entries back to each peer.
    • Harper CPU on that node went to 13.7 of 16 cores, and to 1–3.5 cores on every other node (normal is 0.1–0.4).
    • The table involved holds about 1,500 rows.
  • Plain restarts before any removal replayed nothing. A rolling restart of all 20 nodes before any removals applied zero entries older than an hour.
  • Adding 10 cloned nodes made every existing node re-apply about 1.1M entries, up to 64 days old. The new nodes pulled the removed nodes' logs starting at position 0 and relayed them back out.
  • Removed nodes' logs are still being appended to days after removal, past a 3-day auditRetention.
  • The receiver keeps 32 orphaned resume-cursor rows (seq) for removed nodes, frozen at their removal times. Nothing reads them (see below).
  • Cursors for live peers with no writes of their own stayed frozen for hours while connected.

The restart-triggered replay of dead-node records described on #686 is very likely the same mechanism.

Mechanism

References are to harper-pro 0308c698 and core 5230512f9.

  1. The receiver's subscription request covers only hdb_nodes members.
    • It sends one cursor per subscribed node (replicationConnection.ts around line 8128 and lines 8199–8259).
    • It sends an exclusion list: itself plus computeExclusionOrigins() (subscriptionManager.ts:979), which reads only hdb_nodes.
    • A removed node R is in neither list.
  2. The sender bounds only the peer's own log.
    • It opens one aggregate range with startByLog: new Map([[logName, currentSequenceId]]) and excludeLogs: excludedNodes (replicationConnection.ts:6517-6526).
    • In core, any log that is not excluded and not named in startByLog gets start: 0 (RocksTransactionLogStore.ts:431).
  3. Nothing filters those origins later.
    • matchesSubscription passes any origin that is neither listed nor excluded without checking a time range (replicationConnection.ts:5740).
    • After remove_node, the includeNodes handling also re-adds R's log to a running iterator, again from 0 (addLog, replicationConnection.ts:5494).
  4. The receiver has no cursor for R that could bound the range.
    • The [seq, peerId] row in __dbis__ advances seqId for the directly connected peer.
    • Its nodes[] list advances only for remoteNodeIds.slice(1), i.e. subscribed proxied nodes (core/resources/Table.ts:2111-2149).
    • R's own row was last written while R was a direct peer, and nothing reads it after removal.
  5. The replay feeds itself.
    • A replayed entry that misses dedup is written again into R's log on the receiver (Table.ts around lines 4734–4836).
    • That covers a record whose stored version is older or missing, and a superseded entry that gets an audit-only append.
    • The receiver then relays those entries on its next resubscribe, and the re-written entries outlive the time-based purge.

Idle peers.

  • The sender sends SEQUENCE_ID_UPDATE only from skipAuditRecord's timer (replicationConnection.ts:5826-5840), and REMOTE_SEQUENCE_UPDATE only after it has sent data.
  • A peer with no writes of its own, whose other origins are all excluded, does neither. Its receivers' cursor for it never moves.
  • shouldForceBaseCopyForRetention (replicationConnection.ts:549, :6111-6137) then turns the next reconnect after auditRetention into a full base copy.

Fix

This is the first scope item of W4 (#434), "make nodes[] per-source cursors the primary resume mechanism (per-origin cursor vector via startByLog)", limited to what this defect needs.

  1. Receiver: keep a cursor for every origin it applies.
    • Record each origin it applies in the per-peer seq row's nodes[], not just subscribed proxied ones. That includes relayed origins and origins no longer in hdb_nodes.
    • The cursor must be the entry's origin log key (txnLogKey in the origin's log), not the receiver's local time.
  2. Send the cursors as the W4 per-origin cursor vector.
    • Add an { originName: cursor } map to the subscription request.
    • Negotiate it through protocolCapabilities.ts. Peers without the capability keep today's behaviour.
    • This must be the field that W4 builds on, not a one-off.
  3. Sender: use the vector everywhere the range starts.
    • Fill startByLog from the vector for every origin that is neither listed nor excluded, starting a fixed overlap window below the cursor (60 s to start).
    • Apply the same bound in matchesSubscription and in the includeNodes/addLog path.
    • Duplicates inside the overlap are acceptable. Skipping an entry is not: logs are written out of key order (Replication resume cursors need an append-order mode to survive a reconnect or a relayed origin #879, harper#2928), and today's start at 0 never skips.
    • An origin with no cursor still starts at 0, which is correct for an origin never seen before.
  4. Upgrade: seed the vector from the orphaned rows. When a removed origin has no vector entry but has its old direct-peer seq row, use that row's seqId. Otherwise the first restart after upgrading replays one more time.
  5. Idle peers: advance the cursor when the sender has nothing to send.
  6. Retention: once replay from 0 stops, check whether removed-node logs purge on their own.

Tests

  • Removed-node replay.
    • Set up a 3+ node mesh with writes from each node, then remove_node one of them.
    • Restart a survivor.
    • Assert that each peer sends near-zero entries older than the overlap window, and that the removed node's log on the survivors does not grow.
    • Repeat with a node added after the removal; this is the clone case.
  • Removed node's last writes. Assert that writes R made just before removal, which had not reached every node yet, still arrive at every node.
  • Idle peer.
    • Use one peer with no writes of its own, a short auditRetention, and wait past it.
    • Reconnect, and assert an incremental resume with no base copy.
  • Mixed versions. A peer without the capability falls back to today's behaviour.

🤖 Investigated and filed by Claude (Opus 5.5) on Kris's behalf

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P1

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions