You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Removed node's transaction log is replayed from position 0 on every resubscribe; idle peers' cursors never advance and force a base copy #989
After a node is removed with remove_node, every later resubscribe on every surviving node (a restart, a deploy, a reconnect) replays that removed node's entire retained per-origin transaction log to the subscriber, starting at position 0. The receiver applies those entries and passes them on to its own subscribers, so each restart triggers a CPU spike across the whole cluster.
A second, related cursor gap: a live peer that has no writes of its own never advances its receivers' resume cursor. Once that cursor is older than auditRetention, the next reconnect is upgraded to a full base copy.
Both are cursor-tracking defects. The fix is per-origin cursor tracking, not deleting log data. A removed member's earlier writes stay valid and must keep replicating (see the discussion on #686).
Production evidence
20-node mesh cluster on RocksDB, after two region-removal rounds:
One node restart (about 2 minutes offline).
Each of the 19 peers sent it about 35k old entries, 5 days old on average.
It sent 35–43k entries back to each peer.
Harper CPU on that node went to 13.7 of 16 cores, and to 1–3.5 cores on every other node (normal is 0.1–0.4).
The table involved holds about 1,500 rows.
Plain restarts before any removal replayed nothing. A rolling restart of all 20 nodes before any removals applied zero entries older than an hour.
Adding 10 cloned nodes made every existing node re-apply about 1.1M entries, up to 64 days old. The new nodes pulled the removed nodes' logs starting at position 0 and relayed them back out.
Removed nodes' logs are still being appended to days after removal, past a 3-day auditRetention.
The receiver keeps 32 orphaned resume-cursor rows (seq) for removed nodes, frozen at their removal times. Nothing reads them (see below).
Cursors for live peers with no writes of their own stayed frozen for hours while connected.
The restart-triggered replay of dead-node records described on #686 is very likely the same mechanism.
Mechanism
References are to harper-pro 0308c698 and core 5230512f9.
The receiver's subscription request covers only hdb_nodes members.
It sends one cursor per subscribed node (replicationConnection.ts around line 8128 and lines 8199–8259).
It sends an exclusion list: itself plus computeExclusionOrigins() (subscriptionManager.ts:979), which reads only hdb_nodes.
A removed node R is in neither list.
The sender bounds only the peer's own log.
It opens one aggregate range with startByLog: new Map([[logName, currentSequenceId]]) and excludeLogs: excludedNodes (replicationConnection.ts:6517-6526).
In core, any log that is not excluded and not named in startByLog gets start: 0 (RocksTransactionLogStore.ts:431).
Nothing filters those origins later.
matchesSubscription passes any origin that is neither listed nor excluded without checking a time range (replicationConnection.ts:5740).
After remove_node, the includeNodes handling also re-adds R's log to a running iterator, again from 0 (addLog, replicationConnection.ts:5494).
The receiver has no cursor for R that could bound the range.
The [seq, peerId] row in __dbis__ advances seqId for the directly connected peer.
Its nodes[] list advances only for remoteNodeIds.slice(1), i.e. subscribed proxied nodes (core/resources/Table.ts:2111-2149).
R's own row was last written while R was a direct peer, and nothing reads it after removal.
The replay feeds itself.
A replayed entry that misses dedup is written again into R's log on the receiver (Table.ts around lines 4734–4836).
That covers a record whose stored version is older or missing, and a superseded entry that gets an audit-only append.
The receiver then relays those entries on its next resubscribe, and the re-written entries outlive the time-based purge.
Idle peers.
The sender sends SEQUENCE_ID_UPDATE only from skipAuditRecord's timer (replicationConnection.ts:5826-5840), and REMOTE_SEQUENCE_UPDATE only after it has sent data.
A peer with no writes of its own, whose other origins are all excluded, does neither. Its receivers' cursor for it never moves.
shouldForceBaseCopyForRetention (replicationConnection.ts:549, :6111-6137) then turns the next reconnect after auditRetention into a full base copy.
Fix
This is the first scope item of W4 (#434), "make nodes[] per-source cursors the primary resume mechanism (per-origin cursor vector via startByLog)", limited to what this defect needs.
Receiver: keep a cursor for every origin it applies.
Record each origin it applies in the per-peer seq row's nodes[], not just subscribed proxied ones. That includes relayed origins and origins no longer in hdb_nodes.
The cursor must be the entry's origin log key (txnLogKey in the origin's log), not the receiver's local time.
Send the cursors as the W4 per-origin cursor vector.
Add an { originName: cursor } map to the subscription request.
Negotiate it through protocolCapabilities.ts. Peers without the capability keep today's behaviour.
This must be the field that W4 builds on, not a one-off.
Sender: use the vector everywhere the range starts.
Fill startByLog from the vector for every origin that is neither listed nor excluded, starting a fixed overlap window below the cursor (60 s to start).
Apply the same bound in matchesSubscription and in the includeNodes/addLog path.
An origin with no cursor still starts at 0, which is correct for an origin never seen before.
Upgrade: seed the vector from the orphaned rows. When a removed origin has no vector entry but has its old direct-peer seq row, use that row's seqId. Otherwise the first restart after upgrading replays one more time.
Idle peers: advance the cursor when the sender has nothing to send.
When the sender has caught up and is waiting for new transactions (replicationConnection.ts:6575-6585), send SEQUENCE_ID_UPDATE carrying the latest committed key of its own log, whenever that key is past sentSequenceId. Use the log key, not Date.now().
Summary
After a node is removed with
remove_node, every later resubscribe on every surviving node (a restart, a deploy, a reconnect) replays that removed node's entire retained per-origin transaction log to the subscriber, starting at position 0. The receiver applies those entries and passes them on to its own subscribers, so each restart triggers a CPU spike across the whole cluster.A second, related cursor gap: a live peer that has no writes of its own never advances its receivers' resume cursor. Once that cursor is older than
auditRetention, the next reconnect is upgraded to a full base copy.Both are cursor-tracking defects. The fix is per-origin cursor tracking, not deleting log data. A removed member's earlier writes stay valid and must keep replicating (see the discussion on #686).
Production evidence
20-node mesh cluster on RocksDB, after two region-removal rounds:
auditRetention.seq) for removed nodes, frozen at their removal times. Nothing reads them (see below).The restart-triggered replay of dead-node records described on #686 is very likely the same mechanism.
Mechanism
References are to harper-pro
0308c698and core5230512f9.hdb_nodesmembers.replicationConnection.tsaround line 8128 and lines 8199–8259).computeExclusionOrigins()(subscriptionManager.ts:979), which reads onlyhdb_nodes.startByLog: new Map([[logName, currentSequenceId]])andexcludeLogs: excludedNodes(replicationConnection.ts:6517-6526).startByLoggetsstart: 0(RocksTransactionLogStore.ts:431).matchesSubscriptionpasses any origin that is neither listed nor excluded without checking a time range (replicationConnection.ts:5740).remove_node, theincludeNodeshandling also re-adds R's log to a running iterator, again from 0 (addLog,replicationConnection.ts:5494).[seq, peerId]row in__dbis__advancesseqIdfor the directly connected peer.nodes[]list advances only forremoteNodeIds.slice(1), i.e. subscribed proxied nodes (core/resources/Table.ts:2111-2149).Table.tsaround lines 4734–4836).Idle peers.
SEQUENCE_ID_UPDATEonly fromskipAuditRecord's timer (replicationConnection.ts:5826-5840), andREMOTE_SEQUENCE_UPDATEonly after it has sent data.shouldForceBaseCopyForRetention(replicationConnection.ts:549,:6111-6137) then turns the next reconnect afterauditRetentioninto a full base copy.Fix
This is the first scope item of W4 (#434), "make
nodes[]per-source cursors the primary resume mechanism (per-origin cursor vector viastartByLog)", limited to what this defect needs.seqrow'snodes[], not just subscribed proxied ones. That includes relayed origins and origins no longer inhdb_nodes.txnLogKeyin the origin's log), not the receiver's local time.{ originName: cursor }map to the subscription request.protocolCapabilities.ts. Peers without the capability keep today's behaviour.startByLogfrom the vector for every origin that is neither listed nor excluded, starting a fixed overlap window below the cursor (60 s to start).matchesSubscriptionand in theincludeNodes/addLogpath.seqrow, use that row'sseqId. Otherwise the first restart after upgrading replays one more time.replicationConnection.ts:6575-6585), sendSEQUENCE_ID_UPDATEcarrying the latest committed key of its own log, whenever that key is pastsentSequenceId. Use the log key, notDate.now().Tests
remove_nodeone of them.auditRetention, and wait past it.🤖 Investigated and filed by Claude (Opus 5.5) on Kris's behalf