Repository navigation
Replication full-copy OOM: per-record audit-chain walk in Table.commit scales with audit-log depth #1114
Description
Activity
The unbounded system-DB log growth referenced under "Contributing factor" now has its own issue: #1115 (
hdb_analyticsisaudit: true, replicating per-node telemetry cluster-wide). That growth is the volume that makes this audit-walk OOM fatal during full-copy re-sync.- added a commit that references this issue
on Jun 5, 2026 Proposal: move write de-duplication to the replication ingest layer; simplify full-copy to LWW-merge
Context: investigation with @kriszyp into whether the duplicate-detection burden in
Table.commit(the capped out-of-order walk from #1116/#1122, plus the #1137/#1147 follow-ups) can move to the replication layer, leavingTable.tsto focus on CRDT reconciliation. Trace ofharper-pro/replication/*+harper/resources/Table.tsbelow, then a concrete design.What the trace established
-
Full-copy sends only synthetic
puts, never patches — it iteratesprimaryStore.getRange({ versions: true })and emitstype:'put',previousVersion:null, with origin rewritten to the sender (replicationConnection.ts:1461-1516). Puts don't accumulate, so full-copy itself can't double-apply a commutative op. The re-delivered commutative patch in Unit test red on main: 're-delivered duplicate must not double-apply the commutative op' (v22 + v26) #1137 arrives via the steady-state audit stream that follows (overlap/resume window) or via transitive relay — both carry(originNodeId, version)at the receiver before dispatch (replicationConnection.ts:1646-1696). -
Replication full-copy OOM: per-record audit-chain walk in Table.commit scales with audit-log depth #1114's OOM is the out-of-order fold walking the receiver's deep local audit chain when a full-copy put lands older than local head (
Table.ts:1738-1898; the walk seeds fromexistingEntry.localTimeand scans backward). -
A durability asymmetry constrains where dedup state may live.
dbisDbis RocksDB-WAL-durable (immediate); the primary store's durability is Harper's txn log (separate mechanism). So a dedup mark placed indbisDbcan outlive the data it guards — on crash the mark says "seen" while the write is gone and is never re-requested → silent loss. Dedup state must be no more durable than the data: derived from the primary/txn-log domain and co-committed (the atomic-CF-write flag gives the atomicity), or read-your-writes against the record. This is also why fix(table): reliably skip a re-delivered out-of-order commutative op (#1137) #1147's record-basedadditionalAuditRefscheck works where theauditStore.get(txnTime)lookup lags (Transaction-log point reads (RocksTransactionLogStore) intermittently miss visible entries, breaking out-of-order duplicate detection #1148). -
Exclusion is node-granular, not a windowed partition (
replicationConnection.ts:2110-2137, filter at:1151-1158). The transitive case (A⊂B,C; B,C⊂D) yields duplicates, not gaps — A receives D's full stream from both B and C, each gap-free and in D's order. Genuine gaps only on (a) time-window boundaries (not currently constructed) and (b) topology-change handoffs, where gap-freeness rests on the proxy seq-tracking (nodes[]in the seq entry —Table.ts:425-438,replicationConnection.ts:2069-2089).
Proposal
A. Ingest-layer de-duplication, keyed on
(originNodeId, version).
At the receiver, before dispatch (replicationConnection.ts:1646-1696), drop already-applied writes. State lives in the primary/txn-log durability domain, co-committed with the record (notdbisDb). Because mark and value revert together on crash, this is safe even for non-idempotent commutative ops.- A bare per-origin max high-water mark is unsafe under any gap (it would drop an in-gap late arrival as a false duplicate). Since handoff gap-freeness isn't provable, track contiguity instead — collapse the durable contiguous prefix to one number, keep a bounded seen-set over the overlap/reconnect window above it.
- Exempt full-copy puts (origin-rewritten, key-ordered → non-monotonic); they flow to LWW (B).
- This implements Resync re-delivery (~6.7×) + cleanup starvation balloon a far-behind node's transaction logs (compounds #1114 OOM) #1115's suggested fix Add issue templates #1 ("make resync idempotent: per-peer received position / dedup"), subsumes Unit test red on main: 're-delivered duplicate must not double-apply the commutative op' (v22 + v26) #1137/fix(table): reliably skip a re-delivered out-of-order commutative op (#1137) #1147, and routes around Transaction-log point reads (RocksTransactionLogStore) intermittently miss visible entries, breaking out-of-order duplicate detection #1148 (no dependence on the lagging point reads).
B. Full-copy: lose-on-tie + merge-when-older (keep the merge).
- tie (
version == existing) → drop early, before the walk, so repeated full-copy is a no-op on copied records (no walk, no Replication full-copy OOM: per-record audit-chain walk in Table.commit scales with audit-log depth #1114 cost). - strictly older → still fold/merge to honor local writes the snapshot hasn't seen (cap-bounded, as today).
- newer → normal apply.
precedesExistingVersionstays only for genuine same-version different-origin ties (LWW), which is conflict resolution, not dedup.
C. Drop the per-record receiver-side audit entry for full-copy; leave a single boundary marker.
Full-copy records become durable via flush + re-fetch-from-leader (lose-on-tie makes re-sync idempotent), so the per-record txn-log entry is unnecessary — a large write-amplification cut on catch-up (#1114). But incremental relay reads the audit log (replicationConnection.ts:1547-1557), so dropping the entries outright would silently gap a downstream resuming from before the full-copy. Mitigate with one full-copy boundary marker in the audit log: a downstream whose resume point is below it must re-bootstrap (full-copy) rather than incremental; above it streams normally. Keeps the savings; keeps relay correct.Net
- The
Table.tsout-of-order block reduces to genuine CRDT patch reconciliation (shallow in steady state). The depth cap (fix(table): bound out-of-order audit-chain walk to prevent full-copy OOM #1116/fix(table): bound out-of-order audit-chain walk to prevent full-copy OOM #1122) may become unnecessary once full-copy ties short-circuit and patch dedup moves to ingest. - Unit test red on main: 're-delivered duplicate must not double-apply the commutative op' (v22 + v26) #1137/fix(table): reliably skip a re-delivered out-of-order commutative op (#1137) #1147 resolved at the layer that owns delivery; Transaction-log point reads (RocksTransactionLogStore) intermittently miss visible entries, breaking out-of-order duplicate detection #1148's read-reliability is no longer load-bearing for correctness.
Open questions / follow-ups
- Handoff gap-freeness — confirm the proxy
nodes[]seq-tracking guarantees a new direct path resumes at-or-before where the relayed path stopped. This is the one place a gap (→ data loss) could hide, and it's the same invariant the existing resume watermark already trusts. - Boundary-marker semantics for relay re-bootstrap (per-table vs per-db; interaction with residency).
- Co-commit mechanics for the dedup contiguity state in the primary domain.
Relationship to existing work: builds on #1116/#1122 (the cap), supersedes the #1147 approach for #1137, implements #1115's primary fix, and makes #1148 non-blocking for correctness.
— Claude (investigation with Kris)
-
- added 5 commits that reference this issue
on Jun 9, 2026 Live-production confirmation on
eh-prod.gend+ root cause of the event-loop stallsInvestigated this on the
eh-prod.gendcluster (node7rg-us-west-1, harper-pro 5.1.1), where the symptom presented as replication ping-timeouts / wedged subscriptions rather than OOM. The depth cap added for this issue stops the OOM, but the synchronous walk itself stalls the event loop (steadyJavaScript execution has taken too longwarnings +Timeout waiting for ping→ terminated replication connections). Captured the live mechanism with a non-pausing logpoint at the depth-cap site (Table.js:1902).What the walk is actually doing (logpoint evidence)
Every capped event looked like:
id=<...norton.com/blog/...> depth=1001 type=patch fullUpdate=false addRefs=0 toVisit=0 succ=858–1000 srcNode=4|8|9 viaNodeId=<set> dupFound=true dupVerMatch=true dupNode===srcNodeReading that off:
- It's a genuine deep linear chain, not ref fan-out.
additionalAuditRefs/auditRefsToVisitare empty (addRefs=0,toVisit=0); the depth is almost entirelysucceedingUpdates= 858–1000. The records are high-churn pages updated exclusively viapatch(scheduledlastRefresh/nextRefreshbumps), so there's noputin history to short-circuit the walk on. - The triggering writes are out-of-order re-deliveries via transitive replication.
type=patch, older than the current record head (they reach theprecedesExisting <= 0block and walk ~1000 steps without finding a head-tie), and every one carriesviaNodeId— relayed through a proxy node. - They are pure duplicates. The post-cap keyed lookup
auditStore.get(txnTime, tableId, id, nodeId)returns a matching(version, nodeId)entry (dupVerMatch=true,dupNode===srcNode). So the write was already applied; the walk burns ~1000 synchronous steps only to discard it.
Why PR #370 (leading-duplicate fast-skip) doesn't cover it
Two reasons:
- It deliberately doesn't arm for proxied/indirect subscriptions.
hasPersistedResumeCursorchecks the direct sequence cursor only, and the arming code explicitly leaves a proxy-derived resume un-armed. All the re-deliveries here arrive viaviaNodeId(proxy), so the fast-skip never engages. - Even if armed, its tie check is against the record head (
existing.version === incomingVersion). These re-deliveries are older than the head (the head advanced via other paths before the proxied copy arrived), so they're buried-history matches, not head-ties — the head-tie check can't catch them regardless of arming.
Proposed fix (primary)
Hoist the keyed duplicate lookup to before the resequencing walk. The cap block already does
auditStore.get(txnTime, tableId, id, options?.nodeId)and we've confirmed it identifies these as exact(version, nodeId)duplicates — it's just invoked after the ~1000-step walk. Doing that O(1) keyed lookup up front (for theprecedesExisting <= 0replicated path, with the sameprecedesExistingVersion(...) === 0identity-tie guard) short-circuits the re-delivery before the walk. It is keyed bynodeId, so it is inherently multi-source — no single node can "break" it. A miss simply falls through to today's walk, so there is no correctness change; the existing keyed-lookup-can-miss-under-load caveat only costs us the optimization, never correctness.Follow-ups (separate tickets)
- Reduce the volume of transitive re-deliveries. The proxy-resume start-time derivation re-streams an already-applied tail when the proxied seq cursor isn't relayed/persisted tightly. Worth a separate investigation to cut re-deliveries at the source (the dedup fix above cheaply absorbs them; this would stop generating them).
- Extend Add loadComponent config option for conditional package loading #370 arming to proxied subscriptions — a cheaper still earlier skip for the in-order proxied subset (complementary; does not replace the core hoist, which is what covers the out-of-order case observed here).
- Record-level LWW for full-copy base frames — separate ticket (ref harper-pro feat: expose urlPath in deploy_component operation and CLI #1113); orthogonal to this, since the cap walks here are relayed patches, not copy base puts.
Implementing the primary fix now.
Investigation + writeup by Claude (Opus 4.8) via the fabric-investigation workflow.
- It's a genuine deep linear chain, not ref fan-out.
- added a parent issue
on Jul 7, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
During replication full-copy catch-up, a receiving node can OOM because
Table.commit's out-of-order-write reconciliation walks the entire backward audit-history chain for every record synchronously, buffering intermediate records in the JS heap. The cost scales with audit-history depth, and the walk runs concurrently on all HTTP workers, so a database whose replication transaction log has accumulated a large history becomes effectively un-ingestable under a fixed container memory limit.Observed on
harper-pro5.0.21; the same code path is present onmain(v5.1.0-beta.1).Symptom
Killed process … (MainThread) anon-rss ~5.4GBat a 6 GiB limit).JavaScript execution has taken too long and is not allowing proper event queue cyclingduring ingestion.docker psshows "Up" andRestartCountstays frozen — the kills are only visible viadmesg.Root cause
Table.commitreconciles out-of-order / incremental writes by walking the record's audit chain (resources/Table.ts:1769-1837):Each
auditStore.get(localTime, …)lands inRocksTransactionLogStore.getSync, which does a fresh range scan of the transaction log per step (resources/RocksTransactionLogStore.ts:100-112):Why this OOMs during full-copy:
auditRefsToVisit), with a msgpackr-decode + txnlog scan at each step. During a full-copy of a database with a large accumulated history, this repeats for many records.succeedingUpdatesaccumulates audit records, andgetValue(primaryStore)is later materialized for each.Evidence (live instance)
process.memoryUsage()): memory is JS heap (heapUsed0.6–1.2 GB on each of 8 workers), not native buffers (external/arrayBuffers≤ 330 MB). This rules out a single bad-allocation (e.g. txnlog framing) cause.readAuditEntry/onWSMessage.Table.commit→RocksTransactionLogStore.get/getSync→ transaction-log-reader iteration, underDatabaseTransaction.commit.Contributing factor: unbounded system-DB replication log growth
The trigger in the observed case was the system database's per-peer replication transaction log growing to ~2 GB on the follower (and ~1 GB on the leader) while the actual system data was only ~250 MB. The leader's own-origin (
local) system log was only ~25 MB, so the growth is in the per-peer replication logs, not genuine local writes — suggesting a retention/pruning gap and/or a write-amplification feedback loop (a persistentinvalid user role founderror storm accompanied it). Even if the audit-walk is made memory-safe, the unbounded replication-log growth is worth investigating separately.Suggested fix directions
Table.commitmemory-safe: yield to the event loop and/or cap/streamsucceedingUpdatesso a deep history doesn't pin the heap, and let GC run between commits.Workaround
Temporarily raising the container memory limit (e.g. 6 → 12 GiB) lets the full-copy complete once; steady-state then fits. This is mitigation, not a fix.