Repository navigation
Replicated hdb_analytics floods system transaction logs and spins a worker in native code; system base-copy wedges until worker recycle #480
Description
Activity
Why it never converges: the base-copy cursor never finalizes → forces full re-copy on restart (self-perpetuating)
Traced the copy-resume cursor mechanism in
dist/replication/replicationConnection.js:- The resume cursor is persisted at
dbisDB[Symbol.for('copyCursor'), nodeId]as{copyStartTime, currentTable, afterKey, copyOrder}, advanced per copied record (flushDurableCopyCursor), and cleared only onCOPY_COMPLETE(maybeFinishCopy). The code comment is explicit: if the cursor loops, "the cluster never converges, and COPY_COMPLETE never arrives to clear the cursor." - Because the
systemcopy can't get pasthdb_analytics,COPY_COMPLETEnever arrives → the cursor never finalizes → the socket staysconnected:true/Receiving/lastReceivedVersion:0indefinitely (the seemingly-"cosmetic" ver=0 socket). - On reconnect/restart, an initial copy that never completed deliberately triggers a full re-copy (not an incremental resume) — see the
#426data-loss safeguard at ~L3120-3196: "when an initial copy never completed … the node then resumes from now-60s and permanently loses the gap … A full copy is idempotent … the cost of an occasional extra copy is acceptable versus silent data loss" →startTime = 0(full copy).
So the failure is self-perpetuating: replicating
hdb_analyticsmakes thesystembase-copy too big to finish → cursor never finalizes → any restart/reconnect re-requests a full copy of the wholesystemDB (incl. the 3.1M-row analytics table) → native txn-log scan spins the worker again → still never completes → cursor still never finalizes. Each restart re-pays the full analytics-copy cost.This is why
analytics_replicate: falseis the clean fix: with analytics out of replication, thesystembase-copy is tiny (control-plane only),COPY_COMPLETEarrives, the cursor finalizes/clears, and subsequent reconnects resume incrementally instead of full-recopying.The cursor is inspectable live:
databases.system.<anyTable>.dbisDB.getRange()filtered to keys whose first element isSymbol.for('copyCursor')returns{node, cursor:{currentTable, afterKey, copyOrder, copyStartTime}}per in-flight source. (Present only while a copy is mid-flight; cleared on completion.)- The resume cursor is persisted at
- added 6 commits that reference this issue
on Jun 25, 2026 Fixed by #486 (merged 2026-06-25) — copyApply snapshots base-copy rows without audit entries (RocksDB-only, durability-gated).
— Claude (Opus 4.8), issue-backlog triage on Kris's behalf
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
On a 4-node v5.1.12 cluster, the
systemdatabase base-copy between some node pairs never converges. A replication worker pins ~100% CPU in native code (invisible to JS/V8 profilers) while the affectedsystemsocket makes zero forward progress (connected: true,status: Receiving,lastReceivedVersion: 0,lastReceivedRemoteTimefrozen). It recovers only when the http worker pool is recycled — i.e. it clears episodically and then re-wedges. The driver is replicatinghdb_analytics(on by default), which floods the per-peertransaction_logsand the system base-copy.No data loss: the affected node stays current on
systemvia its other peers; it's a redundant copy direction that wedges while burning a core.Environment
analytics_replicateunset → default (analytics replicates)Evidence
systemdirection showsconnected: true,status: Receiving,lastReceivedVersion: 0,lastReceivedRemoteTimefrozen ~23h (pre-restart),backPressurePercent: 0. The othersystemsockets on that node are current → the node's data is converged; only the redundant direction is wedged.top -Hshows onehttpworker (intermittently also the main thread) pegged ~100% CPU, sustained >1h. Yet V8 CPU profiles of all worker threads show ~100% idle. The built-in analytics self-profiler (analytics/profile.ts) corroborates: of the ~120 CPU-sec/period actually consumed, only ~5 are attributable to JS (harper~4.8s,user/app ~0.2s) — ~96% is native (RocksDB / native addon), invisible to both V8 and the datadog pprof profiler.RocksTransactionLogStore.updateIterators/next,auditStore.Decoder/readAuditEntry,transactionBroadcast.notifyFromTransactionData→ transaction/audit-log iteration + mesh re-broadcast.Resuming interrupted copy of database system … at table hdb_analytics(~3.1M rows); never completes.transaction_logs(per-peer, dominated by replicated analytics audit entries; largest single peer ~684 MB) + ~2.3 GBhdb_deploymentpayload blobs (only ~100 live rows → mostly orphaned superseded payloads).connected:true/Receiving/ver=0direction that is never re-driven.Root cause (assessment)
hdb_analyticsreplicates by default —databases.tsonly adds it toNON_REPLICATING_SYSTEM_TABLESwhenanalytics_replicate === false. On a busy/long-lived cluster the analytics audit stream dominates the per-peertransaction_logs(~1 GB) and thesystembase-copy. The replication path iterates this natively (RocksTransactionLogStore); a worker spins ~100% in native code without making forward progress on the affected copy direction, wedging it until the worker is recycled. No data loss, but a node burns ≥1 core indefinitely and a copy direction never converges.Mitigation (confirmed lever)
Set
analytics_replicate: false— removeshdb_analyticsfrom replication, collapsing the per-peertransaction_logsand the system base-copy to control-plane only. Strongly recommended for large / many-node clusters (cost scales with node count). The product already anticipates this mode —analytics/read.tsdoes cross-node fan-out reads precisely when analytics is not replicated.Suggested fixes / follow-ups
analytics_replicate: falsethe default (or prominently document the cost) — cluster-wide replication of per-node observability data is expensive and rarely needed.systemcopy direction and only recovers on worker recycle (connection-recovery family —connected:true/Receiving/ver=0never re-driven).transaction_logsso a lagging direction can't accumulate ~1 GB of (mostly analytics) history that must be re-scanned.hdb_deploymentpayload blobs vs ~100 live rows suggests superseded deploy payloads aren't being GC'd — bloats system base-copies.Notes
The exact native frame is not named here: the affected hosts have no
perf/gdb/eu-stackandperf_event_paranoid=4, and the spin is episodic. To pin the native function, reproduce in a lab with--perf-basic-prof+perf(or rocksdb-js native instrumentation) while replicating a largehdb_analyticsacross a multi-node cluster.