Skip to content

Replicated hdb_analytics floods system transaction logs and spins a worker in native code; system base-copy wedges until worker recycle #480

Description

@kriszyp

Summary

On a 4-node v5.1.12 cluster, the system database base-copy between some node pairs never converges. A replication worker pins ~100% CPU in native code (invisible to JS/V8 profilers) while the affected system socket makes zero forward progress (connected: true, status: Receiving, lastReceivedVersion: 0, lastReceivedRemoteTime frozen). It recovers only when the http worker pool is recycled — i.e. it clears episodically and then re-wedges. The driver is replicating hdb_analytics (on by default), which floods the per-peer transaction_logs and the system base-copy.

No data loss: the affected node stays current on system via its other peers; it's a redundant copy direction that wedges while burning a core.

Environment

  • harper-pro + harper v5.1.12 (current latest), 4-node mesh, RocksDB storage
  • analytics_replicate unset → default (analytics replicates)

Evidence

  • Stuck socket: the affected system direction shows connected: true, status: Receiving, lastReceivedVersion: 0, lastReceivedRemoteTime frozen ~23h (pre-restart), backPressurePercent: 0. The other system sockets on that node are current → the node's data is converged; only the redundant direction is wedged.
  • CPU is native: top -H shows one http worker (intermittently also the main thread) pegged ~100% CPU, sustained >1h. Yet V8 CPU profiles of all worker threads show ~100% idle. The built-in analytics self-profiler (analytics/profile.ts) corroborates: of the ~120 CPU-sec/period actually consumed, only ~5 are attributable to JS (harper ~4.8s, user/app ~0.2s) — ~96% is native (RocksDB / native addon), invisible to both V8 and the datadog pprof profiler.
  • Hot JS entry points (the small non-idle slice on main): RocksTransactionLogStore.updateIterators/next, auditStore.Decoder/readAuditEntry, transactionBroadcast.notifyFromTransactionData → transaction/audit-log iteration + mesh re-broadcast.
  • Copy parked on analytics: logs repeatedly Resuming interrupted copy of database system … at table hdb_analytics (~3.1M rows); never completes.
  • On-disk system DB: ~1.8 GB data + ~990 MB transaction_logs (per-peer, dominated by replicated analytics audit entries; largest single peer ~684 MB) + ~2.3 GB hdb_deployment payload blobs (only ~100 live rows → mostly orphaned superseded payloads).
  • Disproportionate to real writes: analytics flush is periodic and user/app CPU is ~0.2s — the work grinds accumulated log history, not live load.
  • Recovery: when the http worker pool is recycled (worker thread ids change; container not restarted), the affected socket reconnects on a fresh worker and goes current. Re-wedges over time. Matches the connection-recovery family (Replication wedges permanently after simultaneous cluster restart (reconciler skips open-but-idle sockets) → blocks replicated deploys #420 / fix(replication): recover open-but-idle wedged subscriptions via watchdog-driven reconnect #424 / Replication: outbound subscription to a restarted peer wedges connected:false with no reconnect attempt (no SYN), reconcile never re-drives — base-copy + post-restart TLS window #466) — a connected:true / Receiving / ver=0 direction that is never re-driven.

Root cause (assessment)

hdb_analytics replicates by default — databases.ts only adds it to NON_REPLICATING_SYSTEM_TABLES when analytics_replicate === false. On a busy/long-lived cluster the analytics audit stream dominates the per-peer transaction_logs (~1 GB) and the system base-copy. The replication path iterates this natively (RocksTransactionLogStore); a worker spins ~100% in native code without making forward progress on the affected copy direction, wedging it until the worker is recycled. No data loss, but a node burns ≥1 core indefinitely and a copy direction never converges.

Mitigation (confirmed lever)

Set analytics_replicate: false — removes hdb_analytics from replication, collapsing the per-peer transaction_logs and the system base-copy to control-plane only. Strongly recommended for large / many-node clusters (cost scales with node count). The product already anticipates this mode — analytics/read.ts does cross-node fan-out reads precisely when analytics is not replicated.

Suggested fixes / follow-ups

  1. Consider making analytics_replicate: false the default (or prominently document the cost) — cluster-wide replication of per-node observability data is expensive and rarely needed.
  2. Investigate why native txn-log iteration spins without forward progress on a redundant system copy direction and only recovers on worker recycle (connection-recovery family — connected:true/Receiving/ver=0 never re-driven).
  3. Bound/age the per-peer transaction_logs so a lagging direction can't accumulate ~1 GB of (mostly analytics) history that must be re-scanned.
  4. Separately: ~2.3 GB of hdb_deployment payload blobs vs ~100 live rows suggests superseded deploy payloads aren't being GC'd — bloats system base-copies.

Notes

The exact native frame is not named here: the affected hosts have no perf/gdb/eu-stack and perf_event_paranoid=4, and the spin is episodic. To pin the native function, reproduce in a lab with --perf-basic-prof + perf (or rocksdb-js native instrumentation) while replicating a large hdb_analytics across a multi-node cluster.

Activity

  1. kriszyp commented on Jun 24, 2026

    @kriszyp
    MemberAuthor

    Why it never converges: the base-copy cursor never finalizes → forces full re-copy on restart (self-perpetuating)

    Traced the copy-resume cursor mechanism in dist/replication/replicationConnection.js:

    • The resume cursor is persisted at dbisDB[Symbol.for('copyCursor'), nodeId] as {copyStartTime, currentTable, afterKey, copyOrder}, advanced per copied record (flushDurableCopyCursor), and cleared only on COPY_COMPLETE (maybeFinishCopy). The code comment is explicit: if the cursor loops, "the cluster never converges, and COPY_COMPLETE never arrives to clear the cursor."
    • Because the system copy can't get past hdb_analytics, COPY_COMPLETE never arrives → the cursor never finalizes → the socket stays connected:true / Receiving / lastReceivedVersion:0 indefinitely (the seemingly-"cosmetic" ver=0 socket).
    • On reconnect/restart, an initial copy that never completed deliberately triggers a full re-copy (not an incremental resume) — see the #426 data-loss safeguard at ~L3120-3196: "when an initial copy never completed … the node then resumes from now-60s and permanently loses the gap … A full copy is idempotent … the cost of an occasional extra copy is acceptable versus silent data loss" → startTime = 0 (full copy).

    So the failure is self-perpetuating: replicating hdb_analytics makes the system base-copy too big to finish → cursor never finalizes → any restart/reconnect re-requests a full copy of the whole system DB (incl. the 3.1M-row analytics table) → native txn-log scan spins the worker again → still never completes → cursor still never finalizes. Each restart re-pays the full analytics-copy cost.

    This is why analytics_replicate: false is the clean fix: with analytics out of replication, the system base-copy is tiny (control-plane only), COPY_COMPLETE arrives, the cursor finalizes/clears, and subsequent reconnects resume incrementally instead of full-recopying.

    The cursor is inspectable live: databases.system.<anyTable>.dbisDB.getRange() filtered to keys whose first element is Symbol.for('copyCursor') returns {node, cursor:{currentTable, afterKey, copyOrder, copyStartTime}} per in-flight source. (Present only while a copy is mid-flight; cleared on completion.)

  2. kriszyp commented on Jul 7, 2026

    @kriszyp
    MemberAuthor

    Fixed by #486 (merged 2026-06-25) — copyApply snapshots base-copy rows without audit entries (RocksDB-only, durability-gated).

    — Claude (Opus 4.8), issue-backlog triage on Kris's behalf

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions