Skip to content

Clone sync monitor is vacuous on RocksDB: clone marks itself Available mid-copy (targets all 0) #655

Description

@kriszyp

Summary

clone_node's sync monitor is vacuous on RocksDB: it declares "All databases synchronized" on its first poll — seconds into a multi-GB base copy — and the clone marks itself availability: Available and cloned with an arbitrarily small fraction of the leader's data. This is long-standing (not a regression from the #649 monitor rework); it reproduces deterministically locally.

Evidence

CI: Large-Data stress run 31006315068 failed Clone row count 31393 != leader 104858; the clone log shows "All databases synchronized" 2.2s after requesting a full copy of a 10 GB database. Green runs (30982928212, 30792466356) show the identical instant-synchronized — they passed only because the copy happened to finish inside the test's row-count polling window. Aug 3's run declared synced 59s into a still-running copy.

Local repro (1 GB, HARPER_RUN_STRESS_TESTS=1 HARPER_STRESS_LARGE_DATA_GB=1 … largeClone.test.mjs) fails identically, and with instrumentation shows the exact decision:

[cloneNode]: [KR-WM] targets={"system":0,"data":0} sockets=[{"db":"system","v":1785939110547},{"db":"data","v":0}]
[cloneNode]: [KR-WM] Database system: No target timestamp, skipping sync check
[cloneNode]: [KR-WM] Database data: No target timestamp, skipping sync check
[cloneNode]: All databases synchronized        <-- first poll, data copy just started

Root cause chain

  1. Describe reports no last_updated_record on RocksDB. cloneNode.getLastUpdatedRecord() derives each database's sync target from describe_database/describe_all → last_updated_record. Core's schemaDescribe.ts computes that from auditStore.getKeys({ reverse: true, limit: 1 }) — but RocksTransactionLogStore.getKeys() is an unimplemented stub (return []; // TODO: implement this), and the indices.__updatedtime__ fallback doesn't exist on these tables. So every database's target is 0.
  2. The monitor treats a falsy target as "skip this database". checkSyncStatus (cloneNode/syncMonitor.ts) does if (!targetTime) continue; — with every database skipped, syncComplete stays true and the first poll succeeds with zero verification.

The receive-side copy machinery is doing its part correctly: the per-database received-version watermark (RECEIVED_VERSION_POSITION) is deliberately held at 0 for the whole bulk copy and advanced to copyStartTime only by the single end_txn the sender always emits after the copy — the code calls this "the sole signal that the copy is synced". The monitor just never looks at it when targets are missing.

Impact

Fix (this issue → harper-pro)

In checkSyncStatus, a database without a target timestamp must not be skipped: require its received-version watermark to be positive (receivedVersion >= (targetTime || 1)). The watermark can only become positive via the sender's final copy end_txn (or later live traffic), so this keys completion to the copy's own completion signal — correct against any leader version, including leaders whose describe cannot report last_updated_record, and for empty databases (their copy still emits the final end_txn). Additionally, a database present in the targets but missing from database_sockets must hold syncComplete false, so a not-yet-registered socket can't produce a vacuous pass either.

Follow-up (separate issue → core)

last_updated_record is silently absent from describe_table/describe_all on RocksDB — a user-visible describe regression vs LMDB — because RocksTransactionLogStore.getKeys() is a TODO stub and the transaction log has no tail-read API to implement it cheaply (rocksdb-js TransactionLog exposes only forward query(); last-committed position exists but carries no timestamp). Needs a small rocksdb-js tail API (e.g. getLastEntry()), then core can restore the field.

Activity

  1. self-assigned this
    on Aug 5, 2026
  2. added this to the v5.2 milestone on Aug 5, 2026
  3. added theissue type on Aug 5, 2026
  4. kriszyp commented on Aug 6, 2026

    @kriszyp
    MemberAuthor

    Carrying over the field evidence from #611, which is being closed as a duplicate of this issue. It observed the same defect earlier and on a different engine/version, and it records one consequence this issue doesn't:

    • harper-pro 5.1.14, 2-node docker rig, no Fabric — so the failure is not RocksDB-specific and not new in 5.1.2x.
    • The follower logged All databases synchronized → Clone from leader node … complete, set cloned: true, and published availability: Available while holding 56,071 of 500,000 records (leader count verified at 500,000).
    • Replication then died (the credential's user had been removed on the leader as part of that test), and the node stayed Available with ~11% of the data indefinitely. get_status showed every component healthy and nothing retried.

    That last point is the part worth keeping: this issue frames the window as bounded — minutes, until the copy finishes — and that's what keeps it at P1 rather than P0. #611 shows the window is only bounded if the copy actually completes. If replication dies mid-copy, a node that has already advertised itself Available stays that way, serving an arbitrary fraction of the data, with no self-heal and no signal.

    So the fix probably needs to cover both halves: don't declare synchronized against absent targets (this issue, PR #657), and don't leave Available latched when the copy that justified it never finished.

    — KrAIs (Claude Opus 5)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

Fields

Priority

P1

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions