Repository navigation
Interrupted bulk copies can persist a resume cursor over an undelivered range #537
Description
Activity
For reproduction: the direct simulation is the easiest way to see it. Write a
[Symbol.for('copyCursor'), nodeId]row into a database's__dbis__store with an afterKey deep in a big table's keyspace, then reconnect the subscription. The copy resumes at that key, delivers only the tail, sends COPY_COMPLETE, and the subscription considers itself current with everything before the afterKey never delivered.Organically we hit the symptom on a 4-node stage cluster during v4 to v5 rolling-upgrade testing: a ~30k row blob table (~50KB blobs), full copies running from multiple sources concurrently, and repeated mid-copy interruptions (restart churn plus the 300s copy watchdog cycling legs stalled on the apply bug fixed in HarperFast/harper#1641). After enough interrupt-resume cycles the laggards plateaued thousands of rows short while every subscription believed itself current; only forced fresh copies converged them. The decode-drop mechanism laid out on the PR fits the field cases better than anything we managed to pin at the time, including one where rows landed on two peers but never on a third.
Lavinia, via Claude
- added a commit that references this issue
on Jul 7, 2026 Thanks @ldt1996 — the Part-1 direct simulation is exactly the handle we needed: it makes the symptom (a resume cursor ahead of delivered data → tail-only copy → sealed hole) reproducible on demand, independent of whatever poisoned the cursor. Really useful.
Want to be straight about what's validated so far vs. what's still owed, since I don't want to overclaim a reproduction:
Validated (component-level): the replication receive decoder is a raw
StructonPackr, so an absent shared structure throws ("Could not find typed structure N") rather than decoding to null — and the fix on #545 classifies that as permanent, skips it, and fires the samedecode-missing-structuremetric core's local-read path uses (so it's alertable instead of silent). I also confirmed the sender always emitsTABLE_FIXED_STRUCTUREbefore the first record of a table over the ordered socket — so a throw means the structure is genuinely absent from the sender's synced set, i.e. old-version/corrupt bytes no re-copy heals. That fits your v4→v5 upgrade context and the "landed on two peers but never a third" case (per-node structure divergence) well.Not yet done (and I'd previously overstated this): an actual end-to-end reproduction — a real interrupted copy producing the hole and rows going permanently missing. So far I've only exercised the mechanism's pieces and the happy path. We're adding two tests: (1) your Part-1 poisoned-cursor sim as a characterization test, and (2) a regression test that induces a real decode-drop mid-copy and asserts the metric fires + the row is absent (silent before the fix).
One framing note: the fix makes the permanent drop observable, it doesn't prevent the loss — the records are genuinely unrecoverable, so the operational remedy stays "metric fires → re-clone," which matches your "only forced fresh copies converged." Where the fix does prevent loss is the adjacent unknown-
table-idpath (a missed structure sync / un-propagated schema), which #545 now holds on and resyncs instead of dropping.The question that would nail it: do the field logs from the stage cluster show decode errors on the laggards — anything like
"Could not find typed structure"or"Error decoding replication message"around when they plateaued? That would confirm this is the decode-drop path (mode A) rather than the interrupt + apply-bug (harper#1641) poisoning the cursor by some route that never touches decode. If it's the latter, that's a separate fix from #545 and worth its own issue.— KrAIs (Claude Opus 4.8), with Kris
- added a commit that references this issue
on Jul 7, 2026 - added a commit that references this issue
on Jul 30, 2026 - added 9 commits that reference this issue
on Aug 12, 2026 - added a commit that references this issue
on Aug 15, 2026 - added a commit that references this issue
on Aug 17, 2026 - added 5 commits that reference this issue
on Aug 19, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Bug Summary
An interrupted bulk copy can persist a resume cursor over a range that was never delivered; every later resume then skips to the cursor's tail, completes, and permanently seals the hole.
Steps to Reproduce
[copyCursor, nodeId]rows.afterKey, delivers the tail, sends COPY_COMPLETE, and the cursor clears.Expected Behavior
A resumed copy only trusts the cursor if the range before
afterKeywas actually delivered, or the shortfall is detected and the copy restarts from scratch.Actual Behavior
The copy-order guard validates cursor compatibility but never delivery. Observed on a 4-node staging cluster: laggards plateaued thousands of rows below the source with no active copy running and their subscriptions believing themselves current; forced fresh copies were the only way to converge. Multi-source copy interruption during restart churn makes this reachable in practice.
Platform
Harper 5.1.x, Node 24, 4-node cluster
Lavinia, via Claude