Skip to content

Interrupted bulk copies can persist a resume cursor over an undelivered range #537

Description

@ldt1996

Bug Summary

An interrupted bulk copy can persist a resume cursor over a range that was never delivered; every later resume then skips to the cursor's tail, completes, and permanently seals the hole.

Steps to Reproduce

  1. Run bulk copies of a large blob table while the copy legs are being interrupted (restart churn, watchdog cycles).
  2. Let the interrupted copies persist their [copyCursor, nodeId] rows.
  3. Reconnect: the copy resumes at afterKey, delivers the tail, sends COPY_COMPLETE, and the cursor clears.

Expected Behavior

A resumed copy only trusts the cursor if the range before afterKey was actually delivered, or the shortfall is detected and the copy restarts from scratch.

Actual Behavior

The copy-order guard validates cursor compatibility but never delivery. Observed on a 4-node staging cluster: laggards plateaued thousands of rows below the source with no active copy running and their subscriptions believing themselves current; forced fresh copies were the only way to converge. Multi-source copy interruption during restart churn makes this reachable in practice.

Platform

Harper 5.1.x, Node 24, 4-node cluster

Lavinia, via Claude

Activity

  1. self-assigned this
    on Jul 7, 2026
  2. ldt1996 commented on Jul 7, 2026

    @ldt1996
    ContributorAuthor

    For reproduction: the direct simulation is the easiest way to see it. Write a [Symbol.for('copyCursor'), nodeId] row into a database's __dbis__ store with an afterKey deep in a big table's keyspace, then reconnect the subscription. The copy resumes at that key, delivers only the tail, sends COPY_COMPLETE, and the subscription considers itself current with everything before the afterKey never delivered.

    Organically we hit the symptom on a 4-node stage cluster during v4 to v5 rolling-upgrade testing: a ~30k row blob table (~50KB blobs), full copies running from multiple sources concurrently, and repeated mid-copy interruptions (restart churn plus the 300s copy watchdog cycling legs stalled on the apply bug fixed in HarperFast/harper#1641). After enough interrupt-resume cycles the laggards plateaued thousands of rows short while every subscription believed itself current; only forced fresh copies converged them. The decode-drop mechanism laid out on the PR fits the field cases better than anything we managed to pin at the time, including one where rows landed on two peers but never on a third.

    Lavinia, via Claude

  3. kriszyp commented on Jul 7, 2026

    @kriszyp
    Member

    Thanks @ldt1996 — the Part-1 direct simulation is exactly the handle we needed: it makes the symptom (a resume cursor ahead of delivered data → tail-only copy → sealed hole) reproducible on demand, independent of whatever poisoned the cursor. Really useful.

    Want to be straight about what's validated so far vs. what's still owed, since I don't want to overclaim a reproduction:

    Validated (component-level): the replication receive decoder is a raw StructonPackr, so an absent shared structure throws ("Could not find typed structure N") rather than decoding to null — and the fix on #545 classifies that as permanent, skips it, and fires the same decode-missing-structure metric core's local-read path uses (so it's alertable instead of silent). I also confirmed the sender always emits TABLE_FIXED_STRUCTURE before the first record of a table over the ordered socket — so a throw means the structure is genuinely absent from the sender's synced set, i.e. old-version/corrupt bytes no re-copy heals. That fits your v4→v5 upgrade context and the "landed on two peers but never a third" case (per-node structure divergence) well.

    Not yet done (and I'd previously overstated this): an actual end-to-end reproduction — a real interrupted copy producing the hole and rows going permanently missing. So far I've only exercised the mechanism's pieces and the happy path. We're adding two tests: (1) your Part-1 poisoned-cursor sim as a characterization test, and (2) a regression test that induces a real decode-drop mid-copy and asserts the metric fires + the row is absent (silent before the fix).

    One framing note: the fix makes the permanent drop observable, it doesn't prevent the loss — the records are genuinely unrecoverable, so the operational remedy stays "metric fires → re-clone," which matches your "only forced fresh copies converged." Where the fix does prevent loss is the adjacent unknown-table-id path (a missed structure sync / un-propagated schema), which #545 now holds on and resyncs instead of dropping.

    The question that would nail it: do the field logs from the stage cluster show decode errors on the laggards — anything like "Could not find typed structure" or "Error decoding replication message" around when they plateaued? That would confirm this is the decode-drop path (mode A) rather than the interrupt + apply-bug (harper#1641) poisoning the cursor by some route that never touches decode. If it's the latter, that's a separate fix from #545 and worth its own issue.

    — KrAIs (Claude Opus 4.8), with Kris

  4. added this to the v5.2 milestone on Aug 7, 2026
  5. added theissue type on Aug 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

Fields

Priority

P0

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions