Motivation
A cluster of production incidents on the JJill preprod Fabric cluster (and others) all share one shape that our test suite does not cover: blobs × TTL/eviction × replication/base-copy × failure-and-recovery. Each was found in the field, not CI:
Today integrationTests/cluster has blob fault-injection fixtures (fixture-blob-fail-injector, fixture-blob-fail-transient, fixture-large-blob-authoritative) driving blobSaveRejectionContainment.test.mjs. That proves containment (a failed inbound save doesn't crash the receiver). It never asserts convergence or no-orphan. This issue closes that gap with deterministic, per-PR regression tests; harper-pro#412 covers the randomized chaos/stress tier.
The invariant oracle (shared by both tiers)
After an op sequence + fault injection, quiesce, then on every node assert:
- No orphaned references — for every record carrying a blob, the blob is readable and its content-hash equals the source's.
- Convergence / liveness — all nodes reach an identical
(primaryKey → blobHash) set within a bound; cluster_status shows blobReplicationFailures stabilized and the resume cursor advancing (never permanently pinned).
- Bounded orphan files — leaked blob files (the safe direction of #1364: a skipped unlink) are reclaimed by
cleanupOrphans; blob-store file count does not grow unbounded across cycles.
- Isolation —
system/deploy replication stays live while a data-DB blob copy is backpressured/wedged.
A small shared assertion helper implementing (1)–(4) should be the first deliverable; every scenario below ends by calling it.
Scope — deterministic cluster integration tests (one per failure mode)
Small, seeded, fixed fault points (in the style of the existing .test.mjs). Extend the injector as needed beyond receive-side write-fail to also: skip the blob transfer entirely (commit metadata, never send the blob), delete a blob file out from under a live record, and drop blob chunks mid-stream.
Where it lives / reuse
harper-pro/integrationTests/cluster/ — multi-node harness + the existing blob fixtures.
- Extend
fixture-blob-fail-injector with the new injection modes (env-toggled, same pattern).
- Engines: parameterize across RocksDB and LMDB where feasible (the orphan paths differ subtly by engine).
Acceptance criteria
Related: HarperFast/harper#1364, HarperFast/harper#1353, HarperFast/harper#1369, harper-pro#409, harper-pro#403. Chaos/stress tier: harper-pro#412.
Motivation
A cluster of production incidents on the JJill preprod Fabric cluster (and others) all share one shape that our test suite does not cover: blobs × TTL/eviction × replication/base-copy × failure-and-recovery. Each was found in the field, not CI:
Receivingwhen sender has records referencing blob files missing on disk harper#1337 — an orphaned receive-side blob stream crashing the process.Today
integrationTests/clusterhas blob fault-injection fixtures (fixture-blob-fail-injector,fixture-blob-fail-transient,fixture-large-blob-authoritative) drivingblobSaveRejectionContainment.test.mjs. That proves containment (a failed inbound save doesn't crash the receiver). It never asserts convergence or no-orphan. This issue closes that gap with deterministic, per-PR regression tests; harper-pro#412 covers the randomized chaos/stress tier.The invariant oracle (shared by both tiers)
After an op sequence + fault injection, quiesce, then on every node assert:
(primaryKey → blobHash)set within a bound;cluster_statusshowsblobReplicationFailuresstabilized and the resume cursor advancing (never permanently pinned).cleanupOrphans; blob-store file count does not grow unbounded across cycles.system/deploy replication stays live while a data-DB blob copy is backpressured/wedged.A small shared assertion helper implementing (1)–(4) should be the first deliverable; every scenario below ends by calling it.
Scope — deterministic cluster integration tests (one per failure mode)
Small, seeded, fixed fault points (in the style of the existing
.test.mjs). Extend the injector as needed beyond receive-side write-fail to also: skip the blob transfer entirely (commit metadata, never send the blob), delete a blob file out from under a live record, and drop blob chunks mid-stream.storage.debugLongTransactions: trueto observe) → assert blob files are not unlinked under records that still exist.deploy_component→ assert it still replicates to peers (system-table replication isolated from data-DB blob backpressure).Where it lives / reuse
harper-pro/integrationTests/cluster/— multi-node harness + the existing blob fixtures.fixture-blob-fail-injectorwith the new injection modes (env-toggled, same pattern).Acceptance criteria
main(verify the regression value — e.g. base-copy-past-missing-blob should hang/fail without fix(blob): tolerate source-unavailable blobs in pre-commit so a missing blob can't wedge replication harper#1353).Related: HarperFast/harper#1364, HarperFast/harper#1353, HarperFast/harper#1369, harper-pro#409, harper-pro#403. Chaos/stress tier: harper-pro#412.