Repository navigation
Prevent copy recovery stalls after blob faults - #732
Conversation
|
Reviewed; no blockers found. |
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
06901e2 to
d707f8a
Compare
|
placeholder Generated by Claude Code |
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
…r-stall # Conflicts: # core
Correction: the regression test does pin this fixI reported earlier that the control passed and therefore the test did not discriminate. That was drawn from two samples of an intermittent race and was wrong. A third control run on Node 22 reproduced the original failure exactly. Control = this PR with the ownership guard's condition hardcoded
Run 3's fingerprint against the original nightly red (32226893983): Same assertion, same shape, same magnitude — 362.7s here against the original 363s, and against 61s on Node 24 and 70s on Node 26.5 in the same control run. The cursor pins at 13 and re-walks it thirteen times; So on Node 22 the base reproduces at roughly 1 in 3, matching this PR's own account ("two exact-Node-22 pre-fix stalls... a third run escaped"). The test is a genuine regression anchor; it is just probabilistic, and two green runs are not enough to conclude anything from it. Caveat worth keeping on the record: with base failing ~1/3, three green runs on the fix side (1 CI × 3 runtimes, plus 3 local on Node 22) leave roughly a 30% chance of coincidence. The fix rests on the invariant argument — Closing #741 now that this is recorded. — Claude Opus 5 |
A queued replication record handler could begin a blob repair after connection teardown had already swept
blobsInFlight. That orphaned receive stream had no sender or connection timer left to finish it, so it retained the blob lock and every reconnect declined repair until core's long source-idle timeout. Runtime scheduling changed how often the race appeared; this is not intrinsically a Node 22 defect.This change aborts a newly attached blob receive when its closure no longer owns the websocket. The companion core change makes failed-save rejection a cleanup barrier, ensuring PENDING-marker completion and blob unlock cannot lag promise settlement.
The regression now runs pre-merge on Node 22, 24, and 26.5.0. It still requires faults across two advancing copy cursors, but injects them every five blob saves so the resumed repair pass reliably encounters another fault. Previously, two faults could both occur during the initial pass, leaving the strong two-resume assertion waiting despite complete data recovery.
Depends-on: HarperFast/harper#2228
Companion: HarperFast/harper#2228 — Release blob locks before failed saves settle.
Fixes #699
Issue: HarperFast/harper-pro#699 — 5.2.2 blob-gap wedge is bounded but banks zero copy progress — the copy cursor never persists during a copy body, and the watchdog's reconnect mints the next cycle's faults.
For the human reviewer
Verification
[3, 16], 40/40 records, and no missing payloads.ordered-binarydecoded a key as a fractional number and threw while converting it toBigInt, repeatedly closing the audit subscription. The production and regression changes in this PR do not run in that process. A targeted rerun passed.npm run lint:required, Prettier, and workflow YAML parsing passed.dist.Co-Authored-By: GPT-5 Codex noreply@openai.com
Review-Coverage: authored=unknown; ran=none; rounds=1 @ 5e171f2
Human-Review-Need: 4 @ 5e171f2