Repository navigation
a customer cluster preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8) #453
Description
Activity
Correction / refined analysis after tracing the existing blob timeouts.
I initially attributed this to a stalled blob send (#450). On closer reading that's almost certainly not the a customer cluster mechanism, because the local-blob read path is already bounded:
- Sender:
sendBlobs→blob.stream()is bounded bygetBlobReadTimeout(storage_blobReadTimeout, 20s) at every failure mode — missing file → immediate 404 or ≤20s (blob.ts:351), stalled mid-read → ≤20s (blob.ts:514), corrupt → immediate 500. It cannot hang for hours. - Receiver:
blobsTimersweepsblobsInFlighteveryREPLICATION_BLOBTIMEOUT(120s) and destroys stalled incoming streams (replicationConnection.ts:3601).
And live evidence agrees: the missing
systemblobs (110/1cc) were classified and advanced within seconds (#1425/#443 working). A blob stall would clear in ≤120s, yet a customer cluster has been wedged for hours.So the real wedge is the
systembase-copy not being re-driven toCOPY_COMPLETEafter the rolling upgrade restart (followersReceiving/ver=0; leader not pushing the copy) — a copy/reconnect orchestration issue in the #420/#289/#241 family, not a blob hang. The blob-timeout enablement (PRs #451/#1444) is still valid hardening (those watchdogs were shipping disabled), but it likely does not fix this wedge.Next: reproduce against a copy of a customer cluster's actual
systemDB + cursor state and pin/prove the convergence fix before 5.1.8, plus an integration test that reproduces rolling-restart-during-system-copy. Keeping this issue open as the tracking item; will repoint the root-cause once the repro confirms it.— Claude (Opus 4.8)
- Sender:
Fix PR up (draft): #454 — the copy-progress watchdog that recovers the connected:true copy-stall wedge. Proven with a new integration test (fails without the watchdog, passes with it). — Claude (Opus 4.8)
- changed the title
[-]JJill preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8)[/-][+]a customer cluster preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8)[/+]on Jun 22, 2026 - added 11 commits that reference this issue
on Jun 22, 2026 Fixed by #454 (copy-progress watchdog to recover ping-alive copy-stall wedges), merged 2026-06-25.
— Claude (Opus 4.8), issue-backlog triage on Kris's behalf
- added a commit that references this issue
on Sep 8, 2026 - added a commit that references this issue
on Sep 8, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
a customer cluster preprod (4-node Fabric) upgraded to 5.1.7 today (2026-06-22 ~15:46 UTC). The headline 5.1.7 blob-advance fix (#443/harper#1425) works —
customerCacheconverged past 1500+ missing-blob references. But replicated deploys are still wedged: thesystem-DB base copy never reachesCOPY_COMPLETEon any of the 3 follower nodes, so newhdb_deploymentrows can't replicate and deploys fail the 120s timeout.This is a production confirmation of the stalled-blob-send wedge (#450 / harper#1443) — and crucially, neither fix is in 5.1.7, so there is no rescue. Filing as a customer-incident / release-tracking issue to drive a 5.1.8 patch.
Customer impact
hdb_deployment row did not replicate), but a different underlying cause.Environment
<customer-host>(leaderleader, edgeswlb/jbd/iid).systembase-copy to run uninterrupted toCOPY_COMPLETE. (Jeff's earlier manual staggered restart, waiting per-node, worked as a band-aid — the difference is whether the stagger waits for convergence.)Evidence
Live
cluster_status, stable for 2h+:system connected:true status:Receiving lastReceivedVersion:0— connected, received nothing, parked in copy mode.system connected:false sendingMessage:Copyingwith a frozenlastReceivedVersion(byte-identical across many samples).for awaitis silent).systemdeployment-payload blobs do get classified+advanced on the followers (Blob 110 / 1cc for record 10be229e-… unrecoverable at source … advancing the resume cursor past it (#388)), proving #1425 works — but the copy still never completes, consistent with a different blob whose source stream stalls without throwing (the Blob send loop hangs on a stuck iterator when the local blob file is missing or corrupt #450 failure mode), hangingsendBlobswith no finishing frame.Root cause (leading, high-confidence)
sendBlobs(replication/replicationConnection.ts:3006) uses a plainfor await (const buffer of blob.stream())with no per-chunk deadline. When a local blob read stalls (ENOENT/corrupt manifesting as a stuck iterator rather than a throw), the loop parks forever, the finishingBLOB_CHUNKis never sent, and the receiver's copy-modeawait Promise.all(outstandingBlobsToFinish)(:2873) never resolves → end_txn never commits →outstandingCommitsnever drains →maybeFinishCopy()(:973) never fires → follower stuck in copy mode →COPY_COMPLETEnever reached → live audit replay (which carries newhdb_deploymentrows) never starts → deploy timeout.5.1.7 has neither rescue:
resources/blob.ts).So a single stalled system blob is an unrecoverable, silent, deploy-blocking wedge in 5.1.7.
Secondary contributing factor
Even absent a blob stall, the rolling restart is hostile to convergence: the
systemresume cursor from the pre-5.1.7 copy carriescopyOrder=undefined, so the #421COPY_ORDER_VERSIONguard (:2392) correctly forces a full restart from scratch on every reconnect — and with no mid-copy checkpoint surviving (#241), a copy must run fully uninterrupted to persistcopyOrder=1and converge. The ~30s rolling-restart spacing + reconnect churn prevents that window. Worth considering whether Fabric's automated upgrade restart should wait for per-node replication convergence between steps (central-manager side).Proposed resolution
Cross-refs
Root cause / fixes: #450, PR #451, harper#1443, harper#1444. Family: #420 (closed), #424 (merged, insufficient here), #289, #241, #426, #388.
cc @kriszyp
Filed by Claude (Opus 4.8) from live investigation of the a customer cluster preprod cluster.