Skip to content

a customer cluster preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8) #453

Description

@kriszyp

Summary

a customer cluster preprod (4-node Fabric) upgraded to 5.1.7 today (2026-06-22 ~15:46 UTC). The headline 5.1.7 blob-advance fix (#443/harper#1425) works — customerCache converged past 1500+ missing-blob references. But replicated deploys are still wedged: the system-DB base copy never reaches COPY_COMPLETE on any of the 3 follower nodes, so new hdb_deployment rows can't replicate and deploys fail the 120s timeout.

This is a production confirmation of the stalled-blob-send wedge (#450 / harper#1443) — and crucially, neither fix is in 5.1.7, so there is no rescue. Filing as a customer-incident / release-tracking issue to drive a 5.1.8 patch.

Customer impact

Environment

  • harper-pro 5.1.7 (harper core 5.1.7), RocksDB, 4-node Fabric <customer-host> (leader leader, edges wlb/jbd/iid).
  • Trigger: Fabric's automated upgrade did a rolling/staggered restart with only ~20-40s spacing (container StartedAt: iid 15:45:37, jbd 15:45:59, leader plr 15:46:39 — 3rd, mid-sequence, wlb 15:47:16) — far shorter than the time for a node's system base-copy to run uninterrupted to COPY_COMPLETE. (Jeff's earlier manual staggered restart, waiting per-node, worked as a band-aid — the difference is whether the stagger waits for convergence.)

Evidence

Live cluster_status, stable for 2h+:

  • All 3 followers' own view: system connected:true status:Receiving lastReceivedVersion:0 — connected, received nothing, parked in copy mode.
  • Leader view of each peer: system connected:false sendingMessage:Copying with a frozen lastReceivedVersion (byte-identical across many samples).
  • Zero replication log activity cluster-wide after 15:49 — fully silent (the discriminator: a restart-from-zero loop, Clone bulk table-copy restarts from zero after ping-timeout disconnect (no mid-copy checkpoint) #241, would log recurring restarts; a hung for await is silent).
  • The missing system deployment-payload blobs do get classified+advanced on the followers (Blob 110 / 1cc for record 10be229e-… unrecoverable at source … advancing the resume cursor past it (#388)), proving #1425 works — but the copy still never completes, consistent with a different blob whose source stream stalls without throwing (the Blob send loop hangs on a stuck iterator when the local blob file is missing or corrupt #450 failure mode), hanging sendBlobs with no finishing frame.

Root cause (leading, high-confidence)

sendBlobs (replication/replicationConnection.ts:3006) uses a plain for await (const buffer of blob.stream()) with no per-chunk deadline. When a local blob read stalls (ENOENT/corrupt manifesting as a stuck iterator rather than a throw), the loop parks forever, the finishing BLOB_CHUNK is never sent, and the receiver's copy-mode await Promise.all(outstandingBlobsToFinish) (:2873) never resolves → end_txn never commits → outstandingCommits never drains → maybeFinishCopy() (:973) never fires → follower stuck in copy mode → COPY_COMPLETE never reached → live audit replay (which carries new hdb_deployment rows) never starts → deploy timeout.

5.1.7 has neither rescue:

So a single stalled system blob is an unrecoverable, silent, deploy-blocking wedge in 5.1.7.

Secondary contributing factor

Even absent a blob stall, the rolling restart is hostile to convergence: the system resume cursor from the pre-5.1.7 copy carries copyOrder=undefined, so the #421 COPY_ORDER_VERSION guard (:2392) correctly forces a full restart from scratch on every reconnect — and with no mid-copy checkpoint surviving (#241), a copy must run fully uninterrupted to persist copyOrder=1 and converge. The ~30s rolling-restart spacing + reconnect churn prevents that window. Worth considering whether Fabric's automated upgrade restart should wait for per-node replication convergence between steps (central-manager side).

Proposed resolution

  1. Land fix(replication): per-chunk send timeout to unwedge stalled blob streams #451 (sender per-chunk timeout) + harper#1444 (receiver idle watchdog) and ship a 5.1.8 patch — this is the load-bearing fix; either half rescues the wedge.
  2. Consider the convergence-aware staggered upgrade restart (CM) as defense-in-depth against the Clone bulk table-copy restarts from zero after ping-timeout disconnect (no mid-copy checkpoint) #241/system base-copy gated on huge hdb_analytics (analytics.replicate:true) blocks hdb_deployment convergence → deploys fail after any copy #421 restart-from-zero amplifier.
  3. Regression: add to the blob×replication suite (Blob × replication regression tests (Tier 1: deterministic cluster integration) #411) a rolling-restart-during-system-copy case with a missing/stalling source blob, asserting deploys converge.

Cross-refs

Root cause / fixes: #450, PR #451, harper#1443, harper#1444. Family: #420 (closed), #424 (merged, insufficient here), #289, #241, #426, #388.

cc @kriszyp

Filed by Claude (Opus 4.8) from live investigation of the a customer cluster preprod cluster.

Activity

  1. kriszyp commented on Jun 22, 2026

    @kriszyp
    MemberAuthor

    Correction / refined analysis after tracing the existing blob timeouts.

    I initially attributed this to a stalled blob send (#450). On closer reading that's almost certainly not the a customer cluster mechanism, because the local-blob read path is already bounded:

    • Sender: sendBlobs → blob.stream() is bounded by getBlobReadTimeout (storage_blobReadTimeout, 20s) at every failure mode — missing file → immediate 404 or ≤20s (blob.ts:351), stalled mid-read → ≤20s (blob.ts:514), corrupt → immediate 500. It cannot hang for hours.
    • Receiver: blobsTimer sweeps blobsInFlight every REPLICATION_BLOBTIMEOUT (120s) and destroys stalled incoming streams (replicationConnection.ts:3601).

    And live evidence agrees: the missing system blobs (110/1cc) were classified and advanced within seconds (#1425/#443 working). A blob stall would clear in ≤120s, yet a customer cluster has been wedged for hours.

    So the real wedge is the system base-copy not being re-driven to COPY_COMPLETE after the rolling upgrade restart (followers Receiving/ver=0; leader not pushing the copy) — a copy/reconnect orchestration issue in the #420/#289/#241 family, not a blob hang. The blob-timeout enablement (PRs #451/#1444) is still valid hardening (those watchdogs were shipping disabled), but it likely does not fix this wedge.

    Next: reproduce against a copy of a customer cluster's actual system DB + cursor state and pin/prove the convergence fix before 5.1.8, plus an integration test that reproduces rolling-restart-during-system-copy. Keeping this issue open as the tracking item; will repoint the root-cause once the repro confirms it.

    — Claude (Opus 4.8)

  2. kriszyp commented on Jun 22, 2026

    @kriszyp
    MemberAuthor

    Fix PR up (draft): #454 — the copy-progress watchdog that recovers the connected:true copy-stall wedge. Proven with a new integration test (fails without the watchdog, passes with it). — Claude (Opus 4.8)

  3. changed the title [-]JJill preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8)[/-] [+]a customer cluster preprod 5.1.7: replicated deploys wedge after rolling upgrade — stalled system blob send (#450/harper#1443) blocks COPY_COMPLETE; fixes absent from 5.1.7 (needs 5.1.8)[/+] on Jun 22, 2026
  4. kriszyp commented on Jul 7, 2026

    @kriszyp
    MemberAuthor

    Fixed by #454 (copy-progress watchdog to recover ping-alive copy-stall wedges), merged 2026-06-25.

    — Claude (Opus 4.8), issue-backlog triage on Kris's behalf

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions