Repository navigation
Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB) #386
Description
Activity
- addedbugSomething isn't workingSomething isn't workingarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connections
on Jun 15, 2026 On-disk timeline pins the loss event to the beta.2 upgrade (Jun 11)
Blob-file mtimes confirm the acute loss event was the 5.1.0-beta.2 upgrade on 2026-06-11 ~20:44 UTC, not today's 5.1.0 release upgrade and not a gradual erosion:
central — data blob mtimes by day: 15 (May 8), 103 (May 12), 290 (May 14), 345 (Jun 11). Of the Jun 11 batch, 264 were written as 8-byte stubs at exactly 20:44 (the beta.2 completion minute). 9 earlier stubs date to May 12 02:26.
east4 — surviving blobs are all dated May 8 / May 12 (none from Jun 11+); the node received nothing at beta.2, and its
blobs/dir was last modified Jun 11 21:32 (the truncation window). The multi-MB documents lived only on east4 and were removed here.So the chronic condition (blob replication failing, east4 not receiving since ~May 12) predates the upgrade, but the document bytes were lost at the beta.2 upgrade — central's good blobs were rewritten to empty stubs and east4's store was truncated in the same window, leaving no complete copy anywhere. Correcting the "did not originate the loss" wording above: the replication degradation predates beta.2; the byte loss coincides with it.
Root cause identified: HarperFast/harper#1302 (RecordEncoder invalidate path unlinks live-referenced blobs → progressive silent loss). The timeline/observations in this issue are correct, but the cause is not a replication-transport failure — it's automatic invalidates (replication revalidation / TTL eviction / peer-sent invalidate) deleting blobs that live records still reference, recurring on a
replicate:truecluster. This issue still stands as the observability/durability gap (the loss was silent — no alert, no self-repair,cluster_statusshowed healthy). Repointing the data-loss mechanism to #1302; please treat this one as the 'detect + surface + repair' follow-up.Splitting the Ask into the two halves:
- Observability (detect + surface): PR feat(replication): surface blob-replication divergence in cluster_status (#386) #387 adds a blob-replication divergence signal — per-peer cumulative blob-failure count + last-failure time recorded at the authoritative durability-gap event (the blob-save
.catchthat setshasBlobGap), surfaced incluster_statusasblobReplicationFailures/lastBlobFailure, plus one sustained-divergence escalation log per connection. This closes the "a node silently diverged with no operator signal" gap (the socket reportingconnected: truewhile blobs are absent). - Self-repair (backfill): fix(replication): watermark-based blob-gap handling (no data loss + no deadlock) #368 (merged) already covers the durability + reconnect-heal path — the resume cursor holds on an in-flight blob gap and re-streams on reconnect, so no new loss. What remains uncovered is proactive backfill of already-committed records whose blobs are missing/corrupt (the cursor already advanced past them — the shape seen here). Filed as Proactive blob backfill: repair already-committed records whose blobs are missing/corrupt (no recovery once the resume cursor advances past them) #388.
Leaving this open as the umbrella; the determined root cause is being handled in harper#1302.
- Observability (detect + surface): PR feat(replication): surface blob-replication divergence in cluster_status (#386) #387 adds a blob-replication divergence signal — per-peer cumulative blob-failure count + last-failure time recorded at the authoritative durability-gap event (the blob-save
- changed the title
[-]Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair (canopy.serent: 344/345 attachments, ~663 MB)[/-][+]Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB)[/+]on Jun 20, 2026 - added a commit that references this issue
on Jun 22, 2026 Closing — the full fix chain for this shipped, and this issue has been an umbrella over closed work for a while.
The reporter repointed the loss mechanism mid-issue to harper#1302 (RecordEncoder/Table.ts revalidation misreading replicated writes as invalidates, shedding still-referenced blobs). Everything in that chain merged and shipped:
PR What it fixed Shipped harper#1304 Root cause — shouldRevalidateEventsnow excludesintermediateSource(Table.ts:642)v5.1.1 #387 Surfaces blob-replication divergence in cluster status ( clusterStatus.ts:68-69)v5.1.2 #405 ENOENT wedge — isPermanentSourceBlobErrorCodeadvances the resume cursor past a missing sourcev5.1.4 harper#1425 Blob descriptor-size validation + bounded read timeout v5.1.7 #443 Incomplete-blob wedge, pairs with harper#1425 v5.1.7 #418 Proactive backfill sweep ( replication/blobRepair.ts)v5.1.13 All of that is confirmed present and unreverted on current
main(5.2.x). harper#1303 was an earlier attempt deliberately closed unmerged in favour of #1304 — an abandon-and-redo, not a gap.No residual was found against this issue's own asks. The data loss it recorded — 344 of 345 attachments, ~663 MB — was unrecoverable at the time and remains a historical fact about that cluster; nothing here reopens it.
Worth preserving from this issue: it is the clearest field example of the false-green pattern now tracked in harper#2095, since
cluster_statusreportedconnected: truethroughout weeks of silent blob loss. That framing outlived the defect.— KrAIs (Claude Opus 5)
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
On
[customer-cluster](2-node Fabric, GCP, harper-pro 5.1.0), 344 of 345data.CompanyDocument.fileDataattachments (~663 MB) are unreadable and unrecoverable from the cluster. Read-only investigation shows the cluster's blob replication (central → east4) has been timing out since at least 2026-05-08, on 5.0.x — over a month before the 5.1.0 upgrade. The upgrade surfaced/worsened it (and added corrupt 8-byte stubs, see #385), but did not originate the loss.Findings (read-only live investigation, 2026-06-15)
fileData; 344 throwIncomplete blobon read (~663 MB declared, bytes absent on disk); 1 readable (862 KB). Surviving blob files are a contiguous low-id range only.<rootPath>/backup/holds only config.yaml.bak).hdb.log:Timeout waiting for blob stream to finish <id> for record <uuid> from upv-us-central1...recurring since 2026-05-08 (pre-upgrade).cleanup_orphan_blobs/ orphan-delete log lines. No forced-base-copy signature in logs (the Replication: force bounded base-copy resync when a peer is behind > audit retention #277 / forced-resync trigger is unconfirmed here).Gap
Blob replication can fail continuously for weeks and end in permanent data loss while:
[warn]timeouts (no escalation / alert / metric that an operator would notice);cluster_statusreports the socketconnected: true(no blob-level divergence signal);So a node can silently diverge to the point of unrecoverable loss with no operational signal.
Ask
Related
Recovery for this cluster requires an external source of truth (app re-upload, app object storage, or a GCP persistent-disk snapshot of the east4 volume from before 2026-06-11 20:37 UTC).