Skip to content

Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB) #386

Description

@kriszyp

Summary

On [customer-cluster] (2-node Fabric, GCP, harper-pro 5.1.0), 344 of 345 data.CompanyDocument.fileData attachments (~663 MB) are unreadable and unrecoverable from the cluster. Read-only investigation shows the cluster's blob replication (central → east4) has been timing out since at least 2026-05-08, on 5.0.x — over a month before the 5.1.0 upgrade. The upgrade surfaced/worsened it (and added corrupt 8-byte stubs, see #385), but did not originate the loss.

Findings (read-only live investigation, 2026-06-15)

Gap

Blob replication can fail continuously for weeks and end in permanent data loss while:

  • only emitting [warn] timeouts (no escalation / alert / metric that an operator would notice);
  • cluster_status reports the socket connected: true (no blob-level divergence signal);
  • there is no backfill/repair pass to reconcile missing or corrupt blobs once a link recovers.

So a node can silently diverge to the point of unrecoverable loss with no operational signal.

Ask

Related

Recovery for this cluster requires an external source of truth (app re-upload, app object storage, or a GCP persistent-disk snapshot of the east4 volume from before 2026-06-11 20:37 UTC).

Activity

  1. added
    bugSomething isn't working
    area:replicationReplication, cluster sync, peer connections
    on Jun 15, 2026
  2. kriszyp commented on Jun 15, 2026

    @kriszyp
    MemberAuthor

    On-disk timeline pins the loss event to the beta.2 upgrade (Jun 11)

    Blob-file mtimes confirm the acute loss event was the 5.1.0-beta.2 upgrade on 2026-06-11 ~20:44 UTC, not today's 5.1.0 release upgrade and not a gradual erosion:

    central — data blob mtimes by day: 15 (May 8), 103 (May 12), 290 (May 14), 345 (Jun 11). Of the Jun 11 batch, 264 were written as 8-byte stubs at exactly 20:44 (the beta.2 completion minute). 9 earlier stubs date to May 12 02:26.

    east4 — surviving blobs are all dated May 8 / May 12 (none from Jun 11+); the node received nothing at beta.2, and its blobs/ dir was last modified Jun 11 21:32 (the truncation window). The multi-MB documents lived only on east4 and were removed here.

    So the chronic condition (blob replication failing, east4 not receiving since ~May 12) predates the upgrade, but the document bytes were lost at the beta.2 upgrade — central's good blobs were rewritten to empty stubs and east4's store was truncated in the same window, leaving no complete copy anywhere. Correcting the "did not originate the loss" wording above: the replication degradation predates beta.2; the byte loss coincides with it.

  3. kriszyp commented on Jun 15, 2026

    @kriszyp
    MemberAuthor

    Root cause identified: HarperFast/harper#1302 (RecordEncoder invalidate path unlinks live-referenced blobs → progressive silent loss). The timeline/observations in this issue are correct, but the cause is not a replication-transport failure — it's automatic invalidates (replication revalidation / TTL eviction / peer-sent invalidate) deleting blobs that live records still reference, recurring on a replicate:true cluster. This issue still stands as the observability/durability gap (the loss was silent — no alert, no self-repair, cluster_status showed healthy). Repointing the data-loss mechanism to #1302; please treat this one as the 'detect + surface + repair' follow-up.

  4. self-assigned this
    on Jun 15, 2026
  5. kriszyp commented on Jun 15, 2026

    @kriszyp
    MemberAuthor

    Splitting the Ask into the two halves:

    Leaving this open as the umbrella; the determined root cause is being handled in harper#1302.

  6. changed the title [-]Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair (canopy.serent: 344/345 attachments, ~663 MB)[/-] [+]Sustained blob-replication timeouts cause silent, unrecoverable blob loss with no operator signal or self-repair ([customer-cluster]: 344/345 attachments, ~663 MB)[/+] on Jun 20, 2026
  7. added this to the v5.2 milestone on Jul 7, 2026
  8. kriszyp commented on Aug 7, 2026

    @kriszyp
    MemberAuthor

    Closing — the full fix chain for this shipped, and this issue has been an umbrella over closed work for a while.

    The reporter repointed the loss mechanism mid-issue to harper#1302 (RecordEncoder/Table.ts revalidation misreading replicated writes as invalidates, shedding still-referenced blobs). Everything in that chain merged and shipped:

    PR What it fixed Shipped
    harper#1304 Root cause — shouldRevalidateEvents now excludes intermediateSource (Table.ts:642) v5.1.1
    #387 Surfaces blob-replication divergence in cluster status (clusterStatus.ts:68-69) v5.1.2
    #405 ENOENT wedge — isPermanentSourceBlobErrorCode advances the resume cursor past a missing source v5.1.4
    harper#1425 Blob descriptor-size validation + bounded read timeout v5.1.7
    #443 Incomplete-blob wedge, pairs with harper#1425 v5.1.7
    #418 Proactive backfill sweep (replication/blobRepair.ts) v5.1.13

    All of that is confirmed present and unreverted on current main (5.2.x). harper#1303 was an earlier attempt deliberately closed unmerged in favour of #1304 — an abandon-and-redo, not a gap.

    No residual was found against this issue's own asks. The data loss it recorded — 344 of 345 attachments, ~663 MB — was unrecoverable at the time and remains a historical fact about that cluster; nothing here reopens it.

    Worth preserving from this issue: it is the clearest field example of the false-green pattern now tracked in harper#2095, since cluster_status reported connected: true throughout weeks of silent blob loss. That framing outlived the defect.

    — KrAIs (Claude Opus 5)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:replicationReplication, cluster sync, peer connectionsbugSomething isn't working

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions