Skip to content

Inter-node blob fetch timeout surfaces as uncaughtException (verify still occurs after May fixes) #158

Description

@kriszyp

Investigative — needs verification on current build.

When a node attempts to read a blob physically stored on a peer that is temporarily unavailable, the resulting "File read timed out" error has been observed surfacing as an uncaughtException in the worker thread rather than being handled gracefully.

Original customer impact

a customer production cluster upgrade: [redacted-customer-host] fetched blobs whose source was on [redacted-customer-host]. During central01's compose cycle, MIA01's inter-node blob fetch timed out and threw:

[http/2] [error]: uncaughtException Error: Blob error: Error: File read timed out reading from /home/harperdb/hdb/blobs/blobs2/Resilience/9/bda/12b for record ... from [redacted-customer-host]
  at _write (node:internal/streams/writable:501:10)
  ...
  at Receiver._write (.../ws/lib/receiver.js:94:10)

Errors were transient and self-resolved when central01 came back online — but they surfaced as uncaughtException in the meantime.

Recent related fixes

Two commits since the ticket was filed (2026-03-16) address blob-related uncaughtException patterns in replication/replicationConnection.ts:

  • 22f8777 (2026-05-01) "fix: handle blob save promise rejection to prevent uncaughtException" — adds .catch() to the blob save promise so rejections are consumed rather than escaping.
  • a686a55 (2026-05-14) "fix(replication): swallow blob save rejection in outstandingBlobsToFinish" — follow-up: the .catch() returned a new promise, but the array still held the raw rejected promise. Promise.all(outstandingBlobsToFinish) in the end_txn onCommit path was still surfacing the rejection — observed in prod as "~35/sec ENOENT spam during catch-up".

Both fix a similar pattern. However, the CORE-3047 stack trace ends in ws/lib/receiver.js:94/_write, suggesting a synchronous throw inside the WebSocket message handler, which is a different mechanism from Promise.all rejection. The May fixes may have reduced or eliminated this symptom, but not necessarily covered the exact path.

To investigate

  • Verify with SRE / a customer whether the original "Blob error: ... File read timed out" uncaughtException is still occurring on builds ≥ 4.7.26 (after both May fixes shipped).
  • The "Blob error:" prefix is constructed at replication/replicationConnection.ts:797 where stream.on('error', () => {}) is intended to swallow:
    stream.on('error', () => {}); // don't treat this as an uncaught error
    stream.destroy(new Error('Blob error: ' + error + ...));
    Confirm whether there is a code path that throws synchronously inside the WS message dispatch before the swallow takes effect.
  • Define the desired peer-unavailable behavior (silent retry / queue / propagate to consumer) and document.

Acceptance criteria (once investigation completes)

  • Inter-node blob fetch failures never reach uncaughtException.
  • Peer-unavailable behavior is defined and documented.
  • Repro recipe captured for regression coverage.

Tracked in Jira: CORE-3047 (a customer escalation)
Related: CORE-3046 (TypeError in incoming replication message handling — likely related)
Status: Not Ready — needs SRE verification first.

🤖 Filed by Claude on behalf of Kris.

Activity

  1. added
    bugSomething isn't working
    area:replicationReplication, cluster sync, peer connections
    from-jiraMigrated or originated from a Jira ticket
    on May 18, 2026
  2. added this to the v5.2 milestone on May 21, 2026
  3. kriszyp commented on Jul 7, 2026

    @kriszyp
    MemberAuthor

    Appears to be a duplicate of / consolidate under #195 — same blob-transfer uncaughtException failure class.

    — Claude (Opus 4.8), issue-backlog triage

  4. added
    duplicateThis issue or pull request already exists
    on Jul 7, 2026
  5. added theissue type on Aug 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:replicationReplication, cluster sync, peer connectionsbugSomething isn't workingduplicateThis issue or pull request already existsfrom-jiraMigrated or originated from a Jira ticket

    Type

    Fields

    Priority

    P3

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions