Skip to content

Replicated read_log hangs when a cluster node is down (Studio times out) #172

Description

@kriszyp

A replicated read_log request hangs (and the Studio times out) when any node in the cluster is down. The expected behavior is to return logs from the reachable nodes and either omit or annotate the unreachable ones.

Reproduction

With one Fabric central manager stopped, the Studio's read_log call hangs:

{
  "operation": "read_log",
  "start": 0,
  "replicated": true,
  "limit": 100,
  "order": "desc"
}

Ask

  • The replicated read_log fan-out should treat a node-down as a normal partial response, not a hang.
  • Per-node response should include either the log entries or an error/timeout marker.
  • A bounded per-node timeout so a slow/unreachable node can't stall the aggregate response.

Acceptance criteria

  • With one node down, read_log replicated=true returns within a bounded time with logs from the reachable nodes and a clear indicator that the down node didn't respond.
  • Studio renders the partial result rather than timing out the whole request.

Tracked in Jira: CORE-2960

🤖 Filed by Claude on behalf of Kris.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:replicationReplication, cluster sync, peer connectionsbugSomething isn't workingfrom-jiraMigrated or originated from a Jira ticket

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions