Found while answering how Pattern 2 is meant to fit together. Related to #850, but a distinct defect and, unlike #850, it affects the shared-storage pattern.
The gate. In shared-storage mode Coordinator.IsPrimaryWriter() is "I am the Raft leader AND my role is writer". The role half is deliberate and correct: a reader must never run retention, continuous queries or deletes against shared storage.
The hole. handleJoinRequest calls AddVoter for every node that joins, whatever its role, and the code comment at the gate acknowledges it ("every joining node becomes a Raft voter regardless of role"). So a reader or the compactor can be elected Raft leader. When that happens the leader fails the role half and every writer fails the leader half, so no node in the cluster passes the gate: retention, continuous queries and the non-dry-run retention, CQ and delete endpoints all stop until leadership happens to return to a writer. Nothing forces that to happen; there is no leadership transfer anywhere in the tree.
Compaction is unaffected, because it is gated on the compactor lease rather than on writer state.
Fix: only nodes that take ingest should vote. hashicorp/raft supports this directly through AddNonvoter, which we never use. A non-voter replicates the log and serves reads but never campaigns and cannot be elected. Join should dispatch on role: writers (and standalone, which ingests) join as voters; readers and compactors join as non-voters.
Consequence to accept deliberately. Quorum is then computed among writers only. Three writers survive losing one, which is the documented Pattern 2 shape. A cluster with a single writer makes that writer the only voter: it is always leader, but its loss takes Raft's quorum with it, so token and permission writes and cluster reconfiguration stop too, where today the readers would keep Raft alive. That is the correct trade — a cluster that cannot ingest should not be pretending to be healthy — but it means the documented Pattern 1 topology of one writer plus readers is not HA, which is the same conclusion #856 reaches from the promotion side. The guidance for both patterns should become three writer-role nodes.
Work involved.
- Wrap
AddNonvoter on the raft node alongside AddVoter, and dispatch on role in handleJoinRequest.
- Refuse to bootstrap on a non-ingest role, or the first node of a cluster is an ineligible leader by construction.
- Decide the migration for existing clusters: a node already added as a voter stays one until it is removed and re-added, so either accept that this applies to new joins only, or demote on the next join.
- Tests: a reader joins as a non-voter and never appears in the voter configuration; a cluster of writers plus readers always elects a writer; a single-writer cluster reports the quorum consequence clearly rather than silently.
- Docs and the Helm chart: three writers, in both patterns.
Found while answering how Pattern 2 is meant to fit together. Related to #850, but a distinct defect and, unlike #850, it affects the shared-storage pattern.
The gate. In shared-storage mode
Coordinator.IsPrimaryWriter()is "I am the Raft leader AND my role is writer". The role half is deliberate and correct: a reader must never run retention, continuous queries or deletes against shared storage.The hole.
handleJoinRequestcallsAddVoterfor every node that joins, whatever its role, and the code comment at the gate acknowledges it ("every joining node becomes a Raft voter regardless of role"). So a reader or the compactor can be elected Raft leader. When that happens the leader fails the role half and every writer fails the leader half, so no node in the cluster passes the gate: retention, continuous queries and the non-dry-run retention, CQ and delete endpoints all stop until leadership happens to return to a writer. Nothing forces that to happen; there is no leadership transfer anywhere in the tree.Compaction is unaffected, because it is gated on the compactor lease rather than on writer state.
Fix: only nodes that take ingest should vote.
hashicorp/raftsupports this directly throughAddNonvoter, which we never use. A non-voter replicates the log and serves reads but never campaigns and cannot be elected. Join should dispatch on role: writers (andstandalone, which ingests) join as voters; readers and compactors join as non-voters.Consequence to accept deliberately. Quorum is then computed among writers only. Three writers survive losing one, which is the documented Pattern 2 shape. A cluster with a single writer makes that writer the only voter: it is always leader, but its loss takes Raft's quorum with it, so token and permission writes and cluster reconfiguration stop too, where today the readers would keep Raft alive. That is the correct trade — a cluster that cannot ingest should not be pretending to be healthy — but it means the documented Pattern 1 topology of one writer plus readers is not HA, which is the same conclusion #856 reaches from the promotion side. The guidance for both patterns should become three writer-role nodes.
Work involved.
AddNonvoteron the raft node alongsideAddVoter, and dispatch on role inhandleJoinRequest.