Found while fixing #850 (no initial primary is elected). Separate defect, same feature.
The clustering docs describe Pattern 1 as "single-writer + multi-reader", say the readers are the failover pool ("Deploy 1 writer + 2+ readers... the readers are the failover pool"), and promise that on writer loss "the Raft leader selects the most caught-up reader (by replication LSN)" and promotes it. The Enterprise Helm chart's Pattern 1 example is one writer and three readers with failover.enabled: true.
The code cannot do that. WriterFailoverManager.selectNewPrimary only considers Registry.GetStandbyWriters() and Registry.GetWriters(), and both filter on Role == RoleWriter. A RoleReader node is never a candidate, and nothing anywhere changes a node's role. So in the documented topology, when the single writer dies there is nothing to promote and the cluster has no writer at all. Nor is any LSN comparison implemented; selection is first-match over the writer list.
With #850 fixed, that single writer is now elected primary at startup so retention and CQ run, but its loss is still terminal for ingest until an operator intervenes.
Fix shape (needs design): decide whether Pattern 1 HA means promoting a reader to writer at runtime, which requires a role change through Raft plus a defined WAL catch-up and cutover, or whether Pattern 1 should require two or more writer-role nodes and the docs and chart example should say so. The second is much cheaper and matches what the code already supports. Either way the docs and the chart must stop promising reader promotion until it exists.
Found while fixing #850 (no initial primary is elected). Separate defect, same feature.
The clustering docs describe Pattern 1 as "single-writer + multi-reader", say the readers are the failover pool ("Deploy 1 writer + 2+ readers... the readers are the failover pool"), and promise that on writer loss "the Raft leader selects the most caught-up reader (by replication LSN)" and promotes it. The Enterprise Helm chart's Pattern 1 example is one writer and three readers with
failover.enabled: true.The code cannot do that.
WriterFailoverManager.selectNewPrimaryonly considersRegistry.GetStandbyWriters()andRegistry.GetWriters(), and both filter onRole == RoleWriter. ARoleReadernode is never a candidate, and nothing anywhere changes a node's role. So in the documented topology, when the single writer dies there is nothing to promote and the cluster has no writer at all. Nor is any LSN comparison implemented; selection is first-match over the writer list.With #850 fixed, that single writer is now elected primary at startup so retention and CQ run, but its loss is still terminal for ingest until an operator intervenes.
Fix shape (needs design): decide whether Pattern 1 HA means promoting a reader to writer at runtime, which requires a role change through Raft plus a defined WAL catch-up and cutover, or whether Pattern 1 should require two or more writer-role nodes and the docs and chart example should say so. The second is much cheaper and matches what the code already supports. Either way the docs and the chart must stop promising reader promotion until it exists.