Skip to content

feat(cluster): report and converge a Raft voter set that disagrees with node roles (#880) - #896

Merged
xe-nvdk merged 1 commit into
mainfrom
feat/cluster-voter-reconcile
Sep 16, 2026
Merged

xe-nvdk merged 1 commit into
mainfrom
feat/cluster-voter-reconcile

Conversation

@xe-nvdk

@xe-nvdk xe-nvdk commented Sep 16, 2026

Copy link
Copy Markdown
Member

Closes #880. Follow-up to #862, which made suffrage follow the role at join time and deliberately left existing clusters alone.

Why an upgraded cluster gets stuck

Adding a server that is already a voter as a non-voter updates its address and leaves suffrage untouched — hashicorp/raft's documented behaviour, pinned by an existing Arc test. Most clusters converge anyway via graceful leave → RemoveServer → clean re-add. What does not: a node killed ungracefully, a node whose leave found no leader, the departing leader itself, and any node re-added while still in the configuration.

And there was nothing to look at. A server's suffrage appeared only inside raft.stats.latest_configuration, an opaque fmt.Sprintf dump.

What ships

  • GET /api/v1/cluster → raft.membership — per-server suffrage, the role on record, the suffrage that role calls for, whether they agree, plus counts and the node the view came from.
  • A rate-limited leader-only warning when mismatches persist, with its own slot.
  • POST /api/v1/cluster/voters/converge — admin-only, dry-run by default, demote-only.

The two things worth a reviewer's time

Demote-only is not automatically safe. A configuration change takes effect the moment it is dispatched and commitment is recomputed over the new voter set, so revoking a live voter out of a set whose survivors are dead strands the entry: it can never commit, no further membership change is possible, the leader steps down, and Arc has no RecoverCluster. Every revocation is gated on the liveness of the set it would leave behind, re-checked against what actually remains before each one, and the loop stops at the first failure — deep review caught that the original guard validated only prefixes of the plan, so a mid-plan failure produced a subset nothing had checked. 2→1 voters needs explicit opt-in.

It never promotes: that is the unrecoverable direction, and it is unnecessary because AddVoter on an existing non-voter does promote, so a wrongly-demoted node repairs itself on its next join.

Self is the headline case. Before #862 every joiner was AddVoter'd regardless of role, so a legacy cluster routinely has a reader as Raft leader. Converging the others while the leader is itself a mismatched voter would revoke everyone else's vote, report success, and leave the cluster exactly as it was — a reader leading, IsPrimaryWriter matching nobody, retention/CQ/deletes dead. It now hands leadership to a node whose role votes and returns 409 to re-run.

Live verification — the real upgrade path

A pre-#862 binary was built in a detached worktree, used to form the cluster (where the reader genuinely becomes a voter), then everything SIGKILLed — no leave broadcast — and restarted on this binary over the same Raft data.

legacy:  [{Suffrage:Voter ID:arc-w1} {Suffrage:Voter ID:arc-w2} {Suffrage:Voter ID:arc-r1}]
after:   arc-r1 reader voter matches=false   mismatches: 1
dry run: would_demote [arc-r1]   3 -> 2      (real config untouched)
real:    demoted [arc-r1]        mismatches: 0
         note: "this cluster now has 2 Raft voters; 3 is the documented minimum..."
re-join: arc-r1 SIGKILLed and re-joined -> still a non-voter

And the reader-as-leader case, arranged by bootstrapping a second legacy cluster on the reader: dry run reported would_move_leadership_to: arc-w1 (nothing moved); the real run moved leadership and returned 409; the re-run demoted the reader; and arc-w1 writer_state=primary — that cluster had had no primary writer at all.

The warning fired on a third run left unconverged: 2:06 after upgrade, matching the sustain period, only on the leader.

Test plan

  • go build ./..., gofmt -l empty, go vet ./internal/... clean
  • go test -race ./internal/cluster/... ./internal/api/... green
  • 13 design decisions revert-run individually; all 13 fail without their fix — including the prefix-validation blocker, the dry-run gate, the dedicated warn slot, and admin auth
  • Two tests were found passing pre-fix and repaired: voterMismatchCount reads the real Raft config, so a test feeding it a synthetic one asserted nothing; and the partial-failure response had no test
  • Configuration matrix written before review, incl. the unauthenticated-status row (no new exposure — the node list and raft.stats already publish roles and suffrage)
  • Live: legacy→upgrade→converge, re-join durability, reader-as-leader, sustained warning

Two issues filed and closed during this work — #894 and #892 — both on a false "nothing warns" premise. Arc does warn in both cases; my local rig had ARC_CLUSTER_REPLICATION_ENABLED=false, which disarms WarnIfNoCompactor.

https://claude.ai/code/session_01So2gKWp5TzF9gu3QKNqdeV

…th node roles (#880)

#862 made Raft suffrage follow the node's role at JOIN time and deliberately
left existing clusters alone: adding a server that is already a voter as a
non-voter updates its address and leaves suffrage untouched, which is
hashicorp/raft's documented behaviour. Most clusters converge anyway through
the graceful-leave path, but a node killed ungracefully, one whose leave found
no leader, and the departing leader itself do not — and nothing reported it.
The only record of a server's suffrage was inside raft.stats'
latest_configuration, an opaque fmt.Sprintf dump.

Verified on the real upgrade path rather than a simulation: a cluster formed on
a pre-#862 binary, SIGKILLed so no leave was broadcast, restarted on this one
over the same Raft data.

- GET /api/v1/cluster gains raft.membership: per-server suffrage, the role on
  record, the suffrage that role calls for, whether they agree, and counts.
  Roles resolve from the FSM node table first and the registry second, because
  nodeFromRaftInfo runs ParseRole and so LAUNDERS an unrecognised role into
  standalone — which votes. Registry-first would report the exact record this
  exists to find as correct. Unrecognised and empty roles are unresolved, never
  a mismatch.
- A rate-limited leader-only warning when mismatches persist, with its OWN
  slot: the existing compactor warnings have separate timers for a stated
  reason, and sharing one would let a voter mismatch mute #856's
  writer-redundancy warning.
- POST /api/v1/cluster/voters/converge, admin-only, dry-run by default,
  demote-only.

Demote-only is not automatically safe, which is the part worth reviewing. A
configuration change takes effect the moment it is dispatched and commitment is
recomputed over the NEW voter set, so revoking a LIVE voter out of a set whose
survivors are dead strands the entry: it can never commit, no further
membership change is possible, the leader steps down, and Arc has no
RecoverCluster. So every revocation is gated on the liveness of the set it
would leave behind, re-checked against what actually remains before each one,
and the loop stops at the first failure rather than continuing into a set it
can no longer predict. Collapsing to a single voter needs explicit opt-in.

It never promotes. Promotion raises the quorum requirement and is the
unrecoverable direction; it is also unnecessary, because AddVoter on an
existing non-voter does promote, so a node wrongly demoted repairs itself on
its next join. That asymmetry is what makes demote-only tolerable.

If the node you call is itself voting when its role says otherwise — routine on
a cluster upgraded from before #862, where every joiner was AddVoter'd
regardless of role, so a reader can be the leader — it hands leadership to a
node whose role votes and returns 409 to re-run. Without that it would revoke
everyone else's vote, report success, and leave the cluster exactly as it was:
a reader leading, and IsPrimaryWriter matching nobody, so retention, continuous
queries and deletes stay dead.

Nothing converges automatically. Doing it on leadership acquisition was
considered and rejected for this release.

Claude-Session: https://claude.ai/code/session_01So2gKWp5TzF9gu3QKNqdeV
@xe-nvdk
xe-nvdk merged commit 6d611cd into main Sep 16, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cluster: converge an existing cluster's Raft voters onto the role-based rule

1 participant