Repository navigation
feat(cluster): report and converge a Raft voter set that disagrees with node roles (#880) - #896
Merged
Merged
Conversation
…th node roles (#880) #862 made Raft suffrage follow the node's role at JOIN time and deliberately left existing clusters alone: adding a server that is already a voter as a non-voter updates its address and leaves suffrage untouched, which is hashicorp/raft's documented behaviour. Most clusters converge anyway through the graceful-leave path, but a node killed ungracefully, one whose leave found no leader, and the departing leader itself do not — and nothing reported it. The only record of a server's suffrage was inside raft.stats' latest_configuration, an opaque fmt.Sprintf dump. Verified on the real upgrade path rather than a simulation: a cluster formed on a pre-#862 binary, SIGKILLed so no leave was broadcast, restarted on this one over the same Raft data. - GET /api/v1/cluster gains raft.membership: per-server suffrage, the role on record, the suffrage that role calls for, whether they agree, and counts. Roles resolve from the FSM node table first and the registry second, because nodeFromRaftInfo runs ParseRole and so LAUNDERS an unrecognised role into standalone — which votes. Registry-first would report the exact record this exists to find as correct. Unrecognised and empty roles are unresolved, never a mismatch. - A rate-limited leader-only warning when mismatches persist, with its OWN slot: the existing compactor warnings have separate timers for a stated reason, and sharing one would let a voter mismatch mute #856's writer-redundancy warning. - POST /api/v1/cluster/voters/converge, admin-only, dry-run by default, demote-only. Demote-only is not automatically safe, which is the part worth reviewing. A configuration change takes effect the moment it is dispatched and commitment is recomputed over the NEW voter set, so revoking a LIVE voter out of a set whose survivors are dead strands the entry: it can never commit, no further membership change is possible, the leader steps down, and Arc has no RecoverCluster. So every revocation is gated on the liveness of the set it would leave behind, re-checked against what actually remains before each one, and the loop stops at the first failure rather than continuing into a set it can no longer predict. Collapsing to a single voter needs explicit opt-in. It never promotes. Promotion raises the quorum requirement and is the unrecoverable direction; it is also unnecessary, because AddVoter on an existing non-voter does promote, so a node wrongly demoted repairs itself on its next join. That asymmetry is what makes demote-only tolerable. If the node you call is itself voting when its role says otherwise — routine on a cluster upgraded from before #862, where every joiner was AddVoter'd regardless of role, so a reader can be the leader — it hands leadership to a node whose role votes and returns 409 to re-run. Without that it would revoke everyone else's vote, report success, and leave the cluster exactly as it was: a reader leading, and IsPrimaryWriter matching nobody, so retention, continuous queries and deletes stay dead. Nothing converges automatically. Doing it on leadership acquisition was considered and rejected for this release. Claude-Session: https://claude.ai/code/session_01So2gKWp5TzF9gu3QKNqdeV
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #880. Follow-up to #862, which made suffrage follow the role at join time and deliberately left existing clusters alone.
Why an upgraded cluster gets stuck
Adding a server that is already a voter as a non-voter updates its address and leaves suffrage untouched — hashicorp/raft's documented behaviour, pinned by an existing Arc test. Most clusters converge anyway via graceful leave →
RemoveServer→ clean re-add. What does not: a node killed ungracefully, a node whose leave found no leader, the departing leader itself, and any node re-added while still in the configuration.And there was nothing to look at. A server's suffrage appeared only inside
raft.stats.latest_configuration, an opaquefmt.Sprintfdump.What ships
GET /api/v1/cluster→raft.membership— per-server suffrage, the role on record, the suffrage that role calls for, whether they agree, plus counts and the node the view came from.POST /api/v1/cluster/voters/converge— admin-only, dry-run by default, demote-only.The two things worth a reviewer's time
Demote-only is not automatically safe. A configuration change takes effect the moment it is dispatched and commitment is recomputed over the new voter set, so revoking a live voter out of a set whose survivors are dead strands the entry: it can never commit, no further membership change is possible, the leader steps down, and Arc has no
RecoverCluster. Every revocation is gated on the liveness of the set it would leave behind, re-checked against what actually remains before each one, and the loop stops at the first failure — deep review caught that the original guard validated only prefixes of the plan, so a mid-plan failure produced a subset nothing had checked. 2→1 voters needs explicit opt-in.It never promotes: that is the unrecoverable direction, and it is unnecessary because
AddVoteron an existing non-voter does promote, so a wrongly-demoted node repairs itself on its next join.Self is the headline case. Before #862 every joiner was
AddVoter'd regardless of role, so a legacy cluster routinely has a reader as Raft leader. Converging the others while the leader is itself a mismatched voter would revoke everyone else's vote, report success, and leave the cluster exactly as it was — a reader leading,IsPrimaryWritermatching nobody, retention/CQ/deletes dead. It now hands leadership to a node whose role votes and returns 409 to re-run.Live verification — the real upgrade path
A pre-#862 binary was built in a detached worktree, used to form the cluster (where the reader genuinely becomes a voter), then everything SIGKILLed — no leave broadcast — and restarted on this binary over the same Raft data.
And the reader-as-leader case, arranged by bootstrapping a second legacy cluster on the reader: dry run reported
would_move_leadership_to: arc-w1(nothing moved); the real run moved leadership and returned 409; the re-run demoted the reader; andarc-w1 writer_state=primary— that cluster had had no primary writer at all.The warning fired on a third run left unconverged: 2:06 after upgrade, matching the sustain period, only on the leader.
Test plan
go build ./...,gofmt -lempty,go vet ./internal/...cleango test -race ./internal/cluster/... ./internal/api/...greenvoterMismatchCountreads the real Raft config, so a test feeding it a synthetic one asserted nothing; and the partial-failure response had no testraft.statsalready publish roles and suffrage)Two issues filed and closed during this work — #894 and #892 — both on a false "nothing warns" premise. Arc does warn in both cases; my local rig had
ARC_CLUSTER_REPLICATION_ENABLED=false, which disarmsWarnIfNoCompactor.https://claude.ai/code/session_01So2gKWp5TzF9gu3QKNqdeV