Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 22 additions & 2 deletions RELEASE_NOTES_2026.09.2.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ literals (e.g. a log-search `LIKE '%INTERVAL 5 MINUTE%'`) and identifiers such a

### The Helm chart now deploys three writers by default

Readers replicate the write-ahead log and serve queries, but they are never promoted to writer, so the failover pool is made of writer-role nodes. A deployment with one writer therefore cannot fail over, and two absorbs exactly one failure before it is down to a single writer with no spare left and no pod that can be drained for a rolling upgrade, which is why the chart already refused it. The writer default moves from one to three, which is the minimum for high availability in both deployment patterns: on shared storage all three take ingest behind the load balancer, and on local storage, with `cluster.failover_enabled` on, one is elected primary while the other two stand by. Leave failover off on local storage and there is no primary election at all: every writer-role node treats itself as the primary for retention, continuous queries and deletes, which is why Arc now warns about that shape too. The shared-storage example overlay moves to three as well, so following the Pattern 2 guide no longer silently overrides the new default back to one.
Readers replicate the write-ahead log and serve queries, but they are never promoted to writer, so the failover pool is made of writer-role nodes. A deployment with one writer therefore cannot fail over, and two absorbs exactly one failure before it is down to a single writer with no spare left and no pod that can be drained for a rolling upgrade, which is why the chart already refused it. The writer default moves from one to three, which is the minimum for high availability in both deployment patterns: on shared storage all three take ingest behind the load balancer, and on local storage, with `cluster.failover_enabled` on, one is elected primary while the other two stand by. Leaving failover off on local storage used to mean no primary election at all, so every writer-role node treated itself as the primary for retention, continuous queries and deletes. A primary is now elected either way; what failover adds is a replacement chosen automatically when that primary dies. Arc still warns about a cluster running below three writers, because losing the one it has is then a manual recovery. The shared-storage example overlay moves to three as well, so following the Pattern 2 guide no longer silently overrides the new default back to one.

Existing installs are unaffected until they next apply their own values. An install that deliberately wants a single writer, for development or a single-node deployment, can still set the count to one.

Expand Down Expand Up @@ -388,6 +388,26 @@ Check `cluster.role` on every node before upgrading a cluster. The accepted valu

## Bug fixes

### Every writer ran retention and continuous queries when automatic failover was off ([#872](https://github.com/Basekick-Labs/arc/issues/872))

On local storage, retention, continuous queries and deletes are meant to run on one writer. Deciding which one requires a promotion through Raft, and nothing issued one unless writer failover was both enabled and licensed. With no promotion, every writer-role node considered itself the primary and ran all of it. Three writers meant three nodes executing the same continuous queries and the same deletes.

Arc now elects a primary writer on any local-storage cluster with Raft, whatever `cluster.failover_enabled` says and whatever the licence contains. Having one writer in charge is not a paid capability; it is the difference between a working cluster and one doing everything three times.

What the flag and the licence gate is unchanged and is what the feature is named for: **automatic** failover, meaning a replacement chosen for you when the primary dies. A cluster without it elects a first primary and keeps it. If that primary goes away, nothing promotes a successor, and Arc says so in a rate-limited warning rather than leaving you to infer it.

The line between the two is drawn on durable cluster state rather than on anything a process remembers. A cluster that has never designated a primary needs one, and that is bootstrapping. A cluster whose designated primary is unhealthy needs a replacement, and that is failover. The old code made this distinction with an in-memory field, which is empty on a freshly started process — so a leader restart or a leadership change looked like a cluster that had never had a primary, and it would elect one. That was a replacement under another name.

**New: `POST /api/v1/cluster/writers/{id}/demote`.** Hands the primary-writer role off the named node so the cluster elects a new one. Admin-only, and not gated on the failover licence, because it is how a cluster without automatic failover recovers: the record still names the node that died, so nothing elects until an operator clears it. On a cluster that does have automatic failover, it is also how you drain a writer deliberately. It refuses a node that is not the current primary rather than appearing to succeed, and it names which node the cluster actually has on record.

Handing over is a promotion of somebody else rather than a demotion of the node named, because a promotion already carries the demotion with it and announces both sides at once. Demoting and letting the cluster elect would leave it free to choose the same node straight back, which is not a hand-over. When there is nobody else to promote, the designation is released anyway so the cluster is not pinned to a writer that may never return, and the response says plainly that nobody took it.

Two things were fixed that this endpoint would otherwise have exposed. A demotion announced nothing, so the node being demoted kept believing it was the primary and kept running the singleton work, while the cluster still saw a live primary and never elected anyone. And removing the designated primary left the record naming a node that no longer existed, which meant no election could ever run again — no primary, no retention, no continuous queries, no deletes, and nothing reporting why. Removing the dead primary is the obvious operator move, so that was the likely path.

**Behaviour change worth knowing about.** On local storage, `/ready/write` now reports not-ready on a writer that is not the primary. Clusters with automatic failover already behaved this way; clusters without it reported every writer ready, which was consistent with every writer also running the singleton work, and both were wrong. A load balancer pool built on that endpoint will drop to the primary alone. Kubernetes pod readiness is unaffected, since the chart probes `/ready`.

A crash was fixed along the way. The failover path derived its timeout from a context that only exists once the manager has started, while the callback that can reach it is wired one line earlier, with health checks already running. Hitting that window panicked the goroutine and took the process down.

### A node that left gracefully and restarted was never listed again ([#858](https://github.com/Basekick-Labs/arc/issues/858))

Restart a cluster's Raft leader with a normal shutdown and it came back healthy, serving, and invisible. It listed every node; every other node listed everything except it. Nothing recovered from that. On local storage the detached node also kept its primary-writer designation across the restart and went on running retention and continuous queries the rest of the cluster had moved past.
Expand Down Expand Up @@ -460,7 +480,7 @@ A cluster running below three writer-role nodes now logs a rate-limited warning

The deficit has to hold for two minutes before anything is logged. That covers the two ways a healthy cluster passes through a low writer count on its way somewhere else. A cluster is short of writers for the first seconds of its life while peers are still joining. And a rolling upgrade cycles one writer at a time, and a leaving node tells its peers to drop it, so a correct three-writer cluster walks through two writers on every node, every time. A single-node install never warns at all: it has no redundancy of any kind and its operator knows that. The shape this exists for is the one that looks highly available, several nodes of which exactly one is a writer.

Every cluster mode gets the warning, with the right explanation for each. With local storage and failover enabled, the message explains that readers are never promotion candidates. With shared storage, that writer promotion is deliberately suppressed in that pattern, so the load balancer's backend count is the only thing between a writer crash and an ingest outage. With local storage and no failover manager, because the flag is off or the license does not carry the feature, it says that nothing promotes anything, and it warns against adding writers without first enabling failover: with no failover manager every writer-role node considers itself the primary for retention, continuous queries and deletes.
Every cluster mode gets the warning, with the right explanation for each. With local storage and failover enabled, the message explains that readers are never promotion candidates. With shared storage, that writer promotion is deliberately suppressed in that pattern, so the load balancer's backend count is the only thing between a writer crash and an ingest outage. With local storage and no automatic failover, because the flag is off or the license does not carry the feature, it says that a primary is elected but no replacement will be chosen when it goes, and it names the endpoint that hands the role over by hand.

### Shutdown now stops replication, and stopping it cannot deadlock ([#853](https://github.com/Basekick-Labs/arc/issues/853))

Expand Down
85 changes: 85 additions & 0 deletions internal/api/cluster.go
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
package api

import (
"errors"
"sort"
"strconv"

Expand Down Expand Up @@ -80,6 +81,90 @@ func (h *ClusterHandler) RegisterRoutes(app *fiber.App) {
removeGroup.Use(auth.RequireAdmin(h.authManager))
}
removeGroup.Delete("", h.handleRemoveNode)

// Admin-only: hands the primary-writer role off a node so the cluster
// elects a new one. Deliberately NOT gated on the writer_failover licence
// — that feature is AUTOMATIC failover, and this is the manual recovery a
// cluster without it needs (#872).
writerGroup := app.Group("/api/v1/cluster/writers/:id")
if h.authManager != nil {
writerGroup.Use(auth.RequireAdmin(h.authManager))
}
writerGroup.Post("/demote", h.handleDemoteWriter)
}

// handleDemoteWriter hands the primary-writer role off the named node.
//
// The caller does not choose a successor, and should not: clearing the
// designation is what lets the cluster's own election run, and it already
// applies the selection rules. On a cluster with automatic failover this is a
// way to drain a writer deliberately; on one without, it is the only way to
// recover after the primary is gone for good.
func (h *ClusterHandler) handleDemoteWriter(c *fiber.Ctx) error {
// Validate the input before anything else, so a malformed request is
// rejected the same way whether or not this node is clustered.
nodeID := c.Params("id")
if len(nodeID) == 0 || len(nodeID) > maxNodeIDLength {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"success": false,
"error": "invalid node ID",
})
}

// Deliberately NOT respondNotEnabled, which the read endpoints use: it
// answers 200 with enabled=false, which is right for "describe yourself"
// and wrong for "do this". A caller checking the status code would read it
// as a hand-over that happened.
if h.coordinator == nil {
return c.Status(fiber.StatusConflict).JSON(fiber.Map{
"success": false,
"error": "clustering is not enabled on this node, so there is no primary writer to hand over",
})
}

newPrimary, err := h.coordinator.DemoteWriterViaRaft(nodeID)
if err != nil {
// A wrong target is the operator's mistake to see and correct, not a
// server fault: it names which node the cluster actually has on
// record. Not being the leader is a 409 for the same reason — retry
// against the leader.
status := fiber.StatusInternalServerError
switch {
case errors.Is(err, cluster.ErrNotPrimaryWriter):
status = fiber.StatusConflict
case errors.Is(err, cluster.ErrNotLeaderForTopology):
status = fiber.StatusConflict
}
if status == fiber.StatusInternalServerError {
h.logger.Error().Err(err).Str("node_id", nodeID).Msg("Failed to hand over the primary writer")
}
return c.Status(status).JSON(fiber.Map{
"success": false,
"error": err.Error(),
})
}

h.logger.Info().
Str("node_id", nodeID).
Str("new_primary", newPrimary).
Msg("Primary writer handed over via API")

// Report who took it, so the operator does not have to go looking — and
// say plainly when nobody did, which is the single-writer case.
if newPrimary == "" {
return c.JSON(fiber.Map{
"success": true,
"node_id": nodeID,
"new_primary": nil,
"message": "primary writer designation released, but no other writer was available to take it — this cluster has no primary until one is",
})
}
return c.JSON(fiber.Map{
"success": true,
"node_id": nodeID,
"new_primary": newPrimary,
"message": "primary writer handed over",
})
}

// handleGetStatus returns the overall cluster status.
Expand Down
83 changes: 83 additions & 0 deletions internal/api/cluster_demote_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
package api

import (
"net/http/httptest"
"testing"

"github.com/gofiber/fiber/v2"
"github.com/rs/zerolog"
)

// #872: the manual hand-over endpoint. It is what a cluster WITHOUT automatic
// writer failover recovers with — the FSM still names a dead primary, so
// nothing elects until an operator clears the designation. It is therefore
// deliberately not gated on the writer_failover licence, only on admin auth.

func demoteTestApp(t *testing.T) *fiber.App {
t.Helper()
// A nil coordinator is the OSS/standalone shape. It exercises routing and
// the not-enabled response without needing a Raft cluster; the licence
// boundary itself is tested in internal/cluster.
h := NewClusterHandler(nil, nil, nil, zerolog.Nop())
app := fiber.New()
h.RegisterRoutes(app)
return app
}

// The route has to exist and be a POST. A typo here means an operator with a
// dead primary has no recovery path at all, which is the whole reason the
// endpoint was added.
func TestDemoteWriterRouteIsRegistered(t *testing.T) {
app := demoteTestApp(t)

resp, err := app.Test(httptest.NewRequest("POST", "/api/v1/cluster/writers/writer-1/demote", nil))
if err != nil {
t.Fatalf("app.Test: %v", err)
}
if resp.StatusCode == fiber.StatusNotFound || resp.StatusCode == fiber.StatusMethodNotAllowed {
t.Fatalf("POST /api/v1/cluster/writers/:id/demote is not routed (status %d)", resp.StatusCode)
}
}

// Without clustering the endpoint must say so rather than 404 or 500.
func TestDemoteWriterWithoutClustering(t *testing.T) {
app := demoteTestApp(t)

resp, err := app.Test(httptest.NewRequest("POST", "/api/v1/cluster/writers/writer-1/demote", nil))
if err != nil {
t.Fatalf("app.Test: %v", err)
}
if resp.StatusCode == fiber.StatusOK {
t.Error("a node with no coordinator reported a successful hand-over")
}
}

// A GET must not perform it. Guards against someone widening the group later.
func TestDemoteWriterRejectsGet(t *testing.T) {
app := demoteTestApp(t)

resp, err := app.Test(httptest.NewRequest("GET", "/api/v1/cluster/writers/writer-1/demote", nil))
if err != nil {
t.Fatalf("app.Test: %v", err)
}
if resp.StatusCode != fiber.StatusMethodNotAllowed && resp.StatusCode != fiber.StatusNotFound {
t.Errorf("GET on the hand-over route returned %d; it must not be reachable by GET", resp.StatusCode)
}
}

// An over-long node ID is rejected before it reaches Raft.
func TestDemoteWriterRejectsAnOverlongID(t *testing.T) {
app := demoteTestApp(t)

long := make([]byte, maxNodeIDLength+10)
for i := range long {
long[i] = 'a'
}
resp, err := app.Test(httptest.NewRequest("POST", "/api/v1/cluster/writers/"+string(long)+"/demote", nil))
if err != nil {
t.Fatalf("app.Test: %v", err)
}
if resp.StatusCode == fiber.StatusOK {
t.Error("an over-long node ID was accepted")
}
}
Loading
Loading