Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion content/docs/arc-enterprise/configuration/clustering.md
Original file line number Diff line number Diff line change
Expand Up @@ -288,7 +288,7 @@ When each node has its own local storage, one writer at a time takes ingest, bec

**Deploy three writer-role nodes.** The failover pool is made of writer-role nodes, not readers. One is elected primary and takes ingest; the other two replicate and stand by. Readers serve queries and replicate the WAL, but they are not promotion candidates, so a deployment with a single writer cannot fail over at all, and two absorbs exactly one failure before it is back to a single writer with nothing left to promote. Arc logs a rate-limited warning while a cluster is below three writer-role nodes.

Leave `cluster.failover_enabled` off and there is no primary election at all: nothing promotes a replacement, and every writer-role node treats itself as the primary for retention, continuous queries and deletes. Arc warns about that shape specifically. Enable failover before adding writers to a local-storage cluster.
Leave `cluster.failover_enabled` off and there is no primary election at all: nothing promotes a replacement, and every writer-role node treats itself as the primary for retention, continuous queries, deletes and tiered-storage migration. Arc warns about that shape specifically. Enable failover before adding writers to a local-storage cluster.

**Key characteristics:**

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -294,7 +294,9 @@ On a node that syncs but never migrates, `GET /api/v1/tiering/status` reports `"

A node in shared-storage mode must have `cluster.role` set to `writer`, `reader` or `compactor`. A node left at the default `standalone` refuses to start: it could win Raft leadership without ever passing the primary-writer gate, which would stop every singleton task on the cluster.

**Per-node local storage (Pattern 1).** Tiering is not gated: each node migrates its own local copy, and its own metadata is authoritative for its own disk. Tiering does not yet coordinate with the file-replication manifest, so a file a node migrates and deletes locally can be pulled back from a peer. Until that is addressed, prefer shared storage when combining clustering with cold tiering.
**Per-node local storage (Pattern 1).** A per-node cluster whose nodes keep each other in step with file replication (`cluster.replication_enabled = true`) is gated the same way: only the primary writer migrates, every node syncs its tier metadata from the cold tier each cycle, and every file the primary moves to cold is removed from the cluster file manifest before its hot copy goes — which is what unlinks the replica on every other node and stops replication from pulling it back. A per-node cluster without replication shares nothing and keeps migrating per node.

Two requirements follow. **Every replicating node must run tiering with the same cold backend**: once the primary migrates a file, the only way a node can read it is through its own cold row and cold backend; a node with tiering off or no cold tier cannot see that data any more, and Arc says so at startup on such a node. And a node that has no cold row yet for a measurement does not read that measurement's first migrated file until its next cold sync — at the default `0 2 * * *` schedule up to a day; shorten `tiered_storage.migration_schedule` on readers if that matters.

## Best practices

Expand Down