Skip to content

docs(arc-enterprise): tiering keeps the cluster manifest in step - #85

Merged
xe-nvdk merged 1 commit into
mainfrom
docs/tiering-manifest-integration
Sep 28, 2026
Merged

xe-nvdk merged 1 commit into
mainfrom
docs/tiering-manifest-integration

Conversation

@xe-nvdk

@xe-nvdk xe-nvdk commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Replaces the Per-node local storage (Pattern 1) paragraph of the tiered-storage page: a replicating per-node cluster is now gated like shared storage (only the primary writer migrates, every node syncs its tier metadata from the cold tier), and every file the primary moves to cold is removed from the cluster file manifest before its hot copy goes — which unlinks the replica on every other node and stops replication from pulling it back. A per-node cluster without replication keeps migrating per node.
  • Documents the two requirements that follow: every replicating node must run tiering with the same cold backend (a node with tiering off or no cold tier cannot read migrated data any more; Arc warns at startup), and a node with no cold row yet for a measurement sees that measurement's first migrated file only after its next cold sync — up to a day at the default schedule.
  • Lists tiered-storage migration among the primary-writer-gated tasks in the clustering page's Pattern 1 text.

Companion to Basekick-Labs/arc#953 (which stacks on Basekick-Labs/arc#952).

…tern 1 with replication is gated

Replicating per-node clusters now migrate on the primary only, sync cold
rows on every node, and remove migrated files from the file manifest
before deleting hot copies. Document the requirement that every
replicating node run tiering with the same cold backend, and the
first-cold-file visibility lag.

Companion to the Arc PR on branch fix/tiering-manifest-integration (link to follow).
@xe-nvdk
xe-nvdk merged commit a1bb7be into main Sep 28, 2026
xe-nvdk added a commit to Basekick-Labs/arc that referenced this pull request Sep 28, 2026
…replicating per-node clusters

Tiering never told the cluster file manifest when it moved a file to cold.
On a per-node-storage cluster with file replication the puller pulled the
migrated file straight back from a peer, every node re-did the same
migration, and once every peer had migrated a restarted node failed
startup catch-up on the entry (503 on reads with query_gate_on_catchup).
On a shared bucket stale entries accumulated unless the reconciler ran out
of dry-run mode.

A ManifestCoordinator seam in tiering (adapter in cmd/arc, mirroring
retention's JSON DeleteFilePayload batches; no cluster/raft import in
tiering) lets migration remove a file from the manifest before its hot
copy goes: MigrateBatch copies and flips concurrently, then releases hot
copies in chunks of at most 200, each chunk manifest-first, with the
manager pacing proposals from every source at least a second apart so
each node's bounded unlink queue drains between them. A manifest failure
(after four attempts on leader-election blips) keeps the remaining hot
copies; the rows already say cold and orphan reconciliation — now also
manifest-first, with the edge-sync receipts hook batched ahead — finishes
them within its window. A sweep on the primary removes entries for files
that were already in cold before this change, once a row has settled for
an hour and the cold object exists with the manifest's recorded size, so
upgraded clusters stop re-pulling replicas and restarts catch up. A copy
whose tier flip fails no longer deletes the cold object it wrote.

Replicating per-node clusters are now gated like shared storage: only the
primary writer migrates and every node syncs cold rows. Every such node
must run tiering with the same cold backend — the manifest delete unlinks
every replica, and a node reads the file only through its cold row — and
Arc warns at startup on a replicating node with no usable cold tier.

Live on a 3-writer + reader per-node rig with replication: compaction
registered a daily file that replicated to the reader; the primary's next
tick migrated it, the FSM removed the manifest entry
(reason=tiering:migrated), and the reader's delete worker unlinked its
replica within the same second; no re-pull attempts followed, and the
reader learned the cold row on its next sync.

Docs: Basekick-Labs/docs.basekick.net#85
xe-nvdk added a commit to Basekick-Labs/arc that referenced this pull request Sep 28, 2026
…replicating per-node clusters (#953)

* fix(tiering): keep the cluster manifest in step with migration; gate replicating per-node clusters

Tiering never told the cluster file manifest when it moved a file to cold.
On a per-node-storage cluster with file replication the puller pulled the
migrated file straight back from a peer, every node re-did the same
migration, and once every peer had migrated a restarted node failed
startup catch-up on the entry (503 on reads with query_gate_on_catchup).
On a shared bucket stale entries accumulated unless the reconciler ran out
of dry-run mode.

A ManifestCoordinator seam in tiering (adapter in cmd/arc, mirroring
retention's JSON DeleteFilePayload batches; no cluster/raft import in
tiering) lets migration remove a file from the manifest before its hot
copy goes: MigrateBatch copies and flips concurrently, then releases hot
copies in chunks of at most 200, each chunk manifest-first, with the
manager pacing proposals from every source at least a second apart so
each node's bounded unlink queue drains between them. A manifest failure
(after four attempts on leader-election blips) keeps the remaining hot
copies; the rows already say cold and orphan reconciliation — now also
manifest-first, with the edge-sync receipts hook batched ahead — finishes
them within its window. A sweep on the primary removes entries for files
that were already in cold before this change, once a row has settled for
an hour and the cold object exists with the manifest's recorded size, so
upgraded clusters stop re-pulling replicas and restarts catch up. A copy
whose tier flip fails no longer deletes the cold object it wrote.

Replicating per-node clusters are now gated like shared storage: only the
primary writer migrates and every node syncs cold rows. Every such node
must run tiering with the same cold backend — the manifest delete unlinks
every replica, and a node reads the file only through its cold row — and
Arc warns at startup on a replicating node with no usable cold tier.

Live on a 3-writer + reader per-node rig with replication: compaction
registered a daily file that replicated to the reader; the primary's next
tick migrated it, the FSM removed the manifest entry
(reason=tiering:migrated), and the reader's delete worker unlinked its
replica within the same second; no re-pull attempts followed, and the
reader learned the cold row on its next sync.

Docs: Basekick-Labs/docs.basekick.net#85

* docs(release-notes): link the tiering manifest entry to #953
@xe-nvdk
xe-nvdk deleted the docs/tiering-manifest-integration branch October 1, 2026 19:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant