Skip to content

backup (stage B): named remote targets (local, S3, Azure) with per-database routing #1085

Description

@xe-nvdk

⚠️ Internal work only. Community contributions are not being accepted for this issue.

This is part of a four-stage piece of planned work on backups (per-database scoping, remote targets, cluster safety, cold tier) that Basekick Labs is implementing internally. Please do not open PRs against it; they will be closed without review. Discussion and questions in the comments are welcome.

Stages: 0 cluster safety #1083 → A database scoping #1084 → B remote targets #1085 → C cold tier #1086. Order is fixed; A and B can ship to OSS before 0 only if the docs say "not for clusters yet", and the node gate from 0 lands before any of them is advertised to a cluster. Origin: a prospect asked for a different backup target per database (their audit database to its own bucket). For an audit database, long retention and mostly cold, the answer is complete only with C.

Summary

The only backup target today is backup.local_path, a directory on the server itself; NewManager builds one storage.NewLocalBackend from it (internal/backup/manager.go:78). Everything else in the package talks to the storage.Backend interface, so a remote target is a factory and config change. This stage adds named targets (local, S3, Azure) with per-database routing, the second half of the prospect's ask.

Config shape

[backup]
enabled = true
local_path = "./data/backups"          # stays: the default target

[backup.targets.audit-bucket]
type = "s3"                             # local | s3 | azure
bucket = "acme-arc-audit-backups"
region = "eu-west-1"
prefix = "arc/"                         # optional key prefix inside the bucket
databases = ["audit"]                   # routing: these databases land here

Changes

  • One target is one storage.Backend. The S3 and Azure field sets are tiered_storage.cold's; the inline factory in cmd/arc/main.go (~3800-3845) is extracted into a shared constructor used by primary storage, tiering and backup, including the credential refresher wiring (IRSA: DuckDB query path fails with ExpiredToken ~1h after start (credential_chain never refreshes) #600/Extend #600 credential refresher to all expiring credential sources (instance roles, Pod Identity) via a CanExpire gate #601) and Close on shutdown. backup.targets.<name>.* is a map: specify how ARC_BACKUP_TARGETS_<NAME>_… binds, since viper's AutomaticEnv does not discover map keys absent from the file.
  • Routing validated at start: a database in two targets is a startup error; a database in no target goes to the default. backup.default_target may name a remote target, which means NewManager stops requiring a local path, the Iceberg warehouse containment check (warehouse.go:60-63) runs per local target, and ListBackups lists under a backup- prefix rather than the whole target (manager.go:160 is List(ctx, ""), O(all objects) on S3).
  • Overlap refusal at a boundary, both directions, local paths included. A target overlapping primary storage, the cold tier or another target is a startup error when (endpoint, bucket) are equal and one prefix equals the other or is a parent of it after trailing-slash normalisation; for local paths reuse pathWithin (warehouse.go:100-106). Same bucket with disjoint prefixes is legitimate. Apply it to today's local_path too: ./data/arc and ./data/backups are one typo apart, and a backup directory inside the storage root is both swept by reconciliation (looksLikeManagedPath, internal/reconciliation/sweep_storage.go:297-305) and backed up by the next backup. Matrix rows for prefix spelled "", "arc", "arc/" (the fix(iceberg): warehouseRelKey must trim the storage root, not the warehouse #534 corollary).
  • One backup ID per run (backup-<ts>-<8hex>); each target receives its routed files and its own <id>/manifest.json; the default target also holds an index naming every target touched. Metadata and config go to the default target; include_config defaults to false when any target is remote, since arc.toml carries the targets' own credentials (backup.go:850-856 already calls backup storage sensitive). Two instances sharing a bucket and prefix would merge listings and could restore each other's backups: the manifest records instance identity and listing filters on it.
  • GET /api/v1/backup merges across targets by ID and marks an unreachable target; DELETE removes the ID everywhere it landed; status and Progress gain per-target fields; the 2 h operation timeout (backup_routes.go:85,258) becomes configurable.
  • Restore reads each file from the target its manifest names. No one-off "restore from another bucket" override.
  • The A long but legal source key makes every backup fail permanently, because the destination key overruns the limit #761 key-length reservation and max_source_key_bytes are per target (prefix lengths differ).
  • Cold rows mark the backup INCOMPLETE. Cold objects live in a separate backend and tiering deletes a migrated file's manifest entry (cmd/arc/main.go:4533-4558), so a backup cannot see them. For each scoped database one indexed query (idx_tier_files_database_tier) counts cold rows; cold_files_excluded goes into the manifest and the backup is INCOMPLETE when it is non-zero. The release note says cold objects are excluded until the cold-tier stage, but the marker is the mechanism: it makes the gap visible at backup time instead of restore time.
  • OSS, not enterprise: the S3 and Azure backends already ship as OSS primary storage. Update the enterprise product definition at implementation time.

Matrix rows

Configuration Expected
Standalone, backup.targets.x s3 databases=["audit"], default local audit→s3, rest→local; two manifests + index
Standalone with tiering, audit db partly cold INCOMPLETE with cold_files_excluded
backup.default_target remote, local_path unset allowed; backup- prefixed listing; Iceberg containment per local target
Target prefix "" / "arc" / "arc/" overlap test and #761 reservation at each spelling
Target inside the storage root (local) or same bucket + parent prefix refused at start
Same bucket, disjoint prefix allowed
Target unreachable mid-run run fails; partial manifest absent; index not written
Two instances, one bucket + prefix listings filtered by instance identity
Pattern 1 / Pattern 2 the node gate from the cluster-safety stage applies; targets are per instance, not per node

Acceptance

Regression tests for routing validation, overlap refusal at every spelling, per-target manifests and merged listing, deletion across targets, the cold-row marker. A binary run with a target configured at a non-default value (a real S3 or SeaweedFS bucket) and a database routed to it. Release note in RELEASE_NOTES_2027.01.1.md; docs page gains a "Targets" section; arcli shows the target per backup.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority: mediumCorrectness or operability gap with a workaround

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions