Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
81 changes: 81 additions & 0 deletions RELEASE_NOTES_2027.01.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,87 @@

> **Status:** Planned — January 2027 release.

## New features

### Backups can be scoped to one or more databases ([#1084](https://github.com/Basekick-Labs/arc/issues/1084))

`POST /api/v1/backup` was whole-instance: its body took only `include_metadata`
and `include_config`. It now also takes `databases`:

```json
{"databases": ["audit"]}
```

A scoped backup copies those databases and nothing else: their data files,
their schema anchors under `_schema/` and their compaction state under
`_compaction_state/` (parked `.quarantined` manifests included). The manifest
records the list in `scope`, the backup list and the status endpoint show it,
and the 202 echoes `databases`. `backup_type` stays `full`. An empty or absent
list is the whole-instance backup, unchanged.

Rules a scoped backup applies:

- `include_metadata` and `include_config` default to **false**. An explicit
`include_metadata: true` is refused with 400: the SQLite database holds every
database's tier rows, the tokens, the continuous queries and the audit log,
so it cannot ride along with one database. `include_config: true` is allowed.
- Every name must pass the storage-segment rule (the one that names an existing
database; no `..`, no separators, at most 256 names, no duplicates), and must
be a database this node knows: its hot prefix has a file, or `_schema/<db>/`
has an anchor, or the tier metadata has rows for it. The third rule is what
accepts a fully cold database (an audit database with a long retention, the
case that started this work) instead of answering 400. Its backup completes
with zero data files: the cold tier is not copied yet (#1086), and the
INCOMPLETE marker for that arrives with the remote-targets stage (#1085).
Unknown names answer 400 naming them. The probes are bounded (one indexed
query, then a listing that stops at the first object the storage would
return), so a large database does not hold the request. A database whose only files have keys no listing
can return is still recognised, by a fourth check that runs only after the
three above said no, so the run fails with the rename advice rather than
calling the database unknown.
- The databases' Iceberg namespace directories (`<prefix>_<db>.db/`) are
excluded and counted (`iceberg_namespace_files_excluded`,
`iceberg_namespaces_excluded` on the manifest): the Iceberg SQL catalog is
instance-wide and rides with the metadata a scoped backup refuses, so
restoring those files would land tables no catalog can resolve.
- Scope is the storage-root segment. Edge-sync spoke data lives under
`<spoke>/<db>/…`, so `databases: ["prod"]` does not include `spoke1/prod`
and `databases: ["spoke1"]` takes the whole spoke, anchors and compaction
state included. A spoke namespace below the root segment cannot be named.

Restoring a scoped backup restores only its databases, in either mode. On a
cluster node, a restore of a scoped backup that does not name a `mode` runs in
`replace` — the mode a per-database backup is for (an additive restore of an
audit database resurrects everything retention removed since) — and the 202
echoes the effective mode. An explicit `"mode": "merge"` is honoured; a
standalone node stays additive, as before; an unscoped backup still defaults
to `merge`. In replace mode the current files to remove are selected by the
storage path's first segment when the backup is scoped (a spoke file's
manifest entry carries the canonical database, not the spoke). Replace is
refused, with 400, when the scoped backup holds no data files for one of its
databases: a fully cold database, or one dropped and re-created since, has
nothing in the backup to put back, so replace would only remove its current
files. The message points at `merge`, or at a fresh backup once the database
has hot files again. The cluster-wide compaction pause (#1087) is taken for a
scoped cluster restore exactly as for any other.

A request body that is not JSON (curl's `-d` default is form-encoded) now
answers 400 instead of being read as an empty request: for a scoping feature
the worst outcome is a malformed `databases` silently becoming a
whole-instance backup. An empty body still means the defaults.

Storage backends gained a bounded existence probe (`PrefixProber`) for the
known-database check; the local, S3 and Azure backends implement it and a
backend without it falls back to a listing.

A backup is a copy of the files on disk at the moment each is read: rows
still in the ingest buffers are not in it. "Point in time" means the last
flush.

Not in this stage: a different backup target per database (#1085), the cold
tier (#1086), `arcli backup create --database` (tracked in Basekick-Labs/arcli#44),
and scoping below the storage-root segment.

## Bug fixes

### Backup and restore are cluster-safe ([#1083](https://github.com/Basekick-Labs/arc/issues/1083))
Expand Down
4 changes: 4 additions & 0 deletions cmd/arc/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -4212,12 +4212,16 @@ func main() {
}
if tieringManager != nil {
backupManager.SetTierRecorder(tieringManager)
// The known-database check of a scoped backup (#1084) asks
// the same tier metadata whether a database is fully cold.
backupManager.SetTierLookup(tieringManager)
}
backupHandler.RegisterRoutes(server.GetApp())
log.Info().
Str("backup_path", cfg.Backup.LocalPath).
Bool("cluster_gate", clusterCoordinator != nil).
Bool("tier_recorder", tieringManager != nil).
Bool("tier_lookup", tieringManager != nil).
Msg("Backup/restore enabled")
}
}
Expand Down
200 changes: 185 additions & 15 deletions internal/api/backup_routes.go
Original file line number Diff line number Diff line change
Expand Up @@ -5,11 +5,13 @@ import (
"errors"
"fmt"
"regexp"
"sort"
"sync/atomic"
"time"

"github.com/basekick-labs/arc/internal/auth"
"github.com/basekick-labs/arc/internal/backup"
"github.com/basekick-labs/arc/internal/storage"
"github.com/gofiber/fiber/v2"
"github.com/rs/zerolog"
)
Expand Down Expand Up @@ -98,16 +100,83 @@ func (h *BackupHandler) RegisterRoutes(app fiber.Router) {

// CreateBackupRequest is the request body for POST /api/v1/backup.
type CreateBackupRequest struct {
IncludeMetadata *bool `json:"include_metadata"` // default: true
IncludeConfig *bool `json:"include_config"` // default: true
IncludeMetadata *bool `json:"include_metadata"` // default: true; false and refused when scoped
IncludeConfig *bool `json:"include_config"` // default: true; false when scoped
// Databases scopes the backup to these databases (#1084): storage-root
// segments, so an edge-sync spoke is named as the spoke. Empty or absent
// is a whole-instance backup. At most maxScopeDatabases names, each a
// safe storage path segment that names an existing database.
Databases []string `json:"databases"`
}

// maxScopeDatabases bounds how many databases one backup may be scoped to:
// each is checked synchronously, with bounded work, inside this request.
const maxScopeDatabases = 256

// scopedMetadataRefusal is the 400 for include_metadata: true on a scoped
// backup. No apostrophe on purpose: the log masker truncates error text at
// one.
const scopedMetadataRefusal = "include_metadata is not available on a scoped backup: the SQLite database holds the tier rows of every database, the tokens, the continuous queries and the audit log, so it cannot ride along with one database; take an unscoped backup for it"

// normalizeBackupScope validates the databases of a scoped backup request and
// returns them sorted and de-duplicated, or an error naming the offending
// value. Names are matched exactly (storage segments are case-sensitive).
// Each must pass isSafeStoragePathSegment, the rule for a name that NAMES
// something existing in the storage root (not the create-time rule, see its
// doc comment), because the manager turns each into a ListObjects prefix and
// a PrefixProber prefix; and none may be a reserved root (_schema,
// _compaction_state), which hold Arc's own state and not a database, and
// which the hot-prefix probe would otherwise call known.
func normalizeBackupScope(names []string) ([]string, error) {
if len(names) == 0 {
return nil, nil
}
if len(names) > maxScopeDatabases {
return nil, fmt.Errorf("databases lists %d names; a backup can be scoped to at most %d", len(names), maxScopeDatabases)
}
seen := make(map[string]struct{}, len(names))
out := make([]string, 0, len(names))
for _, name := range names {
if name == "" {
return nil, errors.New("databases contains an empty name")
}
if !isSafeStoragePathSegment(name) {
return nil, fmt.Errorf("databases contains %q, which is not a valid database name: a name is one storage path segment (no separators, no backslash, no NUL, not . or .., not dot-prefixed, at most %d bytes)", name, storage.MaxUsableKeySegmentLen)
}
if storage.IsReservedRootDir(name) {
return nil, fmt.Errorf("databases contains %q, which is a reserved storage root, not a database", name)
}
if _, dup := seen[name]; dup {
return nil, fmt.Errorf("databases lists %q more than once", name)
}
seen[name] = struct{}{}
out = append(out, name)
}
sort.Strings(out)
return out, nil
}

// CreateBackup triggers a new backup.
// POST /api/v1/backup
func (h *BackupHandler) CreateBackup(c *fiber.Ctx) error {
var req CreateBackupRequest
if err := c.BodyParser(&req); err != nil {
// Empty body is fine — use defaults
// An empty body means the defaults. A non-empty body must be JSON and
// must parse: for a scoping feature the worst outcome is a `databases`
// that is silently dropped and becomes a whole-instance backup, which is
// exactly what a form-typed body does (curl -d defaults to
// application/x-www-form-urlencoded, and Fiber form-parses that into an
// empty request with every unknown key ignored).
if len(c.Body()) > 0 {
if !c.Is("json") {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": "Invalid request body: send JSON (Content-Type: application/json)",
})
}
if err := c.BodyParser(&req); err != nil {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": "Invalid request body",
})
}
}

// Node gate first (#1083): a node that may not run the backup does no
Expand All @@ -116,16 +185,58 @@ func (h *BackupHandler) CreateBackup(c *fiber.Ctx) error {
return nil
}

// Scope (#1084): validated before any work and before the slot is taken.
databases, err := normalizeBackupScope(req.Databases)
if err != nil {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": err.Error(),
})
}
scoped := len(databases) > 0

// Defaults: a whole-instance backup carries the SQLite metadata and the
// config; a scoped one carries neither unless asked, and metadata cannot
// be asked for (it is every database's state, not one database's).
opts := backup.BackupOptions{
IncludeMetadata: true,
IncludeConfig: true,
IncludeMetadata: !scoped,
IncludeConfig: !scoped,
Databases: databases,
}
if req.IncludeMetadata != nil {
opts.IncludeMetadata = *req.IncludeMetadata
}
if req.IncludeConfig != nil {
opts.IncludeConfig = *req.IncludeConfig
}
if scoped && opts.IncludeMetadata {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": scopedMetadataRefusal,
})
}

// Known-database check, synchronous and bounded per name (one indexed
// tier-metadata query, then two prefix probes, and only when all three
// say no an enumeration of the hidden keys under <name>/, which then
// holds nothing listable; never a listing of a real database), so a typo
// is a 400 now rather than a failed run later. A probe that cannot be
// answered is a 500, not a guess. See Manager.CheckDatabasesKnown.
if scoped {
ctx, cancel := context.WithTimeout(c.Context(), 30*time.Second)
defer cancel()
if err := h.manager.CheckDatabasesKnown(ctx, databases); err != nil {
var unknown *backup.UnknownDatabasesError
if errors.As(err, &unknown) {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": err.Error(),
"unknown_databases": unknown.Names,
})
}
h.logger.Error().Err(err).Strs("databases", databases).Msg("Could not check whether the scoped databases exist")
return c.Status(fiber.StatusInternalServerError).JSON(fiber.Map{
"error": "Could not check whether the requested databases exist: " + err.Error(),
})
}
}

acquired, err := h.acquireOperation(c, "backup")
if !acquired {
Expand All @@ -144,10 +255,14 @@ func (h *BackupHandler) CreateBackup(c *fiber.Ctx) error {
}()

// Return immediately — client polls /status for progress
return c.Status(fiber.StatusAccepted).JSON(fiber.Map{
resp := fiber.Map{
"message": "Backup started",
"status": "running",
})
}
if scoped {
resp["databases"] = databases
}
return c.Status(fiber.StatusAccepted).JSON(resp)
}

// ListBackups returns all available backups.
Expand Down Expand Up @@ -261,9 +376,13 @@ type RestoreRequest struct {
RestoreMetadata *bool `json:"restore_metadata"` // default: true standalone, false on a cluster node (where true is refused)
RestoreConfig *bool `json:"restore_config"` // default: false (refused on a cluster node)
Confirm bool `json:"confirm"` // must be true
// Mode is "merge" (default: additive, resurrects files deleted since the
// backup) or "replace" (cluster nodes only: the current files of each
// restored database are removed through the cluster manifest first).
// Mode is "merge" (additive, resurrects files deleted since the backup)
// or "replace" (cluster nodes only: the current files of each restored
// database are removed through the cluster manifest first). Absent, it
// is merge, except that a scoped backup (#1084) restored on a cluster
// node defaults to replace; the response echoes the effective mode.
// Replace of a scoped backup that holds no data file for one of its
// databases is refused (400), implicit or explicit: it would only delete.
Mode string `json:"mode"`
}

Expand Down Expand Up @@ -338,12 +457,58 @@ func (h *BackupHandler) RestoreBackup(c *fiber.Ctx) error {
"error": "restore_config is not available on a cluster node: arc.toml holds this node identity (cluster.node_id, role, seeds, raft_bootstrap, shared secret), and a config taken on another node would boot this one as that node",
})
}
if !clustered && mode == backup.RestoreModeReplace {
// Replace is keyed on the Raft manifest being wired, which is what the
// manager runs from, not on the coordinator: a cluster node without
// cluster.raft_data_dir has the coordinator and not the manifest, and the
// manager would refuse the mode anyway.
if !h.manager.ClusterManifestWired() && mode == backup.RestoreModeReplace {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": "mode \"replace\" is only available on a cluster node, where the current files are removed through the cluster manifest; a standalone restore is additive (mode \"merge\")",
})
}

// Effective mode (#1084): a request that named none restores a scoped
// backup in replace mode on a cluster node. Resolved here from the
// manifest's scope and from whether the manager has the Raft manifest
// wired, with the same function the manager applies after it reads the
// manifest, so the echo is what will run. A manifest that cannot be read
// here (unknown id, or a transient backup-storage error) resolves as
// unscoped, is echoed that way and is admitted as before; the manager
// then reads the manifest itself and fails the run with "backup not
// found", or runs from what it read.
var manifest *backup.Manifest
{
ctx, cancel := context.WithTimeout(c.Context(), 30*time.Second)
mf, err := h.manager.GetBackup(ctx, req.BackupID)
cancel()
if err != nil {
h.logger.Debug().Err(err).Str("backup_id", req.BackupID).Msg("Could not read the backup manifest before the restore; resolving the mode as for an unscoped backup")
} else {
manifest = mf
}
}
var scope []string
if manifest != nil {
scope = manifest.Scope
}
mode, err = backup.ResolveRestoreMode(req.Mode, scope, h.manager.ClusterManifestWired())
if err != nil {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": err.Error(),
})
}
// A replace of a scoped backup that holds no data file for one of its
// databases would only delete (#1084); refused here with the manager's
// text, whether the mode was defaulted or named.
if manifest != nil {
if err := backup.CheckScopedReplace(mode, manifest); err != nil {
return c.Status(fiber.StatusBadRequest).JSON(fiber.Map{
"error": err.Error(),
})
}
}
opts.Mode = mode

acquired, err := h.acquireOperation(c, "restore")
if !acquired {
return err
Expand All @@ -367,11 +532,16 @@ func (h *BackupHandler) RestoreBackup(c *fiber.Ctx) error {
}
// Restored databases are STAGED and applied at the next boot (#635), and
// a restored config only takes effect on reload — both need a server
// restart to take effect.
if opts.RestoreMetadata || opts.RestoreConfig {
// restart to take effect. Only when the backup holds them: a scoped
// backup never carries the metadata, so a default standalone restore of
// one stages nothing. When the manifest could not be read the flags
// follow the request, as before.
stagesMetadata := opts.RestoreMetadata && (manifest == nil || manifest.HasMetadata)
restoresConfig := opts.RestoreConfig && (manifest == nil || manifest.HasConfig)
if stagesMetadata || restoresConfig {
resp["restart_required"] = true
}
if opts.RestoreMetadata {
if stagesMetadata {
resp["staged"] = true
}
return c.Status(fiber.StatusAccepted).JSON(resp)
Expand Down
Loading
Loading