Skip to content

Edge sync (spoke→hub replication): phase 1 manual core #569

Description

@xe-nvdk

Tracking issue for edge-to-cloud sync — promised for 26.09.1, currently design-only.

Design: docs/progress/2026-06-04-edge-sync-architecture-converged.md (converged from two independent designs; read that one, it supersedes the two siblings). Re-validated 2026-08-06: all eight §14 reuse targets still exist and the cited primitives are current.

Why

Arc runs at the edge today — a standalone binary with local storage in a rocket, a tractor, a factory cell, a forward operating base — with no first-class way to get data out to a central Arc. Backup is a DR snapshot, not incremental sync. Import re-ingests rows through the Arrow buffer and writes new files with new checksums, which breaks end-to-end integrity and double-counts on retry, so it can't be the receive door.

The field differs from the datacenter in one decisive way: connectivity is the exception, not the rule. The transport must treat "disconnected" as steady state and "connected" as the event.

Shape

Sync immutable parquet files, not rows. Spoke initiates (edges are behind NAT). At-least-once delivery + idempotent receive = exactly-once effect.

Two phases with opposite round-trip economics — the key insight of the merged design:

  • Discovery is batched — one POST /api/v1/sync/reconcile returns missing/present/conflicts for the whole backlog. After a long dark period with 5,000 pending files that is one request, not 5,000.
  • Transfer is per-file — each file its own request so X-Arc-Sync-Offset byte-resume is mechanically real. A contact window will close mid-file, so resume is the common case.

Licensing

Per enterprise product definition line 27: manual export/import is OSS; the automatic scheduled agent is Enterprise. Phase 1 below is entirely OSS. The auto-agent gate and receive-endpoint registration are independent — a box can be an OSS manual spoke, an OSS manual hub, an Enterprise auto spoke, or an Enterprise auto hub.

Phase 1 breakdown (small PRs)

Each row is one PR. Ordered by dependency; each should be independently reviewable and mergeable.

  • 1. Sync ledger (SQLite) — sync_ledger + sync_history tables in the shared DB, keyed UNIQUE(hub_id, path) for multi-hub readiness. bytes_sent on the row as the resume checkpoint. Modeled on internal/tiering/metadata.go. Includes the in_flight → pending revert on startup. Subject to the SQLite Review Checklist — batched deletes, no app mutex across DB I/O, v.SetDefault for every new key.
  • 2. SyncTransport interface + local/no-op impl — Reconcile(ctx, hubID, pending) (*ReconcileResult, error) and PutFile(ctx, hubID, f, body, offset) error. Defining this cleanly now is what keeps the S3-relay and sneakernet transports open later without touching the agent or ledger.
  • 3. ComputeSyncHMAC + validation — binds spoke_id, hub_id, path, content-hash, timestamp. NUL-delimited, modeled directly on ComputeFetchHMAC (internal/cluster/security/auth.go:123) which already binds a path into the signed payload and has a matching constant-time validator with freshness tolerance.
  • 4. Hub receive: POST /api/v1/sync/file — stream body → WriteReader to a temp path, SHA256 tee'd as it streams, atomic-rename into the namespaced final path only on match, then register in the manifest. Verify-before-commit. filereplication/fetch_client.go:190-215 already does exactly this including resume via a caller-supplied prefix hasher.
  • 5. Hub receive: POST /api/v1/sync/reconcile — batch delta from the manifest, O(N) lookups, no I/O on parquet bytes. Must stream both request and response — 100k pending entries is ~20MB and neither side may buffer it whole.
  • 6. Spoke namespacing + 409 conflict — hub rewrites incoming paths to {spoke_id}/{original}; spoke_id is bound into the HMAC so a spoke cannot write into another's namespace. Same path + different SHA is a 409, never an overwrite — this deliberately overrides Arc's internal last-write-wins (raft/fsm.go), which is right for intra-cluster replication and wrong for cross-edge fan-in.
  • 7. Hub registry — sync_spokes table, per-spoke shared secret, POST/GET /api/v1/sync/spokes. Revoking one spoke must not re-key the fleet.
  • 8. One-shot sync trigger + status/ledger endpoints — the manual OSS trigger. No scheduler. Shipped in feat(edgesync): spoke agent and manual sync pass (#569) #579 as POST /api/v1/spoke-sync/run (plus /status and /ledger), not /api/v1/sync/run as planned: Fiber's Group().Use() matches by string prefix, not path segment, so an operator group under /api/v1/sync inherited the hub group's body limit and middleware. A sibling prefix removes the coupling; TestEdgeSyncSpoke_PrefixDoesNotCollideWithHubGroups pins the property.
  • 9. Air-gap export/import bundle — shipped across four PRs: feat(edgesync): add StateExported to the spoke ledger (#569) #580 (ledger StateExported), feat(edgesync): air-gap bundle export (#569) #581 (export), feat(edgesync): hub-side air-gap bundle import (#569) #582 (import + (spoke_id, bundle_id) dedup + Raft batching at 1000 ops), feat(edgesync): bundle acknowledgment, closing the air-gap loop (#569) #583 (acknowledgment). Endpoints are POST /api/v1/spoke-sync/export, POST /api/v1/bundle-import, and POST /api/v1/spoke-sync/ack — not /api/v1/sync/imports as planned: Fiber's Group().Use() matches by string prefix, so anything under /api/v1/sync inherits the spoke-facing HMAC group's middleware and body limit. Bundles are a signed directory (manifest + entries.jsonl + data/), chosen over an archive because resume is free and the contents are auditable with ls and sha256sum.

Non-negotiables

  • Every endpoint gets auth middleware. Receive, reconcile, and the HEAD fallback all carry HMAC validation → license → handler. Even reconcile leaks "what data exists on the hub" if unauthenticated.
  • The hub must not assume clusterCoordinator != nil. "The manifest" is Raft-backed in cluster mode and a local store standalone. This is the classic independently-enabled-subsystem trap.
  • Never delete edge data fire-and-forget. Post-sync retention is phase 3, off by default, and only after a verified acknowledged round-trip.

Open design question (decide before PR 1)

Row-group granularity. Compacted files can be large, and compaction output is bounded by file count not bytes (#568), so a daily partition can produce a multi-hundred-MB file. Over a contested link a whole-file unit means nothing is queryable until the entire file lands, and newest-first ordering loses meaning within one big file.

Parquet is internally chunked at 122,880 rows per row group (set explicitly in the compaction COPY), so row groups are a natural transfer unit that costs nothing — original compressed bytes, no decode, checksum chain intact per group, and each group carries its own statistics for time-based prioritization. This would change the ledger schema and the reconcile protocol, so it should be settled before PR 1, not retrofitted.

Rejected alternative: extracting rows and shipping msgpack. That inflates bandwidth 3–5× (row-major, self-describing per record, no cross-row compression) on the resource the edge has least of, costs edge CPU, and breaks the end-to-end checksum by forcing re-ingestion on the hub.

Later phases

Phase Scope License
2 Connectivity-adaptive agent, FeatureEdgeSyncAuto gate, cluster-manifest receive, bandwidth cap, 429 backpressure, multi-hub Enterprise
3 S3/Azure relay transport, delete_after_sync mixed
4 Metrics, per-spoke dashboard, sync-lag alerting, hub-as-relay Enterprise

🤖 Generated with Claude Code

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions