Skip to content

docs(release-notes): describe both edge-sync transports in the intro - #584

Merged
xe-nvdk merged 1 commit into
mainfrom
docs/release-notes-edge-sync
Aug 9, 2026
Merged

xe-nvdk merged 1 commit into
mainfrom
docs/release-notes-edge-sync

Conversation

@xe-nvdk

@xe-nvdk xe-nvdk commented Aug 8, 2026

Copy link
Copy Markdown
Member

Summary

The edge-sync section of the 26.09.1 notes was written when only the hub receive endpoint existed, and the intro never caught up with what shipped across #570–#576 and #580–#583.

Two drifts fixed:

  • "the spoke initiates" described only the network path. The air-gap transport initiates nothing over a network — an operator carries a drive. The intro now names both transports and says what each is for, since "the edge" is not one thing and a reader choosing between them needs that up front.
  • "The hub receive endpoint is the first piece an operator can turn on" was true when the rest was unshipped. Everything ships now, so it reads as "start with the hub receive endpoint".

Also removes a duplicated "this is the manual form, and it is OSS" sentence — the new intro states it once for both transports.

Verified against the code, not by reading

  • Every documented endpoint (/api/v1/spoke-sync/{run,status,ledger,export,ack}, /api/v1/bundle-import, /api/v1/sync/{file,reconcile}, /api/v1/sync-spokes) matches a registered route.
  • Every edge_sync.* key mentioned has a v.SetDefault in setDefaults.
  • Documented default values match the code: false, empty list, 10000, and 68719476736 = int64(64)<<30.
  • No stale "not yet" / "lands next" language remains except the Enterprise scheduled agent, which is genuinely future work.

The four edge-sync subsections (spoke side, air-gap bundles, import, acknowledgment) were already present and accurate — only the intro and one duplicated sentence needed changing.

Related

Basekick-Labs/docs.basekick.net#24 got the matching audit: all endpoints and keys check out, and Current limitations is renamed Network-transport limitations with a pointer to the air-gap block, since it sits mid-page before the air-gap sections and all five of its entries describe the network path.

🤖 Generated with Claude Code

The edge-sync intro was written when only the hub receive endpoint
existed, and never caught up with what shipped.

Two drifts:

- "the spoke initiates" described only the network path. The air-gap
  transport initiates nothing over a network — an operator carries a
  drive. The intro now names both transports and says what each is for,
  since "the edge" is not one thing and a reader choosing between them
  needs that up front.

- "The hub receive endpoint is the first piece an operator can turn on"
  was true when the rest was unshipped. Everything ships now, so it reads
  as "start with the hub receive endpoint".

Also removes a "this is the manual form, and it is OSS" sentence that the
new intro states once for both transports, rather than repeating it per
section.

Verified against the code rather than by reading: every documented
endpoint matches a registered route, every edge_sync config key has a
v.SetDefault, and the documented default values match (false, empty,
10000, 68719476736 = 64 GiB).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@xe-nvdk
xe-nvdk merged commit 8d100b9 into main Aug 9, 2026
4 checks passed
xe-nvdk added a commit that referenced this pull request Aug 19, 2026
Internal audit of the shipped edge sync subsystem (#576-#584) found four
release blockers; all fixed and live-verified on a two-node rig.

- Auth (B1): the spoke transport never sent an Arc API token while the hub
  mounted admin-level token auth on /api/v1/sync — every request against an
  auth-enabled hub (the default) died with 401. The hub group now gates at
  write level (matching ingest), and the spoke carries a new env-only
  credential ARC_EDGE_SYNC_HUB_TOKEN (same config-file-refusal discipline as
  the spoke secret). Tokenless spokes warn at startup, and the 401/403
  remediation text names the variable.

- Quoted identifiers (B2, general query-layer bug): MaskStringLiterals
  masked double-quoted tokens as string literals, so quoted database and
  measurement names resolved to quote-polluted read_parquet globs and
  returned zero rows — and hyphenated names (every spoke ID) had no working
  syntax at all. Double-quoted tokens now mask into a distinct, deduplicated
  __IDENT__ placeholder class; the table rewriters, the header-db path, and
  RBAC extraction all resolve placeholders to validated unquoted names, so
  extraction matches execution. Deep review of this fix caught that leaving
  an INVALID quoted identifier unrewritten would hand DuckDB a replacement
  scan (net-new cross-tenant read with RBAC off): ValidateSQLRequest now
  rejects invalid quoted identifiers in table position (generalizing the
  GHSA-w8x2 scanner, which also closes the pre-existing double-quoted
  comma-cross-join bypass), and the transform resolves invalid names to an
  inert sentinel as a backstop.

- Reconcile paging (B3): the default batch_size=0 offered the whole backlog
  in one reconcile; a backlog above the hub's cap failed every pass with the
  same 413 forever. batch_size now defaults to 1000, and a refused page is
  split and retried in-pass using the hub's advertised cap (halving when the
  refusal is the byte limit), so no configuration can strand a backlog.

- Vanished files (B4): compaction deleting a discovered-but-unsent file
  permanently wedged air-gap export (the entry was re-selected every time
  with no operator escape) and burned the network retry budget. StateSkipped
  is now wired: the exporter pre-checks existence and skips vanished entries
  (Exists must positively report the file gone — transient errors keep
  today's abort), the network agent skips instead of failing, and skipped
  rows are counted in /status and pruned.

- Dead code wired: SweepStaging (hub staging DoS guard) now runs hourly
  (edge_sync.staging_sweep_max_age_hours, default 72); PruneSynced — also
  never called — and the new PruneSkipped run twice daily on the spoke
  (edge_sync.spoke.ledger_retention_days, default 90).

- Compaction row-duplication on the hub (audit High) is documented with
  loud warnings in the release notes and docs; the hub-side supersede fix
  is tracked in #610. Remaining audit findings filed as #611-#616.

Live gates on a real two-node rig, all previously-broken cells: sync
through an auth-enabled cap-3 hub drains 8/8 at batch_size=0; quoted
hyphenated queries return real rows on the hub; a discovered-then-deleted
file lands skipped; the tokenless 401 names its remedy; export with a
vanished file produces a valid bundle and completes import + ack; the
replacement-scan shapes are rejected over HTTP with explicit errors.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant