You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Edge sync: hub-side compaction of spoke namespaces #619
Follow-up to #610/#618. With defer-until-synced shipped, the hub receives every spoke row exactly once — as raw files, forever. Hub compaction never touches them: spoke data lives at {spoke_id}/{db}/{meas}/{y}/{m}/{d}/{h}/…, one level deeper than the {db}/{meas}/… layout the tiers' partition scanners expect, so candidate discovery finds zero partitions under a spoke namespace (observed live: partition_count: 0 for telemetry/engine_temp under rocket-01). Hub queries over spoke data pay a growing small-files cost.
Sketch
Namespace-aware candidate discovery: compaction learns the registered spoke IDs (from the spoke registry) and scans each {spoke}/{db} as a pseudo-database, so everything downstream (measurement scan, partitions, jobs, manifests) operates on the standard two-level model with a prefixed database name.
The receipt-index interplay — this is the Edge sync: HubIndex.Forget MUST be wired before delete_after_sync or hub-side spoke-namespace retention ships #611 gate firing: hub compaction would DELETE received raws under spoke namespaces. The sync_received index must not "forget" them (confirmPresent checks backend existence and forgets missing files — a spoke whose pruned-then-rediscovered raw is re-offered would then re-upload it, reintroducing the duplicate next to the compacted file). Receipts for compacted-away inputs must be marked (content still delivered, file merged) and treated as present without an existence check; re-delivery of the same path+sha stays a no-op. Requires the compacted-INPUT list to reach the parent process (currently only the output key crosses the subprocess IPC).
Cluster mode: compacted outputs/source deletes flow through the existing Phase-4 completion-manifest machinery (path-based) — verify with the namespace prefix.
Treat #611 (HubIndex.Forget gate) as a blocker on this work per that issue.
Follow-up to #610/#618. With defer-until-synced shipped, the hub receives every spoke row exactly once — as raw files, forever. Hub compaction never touches them: spoke data lives at
{spoke_id}/{db}/{meas}/{y}/{m}/{d}/{h}/…, one level deeper than the{db}/{meas}/…layout the tiers' partition scanners expect, so candidate discovery finds zero partitions under a spoke namespace (observed live:partition_count: 0fortelemetry/engine_tempunderrocket-01). Hub queries over spoke data pay a growing small-files cost.Sketch
{spoke}/{db}as a pseudo-database, so everything downstream (measurement scan, partitions, jobs, manifests) operates on the standard two-level model with a prefixed database name.sync_receivedindex must not "forget" them (confirmPresentchecks backend existence and forgets missing files — a spoke whose pruned-then-rediscovered raw is re-offered would then re-upload it, reintroducing the duplicate next to the compacted file). Receipts for compacted-away inputs must be marked (content still delivered, file merged) and treated aspresentwithout an existence check; re-delivery of the same path+sha stays a no-op. Requires the compacted-INPUT list to reach the parent process (currently only the output key crosses the subprocess IPC).Treat #611 (
HubIndex.Forgetgate) as a blocker on this work per that issue.🤖 Generated with Claude Code