Found in the 2026-08-19 internal audit of edge sync (#569, shipped in #576–#584). Highest-priority open finding after the blocker-fix PR.
Problem
On a spoke with compaction enabled (the default), raw Parquet files sync to the hub, then compaction rewrites them into a _compacted.parquet and deletes the sources locally. Nothing propagates spoke-local deletes — the hub keeps the raw files forever. The compacted file, containing the same rows, is then discovered as a new path and synced too. Hub queries over that partition now double-count every row.
This is the steady state for any spoke that syncs more often than its compaction interval — the more diligent the operator, the worse the duplication. The design doc (docs/progress/2026-06-04-edge-sync-architecture-converged.md §13) calls this "harmless (amplification only)"; for query results that is wrong.
Current mitigation (26.09.1)
Release notes and the docs guide now carry a loud warning: disable compaction on databases a spoke syncs, or account for duplicates in hub queries.
Proper fix
Hub-side supersede: when a compacted file arrives whose partition covers already-received raw files from the same spoke, mark/remove the raws (or exclude them from query resolution). Needs design — interacts with the hub index, the cluster manifest, and HubIndex.Forget (see the separate Forget-gate issue).
🤖 Generated with Claude Code
Found in the 2026-08-19 internal audit of edge sync (#569, shipped in #576–#584). Highest-priority open finding after the blocker-fix PR.
Problem
On a spoke with compaction enabled (the default), raw Parquet files sync to the hub, then compaction rewrites them into a
_compacted.parquetand deletes the sources locally. Nothing propagates spoke-local deletes — the hub keeps the raw files forever. The compacted file, containing the same rows, is then discovered as a new path and synced too. Hub queries over that partition now double-count every row.This is the steady state for any spoke that syncs more often than its compaction interval — the more diligent the operator, the worse the duplication. The design doc (
docs/progress/2026-06-04-edge-sync-architecture-converged.md§13) calls this "harmless (amplification only)"; for query results that is wrong.Current mitigation (26.09.1)
Release notes and the docs guide now carry a loud warning: disable compaction on databases a spoke syncs, or account for duplicates in hub queries.
Proper fix
Hub-side supersede: when a compacted file arrives whose partition covers already-received raw files from the same spoke, mark/remove the raws (or exclude them from query resolution). Needs design — interacts with the hub index, the cluster manifest, and
HubIndex.Forget(see the separate Forget-gate issue).🤖 Generated with Claude Code