Skip to content

Edge sync: spoke compaction duplicates rows in hub queries — needs hub-side supersede #610

Description

@xe-nvdk

Found in the 2026-08-19 internal audit of edge sync (#569, shipped in #576–#584). Highest-priority open finding after the blocker-fix PR.

Problem

On a spoke with compaction enabled (the default), raw Parquet files sync to the hub, then compaction rewrites them into a _compacted.parquet and deletes the sources locally. Nothing propagates spoke-local deletes — the hub keeps the raw files forever. The compacted file, containing the same rows, is then discovered as a new path and synced too. Hub queries over that partition now double-count every row.

This is the steady state for any spoke that syncs more often than its compaction interval — the more diligent the operator, the worse the duplication. The design doc (docs/progress/2026-06-04-edge-sync-architecture-converged.md §13) calls this "harmless (amplification only)"; for query results that is wrong.

Current mitigation (26.09.1)

Release notes and the docs guide now carry a loud warning: disable compaction on databases a spoke syncs, or account for duplicates in hub queries.

Proper fix

Hub-side supersede: when a compacted file arrives whose partition covers already-received raw files from the same spoke, mark/remove the raws (or exclude them from query resolution). Needs design — interacts with the hub index, the cluster manifest, and HubIndex.Forget (see the separate Forget-gate issue).

🤖 Generated with Claude Code

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions