Skip to content

aw-sync: duplicate folders for one device_id silently truncate history on pull #683

Description

@ErikBjare

When one physical device has written under two different folder names in the sync directory, pull_all imports both. Because provenance is derived from the bucket's hostname field rather than the folder name, both folders resolve to the same destination bucket — and the resume logic then silently discards most of the history.

Whether you lose data depends on fs::read_dir ordering, which is arbitrary.

Mechanism

sync_one() (aw-sync/src/sync.rs) picks its resume point from the newest event already in the destination:

let resume_sync_at = most_recent_events
    .first()
    .map(|e| e.timestamp + e.duration)
    .or(sync_spec.start);

and then only fetches events newer than that from the source. So if a folder holding a short, recent slice of history is imported before the folder holding the full history, the resume boundary jumps forward and everything older in the good database is never fetched. Re-running sync does not recover it — the boundary is now baked into the destination bucket. Only deleting the destination bucket and re-pulling recovers the data.

sync_wrapper::pull() already guards against this within a single folder:

if dbs.len() > 1 {
    warn!("More than one db found in sync folder for host, choosing largest db {:?}", dbs);
}

but there is no equivalent check across folders, and nothing anywhere keys on device_id.

Real-world instance

One Android device, two folders, same device_id (41662faa-…), identical hostname on every bucket row inside both databases ("POCO F8 Ultra"):

folder size aw-watcher-android events range
POCO F8 Ultra/ 9.4 MB 4,068 2026-07-02 → 2026-07-18
poco_f8_ultra/ 273 MB 1,027,343 2021-05-18 → 2026-09-14

Both import into aw-watcher-android-synced-from-POCO F8 Ultra. If the 9.4 MB folder is walked first, the resume boundary becomes 2026-07-18 and the entire 2021–2026 history is skipped — ~1M events, silently, with ✓ Already up to date! in the log. aw-watcher-android-unlock (98,035 vs 100,188 events) truncates the same way.

The two folders exist because aw-android started sanitizing the device hostname (see the companion aw-android issue), but this is not Android-specific: any hostname change, re-install, or manual folder rename reproduces it.

Suggested fix

  1. Deduplicate by device_id before importing. Group the discovered remotes by the device_id path component; if a device_id appears under more than one folder, use one (largest, or newest last_updated) and warn! naming the folders that were skipped. This alone removes the hazard and would have caught the case above automatically.
  2. Make device_id the provenance key, not hostname. get_or_create_sync_bucket() currently derives origin from $aw.sync.origin → falling back to bucket_from.hostname; neither is stable across a rename, and both are decoupled from the folder the data actually came from.
  3. Refuse to regress the resume boundary. If a source database's newest event is older than the destination's, that source has nothing to contribute — but if its oldest event is older than the destination's oldest, it holds history the destination lacks. Backfilling that case (or at minimum warning loudly about it) would make the operation safe regardless of ordering.

Discovered alongside the daemon layout bug filed separately; both were hit in the same setup.

cc @TimeToBuildBob

Activity

  1. ErikBjare commented on Sep 15, 2026

    @ErikBjare
    MemberAuthor

    Related, from the same investigation:

  2. TimeToBuildBob commented on Sep 15, 2026

    @TimeToBuildBob
    Contributor

    Took suggested fix 1 (dedupe by device_id before import) plus the loud-warning half of 3.

    PR: #686

    • pull_all now groups {hostname}/{device_id}/*.db by device_id and keeps the largest file, with a warn! naming the skipped folders. That is the POCO F8 Ultra / poco_f8_ultra case.
    • find_remotes_nonlocal does the same collapse so a 3-level walker (aw-sync: daemon (the default subcommand) never pulls — two incompatible sync-folder layouts #682) does not reintroduce the truncation on the daemon path.
    • If a resume-from-newest pull reports up-to-date but the source still has more events than the destination, it now warns instead of only logging ✓ Already up to date!. Recovering an already-truncated dest bucket still requires deleting it and re-pulling — this PR does not backfill.

    Left out: making device_id the provenance key (suggested fix 2 / #649), and automatic backfill of older-than-resume events.

  3. TimeToBuildBob commented on Sep 16, 2026

    @TimeToBuildBob
    Contributor

    #686 merged (876709e). Verified on current master, not from the PR head.

    Hazard (suggested fix 1) is gone:

    • pull_all uses list_remote_dbs + select_remote_dbs_by_device_id — largest file per device_id, order-independent. The POCO F8 Ultra/ vs poco_f8_ultra/ pair no longer both import.
    • find_remotes_nonlocal applies the same collapse, so the advanced/daemon path cannot reintroduce it.
    • Resume-from-newest now warns when the source still has more events than dest, instead of only logging ✓ Already up to date!.
    • list_remote_dbs is 3-level-only on purpose (the aw-sync: daemon (the default subcommand) never pulls — two incompatible sync-folder layouts #682 root orphan is not a pull candidate). Dead get_remotes trio is gone.

    Not done here, and not v0.14.0 work:

    The silent-truncation bug this issue named is fixed. Remaining items have other homes. I cannot close it (pull-only) — maintainer close when this reads as done.

  4. 0xbrayo commented on Sep 24, 2026

    @0xbrayo
    Member

    🤖 Claude, on behalf of @0xbrayo

    The silent-truncation hazard is fixed on master (#686: per-device_id selection in pull_all and find_remotes_nonlocal, plus a warning when a resume pull leaves the source bigger than the destination). The two remaining suggestions already have homes:

    Suggest closing this as fixed, with those two carrying the rest.

  5. TimeToBuildBob commented on Oct 7, 2026

    @TimeToBuildBob
    Contributor

    Fix merged in #686 (skip duplicate device_id folders). Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions