Repository navigation
aw-sync: duplicate folders for one device_id silently truncate history on pull #683
Description
Activity
Related, from the same investigation:
- aw-sync:
daemon(the default subcommand) never pulls — two incompatible sync-folder layouts #682 —aw-sync daemonnever pulls (two incompatible sync-folder layouts); this is why the duplicate folders went unnoticed - aw-sync: sync failures are silent — add
aw-sync status, loud diagnostics, per-device manifest, and surface peers in Raw Data #684 — sync failures are silent;aw-sync status, loud diagnostics, per-device manifest - Sanitized device hostname forks the sync identity and never migrates bucket hostnames aw-android#272 — the aw-android change that produced the duplicate folders in practice
- aw-sync:
Took suggested fix 1 (dedupe by
device_idbefore import) plus the loud-warning half of 3.PR: #686
pull_allnow groups{hostname}/{device_id}/*.dbby device_id and keeps the largest file, with awarn!naming the skipped folders. That is the POCO F8 Ultra / poco_f8_ultra case.find_remotes_nonlocaldoes the same collapse so a 3-level walker (aw-sync:daemon(the default subcommand) never pulls — two incompatible sync-folder layouts #682) does not reintroduce the truncation on the daemon path.- If a resume-from-newest pull reports up-to-date but the source still has more events than the destination, it now warns instead of only logging
✓ Already up to date!. Recovering an already-truncated dest bucket still requires deleting it and re-pulling — this PR does not backfill.
Left out: making
device_idthe provenance key (suggested fix 2 / #649), and automatic backfill of older-than-resume events.#686 merged (
876709e). Verified on current master, not from the PR head.Hazard (suggested fix 1) is gone:
pull_alluseslist_remote_dbs+select_remote_dbs_by_device_id— largest file perdevice_id, order-independent. ThePOCO F8 Ultra/vspoco_f8_ultra/pair no longer both import.find_remotes_nonlocalapplies the same collapse, so the advanced/daemon path cannot reintroduce it.- Resume-from-newest now warns when the source still has more events than dest, instead of only logging
✓ Already up to date!. list_remote_dbsis 3-level-only on purpose (the aw-sync:daemon(the default subcommand) never pulls — two incompatible sync-folder layouts #682 root orphan is not a pull candidate). Deadget_remotestrio is gone.
Not done here, and not v0.14.0 work:
- Suggested fix 2 (
device_idas provenance key) moved to the v2(device_id, name)identity sequenced on Sync: tracking issue and v0.14.0 release triage activitywatch#1445 / Change how hostname/devices work activitywatch#302. - Suggested fix 3 auto-backfill. Detection is in; recovery is still delete the dest bucket and re-pull. On this setup that is a 273 MB / ~1M event re-import.
The silent-truncation bug this issue named is fixed. Remaining items have other homes. I cannot close it (pull-only) — maintainer close when this reads as done.
🤖 Claude, on behalf of @0xbrayo
The silent-truncation hazard is fixed on master (#686: per-
device_idselection inpull_allandfind_remotes_nonlocal, plus a warning when a resume pull leaves the source bigger than the destination). The two remaining suggestions already have homes:- fix 2 (
device_idas the provenance key) → the(device_id, name)identity work on Change how hostname/devices work activitywatch#302 - fix 3 (don't let the resume boundary regress / backfill) → the v2 import, which tracks progress per source by source event id (Sync: tracking issue and v0.14.0 release triage activitywatch#1445), not from the destination's newest event
Suggest closing this as fixed, with those two carrying the rest.
- fix 2 (
Fix merged in #686 (skip duplicate device_id folders). Closing.
When one physical device has written under two different folder names in the sync directory,
pull_allimports both. Because provenance is derived from the bucket'shostnamefield rather than the folder name, both folders resolve to the same destination bucket — and the resume logic then silently discards most of the history.Whether you lose data depends on
fs::read_dirordering, which is arbitrary.Mechanism
sync_one()(aw-sync/src/sync.rs) picks its resume point from the newest event already in the destination:and then only fetches events newer than that from the source. So if a folder holding a short, recent slice of history is imported before the folder holding the full history, the resume boundary jumps forward and everything older in the good database is never fetched. Re-running sync does not recover it — the boundary is now baked into the destination bucket. Only deleting the destination bucket and re-pulling recovers the data.
sync_wrapper::pull()already guards against this within a single folder:but there is no equivalent check across folders, and nothing anywhere keys on
device_id.Real-world instance
One Android device, two folders, same
device_id(41662faa-…), identicalhostnameon every bucket row inside both databases ("POCO F8 Ultra"):aw-watcher-androideventsPOCO F8 Ultra/poco_f8_ultra/Both import into
aw-watcher-android-synced-from-POCO F8 Ultra. If the 9.4 MB folder is walked first, the resume boundary becomes 2026-07-18 and the entire 2021–2026 history is skipped — ~1M events, silently, with✓ Already up to date!in the log.aw-watcher-android-unlock(98,035 vs 100,188 events) truncates the same way.The two folders exist because aw-android started sanitizing the device hostname (see the companion aw-android issue), but this is not Android-specific: any hostname change, re-install, or manual folder rename reproduces it.
Suggested fix
device_idbefore importing. Group the discovered remotes by thedevice_idpath component; if a device_id appears under more than one folder, use one (largest, or newestlast_updated) andwarn!naming the folders that were skipped. This alone removes the hazard and would have caught the case above automatically.device_idthe provenance key, nothostname.get_or_create_sync_bucket()currently derives origin from$aw.sync.origin→ falling back tobucket_from.hostname; neither is stable across a rename, and both are decoupled from the folder the data actually came from.Discovered alongside the daemon layout bug filed separately; both were hit in the same setup.
cc @TimeToBuildBob