Repository navigation
aw-sync: unbounded event fetch in sync_one() causes native OOM crash on Android #630
Description
Activity
@TimeToBuildBob Your "How sync works today" ends with
syncBothAsync() → JNI syncBoth(port, hostname) (SyncInterface.kt:59,86-99, single-flight guarded) → Java_..._syncBoth in aw-server-rust/aw-sync/src/android.rs:162-212 → pull_all() / push_with_hostname() (sync_wrapper.rs) → sync_run() → sync_datastores() → sync_one() in aw-server-rust/aw-sync/src/sync.rsbut doesn't explain how it actually syncs with other devices. On desktop aw-sync created a ~/ActivityWatchSync folder (by default) which users are asked to share across devices however they like (Syncthing/Gdrive/Dropbox, rsync, network mount, etc) since each device only syncs their buckets to a dbfile they own in the sync area. But how does this work on Android? (what I was wondering the whole time) What flow does the user go through to get their data synced between their phone/aw-android app and their desktop app devices running the aw-sync daemon.Syncing generally needs a UX pass (beyond the OOM-cause in this issue), right now it's rather difficult/confusing/non-intuitive to setup aw-sync. Worth a separate UX tracking issue in ActivityWatch/activitywatch. Syncing should probably be disabled by default on Android until this is working (enable via setting somehow?).
Seen. Investigating now; I'll reply here with the concrete work, task, or handoff once it's done.
Android sync directory and flow
Dug through the code. Here's what actually happens on Android:
SyncInterface.ktsetsAW_SYNC_DIRtocontext.getExternalFilesDir(null)/sync(i.e.,Android/data/net.activitywatch.android/files/sync/), whichdirs.rs::get_sync_dir()picks up. The directory structure mirrors desktop:{hostname}/{device_id}/*.db.- Push (
push_with_hostname): works in isolation — Android writes its own events tosync_dir/{device_name}/{device_id}/*.db. This is where the OOM in this issue hits (unbounded fetch from local aw-server, large bucket backlog → all events in one Vec). - Pull (
pull_all): scanssync_dirfor{other_hostname}/{device_id}/*.dbfiles from remote devices. There's nothing to pull on a typical Android device. On desktop, users share~/ActivityWatchSyncvia Syncthing/Dropbox/rsync. On Android, the app's data dir is scoped storage (since Android 11+) — other apps including Syncthing can't write into it — so the desktop sync flow doesn't translate.pull_all()silently no-ops, and the pushed data stays trapped on the phone.
In short: there is no working end-to-end sync flow for Android users today. Push accumulates data in a place nothing else can read, and pull finds nothing to pull.
Tracking
Opened ActivityWatch/activitywatch#1357 to track the UX side (default-off setting, storage access, setup flow).
The Rust OOM fix (chunked fetch in
sync_one()to match the existingBATCH_SIZE = 5000on the write side) is on my task list and is a straightforward change — will open a PR once I have a worktree set up. That's the primary crash fix regardless of the larger sync UX situation.- Push (
PR filed: #631
The fix replaces the unbounded
get_events(limit=None)with a BATCH_SIZE=5000 page loop. Events come back descending (newest first) so pages are fetched oldest→newest using theendparameter in reverse, then processed oldest-page-first to keep the heartbeat merge semantics at the resume boundary.All 4 existing sync tests pass + 1 new test for the resume path.
- added a commit that references this issue
on Jul 12, 2026
Summary
Android sync (
syncBoth) OOM-crashes the nativeaw-server-rustprocess because the event-fetch side ofsync_one()is unbounded, while the write side is already batched.Confirmed as the dominant crash cluster in ActivityWatch/aw-android#176 —
SIGABRT/std::sys::unix::abortinlibaw_server.so, ~22 reports in current vitals, and very likely the root cause behind other unattributed native crashes given how sync-heavy the Android app is.How sync works today (for context — there's no user-facing setting or trigger)
SyncScheduler.ktstarts a Handler loop ~1 min afterBackgroundServiceboots, then re-firessyncBothAsync()every 15 min (mobile/.../SyncScheduler.kt:15,38,49,95-106), with anAlarmManagerfallback (SyncAlarmReceiver.kt:26) in case the process/Handler dies.syncreferences inAWPreferences.kt/AuthSettingsActivity.kt).syncBothAsync()→ JNIsyncBoth(port, hostname)(SyncInterface.kt:59,86-99, single-flight guarded) →Java_..._syncBothinaw-server-rust/aw-sync/src/android.rs:162-212→pull_all()/push_with_hostname()(sync_wrapper.rs) →sync_run()→sync_datastores()→sync_one()inaw-server-rust/aw-sync/src/sync.rs.Root cause
sync_one()fetches events withget_events(bucket_id, resume_sync_at, None, None)—limit: None(sync.rs:321-333). BothDatastore::get_events(local) andAwClient::get_events(HTTP) treatNoneas "return everything," so the entire event range since the last sync point loads into oneVec<Event>in memory, unbounded. There's already a TODO acknowledging this atsync.rs:323:// TODO: Fetch at most ~5,000 events at a time (or so, to avoid timeout from huge buckets)Only the write side is batched —
BATCH_SIZE: usize = 5000(sync.rs:351) is used when inserting intods_to. The read/fetch side is not. A bucket with a large backlog (device offline a while, or a high-frequency AFK/window bucket) bulk-loads unbounded events → native OOM abort, which surfaces on Android as alibaw_server.socrash since aw-sync shares the process/address space with aw-server.Fix
sync_one()'s fetch loop in chunks (e.g. 5,000 events, matching the existing writeBATCH_SIZE) instead of one unboundedget_events()call.Related