Skip to content

aw-sync: unbounded event fetch in sync_one() causes native OOM crash on Android #630

Description

@TimeToBuildBob

Summary

Android sync (syncBoth) OOM-crashes the native aw-server-rust process because the event-fetch side of sync_one() is unbounded, while the write side is already batched.

Confirmed as the dominant crash cluster in ActivityWatch/aw-android#176 — SIGABRT / std::sys::unix::abort in libaw_server.so, ~22 reports in current vitals, and very likely the root cause behind other unattributed native crashes given how sync-heavy the Android app is.

How sync works today (for context — there's no user-facing setting or trigger)

  • SyncScheduler.kt starts a Handler loop ~1 min after BackgroundService boots, then re-fires syncBothAsync() every 15 min (mobile/.../SyncScheduler.kt:15,38,49,95-106), with an AlarmManager fallback (SyncAlarmReceiver.kt:26) in case the process/Handler dies.
  • It is fully automatic — no on-demand button, no foreground/background trigger, no settings screen entry (confirmed no sync references in AWPreferences.kt / AuthSettingsActivity.kt).
  • syncBothAsync() → JNI syncBoth(port, hostname) (SyncInterface.kt:59,86-99, single-flight guarded) → Java_..._syncBoth in aw-server-rust/aw-sync/src/android.rs:162-212 → pull_all() / push_with_hostname() (sync_wrapper.rs) → sync_run() → sync_datastores() → sync_one() in aw-server-rust/aw-sync/src/sync.rs.

Root cause

sync_one() fetches events with get_events(bucket_id, resume_sync_at, None, None) — limit: None (sync.rs:321-333). Both Datastore::get_events (local) and AwClient::get_events (HTTP) treat None as "return everything," so the entire event range since the last sync point loads into one Vec<Event> in memory, unbounded. There's already a TODO acknowledging this at sync.rs:323:

// TODO: Fetch at most ~5,000 events at a time (or so, to avoid timeout from huge buckets)

Only the write side is batched — BATCH_SIZE: usize = 5000 (sync.rs:351) is used when inserting into ds_to. The read/fetch side is not. A bucket with a large backlog (device offline a while, or a high-frequency AFK/window bucket) bulk-loads unbounded events → native OOM abort, which surfaces on Android as a libaw_server.so crash since aw-sync shares the process/address space with aw-server.

Fix

  • Primary: implement the TODO — page sync_one()'s fetch loop in chunks (e.g. 5,000 events, matching the existing write BATCH_SIZE) instead of one unbounded get_events() call.
  • Short-term Android mitigation (if the primary fix takes longer): cap sync lookback window (e.g. last N days) via a JNI param, to bound worst-case allocation on mobile specifically.

Related

Activity

  1. ErikBjare commented on Jul 12, 2026

    @ErikBjare
    Member

    @TimeToBuildBob Your "How sync works today" ends with syncBothAsync() → JNI syncBoth(port, hostname) (SyncInterface.kt:59,86-99, single-flight guarded) → Java_..._syncBoth in aw-server-rust/aw-sync/src/android.rs:162-212 → pull_all() / push_with_hostname() (sync_wrapper.rs) → sync_run() → sync_datastores() → sync_one() in aw-server-rust/aw-sync/src/sync.rs but doesn't explain how it actually syncs with other devices. On desktop aw-sync created a ~/ActivityWatchSync folder (by default) which users are asked to share across devices however they like (Syncthing/Gdrive/Dropbox, rsync, network mount, etc) since each device only syncs their buckets to a dbfile they own in the sync area. But how does this work on Android? (what I was wondering the whole time) What flow does the user go through to get their data synced between their phone/aw-android app and their desktop app devices running the aw-sync daemon.

    Syncing generally needs a UX pass (beyond the OOM-cause in this issue), right now it's rather difficult/confusing/non-intuitive to setup aw-sync. Worth a separate UX tracking issue in ActivityWatch/activitywatch. Syncing should probably be disabled by default on Android until this is working (enable via setting somehow?).

  2. TimeToBuildBob commented on Jul 12, 2026

    @TimeToBuildBob
    ContributorAuthor

    Seen. Investigating now; I'll reply here with the concrete work, task, or handoff once it's done.

  3. TimeToBuildBob commented on Jul 12, 2026

    @TimeToBuildBob
    ContributorAuthor

    Android sync directory and flow

    Dug through the code. Here's what actually happens on Android:

    SyncInterface.kt sets AW_SYNC_DIR to context.getExternalFilesDir(null)/sync (i.e., Android/data/net.activitywatch.android/files/sync/), which dirs.rs::get_sync_dir() picks up. The directory structure mirrors desktop: {hostname}/{device_id}/*.db.

    • Push (push_with_hostname): works in isolation — Android writes its own events to sync_dir/{device_name}/{device_id}/*.db. This is where the OOM in this issue hits (unbounded fetch from local aw-server, large bucket backlog → all events in one Vec).
    • Pull (pull_all): scans sync_dir for {other_hostname}/{device_id}/*.db files from remote devices. There's nothing to pull on a typical Android device. On desktop, users share ~/ActivityWatchSync via Syncthing/Dropbox/rsync. On Android, the app's data dir is scoped storage (since Android 11+) — other apps including Syncthing can't write into it — so the desktop sync flow doesn't translate. pull_all() silently no-ops, and the pushed data stays trapped on the phone.

    In short: there is no working end-to-end sync flow for Android users today. Push accumulates data in a place nothing else can read, and pull finds nothing to pull.

    Tracking

    Opened ActivityWatch/activitywatch#1357 to track the UX side (default-off setting, storage access, setup flow).

    The Rust OOM fix (chunked fetch in sync_one() to match the existing BATCH_SIZE = 5000 on the write side) is on my task list and is a straightforward change — will open a PR once I have a worktree set up. That's the primary crash fix regardless of the larger sync UX situation.

  4. TimeToBuildBob commented on Jul 12, 2026

    @TimeToBuildBob
    ContributorAuthor

    PR filed: #631

    The fix replaces the unbounded get_events(limit=None) with a BATCH_SIZE=5000 page loop. Events come back descending (newest first) so pages are fetched oldest→newest using the end parameter in reverse, then processed oldest-page-first to keep the heartbeat merge semantics at the resume boundary.

    All 4 existing sync tests pass + 1 new test for the resume path.

  5. added a commit that references this issue on Jul 12, 2026
    d87bd93
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions