Repository navigation
Blank white screen and app not working #261
Description
Activity
Root cause found and reproduced; fix is up at ActivityWatch/aw-server-rust#679.
What happens. v0.14.0 is the first build that runs the
aw-watcher-android-test_<host>→aw-watcher-android_<host>bucket migration (#244, merge logic from aw-server-rust#661). Anyone who ran an older release (wrote to the-testbucket) and then v0.14.0b2 (created the new bucket) has both, so the merge path runs on first start. That merge used a correlatedNOT EXISTSoverlap subquery per legacy event with no lower bound onstarttime, so it's O(n²) over the whole history. It runs inside the single datastore worker thread, so while it's running:- every web UI API call queues behind it →
GET /api/0/settings/never answers → blank white WebView (static assets load fine, which is why the app shell still works) - every main-thread JNI datastore call queues too → widget refresh,
WebWatcher, heartbeat alarm → "ActivityWatch isn't responding"
That matches the widget observation exactly: the widget queries the datastore directly and works until the app is opened, which starts
BackgroundServiceand the migration. And because one overlapping cutover heartbeat leaves the merge "partial", the app re-runs the whole thing on every start.Evidence.
- The shipped SQL on synthetic disjoint data: 5k events/bucket 0.65 s, 10k 2.6 s, 20k 10.4 s, 40k 41 s on a desktop CPU. Quadruples per doubling; a phone with a couple of years of history is in the hours range.
- Play vitals: the top ANR clusters on the current build are the main thread parked in
aw_datastore::worker::Datastore::get_buckets→crossbeam recv, from the widget refresh broadcast,WebWatcher,SystemJobService, and theLOG_DATAalarm receiver. - Reproduced on an Android 16 emulator with the v0.14.0 release APK and a 150k+15k event fixture: migration logged at service start, first API request never answered,
/api/0/infotimes out, WebView white. Still wedged after 5 minutes.
Fix. One sorted scan over both buckets plus an endtime min-heap sweep (O(n log n)), batched UPDATEs, identical overlap semantics. 100k+10k events merge in well under a second in a debug build. Includes a wall-clock-bounded regression test with a six-figure fixture.
Ship path. Merge aw-server-rust#679 → I bump the submodule in aw-android → v0.14.1. Once the app has the fixed migration it finishes in seconds on the next start, so affected users recover without any data action. I'll verify the CI-built APK against the same fixture on the emulator before the release.
Why CI didn't catch it. The merge has six unit tests, all with 1–3 events; nothing in either repo runs a migration against a realistically sized database, and the aw-android instrumented tests start from a fresh install so the legacy-bucket path is never exercised. Follow-ups for that in a separate comment once the fix is out.
- every web UI API call queues behind it →
Fix verified end to end on an Android 16 emulator: the CI-built library from ActivityWatch/aw-server-rust#679 packaged into a v0.14.0 debug APK, same 150k+15k event fixture that leaves the shipped v0.14.0 stuck indefinitely. Migration now completes in 2.3 s (unoptimized build) and the server answers within 5 s of launch. Details in the PR comment.
Remaining ship steps: merge aw-server-rust#679 → submodule bump PR in this repo (I'll open it the moment #679 lands) → v0.14.1. Affected users recover automatically on the next start after updating; no data action needed (the partial-merge leftover is two overlapping cutover events, which stay in the legacy bucket by design).
Why CI didn't catch this, and what stops a recurrence
Why nothing caught it
- The merge had tests, but only for correctness. aw-server-rust#661 shipped with six unit tests, each with 1–3 events. A correlated subquery is invisible at that size; it only hurts at 10⁵ events. Nothing in either repo runs any migration against a realistically sized database.
- The instrumented tests start from a fresh install. aw-android's emulator E2E (
ScreenshotTest,BasicTest) has no legacyaw-watcher-android-test_*bucket and no history, so the exact branch that hurts (both buckets exist, large history) is never executed.WatcherAndroidBucketMigrationTestonly tests result-string parsing. - Review looked at semantics, not scaling. #661 was reviewed as "merge disjoint events on name collision". The SQL reads fine; nobody asked how it behaves at n=300k. Greptile passed it too.
- "Small bump over the beta" was measured by intent, not by diff. v0.14.0b2 → v0.14.0 was 12 aw-android PRs plus a submodule bump of 25 aw-server-rust commits (~3k lines). The migration wiring (fix(android): migrate aw-watcher-android-test buckets on startup #244) was merged after b2, so no beta build ever contained it, and the final build was never installed on a device with real history before promotion to production.
Prevention, concrete
- Scale tests for migrations (done in #679): 100k-event and fully-overlapping fixtures with a wall-clock bound. Any future datastore migration or startup pass gets the same treatment; a correlated subquery over
eventswithout a bounded range should not pass review. - Upgrade-with-history instrumented test (next PR here): seed a two-bucket, six-figure-event
sqlite.dbintofilesDirbefore the app launches and assert/api/0/infoanswers within ~15 s and the migration log completes. That exercises the real startup path on the real library, which is the only thing that would have caught this class. - Beta soak rule: every commit that is not already in the last beta build goes through the internal/beta track on a device with real history (Erik's phone is the best upgrade test we have) before promotion. "Small bump" is decided by
git diff --stat <beta>..<release>including the submodule range, not by how it feels. - Post-release vitals watch (Bob side): the ANR clusters for build 41 were visible in Play vitals within hours. I'll add a release-watch that queries crash/ANR clusters for the new versionCode daily for the first three days after a release and flags new clusters, so the next regression is caught before a user report.
- Graceful degradation (fix(android): harden startup paths that turned a slow datastore into a hang #262): the app-side changes so a busy datastore is a slow screen, not a white screen plus ANRs.
- added a commit that references this issue
on Sep 14, 2026 - added a commit that references this issue
on Sep 14, 2026 Fix merged — v0.14.1 release pending
Both fixes have landed:
- fix(datastore): make legacy Android bucket merge linear instead of quadratic aw-server-rust#679 (O(n log n) migration fix) — merged 2026-09-14
- build(deps): bump aw-server-rust for the linear legacy-bucket merge (v0.14.1) #263 (submodule bump for v0.14.1) — merged 2026-09-14
The fix is in `master`; v0.14.1 needs to be tagged to reach affected users. Keeping this open until the release is published. No data action required for users once they update — the fixed migration finishes in seconds on next launch and the partial-merge sentinel is cleared.
What to do now: Erik tags v0.14.1 from current `master`. @ErikBjare
- added a commit that references this issue
on Sep 14, 2026 v0.14.1 is live. Play production now serves versionCode 42 (100% rollout), GitHub release https://github.com/ActivityWatch/aw-android/releases/tag/v0.14.1 has the apk and aab, and the tagged commit pins aw-server-rust at 5e67ac8 (the #679 fix). Affected devices recover on the first start after updating: the migration that used to run for hours finishes in seconds, no data action needed.
The tag build's E2E job went red on
NativeWindowInsetsTest.syncToggleReceivesRealTap("Sync switch must be visible"); that test passed on the identical code in #263 minutes earlier, so it is a flaky UI assertion, unrelated to this fix, and it does not gate publishing.Leaving this open until it is confirmed fixed on a real upgraded device (Erik's phone is the definitive check). Follow-ups: #262 (hardening), #264 (upgrade-with-history CI test), #265 (release discipline: device gate + staged rollout), aw-server-rust#680.
Correction to the above: "live" was too strong. The publish succeeded and the Play API shows the release configured on the production track at 100%, but Play holds production changes for its pre-checks and review before they reach users, so v0.14.1 lands on phones once Google's review clears. The "Change app icon" entry that shows up under Changes in review on every release is the publish step re-uploading the listing icon each time; #265 now skips listing images on tag publishes.
Now! See #267 (comment)
Confirmed. Play production is versionCode 42 (v0.14.1) at 100% — Google review has cleared, so this is actually reaching users (yesterday's publish was API-configured only).
Your phone on v0.14.2b1 running smoothly is the real-device check this was waiting on.
v0.14.0 installs recover on first start after updating; no data action needed. Remaining work stays on #267 (crash-rate as the rollout lands), #262, #264, and #265.
I cannot close issues in this repo (
closeIssueis not granted). @ErikBjare please close.- added a commit that references this issue
on Oct 8, 2026 - added a commit that references this issue
on Oct 10, 2026
Side-menu is working, but "Home" there gets me to "Welcome, early user!" but clicking on "Activity" is a blank web view. Intermittently getting "ActivityWatch not responding" popups from OS.