Skip to content

Play: user-perceived crash rate 3.64% exceeds the 1.09% bad-behavior threshold #267

Description

@TimeToBuildBob

Play Console flags the app under Technical quality → Bad behavior: user-perceived crash rate 3.64% against Google's 1.09% threshold, which reduces discoverability. Per-version crash clusters from the Play Developer Reporting API, last 7 days (2026-09-07 → 09-14):

versionCode 40 (v0.14.0b2, the bulk of installs)

reports users cluster status
532 91 BackgroundService.onCreate → ForegroundServiceStartNotAllowedException fixed by #257, shipped in v0.14.1
9 6 AlarmReceiver.onReceive → ForegroundServiceStartNotAllowedException fixed by #257, shipped in v0.14.1
57 + 26 1 + 3 libaw_sync.so Rust panic → SIGABRT open, #220
31 + 6 1 Java_net_activitywatch_android_… JNI entry in libaw_server.so → SIGABRT open (a Rust panic unwinding across the extern "C" JNI boundary aborts the process)
22 3 libaw_server.so → SIGABRT open, same class
16 8 AccessibilityNodeTraversal.findNode → StackOverflowError open (recursive traversal, no depth bound)

versionCode 41 (v0.14.0): no crash clusters; that release's problem was ANRs (#261).

versionCode 33 (v0.12.1, 2023): 324 reports / 61 users on ChromeWatcher.onAccessibilityEvent NPE plus 8 on ChromeWatcher.onInterrupt. The class no longer exists; only the update to 0.14.x fixes those users.

What this means for the metric: the largest cluster by far (541 reports, 97 users) is already fixed in v0.14.1, so the rate should drop substantially as it rolls out. The v0.12.1 stragglers will keep a floor until they update. The remaining actionable work:

  1. Native panics must not abort the process: wrap every JNI entry point in libaw_sync and libaw_server with catch_unwind and return an error object (a Rust panic reaching extern "C" is an abort since Rust 1.81). Fixing the underlying #220 panic is the second half.
  2. findNode / forEachNode: iterative traversal with a depth and node budget instead of unbounded recursion.
  3. Re-check this metric two weeks after v0.14.1 is at 100%.

Activity

  1. ErikBjare commented on Sep 15, 2026

    @ErikBjare
    Member

    v0.14.1 just went live on prod track in Play Store, will be interesting to see how it does!

    I am personally running the v0.14.2b1 on my device, which is running smoothly!

  2. TimeToBuildBob commented on Sep 15, 2026

    @TimeToBuildBob
    ContributorAuthor

    Confirmed: production is versionCode 42 (v0.14.1) at 100%. Internal is 43 (v0.14.2b1). Glad 43 is smooth on-device — that build already has #269 (bounded findNode); prod 42 does not.

    Fresh Play Reporting snapshot (2026-09-15, API latest complete day is 2026-09-13 so today's rollout is not in the series yet):

    • 7-day userPerceivedCrashRate avg: 1.37% (was 3.64% when this issue opened). Still above the 1.09% bad-behavior line.
    • 28-day simple avg of the same daily series is still ~3.2% — mid-August days were 7–14%. If Console's headline is the longer window, it will not clear until those days roll off (~end of September), even if recent days stay low.
    • Last complete day (2026-09-13): 0%.

    7-day crash clusters, same ranking as the opener, counts still leftover-dominated:

    reports cluster status
    426 BackgroundService.onCreate FGS fix in 42; draining from vc40
    288 ChromeWatcher.onAccessibilityEvent NPE v0.12.1 only; updates are the fix
    ~130 libaw_sync / libaw_server SIGABRT ActivityWatch/aw-server-rust#681 (catch_unwind on JNI) is merged but not in the aw-android submodule — v0.14.1 and v0.14.2b1 both still pin 5e67ac8 (#679); #681 is 2 commits ahead
    15 findNode StackOverflowError fixed in 43 via #269, not in prod 42

    So the rate drop so far is mix-shift + FGS drain, not the remaining native/traversal work. Next Android cut should bump aw-server-rust past ActivityWatch/aw-server-rust#681. Rechecking this metric in two weeks, once 42 has a full vitals window.

  3. TimeToBuildBob commented on Sep 29, 2026

    @TimeToBuildBob
    ContributorAuthor

    Two-week readout from the latest successful Play collector snapshot (collected 2026-09-28; daily rate series complete through 2026-09-26):

    • 7-day userPerceivedCrashRate simple avg (Sep 20–26): 0.786% — down from 1.37% on Sep 15 and now below the 1.09% bad-behavior threshold by 0.30pp.
    • 28-day simple avg (Aug 30–Sep 26): 0.945% — down from ~3.2% and also below the threshold by 0.15pp.
    • Latest complete day (Sep 26): 0.93%.

    The app-wide threshold problem has therefore cleared in both windows. The comparison is no longer a clean versionCode 42-only readout: v0.14.2 (versionCode 46) started a 50% production rollout on Sep 25, so the last two days are a 42/46 mix.

    Latest seven-day crash clusters (Sep 21 16:00 → Sep 28 16:00 UTC):

    reports sample version cluster readout
    200 33 ChromeWatcher.onAccessibilityEvent NPE legacy v0.12.1 tail
    82 46 CategoryTimeWidgetProvider.onUpdate IllegalStateException new current-rollout cluster; needs separate triage
    31 40 BackgroundService.onCreate FGS old build draining; fix is in 42+
    15 42 findNode StackOverflowError fixed in 43+ by #269
    6 42 libaw_server SIGABRT (two grouped issues) old native boundary remains visible in 42

    I am leaving this open through the v0.14.2 rollout rather than calling the app healthy from the aggregate alone: the original discovery penalty is gone, but the new versionCode 46 widget cluster is material.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions