Skip to content

bug(mobile): iPad thread open hangs on 'Syncing messages...' until force-quit #3

Description

@nohat

Problem

On iPad, tapping a thread sometimes leaves the detail pane empty with only the "Syncing messages..." pill above the composer. The UI is then effectively locked up until the app is force-quit. Reopening the same thread after a restart often reproduces it. Started around 2026-10-02 and has been frequent since. Seen on iPad only so far; web and desktop not checked.

What I know

  • The pill is the thread sync label (apps/mobile/src/features/threads/ThreadDetailScreen.tsx ~358-370). It shows when the thread state is empty, cached, or synchronizing while content presentation is ready. So the screen believes it has content, yet the message list renders blank.
  • The server does not look slow for this. In the server trace, the sql.execute spans under ws.rpc.orchestration.subscribeThread peak at 34 ms. Caveat: the trace rotates about every minute at 10 MB, so it only covers roughly the last 12 minutes and nothing from yesterday or from a failing open.
  • The affected thread (id 714471f7-67b1-4d97-80ae-53b9fdf511d6) is not extreme: 153 messages (about 100 KB of text) and 826 activities (about 2.2 MB of payload JSON). Other threads in the list are larger and I have not heard that they fail.
  • "Same thread fails again after restart" points at state that survives a restart: the persisted thread cache (cache.loadThread in packages/client-runtime/src/state/threads.ts ~207) or the server's snapshot for that thread, rather than a one-off connection hiccup.
  • Force-quit being the only exit suggests the JS thread is blocked or the thread state never leaves synchronizing. The spinner still animates, which fits a native animation on a blocked JS thread, but that is a guess.

Not yet known

  • Whether the JS thread is hung (decode or render of a large snapshot or cache entry) or the stream is just never completing.
  • Whether clearing the thread's cache on the iPad fixes it.
  • Which iPad build is installed, and whether the start date matches a build change.

Next steps

  1. On the iPad, reproduce with a thread that fails, then check whether another thread still opens without force-quitting. That separates a blocked JS thread from a stuck stream.
  2. Try the same thread from web or desktop on the same server. If it loads there, the problem is iPad-side cache or decode.
  3. If cache is implicated, add a size or time bound to cache restore and a timeout that falls back to a fresh fetch.
  4. Add a timeout on the sync pill so a sync that never completes shows an error with a retry, not an indefinite spinner.

This is the kind of report the papercut capture (docs/fork/papercuts.md) is meant to make cheap: build SHA, sync phase, and a client event ring buffer would answer the first two unknowns directly.

Related: #2 (client stuck states with no timeout).

Activity

  1. added
    bugSomething isn't working
    on Oct 3, 2026
  2. nohat commented on Oct 3, 2026

    @nohat
    OwnerAuthor

    Mitigation on branch fix/mobile-thread-sync-stall (commit 9ad46d7, not pushed): bounded cache restore plus a tap-to-retry pill after 20 s. It cannot help if the JS thread is blocked, which the report suggests (the whole UI is unresponsive). Finding the cause needs an instrumentable iPad simulator: #4.

  3. nohat commented on Oct 3, 2026

    @nohat
    OwnerAuthor

    Triage 2026-10-03T23:46Z: real symptom, cause still an open question. The server is not the bottleneck (subscribeThread sql.execute spans peak at 34 ms), so this is client-side: cache restore/decode or a stream that never completes. Cannot be root-caused from the Mac.

    Blocked on #4 — no way to see a blocked JS thread or read the iPad cache.

    Mitigation, not a fix: fix/mobile-thread-sync-stall (9ad46d7618, unmerged) skips a saved thread copy that takes over 5 s or exceeds 8 MB, and turns a 20 s sync into tap-to-retry. It does not establish a cause, so it would ship as 'mitigated, cause unknown'.

    Not a duplicate of upstream pingdotgg#13994: that is a SwiftUI/account device-collision after a host restart; different symptom and trigger.

  4. nohat commented on Oct 4, 2026

    @nohat
    OwnerAuthor

    2026-10-04: the bounded mitigation landed. fix/mobile-thread-sync-stall is merged to fork/prod (release 7345a87552): a saved thread copy over 5 s or 8 MB is skipped, and a 20 s sync becomes tap-to-retry. The cause is still unknown, so this stays open as 'mitigated, cause unknown' and Blocked on #4.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions