Skip to content

[Bug]: SwiftUI app stays stuck syncing/reconnecting until the other T3 app's device entry on the same iPhone is removed #13994

Description

@gmelvill

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/mobile

Steps to reproduce

  1. On one iPhone, install both the iOS app (1.3.0, build 85) and the SwiftUI TestFlight app (0.1.0, build 52), signed into the same T3 account.
  2. Pair both apps directly with the same Windows desktop environment over a private HTTPS address (Tailscale Serve). Both appear as separate, valid sessions on the host.
  3. Let the desktop install a nightly update and restart (backend unreachable for about 2–3 minutes).
  4. Reopen the iOS app. It reconnects normally.
  5. Open the SwiftUI app.

Expected behavior

Both apps reconnect on their own once the host is back.

Actual behavior

The SwiftUI app stays stuck on "Syncing threads" or loops on reconnecting. It only recovers after opening Devices in the SwiftUI app and removing the other entry, which is the 1.3.0 app on the same iPhone. This has happened repeatedly after desktop nightly updates, and the entry has to be removed again each time.

On the Devices screen both entries are named "iPhone". The other entry showed "last seen 5 hours ago" even though the 1.3.0 app had talked to the host a minute earlier. Removing it did not revoke the 1.3.0 app's session on the host, and afterwards both apps worked side by side. That suggests the Devices list is account-level and that two T3 apps on one physical device are treated as the same device, or one replaces the other.

Impact

Major degradation or frequent failure

Version or commit

iOS app 1.3.0 (85); SwiftUI TestFlight 0.1.0 (52); desktop 0.0.43-nightly.20260927.2344

Environment

iPhone, iOS 27; Windows 11 desktop host; direct connection over Tailscale Serve HTTPS

Logs or stack traces

The host's auth_sessions table showed one live, unrevoked session each for the SwiftUI app (0.1.0) and the iOS app (1.3.0), both connected within minutes of each other, before and after the device entry was removed from the Devices screen. So the stuck state does not appear to come from host-side session auth.

Workaround

Remove the other app's "iPhone" entry from Devices in the SwiftUI app.

Activity

  1. juliusmarminge commented on Sep 27, 2026

    @juliusmarminge
    Member

    Triage

    The two apps are not being treated as one device. The stuck reconnect and the Devices row are different systems, and removing that row does not touch the host session you checked.

    apps/swift-ios is only on t3code/rebuild-mobile-app-swift (157476f1fb, #5178), not on main. With a T3 account signed in, Devices does not list host auth_sessions. managesServerSessions is false in that case, and the screen loads relay push registrations instead.

    var managesServerSessions: Bool {
        !t3ConnectDeviceManager.hasActiveAccount
    }
    
    func loadDeviceSessions() async throws -> [FeatureDeviceSession] {
        if t3ConnectDeviceManager.hasActiveAccount {
            let devices = try await t3ConnectDeviceManager.registeredDevices()
            relayDeviceSessionIDs = Set(devices.map(\.deviceId))
            return devices.map {
                FeatureDeviceSession(
                    relayDevice: $0,
                    currentDeviceID: t3ConnectDeviceManager.currentRegisteredDeviceID
                )
            }
        }

    Remove then calls DELETE /v1/mobile/devices/:deviceId. That deletes the relay row only. It does not revoke the desktop pairing, which matches auth_sessions staying live for both 0.1.0 and 1.3.0. Host session revoke is the other branch, used only when there is no signed-in T3 account. The confirmation copy in that account-signed-in path says the device will stop receiving T3 Connect notifications.

    The two "iPhone" rows are two installations. SwiftUI stores its id as a UUID in the keychain service com.t3tools.t3code.swiftui.installation (PlatformInstallationIdentity). The iOS app stores a different UUID in its own secure storage (loadOrCreateAgentAwarenessDeviceId). The relay primary key is (userId, deviceId), so those stay two rows. Both titles are the platform device name (UIDevice.current.name and Constants.deviceName). On a current iPhone that string is often just "iPhone" for every app from this vendor. The subtitle is where the version shows up (iOS 27 · T3 Code 1.3.0 versus 0.1.0), when the registration included appVersion.

    "Last seen 5 hours ago" is also that relay row, not the host connection. The mapping copies updatedAt into lastConnectedAt and hardcodes isConnected: false, so the row can never say "Active now":

    init(relayDevice: T3ConnectRelayDevice, currentDeviceID: String?) {
        let updatedAt = Self.t3ConnectRelayDate(relayDevice.updatedAt)
        self.init(
            sessionID: relayDevice.deviceId,
            label: relayDevice.label,
            ...
            lastConnectedAt: updatedAt,
            isConnected: false,
            isCurrent: relayDevice.deviceId == currentDeviceID
        )
    }

    updatedAt moves only when a registration is written. The iOS app skips that request when the payload fingerprint is unchanged, including across launches that are actively talking to the desktop:

    if (
      persisted &&
      persisted.identity === identity &&
      persisted.signature === signature &&
      !needsAndroidReplay
    ) {
      setRegistrationStatus("registered");
      logRegistrationDebug("relay device registration skipped; already registered for account", {
        expectedGeneration,
      });
      return;
    }

    So a minute-old auth_sessions connection next to a hours-old Devices "last seen" is what this screen shows today. It is not evidence that one install replaced the other. After a successful remove, the 1.3.0 app will not put the row back until that fingerprint changes (app version, push token, notification prefs, relay URL). If the row is present again on the next nightly with none of those changed, that part is new.

    The reconnect stall is not explained by this list. Nothing in the direct Tailscale socket path reads relay devices. SwiftUI also has no "Syncing threads" string. That copy is the React Native home, when a shell snapshot is still synchronizing and one is already cached (workspace-connection-status.ts). On SwiftUI the matching UI is the sidebar "{environment} reconnecting", or "Reconnecting..." / "Catching up..." on a thread. HTTP fallback loads a shell without leaving that state. consumeFallbackShell publishes with markSourceConnected: false, and only a websocket shell-event batch calls emitConnection(.connected). A socket that never delivers a batch after the desktop restart will sit on "reconnecting" even though HTTP already has threads.

    Please reply with three things so we can separate a socket bug from the Devices row:

    1. The exact SwiftUI string while it is stuck (sidebar "reconnecting", "Catching up...", "Reconnecting...", or something else), and whether the thread list is already filled in behind it.
    2. Whether it recovers if you open Devices and pull to refresh without removing the other row. Removing it unregisters the 1.3.0 app from T3 Connect pushes, so it is worth not using that as the test.
    3. A host trace slice for the 0.1.0 client during the stuck window: websocket subscribe versus shellSnapshot HTTP, and whether those requests are hanging or failing.
  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 27, 2026
  3. gmelvill commented on Sep 27, 2026

    @gmelvill
    Author

    Thanks for the detailed triage. On point 3, the host trace:

    The stuck window has already rotated out. This host keeps 11 × 10 MB server.trace.ndjson files, which is about 40 minutes at current volume, and the oldest surviving span starts after the recovery. I'll capture the next occurrence promptly. The user agent separates the clients cleanly (T3Code/52 = SwiftUI 0.1.0, T3Code/85 = iOS 1.3.0), so I can give you websocket vs /api/orchestration/shell timing for the SwiftUI client alone.

    The only SwiftUI data left is from just after recovery (UTC+10), which may still be useful:

    • 05:39:09 GET /api/orchestration/shell 200, 10.4 s
    • 05:39:19–05:39:21: six POST /api/auth/websocket-ticket + GET /ws pairs within about two seconds; five of those sockets closed within 0.8–8.4 s
    • 05:39:35 a /ws that then stayed up for 558 s

    So even in a working state, the SwiftUI client opened a burst of short-lived sockets after a slow shell read. All requests returned 200/204; nothing failed on the host side.

    Points 1 and 2 need a check on the phone at the next occurrence. I'll report the exact SwiftUI wording, whether the list is populated, and whether Devices pull-to-refresh alone (without removing the row) recovers it. You're probably right that "Syncing threads" was the React Native app.

  4. gmelvill commented on Sep 28, 2026

    @gmelvill
    Author

    Follow-up with host-side data from two more occurrences (times UTC+10).

    Point 2: pull-to-refresh on the Devices screen (Settings → Devices) in the SwiftUI app recovered it, without removing the other device row.

    Point 1: I didn't catch the exact wording this time; I'll note it next time.

    Point 3, host trace: on both occasions, GET /api/orchestration/shell from the SwiftUI client (T3Code/52) was cancelled by the client (HTTP 499), and the iOS client (T3Code/85) did the same:

    • 08:05 and 08:52: shell requests cancelled after 1.2–6.1 s, followed by bursts of /ws reconnects.
    • 08:54:47: a shell request that did complete took 52.9 s. Almost all of it was ProjectionSnapshotQuery.resolveRepositoryIdentitiesForProjects (55–64 s per run).
    • After a desktop update at 09:05: both apps' shell requests were cancelled at exactly 6.0 s every time; reconnect took about 2 min (SwiftUI) and 3½ min (iOS).

    A host-side factor that is probably part of this: Windows process cleanup on this machine had degraded badly. A trivial WMI query took 15–79 s and taskkill /T took 15–63 s. The ~19 git/gh/jj/glab/az subprocesses inside resolveRepositoryIdentitiesForProjects each ran in under 1 s, with 11–32 s gaps between them, so the delay was process cleanup, not Git. Signing out of Windows and back in cleared it: WMI 0.2 s, cleanup 0.3 s, resolveRepositoryIdentitiesForProjects under 0.6 s.

    So on this machine the stall is at least partly "host too slow for the client's shell timeout". Two things still look like client behaviour worth checking:

    • The shell request is abandoned after about 6 s, and the client then churns websockets instead of waiting for, or reusing, the in-flight request.
    • The client doesn't recover on its own once the host is fast again; it needs a manual refresh (here, on the Devices screen).

    — Geoff. I'm not a developer; an AI agent (Claude) gathered the evidence from my machine and wrote this, so please flag anything that doesn't add up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions