Skip to content

[Bug]: T3 Connect reports "Your cloud sign-in changed" when the host is just overloaded #11004

Description

@Project516

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/web

(The failing logic is in packages/client-runtime, which the dropdown does not list. The message is surfaced by the web client, and the timing that triggers it comes from apps/web/src/cloud/managedAuth.tsx.)

Steps to reproduce

  1. Sign in to T3 Connect and connect to a relay environment from the web client.
  2. Put the host under heavy CPU/memory load. On a low-power arm64 single-board machine this happens on its own: two or three agent turns plus a build is enough to saturate it.
  3. Watch the connection status while the machine is pegged.

Expected behavior

While the machine is overloaded the client should report a transient condition, something along the lines of "Reconnecting, the environment is slow to respond", or a timeout, and keep retrying. The sign-in is still valid and nothing about it changed.

Actual behavior

The connection status reads:

Failed to connect. Reconnecting... Reason: Your cloud sign-in changed. Sign in again to authorize the environment.

The sign-in did not change. Signing out and back in does nothing, because there is nothing wrong with the session. The connection recovers on its own once load drops, which is the tell that the message is describing the wrong cause. The practical cost is that the message sends you to fix an auth problem that does not exist instead of the load problem that does.

Impact

Minor bug or occasional failure (misleading diagnosis, self-recovering)

Version or commit

0.0.40. The string is present in the shipped bundle for that release, and the code below was read at main @ de37964db, whose package.json is also 0.0.40.

Environment

Low-power arm64 single-board Linux host, recent Node. Client is the T3 Code web app, connecting through T3 Connect. Nothing here looks hardware specific: any host slow enough to delay Clerk load or the activation promise chain should reproduce it.

Logs or stack traces

Failed to connect. Reconnecting... Reason: Your cloud sign-in changed. Sign in again to authorize the environment.

No stack trace: this is a handled ConnectionBlockedError rendered into the connection status line, not a throw.

Screenshots, recordings, or supporting files

None. The full text of the status line is quoted above.

Workaround

Wait for the load to drop. The connection recovers without any sign-in action.

Diagnosis

The message comes from sessionChanged() in packages/client-runtime/src/authorization/service.ts:224:

const sessionChanged = () =>
  new ConnectionBlockedError({
    reason: "authentication",
    detail: "Your cloud sign-in changed. Sign in again to authorize the environment.",
  });

It is raised by assertSession and by getDpopToken whenever cloudSession.identity is None, or is a different object reference than the identity captured when the attempt started. On web, identity is the managedRelaySessionAtom value (apps/web/src/connection/platform.ts:185), and that atom is null until ManagedRelayAuthProvider activates it.

Two things make "identity is not available yet" indistinguishable from "the user signed in as somebody else":

  1. ManagedRelayAuthProvider (apps/web/src/cloud/managedAuth.tsx) bails out while !isLoaded (line 51), and even on the signed-in path it defers activateManagedRelayAuthentication behind a promise chain (activateAfterTransition, line 92). On a loaded machine, Clerk taking longer to load and that promise chain taking longer to drain both widen the window where the atom is still null while the connection supervisor is already attempting. Every attempt in that window fails with sessionChanged().
  2. identity is compared by reference (current.value !== identity). Anything that replaces the session object rather than updating it in place reads as an account change. setManagedRelaySession deliberately keeps the object stable across same-account token refreshes, so this is handled for the common case, but the None case has no such guard and produces the same message.

Because sessionChanged() is a ConnectionBlockedError, the supervisor moves to blocked and parks. When activation finally lands, managedRelayAccountChanges fires the credentials-changed signal and the supervisor retries, carrying the stale failure into the next connecting state as lastFailure. That is exactly the "Failed to connect. Reconnecting... Reason: ..." shape from connectionStatusText in packages/client-runtime/src/connection/presentation.ts:68, and it explains why the message is shown while the client is in fact recovering.

Likely making it worse on a slow host: CACHED_ENDPOINT_SOCKET_TIMEOUT_MS is 3000 ms (service.ts:82). On an overloaded host the /api/auth/websocket-ticket round trip against a cached endpoint routinely exceeds that, so authorizeDpop discards a perfectly good access token, re-runs the full relay bootstrap and token exchange, and adds load to the host that is already the bottleneck. Each of those extra passes is another chance to hit the None-identity window above.

Suggested fix

Separate "the cloud session is not available yet" from "the cloud session changed":

  • When cloudSession.identity is None and there is no evidence of an account switch, fail with a ConnectionTransientError (reason: "network", or a new session-pending) so the supervisor uses normal backoff and the UI says reconnecting rather than telling the user to sign in.
  • Keep sessionChanged() for the real case: a captured identity whose accountId differs from the current one.
  • Consider comparing accountId rather than object reference in assertSession, so an object replacement for an unchanged account cannot produce the message.
  • Consider whether the 3 s cached-endpoint socket timeout should scale, or whether a timeout there should avoid throwing away the cached access token.

Activity

  1. juliusmarminge commented on Sep 9, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on current main. This is a real client-runtime bug: a pending or slow-to-activate cloud session is reported as an account switch.

    What the code does

    sessionChanged() in packages/client-runtime/src/authorization/service.ts is a ConnectionBlockedError with Your cloud sign-in changed. Sign in again to authorize the environment. assertSession and getDpopToken raise it whenever cloudSession.identity is None, or is a different object than the identity captured when the attempt started.

    On web and mobile, that identity is the managedRelaySessionAtom value. The atom stays null until ManagedRelayAuthProvider / CloudAuthProvider finishes Clerk load and the activateAfterTransition promise chain. CloudSession has no pending state — None means both “not ready yet” and “signed out / switched.”

    The supervisor parks on blocked errors. When activation finally lands, managedRelayAccountChanges retries and connectionStatusText carries the stale reason into Failed to connect. Reconnecting... Reason: …. That matches the report, including recovery once load drops.

    setManagedRelaySession already keeps the session object stable across same-account token refreshes (CloudSessionIdentity is documented as stable for one signed-in session). Object-ref compare is intentional for that contract. The hole is treating None as “changed.”

    The 3s CACHED_ENDPOINT_SOCKET_TIMEOUT_MS on the cached websocket-ticket path (vs 10s for other remote calls) is a likely amplifier: a slow-but-valid ticket fails transiently, authorizeDpop discards the cached access token, re-bootstraps, and can hit the None window again.

    This logic landed in #9594. Existing tests correctly treat logout during renewal as blocked; they do not cover authorize while identity is still None because activation has not finished.

    Not a duplicate of #5031

    #5031 is the same reconnect chrome on a different failure (GET /ws → 204 on a nightly relay). Leave both open.

    Suggested fix

    1. Separate “session not available yet” from “session changed.” None with no captured identity and no account-switch evidence should be a ConnectionTransientError (or a dedicated pending state on CloudSession), not sessionChanged().
    2. Keep sessionChanged() for a captured identity whose accountId differs, and for logout mid-renewal.
    3. Consider not discarding a cached access token when the cached-endpoint ticket times out, and/or scaling that 3s cap under load.

    Do not make every None transient — a real sign-out should still block with a sign-in message (clerkToken already has Sign in to T3 Connect…).

    No open PR for this. Accepting as a minor, self-recovering bug with a misleading diagnosis.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 9, 2026
  3. added a commit that references this issue on Sep 23, 2026
    fd75b10
  4. added 3 commits that reference this issue on Oct 1, 2026
    9f4e1c4
    9389164
    e06fab3
  5. j-gaertig commented on Oct 9, 2026

    @j-gaertig

    Confirming the triage from a mobile angle and adding a repro without host overload.

    Device: Samsung Galaxy A55. App: self built APK from current main, Play Store version uninstalled first. Exact commit on request.

    I see the same string from this issue:

    Failed to connect. Reconnecting... Reason: Your cloud sign-in changed. Sign in again to authorize the environment.

    My sign in did not change. The account screen still shows me as signed in and the environments stay in the list. It is only a connection error, no deleted environments.

    Repro without reboot:

    1. Open the app once so sessions exist.
    2. Minimize the app with the Home gesture. Do not swipe it away in Recents, only send it to background.
    3. Turn airplane mode on.
    4. Force stop the app via Android Settings.
    5. Turn airplane mode off.
    6. Open the app again. The error appears.

    Counter tests:

    Normal reboot, wait until the phone is idle and WiFi is stable, then open the app: works. Forced reboot with Power plus Volume Down, same waiting: works. So the trigger looks like a cold start while network or Clerk activation is not ready yet, not reboot corruption.

    Healing note for the airplane mode case only: the error does not heal while the app stays open with network back. Full close and reopen heals it, then the retry has a present session and connects.

    This matches the None identity window described here.

  6. added a commit that references this issue on Oct 9, 2026
    97efee6
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions