Skip to content

[Bug]: Usage page loads forever when a remembered remote environment is offline #6045

Description

@Gabriel-Rivera-Work

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/web

Steps to reproduce

  1. Configure T3 Code with two desktop environments/devices.
  2. Leave both environments remembered by the client.
  3. Put one device to sleep (or otherwise make its remote environment unreachable).
  4. On the online device, open Usage.
  5. Wait beyond the remote connection timeout.

Expected behavior

The Usage page should eventually show totals from the available environment and clearly mark the offline environment as unavailable/incomplete. A disconnected device should reach a terminal error state or be excluded after a bounded timeout.

Actual behavior

The page remains on its loading skeleton indefinitely with “1 device still scanning.” The online environment completes successfully, while the offline environment is reconnecting and reports a 10-second remote endpoint timeout under Settings → Connections.

The remote usage request appears to remain pending instead of becoming a failure. In the current client flow, useUsage() treats an environment with no summary and no error as still reporting, and UsagePage deliberately holds all content while isPending || isPartial. Because the unreachable environment never becomes terminal, the page never displays even the completed device's totals.

Impact

Major degradation or frequent failure

Version or commit

0.0.34-nightly.20260810.1061

Environment

macOS desktop app; two remembered Mac environments connected using Tailscale HTTPS; one remote Mac asleep/unreachable.

Logs or stack traces

Remote environment endpoint https://<redacted-tailnet-host>/.well-known/t3/environment timed out after 10000ms.

The Connections screen continues showing “Failed to connect. Reconnecting...” / “Connecting...” while Usage remains on “1 device still scanning.”

Workaround

Wake the remote Mac so it reconnects, or temporarily remove the offline environment under Settings → Connections → Remote environments.

Image

Activity

  1. CDVolvik commented on Aug 16, 2026

    @CDVolvik
    Contributor

    Your diagnosis holds on main, and I think the reason this one is awkward to fix is worth writing down before anyone picks a direction.

    The client-side chain is exactly as you describe: usageByWindowAtom marks a row terminal only when the query itself fails (apps/web/src/state/usage.ts), useUsage counts an environment with no summary and no error as still reporting, and UsagePage holds all content while isPending || isPartial. A query against an unreachable environment stays pending rather than failing, so the row never becomes terminal.

    The obvious fix looks close at hand. That same atom already reads the environment's presentation — it just uses presentation.entry.target.label and ignores presentation.connection.phase, so the connection layer's verdict never reaches the usage row. I built that version: treat the phases that only appear after a failed attempt as terminal, keep connecting pending, and keep any summary that already arrived. It fixes the hang and needs no new state.

    Then the phase names turn out not to mean what they look like, which I think is the real finding here.

    offline is not "the remote is unreachable". The supervisor enters offlineState only when intent.network === "offline" (packages/client-runtime/src/connection/supervisor.ts, both the initial-state branch and the retry loop), and that is the local client's NetworkStatus. In this report the local machine has working network — the other Mac is asleep — so the environment never reads offline. It cycles connecting → backoff, and presentConnectionState maps both a retried connecting and backoff to reconnecting. blocked → error is a different situation.

    So reconnecting is the only phase that covers this report, and it is also the phase that means a retry is actively in flight. Treating it as terminal fixes the hang but makes a single transient failure flash that device as unreachable and then revise totals the page had just presented as settled. That trade seemed worse than the bug to me, so I did not open a PR for it.

    What is missing is a bounded notion of "retried long enough to stop waiting". SupervisorConnectionState carries attempt and retryAt, but EnvironmentConnectionPresentation deliberately narrows to phase/error/traceId, so nothing in the view layer can tell a first retry from the twentieth.

    Three directions, all of which need a number that is yours to choose:

    1. Widen the presentation with attempt (or a derived "retries exhausted") and treat an environment past that threshold as unavailable for usage. Smallest change, keeps the policy in the connection layer where the state already lives.
    2. Bound the usage query itself so an unanswered request fails terminally after a deadline. Consumers then get one consistent lifecycle instead of this view second-guessing a pending query.
    3. Stop holding everything. Let UsagePage render the environments that answered with an explicit coverage note for the ones still out, rather than gating all content on isPending || isPartial.

    One adjacent thing worth deciding at the same time: an environment that answered once and then went unreachable keeps its cached summary counted in the totals with no staleness marker. That is current behavior on main, not something any of the above introduces, but options 1 and 2 both make it more visible.

    Happy to put up whichever of the three you'd prefer — I have 1 working locally with tests.

  2. Gabriel-Rivera-Work commented on Aug 18, 2026

    @Gabriel-Rivera-Work
    Author

    Option 1 sounds best to me. It allows the device time to reconnect, but prevents it from blocking the Usage page forever. Please go ahead with your implementation and tests.

  3. Mnigos commented on Sep 4, 2026

    @Mnigos
    Contributor

    There is a PR for this: #5920 takes the "bound the query" route from the options above (a 30 second per-environment timeout, the page renders as soon as one environment answers, and the unreachable one is listed as unable to report). I verified it on a real two-environment setup with one of them powered off and it fixes exactly this hang. The PR currently references the closed duplicate #6363 rather than this issue, so it will not close this one automatically.

  4. juliusmarminge commented on Sep 5, 2026

    @juliusmarminge
    Member

    Fixed by #9860 (usage renders as environments respond instead of waiting for every environment).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions