Repository navigation
[Bug]: Usage page loads forever when a remembered remote environment is offline #6045
Description
Activity
Your diagnosis holds on
main, and I think the reason this one is awkward to fix is worth writing down before anyone picks a direction.The client-side chain is exactly as you describe:
usageByWindowAtommarks a row terminal only when the query itself fails (apps/web/src/state/usage.ts),useUsagecounts an environment with no summary and no error as still reporting, andUsagePageholds all content whileisPending || isPartial. A query against an unreachable environment stays pending rather than failing, so the row never becomes terminal.The obvious fix looks close at hand. That same atom already reads the environment's
presentation— it just usespresentation.entry.target.labeland ignorespresentation.connection.phase, so the connection layer's verdict never reaches the usage row. I built that version: treat the phases that only appear after a failed attempt as terminal, keepconnectingpending, and keep any summary that already arrived. It fixes the hang and needs no new state.Then the phase names turn out not to mean what they look like, which I think is the real finding here.
offlineis not "the remote is unreachable". The supervisor entersofflineStateonly whenintent.network === "offline"(packages/client-runtime/src/connection/supervisor.ts, both the initial-state branch and the retry loop), and that is the local client'sNetworkStatus. In this report the local machine has working network — the other Mac is asleep — so the environment never readsoffline. It cyclesconnecting→backoff, andpresentConnectionStatemaps both a retriedconnectingandbackofftoreconnecting.blocked→erroris a different situation.So
reconnectingis the only phase that covers this report, and it is also the phase that means a retry is actively in flight. Treating it as terminal fixes the hang but makes a single transient failure flash that device as unreachable and then revise totals the page had just presented as settled. That trade seemed worse than the bug to me, so I did not open a PR for it.What is missing is a bounded notion of "retried long enough to stop waiting".
SupervisorConnectionStatecarriesattemptandretryAt, butEnvironmentConnectionPresentationdeliberately narrows tophase/error/traceId, so nothing in the view layer can tell a first retry from the twentieth.Three directions, all of which need a number that is yours to choose:
- Widen the presentation with
attempt(or a derived "retries exhausted") and treat an environment past that threshold as unavailable for usage. Smallest change, keeps the policy in the connection layer where the state already lives. - Bound the usage query itself so an unanswered request fails terminally after a deadline. Consumers then get one consistent lifecycle instead of this view second-guessing a pending query.
- Stop holding everything. Let
UsagePagerender the environments that answered with an explicit coverage note for the ones still out, rather than gating all content onisPending || isPartial.
One adjacent thing worth deciding at the same time: an environment that answered once and then went unreachable keeps its cached summary counted in the totals with no staleness marker. That is current behavior on
main, not something any of the above introduces, but options 1 and 2 both make it more visible.Happy to put up whichever of the three you'd prefer — I have 1 working locally with tests.
- Widen the presentation with
Option 1 sounds best to me. It allows the device time to reconnect, but prevents it from blocking the Usage page forever. Please go ahead with your implementation and tests.
- added a commit that references this issue
on Sep 2, 2026 There is a PR for this: #5920 takes the "bound the query" route from the options above (a 30 second per-environment timeout, the page renders as soon as one environment answers, and the unreachable one is listed as unable to report). I verified it on a real two-environment setup with one of them powered off and it fixes exactly this hang. The PR currently references the closed duplicate #6363 rather than this issue, so it will not close this one automatically.
Fixed by #9860 (usage renders as environments respond instead of waiting for every environment).
Before submitting
Area
apps/web
Steps to reproduce
Expected behavior
The Usage page should eventually show totals from the available environment and clearly mark the offline environment as unavailable/incomplete. A disconnected device should reach a terminal error state or be excluded after a bounded timeout.
Actual behavior
The page remains on its loading skeleton indefinitely with “1 device still scanning.” The online environment completes successfully, while the offline environment is reconnecting and reports a 10-second remote endpoint timeout under Settings → Connections.
The remote usage request appears to remain pending instead of becoming a failure. In the current client flow,
useUsage()treats an environment with no summary and no error as still reporting, andUsagePagedeliberately holds all content whileisPending || isPartial. Because the unreachable environment never becomes terminal, the page never displays even the completed device's totals.Impact
Major degradation or frequent failure
Version or commit
0.0.34-nightly.20260810.1061
Environment
macOS desktop app; two remembered Mac environments connected using Tailscale HTTPS; one remote Mac asleep/unreachable.
Logs or stack traces
The Connections screen continues showing “Failed to connect. Reconnecting...” / “Connecting...” while Usage remains on “1 device still scanning.”
Workaround
Wake the remote Mac so it reconnects, or temporarily remove the offline environment under Settings → Connections → Remote environments.