Skip to content

[Bug]: Remote environment can't connect while unknown-remote probes saturate VcsProcess ("did not respond during connection setup") #16503

Description

@IvanKhramov

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Run t3 serve on a Linux VM as a remote environment (reverse proxy in front, desktop connects over wss://).
  2. Open several projects whose origin is a self-hosted GitLab with a hostname that has no gitlab DNS label (e.g. git.example.com), so it is classified unknown.
  3. Have glab installed on the server and authenticated for gitlab.com only, not for that host.
  4. Reload the desktop window (or restart the app) so it opens a new WebSocket to the remote environment.

Expected behavior

The connection is established in well under a second. Slow or failing source-control probes may delay PR/VCS status, but must not block connection setup.

Actual behavior

The desktop stays on "Reconnecting: did not respond during connection setup." for minutes, and restarting the desktop app does not help.

What the server trace shows for each attempt:

  • nginx: GET /ws?wsTicket=... returns 101. The server sends ~110 bytes and the client closes after ~7s.
  • EnvironmentAuth.authenticateWebSocketUpgrade succeeds.
  • ws.rpc.subscribeServerConfig is then Interrupted after 7.2s. In the same trace:
    • VcsDriverRegistry.detect 7075 ms
    • detectRepository 6983 ms
    • VcsProcess.run 6982 ms
  • At the same time, VcsProcess.run averages 7-9 s across ~200 runs/min. 1829 runs over 25 minutes were VcsProcessTimeoutError ... source-control.discovery.refine-unknown-remote: glab ... after 5000ms.
  • Sampling the children of t3 serve shows only glab auth status, repeated in every project dir.

So the connection-setup path (subscribeServerConfig → VcsDriverRegistry.detect) waits behind the unknown-remote refinement storm described in #12060. Each refinement there costs a full 5s glab auth status timeout rather than a fast ENOENT. The client gives up at ~7s and retries, and each retry lands behind the same queue.

On 0.0.45 the cause matches #12060. resolveHostingProvider and the RepositoryIdentityResolver refine pass an explicit context, so resolveHandle skips providerContextCache (TTL 5s, failures not cached) and calls refineUnknownRemoteProvider on every read.

Impact

Major degradation or frequent failure. The remote environment is unusable until the queue happens to drain. A connection that succeeds is luck: one detect call has to finish under the client's ~7s setup timeout.

Version or commit

0.0.45 (server and desktop)

Environment

  • Server: Linux x64 VM, t3 serve as a systemd user service, 4 vCPU / 8 GB.
  • Desktop: macOS, connecting through an nginx reverse proxy over a WireGuard mesh.
  • glab 1.119.0 on the server PATH.

Logs or stack traces

ws.rpc.subscribeServerConfig     7197 ms  Interrupted  InterruptError: All fibers interrupted without error
  VcsDriverRegistry.detect       7075 ms
    detectRepository             6983 ms
      VcsProcess.run             6982 ms

VcsProcessTimeoutError: VCS process timed out in source-control.discovery.refine-unknown-remote: glab (<project dir>) after 5000ms

Per minute on the server, before and after the workaround:

before: VcsProcess.run n≈200/min avg 7-9 s   detectRepository 7-15 s
after:  VcsProcess.run n≈380/min avg 0.9-1.4 s   detectRepository ~1 s

Workaround

Run glab auth login --hostname <that host> on the server. The refinement still runs on every read, but each call returns in about a second instead of timing out, and connection setup goes through again. This is fragile: once the token expires the storm and the connection failures come back.

Suggested fixes:

  1. Don't make connection setup (subscribeServerConfig) wait on VCS detection, or give it a separate, unshared process budget.
  2. Route explicit-context calls through a cache keyed by remote host, with a TTL in minutes that includes negative results (the Unrecognised GitHub Enterprise host re-probes fj, tea and glab once per project on every sweep #12060 fix).
  3. Allow declaring self-hosted GitLab/GitHub hosts in settings so they never enter the unknown bucket ([Bug]: Bare pull request numbers can't be linked on a self-hosted GitLab whose host name doesn't say "gitlab" #15390 and [Feature]: GitHub Enterprise support in the source control connector #5087 would also benefit).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions