Skip to content

[Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258

Description

@leeonfield

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. On macOS, run a full-tunnel VPN client whose DNS answers with session-scoped virtual IPs. Here it is a corporate ZTNA/VPN client: while it is on, region1.v2.argotunnel.com resolves to virtual addresses from the client's own range that only route through its utun interface, instead of Cloudflare's published 198.41.192.x addresses.
  2. Run the T3 Code background service with T3 Connect enabled and confirm a phone can open the environment. cloudflared registers 4 QUIC connections to the virtual addresses it resolved at startup.
  3. Put the Mac to sleep, move it to another network, and wake it. Here the VPN's network extension restarted on wake and rebuilt its address map: region1/region2.v2.argotunnel.com now resolve to different virtual addresses, and the old ones no longer route anywhere.
  4. Try to open the environment from the phone.

I don't have a VPN-free repro yet. The core condition is that the addresses cloudflared resolved at startup stop working while DNS starts returning new ones.

Expected behavior

When the managed relay client is still running but has had no registered tunnel connection for a sustained period, T3 should replace it. A fresh cloudflared re-resolves the edge and reconnects. Doing that by hand restored the tunnel within 3 seconds (see Workaround).

Proposed semantics for maintainers to decide before any PR:

  • Signal: use cloudflared's own readiness endpoint. Spawn it with --metrics 127.0.0.1:0, read the bound address from its Starting metrics server on <addr>/metrics line, and poll /ready. It returns 200 when readyConnections > 0 and 503 otherwise. If that line never appears, do nothing, which keeps today's behavior.
  • Policy: poll every 30 s and count only explicit 503 responses; request errors don't count. After 6 consecutive 503s (about 3 minutes of awake time, since timers pause during sleep), kill the child and let the existing superviseConnector restart it. Double the threshold after each restart that doesn't reconnect, up to about 30 minutes. Reset it once /ready returns 200.
  • Open question: the exit path also queues a managed-tunnel recovery request. A watchdog restart fixes a transport problem, not a credential one, so it may be better to skip that request.
  • Out of scope: reporting running before registration ([Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447). Also the stale QUIC MTU case (T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897), where /ready still reports 4 connections.

Actual behavior

Addresses are replaced with documentation ranges: 192.0.2.x stands for the virtual addresses cloudflared resolved at startup, and 203.0.113.x for the ones DNS returned after the network change. 198.41.192.7 is a real, published Cloudflare edge address.

The tunnel stayed down for 2 h 47 min, until I restarted cloudflared by hand:

  • cloudflared kept dialing only the 20 edge addresses it had resolved at startup, 36 hours earlier (192.0.2.120–192.0.2.139). The log has 461 failed to dial to edge with quic: timeout: no recent network activity errors and no successful registration.
  • curl http://127.0.0.1:20241/ready returned {"status":503,"readyConnections":0,...}.
  • The environment's relay hostname returned HTTP 530, and the phone could not connect.
  • T3 did nothing. The process never exited and no registration was rejected, so neither recovery path ran. t3 connect status reports saved setup, not live state, so nothing on the host showed the outage either.

The network was not blocking Cloudflare Tunnel. I checked from the same Mac at the same time, using a raw QUIC version-negotiation probe and an openssl s_client handshake:

Destination UDP 7844 (QUIC) TCP 7844 (TLS)
Current DNS answer for region1/region2 (e.g. 203.0.113.143) reply in ~10 ms Cloudflare presents its Origin certificate
Published edge IP 198.41.192.7 reply in ~8 ms Cloudflare presents its Origin certificate
Address cloudflared kept dialing (192.0.2.121) no reply connection reset

Why it never recovers:

  1. cloudflared resolves edge addresses only once, in NewSupervisor (supervisor.go). edgediscovery.Edge never refreshes them (edgediscovery.go), so it only rotates through the startup set. With normal DNS that set is Cloudflare's stable published IPs, so this rarely shows. With virtual-IP DNS the set can go stale while the process runs.
  2. T3 restarts the connector only in two cases: when the process exits (superviseConnector), or after 4 rejected registrations (observeConnectorOutput). Transport errors are only logged, so a live connector with zero connections is kept forever.

Any DNS that hands out session-scoped virtual IPs, such as other ZTNA clients or fake-IP modes in proxy tools, should hit the same problem. I have only reproduced it with one corporate VPN client.

Related: #7447 (a spawned but unreachable tunnel is reported as running) and #15897 (stale QUIC MTU after joining a VPN). All three treat a live cloudflared process as a working tunnel. This report covers a tunnel that worked and then went stale.

Impact

Blocks work completely

Version or commit

T3 Code desktop 0.0.45 and t3 service 0.0.45. Code links point to main @ f5eb250.

Environment

macOS 27.0 (Apple Silicon); cloudflared 2026.6.0 from Homebrew (picked up from PATH); corporate full-tunnel ZTNA/VPN client;

Logs or stack traces

# 2026-10-04 06:32Z, home network: initial registration (lines trimmed, IDs redacted)
INF Registered tunnel connection connIndex=0 ip=192.0.2.127 protocol=quic

# 2026-10-05 19:00Z, after sleep + network change: every connection drops
ERR failed to accept incoming stream requests error="failed to accept QUIC stream: timeout: no recent network activity" connIndex=0 ip=192.0.2.127

# 19:00Z-21:47Z: only the startup addresses are retried (461 times)
ERR Failed to dial a quic connection error="failed to dial to edge with quic: timeout: no recent network activity" connIndex=0 ip=192.0.2.121

# Meanwhile DNS returns different addresses
$ dscacheutil -q host -a name region1.v2.argotunnel.com
203.0.113.141 ... 203.0.113.150

# 21:47Z, after `kill <cloudflared pid>`
WARN Relay client exited; restarting
INF Registered tunnel connection connIndex=0 ip=203.0.113.143 protocol=quic
INF Registered tunnel connection connIndex=1 ip=203.0.113.204 protocol=quic
INF T3 Connect managed tunnel recovered

Screenshots, recordings, or supporting files

No response

Workaround

Kill only the managed cloudflared process: kill <pid>, where <pid> comes from pgrep -P <t3 serve pid> cloudflared. T3 restarts it right away, the new process re-resolves the edge, and all 4 connections registered within 3 seconds. t3 service restart also works, but it restarts the whole server.


I'm happy to send a focused PR once the intended behavior is agreed.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 5, 2026
  2. juliusmarminge commented on Oct 5, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the detailed timeline, logs, and the proposed watchdog, @leeonfield! The /ready and readyConnections evidence made this very clear.

    What I found

    • CloudManagedEndpointRuntime (apps/server/src/cloud/ManagedEndpointRuntime.ts) only replaces the managed cloudflared in two cases: the process exits (superviseConnector), or the edge rejects the tunnel four times (observeConnectorOutput: Failed to get tunnel, missing record, invalid secret). A live process whose dials time out is logged as a transport warning and kept, and running only means the child is alive. That matches the 2h47m outage with /ready at 503 and readyConnections: 0.
    • cloudflared resolves the edge once (NewSupervisor → ResolveEdge) and edgediscovery.Edge only rotates within that set. That's fine with published anycast addresses but not with session-scoped virtual IPs from a ZTNA client, so the connector stays stuck until restarted.
    • Killing the child works because the supervisor restarts it, the new process resolves the edge again, and the existing recovery path runs. Your log shows exactly that (T3 Connect managed tunnel recovered).
    • The existing crash backoff wouldn't help here: it resets once uptime passes 30s, so a connector that has been up for hours with zero connections looks stable.
    • Not a duplicate of [Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447 (reporting running before first registration) or T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897 (stale QUIC MTU where /ready stays 200).

    Likely fix area

    One option along the lines you proposed, with details a maintainer may want to weigh:

    • Spawn with --metrics 127.0.0.1:0 (loopback only, since the metrics server also exposes debug endpoints) and read the bound address from that child's Starting metrics server on … line, rather than assuming the default port 20241, which may belong to another process.
    • Poll /ready periodically, count only consecutive HTTP 503s (a 200 resets the streak, request errors leave it unchanged), kill the child after a threshold, and back off across restarts that never reach a 200, tracked outside the child's lifetime.
    • Letting the existing supervisor handle the restart would keep the relay recovery request in the loop, which matters if the reaper has deleted the tunnel during a long outage. A local-only restart could end up dialing a deleted tunnel.
    • Tests could cover the 503 streak, the 200 reset, ignored request errors, and a missing metrics line.

    A maintainer will decide on the fix direction.

  3. added
    via-triageFiled through npx t3 triage
    and removed
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 5, 2026
  4. saif-o99 commented on Oct 6, 2026

    @saif-o99

    @juliusmarminge Kindly pay attention to this one, I tried upgrading the app on both devices, tried to turn it off on both devices and turn it on. nothing helped. i can't figure how to reconnect to my laptop

    Image
  5. notgiorgi commented on Oct 7, 2026

    @notgiorgi

    Another repro, this time with Cloudflare WARP rather than a corporate ZTNA client. Desktop Alpha 0.0.45 on macOS 26.6 (Mac mini, arm64), T3 Connect hosted by the desktop app, no background service. Homebrew cloudflared 2026.10.0.

    Timeline:

    • WARP (Warp mode, MASQUE, always-on) was already connected when the app started. cloudflared registered fine and served requests through the relay for ~4h.
    • WARP's daemon did a disconnected → connecting → connected cycle that took ~1.3s. From that moment no request reached the host.
    • 21h later the connector child is still alive, /ready returns {"status":503,"readyConnections":0}, and the relay hostname answers Cloudflare 530 / error 1033. Other clients get endpoint_request_failed.
    • Metrics over the process lifetime: tunnel_register_success 16, tunnel_register_fail{error="server_error"} 4352 and climbing (+2 per 45s), {error="dup_edge_conn"} 75, quic_client_closed_connections 4539. So it is retrying forever, not stuck on a dead dial.
    • Nothing from cloudflared is on disk ([Bug]: Desktop-hosted T3 Connect keeps no record of the managed cloudflared's output #16331), so I can't see the edge's exact rejection text. QUIC gives an opaque control-stream error anyway (see the analysis in fix(connect): recover a reaped tunnel the connector cannot register #16402), which is why isRejectedRelayClientTunnelOutput never fires.

    Contributing factor: the edge ranges 198.41.192.0/24 and 198.41.200.0/24 route through WARP's utun and are not in its split-tunnel exclude list.

    Same conclusion as above: a live-but-unregistered connector needs a readiness-based restart. Polling /ready as proposed would have caught this within minutes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions