Repository navigation
[Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258
Description
Activity
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 5, 2026 Note
Grok responding on behalf of Julius.
Triage
Thanks for the detailed timeline, logs, and the proposed watchdog, @leeonfield! The
/readyandreadyConnectionsevidence made this very clear.What I found
CloudManagedEndpointRuntime(apps/server/src/cloud/ManagedEndpointRuntime.ts) only replaces the managedcloudflaredin two cases: the process exits (superviseConnector), or the edge rejects the tunnel four times (observeConnectorOutput:Failed to get tunnel, missing record, invalid secret). A live process whose dials time out is logged as a transport warning and kept, andrunningonly means the child is alive. That matches the 2h47m outage with/readyat 503 andreadyConnections: 0.- cloudflared resolves the edge once (
NewSupervisor→ResolveEdge) andedgediscovery.Edgeonly rotates within that set. That's fine with published anycast addresses but not with session-scoped virtual IPs from a ZTNA client, so the connector stays stuck until restarted. - Killing the child works because the supervisor restarts it, the new process resolves the edge again, and the existing recovery path runs. Your log shows exactly that (
T3 Connect managed tunnel recovered). - The existing crash backoff wouldn't help here: it resets once uptime passes 30s, so a connector that has been up for hours with zero connections looks stable.
- Not a duplicate of [Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447 (reporting
runningbefore first registration) or T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897 (stale QUIC MTU where/readystays 200).
Likely fix area
One option along the lines you proposed, with details a maintainer may want to weigh:
- Spawn with
--metrics 127.0.0.1:0(loopback only, since the metrics server also exposes debug endpoints) and read the bound address from that child'sStarting metrics server on …line, rather than assuming the default port 20241, which may belong to another process. - Poll
/readyperiodically, count only consecutive HTTP 503s (a 200 resets the streak, request errors leave it unchanged), kill the child after a threshold, and back off across restarts that never reach a 200, tracked outside the child's lifetime. - Letting the existing supervisor handle the restart would keep the relay recovery request in the loop, which matters if the reaper has deleted the tunnel during a long outage. A local-only restart could end up dialing a deleted tunnel.
- Tests could cover the 503 streak, the 200 reset, ignored request errors, and a missing metrics line.
A maintainer will decide on the fix direction.
- addedvia-triageFiled through npx t3 triageFiled through npx t3 triageand removedneeds-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 5, 2026 @juliusmarminge Kindly pay attention to this one, I tried upgrading the app on both devices, tried to turn it off on both devices and turn it on. nothing helped. i can't figure how to reconnect to my laptop

Another repro, this time with Cloudflare WARP rather than a corporate ZTNA client. Desktop Alpha 0.0.45 on macOS 26.6 (Mac mini, arm64), T3 Connect hosted by the desktop app, no background service. Homebrew cloudflared 2026.10.0.
Timeline:
- WARP (Warp mode, MASQUE, always-on) was already connected when the app started. cloudflared registered fine and served requests through the relay for ~4h.
- WARP's daemon did a disconnected → connecting → connected cycle that took ~1.3s. From that moment no request reached the host.
- 21h later the connector child is still alive,
/readyreturns{"status":503,"readyConnections":0}, and the relay hostname answers Cloudflare 530 / error 1033. Other clients getendpoint_request_failed. - Metrics over the process lifetime:
tunnel_register_success16,tunnel_register_fail{error="server_error"}4352 and climbing (+2 per 45s),{error="dup_edge_conn"}75,quic_client_closed_connections4539. So it is retrying forever, not stuck on a dead dial. - Nothing from cloudflared is on disk ([Bug]: Desktop-hosted T3 Connect keeps no record of the managed cloudflared's output #16331), so I can't see the edge's exact rejection text. QUIC gives an opaque control-stream error anyway (see the analysis in fix(connect): recover a reaped tunnel the connector cannot register #16402), which is why
isRejectedRelayClientTunnelOutputnever fires.
Contributing factor: the edge ranges
198.41.192.0/24and198.41.200.0/24route through WARP's utun and are not in its split-tunnel exclude list.Same conclusion as above: a live-but-unregistered connector needs a readiness-based restart. Polling
/readyas proposed would have caught this within minutes.
Before submitting
Area
apps/server
Steps to reproduce
region1.v2.argotunnel.comresolves to virtual addresses from the client's own range that only route through itsutuninterface, instead of Cloudflare's published198.41.192.xaddresses.region1/region2.v2.argotunnel.comnow resolve to different virtual addresses, and the old ones no longer route anywhere.I don't have a VPN-free repro yet. The core condition is that the addresses cloudflared resolved at startup stop working while DNS starts returning new ones.
Expected behavior
When the managed relay client is still running but has had no registered tunnel connection for a sustained period, T3 should replace it. A fresh cloudflared re-resolves the edge and reconnects. Doing that by hand restored the tunnel within 3 seconds (see Workaround).
Proposed semantics for maintainers to decide before any PR:
--metrics 127.0.0.1:0, read the bound address from itsStarting metrics server on <addr>/metricsline, and poll/ready. It returns 200 whenreadyConnections > 0and 503 otherwise. If that line never appears, do nothing, which keeps today's behavior.superviseConnectorrestart it. Double the threshold after each restart that doesn't reconnect, up to about 30 minutes. Reset it once/readyreturns 200.runningbefore registration ([Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447). Also the stale QUIC MTU case (T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897), where/readystill reports 4 connections.Actual behavior
Addresses are replaced with documentation ranges:
192.0.2.xstands for the virtual addresses cloudflared resolved at startup, and203.0.113.xfor the ones DNS returned after the network change.198.41.192.7is a real, published Cloudflare edge address.The tunnel stayed down for 2 h 47 min, until I restarted cloudflared by hand:
192.0.2.120–192.0.2.139). The log has 461failed to dial to edge with quic: timeout: no recent network activityerrors and no successful registration.curl http://127.0.0.1:20241/readyreturned{"status":503,"readyConnections":0,...}.t3 connect statusreports saved setup, not live state, so nothing on the host showed the outage either.The network was not blocking Cloudflare Tunnel. I checked from the same Mac at the same time, using a raw QUIC version-negotiation probe and an
openssl s_clienthandshake:203.0.113.143)198.41.192.7192.0.2.121)Why it never recovers:
NewSupervisor(supervisor.go).edgediscovery.Edgenever refreshes them (edgediscovery.go), so it only rotates through the startup set. With normal DNS that set is Cloudflare's stable published IPs, so this rarely shows. With virtual-IP DNS the set can go stale while the process runs.superviseConnector), or after 4 rejected registrations (observeConnectorOutput). Transport errors are only logged, so a live connector with zero connections is kept forever.Any DNS that hands out session-scoped virtual IPs, such as other ZTNA clients or fake-IP modes in proxy tools, should hit the same problem. I have only reproduced it with one corporate VPN client.
Related: #7447 (a spawned but unreachable tunnel is reported as
running) and #15897 (stale QUIC MTU after joining a VPN). All three treat a live cloudflared process as a working tunnel. This report covers a tunnel that worked and then went stale.Impact
Blocks work completely
Version or commit
T3 Code desktop 0.0.45 and
t3 service0.0.45. Code links point to main @ f5eb250.Environment
macOS 27.0 (Apple Silicon); cloudflared 2026.6.0 from Homebrew (picked up from PATH); corporate full-tunnel ZTNA/VPN client;
Logs or stack traces
Screenshots, recordings, or supporting files
No response
Workaround
Kill only the managed cloudflared process:
kill <pid>, where<pid>comes frompgrep -P <t3 serve pid> cloudflared. T3 restarts it right away, the new process re-resolves the edge, and all 4 connections registered within 3 seconds.t3 service restartalso works, but it restarts the whole server.I'm happy to send a focused PR once the intended behavior is agreed.