Repository navigation
[Bug]: T3 Connect never recovers a reaped tunnel: cloudflared's rejection is invisible over QUIC and reads Tunnel not found over HTTP/2 #16399
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the cloudflared traces and the matcher write-up, @berghtho. They made this much easier to follow. This looks like its own recovery gap, separate from #16258.
What I found
- A reaped tunnel with an unchanged origin never asks Cloudflare.
reconcileOriginininfra/relay/src/environments/ManagedEndpointProvider.tsreturnsreadywhen the stored host and port still match, so startup registration keeps the old token. That matchesdocs/internals/t3-connect.md: registration only touches Cloudflare when the origin changed. - Replacement happens only if the child exits (
superviseConnectorqueues a recovery request) orisRejectedRelayClientTunnelOutputmatches four times. The matcher inapps/server/src/cloud/ManagedEndpointRuntime.tsrequiresRegister tunnel error from server sideplus one ofFailed to get tunnel,Record for tunnel not found, orInvalid tunnel secret. Killing the process works because the exit path doesn't consult the matcher; restarting the app doesn't, because the same stored token is started again. - On the pinned cloudflared
2026.5.2, the QUIC path can't hit that matcher.quicConnection.Servereturns&ControlStreamError{}(control stream encountered a failure while serving), and the supervisor only printsRegister tunnel error from server sideforServerRegisterTunnelError, so this falls through toServe tunnel error. The managed child is started without--protocol, so it stays on QUIC. Your debug run matches that source. - HTTP/2 does print
Register tunnel error from server side, butUnauthorized: Tunnel not foundfails the reason regex. So forcing HTTP/2 (T3 Connect: let users force the managed cloudflared to HTTP/2 — on QUIC-blocking corporate networks the tunnel stays down because cloudflared won't fall back #16344) alone wouldn't make the matcher fire. - [Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258 is a live connector with no route to the edge (
/readystays 503 without a registration rejection), while this is a deleted tunnel whose rejection text isn't recognized. [Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447 and [Bug]: Desktop-hosted T3 Connect keeps no record of the managed cloudflared's output #16331 are the visibility problems you noted, not this recovery miss.
Likely fix area
Extending the matcher is one narrow option, with a few details worth keeping in mind:
- On HTTP/2,
Tunnel not foundcould join the existing reasons, still requiringRegister tunnel error from server side. - On QUIC, counting only
Serve tunnel errorwithcontrol stream encountered a failure while serving(not thefailed to serve tunnel connectionline that repeats it for the same attempt) would avoid double counting.ControlStreamErroris also what a graceful full-process unregister returns, so the existing four-failure threshold and two-minute cooldown would still matter. - One open question: every unmatched
ERRline is normally logged asRelay client reported a transport warning, and that message is missing from your trace. The--loglevel debughand-run shows what cloudflared prints, but not that the managed--loglevel infochild's stderr actually reachedobserveConnectorOutput. A matcher change would only recover the host if those lines are delivered, so that path may need a look too.
A maintainer will decide on the fix direction.
- A reaped tunnel with an unchanged origin never asks Cloudflare.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 6, 2026 - added a commit that references this issue
on Oct 6, 2026 Also observed on macOS (Apple Silicon), T3 Code nightly
0.0.46-nightly.20261005, with the managed cloudflared2026.5.2, connecting from the iPhone TestFlight app.Live Activities / Dynamic Island and push notifications continued working, but opening a thread failed with
endpoint_request_failed.During diagnosis, cloudflared returned
Unauthorized: Tunnel not found, and automatic recovery did not trigger. Restarting the desktop app had reused the same unavailable tunnel.Terminating the managed cloudflared process caused T3 to provision a replacement within about 4 seconds. The replacement registered 4 connections, and an HTTP request to its public endpoint returned 200 and increased the tunnel request counter.
After this workaround, I confirmed that the iPhone app connects successfully and I can open my chats again.
Confirmed on macOS (Apple Silicon) + Windows 11, both on
0.0.46-nightly.20261005.2702, pinned cloudflared2026.5.2, desktop app as host on both machines.Additional trigger —
releaseManagedTunnelOnShutdownmakes this hit on every clean quit, not just reaped-idle tunnels. The server deletes the provisioned Cloudflare tunnel on shutdown (unless an update restart is pending), butcloud-endpoint-runtime-configwith the old tunnelId/token survives. On the next launch, the recovery path registers the stored tunnelId — reports ready without touching Cloudflare — and the managed cloudflared comes up with a dead token:cloudflared_tunnel_tunnel_register_fail{error="server_error",rpcName="registerConnection"} = 14 and climbing cloudflared_tunnel_ha_connections 1 cloudflared_tunnel_total_requests 0Relay then reports
environment_endpoint_unavailable/endpoint_request_failedto every client ("Reconnecting: Relay could not reach the environment endpoint").Observed sequence on macOS:
t3 connect unlink+t3 connect+ app restart → fresh provision, tunnel healthy (ha_connections 4, requests flowing), remote clients connect fine.- Clean quit → relaunch → new cloudflared spawned with the same runtime config (
cloud-endpoint-runtime-configmtime unchanged from the previous provision;cloud-endpoint-confirmed-originrewritten this startup) →registerConnection server_errorloop, identical to the reaped-tunnel case.reconcileDesiredLinkIfStillDesiredeither didn't run or didn't reachapplyCloudRelayConfig— the config file was never rewritten.
So the release-on-shutdown + stale-runtime-config path makes this deterministic per clean relaunch, not only after hours of downtime. Symmetric failure on both machines: each showed
endpoint_request_failedfor the other until its own link was re-provisioned.Workaround that worked here:
npx t3@nightly connect unlink→npx t3@nightly connect→ full app restart on the affected host (re-provisions a fresh tunnel under the same environment). The lighterkill cloudflaredworkaround from this issue also holds.Also seen on Windows 11 Pro (10.0.26200, x64). The host is the desktop app
0.0.45-nightly.20261002.2572with the pinned cloudflared2026.5.2. The clients are the desktop app on a second Windows laptop and the Android app 1.4.0. So the bug is present from the 2 October nightly at the latest.The trigger was a clean OS restart, the same path that @stilak12 describes. The host server ran without a stop for 14 hours. The last good request from a client through the relay was at 01:55 local time. The OS restarted at 03:37. The app started again at 09:30. After that, both clients showed
Reconnecting: Relay could not reach the environment endpoint (endpoint_request_failed).State 10 minutes after startup:
GET 127.0.0.1:20241/ready -> 503 {"readyConnections":0} cloudflared_tunnel_tunnel_register_fail{error="server_error",rpcName="registerConnection"} 55 cloudflared_rpc_client_latency_secs_bucket{...register_connection,le="0.05"} 55 cloudflared_tunnel_ha_connections 1 cloudflared_tunnel_total_requests 0 GET https://<endpoint>.t3coderelay.com/.well-known/t3/environment -> 530Each registration failed in less than 50 ms. The server trace has no recovery span after the startup
activateManagedTunnel, so the four-failure matcher never fired.Two more details:
- The managed cloudflared output is not on disk anywhere. The desktop app writes backend child output to
server-child.logonly when the child fails, and that file was last written 5 days earlier. The trace file has noRelay client reported a transport warningentries either. This agrees with the open question in the triage comment about whether those lines reachobserveConnectorOutput. On this host, it was not possible to see the rejection text without a hand run. - The workaround from this issue worked. I stopped the managed cloudflared process, and within about 10 seconds a replacement started.
/readythen reportedreadyConnections: 4and the public endpoint returned 200. The pairing and the hostname stayed the same.
Found during a
t3 triagesession by Claude Code with Claude Opus 5.5.- The managed cloudflared output is not on disk anywhere. The desktop app writes backend child output to
Adding a data point on why the stored runtime config survives a clean quit, which @stilak12 described above. The desktop app gives the backend 2 seconds after SIGTERM before it sends SIGKILL (
DEFAULT_BACKEND_TERMINATE_GRACEinapps/desktop/src/backend/DesktopBackendManager.ts), while the server allowsreleaseManagedTunnelOnShutdown10 seconds (apps/server/src/server.ts). Release stops the connector, sends the relay DELETE, and removescloud-endpoint-runtime-configonly after anok: trueresponse.Traces from five quits on one macOS host (Nightly from Oct 6, pinned cloudflared):
- Four releases finished in 1.37 to 1.79 seconds, 0.2 to 0.6 seconds before the SIGKILL. Each next launch ran the full link reconcile and connected.
- In the fifth, release stopped the connector within 3 ms, and the server was SIGKILLed 2.04 seconds after quit began, with the release span never completed. The next launch took the recovery path with the stored tunnel, registration came back ready, and cloudflared sat at zero connections for 46 minutes until I killed it. Recovery then handed back a different tunnel ID.
So whenever the DELETE round trip runs past about 2 seconds, the relay most likely deletes the tunnel while the local config survives, and the next start reuses it. #16402 would let the host recover from that, but the race seems worth closing too, for example with a backend terminate grace longer than the release budget, or by marking the stored config stale before sending the DELETE so the next start reconciles.
Sent by Mike's agent (Claude Opus 5.5 in T3 Code)
Follow-up: #16649 closes this race on the relay side (registration now checks
Cloudflare, so the next start gets a replacement right away), and #16648 adds
the three-minute readiness fallback, so the fix suggestions at the end of my
comment above are moot. The traces there match the race #16649 describes.
Sent by Mike's agent (Claude Opus 5.5 in T3 Code)
Before submitting
Area
apps/server
Steps to reproduce
Tunnel not foundfor the stored tunnel ID.registerwith the stored tunnel ID) reports ready, because registration does not touch Cloudflare, and the stored connector config starts cloudflared with the old token. I restarted the app three times; every start behaved the same.Expected behavior
docs/internals/t3-connect.md: "If the connector exits, orcloudflaredreports repeated tunnel rejections, the host asks the relay for a replacement, at most once every two minutes." After the fourth rejected registration attempt (about 30 seconds with cloudflared's backoff) the host should request a replacement tunnel and the environment should come back without anyone touching the machine.Actual behavior
The managed cloudflared stays alive indefinitely with zero registered connections. Its
/readyendpoint returns 503,quic_client_total_connectionsequalsquic_client_closed_connections, and the server trace never gets pastRelay client process started; waiting for tunnel connection. NoRelay client tunnel connection registered, noRelay client tunnel was rejected; requesting recovery, not evenRelay client reported a transport warning. The environment stays unreachable from other devices until someone kills the connector by hand.Two gaps in
isRejectedRelayClientTunnelOutput(apps/server/src/cloud/ManagedEndpointRuntime.ts) keep the recovery from firing:quicConnection.Serve(connection/quic_connection.go) runs the control stream in an errgroup and returns the typeless&ControlStreamError{}for any control-stream failure, so theServerRegisterTunnelErrornever reaches the supervisor's type switch insupervisor/tunnel.goand theRegister tunnel error from server sideline is never logged. The supervisor falls into its default branch and logsServe tunnel error error="control stream encountered a failure while serving". (Thefailed to serve the control streamlog insideServecannot fire either: theerrinif err := q.serveControlStream(...)shadows the outererr, which is nil.) The matcher requires theRegister tunnel error from server sideprefix, so nothing counts. cloudflared also does not fall back to HTTP/2 on registration failures, only on dial and transport errors, so the connector never leaves QUIC.Unauthorized: Tunnel not found. The matcher acceptsFailed to get tunnel,Record for tunnel not foundandInvalid tunnel secret. fix(connect): remove tunnels after hosts go offline #9386 verified the matcher against a missing tunnel ID, which yieldsFailed to get tunnel; a tunnel that existed and was deleted yields a different message.Impact
Blocks work completely
Version or commit
Desktop Alpha 0.0.45 (
service.versionin the trace). The matcher is unchanged onmain@ 4ae976d.Environment
Windows 11 Pro 10.0.26340, desktop app as host, pinned cloudflared 2026.5.2, home network without proxy; QUIC and HTTP/2 to
region1/region2.v2.argotunnel.com:7844both pass cloudflared's pre-checks. Second desktop on another network could not connect.Logs or stack traces
Workaround
Kill the managed
cloudflaredprocess. The supervisor logsRelay client exited; restarting, queues a recovery request, the relay provisions a replacement under the same allocation (new tunnel ID), and the new connector registers four connections within seconds;/readyreturns 200. Restarting the desktop app does not help, because startup registration with the stored tunnel ID reports ready without checking Cloudflare.Related
/readywatchdog. This report is about the rejection matcher: the tunnel is gone and cloudflared does say so, just not in the form the matcher expects.runningbefore any registration) and [Bug]: Desktop-hosted T3 Connect keeps no record of the managed cloudflared's output #16331 (no record of the connector's output on desktop hosts) made this harder to see; neither covers the missing recovery.Proposed fix
Extend
isRejectedRelayClientTunnelOutputso that (a)Tunnel not foundcounts alongside the existing reasons and (b) the QUIC supervisor lineServe tunnel error error="control stream encountered a failure while serving"counts as a failed registration attempt. The control stream is the first error only while a connection is still registering; a connection lost after registration fails its stream listener or datagram handler first. The existing threshold of four failures and the two-minute recovery cooldown bound the cost of a transient failure that happens to look the same. Only theServe tunnel errorline should count, not thefailed to serve tunnel connectionline that repeats the same error for the same attempt.Investigated on the affected machine and written with Claude Fable 5.1 in Claude Code.