Skip to content

[Bug]: T3 Connect never recovers a reaped tunnel: cloudflared's rejection is invisible over QUIC and reads Tunnel not found over HTTP/2 #16399

Description

@berghtho

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Host an environment with the desktop app (Windows 11, Desktop Alpha 0.0.45) and link it with T3 Connect. The managed connector runs the pinned cloudflared 2026.5.2 over QUIC (all connectivity pre-checks pass).
  2. Leave the host offline long enough for the relay's tunnel reaper to delete the managed tunnel. Here the PC was off for a few hours; the relay now reports Tunnel not found for the stored tunnel ID.
  3. Start the desktop app again. Startup registration (register with the stored tunnel ID) reports ready, because registration does not touch Cloudflare, and the stored connector config starts cloudflared with the old token. I restarted the app three times; every start behaved the same.
  4. Open the environment from another device (second desktop or mobile).

Expected behavior

docs/internals/t3-connect.md: "If the connector exits, or cloudflared reports repeated tunnel rejections, the host asks the relay for a replacement, at most once every two minutes." After the fourth rejected registration attempt (about 30 seconds with cloudflared's backoff) the host should request a replacement tunnel and the environment should come back without anyone touching the machine.

Actual behavior

The managed cloudflared stays alive indefinitely with zero registered connections. Its /ready endpoint returns 503, quic_client_total_connections equals quic_client_closed_connections, and the server trace never gets past Relay client process started; waiting for tunnel connection. No Relay client tunnel connection registered, no Relay client tunnel was rejected; requesting recovery, not even Relay client reported a transport warning. The environment stays unreachable from other devices until someone kills the connector by hand.

Two gaps in isRejectedRelayClientTunnelOutput (apps/server/src/cloud/ManagedEndpointRuntime.ts) keep the recovery from firing:

  1. Over QUIC, the default and auto-selected protocol, cloudflared never prints the rejection. In cloudflared 2026.5.2, quicConnection.Serve (connection/quic_connection.go) runs the control stream in an errgroup and returns the typeless &ControlStreamError{} for any control-stream failure, so the ServerRegisterTunnelError never reaches the supervisor's type switch in supervisor/tunnel.go and the Register tunnel error from server side line is never logged. The supervisor falls into its default branch and logs Serve tunnel error error="control stream encountered a failure while serving". (The failed to serve the control stream log inside Serve cannot fire either: the err in if err := q.serveControlStream(...) shadows the outer err, which is nil.) The matcher requires the Register tunnel error from server side prefix, so nothing counts. cloudflared also does not fall back to HTTP/2 on registration failures, only on dial and transport errors, so the connector never leaves QUIC.
  2. Over HTTP/2, the reason for a deleted tunnel is Unauthorized: Tunnel not found. The matcher accepts Failed to get tunnel, Record for tunnel not found and Invalid tunnel secret. fix(connect): remove tunnels after hosts go offline #9386 verified the matcher against a missing tunnel ID, which yields Failed to get tunnel; a tunnel that existed and was deleted yields a different message.

Impact

Blocks work completely

Version or commit

Desktop Alpha 0.0.45 (service.version in the trace). The matcher is unchanged on main @ 4ae976d.

Environment

Windows 11 Pro 10.0.26340, desktop app as host, pinned cloudflared 2026.5.2, home network without proxy; QUIC and HTTP/2 to region1/region2.v2.argotunnel.com:7844 both pass cloudflared's pre-checks. Second desktop on another network could not connect.

Logs or stack traces

# server.trace.ndjson: three connector starts, nothing after this line for any of them
"Relay client process started; waiting for tunnel connection", { "pid": 13996, "tunnelId": "ff96ecf3-...", "tunnelName": "t3coderelay-managedendpoint-prod-..." }

# managed cloudflared, metrics endpoint 127.0.0.1:20241 after ~8 minutes
GET /ready -> 503
cloudflared_tunnel_total_requests 0
quic_client_total_connections 5
quic_client_closed_connections 5

# pinned cloudflared run by hand with the same TUNNEL_TOKEN, default protocol (quic), --loglevel debug
DBG Registering tunnel connection connIndex=0 event=0 ip=198.41.200.43 protocol=quic
ERR failed to run the datagram handler error="context canceled" connIndex=0 event=0 ip=198.41.200.43
ERR failed to serve tunnel connection error="control stream encountered a failure while serving" connIndex=0 event=0 ip=198.41.200.43
ERR Serve tunnel error error="control stream encountered a failure while serving" connIndex=0 event=0 ip=198.41.200.43
INF Retrying connection in up to 2s connIndex=0 event=0 ip=198.41.200.43
(repeats with 4s, 8s, 16s)

# same token with --protocol http2: the edge's reason becomes visible
DBG Registering tunnel connection connIndex=0 event=0 ip=198.41.200.43 protocol=http2
ERR failed to serve incoming request error="Unauthorized: Tunnel not found"
ERR Register tunnel error from server side error="Unauthorized: Tunnel not found" connIndex=0 event=0 ip=198.41.200.43
INF Retrying connection in up to 2s connIndex=0 event=0 ip=198.41.200.43

Workaround

Kill the managed cloudflared process. The supervisor logs Relay client exited; restarting, queues a recovery request, the relay provisions a replacement under the same allocation (new tunnel ID), and the new connector registers four connections within seconds; /ready returns 200. Restarting the desktop app does not help, because startup registration with the stored tunnel ID reports ready without checking Cloudflare.

Related

Proposed fix

Extend isRejectedRelayClientTunnelOutput so that (a) Tunnel not found counts alongside the existing reasons and (b) the QUIC supervisor line Serve tunnel error error="control stream encountered a failure while serving" counts as a failed registration attempt. The control stream is the first error only while a connection is still registering; a connection lost after registration fails its stream listener or datagram handler first. The existing threshold of four failures and the two-minute recovery cooldown bound the cost of a transient failure that happens to look the same. Only the Serve tunnel error line should count, not the failed to serve tunnel connection line that repeats the same error for the same attempt.

Investigated on the affected machine and written with Claude Fable 5.1 in Claude Code.

Activity

  1. juliusmarminge commented on Oct 6, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the cloudflared traces and the matcher write-up, @berghtho. They made this much easier to follow. This looks like its own recovery gap, separate from #16258.

    What I found

    Likely fix area

    Extending the matcher is one narrow option, with a few details worth keeping in mind:

    • On HTTP/2, Tunnel not found could join the existing reasons, still requiring Register tunnel error from server side.
    • On QUIC, counting only Serve tunnel error with control stream encountered a failure while serving (not the failed to serve tunnel connection line that repeats it for the same attempt) would avoid double counting. ControlStreamError is also what a graceful full-process unregister returns, so the existing four-failure threshold and two-minute cooldown would still matter.
    • One open question: every unmatched ERR line is normally logged as Relay client reported a transport warning, and that message is missing from your trace. The --loglevel debug hand-run shows what cloudflared prints, but not that the managed --loglevel info child's stderr actually reached observeConnectorOutput. A matcher change would only recover the host if those lines are delivered, so that path may need a look too.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 6, 2026
  3. added a commit that references this issue on Oct 6, 2026
    34e8ddf
  4. Heggtor commented on Oct 6, 2026

    @Heggtor

    Also observed on macOS (Apple Silicon), T3 Code nightly 0.0.46-nightly.20261005, with the managed cloudflared 2026.5.2, connecting from the iPhone TestFlight app.

    Live Activities / Dynamic Island and push notifications continued working, but opening a thread failed with endpoint_request_failed.

    During diagnosis, cloudflared returned Unauthorized: Tunnel not found, and automatic recovery did not trigger. Restarting the desktop app had reused the same unavailable tunnel.

    Terminating the managed cloudflared process caused T3 to provision a replacement within about 4 seconds. The replacement registered 4 connections, and an HTTP request to its public endpoint returned 200 and increased the tunnel request counter.

    After this workaround, I confirmed that the iPhone app connects successfully and I can open my chats again.

  5. stilak12 commented on Oct 6, 2026

    @stilak12

    Confirmed on macOS (Apple Silicon) + Windows 11, both on 0.0.46-nightly.20261005.2702, pinned cloudflared 2026.5.2, desktop app as host on both machines.

    Additional trigger — releaseManagedTunnelOnShutdown makes this hit on every clean quit, not just reaped-idle tunnels. The server deletes the provisioned Cloudflare tunnel on shutdown (unless an update restart is pending), but cloud-endpoint-runtime-config with the old tunnelId/token survives. On the next launch, the recovery path registers the stored tunnelId — reports ready without touching Cloudflare — and the managed cloudflared comes up with a dead token:

    cloudflared_tunnel_tunnel_register_fail{error="server_error",rpcName="registerConnection"} = 14 and climbing
    cloudflared_tunnel_ha_connections 1
    cloudflared_tunnel_total_requests 0
    

    Relay then reports environment_endpoint_unavailable / endpoint_request_failed to every client ("Reconnecting: Relay could not reach the environment endpoint").

    Observed sequence on macOS:

    1. t3 connect unlink + t3 connect + app restart → fresh provision, tunnel healthy (ha_connections 4, requests flowing), remote clients connect fine.
    2. Clean quit → relaunch → new cloudflared spawned with the same runtime config (cloud-endpoint-runtime-config mtime unchanged from the previous provision; cloud-endpoint-confirmed-origin rewritten this startup) → registerConnection server_error loop, identical to the reaped-tunnel case. reconcileDesiredLinkIfStillDesired either didn't run or didn't reach applyCloudRelayConfig — the config file was never rewritten.

    So the release-on-shutdown + stale-runtime-config path makes this deterministic per clean relaunch, not only after hours of downtime. Symmetric failure on both machines: each showed endpoint_request_failed for the other until its own link was re-provisioned.

    Workaround that worked here: npx t3@nightly connect unlink → npx t3@nightly connect → full app restart on the affected host (re-provisions a fresh tunnel under the same environment). The lighter kill cloudflared workaround from this issue also holds.

  6. Hazzajenko commented on Oct 6, 2026

    @Hazzajenko

    Also seen on Windows 11 Pro (10.0.26200, x64). The host is the desktop app 0.0.45-nightly.20261002.2572 with the pinned cloudflared 2026.5.2. The clients are the desktop app on a second Windows laptop and the Android app 1.4.0. So the bug is present from the 2 October nightly at the latest.

    The trigger was a clean OS restart, the same path that @stilak12 describes. The host server ran without a stop for 14 hours. The last good request from a client through the relay was at 01:55 local time. The OS restarted at 03:37. The app started again at 09:30. After that, both clients showed Reconnecting: Relay could not reach the environment endpoint (endpoint_request_failed).

    State 10 minutes after startup:

    GET 127.0.0.1:20241/ready -> 503 {"readyConnections":0}
    cloudflared_tunnel_tunnel_register_fail{error="server_error",rpcName="registerConnection"} 55
    cloudflared_rpc_client_latency_secs_bucket{...register_connection,le="0.05"} 55
    cloudflared_tunnel_ha_connections 1
    cloudflared_tunnel_total_requests 0
    GET https://<endpoint>.t3coderelay.com/.well-known/t3/environment -> 530
    

    Each registration failed in less than 50 ms. The server trace has no recovery span after the startup activateManagedTunnel, so the four-failure matcher never fired.

    Two more details:

    • The managed cloudflared output is not on disk anywhere. The desktop app writes backend child output to server-child.log only when the child fails, and that file was last written 5 days earlier. The trace file has no Relay client reported a transport warning entries either. This agrees with the open question in the triage comment about whether those lines reach observeConnectorOutput. On this host, it was not possible to see the rejection text without a hand run.
    • The workaround from this issue worked. I stopped the managed cloudflared process, and within about 10 seconds a replacement started. /ready then reported readyConnections: 4 and the public endpoint returned 200. The pairing and the hostname stayed the same.

    Found during a t3 triage session by Claude Code with Claude Opus 5.5.

  7. mwolson commented on Oct 7, 2026

    @mwolson
    Contributor

    Adding a data point on why the stored runtime config survives a clean quit, which @stilak12 described above. The desktop app gives the backend 2 seconds after SIGTERM before it sends SIGKILL (DEFAULT_BACKEND_TERMINATE_GRACE in apps/desktop/src/backend/DesktopBackendManager.ts), while the server allows releaseManagedTunnelOnShutdown 10 seconds (apps/server/src/server.ts). Release stops the connector, sends the relay DELETE, and removes cloud-endpoint-runtime-config only after an ok: true response.

    Traces from five quits on one macOS host (Nightly from Oct 6, pinned cloudflared):

    • Four releases finished in 1.37 to 1.79 seconds, 0.2 to 0.6 seconds before the SIGKILL. Each next launch ran the full link reconcile and connected.
    • In the fifth, release stopped the connector within 3 ms, and the server was SIGKILLed 2.04 seconds after quit began, with the release span never completed. The next launch took the recovery path with the stored tunnel, registration came back ready, and cloudflared sat at zero connections for 46 minutes until I killed it. Recovery then handed back a different tunnel ID.

    So whenever the DELETE round trip runs past about 2 seconds, the relay most likely deletes the tunnel while the local config survives, and the next start reuses it. #16402 would let the host recover from that, but the race seems worth closing too, for example with a backend terminate grace longer than the release budget, or by marking the stored config stale before sending the DELETE so the next start reconciles.


    Sent by Mike's agent (Claude Opus 5.5 in T3 Code)

  8. mwolson commented on Oct 7, 2026

    @mwolson
    Contributor

    Follow-up: #16649 closes this race on the relay side (registration now checks
    Cloudflare, so the next start gets a replacement right away), and #16648 adds
    the three-minute readiness fallback, so the fix suggestions at the end of my
    comment above are moot. The traces there match the race #16649 describes.


    Sent by Mike's agent (Claude Opus 5.5 in T3 Code)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions