Skip to content

T3 Connect: let users force the managed cloudflared to HTTP/2 — on QUIC-blocking corporate networks the tunnel stays down because cloudflared won't fall back #16344

Description

@mats16

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Summary

On networks that drop QUIC (UDP 7844) but allow TCP 7844, T3 Connect goes down for long periods. The managed cloudflared keeps dialing QUIC instead of using HTTP/2, which works on the same networks. This is common on corporate networks: full-tunnel VPNs that fall back from IPsec to an SSL/TLS (TCP 443) tunnel often stop passing QUIC, and some office and tethered networks behave the same way.

T3 always spawns cloudflared tunnel --no-autoupdate --loglevel info --output default run with the default --protocol auto, and it offers no way to choose the transport. A user who knows the network blocks QUIC can't do anything about it. A watchdog that restarts cloudflared (as proposed in #16258) helps less than you'd expect here, for the reasons below.

Why cloudflared doesn't recover on its own (cloudflared 2026.5.2)

  1. Fallback is permanently suppressed once QUIC has ever connected. In supervisor/tunnel.go, Serve() returns before choosing a fallback when e.tracker.HasConnectedWith(e.config.ProtocolSelector.Current()) is true. In tunnelstate/conntracker.go, HasConnectedWith only checks ci.Protocol == protocol and ignores IsConnected. Disconnect events only set IsConnected = false, so the protocol record stays. A connector that used QUIC once, for example at home, never falls back to HTTP/2 after the laptop moves to a QUIC-blocking network. It sits at readyConnections: 0 until it is restarted.
  2. Even a fresh process falls back slowly. For EdgeQuicDialError, IpAddrFallback.ShouldGetNewAddress only reports a fallback-worthy connectivity error after maxRetries edge-address rotations per connection, and the backoff between attempts grows exponentially. In our logs the retry gaps grew from about 20 s to 2 min to 4.5 min. A fresh connector that had never connected over QUIC took about 18 minutes to fall back.
  3. The fallback doesn't stick. One long-lived connector registered over http2 at startup, then registered over quic again on a later reconnect within the same process. It then hit (1) on the next QUIC-blocking network.

So restarting cloudflared resets (1), but on a QUIC-blocking network the new process still starts with QUIC and needs many minutes to fall back.

Steps to reproduce

  1. On a network where QUIC works, start the desktop app with T3 Connect on. Confirm a phone can connect. cloudflared registers 4 connections with protocol=quic.
  2. Move the host to a network that drops UDP 7844 but allows TCP 7844. In our case this was a corporate full-tunnel VPN running in SSL/TLS fallback mode, on both office Wi-Fi and a phone hotspot.
  3. Observe curl -s http://127.0.0.1:<metrics-port>/ready. It returns {"status":503,"readyConnections":0,...} and stays there. The log shows repeated failed to dial to edge with quic: timeout: no recent network activity or handshake did not complete in time.
  4. Try to connect from the phone. It shows Remote environment endpoint https://<relay-host>/oauth/token timed out after 10000ms or a similar error.

Expected behavior

Actual behavior

What we measured on the same host (macOS, Apple silicon):

Host network VPN transport QUIC to tunnel edge (UDP 7844) TCP 7844 T3 Connect
Home broadband IPsec works works fine
Phone hotspot IPsec failed → SSL/TLS dropped or degraded works flaky: 22 of 29 relay-minted pairing credentials were never exchanged, and the phone timed out on /oauth/token
Office Wi-Fi IPsec failed → SSL/TLS no reply works down: 0 ready connections for ~45 min, until the app was restarted with HTTP/2 forced
Office Wi-Fi, socket bound to the Wi-Fi interface (bypassing the VPN) – version-negotiation reply in 6–7 ms works –
  • QUIC was checked with a raw QUIC version-negotiation probe (reserved version 0x0a0a0a0a). TCP 7844 was checked with nc -z.
  • On the blocked path, the connector's metrics showed quic_client_total_connections climbing (45 → 115) with quic_client_latest_rtt 0 and no received bytes.
  • After forcing HTTP/2 (see Workaround), the connector registered 4 connections with protocol=http2 within a minute on the same office network. /.well-known/t3/environment through the relay hostname answered 200 in about 0.1 s.

Impact

Major degradation or frequent failure

Version or commit

Desktop 0.0.46-nightly.20261005.2702 (macOS "T3 Code (Nightly)")

Environment

macOS 26.6.2 (Apple silicon), managed cloudflared 2026.5.2, iOS app (build 105), corporate full-tunnel VPN with IPsec → SSL/TLS fallback

Logs or stack traces

ERR Failed to dial a quic connection error="failed to dial to edge with quic: timeout: no recent network activity" connIndex=0 event=0 ip=198.41.200.113
WRN Failed to dial a quic connection error="failed to dial to edge with quic: timeout: handshake did not complete in time" connIndex=0 event=0 ip=198.41.200.113
# /ready on the stuck connector:
{"status":503,"readyConnections":0,"connectorId":"<redacted>"}
# after forcing http2, same network:
INF Registered tunnel connection connIndex=0 ... location=<redacted> protocol=http2

Workaround

T3 spawns cloudflared with env: { ...process.env, TUNNEL_TOKEN }, so cloudflared's own env var for --protocol is inherited:

launchctl setenv TUNNEL_TRANSPORT_PROTOCOL http2
# then quit T3 Code and relaunch it from the Dock or Finder

This is undocumented, applies to every app launched afterwards, and is lost on reboot or logout. Turning the VPN off also works, but that's often not an option on a work machine.

Related

Activity

  1. juliusmarminge commented on Oct 6, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the measurements and the detailed cloudflared walkthrough, @mats16! This lines up with what we see in the code, and it's separate from the nearby tunnel issues.

    What I found

    • The managed connector starts with a fixed argv and no transport flag. CloudManagedEndpointRuntime in apps/server/src/cloud/ManagedEndpointRuntime.ts spawns cloudflared tunnel --no-autoupdate --loglevel info --output default run and adds only TUNNEL_TOKEN on top of process.env. There's no --protocol, server flag, or setting. The pinned binary is cloudflared 2026.5.2 (CLOUDFLARED_VERSION in packages/shared/src/relayClient.ts), and running means the child is alive, not that a tunnel connection registered.
    • TUNNEL_TRANSPORT_PROTOCOL is cloudflared's own env var for its hidden --protocol flag. Because the spawn spreads process.env, your launchctl setenv workaround reaches the child after a relaunch. That works through inheritance rather than a supported setting.
    • On auto with a tunnel token, cloudflared 2026.5.2 starts on QUIC and treats HTTP/2 as a fallback. The three behaviors you described are all in that tag:
      • tunnelstate.ConnTracker.HasConnectedWith matches on protocol only, and a disconnect leaves the protocol record in place, so one QUIC registration suppresses HTTP/2 fallback for the rest of the process.
      • A fresh process can't fall back until IpAddrFallback has rotated edge addresses MaxEdgeAddrRetries times (default 8). With a 5s QUIC handshake idle timeout and backoff capped around 30s, that's a multi-minute stall. Startup connectivity pre-checks are diagnostic only.
      • A successful registration calls protocolFallback.reset(), which clears inFallback but leaves the selector on QUIC, so the next retry goes back to QUIC ("registered http2, then quic again").
    • superviseConnector only restarts the child after it exits. A zero-connection watchdog that respawns with the same auto args (like the one proposed on [Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258) would reset the QUIC memory and start on QUIC again, and could restart before the edge-address rotations finish.
    • Related but distinct: [Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258 (stuck at zero connections with no restart), T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897 (stale QUIC MTU; /ready stays 200), [Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447 (status says running before the first registration).

    Likely fix area

    Some options:

    • Pass a transport choice (auto default, http2, quic) through to the spawn as --protocol, reachable from both t3 serve and the desktop app. Since the Connect origin is loopback HTTP rather than Cloudflare private routing, the ICMP/UDP limitation cloudflared prints for HTTP/2 shouldn't matter here.
    • If a zero-connection watchdog lands, let the respawn switch to --protocol http2 after QUIC fails to register, since a plain auto restart won't stick on these networks.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions