Repository navigation
[Bug]: applyCloudRelayConfig is not atomic; a failed secret write after the relay link is committed leaves the environment stuck on endpoint_provider_not_managed #11898
Description
Activity
UI-side companion (toggle cannot detect or repair the drift once it exists): #11899
Triage
Confirmed against
main. This is a real server defect and a specific root cause of theendpoint_provider_not_managedstuck state in #6568. Keeping this issue separate from that umbrella (and from the companion UI issue #11899).What the code does
linkPrimaryEnvironmentToCloud(apps/web/src/cloud/linkEnvironment.ts) commits the relay link first (POST /v1/client/environment-links→EnvironmentLinks.upsert, includingendpointProviderKind), then callsPOST /api/connect/relay-config.applyCloudRelayConfig(apps/server/src/cloud/http.ts) then:- Applies the managed-endpoint runtime (
endpointRuntime.applyConfig) before any secret writes. On a managed → publish-only relink that stops cloudflared immediately. - Writes six secrets one at a time via
ServerSecretStore.set/remove:cloud-relay-url,cloud-relay-issuer,cloud-linked-user-id,cloud-relay-environment-credential,cloud-mint-ed25519-public-key, thencloud-endpoint-runtime-config(set or removed).
ServerSecretStore.setis atomic per file (write tmp, rename). There is nosetMany, no staging directory, and no rollback if a later write fails. The CLI reconcile path (reconcileDesiredCloudLinkWith) uses the same helper, so it has the same split-brain.If the second write fails (
SecretStorePersistError, e.g. WindowsEPERMwhile AV holds the rename), the relay already hasendpointProviderKind = "manual"while local state is mixed: newcloud-relay-url, leftover managed credential / mint key / runtime-config.readCloudLinkStatereportsmanagedTunnelActive: truesolely from the leftover runtime-config file. A desktop restart will respawn the old tunnel from that file.The relay then rejects every connect/status with
endpoint_provider_not_managed(EnvironmentConnector.resolveManagedEndpoint). Settings only relinks when localmanagedTunnelActive !== desired.managedTunnel, so toggling T3 Connect is a no-op while Publish stays on. That recovery hole is #11899.Not already fixed
No open/merged PR makes this apply transactional. #10579 only hardens per-file
createcompleteness inServerSecretStore; it does not group these six writes.Existing tests (
apps/server/src/cloud/http.test.ts,linkEnvironment.test.ts) cover the happy-path order, not a mid-set persist failure.Suggested fix
Either the whole relink applies or none of it does:
- Stage all six secrets, then promote (or persist local config before the relay upsert).
- And/or roll back the relay link (re-upsert previous endpoint, or unlink) if
relay-configfails. - Move
applyConfigto after a successful secret promote so runtime and disk cannot diverge. - Regression: fail the second
secrets.setand assert previous local secrets + runtime-config are unchanged.
#11899 is still worth doing so an already-drifted install can be detected and repaired, but it does not replace making the apply atomic.
Workaround (verified by reporter)
Settings > Connections: Publish off, T3 Connect off (true unlink), T3 Connect on, then publish back on.
Related
- [Bug]: T3 Connect environment is discoverable but relay connection is unauthorized (endpoint_provider_not_managed) #6568 — symptom umbrella; do not close this as a duplicate
- [Bug]: Connections settings cannot detect or repair a relay link that drifted to publish-only (toggle shows on, relink is a no-op, mobile gets endpoint_provider_not_managed) #11899 — Connections UI cannot see or repair the drift
- T3 Connect: /oauth/token returns EnvironmentAuthInvalidError after relinking environment, making Desktop and iOS unable to connect #5712 — closed; related relink inconsistency, different error
- Applies the managed-endpoint runtime (
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 15, 2026
Before submitting
Area
apps/server
Steps to reproduce
Context: this is the root cause behind the
endpoint_provider_not_managedconnection failures in #6568. Filing it separately because it is one specific defect with a specific fix, and the other issue has grown to cover several symptoms.What the relink flow does today
linkPrimaryEnvironmentToCloud(apps/web/src/cloud/linkEnvironment.ts) runs three steps in order:POST /v1/client/environment-link-challengeson the relayPOST /v1/client/environment-linkson the relay. This upsertsrelay_environment_linksfor(userId, environmentId), includingendpointProviderKind(infra/relay/src/environments/EnvironmentLinks.ts,upsert). At this point the relay's copy of the link is already committed.POST /api/connect/relay-configon the local environment server, handled byapplyCloudRelayConfig(apps/server/src/cloud/http.ts). That handler writes six secrets one at a time throughServerSecretStore.set, in this order:cloud-relay-url,cloud-relay-issuer,cloud-linked-user-id,cloud-relay-environment-credential,cloud-mint-ed25519-public-key, thencloud-endpoint-runtime-config(set or removed depending on the mode).There is no transaction and no rollback around step 3.
ServerSecretStore.setis atomic per file (write tmp, rename), but the six writes together are not. If any write after the first fails, the relay has the new link and the local server keeps the previous credential, user id, mint key, and runtime config.How to reproduce deterministically
cloud-endpoint-runtime-config.binexists and cloudflared is running.~/.t3/userdata/secrets/cloud-relay-issuer.binread-only, hold it open with a tool that takes an exclusive lock, or on Windows let an on-access antivirus scanner hold the file during the rename. (renameover an open file fails withEPERMon Windows; the store maps that toSecretStorePersistErrorand the handler returns "Could not persist environment relay configuration.")publish_onlymode (useCloudLinkController.ts,mode: desired.managedTunnel ? "managed" : "publish_only"), so step 2 storesendpointProviderKind = "manual"on the relay.cloud-relay-url.binrewritten, everything else from the old managed link, cloudflared still running against the old tunnel.Expected behavior
Either the whole relink applies or none of it does. Concretely, one of:
relay-configfails, orapplyCloudRelayConfigwrite all six secrets to temp files first and only rename them into place once every temp write has succeeded, so local state cannot end up half-updated.Actual behavior
Relay and desktop disagree about the link and nothing reconciles them:
endpointProviderKind = "manual".EnvironmentConnector.resolveManagedEndpoint(infra/relay/src/environments/EnvironmentConnector.ts) rejects every connect and status call withendpoint_provider_not_managed.readCloudLinkStatereportsmanagedTunnelActive: truebecausecloud-endpoint-runtime-configstill exists. cloudflared keeps the old tunnel up, so the environment stays discoverable and health checks pass, which is why the mobile app lists it and then fails on connect.managedTunnelActive !== desired.managedTunnel, so turning the toggle on again is a no-op (that part is tracked separately, see the linked UI issue).Real occurrence on my machine: on 2026-09-11 11:30:26 local,
cloud-relay-url.binwas rewritten and the other five secrets kept their 2026-07-31 timestamps. The mobile app reportedConnection failed. Reason: Relay rejected the environment connection request (endpoint_provider_not_managed). trace id: ad7746aff889c11de16074675a73acdafrom then until I forced a full relink on 2026-09-15.Impact
Blocks work completely
Version or commit
Desktop 0.0.40 (release channel). Code references are against
mainas of 2026-09-15.Environment
Windows 11 Home 10.0.26200, T3 Code desktop 0.0.40, iOS T3 Code app, McAfee real-time scanning enabled at the time of the failed write.
Logs or stack traces
Workaround
Force a full relink so the relay record is rewritten as
cloudflare_tunnel: Settings > Connections, turn Publish agent activity off, turn T3 Connect off (this is the only combination that actually unlinks), then turn T3 Connect on, then publish back on. Turning only the T3 Connect toggle off and on does not help while publishing is on, because "off" becomes a publish-only relink rather than an unlink.