Skip to content

[Bug]: applyCloudRelayConfig is not atomic; a failed secret write after the relay link is committed leaves the environment stuck on endpoint_provider_not_managed #11898

Description

@ElliotDrel

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

Context: this is the root cause behind the endpoint_provider_not_managed connection failures in #6568. Filing it separately because it is one specific defect with a specific fix, and the other issue has grown to cover several symptoms.

What the relink flow does today

linkPrimaryEnvironmentToCloud (apps/web/src/cloud/linkEnvironment.ts) runs three steps in order:

  1. POST /v1/client/environment-link-challenges on the relay
  2. POST /v1/client/environment-links on the relay. This upserts relay_environment_links for (userId, environmentId), including endpointProviderKind (infra/relay/src/environments/EnvironmentLinks.ts, upsert). At this point the relay's copy of the link is already committed.
  3. POST /api/connect/relay-config on the local environment server, handled by applyCloudRelayConfig (apps/server/src/cloud/http.ts). That handler writes six secrets one at a time through ServerSecretStore.set, in this order: cloud-relay-url, cloud-relay-issuer, cloud-linked-user-id, cloud-relay-environment-credential, cloud-mint-ed25519-public-key, then cloud-endpoint-runtime-config (set or removed depending on the mode).

There is no transaction and no rollback around step 3. ServerSecretStore.set is atomic per file (write tmp, rename), but the six writes together are not. If any write after the first fails, the relay has the new link and the local server keeps the previous credential, user id, mint key, and runtime config.

How to reproduce deterministically

  1. Link an environment normally in managed mode so cloud-endpoint-runtime-config.bin exists and cloudflared is running.
  2. Make the second secret write fail. Any of these works: make ~/.t3/userdata/secrets/cloud-relay-issuer.bin read-only, hold it open with a tool that takes an exclusive lock, or on Windows let an on-access antivirus scanner hold the file during the rename. (rename over an open file fails with EPERM on Windows; the store maps that to SecretStorePersistError and the handler returns "Could not persist environment relay configuration.")
  3. In Settings > Connections, with Publish agent activity on, turn T3 Connect off. The controller relinks in publish_only mode (useCloudLinkController.ts, mode: desired.managedTunnel ? "managed" : "publish_only"), so step 2 stores endpointProviderKind = "manual" on the relay.
  4. Step 3 fails after the first write. Local state is now: cloud-relay-url.bin rewritten, everything else from the old managed link, cloudflared still running against the old tunnel.
  5. Try to connect from the mobile app.

Expected behavior

Either the whole relink applies or none of it does. Concretely, one of:

  • Persist the local relay config before committing the relay-side link (obtain everything, write all secrets to staging, then upsert on the relay, then promote the staged secrets), or
  • Keep the current order but roll back the relay link (re-upsert the previous endpoint, or unlink) when relay-config fails, or
  • At minimum, make applyCloudRelayConfig write all six secrets to temp files first and only rename them into place once every temp write has succeeded, so local state cannot end up half-updated.

Actual behavior

Relay and desktop disagree about the link and nothing reconciles them:

  • Relay: endpointProviderKind = "manual". EnvironmentConnector.resolveManagedEndpoint (infra/relay/src/environments/EnvironmentConnector.ts) rejects every connect and status call with endpoint_provider_not_managed.
  • Desktop: readCloudLinkState reports managedTunnelActive: true because cloud-endpoint-runtime-config still exists. cloudflared keeps the old tunnel up, so the environment stays discoverable and health checks pass, which is why the mobile app lists it and then fails on connect.
  • The Connections UI only relinks when managedTunnelActive !== desired.managedTunnel, so turning the toggle on again is a no-op (that part is tracked separately, see the linked UI issue).

Real occurrence on my machine: on 2026-09-11 11:30:26 local, cloud-relay-url.bin was rewritten and the other five secrets kept their 2026-07-31 timestamps. The mobile app reported Connection failed. Reason: Relay rejected the environment connection request (endpoint_provider_not_managed). trace id: ad7746aff889c11de16074675a73acda from then until I forced a full relink on 2026-09-15.

Impact

Blocks work completely

Version or commit

Desktop 0.0.40 (release channel). Code references are against main as of 2026-09-15.

Environment

Windows 11 Home 10.0.26200, T3 Code desktop 0.0.40, iOS T3 Code app, McAfee real-time scanning enabled at the time of the failed write.

Logs or stack traces

# secrets dir after the half-applied relink (2026-09-11), everything else untouched since the original link
-rw-r--r-- 411 2026-07-31 23:49:58 cloud-endpoint-runtime-config.bin   # still providerKind cloudflare_tunnel
-rw-r--r--  32 2026-07-31 23:49:58 cloud-linked-user-id.bin
-rw-r--r-- 112 2026-07-31 23:49:58 cloud-mint-ed25519-public-key.bin
-rw-r--r-- 167 2026-07-31 23:49:58 cloud-relay-environment-credential.bin
-rw-r--r--  22 2026-07-31 23:49:58 cloud-relay-issuer.bin
-rw-r--r--  22 2026-09-11 11:30:26 cloud-relay-url.bin                 # only file the relink touched

# after a successful relink (2026-09-15 09:45:45) all six land within 10 ms of each other
-rw-r--r-- 411 2026-09-15 09:45:45.291 cloud-endpoint-runtime-config.bin
-rw-r--r--  32 2026-09-15 09:45:45.285 cloud-linked-user-id.bin
-rw-r--r-- 112 2026-09-15 09:45:45.289 cloud-mint-ed25519-public-key.bin
-rw-r--r-- 167 2026-09-15 09:45:45.287 cloud-relay-environment-credential.bin
-rw-r--r--  22 2026-09-15 09:45:45.283 cloud-relay-issuer.bin
-rw-r--r--  22 2026-09-15 09:45:45.281 cloud-relay-url.bin

Workaround

Force a full relink so the relay record is rewritten as cloudflare_tunnel: Settings > Connections, turn Publish agent activity off, turn T3 Connect off (this is the only combination that actually unlinks), then turn T3 Connect on, then publish back on. Turning only the T3 Connect toggle off and on does not help while publishing is on, because "off" becomes a publish-only relink rather than an unlink.

Activity

  1. ElliotDrel commented on Sep 15, 2026

    @ElliotDrel
    Author

    UI-side companion (toggle cannot detect or repair the drift once it exists): #11899

  2. juliusmarminge commented on Sep 15, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed against main. This is a real server defect and a specific root cause of the endpoint_provider_not_managed stuck state in #6568. Keeping this issue separate from that umbrella (and from the companion UI issue #11899).

    What the code does

    linkPrimaryEnvironmentToCloud (apps/web/src/cloud/linkEnvironment.ts) commits the relay link first (POST /v1/client/environment-links → EnvironmentLinks.upsert, including endpointProviderKind), then calls POST /api/connect/relay-config.

    applyCloudRelayConfig (apps/server/src/cloud/http.ts) then:

    1. Applies the managed-endpoint runtime (endpointRuntime.applyConfig) before any secret writes. On a managed → publish-only relink that stops cloudflared immediately.
    2. Writes six secrets one at a time via ServerSecretStore.set / remove: cloud-relay-url, cloud-relay-issuer, cloud-linked-user-id, cloud-relay-environment-credential, cloud-mint-ed25519-public-key, then cloud-endpoint-runtime-config (set or removed).

    ServerSecretStore.set is atomic per file (write tmp, rename). There is no setMany, no staging directory, and no rollback if a later write fails. The CLI reconcile path (reconcileDesiredCloudLinkWith) uses the same helper, so it has the same split-brain.

    If the second write fails (SecretStorePersistError, e.g. Windows EPERM while AV holds the rename), the relay already has endpointProviderKind = "manual" while local state is mixed: new cloud-relay-url, leftover managed credential / mint key / runtime-config. readCloudLinkState reports managedTunnelActive: true solely from the leftover runtime-config file. A desktop restart will respawn the old tunnel from that file.

    The relay then rejects every connect/status with endpoint_provider_not_managed (EnvironmentConnector.resolveManagedEndpoint). Settings only relinks when local managedTunnelActive !== desired.managedTunnel, so toggling T3 Connect is a no-op while Publish stays on. That recovery hole is #11899.

    Not already fixed

    No open/merged PR makes this apply transactional. #10579 only hardens per-file create completeness in ServerSecretStore; it does not group these six writes.

    Existing tests (apps/server/src/cloud/http.test.ts, linkEnvironment.test.ts) cover the happy-path order, not a mid-set persist failure.

    Suggested fix

    Either the whole relink applies or none of it does:

    • Stage all six secrets, then promote (or persist local config before the relay upsert).
    • And/or roll back the relay link (re-upsert previous endpoint, or unlink) if relay-config fails.
    • Move applyConfig to after a successful secret promote so runtime and disk cannot diverge.
    • Regression: fail the second secrets.set and assert previous local secrets + runtime-config are unchanged.

    #11899 is still worth doing so an already-drifted install can be detected and repaired, but it does not replace making the apply atomic.

    Workaround (verified by reporter)

    Settings > Connections: Publish off, T3 Connect off (true unlink), T3 Connect on, then publish back on.

    Related

  3. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions