Skip to content

Node delivery: engine-side follow-ups surfaced by #1216 (409-not-500, /v1/node/ws 401 churn, up fresh-workspace) #1217

Description

@khaliqgant

Follow-ups split out of #1216 (workspace-scoped node id). The root injection bug is fixed there; these are separate, mostly engine-side, and tracked here so the PR stays scoped.

1. create_node returns 500 (raw DB error) instead of 409 for a node-id conflict

When a node_id already exists (owned by another workspace), the POST /v1/nodes INSERT hits a unique-constraint violation and the engine returns HTTP 500 internal_error with a raw insert into "nodes" … body, instead of a clean 409 Conflict. Because 500 is retryable, the broker's retry wrapper loops and collapses it to Max retries exceeded, hiding the real cause.

2. Root-cause the /v1/node/ws 401 → re-mint churn

After a node token mints and node-control connects, the WS handshake intermittently 401s, triggering re-mint → reconnect churn (agents undeliverable during the gap).

  • Confirmed pre-existing / engine-side, not a regression: A/B in fix(broker): workspace-scope auto node id so up injection works #1216 reproduced it identically on the branch build and the published agent-relay-broker 9.1.7, same clean cwd + durable workspace key.
  • Confirmed not up-specific: the active pear/fleet broker shows the same 401 + re-mint cycle for the lead/worker fleet node node_641bf7acf3e31604ed374f5fca074f53 in workspace 197733757872058368.

3. agent-relay up mints a fresh workspace on every run (source=fresh)

Each up invocation appears to mint a new workspace (the workspace id changes across restarts), which is why a project directory's node identity kept shifting. Intent vs. bug TBD — if unintended, up should reuse a durable per-project workspace.

Reference: #1216 (root fix + the A/B evidence and reproduction context).

Activity

  1. khaliqgant commented on Jun 30, 2026

    @khaliqgant
    MemberAuthor

    Broker-side evidence for item #2 (/v1/node/ws 401 churn) — points to a ~60s engine-side token/session TTL

    Mechanics (broker, crates/broker/src/node_control.rs): a node token is valid at connect, the connection runs ~60s, then the engine 401s it on /v1/node/ws. The broker re-mints (create_node returns a fresh token for its own node) and reconnects. The consecutive-401 counter resets on every successful connect (should_attempt_remint, node_control.rs:44), so it never trips the give-up cap (MAX_UNAUTHORIZED_BEFORE_GIVING_UP = 5) — it churns indefinitely by design, with a ~6s reconnect gap per cycle (INITIAL_RECONNECT_DELAY 1s + create_node backoffs 200/400/800ms). That ~6s window every ~60s is the spawn/delivery outage — spawn:* rides the same socket, so a spawn attempt during the gap finds no spawn-capable node.

    Connected→disconnect intervals (worker, from broker logs):

    run connected→disconnect 401 after connect
    clean e2e (branch) 59.857s 61.686s
    echo e2e (branch) 62.317s 63.872s
    node-delivery samples 60.081 / 60.166 / 61.572 / 71.373 / 76.125s —
    published 9.1.7 (A/B) 37.960s (outlier) 39.875s
    fleet node node_641bf7… 61.799 / 68.122 / 137.999s + short reconnect failures —

    Median of filtered samples: 60.166s. Reproduces identically on the branch broker and published agent-relay-broker 9.1.7 (so it is not the #1216 change), and affects the lead/worker fleet node node_641bf7…, not just up-path nodes.

    Token format: opaque nt_live… (len 56, single dot-part, not a JWT); cached token JSON has only node_id / workspace_id / base_url — no exp/ttl/expires_at. So the exact TTL is engine-side and needs engine storage/logs to confirm.

    Likely fixes (engine): issue node tokens with an appropriate (long) TTL; OR make /v1/node/ws auth session-lifetime so an authenticated socket isn't dropped on token expiry; OR support in-band token refresh without tearing down the WS. Broker-side we could shrink the reconnect gap (proactive re-mint just before TTL + faster reconnect) as a stopgap, but that doesn't remove the outage.

  2. khaliqgant commented on Jun 30, 2026

    @khaliqgant
    MemberAuthor

    Production verification (via wrangler / prod D1)

    The node-heartbeat liveness fix (relaycast-cloud#19, deployed) is confirmed working in production.

    Authoritative signal — watched our live fleet node node_641bf7… in the prod relaycast-cloud D1 nodes table:

    • Before the fix (earlier today): this node showed the ~60s churn.
    • After deploy: status=online with last_heartbeat_at refreshing every ~12s — (now - last_heartbeat_at) stayed bounded at 1–12s across repeated samples, never approaching the 45s liveness TTL. No reap.

    Root fix: NodeDO stamps nodes.last_heartbeat_at (Unix seconds) in the swept D1 store on each valid node.heartbeat, plus an open-socket self-heal/re-arm guard so a node with a live socket is never reaped on a stale DB timestamp.

    Notes / record accuracy:

    • relaycast-cloud#21 (normalize id:null heartbeat frames) is defensive hardening, not the root fix — the Rust broker omits id (skip_serializing_if), so it never sends id:null. docs: add comprehensive monetization strategy spec #21 is harmless tolerance for non-Rust/older clients.
    • The earlier "still churning" reading was deploy/Durable-Object propagation lag (tested ~5 min post-deploy).
    • An idle (traffic-free) node is covered by the open-socket self-heal guard regardless of heartbeat-write timing.

    Closing as resolved — the original injection + spawn churn (delivery dead-lettering, spawn:* failures) is fixed end-to-end across: relay#1216 (workspace-scoped node id, released v9.1.8) + relaycast-cloud#19/#21 (heartbeat liveness).

  3. khaliqgant commented on Jun 30, 2026

    @khaliqgant
    MemberAuthor

    Closing as completed — fix verified live in prod (relaycast-cloud#19 deployed; relay#1216 released in v9.1.8). Real fleet node node_641bf7 confirmed stable in prod D1 (heartbeat age bounded ~1–12s, no ~60s reap). Passive monitoring only for any recurrence of 401 churn / delivery TTL expiry over the next normal traffic window; reopen if it resurfaces.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions