Repository navigation
Node delivery: engine-side follow-ups surfaced by #1216 (409-not-500, /v1/node/ws 401 churn, up fresh-workspace) #1217
Description
Activity
Broker-side evidence for item #2 (/v1/node/ws 401 churn) — points to a ~60s engine-side token/session TTL
Mechanics (broker,
crates/broker/src/node_control.rs): a node token is valid at connect, the connection runs ~60s, then the engine 401s it on/v1/node/ws. The broker re-mints (create_nodereturns a fresh token for its own node) and reconnects. The consecutive-401 counter resets on every successful connect (should_attempt_remint, node_control.rs:44), so it never trips the give-up cap (MAX_UNAUTHORIZED_BEFORE_GIVING_UP = 5) — it churns indefinitely by design, with a ~6s reconnect gap per cycle (INITIAL_RECONNECT_DELAY1s + create_node backoffs 200/400/800ms). That ~6s window every ~60s is the spawn/delivery outage —spawn:*rides the same socket, so a spawn attempt during the gap finds no spawn-capable node.Connected→disconnect intervals (worker, from broker logs):
run connected→disconnect 401 after connect clean e2e (branch) 59.857s 61.686s echo e2e (branch) 62.317s 63.872s node-delivery samples 60.081 / 60.166 / 61.572 / 71.373 / 76.125s — published 9.1.7 (A/B) 37.960s (outlier) 39.875s fleet node node_641bf7…61.799 / 68.122 / 137.999s + short reconnect failures — Median of filtered samples: 60.166s. Reproduces identically on the branch broker and published
agent-relay-broker 9.1.7(so it is not the #1216 change), and affects the lead/worker fleet nodenode_641bf7…, not justup-path nodes.Token format: opaque
nt_live…(len 56, single dot-part, not a JWT); cached token JSON has onlynode_id/workspace_id/base_url— noexp/ttl/expires_at. So the exact TTL is engine-side and needs engine storage/logs to confirm.Likely fixes (engine): issue node tokens with an appropriate (long) TTL; OR make
/v1/node/wsauth session-lifetime so an authenticated socket isn't dropped on token expiry; OR support in-band token refresh without tearing down the WS. Broker-side we could shrink the reconnect gap (proactive re-mint just before TTL + faster reconnect) as a stopgap, but that doesn't remove the outage.Production verification (via wrangler / prod D1)
The node-heartbeat liveness fix (relaycast-cloud#19, deployed) is confirmed working in production.
Authoritative signal — watched our live fleet node
node_641bf7…in the prodrelaycast-cloudD1nodestable:- Before the fix (earlier today): this node showed the ~60s churn.
- After deploy:
status=onlinewithlast_heartbeat_atrefreshing every ~12s —(now - last_heartbeat_at)stayed bounded at 1–12s across repeated samples, never approaching the 45s liveness TTL. No reap.
Root fix: NodeDO stamps
nodes.last_heartbeat_at(Unix seconds) in the swept D1 store on each validnode.heartbeat, plus an open-socket self-heal/re-arm guard so a node with a live socket is never reaped on a stale DB timestamp.Notes / record accuracy:
- relaycast-cloud#21 (normalize
id:nullheartbeat frames) is defensive hardening, not the root fix — the Rust broker omitsid(skip_serializing_if), so it never sendsid:null. docs: add comprehensive monetization strategy spec #21 is harmless tolerance for non-Rust/older clients. - The earlier "still churning" reading was deploy/Durable-Object propagation lag (tested ~5 min post-deploy).
- An idle (traffic-free) node is covered by the open-socket self-heal guard regardless of heartbeat-write timing.
Closing as resolved — the original injection + spawn churn (delivery dead-lettering,
spawn:*failures) is fixed end-to-end across: relay#1216 (workspace-scoped node id, released v9.1.8) + relaycast-cloud#19/#21 (heartbeat liveness).Closing as completed — fix verified live in prod (relaycast-cloud#19 deployed; relay#1216 released in v9.1.8). Real fleet node node_641bf7 confirmed stable in prod D1 (heartbeat age bounded ~1–12s, no ~60s reap). Passive monitoring only for any recurrence of 401 churn / delivery TTL expiry over the next normal traffic window; reopen if it resurfaces.
Follow-ups split out of #1216 (workspace-scoped node id). The root injection bug is fixed there; these are separate, mostly engine-side, and tracked here so the PR stays scoped.
1.
create_nodereturns 500 (raw DB error) instead of 409 for a node-id conflictWhen a
node_idalready exists (owned by another workspace), thePOST /v1/nodesINSERT hits a unique-constraint violation and the engine returns HTTP 500internal_errorwith a rawinsert into "nodes" …body, instead of a clean 409 Conflict. Because 500 is retryable, the broker's retry wrapper loops and collapses it toMax retries exceeded, hiding the real cause.409for node-id conflicts; ideally makecreate_nodeidempotent for the same owner.upinjection works #1216 (directPOST https://cast.agentrelay.com/v1/nodes).2. Root-cause the
/v1/node/ws401 → re-mint churnAfter a node token mints and node-control connects, the WS handshake intermittently 401s, triggering re-mint → reconnect churn (agents undeliverable during the gap).
upinjection works #1216 reproduced it identically on the branch build and the publishedagent-relay-broker 9.1.7, same clean cwd + durable workspace key.up-specific: the active pear/fleet broker shows the same 401 + re-mint cycle for the lead/worker fleet nodenode_641bf7acf3e31604ed374f5fca074f53in workspace197733757872058368.3.
agent-relay upmints a fresh workspace on every run (source=fresh)Each
upinvocation appears to mint a new workspace (the workspace id changes across restarts), which is why a project directory's node identity kept shifting. Intent vs. bug TBD — if unintended,upshould reuse a durable per-project workspace.Reference: #1216 (root fix + the A/B evidence and reproduction context).