Skip to content

Dispatch is acknowledged or declined by the worker - #2808

Merged
amankrx merged 4 commits into
TraceMachina:mainfrom
amankrx:pr/dispatch-ack
Sep 29, 2026
Merged

amankrx merged 4 commits into
TraceMachina:mainfrom
amankrx:pr/dispatch-ack

Conversation

@amankrx

@amankrx amankrx commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

What and why

A dispatch was fire-and-forget: the scheduler charged the worker's budget and moved on, and a worker that was full, short of memory or shutting down ran the action anyway or dropped it. The worker now answers each dispatch with ExecuteAccepted or ExecuteDeclined (at capacity, load, shutting down). A declined action goes back to the queue untried, without counting an attempt, and the worker is paused until a keepalive; the decline itself only refreshes liveness and does not lift the pause. With dispatch_ack_timeout_s set, a dispatch never acknowledged within that many seconds of the send is requeued the same way, and a worker that acknowledges after that is told to stop it. ConnectionResult tells the worker the scheduler speaks this and which platform property the memory veto reads, so a decline for load and the scheduler's veto look at the same number.

How was this verified?

dispatch_ack_test.rs covers decline with the pause and requeue, the channel-full path, the unacknowledged sweep on and off, and a late decline for an operation the worker no longer holds leaving it unpaused. Service tests show an acknowledgement arriving at the timeout is recorded before the liveness refresh that runs the sweep, the timeout counts from the scheduler clock at the send, a late acknowledgement is answered with a kill, a decline keeps the worker paused until a keepalive with room, and a kill that does not fit the worker's channel is deferred to the next sweep instead of evicting the worker. Worker tests show a worker at max_inflight_tasks declining the second dispatch, and a single-use worker declining for load and still taking the next dispatch. Each fails without its fix. On a test cluster a worker holding a 40 GiB resident process declined a 16 GiB action with Load, and the action ran elsewhere without a failed attempt.

Risk

Three new proto messages and two optional fields on ConnectionResult; an older worker never sends them and an older scheduler never sets them, so mixed versions fall back to today's behaviour. dispatch_ack_timeout_s is off by default. The worker's channel to the scheduler is bounded now, so a worker that stops reading is reported as channel-full instead of growing without bound.

AI assistance

An agent drafted the change and I reviewed every line.


This change is Reviewable

@vercel

vercel Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
nativelink Ready Ready Preview Sep 29, 2026 3:38am UTC
nativelink-aidm Ready Ready Preview Sep 29, 2026 3:38am UTC

Request Review

(cherry picked from commit a52c0a19ad03c7cab606fdb1c087ca3b76d45213)
…ate decline does not pause, counters count requeues, a dead channel is a decline
…, and a late acknowledgement is answered with a kill
Comment thread nativelink-service/src/worker_api_server.rs
Comment thread nativelink-worker/src/local_worker.rs Outdated
Comment thread nativelink-worker/src/local_worker.rs Outdated
Comment thread nativelink-scheduler/src/api_worker_scheduler.rs
@amankrx
amankrx merged commit a21d82c into TraceMachina:main Sep 29, 2026
48 checks passed

This branch was successfully deployed

2 active deployments
Preview – nativelink — d6986255 Deployed Sep 29, 2026 by vercel[bot]
Preview – nativelink-aidm — d6986255 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants