Skip to content

Scheduler: memory escalation with retry budgets by cause, the whole worker as the last step - #2829

Merged
amankrx merged 2 commits into
TraceMachina:mainfrom
amankrx:pr/memory-escalation
Sep 30, 2026
Merged

amankrx merged 2 commits into
TraceMachina:mainfrom
amankrx:pr/memory-escalation

Conversation

@amankrx

@amankrx amankrx commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

What and why

An action the worker killed for memory failed outright or burned max_job_retries on the same reservation. With memory_escalation the scheduler requeues it with a larger one, the reservation times percent (default 200) or the next class of ladder_kb, capped at the largest connected worker that is not draining, and by max_kb below that; the last step reserves that worker whole, memory and through cpu_property CPU, and only a kill there fails the action. A whole-worker ask is not vetoed by live_memory_veto, since no worker reports its whole memory free; percent at or below 100 is refused at load, and a kill with no reservation and no sampled peak is not escalated. Retries are budgeted by cause: a lost worker spends max_worker_loss_retries (default 10) and an escalation spends max_steps (default 8), so neither eats max_job_retries, and the execution histograms carry an execution.outcome attribute.

How was this verified?

escalation_test: a memory kill requeues with double the reservation, an escalation and a lost worker each leave the retry budget alone, max_steps ends the action at the third kill, and a ladder steps class by class, then takes the whole worker, then fails. Escalations count under the shared dispatch_requeues metric, and the metrics reference is regenerated for execution.outcome. The stored record's two new counters default to zero, so older records load unchanged. The config reference is regenerated and the snippet lint passes.

Risk

Escalation is off unless memory_escalation is set. One default changes: a lost worker no longer counts against max_job_retries, so an action whose worker keeps dying is retried up to ten times before it fails, for the release notes.

AI assistance

An agent drafted the change and I reviewed every line.


This change is Reviewable

@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
nativelink Ready Ready Preview Sep 30, 2026 2:12am UTC
nativelink-aidm Ready Ready Preview Sep 30, 2026 2:12am UTC

Request Review

Comment thread nativelink-scheduler/src/api_worker_scheduler.rs
Comment thread nativelink-scheduler/src/api_worker_scheduler.rs
Comment thread nativelink-scheduler/src/api_worker_scheduler.rs
Comment thread nativelink-scheduler/src/api_worker_scheduler.rs
Comment thread nativelink-scheduler/src/api_worker_scheduler.rs Outdated
Comment thread nativelink-util/src/metrics.rs
Comment thread nativelink-scheduler/tests/escalation_test.rs
Comment thread nativelink-scheduler/src/simple_scheduler_state_manager.rs Outdated
…raining workers left out, percent validated, a zero base refused, a max_steps test, the shared requeue metric
@amankrx
amankrx merged commit 3c52ee1 into TraceMachina:main Sep 30, 2026
47 of 50 checks passed

This branch was successfully deployed

2 active deployments
Preview – nativelink — 73a8b1fe Deployed Sep 30, 2026 by vercel[bot]
Preview – nativelink-aidm — 73a8b1fe Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants