Repository navigation
Scheduler: memory escalation with retry budgets by cause, the whole worker as the last step - #2829
Merged
Merged
Conversation
…orker as the last step
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
…raining workers left out, percent validated, a zero base refused, a max_steps test, the shared requeue metric
corcillo
approved these changes
Sep 30, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
An action the worker killed for memory failed outright or burned
max_job_retrieson the same reservation. Withmemory_escalationthe scheduler requeues it with a larger one, the reservation timespercent(default 200) or the next class ofladder_kb, capped at the largest connected worker that is not draining, and bymax_kbbelow that; the last step reserves that worker whole, memory and throughcpu_propertyCPU, and only a kill there fails the action. A whole-worker ask is not vetoed bylive_memory_veto, since no worker reports its whole memory free;percentat or below 100 is refused at load, and a kill with no reservation and no sampled peak is not escalated. Retries are budgeted by cause: a lost worker spendsmax_worker_loss_retries(default 10) and an escalation spendsmax_steps(default 8), so neither eatsmax_job_retries, and the execution histograms carry anexecution.outcomeattribute.How was this verified?
escalation_test: a memory kill requeues with double the reservation, an escalation and a lost worker each leave the retry budget alone,max_stepsends the action at the third kill, and a ladder steps class by class, then takes the whole worker, then fails. Escalations count under the shareddispatch_requeuesmetric, and the metrics reference is regenerated forexecution.outcome. The stored record's two new counters default to zero, so older records load unchanged. The config reference is regenerated and the snippet lint passes.Risk
Escalation is off unless
memory_escalationis set. One default changes: a lost worker no longer counts againstmax_job_retries, so an action whose worker keeps dying is retried up to ten times before it fails, for the release notes.AI assistance
An agent drafted the change and I reviewed every line.
This change is