Skip to content

feat(scheduler): attribute execution.active.count by platform properties - #2918

Open
rohnnyjoy wants to merge 1 commit into
TraceMachina:mainfrom
rohnnyjoy:active-count-platform-attributes
Open

rohnnyjoy wants to merge 1 commit into
TraceMachina:mainfrom
rohnnyjoy:active-count-platform-attributes

Conversation

@rohnnyjoy

Copy link
Copy Markdown
Contributor

What and why

execution.active.count carried only execution.stage, so one scheduler
serving two pools that differ by a platform property (a Linux and a macOS pool
on OSFamily) produced one queue depth for both, and a backlog in one pool
scaled the other. This adds active_action_count_platform_properties to
SimpleSpec: each listed key becomes an attribute execution.platform.<key>
on the gauge with the action's value ("" when unset), so each pool's queue
depth is its own series. Default empty keeps today's output exactly.

Fixes #2917.

Every add and subtract on the gauge goes through one attribute set built from
the configuration (ActiveCountAttributes in nativelink-util; the factory
refuses two keys that the collector's Prometheus exporter would fold into one
label, as it rewrites every character outside [A-Za-z0-9_] to _): the queue
insert, the stage transition and the client-gone removal in the memory DB, and
the periodic recount in the store DB. The memory DB moves an action between
series whenever its attributes change, not only its stage, because a requeue
can rewrite platform properties (memory escalation does), and a +1 and -1 with
different attributes would leave the gauge drifted for good. The store DB,
when keys are configured, reads each stage's actions and counts by value
instead of asking the store for a total, and records a series that emptied
down to zero. A series exists only once an action has carried its values, so an
idle pool reads absent rather than 0 until its first action: the field doc and
both autoscaling pages say so and guard every example recording rule with
or vector(0). The store DB warns at construction when keys are set while
enable_active_action_count_metric is off, since the keys then do nothing.

Docs: the autoscaling pages now show the per-pool recording rules and the
config to enable them, keep instance_name in the fast rule and the per-instance
max guidance, and state the Prometheus label rewriting. The configuration
reference (reference/nativelink-config/main.mdx) is regenerated so the snippet
lint knows the field; it was generated on this branch at v1.7.5, so its header
sha is that one, and the usual regeneration on main supersedes it.

How was this verified?

Unit tests, in nativelink-scheduler/tests/execution_active_count_test.rs:

  • platform_property_attributes_follow_each_action (memory DB): with
    ["OSFamily"], a linux, a macos and a key-less action are counted as
    queued{OSFamily=linux|macos|""}, each moves to executing and then
    completed under its own value while the others stay put, no stage-only
    series appears, and once the clients expire every series is back at zero.
  • store_backend_counts_each_platform_property_value (store DB, paused
    time): a fake store indexed by stage holds two linux queued, one macos
    queued, one macos executing, one bare executing; after the first recount the
    per-value series match, there is no stage-only series, after a rewrite the
    deltas move the right series, and after the store empties every series reads
    0.
  • The existing dropping_executing_action_decrements_active_count is kept and
    reads the same recorder.
  • nativelink-util/tests/metrics_test.rs: the attribute set is the stage then
    each key in order with "" for a key the action lacks, and new refuses
    ["OSFamily","OSFamily"], ["container-image","container.image"] and
    ["gpu_type","gpu-type"] with InvalidArgument naming both keys, while
    ["container-image","container_tag"] is accepted.

Without the change the per-value series never exist (the lookups read 0) and
the memory DB does not accept keys at all, so both new tests fail to build or
fail on their first assertion.

Ran, on rust 1.97.1 (the workspace's rust-version):

cargo +1.97.1 test -p nativelink-scheduler --test execution_active_count_test \
  --test store_awaited_action_db_test --test queue_order_test --test simple_scheduler_test
test store_backend_counts_each_platform_property_value ... ok
test dropping_executing_action_decrements_active_count ... ok
test platform_property_attributes_follow_each_action ... ok
test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.09s
test result: ok. 38 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
test result: ok. 14 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s

cargo +1.97.1 test -p nativelink-util --test metrics_test
test active_count_attributes_carry_the_stage_then_each_key_in_order ... ok
test active_count_attributes_refuse_keys_that_are_one_prometheus_label ... ok
test result: ok. 19 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out

cargo +1.97.1 test -p nativelink-scheduler        # whole crate, 26 test binaries, all ok
cargo +1.97.1 test -p nativelink-config   # all ok (34 + 11 + 1 + 1)

cargo +1.97.1 clippy --keep-going -p nativelink-scheduler -p nativelink-config \
  -p nativelink-util --all-targets
# clean on every touched target; the one finding is pre-existing and untouched:
# nativelink-util/tests/connection_manager_test.rs:145 semicolon_if_nothing_returned

cargo +nightly fmt --all                  # no diff after
cd web/apps/docs && node scripts/lint-snippets.mjs
# Checked 98 config snippet(s) across 83 page(s) against 318 known keys; No drift found.
node scripts/gen-config-reference.mjs main   # regenerated reference: the new row + header sha

Not verified: against a live Redis deployment, or an HPA end to end. The store
path is covered by the fake store only.

Risk

Low for anyone who does not set the new field: the default is the empty list
and the emitted series are byte-identical to before. The memory DB now
compares the full attribute vector rather than the stage discriminant to decide
whether to move the count; with no keys that is the same decision.

With keys set: (1) the store backend's recount reads every action in every
stage each pass instead of asking for four totals, so on a large Redis-backed
queue that is a heavier query every 15 s per replica, which the doc comment
says; (2) cardinality is the operator's: each distinct value combination is a
series kept at 0 for the life of the process; (3) a pool with no action yet has
no series at all, so an unguarded per-pool rule reads absent, which the docs
and the field doc call out. StoreAwaitedActionDb::new and
memory_awaited_action_db_factory gain a parameter, which touches every
scheduler test's construction site (the mechanical
ActiveCountAttributes::default() in the diff). A colliding key list is now a
startup error rather than a silent overwrite.

AI assistance

An agent drafted the change and the tests from a written design; I reviewed
every line and ran the verification above.

@vercel

vercel Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
nativelink Ready Ready Preview Oct 9, 2026 7:34pm UTC
nativelink-aidm Ready Ready Preview Oct 9, 2026 7:34pm UTC

Request Review

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

`execution.active.count` is the gauge an autoscaler reads for queue
depth, and it carried only `execution.stage`. One scheduler serving two
pools that differ by a platform property (a Linux pool and a macOS pool
on `OSFamily`, say) therefore produced one queue depth for both, and a
backlog of macOS actions scaled the Linux pool.

Add `active_action_count_platform_properties` to `SimpleSpec`: a list of
platform property keys, each of which becomes an attribute
`execution.platform.<key>` on `execution.active.count` with the action's
value for that key, or `""` when the action does not set it. The
operator chooses the keys, so cardinality is theirs to bound; every
series carries every listed key, so summing over them still gives the
per-stage totals. Empty, the default, is today's behaviour exactly.

The attribute set for an action is built once from the configuration
(`ActiveCountAttributes` in nativelink-util; the scheduler factory
refuses two keys the collector's Prometheus exporter would fold into one
label, since it rewrites every character outside `[A-Za-z0-9_]` to `_`),
and every add and subtract on the gauge goes through it: the queue insert, the
stage transition, and the client-gone removal in the memory DB, and the
periodic recount in the store DB. The memory DB moves an action between
series whenever its attributes change, not only its stage, since a
requeue may rewrite its platform properties (a memory escalation does).
The store DB, when keys are configured, reads each stage's actions and
counts them by value instead of asking the store for a total, reports
each series as the change since its last pass, and records a series
that emptied down to zero. A series exists only once an action has
carried its values, so an idle pool reads absent rather than 0 until its
first action; the field doc and the autoscaling pages say so and guard
every example recording rule with `or vector(0)`. The store DB warns at
construction when keys are set but the metric is off, since the keys
then do nothing.

Tests cover both backends: per-value counts through queued, executing
and completed, an action without the key under `""`, the attribute
following an action across stages, and every series back at zero once
the actions are gone. The autoscaling docs now show the per-pool
recording rule this makes possible, and the configuration reference is
regenerated for the new field.

This branch was successfully deployed

2 active deployments
Preview – nativelink — e41fb2f6 Deployed Oct 9, 2026 by vercel[bot]
Preview – nativelink-aidm — e41fb2f6 Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Attribute execution.active.count by platform properties so pools sharing a scheduler can autoscale separately

3 participants