Problem statement
execution.active.count is the gauge the autoscaling docs
(operate/autoscaling-reference, operate/scaling-workers) tell you to scale
on, and it is attributed by execution.stage only. One scheduler that serves
two pools told apart by a platform property, a Linux pool and a macOS pool on
OSFamily for instance, therefore emits one queue depth for both. A backlog of
macOS actions scales the Linux pool and changes nothing; the macOS pool never
sees its own demand. The docs admit this ("attributed by stage only") and
suggest splitting the recording rule "by the property that distinguishes the
pools", which the metric does not carry, or running a scheduler per pool.
Proposed solution
A SimpleSpec option, active_action_count_platform_properties: Vec<String>,
default empty (today's behaviour). Each listed platform property key becomes an
attribute execution.platform.<key> on execution.active.count with the
action's value for that key, "" when the action does not set it. Every series
carries every listed key, so sum by (execution_stage) still gives the old
totals, and sum by (execution_platform_OSFamily) gives one queue depth per
pool. Cardinality is bounded by what the operator lists. Both awaited-action
backends honour it: the in-memory one per transition, the store one in its
periodic recount.
Alternatives considered
- A scheduler instance per pool. Works, but doubles the deployment and the
REAPI endpoints for a label.
- Deriving the pool from
execution.worker_id on the transition counters.
Queued actions have no worker yet, and queue depth is the signal.
- Attributing by every platform property automatically. Unbounded cardinality
(memory_kb, per-build values); the operator must choose.
Additional context
I have a patch ready (both backends, tests, docs) and would open the PR once
this is agreed.
Problem statement
execution.active.countis the gauge the autoscaling docs(
operate/autoscaling-reference,operate/scaling-workers) tell you to scaleon, and it is attributed by
execution.stageonly. One scheduler that servestwo pools told apart by a platform property, a Linux pool and a macOS pool on
OSFamilyfor instance, therefore emits one queue depth for both. A backlog ofmacOS actions scales the Linux pool and changes nothing; the macOS pool never
sees its own demand. The docs admit this ("attributed by stage only") and
suggest splitting the recording rule "by the property that distinguishes the
pools", which the metric does not carry, or running a scheduler per pool.
Proposed solution
A
SimpleSpecoption,active_action_count_platform_properties: Vec<String>,default empty (today's behaviour). Each listed platform property key becomes an
attribute
execution.platform.<key>onexecution.active.countwith theaction's value for that key,
""when the action does not set it. Every seriescarries every listed key, so
sum by (execution_stage)still gives the oldtotals, and
sum by (execution_platform_OSFamily)gives one queue depth perpool. Cardinality is bounded by what the operator lists. Both awaited-action
backends honour it: the in-memory one per transition, the store one in its
periodic recount.
Alternatives considered
REAPI endpoints for a label.
execution.worker_idon the transition counters.Queued actions have no worker yet, and queue depth is the signal.
(
memory_kb, per-build values); the operator must choose.Additional context
I have a patch ready (both backends, tests, docs) and would open the PR once
this is agreed.