Repository navigation
CI resilience: single self-hosted runner 'Little-Slave' is a SPOF that blocks ALL PR checks when offline #11380
Description
Activity
- addedenhancementNew feature or requestNew feature or request
on Jul 9, 2026 Measured today — the SPOF is now also a hard throughput ceiling
Snapshot taken 2026-08-03 during ordinary PR activity (13 open PRs: 9 dependabot, 4 hand-authored):
self-hosted runners : 1 (online, busy) queued runs : 270 in_progress runs : 3The pool is not degraded and nothing is stuck — that is the steady state. Every one of those 270 runs is waiting on a single machine.
Where the 270 comes from
A single PR head movement creates roughly 25 workflow runs. Observed on one branch in one push:
AI Security Review · PR Template Check · API Wiring Audit · Auto-fix Generated Types CI Dispatch Watchdog · Migration Gate · SLM Frontend Check · Security Scanning PR Blocking Findings · Frontend Testing Suite · Verify Generated Types Hardened Smoke Test · CodeQL Analysis · Code Quality · Enforce Pre-commit Hooks Backend Startup Import Smoke · Auto-fix Code Formatting · frontend-typecheck-regression No Commit Trailers · duplication-guard · Validate PR Issue Link · LLC Contract Gate Docker Smoke Test · SSOT Coverage · AutoBot CI/CD Pipeline25 runs × 13 open PRs ≈ 325, and each
synchronizemultiplies it again.AutoBot CI/CD Pipelinealone now fans out to 12python-suiteshard jobs (#13300), which is a large throughput win per-PR but 12 runner slots rather than 1.So the arithmetic is: sharding made each suite ~6× faster in wall-clock while making it 12× more expensive in runner slots. On a multi-runner pool that trade is straightforwardly good. On a pool of one it converts latency into queue depth.
Consequences observed today
- Every open PR reports
mergeStateStatus=BLOCKED, none on merit — they are waiting for slots. - Merge throughput fell to roughly 6 PRs in a working morning while the queue grew.
- Shard wall-times inflate under contention: the same shard measured 82.49s in a quiet window and 132.55s while the queue was ~270 deep. That in turn fails wall-clock-threshold tests (
system_benchmarks_performance_test.pyasserts fixed millisecond budgets), producing failures that are indistinguishable from real regressions. This actively corrupts the ratchet signal PRs are gated on.
That last point is the part worth escalating: contention is no longer only a delay, it is a source of false test failures.
Cheapest levers, roughly in order
- Move workflows that need no self-hosted resource to
ubuntu-latest. Several in the list above are pure metadata checks (No Commit Trailers,PR Template Check,Validate PR Issue Link,duplication-guard). The dispatch watchdog already does this deliberately and documents why: "a watchdog that queues behind the outage it exists to report is useless." The same reasoning applies to every check that only reads the diff. - Add a second self-hosted runner — the SPOF fix this issue was opened for; it also halves the contention-induced flake.
- Give wall-clock-threshold tests a contention-tolerant budget or a marker, so a loaded runner cannot manufacture failures (overlaps test: file_locking_test asserts a 0.04s wall-clock budget — fails 2 of 13 runs on unmodified base #13366, which found a 0.04s budget failing 2 of 3 runs).
- Path-filter the heavy suites so a frontend-only or docs-only PR does not spend 12 shard slots.
Items 1 and 4 need no hardware and no owner spend.
Related
- ci: runner outage leaves PRs indefinitely pending with a required context that never reports — reads as '19 success, 0 failures' #13045 — runner outage leaves PRs indefinitely pending (same pool, failure mode rather than capacity)
- python-suite aborts after the first pytest: autobot-slm-backend tests and the 70% coverage gate never run #13300 — the 12-way shard split that traded slots for latency
- test: file_locking_test asserts a 0.04s wall-clock budget — fails 2 of 13 runs on unmodified base #13366 — wall-clock budget already shown to be contention-sensitive
- ci: auto-update-pr-branches leaves refreshed PRs with ZERO checks — GITHUB_TOKEN pushes do not trigger workflows #12823 — parking; unrelated cause, but it compounds queue depth
- Every open PR reports
Correction to the comment above — the queue is not the self-hosted runner
I asserted that the 12-way shard split costs 12 self-hosted slots and that several metadata checks sit on the self-hosted runner. Both are wrong, and the fix directions that followed from them would have been wasted work. Checked properly:
Only three workflows have a
self-hostedruns-on:auto-update-pr-branches.yml [self-hosted, linux] code-quality.yml [self-hosted, Linux, X64] frontend-test.yml [self-hosted, Linux, X64]Everything else in the repo — 79 of 84
runs-ondeclarations — isubuntu-latest.ci.ymlis entirely GitHub-hosted, by explicit design.ci.yml:58-61says so in a comment written when the shards were introduced:This runs on ubuntu-latest, NOT the singleton [self-hosted, Linux, X64] … concurrent shards therefore cost the self-hosted queue nothing
So the shard split traded latency for GitHub-hosted slots, not for self-hosted slots.
No Commit Trailersis likewise alreadyubuntu-latestand documents the reason at the top of the file. My "move metadata checks off self-hosted" lever does not exist — that work was already done.What the 124-deep queue actually consists of:
15 Security Scanning (ubuntu-latest) 15 AutoBot CI/CD Pipeline (ubuntu-latest) 9 Frontend Testing Suite (SELF-HOSTED) 6 Code Quality (SELF-HOSTED) 5 Verify Generated Types / Validate PR Issue Link / SSOT Coverage / Backend Startup Import Smoke / API Wiring Audit / AI Security Review 4 SLM Frontend Check / No Commit Trailers / Hardened Smoke Test / Docker Smoke TestRoughly 15 of 124 queued runs want the self-hosted runner. The rest are queued behind the account's GitHub-hosted concurrency limit, which is a different constraint with different remedies.
What stands and what does not
Still true, and unaffected by the above:
- One self-hosted runner remains a SPOF for
frontend-testandcode-quality, andUnit & Integration Testsinsidefrontend-testis a required check — so a single wedged job there still blocks every merge in the repo. That is this issue's original point and today's snapshot supports it. - Contention still inflates wall-clock and manufactures failures in threshold tests: the same shard measured 82.49s quiet vs 132.55s loaded, and
system_benchmarks_performance_test.pyfailed three fixed-millisecond budgets under load. That happens on GitHub-hosted runners too, so it is not evidence about this pool.
No longer supported by evidence:
move metadata checks off self-hosted— already donepath-filter the heavy suites to save self-hosted slots— they never used them
The genuine lever for the 124-deep queue is per-PR run count (~25 workflows per head movement, re-triggered on every
synchronize) against the GitHub-hosted concurrency cap — worth its own issue rather than being tracked here, since it has nothing to do with the SPOF.Apologies for the noise; the original comment's headline numbers (1 runner, 270 queued, 3 in-progress) were real, but the attribution was not.
- One self-hosted runner remains a SPOF for
Filed as #13388 — per-PR run count against GitHub-hosted concurrency, with the path-filtering directions and the threshold-test corruption evidence. Keeping this issue scoped to the self-hosted SPOF (
frontend-test+code-quality, ~15 of 124 queued runs today).Cross-link: #13801 covers a distinct mechanic on the same runner —
required_status_checks.strict: truecombined with a 62-check run that outlasts the ~31min average gap between merges onDev_new_gui, so a green run is routinely invalidated and restarted while the runner is online and healthy. Contention rather than outage, so it is not addressed by this issue; filed separately with measurements.- addedarea: ci-gatesWave 0 · cluster P — CI gates & merge integrityWave 0 · cluster P — CI gates & merge integrity
on Sep 1, 2026
Summary
All required checks (smoke-test, code-quality, startup-import-smoke, Unit & Integration, migration-matrix) run on ONE self-hosted runner ("Little-Slave"). When it went offline mid-session, EVERY open PR's checks were stuck "pending" indefinitely — no PR could pass CI until it was manually restarted. A single runner is a single point of failure for the entire merge pipeline.
Impact observed (this session)
Options (infra decision)
Restart=always+ a watchdog timer that re-registers if the agent dies) so an offline runner self-heals quickly.Also separately observed: a required check (
migration-matrix) was silently RED on the base branch for a period ([[migration_sa_enum_ignores_create_type_2026_07_09]]) — add a base-branch CI health signal so a persistently-red required check on the default branch is surfaced, not discovered PR-by-PR.Recommendation
At minimum (2) auto-restart watchdog + (4) offline alert now; (1) second runner when hardware allows.
Related: [[ci_self_hosted_runner_singleton]]