Skip to content

CI resilience: single self-hosted runner 'Little-Slave' is a SPOF that blocks ALL PR checks when offline #11380

Description

@mrveiss

Summary

All required checks (smoke-test, code-quality, startup-import-smoke, Unit & Integration, migration-matrix) run on ONE self-hosted runner ("Little-Slave"). When it went offline mid-session, EVERY open PR's checks were stuck "pending" indefinitely — no PR could pass CI until it was manually restarted. A single runner is a single point of failure for the entire merge pipeline.

Impact observed (this session)

  • Runner offline → 4 PRs blocked (BEHIND/BLOCKED with pending Unit&Integration) for an extended window; had to admin-merge frontend PRs whose other required checks had passed before the outage, and wait to restart the runner for the rest.

Options (infra decision)

  1. Add a second self-hosted runner (same labels) so jobs fail over — removes the SPOF. Preferred.
  2. Auto-restart / health-watchdog on the runner host (systemd Restart=always + a watchdog timer that re-registers if the agent dies) so an offline runner self-heals quickly.
  3. Ephemeral runners (register-per-job) so a hung agent can't wedge the queue.
  4. Document a fast manual-recovery runbook + alert when the runner goes offline (so it's noticed in minutes, not after PRs pile up).

Also separately observed: a required check (migration-matrix) was silently RED on the base branch for a period ([[migration_sa_enum_ignores_create_type_2026_07_09]]) — add a base-branch CI health signal so a persistently-red required check on the default branch is surfaced, not discovered PR-by-PR.

Recommendation

At minimum (2) auto-restart watchdog + (4) offline alert now; (1) second runner when hardware allows.
Related: [[ci_self_hosted_runner_singleton]]

Activity

  1. mrveiss commented on Aug 3, 2026

    @mrveiss
    OwnerAuthor

    Measured today — the SPOF is now also a hard throughput ceiling

    Snapshot taken 2026-08-03 during ordinary PR activity (13 open PRs: 9 dependabot, 4 hand-authored):

    self-hosted runners : 1   (online, busy)
    queued runs         : 270
    in_progress runs    : 3
    

    The pool is not degraded and nothing is stuck — that is the steady state. Every one of those 270 runs is waiting on a single machine.

    Where the 270 comes from

    A single PR head movement creates roughly 25 workflow runs. Observed on one branch in one push:

    AI Security Review · PR Template Check · API Wiring Audit · Auto-fix Generated Types
    CI Dispatch Watchdog · Migration Gate · SLM Frontend Check · Security Scanning
    PR Blocking Findings · Frontend Testing Suite · Verify Generated Types
    Hardened Smoke Test · CodeQL Analysis · Code Quality · Enforce Pre-commit Hooks
    Backend Startup Import Smoke · Auto-fix Code Formatting · frontend-typecheck-regression
    No Commit Trailers · duplication-guard · Validate PR Issue Link · LLC Contract Gate
    Docker Smoke Test · SSOT Coverage · AutoBot CI/CD Pipeline
    

    25 runs × 13 open PRs ≈ 325, and each synchronize multiplies it again. AutoBot CI/CD Pipeline alone now fans out to 12 python-suite shard jobs (#13300), which is a large throughput win per-PR but 12 runner slots rather than 1.

    So the arithmetic is: sharding made each suite ~6× faster in wall-clock while making it 12× more expensive in runner slots. On a multi-runner pool that trade is straightforwardly good. On a pool of one it converts latency into queue depth.

    Consequences observed today

    • Every open PR reports mergeStateStatus=BLOCKED, none on merit — they are waiting for slots.
    • Merge throughput fell to roughly 6 PRs in a working morning while the queue grew.
    • Shard wall-times inflate under contention: the same shard measured 82.49s in a quiet window and 132.55s while the queue was ~270 deep. That in turn fails wall-clock-threshold tests (system_benchmarks_performance_test.py asserts fixed millisecond budgets), producing failures that are indistinguishable from real regressions. This actively corrupts the ratchet signal PRs are gated on.

    That last point is the part worth escalating: contention is no longer only a delay, it is a source of false test failures.

    Cheapest levers, roughly in order

    1. Move workflows that need no self-hosted resource to ubuntu-latest. Several in the list above are pure metadata checks (No Commit Trailers, PR Template Check, Validate PR Issue Link, duplication-guard). The dispatch watchdog already does this deliberately and documents why: "a watchdog that queues behind the outage it exists to report is useless." The same reasoning applies to every check that only reads the diff.
    2. Add a second self-hosted runner — the SPOF fix this issue was opened for; it also halves the contention-induced flake.
    3. Give wall-clock-threshold tests a contention-tolerant budget or a marker, so a loaded runner cannot manufacture failures (overlaps test: file_locking_test asserts a 0.04s wall-clock budget — fails 2 of 13 runs on unmodified base #13366, which found a 0.04s budget failing 2 of 3 runs).
    4. Path-filter the heavy suites so a frontend-only or docs-only PR does not spend 12 shard slots.

    Items 1 and 4 need no hardware and no owner spend.

    Related

  2. mrveiss commented on Aug 3, 2026

    @mrveiss
    OwnerAuthor

    Correction to the comment above — the queue is not the self-hosted runner

    I asserted that the 12-way shard split costs 12 self-hosted slots and that several metadata checks sit on the self-hosted runner. Both are wrong, and the fix directions that followed from them would have been wasted work. Checked properly:

    Only three workflows have a self-hosted runs-on:

    auto-update-pr-branches.yml   [self-hosted, linux]
    code-quality.yml              [self-hosted, Linux, X64]
    frontend-test.yml             [self-hosted, Linux, X64]
    

    Everything else in the repo — 79 of 84 runs-on declarations — is ubuntu-latest.

    ci.yml is entirely GitHub-hosted, by explicit design. ci.yml:58-61 says so in a comment written when the shards were introduced:

    This runs on ubuntu-latest, NOT the singleton [self-hosted, Linux, X64] … concurrent shards therefore cost the self-hosted queue nothing

    So the shard split traded latency for GitHub-hosted slots, not for self-hosted slots. No Commit Trailers is likewise already ubuntu-latest and documents the reason at the top of the file. My "move metadata checks off self-hosted" lever does not exist — that work was already done.

    What the 124-deep queue actually consists of:

    15  Security Scanning              (ubuntu-latest)
    15  AutoBot CI/CD Pipeline         (ubuntu-latest)
     9  Frontend Testing Suite         (SELF-HOSTED)
     6  Code Quality                   (SELF-HOSTED)
     5  Verify Generated Types / Validate PR Issue Link / SSOT Coverage /
        Backend Startup Import Smoke / API Wiring Audit / AI Security Review
     4  SLM Frontend Check / No Commit Trailers / Hardened Smoke Test / Docker Smoke Test
    

    Roughly 15 of 124 queued runs want the self-hosted runner. The rest are queued behind the account's GitHub-hosted concurrency limit, which is a different constraint with different remedies.

    What stands and what does not

    Still true, and unaffected by the above:

    • One self-hosted runner remains a SPOF for frontend-test and code-quality, and Unit & Integration Tests inside frontend-test is a required check — so a single wedged job there still blocks every merge in the repo. That is this issue's original point and today's snapshot supports it.
    • Contention still inflates wall-clock and manufactures failures in threshold tests: the same shard measured 82.49s quiet vs 132.55s loaded, and system_benchmarks_performance_test.py failed three fixed-millisecond budgets under load. That happens on GitHub-hosted runners too, so it is not evidence about this pool.

    No longer supported by evidence:

    • move metadata checks off self-hosted — already done
    • path-filter the heavy suites to save self-hosted slots — they never used them

    The genuine lever for the 124-deep queue is per-PR run count (~25 workflows per head movement, re-triggered on every synchronize) against the GitHub-hosted concurrency cap — worth its own issue rather than being tracked here, since it has nothing to do with the SPOF.

    Apologies for the noise; the original comment's headline numbers (1 runner, 270 queued, 3 in-progress) were real, but the attribution was not.

  3. mrveiss commented on Aug 3, 2026

    @mrveiss
    OwnerAuthor

    Filed as #13388 — per-PR run count against GitHub-hosted concurrency, with the path-filtering directions and the threshold-test corruption evidence. Keeping this issue scoped to the self-hosted SPOF (frontend-test + code-quality, ~15 of 124 queued runs today).

  4. mrveiss commented on Aug 9, 2026

    @mrveiss
    OwnerAuthor

    Cross-link: #13801 covers a distinct mechanic on the same runner — required_status_checks.strict: true combined with a 62-check run that outlasts the ~31min average gap between merges on Dev_new_gui, so a green run is routinely invalidated and restarted while the runner is online and healthy. Contention rather than outage, so it is not addressed by this issue; filed separately with measurements.

  5. added
    area: ci-gatesWave 0 · cluster P — CI gates & merge integrity
    on Sep 1, 2026
  6. added this to the v0.10.0 milestone on Sep 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions