Skip to content

ci: runner outage leaves PRs indefinitely pending with a required context that never reports — reads as '19 success, 0 failures' #13045

Description

@mrveiss

Sibling of #12823. That issue covers PRs left with zero checks because GITHUB_TOKEN bot pushes do not trigger workflows. This is a second, distinct route to the same dangerous end state: a PR that reads as green but is structurally unmergeable.

Observed 2026-07-28 on PRs #12895 and #12896

While the singleton self-hosted runner was offline:

runner:  Little-Slave  status=offline  busy=false  labels=self-hosted,Linux,X64
run:     status=pending  conclusion=null  jobs=[]        <- zero jobs ever created

Both PRs reported 19 success / 0 failures across every check that ran, while being blocked the whole time: the required context Unit & Integration Tests had produced no check-run at all, so commits/{sha}/status sat at pending.

Three properties made this hard to diagnose, and each is independently worth guarding against:

  1. Absence is not failure. Any tooling that summarises success/failure counts reports these PRs as clean. The only way to see the block is to diff the reported contexts against branches/<base>/protection.required_status_checks.contexts.
  2. The approval sweep does not catch it. The standard remedy searches for conclusion == "action_required". These runs are status=pending, conclusion=null, and POST /actions/runs/{id}/approve returns 403 not waiting for approval.
  3. Runs created during the outage never self-schedule. After the runner came back, both runs went to completed/cancelled with jobs: 0. gh run rerun cannot help a run that never created jobs — it has nothing to re-run.

frontend-test.yml is correctly written for the ordinary case: it deliberately carries no pull_request.paths filter and ships a unit-tests-skip shim (same name:, ubuntu-latest) precisely so the required context always reports. That design is defeated when the runner hosting the real job is offline, because the run never reaches the point of choosing a branch.

Scope

  • Give the required-context shim a path that does not depend on the self-hosted runner, so the context reports even when the runner pool is empty.
  • Or: surface the condition — a check that compares reported contexts against the required list and fails loudly with "required context X never reported", instead of leaving the PR silently at pending.

Done when

  • A runner outage produces a visible, named failure on affected PRs rather than an indefinite pending with no explanation.
  • The diagnosis does not require manually diffing check-runs against branch protection.

Refs #12823

Activity

  1. mrveiss commented on Aug 9, 2026

    @mrveiss
    OwnerAuthor

    Cross-link: #13801 covers a distinct mechanic on the same runner — required_status_checks.strict: true combined with a 62-check run that outlasts the ~31min average gap between merges on Dev_new_gui, so a green run is routinely invalidated and restarted while the runner is online and healthy. Contention rather than outage, so it is not addressed by this issue; filed separately with measurements.

  2. added
    area: ci-gatesWave 0 · cluster P — CI gates & merge integrity
    on Sep 1, 2026
  3. added this to the v0.10.0 milestone on Sep 12, 2026
  4. mrveiss commented on Oct 4, 2026

    @mrveiss
    OwnerAuthor

    Closing — verified against merged main @ 8bd2d70b04.

    The required context can be published by a GitHub-hosted runner. .github/workflows/frontend-required-context.yml:84 declares name: Unit & Integration Tests, matching frontend-test.yml:109 exactly — the required status context is the job name, so the complement reports the same context when the real suite legitimately does not run. Repo-wide there are zero runs-on: values naming a self-hosted runner (control: an unanchored self-hosted grep returns 20+ hits, all in comments, so the search works).

    A runner outage now produces a named failure rather than indefinite pending. pipeline-scripts/ci_dispatch_labels.py:140 starved_verdict classifies it, and it is published as a real status context — ci_dispatch_watchdog.py:209 DEFAULT_STATUS_CONTEXT = "ci-dispatch-watchdog" via set_status at :759-766, with the schedule live at ci-dispatch-watchdog.yml:62.

    Stated limit: I verified the code and the wiring, not that the watchdog cron has actually fired and published. If you want run history on the record, one gh run list --workflow=ci-dispatch-watchdog.yml would settle it — the closure here rests on the mechanism being in place, which is what the criteria name.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions