ci: auto-update-pr-branches leaves refreshed PRs with ZERO checks — GITHUB_TOKEN pushes do not trigger workflows #12823
Description
Activity
Recurred today on PR #12834 (issue #12726) — second confirmed instance, so this is systematic, not a one-off:
head = e483e634ef0ea8120ad7ad04b4f7e47186c9c76b message = Merge branch 'Dev_new_gui' into issue-12726 check-runs = 0 statusCheckRollup = null mergeStateStatus = BLOCKEDNote
statusCheckRollupcomes back null, not an empty list — jq filters of the form.statusCheckRollup[] | select(.conclusion != "SUCCESS")error out rather than returning "no failures", which is at least a loud failure mode. Butgh pr checksstaying silent remains the dangerous one.Interim mitigation that works today:
gh pr close N && gh pr reopen Nre-fires the full suite (29 checks came back on #12834). It needs no new secret and no workflow change, but it is manual and it drops any in-flight run.Decision still open — the two real fixes:
- PAT (
secrets.PR_UPDATE_TOKEN) for theupdate-branchcall. Cleanest: the merge commit is authored by a real user sosynchronizefires and checks run exactly as on a normal push. Cost: a long-lived credential withreposcope that has to be rotated. - Explicit re-dispatch after updating:
workflow_dispatcheach required workflow against the new head. No new secret, but each required workflow needs aworkflow_dispatchtrigger, and the run is not linked to thepull_requestevent — so required-check names may not match what branch protection expects, which risks trading a silent gap for a permanent block.
Recommend option 1 — option 2's check-name mismatch against branch protection is the kind of thing that fails only at merge time.
- PAT (
Same root cause produces a second, separate symptom: "workflows awaiting approval"
This issue tracks bot branch-updates producing zero checks. There is a sibling symptom from the same identity problem: after the auto-fix workflows push to a PR branch, every workflow run on the new SHA lands in
action_requiredand has to be hand-approved (~20 runs per occurrence).Evidence from PR #12888:
head commit author : github-actions[bot] run : Auto-fix Generated Types event : pull_request actor : github-actions[bot] triggering_actor : github-actions[bot] status=completed conclusion=action_required (20 of 23 runs)Correlated across recent PRs — a human-authored head needs no approval, a bot-authored one does:
PR 12889 head author=mrveiss action_required=0/20 PR 12888 head author=github-actions[bot] action_required=20/23 PR 12882 head author=mrveiss action_required=0/21The pushers:
.github/workflows/auto-fix-generated-types.yml:37 token: ${{ secrets.GITHUB_TOKEN }} :84 git config user.name "github-actions[bot]" :97 git push .github/workflows/auto-fix-formatting.yml same shapeSo
github-actions[bot]commits and pushes to the PR branch; the resultingpull_requestevent has the bot as actor, and because that identity is not a repo collaborator every run on that SHA requires approval.Why this matters for the decision here
Option 1 (PAT) fixes both symptoms, not just this issue's. Pushing as a real user means the branch-update event both triggers workflows (this issue) and does not require approval (the sibling symptom). Option 2 (
workflow_dispatchre-dispatch) fixes only the zero-checks half and leaves the approval friction untouched.A third option exists specifically for the approval half: relax Settings → Actions → General → "Approval for running fork pull request workflows from contributors". That is a repo setting rather than a code change, and it weakens a guard globally rather than fixing the identity — so it is worth knowing about but is not my recommendation.
Measured cost of the status quo: I have hand-approved runs on roughly a dozen PRs this session, ~20 runs each.
Related sibling filed as #13045: a second route to the same zero-checks end state, this one caused by the self-hosted runner being offline rather than by
GITHUB_TOKENbot pushes.Shared root symptom worth solving once: a required context that never reports at all leaves the PR at
pendingwith no visible failure, and any tooling that counts success/failure reads it as clean. In the observed case both PRs showed 19 success / 0 failures while being unmergeable.Recurred today at scale, and with a second symptom shape worth pinning to this issue.
Merging three PRs to
Dev_new_guithis afternoon left 92 workflow runs parked inaction_requiredon the one remaining open PR — created, but never dispatched pending manual approval.gh pr view --json statusCheckRollupreports that asnull, which is the same observable as the zero-runs case in the original report but a different underlying state.That distinction matters for the fix: "no runs created" and "runs created but not dispatched" do not necessarily respond the same way to pushing with a PAT/GitHub App token (option 2). Whichever route is taken should be verified against both.
Manual unblock for the parked shape:
gh api "repos/{owner}/{repo}/actions/runs?branch=<branch>&status=action_required" --jq '.workflow_runs[].id' \ | xargs -I{} gh api -X POST "repos/{owner}/{repo}/actions/runs/{}/approve"Not safe to run repo-wide — most parked runs belong to branches with no open PR, and approving them queues ahead of live work.
Cost measured today: three PRs sat several hours reading as "waiting for CI" when nothing had been queued. This is the single largest throughput drag on the PR pipeline right now. Duplicate #13294 closed into this one; its fix options are folded in there.
- added 8 commits that reference this issue
on Aug 2, 2026 5 remaining items
Stale-issue sweep (never auto-closed — Dev_new_gui wasn't GitHub's default branch at merge time). Verified against current
origin/mainand a live check right now; leaving open — the defect is reproducing today.- PR fix(ci): sweep the dispatch watchdog on PR events, align the PR queue limit (#12823) #13318 explicitly disclaims closure: "does not fix ci: auto-update-pr-branches leaves refreshed PRs with ZERO checks — GITHUB_TOKEN pushes do not trigger workflows #12823... Refs, not Closes."
- The foundational mechanism landed via unlisted fix(ci): repair the four ways CI silently never runs (#12823, #13300, #13286, #13045) #13304, whose own body says "treat ci: auto-update-pr-branches leaves refreshed PRs with ZERO checks — GITHUB_TOKEN pushes do not trigger workflows #12823 as unproven."
- Live check during this audit:
gh api .../actions/runs?status=action_requiredreturned 238 parked runs repo-wide right now, including the watchdog workflow's own most recent runs — the exact symptom this issue describes.
No merged PR claims closure, and live evidence confirms the defect persists.
It happened on a merge, 2026-09-18: #17051 merged with zero required checks at its head
- The fix(auth): attribute approval decisions to the verified human caller (#17042) #17051 head moved to
4e3a1cd1a, achore(types): regenerate generated API typescommit bygithub-actions[bot]. The push used the workflow token, so no workflows ran on that SHA.commits/4e3a1cd1a/check-runsholds none of main's 10 required contexts; the only status is CodeRabbit. - The PR was merged at 20:07:25Z as
27acbd1e8. How it passed branch protection with the required contexts missing is being established; it will be recorded here. - A coordinator's CI monitor counted the empty rollup (0 pending, 0 failures) as green. That monitor is fixed: green now requires a completed
successrun of every required context at the exact head SHA. - Main's own post-merge CI on
27acbd1e8is queued and is the first real verification of fix(auth): attribute approval decisions to the verified human caller (#17042) #17051's code.
That moves this issue from a nuisance to a merge-gate hole. Every regen, auto-fix or auto-update bot push can leave a PR head that looks mergeable and has never been tested. The fix direction in this issue (push with a token that triggers workflows, or dispatch CI after the push) is now needed for v0.9.0.
- The fix(auth): attribute approval decisions to the verified human caller (#17042) #17051 head moved to
Option 3 is already implemented — verified against merged
main. What remains is options 1/2, which need an owner decision, plus one duplicate worth collapsing.Option 3 — "tooling must treat zero check-runs as NOT ready" — done
scripts/pr_required_gate.pybuckets required contexts intonever_reported/running/not_green/green, returnsnever_reportedin its result, and prints each one as a blocker.scripts/lib/check_run_status.py:split_by_statecarries the rule explicitly:"
never_reportedis kept apart deliberately. A context nothing published is not a passing context and not a failing one; collapsing it into either is how a merge gate reports a green it did not earn."The gate also guards against the failure mode in itself: it materialises
requiredwithlist(required)because a generator would be exhausted by the first read and "every required context would be reclassified as unrequired — the failure mode this whole tool exists to catch, in the tool."So a PR refreshed by the bot, with zero check-runs, reads as blocked by this gate rather than ready.
What is still open, and it is not option 3
Options 1 and 2 both need a decision I should not make. Option 1 requires creating and storing a PAT or GitHub App installation token — a credential, which in this repo goes through the canonical secrets manager and is never introduced by an agent on its own initiative. Option 2 couples the workflow to the list of required workflows and only reaches
workflow_dispatch-enabled ones, which is a design trade rather than a fix. Flagging both as owner calls, as the issue already says of the first.One real residual: two implementations of "what counts as ready"
split_by_statelives inscripts/lib/check_run_status.pyand nothing in production calls it —pr_required_gate.pyimportsACCEPTABLE,RUNNING,all_pagesandlatest_per_namefrom that module but keeps its own_split_required. The library's own comment is the argument against that state:"keeping a second copy means two implementations that must be kept in sync by hand, which is how the naive query gets written again by whoever reads only one of them."
Same shape as #17468, where a named constant and a hardcoded literal agreed until someone changed one. Both bucketings are currently correct, which is exactly when the duplication is invisible.
I have not collapsed them in this pass, deliberately:
pr_required_gate.pyis the gate every session's merge decision currently depends on, and a behaviour-preserving refactor of it during a serialized-CI freeze is the wrong risk at the wrong time. It wants its own change, with an equivalence test between the two implementations before either is deleted.One correction to my own method, since it bears on the evidence
My first pass grepped
split_by_stateand concluded it had no importer at all. That was wrong —pr_required_gate.pyimports fromcheck_run_statusacross a multi-linefrom ... (block, which a single-line grep cannot see. The conclusion "unwired" would have been a confident answer produced by an instrument that could not see the import. Checked before reporting; recording it because the same grep shape will mislead the next person.Chunking triage — a proposal, not an assignment
- Proposed priority:
priority: low— not applied — and the priority is moot if the recommended move happens first. - triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639 criterion met: none of the five. Placed as misfiled: not blocking and not "landing now". triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639 recommends moving it out of v0.9.0.
- triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639 bucket: 4 — misfiled (not blocking and not "landing now")
- Scope:
ci (secondary hooks) - Umbrella / container: no.
- Pre-filter: clean — no merged commit on
origin/mainsince 2026-08-14 references this issue. - Basis: triage(v0.9.0): release criticality of the 123 open issues — 26 blocking, 36 small, 18 umbrellas, 41 misfiled, 2 undetermined #17639's per-issue row, reused rather than re-derived — "needs ruling on options 1/2 (credential-bearing token vs re-dispatch); option 3 (zero checks read as not ready) is on main in scripts/pr_required_gate.py, closing the merge-gate hole; owner comment 2026-09-18 asked for i"
Nothing was relabelled, moved or closed by this pass.
- Proposed priority:
Pre-filter (read-only, v0.9.5): partial — the fix is conditional on a secret, and the fallback restores the defect silently
Anchored at
origin/mainab05ba3aef. One commit carries(#12823):984b061837 2026-08-03 fix(ci): sweep the dispatch watchdog on PR events, align the PR queue limit (#12823) (#13318)— a mitigation (the watchdog sweeps), not a cause fix.The cause is addressed elsewhere in the same workflow, and conditionally:
.github/workflows/auto-update-pr-branches.yml:133 GH_TOKEN: ${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }} :355 GITHUB_TOKEN: ${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }} :131 "#13791: see auto-fix-generated-types.yml — a human-attributed token"A human-attributed token makes the push trigger workflows, which is the fix. The
||fallback means that whenAUTOBOT_PUSH_TOKENis absent the original defect returns, and nothing says so — the run looks identical and the refreshed PR simply has no checks, which is the failure this issue describes.Whether the secret is configured is host state, and I did not go looking for it — a stated gap beats a guess about a credential. Two things would close this without needing that answer: log which token was used, and fail the job loudly when only
GITHUB_TOKENis available rather than proceeding into the known-broken path. Then the premise is decided by the run's own output instead of by repository archaeology.Host-state half of the premise check, measured (names only, no values):
gh secret liston this repository shows no repository-level secret namedAUTOBOT_PUSH_TOKEN(0 matches, 2026-09-28 20:58 Riga). Organization-level and environment-level secrets were not checked — if one exists there it would satisfy the expression. At repository scope,.github/workflows/auto-update-pr-branches.yml:133and:355(${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }}) therefore resolve toGITHUB_TOKENtoday, i.e. the||fallback silently restores the defect this issue describes. Fix shape: fail loudly when the dedicated token is absent (no fallback), or provision the secret and drop the fallback — a fallback that recreates the bug is not a fallback. Read-only finding by 87, measurement by the coordinating session.Cross-link: same cause and same one-decision fix as #17728 (bot-authored changelog PR held by the external-contributor ruleset). See the comment there; the provisioning decision is raised once, for both.
Consequence of #12801/#12818 making
auto-update-pr-branchesactually work. The workflow now correctly detects and updates stale branches — but the resulting commit gets no CI at all, which is arguably more dangerous than the branch simply staying stale.Evidence
PR #12821 after the workflow refreshed it:
gh pr checksreturns nothing at all for that SHA — not "pending", not "failing". Any tooling that treats "no failures" as "ready" will call this mergeable. My own monitor did exactly that and reportedREADY:#12821; the PR was one step from being merged with zero verification.Cause
GitHub deliberately does not trigger workflow runs from pushes authored with the default
GITHUB_TOKEN— it is loop prevention, and it is working as designed. The workflow currently runs with:so
PUT /pulls/{n}/update-branchproduces a bot-authored merge commit that fires nopull_request: synchronizeevent.Net effect: a PR can end up
behind_by=0, mergeable, and never CI-verified on the code that would actually land. Branch protection cannot save us here — with no check-runs present there is nothing for it to evaluate as failed.Options (needs an owner call on the first)
workflow_dispatch-enabled ones.Option 3 should happen regardless of which of 1/2 is chosen.
Interim workaround
Close/reopen the PR, or push any non-bot commit, to make CI run on the current head.
Related