Skip to content

ci: auto-update-pr-branches leaves refreshed PRs with ZERO checks — GITHUB_TOKEN pushes do not trigger workflows #12823

Description

@mrveiss

Consequence of #12801/#12818 making auto-update-pr-branches actually work. The workflow now correctly detects and updates stale branches — but the resulting commit gets no CI at all, which is arguably more dangerous than the branch simply staying stale.

Evidence

PR #12821 after the workflow refreshed it:

head       = e822b17ff
author     = github-actions[bot]
message    = Merge branch 'Dev_new_gui' into issue-12774
behind_by  = 0
check-runs = 0
workflow runs for head_sha = 0

gh pr checks returns nothing at all for that SHA — not "pending", not "failing". Any tooling that treats "no failures" as "ready" will call this mergeable. My own monitor did exactly that and reported READY:#12821; the PR was one step from being merged with zero verification.

Cause

GitHub deliberately does not trigger workflow runs from pushes authored with the default GITHUB_TOKEN — it is loop prevention, and it is working as designed. The workflow currently runs with:

env:
  GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}

so PUT /pulls/{n}/update-branch produces a bot-authored merge commit that fires no pull_request: synchronize event.

Net effect: a PR can end up behind_by=0, mergeable, and never CI-verified on the code that would actually land. Branch protection cannot save us here — with no check-runs present there is nothing for it to evaluate as failed.

Options (needs an owner call on the first)

  1. Run the update with a token that does trigger CI — a PAT or GitHub App installation token in a secret. Most faithful fix, but it means creating and storing a credential, so it is an owner decision, not a drive-by.
  2. Re-dispatch CI after updating — have the workflow explicitly kick the required workflows for the updated ref. Needs no new secret, but couples the job to the list of required workflows and only works for workflow_dispatch-enabled ones.
  3. Do not auto-merge on absent checks — independent of the above, tooling must treat zero check-runs as NOT ready. "No failures" is not "verified".

Option 3 should happen regardless of which of 1/2 is chosen.

Interim workaround

Close/reopen the PR, or push any non-bot commit, to make CI run on the current head.

Related

Activity

  1. mrveiss commented on Jul 27, 2026

    @mrveiss
    OwnerAuthor

    Recurred today on PR #12834 (issue #12726) — second confirmed instance, so this is systematic, not a one-off:

    head       = e483e634ef0ea8120ad7ad04b4f7e47186c9c76b
    message    = Merge branch 'Dev_new_gui' into issue-12726
    check-runs = 0
    statusCheckRollup = null
    mergeStateStatus  = BLOCKED
    

    Note statusCheckRollup comes back null, not an empty list — jq filters of the form .statusCheckRollup[] | select(.conclusion != "SUCCESS") error out rather than returning "no failures", which is at least a loud failure mode. But gh pr checks staying silent remains the dangerous one.

    Interim mitigation that works today: gh pr close N && gh pr reopen N re-fires the full suite (29 checks came back on #12834). It needs no new secret and no workflow change, but it is manual and it drops any in-flight run.

    Decision still open — the two real fixes:

    1. PAT (secrets.PR_UPDATE_TOKEN) for the update-branch call. Cleanest: the merge commit is authored by a real user so synchronize fires and checks run exactly as on a normal push. Cost: a long-lived credential with repo scope that has to be rotated.
    2. Explicit re-dispatch after updating: workflow_dispatch each required workflow against the new head. No new secret, but each required workflow needs a workflow_dispatch trigger, and the run is not linked to the pull_request event — so required-check names may not match what branch protection expects, which risks trading a silent gap for a permanent block.

    Recommend option 1 — option 2's check-name mismatch against branch protection is the kind of thing that fails only at merge time.

  2. mrveiss commented on Jul 28, 2026

    @mrveiss
    OwnerAuthor

    Same root cause produces a second, separate symptom: "workflows awaiting approval"

    This issue tracks bot branch-updates producing zero checks. There is a sibling symptom from the same identity problem: after the auto-fix workflows push to a PR branch, every workflow run on the new SHA lands in action_required and has to be hand-approved (~20 runs per occurrence).

    Evidence from PR #12888:

    head commit author : github-actions[bot]
    run                : Auto-fix Generated Types
    event              : pull_request
    actor              : github-actions[bot]
    triggering_actor   : github-actions[bot]
    status=completed   conclusion=action_required     (20 of 23 runs)
    

    Correlated across recent PRs — a human-authored head needs no approval, a bot-authored one does:

    PR 12889  head author=mrveiss              action_required=0/20
    PR 12888  head author=github-actions[bot]  action_required=20/23
    PR 12882  head author=mrveiss              action_required=0/21
    

    The pushers:

    .github/workflows/auto-fix-generated-types.yml:37   token: ${{ secrets.GITHUB_TOKEN }}
                                                  :84   git config user.name "github-actions[bot]"
                                                  :97   git push
    .github/workflows/auto-fix-formatting.yml           same shape
    

    So github-actions[bot] commits and pushes to the PR branch; the resulting pull_request event has the bot as actor, and because that identity is not a repo collaborator every run on that SHA requires approval.

    Why this matters for the decision here

    Option 1 (PAT) fixes both symptoms, not just this issue's. Pushing as a real user means the branch-update event both triggers workflows (this issue) and does not require approval (the sibling symptom). Option 2 (workflow_dispatch re-dispatch) fixes only the zero-checks half and leaves the approval friction untouched.

    A third option exists specifically for the approval half: relax Settings → Actions → General → "Approval for running fork pull request workflows from contributors". That is a repo setting rather than a code change, and it weakens a guard globally rather than fixing the identity — so it is worth knowing about but is not my recommendation.

    Measured cost of the status quo: I have hand-approved runs on roughly a dozen PRs this session, ~20 runs each.

  3. mrveiss commented on Jul 30, 2026

    @mrveiss
    OwnerAuthor

    Related sibling filed as #13045: a second route to the same zero-checks end state, this one caused by the self-hosted runner being offline rather than by GITHUB_TOKEN bot pushes.

    Shared root symptom worth solving once: a required context that never reports at all leaves the PR at pending with no visible failure, and any tooling that counts success/failure reads it as clean. In the observed case both PRs showed 19 success / 0 failures while being unmergeable.

  4. mrveiss commented on Aug 2, 2026

    @mrveiss
    OwnerAuthor

    Recurred today at scale, and with a second symptom shape worth pinning to this issue.

    Merging three PRs to Dev_new_gui this afternoon left 92 workflow runs parked in action_required on the one remaining open PR — created, but never dispatched pending manual approval. gh pr view --json statusCheckRollup reports that as null, which is the same observable as the zero-runs case in the original report but a different underlying state.

    That distinction matters for the fix: "no runs created" and "runs created but not dispatched" do not necessarily respond the same way to pushing with a PAT/GitHub App token (option 2). Whichever route is taken should be verified against both.

    Manual unblock for the parked shape:

    gh api "repos/{owner}/{repo}/actions/runs?branch=<branch>&status=action_required" --jq '.workflow_runs[].id' \
      | xargs -I{} gh api -X POST "repos/{owner}/{repo}/actions/runs/{}/approve"
    

    Not safe to run repo-wide — most parked runs belong to branches with no open PR, and approving them queues ahead of live work.

    Cost measured today: three PRs sat several hours reading as "waiting for CI" when nothing had been queued. This is the single largest throughput drag on the PR pipeline right now. Duplicate #13294 closed into this one; its fix options are folded in there.

  5. 5 remaining items

  6. added this to the Backlog milestone on Sep 12, 2026
  7. mrveiss commented on Sep 12, 2026

    @mrveiss
    OwnerAuthor

    Stale-issue sweep (never auto-closed — Dev_new_gui wasn't GitHub's default branch at merge time). Verified against current origin/main and a live check right now; leaving open — the defect is reproducing today.

    No merged PR claims closure, and live evidence confirms the defect persists.

  8. modified the milestones: Backlog, v0.14.0 on Sep 14, 2026
  9. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    It happened on a merge, 2026-09-18: #17051 merged with zero required checks at its head

    • The fix(auth): attribute approval decisions to the verified human caller (#17042) #17051 head moved to 4e3a1cd1a, a chore(types): regenerate generated API types commit by github-actions[bot]. The push used the workflow token, so no workflows ran on that SHA. commits/4e3a1cd1a/check-runs holds none of main's 10 required contexts; the only status is CodeRabbit.
    • The PR was merged at 20:07:25Z as 27acbd1e8. How it passed branch protection with the required contexts missing is being established; it will be recorded here.
    • A coordinator's CI monitor counted the empty rollup (0 pending, 0 failures) as green. That monitor is fixed: green now requires a completed success run of every required context at the exact head SHA.
    • Main's own post-merge CI on 27acbd1e8 is queued and is the first real verification of fix(auth): attribute approval decisions to the verified human caller (#17042) #17051's code.

    That moves this issue from a nuisance to a merge-gate hole. Every regen, auto-fix or auto-update bot push can leave a PR head that looks mergeable and has never been tested. The fix direction in this issue (push with a token that triggers workflows, or dispatch CI after the push) is now needed for v0.9.0.

  10. modified the milestones: v0.14.0, v0.9.0 on Sep 18, 2026
  11. mrveiss commented on Sep 25, 2026

    @mrveiss
    OwnerAuthor

    Option 3 is already implemented — verified against merged main. What remains is options 1/2, which need an owner decision, plus one duplicate worth collapsing.

    Option 3 — "tooling must treat zero check-runs as NOT ready" — done

    scripts/pr_required_gate.py buckets required contexts into never_reported / running / not_green / green, returns never_reported in its result, and prints each one as a blocker. scripts/lib/check_run_status.py:split_by_state carries the rule explicitly:

    "never_reported is kept apart deliberately. A context nothing published is not a passing context and not a failing one; collapsing it into either is how a merge gate reports a green it did not earn."

    The gate also guards against the failure mode in itself: it materialises required with list(required) because a generator would be exhausted by the first read and "every required context would be reclassified as unrequired — the failure mode this whole tool exists to catch, in the tool."

    So a PR refreshed by the bot, with zero check-runs, reads as blocked by this gate rather than ready.

    What is still open, and it is not option 3

    Options 1 and 2 both need a decision I should not make. Option 1 requires creating and storing a PAT or GitHub App installation token — a credential, which in this repo goes through the canonical secrets manager and is never introduced by an agent on its own initiative. Option 2 couples the workflow to the list of required workflows and only reaches workflow_dispatch-enabled ones, which is a design trade rather than a fix. Flagging both as owner calls, as the issue already says of the first.

    One real residual: two implementations of "what counts as ready"

    split_by_state lives in scripts/lib/check_run_status.py and nothing in production calls it — pr_required_gate.py imports ACCEPTABLE, RUNNING, all_pages and latest_per_name from that module but keeps its own _split_required. The library's own comment is the argument against that state:

    "keeping a second copy means two implementations that must be kept in sync by hand, which is how the naive query gets written again by whoever reads only one of them."

    Same shape as #17468, where a named constant and a hardcoded literal agreed until someone changed one. Both bucketings are currently correct, which is exactly when the duplication is invisible.

    I have not collapsed them in this pass, deliberately: pr_required_gate.py is the gate every session's merge decision currently depends on, and a behaviour-preserving refactor of it during a serialized-CI freeze is the wrong risk at the wrong time. It wants its own change, with an equivalence test between the two implementations before either is deleted.

    One correction to my own method, since it bears on the evidence

    My first pass grepped split_by_state and concluded it had no importer at all. That was wrong — pr_required_gate.py imports from check_run_status across a multi-line from ... ( block, which a single-line grep cannot see. The conclusion "unwired" would have been a confident answer produced by an instrument that could not see the import. Checked before reporting; recording it because the same grep shape will mislead the next person.

  12. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Chunking triage — a proposal, not an assignment

    Nothing was relabelled, moved or closed by this pass.

  13. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Pre-filter (read-only, v0.9.5): partial — the fix is conditional on a secret, and the fallback restores the defect silently

    Anchored at origin/main ab05ba3aef. One commit carries (#12823): 984b061837 2026-08-03 fix(ci): sweep the dispatch watchdog on PR events, align the PR queue limit (#12823) (#13318) — a mitigation (the watchdog sweeps), not a cause fix.

    The cause is addressed elsewhere in the same workflow, and conditionally:

    .github/workflows/auto-update-pr-branches.yml:133   GH_TOKEN: ${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }}
                                                :355   GITHUB_TOKEN: ${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }}
                                                :131   "#13791: see auto-fix-generated-types.yml — a human-attributed token"
    

    A human-attributed token makes the push trigger workflows, which is the fix. The || fallback means that when AUTOBOT_PUSH_TOKEN is absent the original defect returns, and nothing says so — the run looks identical and the refreshed PR simply has no checks, which is the failure this issue describes.

    Whether the secret is configured is host state, and I did not go looking for it — a stated gap beats a guess about a credential. Two things would close this without needing that answer: log which token was used, and fail the job loudly when only GITHUB_TOKEN is available rather than proceeding into the known-broken path. Then the premise is decided by the run's own output instead of by repository archaeology.

  14. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Host-state half of the premise check, measured (names only, no values): gh secret list on this repository shows no repository-level secret named AUTOBOT_PUSH_TOKEN (0 matches, 2026-09-28 20:58 Riga). Organization-level and environment-level secrets were not checked — if one exists there it would satisfy the expression. At repository scope, .github/workflows/auto-update-pr-branches.yml:133 and :355 (${{ secrets.AUTOBOT_PUSH_TOKEN || secrets.GITHUB_TOKEN }}) therefore resolve to GITHUB_TOKEN today, i.e. the || fallback silently restores the defect this issue describes. Fix shape: fail loudly when the dedicated token is absent (no fallback), or provision the secret and drop the fallback — a fallback that recreates the bug is not a fallback. Read-only finding by 87, measurement by the coordinating session.

  15. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Cross-link: same cause and same one-decision fix as #17728 (bot-authored changelog PR held by the external-contributor ruleset). See the comment there; the provisioning decision is raised once, for both.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions