Skip to content

page_build cannot fire for a deploy that never completes — the guard is blind to its worst case #14

Description

@mrveiss

Finding, from a live incident rather than reasoning

A real stale-site event is happening right now and the freshness guard has not reported it.

09:36     PR #11 merged -> main gains assets/social-card.png and the new og:image
09:40:09  pages build and deployment created -> queued
09:40:10  freshness [page_build] fires for the PREVIOUS deploy -> success
...
10:09     deploy still queued, 29 minutes, never started
          card URL: HTTP 404
          live og:image: still /AutoBot-AI/screenshots/01-chat.png
          freshness runs since 09:40:10: NONE

The site has served content that does not match main for ~30 minutes, and the guard built to detect exactly that has been silent throughout.

The mechanism

page_build fires when a Pages build completes. A deploy that never completes never fires it. So the event-driven trigger — the one added in #6/#8 specifically because the cron was too sparse — is structurally blind to the failure mode that matters most: a deploy that hangs or is cancelled holds the divergence open indefinitely, and that is precisely the case where no completion event will ever arrive.

The cron: '7,37 * * * *' backstop is the only path that can catch it, and it is measured as shed roughly three in four (observed gaps of 2h02m and 3h49m). It has not fired in this window either.

The irony is exact: check_freshness.py contains a Deploy stuck branch, with a 30-minute threshold, written for this scenario — and it cannot run, because nothing triggers it.

Why this is not #9

#9 is about latency — the guard runs, but takes ~15 minutes typical / ~40 worst. This is about the guard not running at all for an entire class of failure. A faster trigger does not fix it; no completion-driven trigger can.

Acceptance criteria

  • A divergence caused by a deploy that never completes is reported within a bounded, stated time
  • The mechanism does not depend on the deploy emitting a completion event, since the failure case is defined by its absence
  • The existing Deploy stuck branch in check_freshness.py — already written and unit-tested — actually becomes reachable in production
  • Whatever is chosen, its own blind spots are stated in the workflow file, the way page_build's unverified status was stated before it was proven

Options, none yet chosen

  • Make the cron the primary rather than the backstop and accept its shedding — cheap, but the shedding is the reason it was demoted.
  • Trigger on the Pages deploy being created rather than completing, then re-check after a delay. deployment_status may or may not fire for legacy Pages builds — unverified, and the last two trigger guesses cost one inert workflow_run and one round of measurement.
  • Accept the gap explicitly and document it, as Drift detection takes ~40 minutes end to end — the same size as the window it exists to catch #9 concluded for latency.

No trigger should be swapped in before it is observed firing. #3 already burned workflow_run on reasoning ahead of evidence; the cost of guessing here is a guard that looks like coverage and is not.

Provenance

Found while auditing this session's work, by noticing the guard was silent during an outage it should have caught. The incident is still open at time of filing — the deploy has been queued 29 minutes, one minute short of the threshold it would be measured against if anything were running to measure it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions