You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
page_build cannot fire for a deploy that never completes — the guard is blind to its worst case #14
Finding, from a live incident rather than reasoning
A real stale-site event is happening right now and the freshness guard has not reported it.
09:36 PR #11 merged -> main gains assets/social-card.png and the new og:image
09:40:09 pages build and deployment created -> queued
09:40:10 freshness [page_build] fires for the PREVIOUS deploy -> success
...
10:09 deploy still queued, 29 minutes, never started
card URL: HTTP 404
live og:image: still /AutoBot-AI/screenshots/01-chat.png
freshness runs since 09:40:10: NONE
The site has served content that does not match main for ~30 minutes, and the guard built to detect exactly that has been silent throughout.
The mechanism
page_build fires when a Pages build completes. A deploy that never completes never fires it. So the event-driven trigger — the one added in #6/#8 specifically because the cron was too sparse — is structurally blind to the failure mode that matters most: a deploy that hangs or is cancelled holds the divergence open indefinitely, and that is precisely the case where no completion event will ever arrive.
The cron: '7,37 * * * *' backstop is the only path that can catch it, and it is measured as shed roughly three in four (observed gaps of 2h02m and 3h49m). It has not fired in this window either.
The irony is exact: check_freshness.py contains a Deploy stuck branch, with a 30-minute threshold, written for this scenario — and it cannot run, because nothing triggers it.
#9 is about latency — the guard runs, but takes ~15 minutes typical / ~40 worst. This is about the guard not running at all for an entire class of failure. A faster trigger does not fix it; no completion-driven trigger can.
Acceptance criteria
A divergence caused by a deploy that never completes is reported within a bounded, stated time
The mechanism does not depend on the deploy emitting a completion event, since the failure case is defined by its absence
The existing Deploy stuck branch in check_freshness.py — already written and unit-tested — actually becomes reachable in production
Whatever is chosen, its own blind spots are stated in the workflow file, the way page_build's unverified status was stated before it was proven
Options, none yet chosen
Make the cron the primary rather than the backstop and accept its shedding — cheap, but the shedding is the reason it was demoted.
Trigger on the Pages deploy being created rather than completing, then re-check after a delay. deployment_status may or may not fire for legacy Pages builds — unverified, and the last two trigger guesses cost one inert workflow_run and one round of measurement.
No trigger should be swapped in before it is observed firing.#3 already burned workflow_run on reasoning ahead of evidence; the cost of guessing here is a guard that looks like coverage and is not.
Provenance
Found while auditing this session's work, by noticing the guard was silent during an outage it should have caught. The incident is still open at time of filing — the deploy has been queued 29 minutes, one minute short of the threshold it would be measured against if anything were running to measure it.
Finding, from a live incident rather than reasoning
A real stale-site event is happening right now and the freshness guard has not reported it.
The site has served content that does not match
mainfor ~30 minutes, and the guard built to detect exactly that has been silent throughout.The mechanism
page_buildfires when a Pages build completes. A deploy that never completes never fires it. So the event-driven trigger — the one added in #6/#8 specifically because the cron was too sparse — is structurally blind to the failure mode that matters most: a deploy that hangs or is cancelled holds the divergence open indefinitely, and that is precisely the case where no completion event will ever arrive.The
cron: '7,37 * * * *'backstop is the only path that can catch it, and it is measured as shed roughly three in four (observed gaps of 2h02m and 3h49m). It has not fired in this window either.The irony is exact:
check_freshness.pycontains aDeploy stuckbranch, with a 30-minute threshold, written for this scenario — and it cannot run, because nothing triggers it.Why this is not #9
#9 is about latency — the guard runs, but takes ~15 minutes typical / ~40 worst. This is about the guard not running at all for an entire class of failure. A faster trigger does not fix it; no completion-driven trigger can.
Acceptance criteria
Deploy stuckbranch incheck_freshness.py— already written and unit-tested — actually becomes reachable in productionpage_build's unverified status was stated before it was provenOptions, none yet chosen
deployment_statusmay or may not fire for legacy Pages builds — unverified, and the last two trigger guesses cost one inertworkflow_runand one round of measurement.No trigger should be swapped in before it is observed firing. #3 already burned
workflow_runon reasoning ahead of evidence; the cost of guessing here is a guard that looks like coverage and is not.Provenance
Found while auditing this session's work, by noticing the guard was silent during an outage it should have caught. The incident is still open at time of filing — the deploy has been queued 29 minutes, one minute short of the threshold it would be measured against if anything were running to measure it.