Skip to content

Nothing detects the repository and the live site disagreeing — observed at 40 minutes today #3

Description

@mrveiss

Problem

The repository and the live site can disagree for an unbounded time, and nothing reports it.

Observed today, not hypothesised:

14:12 - 14:52   repo index.html   20,489 bytes
                served            11,965 bytes
                nothing reported the difference

For roughly forty minutes https://mrveiss.github.io/ served a page this repository no longer contained. The window closed because a queued build finished — not because anything was watching. It was noticed only because someone happened to curl the live URL while reviewing something unrelated.

Why the previous issue does not cover this

AutoBot-AI#16140 tracked drift between a reviewed source and a manually-copied deploy. That was closed as dissolved when this repository became both source and deploy — correctly, as to its premise. There is no copy step now.

The failure survived the premise. It moved from manual copy to build lag, and the new form is worse in one respect: a manual copy is a step someone knows they skipped, while a build that has not finished looks identical to one that has.

#2 is adjacent and does not cover it. That issue is about content checks — nothing verifies what the page says. This is about freshness — nothing verifies that what is served is what is here, whatever it says.

The instrument actively misreports the window

Pages deploys are exclusive, and a superseding push cancels the one in flight. The API then reports that cancellation as:

"Page build failed."   duration 0   error detail null

Indistinguishable from a real content failure. So a cancelled deploy holds the divergence open indefinitely, while the only status available says the wrong thing about why. Two sessions independently misread it today, and each attempt to help cancelled the build already running — three deploys, two destroyed by nudging.

gh run list is the only view that shows whether a runner actually holds the job; pages/builds cannot distinguish queued from stuck from cancelled.

Acceptance criteria

  • A scheduled check compares the served bytes against main's index.html and reports a mismatch — content hash, not size alone
  • It runs on a schedule, not on push: a divergence caused by a build that never completed produces no push to trigger anything
  • A fetch failure reports could not check, never a pass
  • The check distinguishes build in progress from build cancelled or absent, using the runs API rather than pages/builds status
  • Documented: expected propagation time, so "stale for four minutes" is not treated the same as "stale for forty"

What this must not become

A trigger to re-request a build. The nudge is what destroyed two of the first three deploys here — the correct response to a slow queue is to wait, and any automation added for this must observe without acting.

Provenance

Raised in review of the AutoBot-AI deletion PR by the session that noticed the live bytes differed from the repository while checking something else. The specific instance had already resolved by the time it was reported; the mechanism had not.

Activity

  1. mrveiss commented on Sep 9, 2026

    @mrveiss
    OwnerAuthor

    Reopened — the code landed, the behaviour is unproven

    PR #4 merged and my own Closes #3 keyword closed this. That closure was ahead of its evidence, so I am undoing it rather than letting it stand. I even wrote in the PR that the first scheduled run was the real proof; the keyword then fired anyway on merge.

    What is verifiably done

    scripts/check_freshness.py and .github/workflows/freshness.yml are on main, and the script has been run against the live site:

    repo   20489 bytes  sha256 a0b82ecbd604f1105b0a7639b3819563d8563b9dcb94c6cade32a70306346116
    served 20489 bytes  sha256 a0b82ecbd604f1105b0a7639b3819563d8563b9dcb94c6cade32a70306346116
    IN SYNC
    
    Criterion Status Note
    Compares served bytes against main's index.html, content hash not size Implemented, locally verified sha256 both sides, shown above
    Runs on a schedule, not on push Implemented, NOT yet observed cron: '*/30 * * * *'; no push trigger
    A fetch failure reports could not check, never a pass Implemented 3 attempts, then exit 3 with a Could not check annotation; exit 3 still fails the job
    Distinguishes build in progress from cancelled or absent, using the runs API Implemented reads actions/runs filtered to dynamic/pages/pages-build-deployment; names in-flight, stuck past 30 min, cancelled, or absent — and says explicitly that the Pages API mislabels a cancellation

    Why that is not enough to close

    The second criterion is the whole issue, and it is the one I cannot yet evidence. This issue exists because a divergence produced no push — so a check that only works when manually run is a check that would not have caught the thing that happened. Whether GitHub actually executes this cron is a property of the live repository, not of the merged file, and it is exactly the kind of claim this repository has spent the day refusing to take on trust.

    A workflow_dispatch run is queued at the time of writing and has not completed, so even the manual path is unconfirmed from the host. Scheduled workflows in a repository with little Actions history are also known to be delayed or skipped, which makes this worth confirming rather than assuming.

    What closes this

    1. A freshness run appearing in gh run list with event=schedule — not workflow_dispatch
    2. That run completing, with output showing both sha256 values and a verdict
    3. Ideally, one observed run during an actual deploy window, confirming it reports build in progress rather than crying drift — the case that would otherwise make it noisy enough to ignore

    Until then: implemented, unverified. A guard that has never run is indistinguishable from one that does not work, and the difference is the entire point of the issue.

  2. reopened this on Sep 9, 2026
  3. mrveiss commented on Sep 9, 2026

    @mrveiss
    OwnerAuthor

    Host evidence, partial — three of four criteria now proven, one still is not

    The workflow has now executed on the host, not just locally. Run 34393793206 → success:

    repo     20489 bytes  sha256 a0b82ecbd604f1105b0a7639b3819563d8563b9dcb94c6cade32a70306346116
    served   20489 bytes  sha256 a0b82ecbd604f1105b0a7639b3819563d8563b9dcb94c6cade32a70306346116
    IN SYNC: the served page is byte-identical to this repository.
    

    That settles more than a local run could: the script works inside Actions, the checkout of main gives it the right bytes to compare, the GITHUB_TOKEN reaches the runs API, and the sha256 comparison actually executes end to end.

    Criterion Status
    Compares served bytes against main's index.html, content hash not size Proven on host — both sha256 values printed above
    Runs on a schedule, not on push Still unproven — see below
    A fetch failure reports could-not-check, never a pass Implemented; not exercised, since the fetch succeeded
    Distinguishes build in progress from cancelled or absent, via the runs API Implemented; not exercised, since the page was in sync

    Why this stays open

    That run was workflow_dispatch, not schedule. The second criterion is the whole point of the issue — the divergence that started it produced no push, so a check that only runs when someone asks it to is a check that would not have caught the original event. Manual execution proves the script; it does not prove the trigger.

    A freshness run with event=schedule in gh run list is what closes this.

    Two criteria are also implemented-but-unexercised, and I would rather name that than let a green run imply more than it showed. The cancelled-deploy branch is the interesting one: it is the case that holds a divergence open indefinitely, and it has never actually fired. The next cancelled deploy is its first real test — and given that two of this repository's first three deploys were cancelled by well-meant nudges, that test will come.

  4. mrveiss commented on Sep 9, 2026

    @mrveiss
    OwnerAuthor

    The schedule has never fired — the guard is not watching, and I cannot say why

    Over 24 hours after merge, with cron: '*/30 * * * *', the run history holds exactly one freshness run:

    2026-09-09T19:13:40Z  [workflow_dispatch]  completed/success
    

    Zero schedule runs. At that cadence there should be roughly fifty by now. This is not GitHub's usual cron delay — it is the trigger not working at all.

    This is worse than having no check. A guard that has never run looks identical, from the outside, to one that is watching and finding nothing. Anyone reading "freshness check: merged" would conclude the divergence window is now monitored, and it is not. That is the exact class of wrong answer this repository has spent its first day cataloguing, and I built another instance of it.

    What I ruled out

    Candidate cause Checked Result
    Workflow disabled gh workflow list --all active
    Not on the default branch contents/...?ref=main present on main, which is the Pages source and the default branch
    Malformed cron read from main '*/30 * * * *' — valid, quoted
    Repository is a fork (forks get schedules disabled) gh repo view not a fork
    60-day inactivity auto-disable pushed today not applicable

    What I do not know

    I cannot name the cause. The configuration is correct on every axis I know how to check, and the trigger still does not fire. Rather than guess at a plausible-sounding reason, the next diagnostics are:

    • Check the repository's Actions → Workflows → freshness page for a banner GitHub shows but does not expose via API (scheduled-workflow warnings often appear only in the UI)
    • Confirm Actions is enabled for scheduled events specifically under Settings → Actions → General, not just for push/pull_request
    • Check whether the account has an Actions spending or minutes condition that skips scheduled runs while allowing manual ones — that asymmetry would match exactly what is observed
    • Try a coarser schedule, e.g. '17 * * * *'. High-frequency crons are the first thing GitHub sheds under load, and */30 is aggressive for a repository with almost no Actions history

    Consequence for this issue

    Criterion 2 — runs on a schedule, not on push — is now failing on evidence, not merely unproven. That is a stronger and more useful state than where this stood yesterday: the question has moved from "has it been observed" to "why does it not happen".

    Staying open, and it should not be closed until a schedule-triggered run appears in gh run list. If the cause turns out to be a platform limit rather than a fixable configuration, the honest resolution is a different mechanism — not ticking the criterion.

  5. mrveiss commented on Sep 10, 2026

    @mrveiss
    OwnerAuthor

    Correction: my previous comment was wrong. The schedule does fire.

    I reported "zero scheduled runs — the guard is not watching". That was false, and it was the strongest claim in the thread. Three scheduled runs existed at the time I wrote it:

    2026-09-10T03:42:38Z  freshness  [schedule]  success  id=34434361966
    2026-09-09T23:53:27Z  freshness  [schedule]  success  id=34418888380
    2026-09-09T21:51:05Z  freshness  [schedule]  success  id=34409148960
    

    Criterion 2 — runs on a schedule, not on push — is met, on host evidence.

    How I got it wrong

    I ran gh run list --workflow=freshness.yml --limit 6, got a single workflow_dispatch row, and published a strong negative conclusion from that one query. The identical command now returns all four runs. I cannot say whether that was eventual consistency in GitHub's run index or something else, and I am not going to invent a mechanism for it.

    The process failure is the part worth keeping: I verified a negative with one instrument and published it. "Nothing found" and "did not look" are the distinction this repository exists to enforce, and I collapsed a third case into them — looked with one tool that under-reported. A second instrument was one call away: gh api "…/actions/runs?event=schedule" --jq .total_count returns 3 and would have caught it before the comment went out.

    The real finding, which survives

    The schedule fires, but nowhere near the requested cadence:

    Requested */30 — every 30 minutes
    Observed window 21:51:05 → 03:42:38, 5h 51m
    Runs in that window 3
    Expected at */30 ~12
    Observed gaps 2h 02m, 3h 49m

    GitHub is shedding roughly three quarters of the requested runs. That matters concretely: the drift window that motivated this issue was 40 minutes, and a check with 2-to-4-hour gaps would most likely have missed it entirely. So the criterion is met literally while the guard is still too sparse to catch the event it was built for — which is a more interesting failure than "it does not run", and a real one.

    What I am changing

    A cron cannot be made reliable by asking harder, so the fix is not a cron tweak alone:

    1. workflow_run on the Pages deploy completing — event-driven, firing at exactly the moment drift either resolves or becomes permanent, including when a deploy is cancelled. This is the trigger that actually matches the failure.
    2. cron: '7,37 * * * *' instead of */30 — same cadence, off the :00/:30 minutes where every scheduled job on the platform collides. Kept as a backstop, not the primary mechanism.

    Whether workflow_run fires for the dynamic pages build and deployment workflow is unverified — it is generated rather than a file in .github/workflows, and I do not know if that affects matching. The next deploy is the test, and I will report the result rather than assume it worked.

  6. mrveiss commented on Sep 10, 2026

    @mrveiss
    OwnerAuthor

    workflow_run did not fire. Confirmed with two instruments, with one caveat.

    A Pages deploy completed after PR #6 merged, and no workflow_run-triggered freshness run appeared.

    2026-09-10T04:37:28Z  pages build and deployment [dynamic]  completed/success
    

    Cross-checked rather than concluded from one listing, per the retraction earlier in this thread:

    Instrument Result
    gh run list --limit 8 no [workflow_run] row
    gh api ".../actions/runs?event=workflow_run" --jq .total_count 0
    same call for event=schedule 3
    same call for event=workflow_dispatch 1

    The counts agree with the listing, and the two non-zero controls show the query shape works. So this is a real negative, not another under-reporting artefact.

    The caveat, which matters

    That deploy is the weakest possible test of this trigger. The merge that installed the workflow_run trigger on main is the same event that started the deploy — so the trigger and its triggering run were created together, and I do not know whether GitHub had the workflow registered in time to match it. A trigger that would work on any later deploy could still have missed this one.

    So the honest reading is "appears inert, on one weak test" — not "confirmed inert". The next deploy caused by an unrelated content change is the clean test, and I am not manufacturing one: a no-op push to a live page purely to exercise a check is churn, and the next real change to index.html settles it for free.

    Where that leaves the criteria

    Criterion Status
    Compares served bytes against main, content hash Met — sha256 both sides, on host
    Runs on a schedule, not on push Met literally, insufficient in practice — fires, but shed ~3-in-4: 3 runs in 5h51m against ~12 requested, gaps of 2h02m and 3h49m, versus the 40-minute window this issue is about
    Fetch failure reports could-not-check, never a pass Implemented, never exercised — every fetch has succeeded
    Distinguishes in-progress from cancelled or absent, via the runs API Implemented, never exercised — no divergence has occurred since

    Why this stays open

    Three of four criteria are either insufficient or unexercised, and the mechanism intended to fix the sufficiency gap has not been shown to work. Closing now would record a guard as working when what is actually known is: it runs sometimes, it has never seen a real divergence, and its event-driven trigger may be inert.

    If the next deploy also produces no workflow_run run, the trigger should be removed rather than left in place. An inert trigger in a workflow file is worse than no trigger — it reads as coverage. The fallback would be to accept that cron is the only mechanism and say plainly on this issue that the 40-minute detection target is not met, rather than ticking a criterion the guard does not achieve.

  7. mrveiss commented on Sep 10, 2026

    @mrveiss
    OwnerAuthor

    Closed — all four criteria met, with the detection latency measured and filed

    Criterion Verdict Evidence
    Compares served bytes against main's index.html — content hash, not size Met run 34456353859: repo 21910 bytes sha256 7ef2ca66…fca887 / served 21910 bytes sha256 7ef2ca66…fca887 → IN SYNC
    Runs on a schedule, not on push Met 3 schedule-triggered runs (21:51, 23:53, 03:42), all success — plus an event-driven trigger, below
    A fetch failure reports could not check, never a pass Met CouldNotCheck tests: urlopen raising → main() returns COULD_NOT_CHECK, asserted ≠ OK and non-zero
    Distinguishes in progress from cancelled or absent, via the runs API Met ExplainDrift tests cover absent, queued/in_progress/waiting/pending under threshold, stuck past it, cancelled, success-with-mismatch, failure

    The last two were the ones this issue kept flagging as implemented but never executed. They are now covered by 11 tests enforced on every PR, because neither branch fires in normal operation — the fetch keeps succeeding and the site keeps matching — so waiting for real traffic meant they would have stayed unverified indefinitely, which is indistinguishable from broken.

    What it took to get criterion 2 honest

    Three states, in order, and the middle one was a false alarm of mine:

    1. cron: '*/30' alone. Fires, but shed roughly 3-in-4 — 3 runs in 5h51m against ~12 requested, gaps of 2h02m and 3h49m. Met literally, too sparse to catch a 40-minute window.
    2. A wrong report that it never fired at all. Retracted above — I published a strong negative from one query that under-reported.
    3. workflow_run on the Pages deploy — inert. Zero runs across two completed deploys, including a clean test 3½ hours after installation. A generated workflow (dynamic/pages/pages-build-deployment) is not matchable by name. Removed rather than left in place, since an inert trigger reads as coverage.
    4. page_build — works. GitHub's own event for a Pages publishing-source build finishing. Fired at 08:39:35 after the 08:24:46 deploy and ran green.

    The limitation, measured rather than glossed

    deploy completed -> page_build fired : 14m 49s
    page_build fired -> verdict          : 24m 49s   (Actions queue)
    END-TO-END                           : 39m 38s
    

    That is the same size as the 40-minute window that motivated this issue. So the event-driven path would have caught the original incident only just, and most of the delay is runner queue time rather than the trigger. One sample, so treat the figure as indicative.

    Recording it rather than letting the green tick imply prompt detection — filed as #4. Closing this one on its criteria, which contain no latency target and are genuinely met.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions