Skip to content

[proposal] §5.1 consumes a number without asking whether it was measured — a suite printed 37/37 while 55 checks ran #26

Description

@max-friedman

Area: Round protocol (§5 Verify wider than you changed / §0 Preconditions)

Is this evidence or preference? Evidence — but not mine. See provenance, below.

Provenance, stated first

This did not happen in a project I ran. It was reported by @tonydzi in a comment on #23, as a side finding while independently replicating that issue's carrier result in a different codebase. I have no access to that codebase, its suite, or its runner, so the number below is a report, not a measurement — I cannot reproduce it and I am not claiming to have.

Filing it separately anyway, for two reasons: its blast radius is much wider than the carrier thread it arrived on, where it will read as a footnote; and unlike the counts in that report, its shape is checkable by any project in about five minutes, which is the whole proposal.

The finding

A test suite printed 37/37 while 55 checks had actually executed. The summary line was computed independently of the run rather than derived from the executed set. Every gate above that suite had been reading a number that was simply wrong.

Why this is a loop concern and not one runner's bug

The protocol consumes reported numbers at three separate steps, and none of them asks where the number came from:

  • §5.1 — "Run the project's gate command (tests + lint). Both green. No exceptions."
  • §A step 4 — "Run it. Report the number, especially if bad."
  • §6 — the number goes into the state file, and every later round reads it without re-deriving it. That is the state file's entire purpose.

So a suite that miscounts does not produce one wrong round. It produces a wrong number that is laundered into the memory the loop runs on, and every subsequent round treats it as settled.

The protocol already holds the adjacent lesson. #2 argued a gate that has never failed is not yet known to be a gate, and landed as §0 precondition 5 plus a hard-rules row. Both are about where the gate runs. This is the sharper sibling and it is about whether the gate can count:

A gate whose summary is computed independently of its run is not a gate at all — no matter where it runs, and no matter how often it has gone red.

Note the recursion, which is the part I find hardest to argue away: this also defeats mutation testing, the instrument #23 leans on. A mutation score is a ratio of caught to applied, and both halves are read off the same summary line. Where that line is decorative, "2 of 8 caught" is decorative too. That applies to the report this finding came from — its own 5-of-16 was read off a grid it had just shown can lie — and it would apply to my R102 numbers if this project's runner had the same defect. I have not checked. That I have not checked, after 102 rounds of quoting test counts, is the argument for the precondition.

What it cost

Here, nothing yet — nothing is known to be wrong in this project. Downstream of the reporter, an unknown number of rounds' worth of green.

The honest version: the cost is unquantified, and this proposal is asking for a check on the basis of a plausible mechanism plus one second-hand instance, not on the basis of a measured loss. Rank it accordingly against #19, #21, #22, and #23, all of which can price themselves.

How often

Once, second-hand. Latent for an unknown span before that.

Rounds run: 102 in this project, none of which checked this.

Proposed change

Narrowest version — extend the §0 precondition that already exists for the gate, since it is the same sentence's concern. Current:

  1. Confirm the gate runs somewhere other than this machine. A gate only ever run where it was written is an unverifiable claim, not a check.

Proposed:

  1. Confirm the gate runs somewhere other than this machine, and that it can count: once, check that the number of checks it reports matches the number that actually ran, and that breaking one check moves the number. A gate only ever run where it was written is an unverifiable claim, not a check; a summary line computed independently of the run is not a check at all.

Optionally sharpen the existing hard-rules row rather than adding one, since it is the same idea one level deeper:

A gate that has never failed is not yet known to be a gate — and one whose count is not derived from its run is not a gate even when it fails.

I would rather see the row sharpened than a new row added. The table is read in full every round and this does not deserve its own line.

Blast radius

Near zero for a project on a mainstream runner that reports honestly — a one-time check, a few minutes, never repeated. It costs most where suites are assembled by hand or aggregated across harnesses, which is exactly where the defect lives.

The real cost is not the check, it is the wording: §0 is read in full every round, forever, and this adds a clause to a precondition that only needs satisfying once. That is a recurring read cost for a one-time action, which is the same objection #3 had to raise about placement.

Why this might be wrong

Three arguments, descending strength.

  1. The placement is probably wrong even if the check is right. A one-time check does not belong in a per-round precondition. Two better homes: §B Bootstrap, where the gate command is first identified and the check is free; or §A, whose step 1 already says to prefer "a claim whose supporting test would still pass if the claim became false" — a miscounting suite is the purest instance of that sentence in existence, and §A may already cover this without a single new line. If a maintainer takes only one thing from this issue, I would rather it be a pointer at §A than a clause in §0.

  2. One second-hand instance, from an unverifiable source. The reporter is an agent-authored account with no history in this repository, reporting on a codebase nobody here can see. The mechanism is plausible and the failure mode is real in general, but I would not have filed this from my own project on this evidence, and it should not outrank proposals that measured their own cost.

  3. This may be generic tooling hygiene rather than a loop concern — the same objection #2 had to answer. My counter is the same one, and I think it is stronger here: the loop's distinguishing claim is that agent-run work degrades in ways that look fine, and a green number that was never computed from a run is the exact shape of that claim. But it is a fair reason to decline.


  • I removed repository names, paths, code, and business specifics.
  • This is about the loop itself, not about the project I ran it on.
  • I understand this changes nothing until a maintainer reviews, merges, and releases it.

Activity

  1. added
    proposalSuggested change to the loop, filed via the Loop proposal form
    needs-triageAwaiting maintainer triage
    on Sep 7, 2026
  2. max-friedman commented on Sep 7, 2026

    @max-friedman
    OwnerAuthor

    Checked this project's gate. It does not have the defect — and the check took about two minutes, which is the argument for the precondition.

    The filing admitted I had never verified this after 102 rounds of quoting test counts. Verified now, against scripts/check.py at 405aba9 (current main) and again at the tip of the open round-2 branch, since that is the version about to land:

    • checks_run += 1 happens inside check(), on the same call that prints the ok/FAIL line.
    • failures.append(name) happens inside the same function, in the failing branch.
    • The summary line reads checks_run and len(failures) — both mutated nowhere else.
    • The exit code keys off failures, not off the printed number.

    So the count is derived from the run, and the gate cannot go green on a miscount. First datapoint, and it is a negative one: the mechanism is real in general but absent here. That weakens the case for a §0 clause about this specific project and strengthens the case for whatever form makes the check cheap to run once, anywhere — see the placement argument in the filing.

    The finding this did surface, which narrows the proposal

    The proposal as filed asks for one property. There are actually two, and only the weaker one holds here:

    1. The reported count is derived from the run. ✅ Satisfied. This is what 37/37 vs 55 violates.
    2. The reported count is comparable to an expected total. ❌ Not satisfied, and not satisfiable as the gate is written.

    check_version_bump() prints skip no behavior paths changed vs <base> and returns without ever calling check(). That is correct behavior — the section genuinely does not apply — but it means a section that vanished entirely would produce output indistinguishable from a section that legitimately skipped. The gate has honestly reported 32 checks and 37 checks across recent versions, with no declared total anywhere, so there is no number a reader could compare against to notice a disappearance.

    That is the milder cousin of the reported failure: not a lying summary, but a truthful summary with no denominator. It is worth distinguishing, because the proposed §0 clause as worded ("the number it reports matches the number that actually ran") is satisfied by a gate that silently stopped running a third of its checks.

    Revised proposal

    Property 1 is the one worth a protocol line, and it is a once per project check, not a per-round one — which sharpens the placement objection already in the filing. Property 2 is a stronger and more expensive ask (it means declaring expected counts and maintaining them), and I do not think one second-hand instance justifies it. Recording it here so the distinction survives if this is picked up:

    Once, confirm the gate can count: the number it reports must be produced by the run itself, not asserted alongside it — check that breaking one check moves the number, and that the exit code follows the failures rather than the printed total. A gate that reports a count it did not derive is not a gate. (Separately and more expensively: a gate whose sections can skip silently has an honest count with no denominator. Noted, not proposed.)

    No change to the disposition being asked for. This is one project's answer to the question the issue raises, filed so the issue is not resting entirely on an unverifiable report.


    Generated by Claude Code

  3. added a commit that references this issue on Sep 7, 2026
  4. max-friedman commented on Sep 16, 2026

    @max-friedman
    OwnerAuthor

    Verdict: REJECT

    Criterion 1 (Evidence) fails, and hard disqualifier 7 fires on the same text. The filing states it itself, plainly and to its credit:

    the cost is unquantified, and this proposal is asking for a check on the basis of a plausible mechanism plus one second-hand instance, not on the basis of a measured loss.

    and:

    What it cost — Here, nothing yet — nothing is known to be wrong in this project.

    Disqualifier 7 is "a proposal with no round it actually cost something in is a preference." There is no such round. The instance is real but second-hand, from a codebase nobody here can inspect, reported by an account with no history in this repository — which the filing also says, in its own objection 2.

    The follow-up comment makes this decisive rather than merely thin. The check was run against this project's gate, and it came back negative:

    So the count is derived from the run, and the gate cannot go green on a miscount. First datapoint, and it is a negative one: the mechanism is real in general but absent here.

    That is the right work and the right way to report it. But a proposal whose only first-hand evidence is a negative result is asking every project running the loop to pay a recurring read cost for a defect that has been looked for once and not found.

    Criterion 3 (Necessity) fails independently. The filing's own objection 1 identifies it:

    §A, whose step 1 already says to prefer "a claim whose supporting test would still pass if the claim became false" — a miscounting suite is the purest instance of that sentence in existence, and §A may already cover this without a single new line.

    That is correct, and it is the rubric's "does not restate something the protocol already says." A gate that reports a number it did not derive is a claim whose supporting test would pass if the claim became false. §A already points an audit at exactly that, and §A's verdict vocabulary already includes unmeasurable as stated, which is what such a gate is.

    Criteria that pass. Criterion 6 (Self-criticism) is the strongest section in the filing — three objections, correctly ordered by strength, and the first two are the two this verdict rests on. Criterion 5 (Blast radius) is properly stated, including the recurring-read-cost-for-a-one-time-action objection against its own placement. The provenance disclosure at the top is exactly how a second-hand report should be filed.

    What would change the answer. Either of:

    1. A project that ran the five-minute check and found its gate does miscount — naming the rounds whose numbers were affected and what was concluded from them. That converts a mechanism into a cost. One first-hand positive instance would carry this much further than the argument does.
    2. The narrower finding the follow-up comment actually surfaced, which is not this proposal and is more interesting than it: a gate whose sections can skip silently has an honest count with no denominator, so "a section that vanished entirely would produce output indistinguishable from a section that legitimately skipped." That is a distinct defect, it was found first-hand, and the filing's proposed wording does not catch it — a gate satisfying "the number it reports matches the number that ran" can still have silently stopped running a third of its checks. File that on its own once it has a round behind it.

    Rejections here are cheap and recoverable. This one is "not yet," not "wrong."


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    proposalSuggested change to the loop, filed via the Loop proposal formrejected

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions