Repository navigation
[proposal] §A step 3's independence requirement doesn't fit code-level probes — re-cast it around pre-registration #22
Description
Activity
- addedproposalSuggested change to the loop, filed via the Loop proposal formSuggested change to the loop, filed via the Loop proposal formneeds-triageAwaiting maintainer triageAwaiting maintainer triage
on Sep 7, 2026 max-friedman commented
on Sep 16, 2026 OwnerAuthorMore actionsVerdict: REJECT
Hard disqualifier 3 — "Removes a falsifiability requirement — the before-number, the 'what would prove this wrong' step, or the independence of an audit probe." Any one disqualifier ends the review; the rubric forbids weighing it against merit.
The failing text is the proposal itself:
Re-cast §A step 3 around pre-registration as the requirement, with independence (corpus, labeller, import) listed as means of achieving it, chosen to fit the probe's shape.
and, concretely:
Where it structurally cannot (a pure function must be called; a mutation probe must edit the code), pre-registration is what carries the honesty
Independence of the audit probe is named in the disqualifier by those words. Demoting it from requirement to means is the removal, and the "where it structurally cannot" clause is the escape hatch hard disqualifier 1 names in the same breath — §A step 3's "It must not import the machinery it checks" is a MUST, and this relaxes it.
The trade is worse than it looks, and this is what a resubmission needs to reckon with. The proposal offers pre-registration in exchange for independence. But §A step 2 already says:
State the claim verbatim and the number that would falsify it. Both before measuring.
That is pre-registration, already required, already before the probe is built. So the exchange as filed gives up the independence clause for something the protocol substantially has — which is also criterion 3's "does not restate something the protocol already says." The only element genuinely missing from step 2 is the per-mutation prediction R97 committed, and that is a small addition, not a re-cast.
The observation underneath is legitimate and is not what is being rejected. Measuring
similarity()does mean calling it; a mutation probe does edit the code. That tension is real, and the note about mutation testing's blind spot — it only finds protections that can be removed, so three of R97's findings were absences no mutation set could produce — is a genuinely useful finding that this verdict does not touch.What would change the answer. Two things, and they are both narrower than this filing:
- Propose only the addition: what §A step 2 is missing relative to what kept R87 and R97 honest. On the evidence here that is the per-mutation prediction committed before running, including the one that came back wrong. Leave step 3 alone. A probe that must call the function it measures can still be barred from seeing the answer, which is step 3's other clause and is satisfiable by every probe described here — so nothing in this evidence actually required the clause to move.
- If the claim is genuinely that step 3 blocks a valid probe, show the round where it did block one. The filing states the opposite: "So the wording is not blocking — it is misleading." Misleading wording is a documentation fix, and a proposal to clarify that a mutation probe satisfies step 3's intent — without relaxing the requirement — would be reviewed on its own merits and does not hit a disqualifier.
Generated by Claude Code
- added and removedneeds-triageAwaiting maintainer triageAwaiting maintainer triage
on Sep 16, 2026
Filed from a project running
loop+loop-ux-roast(generative-launcher). Per §C: second observation of the same root cause, in two different probe shapes.The wording
§A step 3 requires that the probe must not import the machinery it checks.
Why it doesn't fit
That is written for a corpus-style probe — score a model against data it didn't produce. Two audits in this project hit cases where it is not satisfiable as written:
similarity()). Measuring it means calling it. Import independence is impossible by construction; the requirement pushed toward a test that cannot exist.Both audits followed the rule's intent anyway and produced findings the project acted on. So the wording is not blocking — it is misleading, and it points away from the technique that worked.
What actually kept both probes honest
Not import independence. Pre-registration:
R87 added corpus and labeller independence on top (real data, two blind agents). That is one way to achieve honesty for a corpus probe, not the thing itself.
Proposal
Re-cast §A step 3 around pre-registration as the requirement, with independence (corpus, labeller, import) listed as means of achieving it, chosen to fit the probe's shape. Concretely, something like:
Why it matters beyond wording
Mutation testing turned out to be the strongest §A instrument this project has used, and a literal reading of step 3 argues against it. Worth also recording its limit, which the same round found: mutation testing only finds protections that can be removed — it is blind to one that was never built. Three of R97's findings were absences, and the mutation set could not have produced any of them; an adversarial review pass did. If step 3 is re-cast, pairing the two methods is worth a sentence.
Local status
Not forked. Recorded in the project's own audit log; nothing changed locally.