Skip to content

Commit 0697586

Browse files
Merge pull request #143 from UnityInFlow/stop26/b11-efficiency
Stop 26 (B11 — efficiency, v1.2): NOT DETECTABLE on both tasks, 847 hook decisions with zero refusals, and three instrument defects fixed
2 parents 70221a7 + 32a7c96 commit 0697586

766 files changed

Lines changed: 43447 additions & 31 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/ci.yml‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -72,6 +72,17 @@ jobs:
7272
run: ./tools/verify-repair-limit.sh
7373
- name: The run state checker still refuses
7474
run: ./tools/verify-run-state-checker.sh
75+
# B11 (v1.2) adds three EXECUTING controls, and the batch cannot test any of them: on a task
76+
# the model passes nearly always, a hook can log ten allows and refuse nothing, and a control
77+
# never shown to reject anything is indistinguishable from one that rejects nothing. These
78+
# three sets are where the refusals are proved. The dedup set builds a real git repository
79+
# because its whole claim is about a code fingerprint; a stubbed git would test the stub.
80+
- name: The retrieval budget still refuses
81+
run: ./tools/verify-retrieval-budget.sh
82+
- name: The file-summary cache still refuses a stale entry
83+
run: ./tools/verify-summary-cache.sh
84+
- name: Command deduplication still refuses
85+
run: ./tools/verify-command-dedup.sh
7586
# A script nothing ShellChecks looks exactly like one that lints clean: both are a
7687
# green run with no output. This compares every tracked *.sh against the directories
7788
# the two steps above scan — pinned in tools/shellcheck-scanned.txt, never parsed out

‎HANDOFF.md‎

Lines changed: 64 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -3,8 +3,12 @@
33
Read `CLAUDE.md` first; it carries the operational facts and is loaded automatically. This
44
file is the *state*: what is in flight, what is blocked, and on whom.
55

6-
**Start at the Position section immediately below. Positions 4–25 are CLOSED — stop 25 (Phase 10)
7-
closed 2026-09-29 — stop 26 (B11) is NOT OPENED, and NOTHING is blocked on the author.**
6+
**Start at the Position section immediately below. Positions 4–26 are CLOSED — stop 26 (B11 —
7+
efficiency) closed 2026-10-05 `NOT DETECTABLE` on both tasks, v1.2 kept and NOT promoted — stop 27
8+
(B12) is NOT OPENED, and NOTHING is blocked on the author.**
9+
10+
*(Superseded pointer, kept:)* "Positions 4–25 are CLOSED — stop 25 (Phase 10) closed 2026-09-29 —
11+
stop 26 (B11) is NOT OPENED, and NOTHING is blocked on the author." **Stop 26 has closed since.**
812

913
*(Superseded pointer, kept — and it had gone stale by FIVE stops, which is the worst this line has
1014
managed:)* "Start at \"The author's decision of 2026-09-26\" immediately below the Position section.
@@ -51,6 +55,64 @@ against 0 of 5, and one sentence of borrowed authority moved it not at all.**
5155

5256
## Position
5357

58+
**Spine 26 of 28. Positions 4–26 CLOSED — 26 (B11 — efficiency, v1.2) closed 2026-10-05
59+
`NOT DETECTABLE` on both tasks at `n = 10` per arm.** `v1.2` is **kept and NOT promoted**, which is
60+
what was registered before the batch: four of the gate's seven clauses have no instrument, so the
61+
version could not be promoted even on an `IMPROVED` row, and it did not get one. **NOTHING is
62+
blocked on the author.** `lab#36` (B11) is **closed** — a B step's issue closes when its deliverable
63+
is decided, and a kept-not-promoted version is a decision. `lab#12`, `lab#11`, `lab#10`, `lab#9`,
64+
`lab#16` and `lab#8` stay open as Phase issues whose gates are not met from measurement. **Stop 27
65+
(B12 — governed self-learning) is NOT OPENED and nothing of it exists**; §6 forbids a future step's
66+
artifacts early, and opening it at §4 step 1 is the next session's first act.
67+
68+
**The headline, both tasks, and neither is detectable.** Context total 346 697 → 321 179 on BE-003
69+
(**−7.36 %**, exact permutation `p = 0.417` over all 184 756 relabellings) and 515 872 → 527 854 on
70+
BE-004 (**+2.32 %**, `p = 0.851`). Acceptance, hidden tests, build and static analysis **identical in
71+
all four arms** — `passed: true` on 10 of 10 everywhere. Decision-rule row 4 fires on both.
72+
73+
**The finding that outlives the verdict: the three executing mechanisms are correct and inert, and
74+
both halves of that are measured.** Across **847 hook decisions** on the registered 20 treated runs
75+
— budget 324, cache 403, dedup 120 — there are **ZERO refusals**. The agent under test never re-read
76+
a file at unchanged content, never repeated a command under unchanged code, and issued **no `Grep`
77+
and no `Glob` at all** (the wiring was checked: the hook is on `PreToolUse Read|Grep|Glob`, and it
78+
logged 207 `Read` lines beside 0 searches). It read **7–13 distinct files against a limit of 15**.
79+
`H₂ = H₃ = H₅ = 20 of 20` says those hooks **ran**; it does not say they **did** anything, and this
80+
is the number that separates the two. **So "was this the agent, or the harness?" answers neither:
81+
the waste the version removes is not present in this agent on these two tasks.**
82+
83+
**§4 step 10, per mechanism, never pooled: three removals and two withdrawn claims.** The task
84+
classifier goes (`H₁ = 0 of 20` — *delivery is not invocation*, now measured twice on this
85+
instrument, with B9's `H = 2 of 10`). The verification planner goes on a **different** ground —
86+
unmeasurable **by registration**, not a measured null — with a condition of re-entry: it returns only
87+
if it writes the sequence it chose. The three L2 hooks **stay**, their efficiency claims are
88+
**withdrawn**, and the cache splits: the *"never trust a stale summary"* branch fired 4 times in paid
89+
runs and §4 step 9 showed that deleting it produces a **correctness** failure, so it is load-bearing;
90+
the *"reuse on hash match"* branch fired **0 times in 403 decisions**.
91+
92+
**Three instrument defects were found at this stop and all three are fixed in its PR.** Two were
93+
recorded during the batch and deliberately not fixed then (§4 step 4 forbids editing a tool mid-run):
94+
a manifest header printing a **hardcoded prediction commit that belonged to stop 20**, and `TREATED_N`
95+
/ `ROW0A` re-zeroing on resume so a resumed batch printed *"over 9 treated run(s)"* above mechanism
96+
counts of 20. The third came from the §4a review: the workbook promised the batch-guard fixture set
97+
*"is re-run before the batch"* and **nothing executed to keep that promise** — the 40-run batch ran
98+
with that set last recorded at 16/17. It is **17 of 17** now, and the driver's gate enforces it.
99+
**That fix then recursed** — the guard set invokes the driver seventeen times — which is itself
100+
recorded, with the reentrancy guard and the two stub cases that prove both directions without any
101+
fixture being able to start a batch.
102+
103+
**And the review found the house failure mode inside the step whose subject is the house failure
104+
mode.** §4 step 9's D4 check — *"the refusal leaks no file body"* — passed **vacuously** when its
105+
subject had no source line to look for. Fixed to refuse at exit 4, with the fixture case that proves
106+
it, and the probe re-run at **45 of 45** under the stricter checks. The first run's record is kept.
107+
108+
**Two registration defects of this stop are named rather than smoothed over, and they share a
109+
shape:** the MDE table said *"re-derived from this batch's control"* without saying **how** (three
110+
readings were therefore reported, and they agreed — recorded as luck), and E-027's decision-rule row
111+
1 and prediction P5, **registered in the same commit**, can both fire on the `change-focus` median.
112+
Both leave a choice to be made after the numbers are seen. The forward rule is in E-027 Amendment 1.
113+
114+
**Position — superseded 2026-10-05 at the stop-26 close**
115+
54116
**Spine 25 of 28. Positions 4–25 CLOSED — 25 (Phase 10 — production observability) closed
55117
2026-09-29 with `n = 0` benchmark runs commissioned.** The stop read the phase, wrote a
56118
second-pass extract, and **ran Lab 10.0** — all three of its checkboxes answered, two of them

0 commit comments

Comments
 (0)