Skip to content

fix(pm): dispatch-gates --ran never reads a killed run as RUN — a claim with a reason wins, a bare kill is UNRUN, a verdict-plus-claim contradiction prints (#18074) - #18108

Merged
os-project-manager merged 4 commits into
mainfrom
claude/issue-18074-ran-record-cap-kill-is-not-run
Sep 14, 2026
Merged

os-project-manager merged 4 commits into
mainfrom
claude/issue-18074-ran-record-cap-kill-is-not-run

Conversation

@claude

@claude claude Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #18074

scripts/pm/dispatch-gates.mjs --ran had two channels for declaring a gate family
unmeasured, and one silently ate the other. A recorded exit 124 — the number a
timeout wrapper reports when it had to signal the child — classified as run,
and so did every code at or above the signal floor. A record that also carried an
explicit NOT-MEASURED ... :: reason for the same command had that declaration
discarded without a word. The most careful record produced the least truthful total.

A kill is not a verdict. The process was ended before it produced a result, so
counting the family as run asserts a measurement that does not exist — the false
green the whole reconciliation exists to refuse.

What changed — scripts/pm/dispatch-gates.mjs only

One named constant, judged in one place. RUN_RECORD_KILL_EXITS holds the
timeout wrapper's 124 and the 128 signal floor, with a docblock saying why each
member is in it; runRecordKillLabel() is the only expression that compares a code
against it, so no reconciliation branch and no rendered line carries a bare number of
its own. 124 is the member the set exists for: it sits below the floor, so a
floor test alone never catches it.

A killed family lands in one of two places, never run:

  • NOT-MEASURED when the record also carries a claim with a stated reason for that
    command. The claim wins over the run line beside it — the runner who recorded
    the code and declared the family has said everything there is to say, and the two
    lines do not actually disagree. It gets its own source, its own block
    (NOT-MEASURED · KILLED) and a row naming both the code and the reason.
  • UNRUN otherwise — the direction that costs a rerun. The row names the recorded
    code and prescribes the spelling that would declare it.

It is counted once. A killed-and-declared family arrives through two doors and is
one family. It is already inside coded, so killClaimed is reported beside the
other evidence counts and added to none of them; accounted stays equal to the number
of families the record accounts for.

A verdict plus a claim is still a genuine contradiction — 0, 1, 2 and every
other code below the floor — run still wins, unchanged, and the contradiction line
still prints. The conflicts rendering now reads the class the reconciliation actually
assigned instead of re-deciding it, so it can no longer report a reading the totals
above it do not share.

Premise readings, taken before writing

P1 — reading D reproduces. 2026-09-14T02:11Z, worktree at a26a114d7, the card's
own reproduction (--repo objectstack-ai/objectstack packages/spec/src/system/metrics.zod.ts,
73 derived families on this tree), exits captured by redirect-then-capture:

Run reconciliation — 73 derived, 1 run, 0 NOT-MEASURED, 72 UNRUN.

The seat's open question, measured on the full output rather than the card's
reconciliation-only filter: the contradiction line does print. It is line 77 of
stdout, and it reads

  ⚠️ 'node packages/lint/scripts/check-reference-carrier-shape.mjs' is recorded BOTH as run and as NOT-MEASURED. Read as run; fix the record so it states one thing.

So conflicts was both computed and rendered; the defect was never that the tool
failed to notice, it was which bucket the family landed in. Since the line already
printed, it is pinned rather than repaired — by a self-test that reads the rendered
text
, not the conflicts array.

P2 — a bare exit 124 classified as run. Same run, reading A:
73 derived, 1 run, 0 NOT-MEASURED, 72 UNRUN — identical to D. Reading B (exit 3)
and reading C (the claim alone) both read 0 run, 1 NOT-MEASURED on the same tree,
which is what makes A and D readings of one instrument rather than four instruments.

P3 — the module argues the opposite way round. Its own parseRunRecord docblock:

a marked line without one leaves its family UNRUN, which is the direction that costs
a rerun instead of a false green.

P4 — hold #14290, declared rider, nothing folded. 2026-09-14T02:20Z: #14290 is
open, domain:devx, priority:p3; its hold comment 5556473243 carries
Restart-touch: scripts/pm/dispatch-gates.mjs and
Restart-when: closed objectstack-ai/objectstack#16132, which has fired. Its subject
is the STAGE-THEN-RUN edge on the derivation side — check:objectui-changeset
inheriting no watch hints from scripts/bump-objectui.sh, an edge neither follow
traverses. This PR touches only the --ran record reconciliation and its
rendering: it adds no follow, reads no program text and changes no hint extraction, so
it does not meet that edge and stops short of it. The hold's own re-pricing stays in
the domain:devx lane. Riders #12797 and #12808 are closed.

Before / after, same command, same derivation

record before after
A CMD :: exit 124 73 derived, 1 run, 0 NOT-MEASURED, 72 UNRUN 73 derived, 0 run, 0 NOT-MEASURED, 73 UNRUN
B CMD :: exit 3 0 run, 1 NOT-MEASURED, 72 UNRUN unchanged
C NOT-MEASURED CMD :: reason 0 run, 1 NOT-MEASURED, 72 UNRUN unchanged
D both A and C 1 run, 0 NOT-MEASURED, 72 UNRUN 0 run, 1 NOT-MEASURED, 72 UNRUN
E CMD :: exit 0 + a claim 1 run, 0 NOT-MEASURED, 72 UNRUN unchanged

Reading D's family row and its conflict line, after:

  NOT-MEASURED · KILLED (1) — your record carries BOTH a kill code and a stated reason for these, so the run line beside it is NOT read as run: a kill is not a verdict. ...
    - node packages/lint/scripts/check-reference-carrier-shape.mjs   [recorded exit 124 on line 1 — a `timeout` wrapper fired and signalled the child, so no verdict was reached; declared on line 2: cap-killed at the foreground ceiling]
  ⚠️ 'node packages/lint/scripts/check-reference-carrier-shape.mjs' is recorded BOTH as run and as NOT-MEASURED. The two AGREE: exit 124 is a kill, not a verdict. Read as NOT-MEASURED with your stated reason — ⭐ this is the record shape to keep, not one to repair.

Reading A's, after:

  ⛔ UNRUN (73) — derived for these paths, and the record does not account for them:
    - node packages/lint/scripts/check-reference-carrier-shape.mjs   [recorded exit 124 on line 1 — a `timeout` wrapper fired and signalled the child, so no verdict was reached; declare it as `NOT-MEASURED CMD :: reason` beside the code to count it NOT-MEASURED]

Tests

node scripts/pm/dispatch-gates.mjs --self-test — before: 1682 cases pass,
after: see the verification section below. Every existing case is kept. The one
re-spelling is the both-ways case: its expectation is byte-for-byte what it always
asserted, and the sentence now says which run lines it covers, because a bare run
line carries no code for a kill to be read out of — it is the control beside the new
branch, not a casualty of it.

New pins: reading A (bare kill → UNRUN, and the row names the code), reading D (kill +
reasoned claim → NOT-MEASURED, and the row carries both halves), the double-count
control on accounted, every member of the kill set driven both ways round
(124, 130, 137, 141, 143, 128, 149), the verdict controls (0, 1, 2
and 127 — unusual-looking, below the floor, and still run), a kill code beside an
unreasoned claim staying UNRUN, the rendered contradiction line for the
verdict-plus-claim case, and readings B and C unchanged.

Verification

All readings below are from real output on this branch head dda4d903e
(origin/main merged in at ca7886047), exits captured by redirect-then-capture.

Self-test, before and after. node scripts/pm/dispatch-gates.mjs --self-test

before:  ✓ dispatch-gates self-test: 1682 cases pass.
after:   ✓ dispatch-gates self-test: 1712 cases pass.   (exit 0)

Falsifiability, one-shot. With the classification reverted on disk — the single call
site const killed = runRecordKillLabel(recorded.code); replaced by const killed = null;,
which restores the pre-fix fall-through exactly — 16 of the new cases fail BY NAME and
every control stays green:

✗ READING A — a bare cap-kill code is UNRUN, never run: the run left no verdict to read
✗ ...and the unrun row NAMES the recorded code, so the runner can see what this tool read
✗ READING D — a kill code PLUS a reasoned claim is NOT-MEASURED: the honest runner's declaration WINS ...
✗ ...and its row carries BOTH halves of the record — the code this tool read and the reason the runner stated
✗ ...and the evidence counts that ONE family once: it is inside `coded`, and `killClaimed` is reported beside ...
✗ a recorded exit 124 / 130 / 137 / 141 / 143 / 128 / 149 is a kill: UNRUN bare, NOT-MEASURED when declared   (7 cases)
✗ a kill code beside an UNREASONED claim stays UNRUN — the claim costs a reason here exactly as it does alone
✗ the NOT-MEASURED · KILLED block prints, and names the code and the reason on the family's own row
✗ ...and reading D is reported as the two channels AGREEING — ⛔ never as a record to repair
✗ reading A names its kill in the UNRUN block and prescribes the spelling that would declare it

Readings B and C, the verdict controls (0, 1, 2, 127) and the rendered
verdict-plus-claim contradiction line all stay ✓ under the ablation, which is what says
the new branch sits beside the two existing channels rather than over either of them.

The mutation was proven on disk before the run (injected text present once, deleted text
absent, and the blob hash moved: 224b8a91 → adf94009), and the restore was proven by
bytes and not by an exit code — git hash-object back to 224b8a91, git diff HEAD
empty, git status --porcelain empty. The script carried
trap restore EXIT INT TERM throughout, since a cap kill landing mid-mutation is the
exact failure this card is about.

Gates. node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack
derives 31 families for this diff on dda4d903e — the tool deriving its own families,
since the changed path is the tool. All 31 were run with each exit captured by
redirect-then-capture, recorded as CMD :: exit CODE, and handed back through --ran,
so the tool judged its own record on the same run that exercises the new classification:

Run reconciliation — 31 derived, 31 run, 0 NOT-MEASURED, 0 UNRUN.
  EXIT CODES — all 31 accounted famil(ies) carry one, so the NOT-MEASURED count above is DERIVED from them. ⛔ This tool ran none of them; it read the codes you recorded.
✓ dispatch-gates --ran: 31 derived famil(ies) accounted for — 31 run, 0 NOT-MEASURED (a DERIVED zero — all 31 recorded an exit code and none of them is 3).

Every family exited 0, check:pm-dispatch-gates included — the family that was
cap-killed at the container ceiling on the round that surfaced this card, and the reason
that round's ledger read 81 run over a gate that never reached a verdict.

npx eslint --no-inline-config scripts/pm/dispatch-gates.mjs — exit 0, no output. The
repo-level pnpm lint scan is CI's run, not this PR's, and is not claimed here.

The changed file is self-scanned for control bytes
(grep -naP '[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]'): none, and check:nul-bytes is
among the 31 green families.

Acceptance notes

  • Rider dispatch-gates: STAGE-THEN-RUN reaches a program by an edge neither follow traverses — check:objectui-changeset inherits nothing from scripts/bump-objectui.sh #14290, not folded. See P4. scripts/pm/dispatch-gates.mjs is that hold's
    Restart-touch: file, so this PR names it; the change does not meet the
    STAGE-THEN-RUN edge it describes and stops short of it. Its re-pricing belongs to
    domain:devx.
  • skip-changeset is the declaration. The changeset-check job in
    .github/workflows/pr-automation.yml counts added .changeset/*.md files with
    git diff --name-only --diff-filter=A MERGE_BASE HEAD -- '.changeset/*.md' and has
    no path exemption — its only two exemptions are the live skip-changeset label
    and the changesets release branch. scripts/pm/dispatch-gates.mjs sits at the repo
    root, whose package is private: true, and no released package's files[] names
    scripts/pm, so this diff publishes nothing.
  • Noted, not filed — the evidence block's "a DERIVED zero" sentence is reachable
    only from the ✓ line, so it can never be printed over a killed family; the comment
    now says so, but the coupling between that sentence and the caller that reaches it
    is implicit rather than asserted. Observation, not a defect. Carrier: the next PR to
    touch runRecordEvidenceLines / notMeasuredEvidenceTerm.

Clause-②: no

🤖 Generated with Claude Code

https://claude.ai/code/session_01DAcomhvR9kKizeYgg89Vo8


Generated by Claude Code

A recorded exit code that is a KILL rather than a verdict — 124 from a
`timeout` wrapper, or any code at or above the 128 signal floor — is a
number the gate never chose: the process was ended before it produced a
result. Counting such a family as `run` asserts a measurement that does
not exist, which is the false green the whole reconciliation is built to
refuse, and it is what the reconciler did.

The set is one named constant, `RUN_RECORD_KILL_EXITS`, judged in one
place (`runRecordKillLabel`), so no branch and no rendered line carries a
bare code of its own. A killed family lands:

  - NOT-MEASURED when the record ALSO carries a claim with a stated
    reason for that command — the claim WINS over the run line beside
    it, because the runner who recorded the code and declared the family
    has said everything there is to say. Its own source, its own block
    (NOT-MEASURED · KILLED), and a row naming both the code and the
    reason. It is counted once: the family is already inside `coded`, so
    `killClaimed` is reported beside the other counts and added to none.
  - UNRUN otherwise — the direction that costs a rerun.

A run line carrying a VERDICT code (0, 1, 2, and anything else below the
floor) plus a claim stays a genuine contradiction, run still wins, and
the contradiction line still prints on the default output. The conflicts
rendering now reads the class the reconciliation actually assigned
instead of re-deciding it, so it cannot report a reading the totals above
it do not share.

Claude-Session: https://claude.ai/code/session_01DAcomhvR9kKizeYgg89Vo8
Co-authored-by: Claude <noreply@anthropic.com>
The battery pins the card's four readings as four readings of ONE
instrument — same command, same derivation, only the record's spelling
differs — plus the controls that say the kill branch was inserted BESIDE
the two existing channels rather than over either of them: exit 3 is
still a DERIVED refusal, a lone reasoned claim is still CLAIMED, and a
verdict code (0, 1, 2 and 127, which is unusual-looking but below the
floor) still reads as run. Every member of the kill set is driven both
ways round, bare and declared, so the floor is a floor rather than the
four signals that have been met so far.

The contradiction line is read off the RENDERED text and not out of
`conflicts`: a non-empty array says the tool noticed, and the card's own
reproduction filtered the output down to the count line, so it could not
tell whether anything was ever printed. Measured before writing: the
line does print, and it prints for a verdict-plus-claim record unchanged.

The existing both-ways case is re-spelled, not weakened — same
expectation, and the sentence now says which run lines it covers, since
a bare line carries no code for a kill to be read out of.

Claude-Session: https://claude.ai/code/session_01DAcomhvR9kKizeYgg89Vo8
Co-authored-by: Claude <noreply@anthropic.com>
The first ablation of the kill classification did not report which cases
the regression moved: the reverted reading leaves `unrun` empty, the pin
read `unrun[0].why`, and the TypeError took the whole battery down at
that line — so 1712 pins produced one named failure and a stack trace.
A pin that crashes on its own subject is a pin that can only be graded
by the author who already knows the answer.

The row-content assertions now read through a helper that yields '' for
an absent entry. The class assertions beside them need no guard: each
one follows a `length === 1` conjunct that short-circuits.

Claude-Session: https://claude.ai/code/session_01DAcomhvR9kKizeYgg89Vo8
Co-authored-by: Claude <noreply@anthropic.com>
@claude claude Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Sep 14, 2026
@claude

claude Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author
  • Served-tier: 1917/1917 CONTRACT_REVIEW_TIER — harness model stamp counted over this seat's own transcript (non-sidechain assistant messages a model served; <synthetic> harness notices excluded) at 2026-09-14T03:55Z and compared to the constant's value outside the repository; get_session external_metadata.last_served_model read equal to the constant at 2026-09-14T00:20Z.

Contract review

Head: dda4d903 (PR #18108, card #18074) — reviewed at 2026-09-14T03:55Z by the skills seat at the contract-review tier. Non-governed path (scripts/pm/dispatch-gates.mjs only) ⇒ the in-seat review lands it: ready + auto-merge through the queue. The final head merges origin/main (ca788604) as a merge commit; the file's content is byte-identical to the pre-read head 2b1949dd (git diff between the two on the file: 0 lines), so the seat's readings on 2b1949dd stand.

① derived judgments (seat-measured on the fetched head, ⛔ not taken from the report):

  1. One named set, one judge: RUN_RECORD_KILL_EXITS (TIMEOUT 124, SIGNAL_FLOOR 128, the named signals 130/137/141/143) and runRecordKillLabel(code) — the only place a code is judged a kill; no bare number in the loop or the rendering.
  2. The classification, seat-reproduced on the head with the card's own four records (--ran, --repo objectstack-ai/objectstack, 73 derived): A (exit 124 alone) → 0 run, 0 NOT-MEASURED, 73 UNRUN (was 1 run); D (exit 124 + a reasoned claim) → 0 run, 1 NOT-MEASURED in a KILLED block that keeps the declaration and names the code and the reason (was 1 run, the claim discarded); C (the claim alone) → 1 NOT-MEASURED · CLAIMED, unchanged; V (exit 0 + a claim) → 1 run with the contradiction line printed 「is recorded BOTH as run and as NOT-MEASURED. Read as run; fix the record」 — and for D the line reads 「The two AGREE: exit 124 is a kill, not a verdict. Read as NOT-MEASURED with your stated reason」. Counted once: killClaimed is reported beside the evidence counts and added to none.
  3. The card's P1 answered: the contradiction line always printed on the full default output; the card's reproduction filtered on reconciliation and could not see it — pinned now by a self-test that reads the rendered text, not the conflicts array.
  4. Self-test: 1682 → 1712 cases pass (exit 0; the seat ran the battery on 2b1949dd in a worktree: 1712 pass); 19 multi-line t( pins added, the one removed pin (「claimed BOTH ways reads as run」) re-spelled as the verdict-plus-claim case with the same expectation for verdict codes; the dev's ablation (call site → null) reds 16 new cases BY NAME (A, D, all seven kill codes both ways, the unreasoned-claim case, the three rendered-text pins) while B, C, the verdict controls 0/1/2/127 stay green; restore proven by blob hash + empty git diff HEAD.
  5. Gates: 31 derived / 31 run / 0 NOT-MEASURED / 0 UNRUN on the final head — the tool deriving its own families and judging its own record; check:pm-dispatch-gates green; eslint on the file exit 0. Checks on the head at 2026-09-14T03:55Z: 0 red at read time.
  6. Scope held: one file; skip-changeset is the declaration; hold dispatch-gates: STAGE-THEN-RUN reaches a program by an edge neither follow traverses — check:objectui-changeset inherits nothing from scripts/bump-objectui.sh #14290's Restart-touch: file touched with nothing folded (its re-pricing is the domain:devx lane's; no STAGE-THEN-RUN interaction met); Clause-②: no holds — no contract path; --pair 18108 on origin/main's reader → exit 0 at 2026-09-14T03:55Z before this record.

② semver: unchanged — PM tooling, nothing published.

③ boundary flags: none binding. Two readings for the file's next editor: self-test pins that index x[0].field crash the battery instead of failing by name when the array empties (fixed for the new pins, the old shape left alone); the 「a DERIVED zero」 sentence is reachable only from the OK line (an implicit coupling, commented).

Implemented-by: claude/issue-18074-ran-record-cap-kill-is-not-run
Reviewed-by: session_01DAcomhvR9kKizeYgg89Vo8

Verdict: PASS — ready + auto-merge by this seat.


Generated by Claude Code

@os-project-manager
os-project-manager marked this pull request as ready for review September 14, 2026 03:55
@os-project-manager
os-project-manager added this pull request to the merge queue Sep 14, 2026
Merged via the queue into main with commit 7c7e76f Sep 14, 2026
40 checks passed
@os-project-manager
os-project-manager deleted the claude/issue-18074-ran-record-cap-kill-is-not-run branch September 14, 2026 04:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/m skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

2 participants