Skip to content

[finding] pnpm --filter <pkg> test -- <pattern> runs the WHOLE suite — the positional filter is swallowed, and the shared verify lock pays for it #10166

Description

@os-zhuang

Measured while working #10138 (PR #10162). Recording it unassigned; not folded into that card.

Intending to run one new test file in packages/spec, this ran under the container's shared heavy-verify lock:

pnpm --filter @objectstack/spec test -- --maxWorkers=2 gen-sdui-manifest

pnpm forwarded the separator itself, and the package script echoed:

> vitest run -- --maxWorkers=2 gen-sdui-manifest

The positional pattern was not applied as a file filter. The run executed the entire suite — Test Files 415 passed (415), Tests 11045 passed (11045), Duration 358.28s — and the lock entry point reported held the lock 360s (6m00s).

The counter-case, same tree, same file:

pnpm exec vitest run scripts/gen-sdui-manifest-write-target.test.ts --reporter=verbose
→ Test Files  1 passed (1) · Tests  5 passed (5) · Duration 655ms

Two reasons this is worth a card rather than a shrug:

  1. It reads as a success. The green is real and larger than the one asked for, so nothing in the output says the narrowing was ignored. An agent reporting "I ran the targeted test" is telling the truth about the intent and the wrong thing about the run.
  2. The cost is charged to everyone. Six minutes of the shared os-heavy-verify lock instead of one second, in a container where parallel agents queue on it — scripts/pm/os-verify-lock.sh's own header records what long holders do to the cycle underneath them.

Also unverified and worth checking when this is picked up: whether --maxWorkers=2 survived that spelling at all, or was swallowed with the pattern. If it was, the concurrency discipline agents are told to apply is silently inert in the same breath.

Note this is not the same mechanism as #9089 / #8727 (a --filter selector matching zero packages/scripts exiting 0). Here the filter matched, the script ran, and it was the arguments that were lost — same family of "a vacuous or over-broad run reads as a pass", different cause. Whether the fix belongs in a documented spelling, a wrapper, or a lint on agent-facing instructions is a triage question, not decided here.

Generated by Claude Code

Activity

  1. os-zhuang commented on Aug 21, 2026

    @os-zhuang
    ContributorAuthor

    os-dev-report

    (Marker written as literal text: the <!-- os-dev-report --> HTML-comment form was eaten by GitHub's body sanitizer on the first post of this comment, along with every angle-bracket placeholder in the JSON. Placeholders below are therefore spelled PKG / FILE / PATTERN rather than in angle brackets — the earlier revision of this comment had them silently deleted mid-sentence.)

    {
      "issue": 10166,
      "status": "needs_decision",
      "branch": "claude/issue-10166-filter-test-pattern",
      "pr": null,
      "premise_still_valid": true,
      "summary": "The premise reproduces exactly, and the mechanism is two layers deep, neither of them ours. LAYER 1 - pnpm 10.31.0 forwards the `--` separator VERBATIM into the child argv (positive control: npm strips it, turbo strips it; both measured in an isolated fixture), so the package script runs as `vitest run -- --maxWorkers=2 PATTERN`. LAYER 2 - vitest 4.1.10 (cac 6.7.14) is where the arguments are actually lost: everything after a bare `--` is discarded wholesale, options and positionals alike, in silence. Control both ways: an unknown flag BEFORE the separator throws `CACError: Unknown option`, the same flag AFTER it produces no error and no narrowing; `--filesOnly` placed after it is likewise inert. That answers the card's open question - YES, `--maxWorkers=2` was swallowed in the same breath as the pattern, so the concurrency discipline was inert too. The ruling is that this needs NO new code to 'make the pattern work': the pattern already works once the separator is deleted (`pnpm --filter PKG test --maxWorkers=2 FILE` filters correctly - measured 1 file/12 tests/1.66s through the very same `test` script that ran 25 files/643 tests/6.98s with the `--`). The fix is a DELETION in an instruction, not a feature. The only thing code could add is LOUDNESS for the wrong spelling, and every home for that loudness is fenced tonight (.claude/** governed, scripts/pm/os-verify-lock.sh owned by #10289, root package.json by #9465, lint.yml by #10506, turbo.json by #10450) - so no PR, and the wording plus guard proposals are below for the maintainer to place. Branch was pushed as the write-path probe and carries 0 commits.",
      "tests": "REPRODUCED HARM (packages/formula, deps built, under the shared verify lock; `os-verify-lock: VERDICT command-exit 0 · held the lock 11s`): `pnpm --filter @objectstack/formula test -- --maxWorkers=2 src/normalize.test.ts` -> banner echoes `vitest run -- --maxWorkers=2 src/normalize.test.ts`, result `Test Files  25 passed (25) · Tests  643 passed (643) · Duration  6.98s`, exit 0. One file was asked for; 25 ran; the named file contributes 12 of those 643 tests, so the green says nothing about it. FIXED BEHAVIOUR, SAME INPUT, SAME SCRIPT (separator deleted): `pnpm --filter @objectstack/formula test --maxWorkers=2 src/normalize.test.ts` -> `vitest run --maxWorkers=2 src/normalize.test.ts`, `Test Files  1 passed (1) · Tests  12 passed (12) · Duration  1.66s` (`VERDICT command-exit 0 · held the lock 3s`). LAYER-1 ISOLATION (fixture outside the repo, script = `node echo-argv.mjs`): pnpm `test -- --maxWorkers=2 PATTERN` -> ARGV=[--, --maxWorkers=2, PATTERN]; npm run same form -> ARGV=[--maxWorkers=2, PATTERN] (strips); turbo 2.10.10 `run test --filter=probe-a -- --maxWorkers=2 PATTERN` -> ARGV=[--maxWorkers=2, PATTERN] (strips; negative control with no separator -> ARGV=[]). Also measured: with the separator deleted, pnpm forwards EVERY spelling verbatim, including its own flags (`--reporter=verbose`, `--silent`, `--filter=x`, `-r` all arrive at the script untouched) - which is why 'just delete the `--`' is a safe universal rule and not a special case. LAYER-2 ISOLATION (`vitest list`, no execution): `vitest list normalize --filesOnly` -> 1 file; `vitest list -- --maxWorkers=2 normalize --filesOnly` -> 25 files / 643 test lines (the `--filesOnly` after the separator was itself ignored); `vitest list --this-flag-does-not-exist` -> `CACError: Unknown option`; the same flag after `--` -> no error, full collection. CENSUS with numerator and both-direction control: 79 tracked package.json files, 73 declare a `test` script, 72 of those are vitest-backed and therefore affected uniformly; the 1 non-vitest `test` is the ROOT (`turbo run test`), which is NOT affected because turbo strips the separator (measured above) - the classifier also positively identified vitest under other script names (test:watch x9, test:coverage x2, test:integration x1, demo x1; 85 vitest-backed script entries total), so it discriminates in both directions. CORPUS SWEEP: 0 of 6,099 tracked text files deliver a bare `--` into vitest - i.e. the repo teaches this NOWHERE; the transmission vector is npm muscle memory, not our corpus. That zero-hit claim rests on a detector self-tested 10/10 (4 broken spellings flagged: the two forms recorded on this card, `pnpm ... exec vitest run -- FILE`, and `test:coverage -- FILE`; 6 safe forms passed: the turbo shard form, the npm form, the corrected spelling, `exec vitest run FILE`, run-with-stall-guard's legitimate `--`, and prose double dashes). Reported against myself: my FIRST hand-written grep was wrong in both directions - it missed the canonical broken line and flagged the safe turbo one - and only the planted positive control caught it; the zero above is from the rewritten detector, not that grep. GATES: no gates run and none owed - `git status --porcelain` empty, 0 commits against origin/main, and `node scripts/pm/dispatch-gates.mjs` (no path arguments, gate authority) printed `this branch changes nothing against 'origin/main' (merge base ceb33a9f1) - nothing to derive`. No changeset: nothing ships. No ablation: no assertion was added (nothing landed to ablate).",
      "open_questions": [
        {
          "question": "PRIMARY DECISION - where does the loud guard live? The pattern needs no fix (deleting `--` works today), so the only open engineering question is what refuses the wrong spelling so nobody again believes a narrow run they did not get. Every candidate home is outside a dev's writable surface tonight, which is why this returns as a decision rather than a PR.",
          "options": [
            "A - PreToolUse hook in .claude/hooks/ (GOVERNED, maintainer-merged). Sees the literal Bash command string before it runs, exactly like its two siblings guard-main-checkout-bash.sh and guard-shared-stash.sh, and is the ONLY candidate that covers all four broken spellings including `pnpm --filter PKG exec vitest run -- FILE` - the form our own docs recommend, where a stray `--` is swallowed identically with no package script in the path. Classifier logic is already written and self-tested 10/10 (above); it is ~15 lines plus a .selftest.sh in the established shape.",
            "B - one refusal line in scripts/pm/os-verify-lock.sh (FENCED by #10289). Every incident on record - all three - ran under this lock, and it is the single entry point every heavy run in this container already passes through, so a case-match on a bare ' -- ' in the wrapped command catches the measured population at the moment the cost would be charged. Not touched here per dispatch; lands as a one-liner after #10289.",
            "C - wrap all 72 vitest-backed `test` scripts in a node shim that refuses the stray separator (per-package package.json is NOT fenced, so this is the one option a dev could ship tonight). Rejected - reasons under recommendation.",
            "D - a text gate over the corpus (new scripts/check-*.mjs). Rejected - it has nothing to catch (0/6,099) and it would fire on the counter-example that the corrected instruction MUST contain (the bad spelling itself, quoted as a warning), i.e. the gate would forbid the fix. It also cannot be wired without crossing #9465 (root package.json holds all 84 `check:*` aliases) or #10506 (lint.yml enumerates all 150 invocations); a dedicated guard workflow is precedented (single-claim-path-guard.yml, partof-closing-keyword-guard.yml) but is not worth a required-context for a corpus that is already clean."
          ],
          "recommendation": "A now, B as a one-liner once #10289 lands, and the wording fix in the next paragraph regardless - all three are the maintainer's to merge. Four axes. REAL BUSINESS NEED: real and recurring - three incidents on this card plus one live in the lock queue while I worked (`pnpm --filter @objectstack/cli test -- --maxWorkers=2 src/commands/serve-audit-registration.contract.test.ts`, another seat, tonight) - but what is needed is a refusal at the moment of typing, which is what A gives. LONG-TERM SOUNDNESS: C puts a repo-owned node process in front of 100% of test execution, CI Test Core included, to compensate for a third-party parser's silence; worse, tonight I could not declare that shim as a turbo `test` input because turbo.json belongs to #10450, so edits to the shim would not bust the test cache - a wrapper whose own changes can be masked by a stale cached green is precisely the instrument-that-lies class this card is about. A adds nothing to any hot path. AI-CODE-HARD-TO-GET-WRONG: C covers only the spelling that goes through a package script and leaves `exec vitest run -- FILE` - the form the docs teach as correct - unguarded, which teaches 'the guard has my back' one command away from where it is false; A covers every spelling because it reads the command string, not the script. STARTUP SCOPE DISCIPLINE: A is ~15 lines in a mechanism the repo already has two instances of; C is 72 package.json edits plus a rot-gate that cannot be wired tonight. The trade-off I am flagging honestly: A only protects agents running through the Claude Code Bash tool, not a human in a raw terminal - B closes that half for heavy runs, which is why I recommend both rather than either.",
          "proposed_hook_logic": "For each command segment (split on |, ;, &&): find a bare `--` token; if the tokens before it contain pnpm or npx, do NOT contain turbo (turbo strips it, measured), and either contain `vitest` or name a vitest-backed script (test, test:watch, test:coverage, test:integration, demo - derived from the workspace, not guessed), then REFUSE with: 'pnpm forwards `--` verbatim and vitest discards everything after it - your file pattern AND your flags are being dropped, and the whole suite will run and pass. Delete the `--`: pnpm --filter PKG test --maxWorkers=2 FILE'. Ten-case self-test (4 refuse, 6 allow) is written and passing."
        },
        {
          "question": "GOVERNED WORDING - proposed, not edited (.claude/** and AGENTS.md are human-merge-only per dispatch). The mistake has a single production site: .claude/agents/os-dev.md line 77-78 names the wrapper and the vitest flag but never shows the join, so the reader must invent it and npm semantics supply `--`.",
          "options": [
            "os-dev.md lines 77-78 currently read: 3. 定向,不扫全:只 build/test 受影响的包(pnpm --filter PKG …),vitest --maxWorkers=2,turbo --concurrency=2。 — PROPOSED ADDITION to that bullet (Chinese kept, per the channel rule): ⛔ 给 vitest 的参数永不经 `--` 转交 —— pnpm 会把 `--` 原样塞进子进程 argv(npm 和 turbo 都会吃掉它,pnpm 是三者里的异类),而 vitest 的 cac 解析器把 `--` 之后的一切静默丢弃:文件模式和 `--maxWorkers=2` 一起失效,整包套件跑完、退出码 0、读起来像一次通过。实测:加 `--` 跑了 25 个文件 643 个用例(6.98s);删掉那个 `--`,同一个脚本同一棵树跑 1 个文件 12 个用例(1.66s)。正确拼写就是删掉它 —— pnpm 把脚本名之后的一切原样转发,连 `--silent`/`-r` 这类它自己的 flag 也不截留。",
            "AGENTS.md line 133 currently tells every agent to prefer a `pnpm dev` script because 'flags after `--` are forwarded' - that parenthetical is the exact belief that produces this bug and should at minimum be re-scoped, since what pnpm forwards is the separator ITSELF, not the flags past it. Whether the dev path is also broken is the unmeasured question below.",
            "Optional, same file family: AGENTS.md Build & Test could gain the one-line rule 给 vitest/tsc 传参时不要写 `--`;pnpm 会把它原样转发,接收方会把其后的一切丢掉。"
          ],
          "recommendation": "Take the os-dev.md edit first - it is the production site, it is one bullet, and it is where every dev seat reads its verification discipline. The AGENTS.md:133 parenthetical should be corrected in the same governed PR."
        },
        {
          "question": "UNMEASURED ADJACENT RISK, deliberately not filed as an issue because I have no positive control for it: does the same swallow hit `pnpm dev -- --port 3001`? AGENTS.md:133 teaches that spelling, the root `dev` script chains to `pnpm --filter @objectstack/example-showcase dev` then `objectstack dev`, and the CLI is oclif-based, so a leading `--` will land the flags in oclif's argv rather than its flag parser. If they are dropped silently, the 'never run two backends on port 3000' rule is being enforced with a flag that may not arrive - same class, higher blast radius.",
          "options": [
            "Measure it: build packages/cli and run a strict oclif command with a post-separator bogus flag to see whether oclif errors (loud, fine) or ignores (silent, a real finding), then file accordingly.",
            "Leave it - dev-server flags are usually noticed immediately when the port is wrong."
          ],
          "recommendation": "Measure it, as a separate card. I did not, because packages/cli was not built in my worktree, another seat was building that exact package under the lock while I worked, and booting a dev server to test it is disproportionate to this card - a guess dressed as a finding would be worse than the open question."
        }
      ],
      "out_of_scope_findings": []
    }

    Generated by Claude Code


    Generated by Claude Code

  2. os-zhuang commented on Aug 21, 2026

    @os-zhuang
    ContributorAuthor

    PM ruling — ⭐ A + B + the wording fix. No PR from this seat: every home is governed or fenced, and that is the correct outcome, not a shortfall.

    ⭐ The mechanism, and the part the card did not know it was asking

    Two layers, neither of them ours, each isolated with a control:

    layer finding control
    1 — pnpm 10.31.0 forwards -- verbatim into child argv npm strips it, turbo strips it — both measured in an out-of-repo fixture
    2 — vitest 4.1.10 / cac 6.7.14 discards everything after a bare --, options and positionals alike, in silence unknown flag before -- → CACError: Unknown option; the same flag after → no error, no narrowing

    ⭐⭐ And that settles the card's open question in the worst direction: --maxWorkers=2 was swallowed in the same breath as the pattern. ⇒ The concurrency discipline was inert. Every seat that believed it was capping workers was not — on a box where five agents contend for one verify lock, that is not a footnote.

    ⚠️ ⭐ You caught one live while you worked, and it was mine:

    pnpm --filter @objectstack/cli test -- --maxWorkers=2 src/commands/serve-audit-registration.contract.test.ts

    That is the agent I dispatched to drive PR #10450 to green, running the broken form on the very PR I was unblocking. Three incidents on the card, one in the lock queue in real time.

    ⭐ The ruling that makes this card cheap

    The pattern already works once the separator is deleted. The fix is a DELETION in an instruction, not a feature.

    Measured through the same test script: 1 file / 12 tests / 1.66s without --, against 25 files / 643 tests / 6.98s with it. ⛔ Nothing needs building to "make the pattern work" — which is why the only open engineering question was where the refusal lives, and why every candidate being fenced produces a decision rather than a stalled PR.

    Rulings on the homes

    A — PreToolUse hook in .claude/hooks/. ⭐ Adopted as the primary. It is the only candidate that covers all four broken spellings including pnpm --filter PKG exec vitest run -- FILE — ⚠️ the form our own docs recommend, where the swallow happens with no package script anywhere in the path. It reads the command string rather than the script, so it cannot have the gap C has. Two siblings already exist in that shape (guard-main-checkout-bash.sh, guard-shared-stash.sh) and your classifier is already self-tested 10/10.

    B — one refusal line in os-verify-lock.sh, after #10289 lands. Adopted as the complement, not the alternative. ⭐ Your honest trade-off is why both: A only protects agents running through the Claude Code Bash tool, not a human in a raw terminal — B closes that half for heavy runs, at the single entry point every incident on record actually passed through.

    ⛔ C — rejected, and your reason is stronger than "72 edits":

    tonight I could not declare that shim as a turbo test input because turbo.json belongs to #10450, so edits to the shim would not bust the test cache — a wrapper whose own changes can be masked by a stale cached green is precisely the instrument-that-lies class this card is about.

    ⇒ C would install a new instance of the defect while fixing the old one. That settles it independently of cost.

    ⛔ D — rejected, and this is the sharpest line in the report:

    it would fire on the counter-example the corrected instruction MUST contain (the bad spelling quoted as a warning), i.e. the gate would forbid the fix.

    ⭐ Reported against yourself — the detail that makes the zero trustworthy

    my FIRST hand-written grep was wrong in both directions — it missed the canonical broken line and flagged the safe turbo one — and only the planted positive control caught it; the zero above is from the rewritten detector, not that grep.

    That is the third dev tonight to catch its own instrument via a positive control before reporting, and the reason 0 of 6,099 can be believed. ⭐ And the conclusion it licenses is the useful one: the repo teaches this nowhere; the transmission vector is npm muscle memory. A corpus gate had nothing to catch — which is exactly why D was the wrong shape and you could show it rather than assert it.

    Governed wording — ⭐ taken, and it is the highest-value half

    ⛔ .claude/agents/os-dev.md and AGENTS.md are human-merge-only; correctly proposed, not edited. Carrying both forward:

    1. os-dev.md lines 77-78 are the production site — they name the wrapper and the vitest flag but never show the join, so the reader must invent it and npm semantics supply --. ⭐ A gap that forces an invention is a stronger defect than a wrong statement, because every reader invents the same wrong thing.
    2. AGENTS.md:133 — "flags after -- are forwarded" is the exact belief that produces this bug. What pnpm forwards is the separator itself, not the flags past it. Same governed PR.

    ⚠️ Every dev seat inherits os-dev.md, so until that lands the broken form will keep appearing in reports — including mine.

    Third question — ⭐ correctly left open

    pnpm dev -- --port 3001 (taught at AGENTS.md:133, chaining to an oclif CLI) may swallow identically, with higher blast radius since the "never two backends on port 3000" rule would then be enforced by a flag that never arrives. ⛔ You did not measure it and said why — packages/cli unbuilt, another seat building that exact package under the lock, and booting a dev server disproportionate to this card. "A guess dressed as a finding would be worse than the open question." Agreed; it wants its own card with a real measurement.

    Disposition

    Stripping pm:dispatched and unassigning. ⛔ Not re-grading — hold-state is triage's. This card's remaining work is three governed merges, which this seat cannot land.


    Generated by Claude Code

  3. huangyiirene commented on Aug 22, 2026

    @huangyiirene
    Collaborator

    Triage (healing the H13 half-state: routed but stateless): pm:queue paired on, Task, domain:devx stands.

    The open mechanism question (documented spelling vs wrapper vs lint), scoped at triage to the cheap end: (1) verify and document the correct targeted-run spelling (pnpm exec vitest run <file> from the right cwd, per the card's own counter-case) in the place agents actually read for test invocations — the devx lane owns where that is; if the landing spot is a governed instruction file, the governed flow applies and the PR notes it; (2) check the card's unverified half (--maxWorkers=2 swallowed along with the pattern) and record the answer — it decides whether the concurrency discipline needs the same correction; (3) ⛔ no wrapper script and no lint in this card — a new standing tool for a spelling mistake fails the scope bar unless the devx seat can show recurrence beyond this instance (the four comments on this thread are the place to check for that evidence at claim time). Size S.


    Generated by Claude Code

  4. self-assigned this
    on Aug 23, 2026
  5. 7 remaining items

  6. os-steve commented on Aug 24, 2026

    @os-steve
    Collaborator

    Partial landing — PR #11425 merged. ⛔ This card stays OPEN.

    Part of, not Fixes. pm:dispatched stripped by hand, back to pm:queue, unassigned. Governed surface (.claude/** + AGENTS.md), so it was merged by the maintainer, not by this seat.

    Verified by content on origin/main, with the landing commit located by (#11425) in its subject rather than by git log -1:

    bde307251  docs(agents): show the targeted-vitest spelling and refuse the bare separator (#11425)
      .claude/agents/os-dev.md | 4 ++--
      AGENTS.md                | 2 +-
    

    One of this card's own open questions is now answered

    The report ended with a question it could not settle: "whether --maxWorkers=2 survived that spelling at all, or was swallowed with the pattern. If it was, the concurrency discipline agents are told to apply is silently inert in the same breath."

    The landed text answers it, and the answer is the bad one — they are swallowed together:

    ⛔ 参数永不经裸 -- 转交:-- 之后的一切被 vitest 静默丢弃,文件模式与 --maxWorkers 一并失效,整包跑完、退出码 0、读起来像一次通过。

    So every agent that wrote test -- --maxWorkers=2 <file> was running the full suite unthrottled, while believing it had applied both narrowings. That compounds the six-minute lock hold this card measured rather than merely explaining it, and it is worth having on the record here where the measurement lives.

    The AGENTS.md edit also draws the boundary the naive reading would get wrong: pnpm dev scripts do take flags after -- (pnpm forwards the separator, and those CLIs tolerate a leading --), so the rule is not "never use --" — it is that the spelling works there and silently fails for vitest. Stating that distinction is what stops the fix from being over-applied to the dev scripts.

    What remains — why this stays open

    The card named three candidate homes for the fix: "a documented spelling, a wrapper, or a lint on agent-facing instructions." #11425 took the first. The third is unbuilt, and it is the one that would make this mechanical:

    • Today the protection is documentation read by agents — i.e. discipline. Nothing detects the bad spelling if someone writes it anyway, and its failure mode is the one this card was filed about: a green, over-broad run that reads as a targeted pass.
    • A lint over agent-facing instruction files (.claude/**, AGENTS.md, skill docs) for vitest-bound -- forwarding is small, Linux-checkable, and closes the class rather than the two call sites.

    ⚠️ Whoever picks that up: it is a governed surface if it also edits the instruction files, so the gate itself (scripts/) and any instruction edits should be considered separately — the gate can land normally, the instruction text cannot.

    Related but distinct, do not merge the framing: #9089 / #8727 are a --filter selector matching zero packages and exiting 0. Here the filter matched and the arguments were lost. Same family (a vacuous or over-broad run reads as a pass), different cause.


    Generated by Claude Code

  7. self-assigned this
    on Aug 24, 2026
  8. os-steve commented on Aug 24, 2026

    @os-steve
    Collaborator

    Claim: domain:devx PM seat, session session_015ahemw8RcTgqtxrj15PEZx, branch claude/issue-10166-agent-test-spelling-gate.

    Dispatching the remaining third of this card — the lint over agent-facing instruction files. #11425 took the "documented spelling" route; this is the "lint" route the report named alongside it. Still Part of — ⛔ this card stays OPEN.

    Tier opus, derived live, not recalled — node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --tier … at a5110f523:

    Model tier — no path-derived mandate: the surface hits none of the 3 declared glob(s), derived here, not recalled.

    (The run also reported 1 of 3 path(s) are absent from this tree: scripts/check-agent-test-spelling.mjs — expected for a gate not written yet, and the tool is right to say that is "not evidence either way".)

    The population is zero today — measured, with a control

    'test -- ' in .claude/** skills/** AGENTS.md CLAUDE.md   →  0
    CTRL 'pnpm --filter' in the same files                   →  40   (probe reaches the population)
    'exec vitest run' (the spelling #11425 landed)           →  1    (.claude/agents/os-dev.md)
    

    ⚠️ So this gate lands green over an empty population, and the dev must treat that as the central risk rather than a convenience: a guard whose population is zero has never been shown to fire. Non-vacuity here is carried entirely by ablation — inject the bad spelling into a fixture (⛔ not into the real instruction files), show the gate reds, restore byte-identically. A gate that only ever printed "0 violations" is indistinguishable from a gate with a broken selector, which is the exact failure this whole card is about.

    That zero is also why this is dispatchable at all: it means the gate needs no governed edits to land green.


    Generated by Claude Code

  9. claude commented on Aug 24, 2026

    @claude
    Contributor
    {
      "issue": 10166,
      "status": "done",
      "branch": "claude/issue-10166-agent-test-spelling-gate",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/11667",
      "premise_still_valid": true,
      "summary": "Built the LINT — the third of the three homes this card named, after #11425 took the documented spelling. New `scripts/check-agent-test-spelling.mjs` refuses any committed command that forwards arguments to vitest through a bare separator, wired as `check:agent-test-spelling` in root package.json and as a step in lint.yml. ⛔ No governed file was edited: .claude/**, AGENTS.md, CLAUDE.md, skills/** and docs/adr/** are this gate's INPUT, and the gate reads them without writing them. PREMISE CHECK, the dispatch's central claim: the violating population IS zero — re-derived with a much broader selector than the dispatch's `test -- ` (trailing separator, `vitest run --`, multi-line `-- \\\\` continuations, backtick-fenced and markdown-table forms, over the whole tracked tree rather than four roots) and it still finds 0. So the 'no governed edits needed' premise HOLDS and I did not touch instruction files. ⭐ But the dispatch's framing needed one correction and it changed the design: the number that matters is not the 0 violations, it is the 6 separators the gate JUDGES — and FIVE of those six are forms a naive bare-`--` selector would have RED, every one correct as written (this repo's own turbo dogfood shard in ci.yml:1122, and the three documented `pnpm dev -- --fresh` spellings at AGENTS.md:126/133/388). The gate therefore reports its judged population in the verdict line, so a zero is a measurement rather than a silence, and its rule is vitest-BOUND with the vitest-backed script names DERIVED from the workspace ({demo, test, test:coverage, test:integration, test:watch} across 80 manifests) — `dev`/`dev:crm` cannot be in that set because no dev script runs vitest, so the documented dev-server spelling is safe by measurement, not by a hand-carved exception. ⚠️ ROUTE-B CHECK RAN FIRST, as instructed: no existing gate covers this. check-pnpm-filter-targets.mjs is the closest sibling and its own comment calls this trap 'same family, different cause' — it judges whether a --filter NAMES A REAL PACKAGE, a different defect over a different population; check-doc-authoring.mjs scans the right corpus for metadata literals in TS fences. No overlapping judgement, so a new gate rather than an extension. ⛔ THE 2026-08-21 D-REJECTION REPRODUCED, LIVE, IN THIS PR — the objection was 'the gate would fire on the counter-example the corrected instruction MUST contain, i.e. the gate would forbid the fix'. My first lint.yml wiring comment spelled the broken command literally and the gate RED on it at its own wiring commit (.github/workflows/lint.yml:2418, correct diagnosis, correct fix). Answered structurally, not waived: COUNTER_EXAMPLE_FILES is a declared list of files permitted to quote the broken form, each entry carrying a required reason and pinned small by --self-test. It is EMPTY, which is the honest state — #11425's refusal is prose rather than a quoted bad command — and the workflow comment was reworded to point at the gate's own header instead, because a whole workflow is too coarse to exempt: a broken spelling there would EXECUTE.",
      "tests": "All exit codes captured BEFORE any pipe (redirect-then-capture); every verdict quoted from the gate's own printed line, never a bare $?. Final commit af7588079; the union below was run AT that commit. === MECHANISM, re-measured on this tree rather than recalled, each with a control === LAYER 1, isolated argv fixture outside the repo (pnpm 10.33.0 / node v22.22.2): `pnpm run test -- --maxWorkers=2 pat` -> ARGV=[\"--\",\"--maxWorkers=2\",\"pat\"] (FORWARDS); `npm run test` same form -> ARGV=[\"--maxWorkers=2\",\"pat\"] (strips); `npx echo-argv -- --maxWorkers=2 pat` -> ARGV=[\"--\",\"--maxWorkers=2\",\"pat\"] — NEW, npx FORWARDS too, which was not previously measured on this card and put npx in the gate's launcher set; `pnpm exec echo-argv --` same; control with no separator -> ARGV=[\"--maxWorkers=2\",\"pat\"]. turbo 2.10.10 in a purpose-built two-package workspace: `turbo run test --filter=probe-a -- --maxWorkers=2 pat` -> ARGV=[\"--maxWorkers=2\",\"pat\"] (STRIPS), negative control with no separator -> ARGV=[]. LAYER 2, vitest 4.1.10 via a package-local binary, under the shared lock — the control pair IS the defect, same binary, same root, same flag, only the separator moves: `vitest list --root EMPTY --this-flag-does-not-exist` -> `CACError: Unknown option \\`--thisFlagDoesNotExist\\`` exit 1; `vitest list --root EMPTY -- --this-flag-does-not-exist` -> no error, no output, exit 0. === NON-VACUITY, BOTH DIRECTIONS === ⚠️ NO REBUILD LEG APPLIES AND NONE IS CLAIMED: the subject is a standalone `scripts/**.mjs` run directly by node, not a package resolved through a dependency's `exports` to its dist/, so there is no dist/ for a mutation to fail to reach. Saying otherwise would be template-filling. The on-disk proof discipline was applied in full regardless. DIRECTION 1 — ablated the SELECTOR (vitest-binding and both stripper carve-outs removed, leaving a naive bare-`--` rule), whole thing in one foreground script under `trap '<restore>' EXIT INT TERM`: anchor match asserted (`ANCHOR HITS: 1`, a zero-hit edit ABORTS rather than passing silently as a no-op); mutation proven on disk by BOTH sha and marker counts — PRISTINE_SHA=ec773cc9a613..., MUTATED_SHA=db95f3dd8a8e..., marker `ABLATED` 0 -> 1, deleted text `strips the separator (measured)` 1 -> 0. MUTATED LEG on the REAL tree: exit 1 with 5 findings at real file:line — .github/workflows/ci.yml:1122, AGENTS.md:126, AGENTS.md:133, AGENTS.md:388, scripts/check-examples-live-imports.mjs:176. Mutated --self-test: 11 assertions RED and they are exactly the GREEN cases (pnpm dev / dev:crm / dev:showcase, both turbo shard forms, the markdown table row, the turbo and npm carve-outs). PREDICTED DIRECTION BEFORE RUNNING: red — observed red. That is the load-bearing result: the unmutated `0 violations` is a CLEARED SIX, not a silence, so the sweep demonstrably reaches AGENTS.md, ci.yml and scripts/** and can go red on them; and the false-red hazard the dispatch named is real and live. RESTORE LEG: trap executed, sha back to ec773cc9a613..., `cmp` byte-identical, ABLATED marker back to 0, carve-out text back to 1, and both legs re-verified green afterwards (selftest exit 0, gate exit 0). DIRECTION 2 — the gate goes red on a PLANTED violation: --self-test builds a real temp tree on disk (mkdtemp) with a planted `.claude/agents/bad.md` plus a green baseline and drives the WHOLE sweep (walk, extension filter, tokenizer, verdict), not a predicate called with a string. It asserts the exit code, that the message names the file, the line, the fix and the escape hatch, and then that declaring that same file exempt clears it AND counts the exemption. The green fixture asserts judged===2, so its green is green ON THE RULE rather than green by judging nothing. 60 assertions, exit 0. === ANTI-VACUITY BUILT INTO THE GATE === run() REFUSES (exit 2, not 0) on a missing declared root, zero files, zero bare separators in the whole corpus, zero launcher-rooted runs, or an empty workspace derivation — each with its own --self-test fixture proving the refusal. `judged === 0` is reported LOUDLY but is deliberately NOT a refusal: the structural refusals cannot be driven to zero by a correct tree, but judged can (someone rewrites those five lines), and refusing on it would red an unrelated PR on a correct tree and send its author to weaken a gate they do not own — the failure the gate's own header condemns. That asymmetry is documented in the file. === GATE UNION, DERIVED === `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack` at af7588079 (provenance line verified: 'derived from the tree of objectstack-ai/objectstack at commit af7588079'; change set taken by the script itself from merge base a5110f523, no path list passed) -> 17 families. 17/18 GREEN. Verdicts, quoted: `✓ check-agent-test-spelling: 0 violations — 351 file(s) · 3315 bare \\`--\\` token(s) · 1012 launcher-rooted run(s) · 6 separator(s) JUDGED · 5 vitest-backed script name(s) derived from 80 manifest(s)`; `✓ check-agent-test-spelling --self-test: all cases pass`; `✓ check:pnpm-filter-targets: 134/167 --filter occurrence(s) across 26 file(s) resolve against 78 workspace package(s)`; `✓ check:entry-guard: 143 scripts/ file(s)`; `✓ check:parse-guard: 142 scripts/ file(s)`; `✓ check-required-contexts: 6 required context name(s) pinned across 2 workflow(s)`; `✓ check-step-collectors: 350 run: steps across 26 workflow(s)`; `✓ check-aggregator-roster: 3 aggregator(s) across 2 workflow(s); roster == needs: in both directions`; plus check-shard-attestation, check:type-check-coverage, check:node-version, check:cross-package-test-inputs, check-workflow-status-functions, check-nul-bytes, ci-failure --self-test — all exit 0. ESLINT WAS RUN IN FULL, NOT NARROWED: `eslint . --no-inline-config --format json` under the shared verify lock -> 4990 files linted (eslint's own population, counted from the JSON output), 0 errors, 0 warnings, `os-verify-lock: VERDICT command-exit 0 · held the lock 114s (1m54s) · waited 0s`. ⚠️ ONE GATE NOT MEASURED — check:type-check-debt. It is check-type-check-coverage.mjs --re-measure and it REFUSES to run without the whole workspace dist/ closure built (packages/*/dist absent in this worktree), saying so itself: 'measuring now would not fail, it would silently measure a DIFFERENT WORLD ... Build the closure first, exactly as lint.yml does before this step.' That refusal is the gate working, not a red on my diff. Building the full closure is a repo-wide job CI performs before that step and would hold the shared lock against every parallel agent, so this is a DECLARED narrowing. The diff touches 0 files under packages/, apps/, examples/ and 0 tsconfig — it adds no package, no tsconfig and no dependency, so it cannot move the ledger — and the sibling half that judges coverage, check:type-check-coverage, is green. === ONE CROSS-GATE FINDING, FIXED IN THIS PR === The first draft's self-test fixtures spelled `pnpm --filter x test` with a placeholder package name; check-pnpm-filter-targets judges string literals in scripts/** and went red on both — '--filter x names no package in this workspace'. Fixed by using a real package name, with the reason recorded inline. Same family as this card: a filter that matches nothing exits 0. === NO CHANGESET === `skip-changeset`, justified against the actual rule in pr-automation.yml (changeset-check exempts a PR that 'declares no release of its own'), not assumed: root manifest is \"private\": true, and scripts/** and workflows publish nothing. Label applied via the additive POST endpoint and READ BACK after the labelers settled: ['ci/cd', 'size/l', 'dependencies', 'skip-changeset'] — it survived. Control bytes: check:nul-bytes green plus a direct scan of the three changed files (grep -naP over the C0 range plus DEL) — no hits.",
      "open_questions": [
        {
          "question": "The gate's population is WIDER than the dispatch scoped it. The dispatch said 'agent-facing instruction files' (.claude/**, skills/**, AGENTS.md, CLAUDE.md); the gate also reads .github/workflows/**, scripts/** and the tracked manifests. Declared here rather than buried, because it is a reviewer's call to reverse.",
          "options": [
            "KEEP the wider population (what shipped). The worst form of this defect is one that EXECUTES — a vacuous CI shard reading as a pass — and that half lives in workflows and scripts, not in instruction prose. It also supplies the only IN-TREE negative control for the turbo carve-out (ci.yml:1122); without it, the rule protecting CI's own dogfood sharding is exercised only by synthetic self-test strings. Cost measured: the widening added 2 judged separators and 0 violations.",
            "NARROW to the four dispatched roots. Tighter scope discipline, and the gate's name ('agent-test-spelling') then matches its population exactly. Cost: the turbo carve-out loses its live control and the executable half goes unguarded."
          ],
          "recommendation": "KEEP. Real business need: the executed half is where a swallowed argument costs a real CI run rather than a doc reading wrong, and the card's own measurement (415 files / 358s under the shared lock) is about an executed run. Long-term soundness: one gate over one defect, rather than a second gate later for the executable half — the two-opinions shape this repo keeps paying for. AI-hard-to-get-wrong: a broken spelling in a workflow is exactly the kind an agent copies from CI into a local command. Startup scope: the widening is the same walker and extractor, zero extra code, and it was measured at +2 judged / +0 violations rather than assumed cheap."
        },
        {
          "question": "⚠️ NOT MEASURED, and it decides whether one of this gate's own clearances is teaching something false: does `pnpm dev -- --fresh -p PORT` actually reach oclif? Layer 1 is settled (pnpm forwards the separator, measured above), so the open half is purely what oclif does with a leading `--`. AGENTS.md teaches that spelling in three places and this gate deliberately clears all three, correctly for ITS rule (`dev` is not vitest-backed).",
          "options": [
            "Filed as its own card with the reproduction and the control pair — done, #11669, unassigned, `finding` label, no pm:queue.",
            "Measure it inside this card instead: build packages/cli and run a strict oclif command with a bogus flag on both sides of the separator."
          ],
          "recommendation": "The filed card. The 2026-08-21 ruling already asked for it as a separate card with a real measurement and it was never filed; I did not measure it because packages/cli was unbuilt here and booting a dev server is disproportionate to this card, and a guess dressed as a finding would be worse than the open question. Worth the PM's attention because the rule it protects — 'never two backends on port 3000' — may be being enforced by a flag that never arrives, on a box running several agents at once."
        }
      ],
      "out_of_scope_findings": [
        "filed as #11669: unverified — does `pnpm dev -- --fresh -p PORT` reach oclif, or does the separator swallow the flags the port rule depends on? Unassigned, `finding` label, carries the measured pnpm/npx/npm/turbo layer-1 results and the exact control pair the measurement needs. Searched open issues first (341 open, keyword + semantic): no existing card; #11317 is a different oclif issue."
      ]
    }
    

    Generated by Claude Code

  10. removed their assignment
    on Aug 24, 2026
  11. os-steve commented on Aug 24, 2026

    @os-steve
    Collaborator

    Second route landed — PR #11667. ⛔ This card stays OPEN.

    Part of, not Fixes. pm:dispatched stripped by hand, back to pm:queue, unassigned.

    Verified by content on origin/main, landing commit located by (#11667) in its subject:

    6f75c2f62  ci(gates): refuse a bare `--` between pnpm and vitest in committed commands (#11667)
    
    scripts/check-agent-test-spelling.mjs            ABSENT -> PRESENT
      ROOT_DIR_WATCH_HINTS in that file                   0 -> 6
    package.json  check:agent-test-spelling               0 -> 1
    CTRL check-pnpm-filter-targets ROOT_DIR_WATCH_HINTS   6 =  6   (non-zero both sides)
    

    ⭐ And the declaration is live, not merely parsed — run against origin/main after landing:

    $ node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack scripts/check-role-word.mjs
      - pnpm check:agent-test-spelling  [lint.yml]  matched via scripts/check-role-word.mjs ⇢ gate source 'scripts/**'
    

    A scripts/** card now derives this gate. That is the assertion the whole bare-root detour existed to make true, confirmed on main rather than on a branch.

    Where this card now stands — two of three routes done

    The report named three homes: "a documented spelling, a wrapper, or a lint on agent-facing instructions."

    What the lint actually established, beyond going green

    The dispatch framed the risk as "a gate green over an empty population." That was the right worry aimed at the wrong number. Violations are 0 — but the gate judges 6 separators, and 5 of those 6 are forms a naive bare--- rule would have RED, every one correct as written:

    ci.yml:1122        turbo dogfood shard      (turbo strips the separator — measured)
    AGENTS.md:126/133/388   pnpm dev -- --fresh -p …    ← the spelling #11425 landed
    check-examples-live-imports.mjs:176   prose in a comment
    

    Three of the five are the very instruction PR #11425 added to this card. A naive gate would have red-flagged its own documentation fix. So the verdict line reports the judged population, which is what makes a zero mean something — and the vitest binding is derived from 80 workspace manifests rather than hand-carved, so dev and dev:crm are safe by measurement, not by an exception someone has to remember.

    Two measurements worth keeping on this card, since they are the mechanism it was filed about:

    • npx forwards the separator too — not previously measured here. npx vitest run -- FILE is broken identically, so the launcher set is {pnpm, pnpx, npx} and {npm, turbo} are the measured strippers.
    • The defect in one control pair: same binary, same root, same flag, only the separator moves — vitest list --root EMPTY --bogus errors and exits 1; vitest list --root EMPTY -- --bogus prints nothing and exits 0.

    The old objection reproduced, and was answered structurally

    A corpus gate here was rejected once on the grounds that "it would fire on the counter-example the corrected instruction MUST contain — the gate would forbid the fix." That objection is correct and it reproduced in this PR's own first draft: the lint.yml comment wiring the gate spelled the broken command literally and the gate red on it at its own wiring commit. Answered with a declared COUNTER_EXAMPLE_FILES list (each entry carrying a required reason, pinned small) — empty today, because #11425's refusal is prose rather than a quoted bad command. An escape hatch that is built, pinned, and unused beats both having none and pre-filling it.

    For whoever takes the wrapper

    The remaining route is the only one that makes this mechanical rather than detected-after-the-fact: today an agent that writes the bad spelling in an ad-hoc shell command is caught by nothing — the lint judges committed text. Whether a wrapper earns that is a real question; it is not answered by this PR.


    Generated by Claude Code

  12. os-support-ai commented on Aug 25, 2026

    @os-support-ai
    Collaborator

    Maintainer ruling — option A: stop at two routes; the wrapper is NOT PLANNED

    Source: maintainer, 2026-08-25, live PM chat (decision-inbox batch 1 review, session session_01KWRU3s15AJz7PGW7a79f8eb.replace — session session_01KWRU3s15AJz7PGW7a7wdCh), verbatim: 「接受你的建议,继续」 — accepting the presented recommendation for this card: A — close here; no wrapper.

    What this settles:

    • The two landed routes stand as the complete remedy: the documented spelling in the instruction files (PR docs(agents): show the targeted-vitest spelling and refuse the bare separator #11425, human-merged) and the CI gate refusing the bare--- spelling in committed text (PR ci(gates): refuse a bare -- between pnpm and vitest in committed commands #11667, live on main, violating population 0).
    • The third route (a wrapper making the bad spelling mechanically impossible in ad-hoc shell commands) is judged not planned: it is a standing tool with a permanent maintenance obligation, serving a tail risk the two landed defenses already narrow, and its own adoption is unmeasurable. Reopening is free if the same swallow reproduces in live sessions despite the landed instruction text — that evidence would be the new card's premise.

    Closing completed (the card's mechanism was measured, documented, and gated; the declined route is recorded above). needs-user-decision off in the same write.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions