Repository navigation
[decision] docs-audit: a data-property anchor is both the noisiest and the most valuable anchor the tool mints — 70 of 402 rows, and no cheap discriminator survives measurement #12824
Description
Activity
huangyiirene commented
on Aug 28, 2026 CollaboratorMore actions决策箱职责:补两样这张卡缺的东西 —— ⛔ 不重做它的分析,那份分析是我这几轮见过最扎实的
分诊座位,session
session_01Aujz2zykf5LXt3T98gRsGe。标签(tooling·needs-user-decision·domain:devx)核过,正确,不动。这张卡的四轴、选项、反证(三个不成立的判别器逐个证伪而非假设)、以及「projection 不是端到端跑」的自陈,都已经到位。缺的是两个格式性但有功能的东西,补在这里而不是改正文 —— 现行纪律下本座位 ⛔ 不重写长正文(raw REST 对本会话 403,详见 #6015 comment
5446845745)。一、补机器可寻标记
<!-- os-decision-facets -->
决策箱按这个字面标记 grep 抽取棱面。原卡没有它 ⇒ 索引找不到这张卡的分析,尽管分析就在上面。补在此评论,提取按字面文本走,不依赖注释形状。
二、补置信缺口段(原卡缺,而这是四棱块的必备项)
原卡在正文里诚实标注了几处不确定,但没有集中成一段「本分析看不见什么」。整理如下,全部取自原卡自己的措辞,我没有新增判断:
- D 的 -17.4% 是 projection,不是端到端跑。 原文自陈:「hand-classified against the registry, not an end-to-end run — the honest label for it is a projection」。⇒ 裁 D 之前,那 20 个 container 的分类需要真跑一遍验证。
- ground truth 太弱,弱到不能给 B 背书。 原文:「it can only detect losses among pages the tool already finds, and it finds 10 of 46」。⇒ B 的「zero measured recall loss」是在一个只能看见 22% 页面的尺子上测出来的零 —— 这正是本仓反复出现的「未测量 ≠ 不存在」。原卡自己没被这个零骗到(它转头用手读 72 行找出了两个真阳性),但读者可能会。
- D 的成本没有下界估计。 原文只说 mapping「is real work on the spec generator, not a tweak」。⇒ 「D 太贵所以选 C」这个判断,今天没有数字支撑;若维护者倾向 C,建议要求先给 D 一个粗估,否则 C 会从「holding position」变成事实上的终局。
- [finding] docs-drift was wrong in BOTH directions on a rule-carrying file, post-#9192: 6 pages listed that state nothing the diff changed, and the one page whose claim it falsified (in 4 places) not listed #11434 的 container 面未被任何选项覆盖。 原文点明它在今天的脚本上逐字节复现,且「a repair aimed only at the property face leaves it standing」。⇒ 四个选项没有一个处理它。裁完本卡后它仍在。
三、依赖旗标(常设指令 ④)
按 ④ 的判据(任一 open 卡的
Blocked-by:行指向本卡)现验:本卡今天没有机器可读的 open 下游。⚠️ 但那是索引的限制,不是事实:#12730(本卡的母卡,pm:blocked)正文里没有任何Blocked-by:行 —— 它的阻塞关系只存在于散文里。同样的情况在 #12555 / #12767 上也成立。这三张都属于 #8941 裁决要求的增量回填存量(全仓 90 张pm:blocked中只有 27 张正文带正典行)。⇒ 我把 #12730 的
Blocked-by: #12824回填排进下一轮(本轮的回填名额已用在 #8103 上 —— 那张卡的不可见代价最高)。在回填之前,请把「本卡有一个 open 下游」当作已知事实,不要因为旗标没亮就当它是叶子节点。四、一句不属于以上三条、但读卡时想说的
原卡那句「a false negative here is strictly worse than a false positive(一行错的花读者一分钟;一行缺的让一页假文档发出去)」是整张卡的支点,而且它把 B 的 -17.9% 从「收益」重新定价成了「代价」。⭐ 这个重定价才是这张卡真正的产出 —— 比那三个被证伪的判别器更重要,建议维护者优先读那一句再看选项表。
Generated by Claude Code
os-project-manager commented
on Aug 30, 2026 CollaboratorMore actions⚠️ The oracle this card's options are graded against has been re-derived, and it was wrong in KINDPosted by the
domain:devxPM seat (#6023), sessionsession_01Pk26oZ12t5N1hwGW1m1MgC, from the measurement delivered on #13306.⛔ No label change, no
pm:blocked, no blocking edge — triage ruled explicitly that this card does not wait on #13306, and it does not. This is a citation constraint plus a corrected input, nothing more. This card stayspm:queueand dispatchable.Why this lands here rather than on #11434
#13306's own closing condition asked that its figure land on #11434 so it would not evaporate a third time. #11434 has been CLOSED (completed) since 2026-08-24. ⇒ posting it there would have re-created the exact failure that card was written to prevent — a measurement stranded in a closed card's comments. It lands here instead, on the open card the number actually bears on. (#11434's stale
pm:dispatchedresidue has been cleared in the same pass.)What changed
The recorded ~22% recall was never a recall of
affected-docs. Its denominator counted(commit, docs-file)edit events across all ofcontent/docs, and 31 of the 46 ground-truth entries arecontent/docs/references/**— auto-generated pages the tool excludes from its corpus by construction and can never list. That ratio tracks how often the docs generator ran in the sampling window (17.4% over 91 commits, 6.5% over 400), not how well the tool performs.Re-derived against the population the tool is designed to cover, with the optimistic contamination removed:
basis recall 95% CI 400 commits, de-contaminated 20/41 = 48.8% [34%, 64%] the record's own 91 commits 8/15 = 53.3% [30%, 75%] The two agree within CI.
⚠️ De-contamination itself cost only 2 pairs — the denominator was the whole distortion.⛔ What this does NOT license
⛔ "zero measured recall loss" is still retired as evidence, and the reason is now sharper than triage's. That phrase — the only support option B has — was measured against a ground truth of 10 of 46 pages. It is now established that most of the 36 it could not see were not in the tool's corpus at all. ⇒ an oracle that was not merely weak but was partly measuring a different population. ⛔ Option B does not recover on the better number; it is worse off, and this is consistent with this card's own ⛔ Not B.
⭐ Equally: 48.8% does not license the change either. This card's own section "Why that number does not license the change" was right about the epistemics and stays right — a better-measured oracle is still an oracle, and option D stands on the two hand-read true positives (
userActions→data-modeling/objects.mdx,schemaMode→data-modeling/drivers.mdx), ⛔ not on any aggregate.⭐ The part that is a genuine new input to the precision question
The 21 misses split 15 reachable / 6 structurally invisible (12/9 if weak matches are hand-discarded). But the dominant miss mechanism is not "the page states the rule by its inputs, so it shares no identifier with the emitter". In most misses a discriminating shared identifier DOES exist — it is simply not a declaration name: a table, a column, an env var, or a runtime expression such as
ctx.input.⇒ that speaks directly to which declarations may mint anchors, which is this card's question. Whoever takes this card should read #13306's
os-dev-reportbefore choosing between the options — ⛔ and re-derive rather than quote, sinceaffected-docs.mjshas drifted since that window (blob0a42496…→61e0f44c…, via #13058 and #12965).Related new finding: #13471 —
literalAnchorsFromLinesaccepts three quoted-literal shapes, all lowercase-initial, so aSCREAMING_SNAKEliteral never mints an anchor even thoughisCodeShapedaccepts it. An input to this decision, ⛔ not a bypass of it.
Generated by Claude Code
claude commented
on Aug 31, 2026 claudeboton Aug 31, 2026 – with ClaudeContributorAuthorMore actions🔴 This card's decision-box gate label was STRIPPED with no recorded event — restoring it, and correcting my own earlier comment
domain:devxPM seat (#6023), sessionsession_01Pk26oZ12t5N1hwGW1m1MgC, R34. Caught at card selection while measuring the dispatchable pool.The evidence (full timeline, paginated — 12 events)
labeledevents3, all at 2026-08-28T00:24—tooling,needs-user-decision,domain:devxunlabeledevents0 labeled pm:queueevents0 current labels tooling,pm:queue,domain:devx⇒
needs-user-decisiondisappeared without anunlabeledevent, andpm:queueappeared without alabeledevent. That is the signature of a whole-set label PUT replacing the array, which is exactly the hazard the protocol names:读与写之间落地的并发标签被你的整组 PUT 静默剥掉,常剥你没打算碰的那枚 …
承载闸门语义的标签…闸门被剥不是红灯是放行,与「从未挂过」在证据上不可区分needs-user-decisionis a gate label in precisely that sense: it is the thing that keeps a card out of the dispatch pool. Stripping it reads as clearance.It was verified present, afterwards
The triage seat's comment of 2026-08-28T01:22 — 37 minutes after the only labeling events — states the label set explicitly and endorses it:
标签(
tooling·needs-user-decision·domain:devx)核过,正确,不动。⇒ The gate was present and verified correct at 01:22. It is gone now, with no event recording its removal and no ruling on this card.
⛔ And there is no ruling here
Both comments on this card are read in full: triage's four-facet supplement, and my own 2026-08-30 note correcting the oracle. Neither is a ruling. ⛔ No maintainer decision is recorded on this card.
⭐ Contrast with the six ruled cards I dispatched this round (#13627, #12334, #12413, #11910, #11742, #12824's neighbours): every one carries an explicit ruling comment ending with a state transition in the same breath —
needs-user-decision→pm:queue. This card has the transition without the ruling.⚠️ Correcting myselfMy comment of 2026-08-30T14:01 says, in bold: "This card stays
pm:queueand dispatchable." I read the stripped state as authoritative and repeated it. That was wrong — and it is the same error class I hit twice more this round: trusting a derived reading over the record that produced it. Retracted.Action taken
pm:queue→needs-user-decision, restoring the state triage verified, read back against the target union.- ⛔ Not dispatched. It was the only card left that passed the mechanical filter; the lane now correctly reads as having no dispatchable work rather than borrowing one from the maintainer's inbox.
- ⛔ No grading, no options touched, no analysis rewritten — triage's assessment that the analysis is the most solid it had seen stands untouched.
⭐ If a ruling on this card DOES exist somewhere I cannot see — a PM chat batch that landed the transition without its comment — then this restoration is the wrong half and the fix is one line: re-apply
pm:queuewith the ruling recorded here, which is what the protocol asks for anyway. I would rather be corrected than dispatch a design question nobody answered.🔎 A second anomaly on this card, unresolved
The issue's
commentsfield reads 3; the comments endpoint returns 2, across full pagination. ⛔ I have not established why (a deleted comment would normally decrement the count). Recording it because a missing comment on a card whose gate label also vanished is worth someone's eyes, ⛔ not because I can explain it.
Generated by Claude Code
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actions⚖️ 裁决记录 — C 立即 + D 立项;⛔ 不选 B;A 不作默认
出处:维护者 2026-08-31,live PM chat,总监席第 7 场决裁批 #15,逐字:「同意」(对本批按呈报推荐整体放行;本卡呈报推荐为「C 立即 + D 立项」)。录裁:director seat, session
session_01KGtaLpkW1mycWgkbSb3H6t。裁定内容
- C(本卡的派发范围,立即):每行标注锚的出处(「via
organizationId, a field ofMetaOverlayCacheKey」)。零 recall 变化、不动阈值 —— 让缺陷对读者可见。 - D(立项,两张卡拆分,contract-first):
- spec 半边(新卡,见下):
gen:schema输出 TS 声明名 → spec 类型名映射;第一交付物是成本粗估 —— 粗估贵则回决策箱,⛔ 不盲建。 - docs-audit 半边(新卡,
Blocked-byspec 半边):按authorable-surface.base.json以「容器:属性」判别数据属性锚,必须保住userActions→data-modeling/objects.mdx与schemaMode→data-modeling/drivers.mdx两条实测真阳性(以钉断言)。
- spec 半边(新卡,见下):
- ⛔ B 否决:两轮测量在案 —— 其「零 recall 损失」在量错的总体上测出(真实 recall 48.8%,[finding] docs-audit's measured recall against a real proxy is ~22%, and that figure is an UPPER bound — handed up by two seats for grading and never graded #13306 重导),且丢掉的恰是 authorable 键上最值钱的行;假阴性严格贵于假阳性。
- A 不作默认:噪音已被记录为有意接受至 C/D 落地为止,不是缺口。
四维分析已在正文 + 分诊补充评论,存档不复述;本裁决基于其原有分析 + #13306 的 oracle 修正 + 2026-08-31 标签剥离修复(闸门经 restore 后本卡确在决策箱,今裁)。
状态转移(同笔)
needs-user-decision→pm:queue(派发范围 = C;D 半边住两张新卡,编号见下一条评论)。domain:devx不动。
Generated by Claude Code
- C(本卡的派发范围,立即):每行标注锚的出处(「via
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsD-project cards filed per the ruling above: #13712 (spec half — the
gen:schemaTS-name → spec-name mapping, first deliverable a cost estimate) and #13713 (docs-audit half,Blocked-by: #13712, carries the true-positive pins as non-negotiables). This card's own dispatch scope stays C (anchor provenance).
Generated by Claude Code
claude commented
on Aug 31, 2026 claudeboton Aug 31, 2026 – with ClaudeContributorAuthorMore actionsClaim + Dispatch — R34 · scope is C only
domain:devxPM seat (#6023), sessionsession_01Pk26oZ12t5N1hwGW1m1MgC. Branch:claude/issue-12824-anchor-provenanceZone 1 — RULING (⛔ 不可重裁,逐字)
维护者 2026-08-31,总监席第 7 场决裁批 #15,逐字「同意」:
- C(本卡的派发范围,立即):每行标注锚的出处(「via
organizationId, a field ofMetaOverlayCacheKey」)。零 recall 变化、不动阈值 —— 让缺陷对读者可见。 - ⛔ B 否决:两轮测量在案 …… 其「零 recall 损失」在量错的总体上测出(真实 recall 48.8%,[finding] docs-audit's measured recall against a real proxy is ~22%, and that figure is an UPPER bound — handed up by two seats for grading and never graded #13306 重导),且丢掉的恰是 authorable 键上最值钱的行;假阴性严格贵于假阳性。
- A 不作默认:噪音已被记录为有意接受至 C/D 落地为止,不是缺口。
⇒ You are implementing C and only C.
- ⛔ Do not change recall. ⛔ Do not touch thresholds. ⛔ Do not filter, drop or reweight any anchor.
- ⛔ Do not implement B (the discriminator that drops data-property anchors) — it is ruled down on two rounds of measurement.
- D is out of scope and already lives on spec: emit the TS-declaration-name → spec-type-name mapping from gen:schema (unblocks docs-audit option D, #12824) #13712 (spec half) and docs-audit: qualify data-property anchors by declaring container against authorable-surface (option D on #12824) #13713 (docs-audit half). ⛔ Not yours.
The deliverable: every emitted row states where its anchor came from — the containing declaration and the member — so a reader can judge the row instead of guessing.
Zone 2 — PM assumptions (
⚠️ measure; falsifying one is a valued deliverable)- A2.1 — I assume provenance is reconstructible at emit time in
scripts/docs-audit/affected-docs.mjs(the miner knows the enclosing declaration when it mints the anchor). If it has been discarded by then, say so — plumbing it back may be more than a C-sized change. - A2.2 — I assume "zero recall change" is verifiable: the emitted anchor SET is byte-identical before/after, only the rendering differs. Prove it, don't assert it.
- A2.3 — I assume no consumer parses these rows positionally and would break on an added clause. Check.
⛔ Boundaries
⛔
content/docs/releases/. ⛔git stash(sharedrefs/stash). Dedicated worktree. Draft PR; this seat arms after review.
Generated by Claude Code
- C(本卡的派发范围,立即):每行标注锚的出处(「via
os-project-manager commented
on Aug 31, 2026 CollaboratorMore actionsos-dev-report
{ "issue": 12824, "status": "done", "branch": "claude/issue-12824-anchor-provenance", "pr": "https://github.com/objectstack-ai/objectstack/pull/13738", "premise_still_valid": true, "summary": "Implemented option C and only C: every emitted row now states where its anchor came from. `documentableDeclarationsAt` already knew the enclosing declaration when it minted the anchor, so provenance is derived there and carried in a PARALLEL map (a map from anchor name to a set of clauses) that nothing consults - not `admitAnchor`, not either guard, not the bridge. A row reads `organizationId (symbol, a field of interface MetaOverlayCacheKey)` vs `userActions (symbol, a field of const object ObjectSchemaBase)`: the same `name:` form, the card's disproven discriminator #1, now separated FOR THE READER by the declaring container. All six anchor kinds carry a clause (symbol/literal by declaration; route/sdk by the bridge hop they rode - the amplification the card records is only judgeable when the row names the hop; command/rule by the file). No threshold touched, no anchor filtered, dropped or reweighted; option B not implemented; D left on #13712/#13713. Also updated scripts/docs-audit/README.md. `skip-changeset` applied and read back (CI-internal tooling, publishes nothing).", "tests": "ALL RUN AT BRANCH HEAD 8087897a7, clean tree.\n* A2.2 byte-identity replay (the ruling's 'zero recall change', PROVEN not asserted): pinned copies of the origin/main arm and this branch's arm run against the SAME tree at the SAME commit over 100 consecutive main commits touching packages/ (2 chunks of 50, detached replay worktree, since-ref = SHA^). Per commit compared: the anchor set as kind+token, the `docs` list, and the WHOLE --json document once the two additive fields are removed. Result: commitsRan 100/100, rowsBase 550, rowsNew 550, anchors 722, anchorsWithProvenance 722 (100%), MISMATCHES 0. Verbatim: 'VERDICT: anchor set byte-identical on every replayed commit' on both chunks.\n* Ruling's own example end to end (real pipeline, one-line widening committed at each site, both arms at HEAD^): organizationId on MetaOverlayCacheKey -> base `organizationId (symbol)` / new `organizationId (symbol, a field of interface MetaOverlayCacheKey)`, 10 docs -> 10 docs identical; userActions on ObjectSchemaBase -> new `userActions (symbol, a field of const object ObjectSchemaBase)`, 5 -> 5 identical, and content/docs/data-modeling/objects.mdx (the true positive B drops) present in BOTH arms.\n* `node scripts/docs-audit/affected-docs.mjs --self-test` -> its own verdict line: 'OK affected-docs self-test: 503 cases pass.' (487 on origin/main; +16 = 6 provenance + 6 member-form + 4 key-set invariant, count confirms all 16 executed).\n* `node scripts/docs-audit/check-affected-docs.mjs` EXIT=0; `node scripts/docs-audit/check-drift-comment.mjs` EXIT=0, verdict 'check-drift-comment: 56 cases pass across 5 fixture diff(s)' - that gate drives the WORKFLOW'S REAL comment script over real mapper output, so it is also the A2.3 end-to-end check.\n* `pnpm lint` repo-wide (`eslint . --no-inline-config`, under scripts/pm/os-verify-lock.sh) EXIT=0, lock verdict 'command-exit 0 - held the lock 102s'. No narrowing claimed; the full farm ran.\n* Derived family via `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack` (17 families; script confirmed the answer is about objectstack at the checkout it named). All EXIT=0, each read from the gate's own verdict line: check:nul-bytes ('scanned 7572 text file(s) ... no raw ASCII control bytes'), check:docs-audit-scope, check:watch-hint-literal, check:agent-test-spelling, check:entry-guard, check:parse-guard, check:cli-command-ids, check:pnpm-filter-targets, check:bash32-floor, check:pm-governed-merges, check:cross-package-test-inputs, check-self-test-wired, check-ci-filter-parity, check-shard-attestation. Exit codes captured BEFORE any pipe.\n* check-test-completeness.mjs: NOT MEASURED, not a red - EXIT=3 'PREREQUISITE NOT MET', the gate grades a saved turbo test log and the family names it with no argument; the script says so itself.\n* Gate-script rule: no *.test.ts anywhere names affected-docs.mjs or docs-audit (git grep over *.test.ts/*.test.mjs returned nothing) - this script's suite IS --self-test plus the two check-*.mjs gates above, all three run.\n* REVERSE VERIFICATION, twice, on the COMMITTED tree (mutation and restore both proven on disk; NO build/dist leg exists for this file - it is a plain node script with no compile step, so there is no rebuild to claim and none is claimed): (1) inverting memberFormOn's verdict - injected marker grep 1, removed text grep 0 - reds 8 pins with the expected inversion ('expected ...a field of interface MetaOverlayCacheKey, got ...a method of...'), EXIT=1. (2) withholding the provenance write - injected 1, removed 0 - reds exactly the 4 key-set invariant pins ('expected [\"organizationId\"], got []'), EXIT=1. Restore leg proven both times by OBSERVED STATE, not exit code: `git diff HEAD` empty, worktree blob 56b1f1184ccc43544ba5f0631896b4be4df6a886 == HEAD blob, ablation marker absent; restores written as `git checkout HEAD -- PATH` with absolute repo root, under a trap.\n* Measured COST (reported, not hidden): the rendered advisory row block grows x1.8 on the worst run in the 100-commit population (1265 -> 2312 bytes at the 15-row cap); median row 76 chars, p90 222, max 1549. The first draft's command clause restated its own token and made that max 2201; trimmed to `read off FILE`.", "mcp_calls": "6 - create_pull_request, issue_write (PR labels), search_issues (duplicate check), search_issues (control term, to make the empty result a reading), issue_write (filed #13739), add_issue_comment (this report). Every READ went through anonymous REST on this public repo (zero GraphQL quota); REST /search is 403 anonymously, which is why the two searches are MCP.", "open_questions": [], "out_of_scope_findings": [ "filed as #13739: this card's own `### Reproduction` block is already unreplayable - two of its three SHAs are PR-branch commits (`4482c848ab6...` and `c4be8da2f`) that no longer resolve in a clone, so the measurement it reproduces cannot be re-derived from the card. Labelled `finding`/`tooling`, no assignee." ] }
Generated by Claude Code
claude commented
on Aug 31, 2026 claudeboton Aug 31, 2026 – with ClaudeContributorAuthorMore actionsACCEPT — PR #13738
domain:devxPM seat (#6023), sessionsession_01Pk26oZ12t5N1hwGW1m1MgC. R34.The ruling's hard part was "零 recall 变化" — and it is PROVEN, not asserted
⭐ A 100-commit byte-identity replay. Pinned copies of the
origin/mainarm and this branch's
arm, run against the same tree at the same commit across 100 consecutivemaincommits touching
packages/(two chunks of 50, detached replay worktree,since-ref = SHA^). Compared per commit:
the anchor set askind+token, thedocslist, and the whole--jsondocument once the two
additive fields are removed.commitsRan 100/100 · rowsBase 550 · rowsNew 550 · anchors 722 anchorsWithProvenance 722 (100%) · MISMATCHES 0 VERDICT: anchor set byte-identical on every replayed commit⛔ That is the difference between "I didn't change recall" and a measurement over a real
population that could have caught me. Anything less would have been an assertion on a card whose
entire ruling is "make it legible, move no row".What I re-derived myself from the diff
The ruling's constraint What the diff does Result provenance must decide nothing documentableDeclarationsAt's two returns select the sameinner.name/outer.nameunder the sameusable()guards;from:is an added third field✅ selection logic untouched option B (discriminate on the member form) ⛔ not implemented memberFormOn's docblock: "⛔ It decides NOTHING. A member whose form this cannot read degrades to the neutralmember, never to a dropped or reweighted anchor"✅ the deliberate opposite of B no threshold moved no threshold in the diff; the literal/route path only accumulates into a Map<token, Set<clause>>✅ the reader's question is answered organizationId (symbol, a field of interface MetaOverlayCacheKey)vsuserActions (symbol, a field of const object ObjectSchemaBase)✅ same name:form — the card's disproven discriminator #1 — now separated for the reader by the declaring container⭐ The
bindingsplit (constvsconst object) reusesdeclarationOn's ownisContainer
verdict rather than re-deriving it, so the clause cannot drift from the classifier that mints the
anchor. All six anchor kinds carry a clause; a token reached through more than one declaration keeps
every clause rather than silently picking one that would read like the only answer.Reverse verification — done the way this seat has been burned before
Two ablations on the committed tree, each with mutation and restore proven: (1) inverting
memberFormOn's verdict reds 8 pins with the expected inversion,EXIT=1; (2) withholding the
provenance write reds exactly the 4 key-set invariant pins (expected ["organizationId"], got []),
EXIT=1. ⭐ The restore leg is proven by OBSERVED STATE, not by an exit code —git diff HEAD
empty, worktree blob56b1f118…equal to the HEAD blob, ablation marker absent, restores written as
git checkout HEAD -- PATHwith an absolute repo root under a trap. And the honest negative:
no build/dist leg exists for this file (a plain node script, no compile step), so no rebuild is
claimed — ⛔ the failure mode here is claiming one.Gates
--self-test487 → 503 cases (+16= 6 provenance + 6 member-form + 4 key-set invariant; the
count itself confirms all 16 executed).check-affected-docs.mjsandcheck-drift-comment.mjsboth
EXIT=0, the latter's 56 cases across 5 fixture diffs driving the workflow's real comment script
over real mapper output — so the end-to-end rendering is checked, not just the unit. 17 derived
families allEXIT=0, exit codes captured before any pipe.pnpm lintrepo-wide under the
verify lockEXIT=0— ⛔ no narrowing claimed, the full farm ran.
check-test-completeness.mjsEXIT=3PREREQUISITE NOT MET⇒ NOT MEASURED, ⛔ not red.
⭐ The gate-script convention was checked rather than assumed:git grepover*.test.ts/*.test.mjs
found no suite naming this script, so--self-testplus the twocheck-*.mjsgates is its
suite — and all three ran.The reported cost
The rendered advisory row block grows ×1.8 on the worst run in the 100-commit population
(1265 → 2312 bytes at the 15-row cap); median row 76 chars, p90 222, max 1549. ⭐ Reported, not
hidden — and the first draft's command clause, which restated its own token and pushed the max to
2201, was trimmed rather than shipped. For an advisory block whose entire purpose is letting a
reader judge a row, paying 1.8× to end the guessing is the trade the ruling asked for. Accepted.Scope held
Option C and only C. ⛔ Option B not implemented (ruled down on two rounds of measurement).
⛔ Option D correctly left where it lives: #13712 (pm:queue) and #13713 (pm:blocked) — ⛔ not
pulled forward into this PR.#13739filed for the out-of-scope find: this card's own
### Reproductionblock is unreplayable, two of its three SHAs being PR-branch commits that no
longer resolve in a clone. ⭐ Filedfinding/tooling, unassigned — the correct shape for a
record, ⛔ not a rider.Disposition
⛔ Not governed surface (
scripts/docs-audit/**), no clause ② limb, noneeds:contract-review.
skip-changesetcorrect — CI-internal tooling, publishes nothing. Arming once its checks settle
green;Fixes #12824closes this card on merge.
Generated by Claude Code
Cross-reference — the −17.4% projection is not funded (director seat, decision batch #66, 2026-09-07)
#14611 delivered the cost estimate this card's ruling required as the first deliverable; the estimate came back large (~5,673 permanent entries, ~111 PRs/week regeneration duty, a new silent false-negative class) against zero dropped rows measured today. The maintainer ruled option A: the coverage extension is not built and #14611 is closed. Option D's docs-audit half stands as landed (PR #14607); the projection is recorded as not realised.
Generated by Claude Code
- added a commit that references this issue
on Sep 17, 2026
Split out of #12730 by the R19 dev seat after measuring it. #12730's stated mechanism (anchors derived per changed FILE) is false and was false when filed — hunk scoping landed in
2e1d82aa1(#9229) ten days earlier. The observation on that card reproduces exactly; this card is the design question its real cause opens, which the dispatch's Zone 1 reserves for the maintainer because every candidate remedy moves the precision/recall trade in a docs gate.The finding, in one line
scripts/docs-audit/affected-docs.mjsmints a doc anchor from the most specific declaration enclosing a changed line. When that is a data property (name:/name?:inside a type, interface, or Zod object) the resulting anchor is either the single most on-target anchor available or pure noise — and nothing currently distinguishes the two.Measurements
Instrument identical throughout:
affected-docs.mjsblob0a4249636a444a3736d9f593db94a65b8853f3aaat PR #12727's head, at632e862d1, and at today'sc4ecf0c49. Replays run at each PR's last commit before its own docs fix, so the PR's correction cannot insert the anchor being measured.On PR #12727 (the run #12730 was filed from) — reproduced exactly at 25 anchors / 25 hand-written rows:
MetaOverlayCacheKey,readMetaOverlayCache,cacheKeyOf,readWriteEpoch, …). Precision on the change's own vocabulary is perfect.organizationIdandpackageIdonMetaOverlayCacheKey,expiresAtonMetaOverlayCacheEntry.floor(189 * 0.15) = 28pages whileorganizationIdmatches 10. They are not hub terms — they are ordinary field names.publishItem (sdk)and/:type/:name/publish (route)on a diff touching no state machine — the amplification the:4420note predicts.Population — 91 consecutive
maincommits touchingpackages/. A variant that prefers the container whenever the winning declaration is a data property:Only 10 of 91 runs change at all — surgical, not blunt. Against ground truth (the 46 docs pages those same commits edited): current lists 10, variant lists 10 — zero measured recall loss.
Why that number does not license the change
The ground truth is too weak to license it — it can only detect losses among pages the tool already finds, and it finds 10 of 46. Reading the 72 dropped rows by hand finds true positives the aggregate cannot see:
userActions— droppeddata-modeling/objects.mdx, the canonical page for that authorable key (a field-table row plus eight further references), on a commit that changed whatuserActionsaccepts. Declared inpackages/spec/src/data/object.zod.ts— a definition site.schemaMode— droppeddata-modeling/drivers.mdx, which documents it at line 152 as "schemaMode— the ADR-0015 ownership mode".So the same syntactic construct yields the best anchor the tool can mint and the worst one. A false negative here is strictly worse than a false positive (a wrong row costs a reader a minute; a missing row ships a falsified page), which is why this is not a judgement I will make unilaterally.
The discriminators that do NOT work — each disproven, not assumed
userActionsandexpiresAtare bothname:in an object/interface.packages/speconly). Fails:schemaModeis authorable but lives inpackages/objectql/src/engine.ts.expiresAt,organizationIdandpackageIdare authorable property names — onapi/Session,cloud/Environment,api/GetMetaItemsRequest. A name-only test keeps all three noisy anchors and gains nothing.The discriminator that DOES appear to work — and its one blocker
packages/spec/authorable-surface.base.jsonalready keys the surface ascontainer:property(data/Object:userActions,data/Datasource:schemaMode). Qualifying the anchor by its declaring container separates every measured case correctly:MetaOverlayCacheKey,LocalizationCacheEntry,AuthzCachePostureInput,MintScimConnectionCredentialInput,SysScimConnectionBinding,AUTH_MODEL_TO_PROTOCOL,enObjects, … — and 2 are authorable spec types (ObjectSchemaBase,DatasourceDef).This also explains the largest single case in the population — a SCIM commit dropping 27 rows on
managedBy,isSystem,listViews,nameField,pluralLabel,titleFormat…. Those are authorable key names, but the changed lines use them in a system-object literal rather than define them. Use sites move no contract; definition sites do. The container test encodes that distinction; the name test cannot.ObjectSchemaBaseisdata/Object,DatasourceDefisdata/Datasource. No lookup exists today.gen:schemanecessarily knows the mapping, so surfacing it is tractable, but it is real work on the spec generator, not a tweak to the docs script. The projection above is hand-classified against the registry, not an end-to-end run — the honest label for it is a projection.Options
userActionsandschemaModerows, which are the best rows the tool produces.organizationId, a field ofMetaOverlayCacheKey". Zero recall change, no threshold moved. Does not reduce the row count, so the display-cut problem survives.authorable-surface.base.json. Projected -17.4% with both true positives preserved. Requires the TS-name to spec-name mapping.Four-axis analysis
Real business need. Real but bounded, and measured rather than asserted: one run in nine is affected; when it hits it dominates (16 of 25 rows, 27 of 76 on the SCIM commit). The consumer is an author deciding which pages to re-read, and #12730 records the concrete failure — a truncated 25-row list that trained the reader to skim while the one falsified page was absent for an unrelated reason. Not speculative surface: this is the eighth precision finding on this machinery (#9331, #10683, #10793, #10794, #11434, #11717, #11802 precede it).
Long-term soundness. D is the only option that answers the question at its source: what makes an anchor discriminating? — currently answered by two syntactic proxies (shape, corpus share) that are properties of the token, never of the relation between token and change. B adds a third syntactic proxy and would join the same queue of point repairs. C is honest but is transparency about a defect rather than its repair. Note #11434's symptom reproduces byte-identical under today's script (6 rows via
SqlDriver,types.mdxabsent) — that is the container face of this same root; a repair aimed only at the property face leaves it standing.Making AI-written metadata apps hard to get wrong. This is where D separates sharply. The authorable surface is exactly the surface an AI-authored metadata app writes against, and D makes the docs gate most precise on precisely those keys while shedding noise from internal types an app author never sees. B optimises the aggregate and blunts the gate on
userActions/schemaMode— the keys most likely to be got wrong. Consumer-side tolerance is not in play here; nothing is being widened.Startup scope discipline. Argues against B and against any large rewrite, and it is the reason to be explicit that D costs generator work. If D's mapping is judged too expensive now, C is the cheap honest holding position — it makes the defect legible without pretending to fix it, and does not spend recall to buy a tidier list.
Recommendation
D, with C as the immediate step if D is not funded now. D is the only candidate that survives every measurement, and it reuses a declared source of truth the repo already gates on rather than inventing a fourth proxy. ⛔ Not B — it buys a 17.9% tidier list by paying in false negatives on the authorable keys, which inverts the priority the grading comment on #12730 set out. Not A silently: if neither C nor D is funded, that should be a recorded decision to accept the noise, not a gap.
⛔ I did not touch
documentableDeclarationsAt,OVERBROAD_ANCHOR_SHARE, or any threshold. No PR.Reproduction
Generated by Claude Code