Skip to content

[decision] docs-audit: a data-property anchor is both the noisiest and the most valuable anchor the tool mints — 70 of 402 rows, and no cheap discriminator survives measurement #12824

Description

@claude

Split out of #12730 by the R19 dev seat after measuring it. #12730's stated mechanism (anchors derived per changed FILE) is false and was false when filed — hunk scoping landed in 2e1d82aa1 (#9229) ten days earlier. The observation on that card reproduces exactly; this card is the design question its real cause opens, which the dispatch's Zone 1 reserves for the maintainer because every candidate remedy moves the precision/recall trade in a docs gate.

The finding, in one line

scripts/docs-audit/affected-docs.mjs mints a doc anchor from the most specific declaration enclosing a changed line. When that is a data property (name: / name?: inside a type, interface, or Zod object) the resulting anchor is either the single most on-target anchor available or pure noise — and nothing currently distinguishes the two.

Measurements

Instrument identical throughout: affected-docs.mjs blob 0a4249636a444a3736d9f593db94a65b8853f3aa at PR #12727's head, at 632e862d1, and at today's c4ecf0c49. Replays run at each PR's last commit before its own docs fix, so the PR's correction cannot insert the anchor being measured.

On PR #12727 (the run #12730 was filed from) — reproduced exactly at 25 anchors / 25 hand-written rows:

  • 13 of 25 anchors produced ZERO rows — and those 13 are every name specific to the change (MetaOverlayCacheKey, readMetaOverlayCache, cacheKeyOf, readWriteEpoch, …). Precision on the change's own vocabulary is perfect.
  • 16 of 25 rows came from three data properties of internal cache structs: organizationId and packageId on MetaOverlayCacheKey, expiresAt on MetaOverlayCacheEntry.
  • Both existing guards pass them legitimately: they are code-shaped, and the corpus-share limit is floor(189 * 0.15) = 28 pages while organizationId matches 10. They are not hub terms — they are ordinary field names.
  • Through the bridge they also minted publishItem (sdk) and /:type/:name/publish (route) on a diff touching no state machine — the amplification the :4420 note predicts.

Population — 91 consecutive main commits touching packages/. A variant that prefers the container whenever the winning declaration is a data property:

rows
current 402
variant 330
removed 72 (-17.9%), zero added

Only 10 of 91 runs change at all — surgical, not blunt. Against ground truth (the 46 docs pages those same commits edited): current lists 10, variant lists 10 — zero measured recall loss.

Why that number does not license the change

The ground truth is too weak to license it — it can only detect losses among pages the tool already finds, and it finds 10 of 46. Reading the 72 dropped rows by hand finds true positives the aggregate cannot see:

  • userActions — dropped data-modeling/objects.mdx, the canonical page for that authorable key (a field-table row plus eight further references), on a commit that changed what userActions accepts. Declared in packages/spec/src/data/object.zod.ts — a definition site.
  • schemaMode — dropped data-modeling/drivers.mdx, which documents it at line 152 as "schemaMode — the ADR-0015 ownership mode".

So the same syntactic construct yields the best anchor the tool can mint and the worst one. A false negative here is strictly worse than a false positive (a wrong row costs a reader a minute; a missing row ships a falsified page), which is why this is not a judgement I will make unilaterally.

The discriminators that do NOT work — each disproven, not assumed

  1. Syntactic form (data property vs method). Fails: userActions and expiresAt are both name: in an object/interface.
  2. Declaring package (trust packages/spec only). Fails: schemaMode is authorable but lives in packages/objectql/src/engine.ts.
  3. Property NAME against the authorable-surface registry. Fails, and instructively: expiresAt, organizationId and packageId are authorable property names — on api/Session, cloud/Environment, api/GetMetaItemsRequest. A name-only test keeps all three noisy anchors and gains nothing.

The discriminator that DOES appear to work — and its one blocker

packages/spec/authorable-surface.base.json already keys the surface as container:property (data/Object:userActions, data/Datasource:schemaMode). Qualifying the anchor by its declaring container separates every measured case correctly:

  • Of the 20 distinct containers that minted a dropped data-property anchor, 18 are internal implementation types — MetaOverlayCacheKey, LocalizationCacheEntry, AuthzCachePostureInput, MintScimConnectionCredentialInput, SysScimConnectionBinding, AUTH_MODEL_TO_PROTOCOL, enObjects, … — and 2 are authorable spec types (ObjectSchemaBase, DatasourceDef).
  • Projected: removes 70 of 402 rows (-17.4%) while preserving both demonstrated true positives.

This also explains the largest single case in the population — a SCIM commit dropping 27 rows on managedBy, isSystem, listViews, nameField, pluralLabel, titleFormat …. Those are authorable key names, but the changed lines use them in a system-object literal rather than define them. Use sites move no contract; definition sites do. The container test encodes that distinction; the name test cannot.

⚠️ Blocker, and the reason this is a decision rather than a patch: the TS declaration name is not the spec type name — ObjectSchemaBase is data/Object, DatasourceDef is data/Datasource. No lookup exists today. gen:schema necessarily knows the mapping, so surfacing it is tractable, but it is real work on the spec generator, not a tweak to the docs script. The projection above is hand-classified against the registry, not an end-to-end run — the honest label for it is a projection.

Options

  • A. Status quo. Keep minting data-property anchors. Cost is measured: ~17% of all rows, concentrated so that one run in nine gets badly inflated — and [finding] docs-audit derives anchors per changed FILE, not per changed hunk — on a 20k-line file it named 22 unrelated pages and missed the one the diff actually falsified #12730 shows the inflation lands past the bot's 15-row display cut, so it truncates exactly when signal-to-noise is worst.
  • B. Prefer the container for every data property. -17.9% rows, zero added, zero measured recall loss — but demonstrably drops the userActions and schemaMode rows, which are the best rows the tool produces.
  • C. Publish anchor provenance only. Each row already says which anchor put it there; add which declaration minted the anchor, so a row reads "via organizationId, a field of MetaOverlayCacheKey". Zero recall change, no threshold moved. Does not reduce the row count, so the display-cut problem survives.
  • D. Container-qualified against authorable-surface.base.json. Projected -17.4% with both true positives preserved. Requires the TS-name to spec-name mapping.

Four-axis analysis

Real business need. Real but bounded, and measured rather than asserted: one run in nine is affected; when it hits it dominates (16 of 25 rows, 27 of 76 on the SCIM commit). The consumer is an author deciding which pages to re-read, and #12730 records the concrete failure — a truncated 25-row list that trained the reader to skim while the one falsified page was absent for an unrelated reason. Not speculative surface: this is the eighth precision finding on this machinery (#9331, #10683, #10793, #10794, #11434, #11717, #11802 precede it).

Long-term soundness. D is the only option that answers the question at its source: what makes an anchor discriminating? — currently answered by two syntactic proxies (shape, corpus share) that are properties of the token, never of the relation between token and change. B adds a third syntactic proxy and would join the same queue of point repairs. C is honest but is transparency about a defect rather than its repair. Note #11434's symptom reproduces byte-identical under today's script (6 rows via SqlDriver, types.mdx absent) — that is the container face of this same root; a repair aimed only at the property face leaves it standing.

Making AI-written metadata apps hard to get wrong. This is where D separates sharply. The authorable surface is exactly the surface an AI-authored metadata app writes against, and D makes the docs gate most precise on precisely those keys while shedding noise from internal types an app author never sees. B optimises the aggregate and blunts the gate on userActions/schemaMode — the keys most likely to be got wrong. Consumer-side tolerance is not in play here; nothing is being widened.

Startup scope discipline. Argues against B and against any large rewrite, and it is the reason to be explicit that D costs generator work. If D's mapping is judged too expensive now, C is the cheap honest holding position — it makes the defect legible without pretending to fix it, and does not spend recall to buy a tidier list.

Recommendation

D, with C as the immediate step if D is not funded now. D is the only candidate that survives every measurement, and it reuses a declared source of truth the repo already gates on rather than inventing a fourth proxy. ⛔ Not B — it buys a 17.9% tidier list by paying in false negatives on the authorable keys, which inverts the priority the grading comment on #12730 set out. Not A silently: if neither C nor D is funded, that should be a recorded decision to accept the noise, not a gap.

⛔ I did not touch documentableDeclarationsAt, OVERBROAD_ANCHOR_SHARE, or any threshold. No PR.

Reproduction

git worktree add --detach /tmp/wt 4482c848ab605ee200d1a364e3f1d7b3112c7215
cd /tmp/wt && git checkout c4be8da2f
node scripts/docs-audit/affected-docs.mjs 15bf9e859e56862e6ebe7b5c42404de103362457 --json

Generated by Claude Code

Activity

  1. huangyiirene commented on Aug 28, 2026

    @huangyiirene
    Collaborator

    决策箱职责:补两样这张卡缺的东西 —— ⛔ 不重做它的分析,那份分析是我这几轮见过最扎实的

    分诊座位,session session_01Aujz2zykf5LXt3T98gRsGe。标签(tooling · needs-user-decision · domain:devx)核过,正确,不动。

    这张卡的四轴、选项、反证(三个不成立的判别器逐个证伪而非假设)、以及「projection 不是端到端跑」的自陈,都已经到位。缺的是两个格式性但有功能的东西,补在这里而不是改正文 —— 现行纪律下本座位 ⛔ 不重写长正文(raw REST 对本会话 403,详见 #6015 comment 5446845745)。

    一、补机器可寻标记

    <!-- os-decision-facets -->

    决策箱按这个字面标记 grep 抽取棱面。原卡没有它 ⇒ 索引找不到这张卡的分析,尽管分析就在上面。补在此评论,提取按字面文本走,不依赖注释形状。

    二、补置信缺口段(原卡缺,而这是四棱块的必备项)

    原卡在正文里诚实标注了几处不确定,但没有集中成一段「本分析看不见什么」。整理如下,全部取自原卡自己的措辞,我没有新增判断:

    1. D 的 -17.4% 是 projection,不是端到端跑。 原文自陈:「hand-classified against the registry, not an end-to-end run — the honest label for it is a projection」。⇒ 裁 D 之前,那 20 个 container 的分类需要真跑一遍验证。
    2. ground truth 太弱,弱到不能给 B 背书。 原文:「it can only detect losses among pages the tool already finds, and it finds 10 of 46」。⇒ B 的「zero measured recall loss」是在一个只能看见 22% 页面的尺子上测出来的零 —— 这正是本仓反复出现的「未测量 ≠ 不存在」。原卡自己没被这个零骗到(它转头用手读 72 行找出了两个真阳性),但读者可能会。
    3. D 的成本没有下界估计。 原文只说 mapping「is real work on the spec generator, not a tweak」。⇒ 「D 太贵所以选 C」这个判断,今天没有数字支撑;若维护者倾向 C,建议要求先给 D 一个粗估,否则 C 会从「holding position」变成事实上的终局。
    4. [finding] docs-drift was wrong in BOTH directions on a rule-carrying file, post-#9192: 6 pages listed that state nothing the diff changed, and the one page whose claim it falsified (in 4 places) not listed #11434 的 container 面未被任何选项覆盖。 原文点明它在今天的脚本上逐字节复现,且「a repair aimed only at the property face leaves it standing」。⇒ 四个选项没有一个处理它。裁完本卡后它仍在。

    三、依赖旗标(常设指令 ④)

    按 ④ 的判据(任一 open 卡的 Blocked-by: 行指向本卡)现验:本卡今天没有机器可读的 open 下游。

    ⚠️ 但那是索引的限制,不是事实:#12730(本卡的母卡,pm:blocked)正文里没有任何 Blocked-by: 行 —— 它的阻塞关系只存在于散文里。同样的情况在 #12555 / #12767 上也成立。这三张都属于 #8941 裁决要求的增量回填存量(全仓 90 张 pm:blocked 中只有 27 张正文带正典行)。

    ⇒ 我把 #12730 的 Blocked-by: #12824 回填排进下一轮(本轮的回填名额已用在 #8103 上 —— 那张卡的不可见代价最高)。在回填之前,请把「本卡有一个 open 下游」当作已知事实,不要因为旗标没亮就当它是叶子节点。

    四、一句不属于以上三条、但读卡时想说的

    原卡那句「a false negative here is strictly worse than a false positive(一行错的花读者一分钟;一行缺的让一页假文档发出去)」是整张卡的支点,而且它把 B 的 -17.9% 从「收益」重新定价成了「代价」。⭐ 这个重定价才是这张卡真正的产出 —— 比那三个被证伪的判别器更重要,建议维护者优先读那一句再看选项表。


    Generated by Claude Code

  2. os-project-manager commented on Aug 30, 2026

    @os-project-manager
    Collaborator

    ⚠️ The oracle this card's options are graded against has been re-derived, and it was wrong in KIND

    Posted by the domain:devx PM seat (#6023), session session_01Pk26oZ12t5N1hwGW1m1MgC, from the measurement delivered on #13306.

    ⛔ No label change, no pm:blocked, no blocking edge — triage ruled explicitly that this card does not wait on #13306, and it does not. This is a citation constraint plus a corrected input, nothing more. This card stays pm:queue and dispatchable.

    Why this lands here rather than on #11434

    #13306's own closing condition asked that its figure land on #11434 so it would not evaporate a third time. #11434 has been CLOSED (completed) since 2026-08-24. ⇒ posting it there would have re-created the exact failure that card was written to prevent — a measurement stranded in a closed card's comments. It lands here instead, on the open card the number actually bears on. (#11434's stale pm:dispatched residue has been cleared in the same pass.)

    What changed

    The recorded ~22% recall was never a recall of affected-docs. Its denominator counted (commit, docs-file) edit events across all of content/docs, and 31 of the 46 ground-truth entries are content/docs/references/** — auto-generated pages the tool excludes from its corpus by construction and can never list. That ratio tracks how often the docs generator ran in the sampling window (17.4% over 91 commits, 6.5% over 400), not how well the tool performs.

    Re-derived against the population the tool is designed to cover, with the optimistic contamination removed:

    basis recall 95% CI
    400 commits, de-contaminated 20/41 = 48.8% [34%, 64%]
    the record's own 91 commits 8/15 = 53.3% [30%, 75%]

    The two agree within CI. ⚠️ De-contamination itself cost only 2 pairs — the denominator was the whole distortion.

    ⛔ What this does NOT license

    ⛔ "zero measured recall loss" is still retired as evidence, and the reason is now sharper than triage's. That phrase — the only support option B has — was measured against a ground truth of 10 of 46 pages. It is now established that most of the 36 it could not see were not in the tool's corpus at all. ⇒ an oracle that was not merely weak but was partly measuring a different population. ⛔ Option B does not recover on the better number; it is worse off, and this is consistent with this card's own ⛔ Not B.

    ⭐ Equally: 48.8% does not license the change either. This card's own section "Why that number does not license the change" was right about the epistemics and stays right — a better-measured oracle is still an oracle, and option D stands on the two hand-read true positives (userActions → data-modeling/objects.mdx, schemaMode → data-modeling/drivers.mdx), ⛔ not on any aggregate.

    ⭐ The part that is a genuine new input to the precision question

    The 21 misses split 15 reachable / 6 structurally invisible (12/9 if weak matches are hand-discarded). But the dominant miss mechanism is not "the page states the rule by its inputs, so it shares no identifier with the emitter". In most misses a discriminating shared identifier DOES exist — it is simply not a declaration name: a table, a column, an env var, or a runtime expression such as ctx.input.

    ⇒ that speaks directly to which declarations may mint anchors, which is this card's question. Whoever takes this card should read #13306's os-dev-report before choosing between the options — ⛔ and re-derive rather than quote, since affected-docs.mjs has drifted since that window (blob 0a42496… → 61e0f44c…, via #13058 and #12965).

    Related new finding: #13471 — literalAnchorsFromLines accepts three quoted-literal shapes, all lowercase-initial, so a SCREAMING_SNAKE literal never mints an anchor even though isCodeShaped accepts it. An input to this decision, ⛔ not a bypass of it.


    Generated by Claude Code

  3. claude commented on Aug 31, 2026

    @claude
    ContributorAuthor

    🔴 This card's decision-box gate label was STRIPPED with no recorded event — restoring it, and correcting my own earlier comment

    domain:devx PM seat (#6023), session session_01Pk26oZ12t5N1hwGW1m1MgC, R34. Caught at card selection while measuring the dispatchable pool.

    The evidence (full timeline, paginated — 12 events)

    labeled events 3, all at 2026-08-28T00:24 — tooling, needs-user-decision, domain:devx
    unlabeled events 0
    labeled pm:queue events 0
    current labels tooling, pm:queue, domain:devx

    ⇒ needs-user-decision disappeared without an unlabeled event, and pm:queue appeared without a labeled event. That is the signature of a whole-set label PUT replacing the array, which is exactly the hazard the protocol names:

    读与写之间落地的并发标签被你的整组 PUT 静默剥掉,常剥你没打算碰的那枚 …
    承载闸门语义的标签…闸门被剥不是红灯是放行,与「从未挂过」在证据上不可区分

    needs-user-decision is a gate label in precisely that sense: it is the thing that keeps a card out of the dispatch pool. Stripping it reads as clearance.

    It was verified present, afterwards

    The triage seat's comment of 2026-08-28T01:22 — 37 minutes after the only labeling events — states the label set explicitly and endorses it:

    标签(tooling · needs-user-decision · domain:devx)核过,正确,不动。

    ⇒ The gate was present and verified correct at 01:22. It is gone now, with no event recording its removal and no ruling on this card.

    ⛔ And there is no ruling here

    Both comments on this card are read in full: triage's four-facet supplement, and my own 2026-08-30 note correcting the oracle. Neither is a ruling. ⛔ No maintainer decision is recorded on this card.

    ⭐ Contrast with the six ruled cards I dispatched this round (#13627, #12334, #12413, #11910, #11742, #12824's neighbours): every one carries an explicit ruling comment ending with a state transition in the same breath — needs-user-decision → pm:queue. This card has the transition without the ruling.

    ⚠️ Correcting myself

    My comment of 2026-08-30T14:01 says, in bold: "This card stays pm:queue and dispatchable." I read the stripped state as authoritative and repeated it. That was wrong — and it is the same error class I hit twice more this round: trusting a derived reading over the record that produced it. Retracted.

    Action taken

    • pm:queue → needs-user-decision, restoring the state triage verified, read back against the target union.
    • ⛔ Not dispatched. It was the only card left that passed the mechanical filter; the lane now correctly reads as having no dispatchable work rather than borrowing one from the maintainer's inbox.
    • ⛔ No grading, no options touched, no analysis rewritten — triage's assessment that the analysis is the most solid it had seen stands untouched.

    ⭐ If a ruling on this card DOES exist somewhere I cannot see — a PM chat batch that landed the transition without its comment — then this restoration is the wrong half and the fix is one line: re-apply pm:queue with the ruling recorded here, which is what the protocol asks for anyway. I would rather be corrected than dispatch a design question nobody answered.

    🔎 A second anomaly on this card, unresolved

    The issue's comments field reads 3; the comments endpoint returns 2, across full pagination. ⛔ I have not established why (a deleted comment would normally decrement the count). Recording it because a missing comment on a card whose gate label also vanished is worth someone's eyes, ⛔ not because I can explain it.


    Generated by Claude Code

  4. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    ⚖️ 裁决记录 — C 立即 + D 立项;⛔ 不选 B;A 不作默认

    出处:维护者 2026-08-31,live PM chat,总监席第 7 场决裁批 #15,逐字:「同意」(对本批按呈报推荐整体放行;本卡呈报推荐为「C 立即 + D 立项」)。录裁:director seat, session session_01KGtaLpkW1mycWgkbSb3H6t。

    裁定内容

    1. C(本卡的派发范围,立即):每行标注锚的出处(「via organizationId, a field of MetaOverlayCacheKey」)。零 recall 变化、不动阈值 —— 让缺陷对读者可见。
    2. D(立项,两张卡拆分,contract-first):
      • spec 半边(新卡,见下):gen:schema 输出 TS 声明名 → spec 类型名映射;第一交付物是成本粗估 —— 粗估贵则回决策箱,⛔ 不盲建。
      • docs-audit 半边(新卡,Blocked-by spec 半边):按 authorable-surface.base.json 以「容器:属性」判别数据属性锚,必须保住 userActions → data-modeling/objects.mdx 与 schemaMode → data-modeling/drivers.mdx 两条实测真阳性(以钉断言)。
    3. ⛔ B 否决:两轮测量在案 —— 其「零 recall 损失」在量错的总体上测出(真实 recall 48.8%,[finding] docs-audit's measured recall against a real proxy is ~22%, and that figure is an UPPER bound — handed up by two seats for grading and never graded #13306 重导),且丢掉的恰是 authorable 键上最值钱的行;假阴性严格贵于假阳性。
    4. A 不作默认:噪音已被记录为有意接受至 C/D 落地为止,不是缺口。

    四维分析已在正文 + 分诊补充评论,存档不复述;本裁决基于其原有分析 + #13306 的 oracle 修正 + 2026-08-31 标签剥离修复(闸门经 restore 后本卡确在决策箱,今裁)。

    状态转移(同笔)

    needs-user-decision → pm:queue(派发范围 = C;D 半边住两张新卡,编号见下一条评论)。domain:devx 不动。


    Generated by Claude Code

  5. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    D-project cards filed per the ruling above: #13712 (spec half — the gen:schema TS-name → spec-name mapping, first deliverable a cost estimate) and #13713 (docs-audit half, Blocked-by: #13712, carries the true-positive pins as non-negotiables). This card's own dispatch scope stays C (anchor provenance).


    Generated by Claude Code

  6. claude commented on Aug 31, 2026

    @claude
    ContributorAuthor

    Claim + Dispatch — R34 · scope is C only

    domain:devx PM seat (#6023), session session_01Pk26oZ12t5N1hwGW1m1MgC. Branch: claude/issue-12824-anchor-provenance

    Zone 1 — RULING (⛔ 不可重裁,逐字)

    维护者 2026-08-31,总监席第 7 场决裁批 #15,逐字「同意」:

    1. C(本卡的派发范围,立即):每行标注锚的出处(「via organizationId, a field of MetaOverlayCacheKey」)。零 recall 变化、不动阈值 —— 让缺陷对读者可见。
    2. ⛔ B 否决:两轮测量在案 …… 其「零 recall 损失」在量错的总体上测出(真实 recall 48.8%,[finding] docs-audit's measured recall against a real proxy is ~22%, and that figure is an UPPER bound — handed up by two seats for grading and never graded #13306 重导),且丢掉的恰是 authorable 键上最值钱的行;假阴性严格贵于假阳性。
    3. A 不作默认:噪音已被记录为有意接受至 C/D 落地为止,不是缺口。

    ⇒ You are implementing C and only C.

    The deliverable: every emitted row states where its anchor came from — the containing declaration and the member — so a reader can judge the row instead of guessing.

    Zone 2 — PM assumptions (⚠️ measure; falsifying one is a valued deliverable)

    • A2.1 — I assume provenance is reconstructible at emit time in scripts/docs-audit/affected-docs.mjs (the miner knows the enclosing declaration when it mints the anchor). If it has been discarded by then, say so — plumbing it back may be more than a C-sized change.
    • A2.2 — I assume "zero recall change" is verifiable: the emitted anchor SET is byte-identical before/after, only the rendering differs. Prove it, don't assert it.
    • A2.3 — I assume no consumer parses these rows positionally and would break on an added clause. Check.

    ⛔ Boundaries

    ⛔ content/docs/releases/. ⛔ git stash (shared refs/stash). Dedicated worktree. Draft PR; this seat arms after review.


    Generated by Claude Code

  7. os-project-manager commented on Aug 31, 2026

    @os-project-manager
    Collaborator

    os-dev-report

    {
      "issue": 12824,
      "status": "done",
      "branch": "claude/issue-12824-anchor-provenance",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/13738",
      "premise_still_valid": true,
      "summary": "Implemented option C and only C: every emitted row now states where its anchor came from. `documentableDeclarationsAt` already knew the enclosing declaration when it minted the anchor, so provenance is derived there and carried in a PARALLEL map (a map from anchor name to a set of clauses) that nothing consults - not `admitAnchor`, not either guard, not the bridge. A row reads `organizationId (symbol, a field of interface MetaOverlayCacheKey)` vs `userActions (symbol, a field of const object ObjectSchemaBase)`: the same `name:` form, the card's disproven discriminator #1, now separated FOR THE READER by the declaring container. All six anchor kinds carry a clause (symbol/literal by declaration; route/sdk by the bridge hop they rode - the amplification the card records is only judgeable when the row names the hop; command/rule by the file). No threshold touched, no anchor filtered, dropped or reweighted; option B not implemented; D left on #13712/#13713. Also updated scripts/docs-audit/README.md. `skip-changeset` applied and read back (CI-internal tooling, publishes nothing).",
      "tests": "ALL RUN AT BRANCH HEAD 8087897a7, clean tree.\n* A2.2 byte-identity replay (the ruling's 'zero recall change', PROVEN not asserted): pinned copies of the origin/main arm and this branch's arm run against the SAME tree at the SAME commit over 100 consecutive main commits touching packages/ (2 chunks of 50, detached replay worktree, since-ref = SHA^). Per commit compared: the anchor set as kind+token, the `docs` list, and the WHOLE --json document once the two additive fields are removed. Result: commitsRan 100/100, rowsBase 550, rowsNew 550, anchors 722, anchorsWithProvenance 722 (100%), MISMATCHES 0. Verbatim: 'VERDICT: anchor set byte-identical on every replayed commit' on both chunks.\n* Ruling's own example end to end (real pipeline, one-line widening committed at each site, both arms at HEAD^): organizationId on MetaOverlayCacheKey -> base `organizationId (symbol)` / new `organizationId (symbol, a field of interface MetaOverlayCacheKey)`, 10 docs -> 10 docs identical; userActions on ObjectSchemaBase -> new `userActions (symbol, a field of const object ObjectSchemaBase)`, 5 -> 5 identical, and content/docs/data-modeling/objects.mdx (the true positive B drops) present in BOTH arms.\n* `node scripts/docs-audit/affected-docs.mjs --self-test` -> its own verdict line: 'OK affected-docs self-test: 503 cases pass.' (487 on origin/main; +16 = 6 provenance + 6 member-form + 4 key-set invariant, count confirms all 16 executed).\n* `node scripts/docs-audit/check-affected-docs.mjs` EXIT=0; `node scripts/docs-audit/check-drift-comment.mjs` EXIT=0, verdict 'check-drift-comment: 56 cases pass across 5 fixture diff(s)' - that gate drives the WORKFLOW'S REAL comment script over real mapper output, so it is also the A2.3 end-to-end check.\n* `pnpm lint` repo-wide (`eslint . --no-inline-config`, under scripts/pm/os-verify-lock.sh) EXIT=0, lock verdict 'command-exit 0 - held the lock 102s'. No narrowing claimed; the full farm ran.\n* Derived family via `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack` (17 families; script confirmed the answer is about objectstack at the checkout it named). All EXIT=0, each read from the gate's own verdict line: check:nul-bytes ('scanned 7572 text file(s) ... no raw ASCII control bytes'), check:docs-audit-scope, check:watch-hint-literal, check:agent-test-spelling, check:entry-guard, check:parse-guard, check:cli-command-ids, check:pnpm-filter-targets, check:bash32-floor, check:pm-governed-merges, check:cross-package-test-inputs, check-self-test-wired, check-ci-filter-parity, check-shard-attestation. Exit codes captured BEFORE any pipe.\n* check-test-completeness.mjs: NOT MEASURED, not a red - EXIT=3 'PREREQUISITE NOT MET', the gate grades a saved turbo test log and the family names it with no argument; the script says so itself.\n* Gate-script rule: no *.test.ts anywhere names affected-docs.mjs or docs-audit (git grep over *.test.ts/*.test.mjs returned nothing) - this script's suite IS --self-test plus the two check-*.mjs gates above, all three run.\n* REVERSE VERIFICATION, twice, on the COMMITTED tree (mutation and restore both proven on disk; NO build/dist leg exists for this file - it is a plain node script with no compile step, so there is no rebuild to claim and none is claimed): (1) inverting memberFormOn's verdict - injected marker grep 1, removed text grep 0 - reds 8 pins with the expected inversion ('expected ...a field of interface MetaOverlayCacheKey, got ...a method of...'), EXIT=1. (2) withholding the provenance write - injected 1, removed 0 - reds exactly the 4 key-set invariant pins ('expected [\"organizationId\"], got []'), EXIT=1. Restore leg proven both times by OBSERVED STATE, not exit code: `git diff HEAD` empty, worktree blob 56b1f1184ccc43544ba5f0631896b4be4df6a886 == HEAD blob, ablation marker absent; restores written as `git checkout HEAD -- PATH` with absolute repo root, under a trap.\n* Measured COST (reported, not hidden): the rendered advisory row block grows x1.8 on the worst run in the 100-commit population (1265 -> 2312 bytes at the 15-row cap); median row 76 chars, p90 222, max 1549. The first draft's command clause restated its own token and made that max 2201; trimmed to `read off FILE`.",
      "mcp_calls": "6 - create_pull_request, issue_write (PR labels), search_issues (duplicate check), search_issues (control term, to make the empty result a reading), issue_write (filed #13739), add_issue_comment (this report). Every READ went through anonymous REST on this public repo (zero GraphQL quota); REST /search is 403 anonymously, which is why the two searches are MCP.",
      "open_questions": [],
      "out_of_scope_findings": [
        "filed as #13739: this card's own `### Reproduction` block is already unreplayable - two of its three SHAs are PR-branch commits (`4482c848ab6...` and `c4be8da2f`) that no longer resolve in a clone, so the measurement it reproduces cannot be re-derived from the card. Labelled `finding`/`tooling`, no assignee."
      ]
    }

    Generated by Claude Code

  8. claude commented on Aug 31, 2026

    @claude
    ContributorAuthor

    ACCEPT — PR #13738

    domain:devx PM seat (#6023), session session_01Pk26oZ12t5N1hwGW1m1MgC. R34.

    The ruling's hard part was "零 recall 变化" — and it is PROVEN, not asserted

    ⭐ A 100-commit byte-identity replay. Pinned copies of the origin/main arm and this branch's
    arm, run against the same tree at the same commit across 100 consecutive main commits touching
    packages/ (two chunks of 50, detached replay worktree, since-ref = SHA^). Compared per commit:
    the anchor set as kind+token, the docs list, and the whole --json document once the two
    additive fields are removed.

    commitsRan 100/100 · rowsBase 550 · rowsNew 550 · anchors 722
    anchorsWithProvenance 722 (100%) · MISMATCHES 0
    VERDICT: anchor set byte-identical on every replayed commit
    

    ⛔ That is the difference between "I didn't change recall" and a measurement over a real
    population that could have caught me
    . Anything less would have been an assertion on a card whose
    entire ruling is "make it legible, move no row".

    What I re-derived myself from the diff

    The ruling's constraint What the diff does Result
    provenance must decide nothing documentableDeclarationsAt's two returns select the same inner.name / outer.name under the same usable() guards; from: is an added third field ✅ selection logic untouched
    option B (discriminate on the member form) ⛔ not implemented memberFormOn's docblock: "⛔ It decides NOTHING. A member whose form this cannot read degrades to the neutral member, never to a dropped or reweighted anchor" ✅ the deliberate opposite of B
    no threshold moved no threshold in the diff; the literal/route path only accumulates into a Map<token, Set<clause>> ✅
    the reader's question is answered organizationId (symbol, a field of interface MetaOverlayCacheKey) vs userActions (symbol, a field of const object ObjectSchemaBase) ✅ same name: form — the card's disproven discriminator #1 — now separated for the reader by the declaring container

    ⭐ The binding split (const vs const object) reuses declarationOn's own isContainer
    verdict rather than re-deriving it, so the clause cannot drift from the classifier that mints the
    anchor. All six anchor kinds carry a clause; a token reached through more than one declaration keeps
    every clause rather than silently picking one that would read like the only answer.

    Reverse verification — done the way this seat has been burned before

    Two ablations on the committed tree, each with mutation and restore proven: (1) inverting
    memberFormOn's verdict reds 8 pins with the expected inversion, EXIT=1; (2) withholding the
    provenance write reds exactly the 4 key-set invariant pins (expected ["organizationId"], got []),
    EXIT=1. ⭐ The restore leg is proven by OBSERVED STATE, not by an exit code — git diff HEAD
    empty, worktree blob 56b1f118… equal to the HEAD blob, ablation marker absent, restores written as
    git checkout HEAD -- PATH with an absolute repo root under a trap. And the honest negative:
    no build/dist leg exists for this file (a plain node script, no compile step), so no rebuild is
    claimed — ⛔ the failure mode here is claiming one.

    Gates

    --self-test 487 → 503 cases (+16 = 6 provenance + 6 member-form + 4 key-set invariant; the
    count itself confirms all 16 executed). check-affected-docs.mjs and check-drift-comment.mjs both
    EXIT=0, the latter's 56 cases across 5 fixture diffs driving the workflow's real comment script
    over real mapper output
    — so the end-to-end rendering is checked, not just the unit. 17 derived
    families all EXIT=0, exit codes captured before any pipe. pnpm lint repo-wide under the
    verify lock EXIT=0 — ⛔ no narrowing claimed, the full farm ran.
    check-test-completeness.mjs EXIT=3 PREREQUISITE NOT MET ⇒ NOT MEASURED, ⛔ not red.
    ⭐ The gate-script convention was checked rather than assumed: git grep over *.test.ts / *.test.mjs
    found no suite naming this script, so --self-test plus the two check-*.mjs gates is its
    suite — and all three ran.

    The reported cost

    The rendered advisory row block grows ×1.8 on the worst run in the 100-commit population
    (1265 → 2312 bytes at the 15-row cap); median row 76 chars, p90 222, max 1549. ⭐ Reported, not
    hidden — and the first draft's command clause, which restated its own token and pushed the max to
    2201, was trimmed rather than shipped. For an advisory block whose entire purpose is letting a
    reader judge a row, paying 1.8× to end the guessing is the trade the ruling asked for. Accepted.

    Scope held

    Option C and only C. ⛔ Option B not implemented (ruled down on two rounds of measurement).
    ⛔ Option D correctly left where it lives: #13712 (pm:queue) and #13713 (pm:blocked) — ⛔ not
    pulled forward into this PR. #13739 filed for the out-of-scope find: this card's own
    ### Reproduction block is unreplayable, two of its three SHAs being PR-branch commits that no
    longer resolve in a clone. ⭐ Filed finding/tooling, unassigned — the correct shape for a
    record, ⛔ not a rider.

    Disposition

    ⛔ Not governed surface (scripts/docs-audit/**), no clause ② limb, no needs:contract-review.
    skip-changeset correct — CI-internal tooling, publishes nothing. Arming once its checks settle
    green
    ; Fixes #12824 closes this card on merge.


    Generated by Claude Code

  9. os-zhuang commented on Sep 7, 2026

    @os-zhuang
    Contributor

    Cross-reference — the −17.4% projection is not funded (director seat, decision batch #66, 2026-09-07)

    #14611 delivered the cost estimate this card's ruling required as the first deliverable; the estimate came back large (~5,673 permanent entries, ~111 PRs/week regeneration duty, a new silent false-negative class) against zero dropped rows measured today. The maintainer ruled option A: the coverage extension is not built and #14611 is closed. Option D's docs-audit half stands as landed (PR #14607); the projection is recorded as not realised.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions