Skip to content

Published-skills factual sweep: verify every behavioral claim in skills/** against the implementation — program anchor #13658

Description

@zhuangjianguo

Program anchor, filed by the skills lane seat (session session_01EXxTW8mvPBhoHxmyPZ63de) on the maintainer's scheduling order.

Mandate (verbatim)

2026-08-31, chat: 「指对外发布的 skills/**,也需要排程」 — confirming the program proposed after PR #13577. Standing context, 2026-08-21: 「对外发布的 skills 是整个平台的最大价值」.

Why a sweep, not more incidents

The published corpus has token/line ratchets (scripts/check-skills-token-ratchet.mjs) and compile-validity gates (check:skill-examples type-checks 260 prose examples), but nothing verifies behavioral truth against the implementation. PR #13577 proved the class: two published rows taught $exists → MongoDB $exists while every engine implements has-a-value (!= null) — found incidentally by an unrelated card, not by any sweep. One measured false row in ~186k published tokens is a floor, not a ceiling.

Method — PR #13577 is the spec

Per behavioral claim (operator/API table rows, key names, behavioral sentences, code examples' asserted outputs):

  1. Locate the implementing code (never verify one document against another document);
  2. verify by reading plus an executed test where the claim is behavior-bearing;
  3. verdict per claim: VERIFIED / FALSE (fixed in the sweep PR itself, byte-neutral-or-shrinking under the token ratchet — ⛔ no ceiling raises by devs; nuance overflow goes to content/docs/** follow-ups, not into ratcheted skill text) / NOT MEASURABLE (recorded with the reason — never silently skipped);
  4. non-vacuity control per sweep: at least one claim proven true by execution — a sweep reporting only-verified with zero executed evidence has measured nothing;
  5. PR body carries the per-item 落点 | before | after list (sweep-packaging五条); zero out-of-list changes.

Inventory and order

# package .md lines note
① objectstack-formula 579 calibration flight — measures cost-per-claim and false-density to size the rest
② objectstack-data 4,935 largest; same semantic family as the proven false row; split-by-file allowed with report
③ objectstack-query 1,550 ⛔ serial behind PR #13577's merge (same files)
④ objectstack-platform 2,633
⑤ objectstack-ui 2,298
⑥ objectstack-automation 1,163
⑦–⑪ api 707 · i18n 807 · upgrade 696 · ai 691 · pm-dispatch 1,077
⑫ objectui skills/objectui 5,686 merge regime pending the ui#6866 half-1 ruling; until ruled, governed posture (the fleet-side default)

Batching: one flight per package, ≤3 parallel, no file overlap. All objectstack skills/** PRs: governed — draft, human merge.

Members are filed one per package at dispatch time and link back here; this anchor tracks the roster. First calibration results re-size the schedule and get posted here.

Refs: PR #13577 (the class proof and the method precedent) · #13539 (the incident card) · #13597 (the separate PM-corpus audit — different corpus, parallel program).

Generated by Claude Code

Activity

  1. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    CollaboratorAuthor

    Roster update: flight ⑫'s merge regime is RESOLVED, not pending — ui#6866 half-1 was already ruled on 2026-08-30 (决裁批 #7, maintainer verbatim 「批 #7 同意」, recorded at ui#6866 comment 5469339478): Option A — published skills/** is governed in objectui too; .md content ⇒ maintainer review, pure code faces ⇒ seat contract-tier self-review. My earlier "pending" note came from a body-only triage read; corrected at the source card (comment 5474282135). Flight ⑫ therefore runs under the same governed-draft posture as flights ①–⑪. The half-1 TEXT implementation in objectui AGENTS.md is a separate card on that ruling, serial behind ui PR #6976.


    Generated by Claude Code

  2. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    CollaboratorAuthor

    Calibration results (flight ① delivered — PR #13674, ACCEPTed, awaiting human merge)

    • 197 behavioral claims across 13 classes in 579 lines (~118 lines non-behavioral, excluded by rule) · 6 FALSE = 3.0% · 26 NOT MEASURABLE (recorded) · 165 VERIFIED, all with executed probes where behavior-bearing.
    • The sizing finding: falsehoods concentrate ~8× in surfaces/mapping tables — 4/38 (10.5%) there vs 2/159 (1.3%) elsewhere. Matches PR docs(skills): $exists means HAS A VALUE — retire the MongoDB $exists claim from objectstack-query #13577's pattern (a mapping column asserting engine behavior). Flights ②–⑫ therefore verify TABLES with execution priority first, prose classes second; the method stays full-inventory (the 1.3% is still real falsehoods).
    • Falsehood flavors worth pattern-matching in later flights: retired-surface rows outliving ADR-0049 removals (2 of 6) · phantom fields with zero repo-wide hits (2) · slot-semantics overclaims (defaultValue not auto-wrapping) (1) · unshipped-path pointers (src/ in a dist-only package) (1).

    Revised schedule: mid-size packages (≤1,200 lines) = one flight each; objectstack-data (4,935) = one card, up to three sequential PRs at whole-file boundaries (Part-of until the last, which closes); first full pass still projected 2–4 days, gated on merge cadence. Flight ② dispatching now.


    Generated by Claude Code

  3. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    CollaboratorAuthor

    Calibration addendum from flight ①'s full report (effort economics, for sizing ②–⑫):

    claim class effort/claim note
    executable-oracle rows (stdlib, formatters, cron forms, translation tables) ~1.5 min they BATCH: one probe file, one run
    spec-surface rows ~4 min schema location + retirement-note chasing — where both the cost and the yield live
    prose behavioral paragraphs ~8 min truth lives in large lint/validate modules
    cross-package runtime semantics NOT MEASURABLE from a package flight needs a kernel harness; ~20 min/claim if a flight is scoped to build one — 26 such claims recorded in flight ① and this class will recur in data/platform; a dedicated harness flight may pay for itself if their count crosses ~50 program-wide

    Wall clock flight ①: ~3.5h incl. builds (~45 sec/claim averaged — misleading, use the class split). A table-first pass would have found 4 of the 6 falsehoods for ~20% of the effort — ②–⑫ briefs already order it so.

    Mechanical follow-on carded: #13679 — a named-identifier liveness gate over skills/** (3 of 6 falsehoods were phantom rows, each detectable by one grep). The sweep fixes today's corpus; the gate keeps tomorrow's ADR-0049 retirements from re-creating phantoms.


    Generated by Claude Code

  4. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    CollaboratorAuthor

    Correction to the addendum above: the named-identifier liveness gate card is #13678 (the comment guessed #13679 before the create returned — read all references to #13679 there as #13678).


    Generated by Claude Code

  5. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    CollaboratorAuthor

    Flight ② closing calibration + program state at shift wind-down

    objectstack-data fully swept: 12/12 files · 4,935 lines · 883 claims · 13 FALSE (1.5%) · ~92 NOT MEASURABLE recorded. Three draft PRs awaiting human merge: #13687 · #13694 · #13698 (the last carries the closing keyword for member card #13675). Program totals after two flights: 1,080 claims · 19 FALSE (1.8%) across 5,514 of ~22,800 lines.

    Method updates BINDING on flights ③–⑫ (from flight ②'s measured deltas):

    1. Target the content class, not the markdown syntax: 11 of 13 falsehoods are surface/liveness claims ("does this key exist / is it retired / is it automatic / what does the engine install"), but only 4 sit in pipe tables — bullet lists, callouts, numbered rules and code comments carry the rest. Flight ①'s "8× in tables" was a proxy for this class, not a shape rule.
    2. Both error directions collapse to one probe: for every surface a document names, PARSE it and read what the schema says back. Flight ② found the deader-than-real direction is now the majority (7/13: maskingRule "pruned" but live-and-enforced, titleFormat "retired" but deprecated, ctx.log 3 of 4 levels, 7 of 10 repo methods, …) — that direction hides shipped capability and no retired-things grep finds it.
    3. Cross-file contradiction inside a package is the best free candidate-finder (found 4 of 13; every one then settled against the implementation).
    4. Planning floor: ~1.5–3% false density across the published catalog.

    Program PAUSED at the maintainer's shift wind-down order (「当前任务处理完,合并后就可以下班了」): flights ③–⑫ are NOT dispatched; the roster, order, method and calibration above are the resume state for the next seat. Flight ③ (objectstack-query) additionally waits on PR #13577's merge (same files). Related open follow-ups: #13678 (phantom-row liveness gate) · #13695 (dead 'status' vocab, spec) · #13699 (master-detail severity judgment, spec).


    Generated by Claude Code

  6. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Program resumed by the incoming skills seat (session session_01EnE7G31tqbxN1rqpQmzurT, takeover audit on seat post #7623). Flight ③ (objectstack-query) is DISPATCHED as member card #13717 — its blocker PR #13577 merged 03:15:39Z, verified before dispatch. Method comments 5474435934 / 5474464789 / 5475032239 are quoted as binding in the flight brief. Flights ④–⑫ remain queued on this roster; next flight dispatches when ③ lands or a batch slot frees.


    Generated by Claude Code

  7. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Flight ③ delivered — member card #13717, draft PR #13740 (17 false landing sites / 13 distinct facts corrected, net −33 lines / −255 tokens, contract-review PASS on the PR, awaiting the maintainer's merge). Program totals after three flights: 1,376 claims inventoried · 37 false landing sites across 7,064 of ~22,800 lines.

    Calibration deltas BINDING on flights ④–⑫, from ③'s measurements:

    1. Planning floor revised up: ③ measured 4.4% (distinct facts) / 5.7% (landing sites) — roughly 2× the 1.5–3% floor set after ②. Working range for remaining flights: 1.5–6%, with query/analytics-adjacent packages at the top of it (their surface claims age fastest against ADR retirements).
    2. Eval rubrics are claim carriers: 4 of ③'s 17 sites were in evals/README.md, where a rubric TAUGHT graders to fail correct answers (having marked "silently dropped" while enforced). Flights ④+ include eval/rubric files in the inventory at full weight, not as appendix.
    3. check:skill-examples population rule: the gate type-checks only os:check-marked fences; a package with zero markers has an empty population (control-verify against a package that has them, then record empty-population — not a skipped gate).
    4. Cross-file contradiction scan stays first — ③: 4 found, 4 false, including a three-documents-agree-all-wrong case (per-aggregation filter), the strongest instance yet of the "never verify one document against another" rule.

    Flight ④ dispatches when a batch slot frees; roster order per the pause-state comment stands.


    Generated by Claude Code

  8. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Flight ④ delivered — member card #13747, draft PR #13760 (6 distinct false facts / 10 landing sites, net lines 0 / tokens −24, contract-review PASS, awaiting the maintainer's merge). Flight ③'s PR #13740 was merged by the maintainer at 09:46Z and its card #13717 auto-closed — the first flight to complete end-to-end through the governed terminal.

    Program totals after four flights: 1,694 claims inventoried · 47 false landing sites across 9,697 of ~22,800 published lines. Density readings per flight: ① 3.0% · ② 1.5% · ③ 4.4–5.7% · ④ 1.9–3.1% — the 1.5–6% working range holds; query/analytics-adjacent surfaces sit at the top of it.

    Harness-flight threshold CROSSED: the cross-package-runtime NOT MEASURABLE class now counts ~130 program-wide (① 26 · ② ~92 · ④ ~12+), far past the ~50 the flight-① addendum named as the point where a dedicated kernel/boot harness flight pays for itself. Added to the roster as an OPTIONAL flight (harness: a driver call-count + real-boot fixture that settles the recurring classes — expand batching, boot ordering, os start exit behavior, two-tier date-bucket agreement). Not dispatched yet: it is a new validation surface, so it queues behind the maintainer's merge cadence like flights ⑤+, and its scope card will be filed for triage before any dispatch.

    Flight ⑤ (objectstack-ui, 2,298 lines) dispatches when the draft backlog clears — six sweep/lane drafts currently await human merges, and stacking more is the addendum's merge-cadence gate, not a stall.


    Generated by Claude Code

  9. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Flight ⑤ delivered — member card #13772, draft PR #13777 (8 distinct false facts / 8 landing sites, all SKILL.md; net −2 lines / −12 tokens; contract-review PASS; both carriers attached same-stroke this time; awaiting the maintainer's merge). Flight ④'s PR #13760 was merged by the maintainer at 10:43Z and #13747 auto-closed — two flights now complete end-to-end.

    Program totals after five flights: ~1,879 claims inventoried · 55 false landing sites across 11,995 of ~22,800 published lines (53% of the catalog swept). Densities: ① 3.0% · ② 1.5% · ③ 4.4–5.7% · ④ 1.9–3.1% · ⑤ 4.3% — the range holds, analytics/UI-adjacent surfaces confirm at the top of it.

    Method note binding on remaining flights: md-vs-generated-contract cross-checks are free candidates (⑤ found 2 of 8 that way — the generated JSON was RIGHT and the prose wrong both times, so the generated artifacts serve as oracles, not just sync targets). The retired-"silently-dropped" class (claims that outlived the protocol-17 strict-unknown-keys cutover) is now measured in three flights — remaining flights should grep their package for "silently"/"ignored"/"dropped" as a candidate seed.

    Remaining roster: ⑥ objectstack-automation (1,163) · ⑦–⑪ api/i18n/upgrade/ai/pm-dispatch · ⑫ objectui skills/objectui (5,686) · optional cross-package harness flight. Cadence: three sweep drafts merged today; next flight dispatches on the next free wave.


    Generated by Claude Code

  10. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Flight ⑥ delivered — member card #13793, draft PR #13808 (11 distinct false facts / 12 landing sites, 5.4–5.9%; net −4 lines / −24 tokens; contract-review PASS; both carriers same-stroke; awaiting the maintainer's merge). Flight ⑤'s PR #13777 merged at 11:56Z, #13772 closed — three flights now complete end-to-end.

    Program totals after six flights: ~2,084 claims inventoried · 67 false landing sites across 13,158 of ~22,800 published lines (58% of the catalog swept). Densities: ① 3.0 · ② 1.5 · ③ 4.4–5.7 · ④ 1.9–3.1 · ⑤ 4.3 · ⑥ 5.4–5.9 — automation joins query at the top of the range.

    Two findings worth the roster's attention:

    1. The wrong-default class has an upstream: ⑥'s sharpest correction (quorum minApprovals "Default 1" vs runtime ALL-resolvable) traces to a false .describe() in packages/spec itself (spec: ApprovalNodeConfigSchema.minApprovals describes "Default 1", but the quorum runtime defaults to ALL resolvable approvers #13809, filed for central triage) — where a skill falsehood has a spec-side describe() twin, fixing only the skill leaves the generated reference docs and Studio property forms still wrong. Remaining flights: when a false claim matches a .describe() string verbatim, file the spec-side twin.
    2. The cross-package-runtime NOT MEASURABLE count keeps climbing (~176 program-wide with ⑥'s 46) — the optional harness flight's case strengthens.

    Remaining roster: ⑦–⑪ api 707 · i18n 807 · upgrade 696 · ai 691 · pm-dispatch 1,077 · ⑫ objectui skills/objectui 5,686 · optional harness. Note for ⑪ (pm-dispatch): that package is ALSO a pm-dispatch governance face — construction stays opus but the flight brief must fence .claude/skills/pm-dispatch/** (separate corpus, separate audit #13597) from skills/objectstack-pm-dispatch/** (the published target), and #13729's published-half scoping already covers its recipe line.


    Generated by Claude Code

  11. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Flights ⑦ + ⑧ delivered; flight ⑥'s PR #13808 merged 14:01Z (#13793 closed — four flights complete end-to-end).

    Program totals after eight flights: ~2,508 claims · 96 false landing sites across 14,672 of ~22,800 published lines (64% swept). Densities: ① 3.0 · ② 1.5 · ③ 4.4–5.7 · ④ 1.9–3.1 · ⑤ 4.3 · ⑥ 5.4–5.9 · ⑦ 3.6–4.0 · ⑧ 8.5–10.

    Method updates BINDING on remaining flights (④–⑫ numbering: ⑨ upgrade 696 · ⑩ ai 691 · ⑪ pm-dispatch 1,077 · ⑫ objectui 5,686 · optional harness):

    1. Working range revised to 1.5–10%.
    2. The dominant class now has a sharper name than "tables": an enumeration that stopped growing when the schema did (a live schema member with no doc row — 15 of ⑧'s 20 sites). Mechanical countermeasure recorded as a second leg on gate card Gate candidate: named-identifier liveness check over skills/** — a skill row citing a schema/field that greps to zero repo-wide is a red, not a sentence #13678 (both directions, one corpus walk). Until that gate exists, flights compare every presented-as-exhaustive enumeration's row set against the schema's member set as a standing probe class.
    3. The spec-.describe()/CLI-help twin rule (spec: ApprovalNodeConfigSchema.minApprovals describes "Default 1", but the quorum runtime defaults to ALL resolvable approvers #13809 / os i18n check --help under-counts what it reports by 9 of 14 key kinds — the CLI-side twin of a skill falsehood corrected in #13833 #13837 pattern) holds: fixing only the skill leaves the same falsehood published one surface over — file the twin.

    Generated by Claude Code

  12. objectstack-fleet commented on Sep 28, 2026

    @objectstack-fleet
    Contributor

    Closed completed: waves ①–⑪ of the published-skills factual sweep are delivered; wave ⑫ (objectui) has no carrier and is not filed

    Triage seat (objectstack-wide, seat post #6015) · session_01AavokzJ5DndAwitDXvKy4U · 2026-09-28T14:37Z.

    Provenance: the maintainer's ruling. In the triage seat's chat (2026-09-28), the long-term-program batch 1 was presented in the director format with this card as item ⑤. Option A was 「按完成关闭;如果 objectui 那一波确实需要做,就在 objectui 另立一张卡」, recommended A with B as the fallback. The maintainer replied, verbatim: 「13597 A 14292 A 13658 A」.

    Delivered. Waves ①–⑪ (the objectstack skills/** packages) shipped. The skills catalog program #14292 records this sweep as complete, 12/12, and it closed under the same ruling. This card's one sub-issue is closed.

    Wave ⑫, objectui skills/objectui.

    • The seat found no card and no PR for it; it may have been done under another name, and that cannot be ruled out.
    • The ruling set no condition to file it, and nothing measured pulls it. Under the filing gate, a card with zero pull is not filed.
    • If the maintainer wants it: one line, and the seat files it in objectui as its own card.

    The assignee is released in this act; the claim ends with the program.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions