Repository navigation
Published-skills factual sweep: verify every behavioral claim in skills/** against the implementation — program anchor #13658
Description
Activity
- addeddocumentationImprovements or additions to documentationImprovements or additions to documentationpriority:p1High: required for production / M2High: required for production / M2
on Aug 31, 2026 zhuangjianguo commented
on Aug 31, 2026 CollaboratorAuthorMore actionsRoster update: flight ⑫'s merge regime is RESOLVED, not pending — ui#6866 half-1 was already ruled on 2026-08-30 (决裁批 #7, maintainer verbatim 「批 #7 同意」, recorded at ui#6866 comment 5469339478): Option A — published
skills/**is governed in objectui too;.mdcontent ⇒ maintainer review, pure code faces ⇒ seat contract-tier self-review. My earlier "pending" note came from a body-only triage read; corrected at the source card (comment 5474282135). Flight ⑫ therefore runs under the same governed-draft posture as flights ①–⑪. The half-1 TEXT implementation in objectui AGENTS.md is a separate card on that ruling, serial behind ui PR #6976.
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorAuthorMore actionsCalibration results (flight ① delivered — PR #13674, ACCEPTed, awaiting human merge)
- 197 behavioral claims across 13 classes in 579 lines (~118 lines non-behavioral, excluded by rule) · 6 FALSE = 3.0% · 26 NOT MEASURABLE (recorded) · 165 VERIFIED, all with executed probes where behavior-bearing.
- The sizing finding: falsehoods concentrate ~8× in surfaces/mapping tables — 4/38 (10.5%) there vs 2/159 (1.3%) elsewhere. Matches PR docs(skills):
$existsmeans HAS A VALUE — retire the MongoDB$existsclaim from objectstack-query #13577's pattern (a mapping column asserting engine behavior). Flights ②–⑫ therefore verify TABLES with execution priority first, prose classes second; the method stays full-inventory (the 1.3% is still real falsehoods). - Falsehood flavors worth pattern-matching in later flights: retired-surface rows outliving ADR-0049 removals (2 of 6) · phantom fields with zero repo-wide hits (2) · slot-semantics overclaims (
defaultValuenot auto-wrapping) (1) · unshipped-path pointers (src/in a dist-only package) (1).
Revised schedule: mid-size packages (≤1,200 lines) = one flight each; objectstack-data (4,935) = one card, up to three sequential PRs at whole-file boundaries (Part-of until the last, which closes); first full pass still projected 2–4 days, gated on merge cadence. Flight ② dispatching now.
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorAuthorMore actionsCalibration addendum from flight ①'s full report (effort economics, for sizing ②–⑫):
claim class effort/claim note executable-oracle rows (stdlib, formatters, cron forms, translation tables) ~1.5 min they BATCH: one probe file, one run spec-surface rows ~4 min schema location + retirement-note chasing — where both the cost and the yield live prose behavioral paragraphs ~8 min truth lives in large lint/validate modules cross-package runtime semantics NOT MEASURABLE from a package flight needs a kernel harness; ~20 min/claim if a flight is scoped to build one — 26 such claims recorded in flight ① and this class will recur in data/platform; a dedicated harness flight may pay for itself if their count crosses ~50 program-wide Wall clock flight ①: ~3.5h incl. builds (~45 sec/claim averaged — misleading, use the class split). A table-first pass would have found 4 of the 6 falsehoods for ~20% of the effort — ②–⑫ briefs already order it so.
Mechanical follow-on carded: #13679 — a named-identifier liveness gate over
skills/**(3 of 6 falsehoods were phantom rows, each detectable by one grep). The sweep fixes today's corpus; the gate keeps tomorrow's ADR-0049 retirements from re-creating phantoms.
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorAuthorMore actionsCorrection to the addendum above: the named-identifier liveness gate card is #13678 (the comment guessed #13679 before the create returned — read all references to #13679 there as #13678).
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorAuthorMore actionsFlight ② closing calibration + program state at shift wind-down
objectstack-data fully swept: 12/12 files · 4,935 lines · 883 claims · 13 FALSE (1.5%) · ~92 NOT MEASURABLE recorded. Three draft PRs awaiting human merge: #13687 · #13694 · #13698 (the last carries the closing keyword for member card #13675). Program totals after two flights: 1,080 claims · 19 FALSE (1.8%) across 5,514 of ~22,800 lines.
Method updates BINDING on flights ③–⑫ (from flight ②'s measured deltas):
- Target the content class, not the markdown syntax: 11 of 13 falsehoods are surface/liveness claims ("does this key exist / is it retired / is it automatic / what does the engine install"), but only 4 sit in pipe tables — bullet lists, callouts, numbered rules and code comments carry the rest. Flight ①'s "8× in tables" was a proxy for this class, not a shape rule.
- Both error directions collapse to one probe: for every surface a document names, PARSE it and read what the schema says back. Flight ② found the deader-than-real direction is now the majority (7/13:
maskingRule"pruned" but live-and-enforced,titleFormat"retired" but deprecated,ctx.log3 of 4 levels, 7 of 10 repo methods, …) — that direction hides shipped capability and no retired-things grep finds it. - Cross-file contradiction inside a package is the best free candidate-finder (found 4 of 13; every one then settled against the implementation).
- Planning floor: ~1.5–3% false density across the published catalog.
Program PAUSED at the maintainer's shift wind-down order (「当前任务处理完,合并后就可以下班了」): flights ③–⑫ are NOT dispatched; the roster, order, method and calibration above are the resume state for the next seat. Flight ③ (objectstack-query) additionally waits on PR #13577's merge (same files). Related open follow-ups: #13678 (phantom-row liveness gate) · #13695 (dead
'status'vocab, spec) · #13699 (master-detail severity judgment, spec).
Generated by Claude Code
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsProgram resumed by the incoming skills seat (session
session_01EnE7G31tqbxN1rqpQmzurT, takeover audit on seat post #7623). Flight ③ (objectstack-query) is DISPATCHED as member card #13717 — its blocker PR #13577 merged 03:15:39Z, verified before dispatch. Method comments 5474435934 / 5474464789 / 5475032239 are quoted as binding in the flight brief. Flights ④–⑫ remain queued on this roster; next flight dispatches when ③ lands or a batch slot frees.
Generated by Claude Code
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsFlight ③ delivered — member card #13717, draft PR #13740 (17 false landing sites / 13 distinct facts corrected, net −33 lines / −255 tokens, contract-review PASS on the PR, awaiting the maintainer's merge). Program totals after three flights: 1,376 claims inventoried · 37 false landing sites across 7,064 of ~22,800 lines.
Calibration deltas BINDING on flights ④–⑫, from ③'s measurements:
- Planning floor revised up: ③ measured 4.4% (distinct facts) / 5.7% (landing sites) — roughly 2× the 1.5–3% floor set after ②. Working range for remaining flights: 1.5–6%, with query/analytics-adjacent packages at the top of it (their surface claims age fastest against ADR retirements).
- Eval rubrics are claim carriers: 4 of ③'s 17 sites were in
evals/README.md, where a rubric TAUGHT graders to fail correct answers (havingmarked "silently dropped" while enforced). Flights ④+ include eval/rubric files in the inventory at full weight, not as appendix. check:skill-examplespopulation rule: the gate type-checks onlyos:check-marked fences; a package with zero markers has an empty population (control-verify against a package that has them, then record empty-population — not a skipped gate).- Cross-file contradiction scan stays first — ③: 4 found, 4 false, including a three-documents-agree-all-wrong case (per-aggregation
filter), the strongest instance yet of the "never verify one document against another" rule.
Flight ④ dispatches when a batch slot frees; roster order per the pause-state comment stands.
Generated by Claude Code
- added a commit that references this issue
on Aug 31, 2026 huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsFlight ④ delivered — member card #13747, draft PR #13760 (6 distinct false facts / 10 landing sites, net lines 0 / tokens −24, contract-review PASS, awaiting the maintainer's merge). Flight ③'s PR #13740 was merged by the maintainer at 09:46Z and its card #13717 auto-closed — the first flight to complete end-to-end through the governed terminal.
Program totals after four flights: 1,694 claims inventoried · 47 false landing sites across 9,697 of ~22,800 published lines. Density readings per flight: ① 3.0% · ② 1.5% · ③ 4.4–5.7% · ④ 1.9–3.1% — the 1.5–6% working range holds; query/analytics-adjacent surfaces sit at the top of it.
Harness-flight threshold CROSSED: the cross-package-runtime NOT MEASURABLE class now counts ~130 program-wide (① 26 · ② ~92 · ④ ~12+), far past the ~50 the flight-① addendum named as the point where a dedicated kernel/boot harness flight pays for itself. Added to the roster as an OPTIONAL flight (harness: a driver call-count + real-boot fixture that settles the recurring classes — expand batching, boot ordering,
os startexit behavior, two-tier date-bucket agreement). Not dispatched yet: it is a new validation surface, so it queues behind the maintainer's merge cadence like flights ⑤+, and its scope card will be filed for triage before any dispatch.Flight ⑤ (objectstack-ui, 2,298 lines) dispatches when the draft backlog clears — six sweep/lane drafts currently await human merges, and stacking more is the addendum's merge-cadence gate, not a stall.
Generated by Claude Code
- added a commit that references this issue
on Aug 31, 2026 huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsFlight ⑤ delivered — member card #13772, draft PR #13777 (8 distinct false facts / 8 landing sites, all SKILL.md; net −2 lines / −12 tokens; contract-review PASS; both carriers attached same-stroke this time; awaiting the maintainer's merge). Flight ④'s PR #13760 was merged by the maintainer at 10:43Z and #13747 auto-closed — two flights now complete end-to-end.
Program totals after five flights: ~1,879 claims inventoried · 55 false landing sites across 11,995 of ~22,800 published lines (53% of the catalog swept). Densities: ① 3.0% · ② 1.5% · ③ 4.4–5.7% · ④ 1.9–3.1% · ⑤ 4.3% — the range holds, analytics/UI-adjacent surfaces confirm at the top of it.
Method note binding on remaining flights: md-vs-generated-contract cross-checks are free candidates (⑤ found 2 of 8 that way — the generated JSON was RIGHT and the prose wrong both times, so the generated artifacts serve as oracles, not just sync targets). The retired-"silently-dropped" class (claims that outlived the protocol-17 strict-unknown-keys cutover) is now measured in three flights — remaining flights should grep their package for "silently"/"ignored"/"dropped" as a candidate seed.
Remaining roster: ⑥ objectstack-automation (1,163) · ⑦–⑪ api/i18n/upgrade/ai/pm-dispatch · ⑫ objectui skills/objectui (5,686) · optional cross-package harness flight. Cadence: three sweep drafts merged today; next flight dispatches on the next free wave.
Generated by Claude Code
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsFlight ⑥ delivered — member card #13793, draft PR #13808 (11 distinct false facts / 12 landing sites, 5.4–5.9%; net −4 lines / −24 tokens; contract-review PASS; both carriers same-stroke; awaiting the maintainer's merge). Flight ⑤'s PR #13777 merged at 11:56Z, #13772 closed — three flights now complete end-to-end.
Program totals after six flights: ~2,084 claims inventoried · 67 false landing sites across 13,158 of ~22,800 published lines (58% of the catalog swept). Densities: ① 3.0 · ② 1.5 · ③ 4.4–5.7 · ④ 1.9–3.1 · ⑤ 4.3 · ⑥ 5.4–5.9 — automation joins query at the top of the range.
Two findings worth the roster's attention:
- The wrong-default class has an upstream: ⑥'s sharpest correction (quorum
minApprovals"Default 1" vs runtime ALL-resolvable) traces to a false.describe()inpackages/specitself (spec:ApprovalNodeConfigSchema.minApprovalsdescribes "Default 1", but the quorum runtime defaults to ALL resolvable approvers #13809, filed for central triage) — where a skill falsehood has a spec-side describe() twin, fixing only the skill leaves the generated reference docs and Studio property forms still wrong. Remaining flights: when a false claim matches a.describe()string verbatim, file the spec-side twin. - The cross-package-runtime NOT MEASURABLE count keeps climbing (~176 program-wide with ⑥'s 46) — the optional harness flight's case strengthens.
Remaining roster: ⑦–⑪ api 707 · i18n 807 · upgrade 696 · ai 691 · pm-dispatch 1,077 · ⑫ objectui
skills/objectui5,686 · optional harness. Note for ⑪ (pm-dispatch): that package is ALSO a pm-dispatch governance face — construction stays opus but the flight brief must fence.claude/skills/pm-dispatch/**(separate corpus, separate audit #13597) fromskills/objectstack-pm-dispatch/**(the published target), and #13729's published-half scoping already covers its recipe line.
Generated by Claude Code
- The wrong-default class has an upstream: ⑥'s sharpest correction (quorum
huangyiirene commented
on Aug 31, 2026 CollaboratorMore actionsFlights ⑦ + ⑧ delivered; flight ⑥'s PR #13808 merged 14:01Z (#13793 closed — four flights complete end-to-end).
- ⑦ objectstack-api — card skills-sweep ⑦: objectstack-api (707 lines, 3 files) — behavioral-claim verification, content-class execution-first #13814, draft PR docs(skills): correct 8 false behavioral facts in objectstack-api (sweep flight ⑦) #13827: 8 distinct / 9 sites (3.6% / 4.0%). Sharpest:
error.codetaught as the numeric status (it is the semantic string; the number ishttpStatus), and a driver catalog teaching the retiredmongoalias while omitting livesqlite-wasm/turso. Spin-off: finding(spec):RestApiEndpointSchema.handlerStatusis authorable but has zero runtime consumers — the 501 it is documented to cause comes from somewhere else #13823 (handlerStatusauthorable with zero runtime consumers — ADR-0049 class). - ⑧ objectstack-i18n — card skills-sweep ⑧: objectstack-i18n (807 lines, 3 files) — behavioral-claim verification, content-class execution-first #13815, draft PR Correct 17 false behavioral claims in skills/objectstack-i18n — sweep flight ⑧ #13833: 17 distinct / 20 sites — 8.5% / 10%, the highest density in eight flights, ABOVE the 1.5–6% working range. Spin-offs: lint:
validate-translation-referenceshas no_tabsleg — an authored key naming a filter-preset tab that does not exist warns nobody #13835 (lint orphan-leg gap for_tabs) ·os i18n check --helpunder-counts what it reports by 9 of 14 key kinds — the CLI-side twin of a skill falsehood corrected in #13833 #13837 (CLI--helpunder-count — the spec-twin pattern on a CLI surface).
Program totals after eight flights: ~2,508 claims · 96 false landing sites across 14,672 of ~22,800 published lines (64% swept). Densities: ① 3.0 · ② 1.5 · ③ 4.4–5.7 · ④ 1.9–3.1 · ⑤ 4.3 · ⑥ 5.4–5.9 · ⑦ 3.6–4.0 · ⑧ 8.5–10.
Method updates BINDING on remaining flights (④–⑫ numbering: ⑨ upgrade 696 · ⑩ ai 691 · ⑪ pm-dispatch 1,077 · ⑫ objectui 5,686 · optional harness):
- Working range revised to 1.5–10%.
- The dominant class now has a sharper name than "tables": an enumeration that stopped growing when the schema did (a live schema member with no doc row — 15 of ⑧'s 20 sites). Mechanical countermeasure recorded as a second leg on gate card Gate candidate: named-identifier liveness check over skills/** — a skill row citing a schema/field that greps to zero repo-wide is a red, not a sentence #13678 (both directions, one corpus walk). Until that gate exists, flights compare every presented-as-exhaustive enumeration's row set against the schema's member set as a standing probe class.
- The spec-
.describe()/CLI-help twin rule (spec:ApprovalNodeConfigSchema.minApprovalsdescribes "Default 1", but the quorum runtime defaults to ALL resolvable approvers #13809 /os i18n check --helpunder-counts what it reports by 9 of 14 key kinds — the CLI-side twin of a skill falsehood corrected in #13833 #13837 pattern) holds: fixing only the skill leaves the same falsehood published one surface over — file the twin.
Generated by Claude Code
- ⑦ objectstack-api — card skills-sweep ⑦: objectstack-api (707 lines, 3 files) — behavioral-claim verification, content-class execution-first #13814, draft PR docs(skills): correct 8 false behavioral facts in objectstack-api (sweep flight ⑦) #13827: 8 distinct / 9 sites (3.6% / 4.0%). Sharpest:
- added a commit that references this issue
on Aug 31, 2026 - added 5 commits that reference this issue
on Sep 1, 2026 objectstack-fleet commented
on Sep 28, 2026 ContributorMore actionsClosed
completed: waves ①–⑪ of the published-skills factual sweep are delivered; wave ⑫ (objectui) has no carrier and is not filedTriage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-09-28T14:37Z.Provenance: the maintainer's ruling. In the triage seat's chat (2026-09-28), the long-term-program batch 1 was presented in the director format with this card as item ⑤. Option A was 「按完成关闭;如果 objectui 那一波确实需要做,就在 objectui 另立一张卡」, recommended A with B as the fallback. The maintainer replied, verbatim: 「13597 A 14292 A 13658 A」.
Delivered. Waves ①–⑪ (the objectstack
skills/**packages) shipped. The skills catalog program #14292 records this sweep as complete, 12/12, and it closed under the same ruling. This card's one sub-issue is closed.Wave ⑫, objectui
skills/objectui.- The seat found no card and no PR for it; it may have been done under another name, and that cannot be ruled out.
- The ruling set no condition to file it, and nothing measured pulls it. Under the filing gate, a card with zero pull is not filed.
- If the maintainer wants it: one line, and the seat files it in objectui as its own card.
The assignee is released in this act; the claim ends with the program.
Program anchor, filed by the skills lane seat (session
session_01EXxTW8mvPBhoHxmyPZ63de) on the maintainer's scheduling order.Mandate (verbatim)
2026-08-31, chat: 「指对外发布的 skills/**,也需要排程」 — confirming the program proposed after PR #13577. Standing context, 2026-08-21: 「对外发布的 skills 是整个平台的最大价值」.
Why a sweep, not more incidents
The published corpus has token/line ratchets (
scripts/check-skills-token-ratchet.mjs) and compile-validity gates (check:skill-examplestype-checks 260 prose examples), but nothing verifies behavioral truth against the implementation. PR #13577 proved the class: two published rows taught$exists→ MongoDB$existswhile every engine implements has-a-value (!= null) — found incidentally by an unrelated card, not by any sweep. One measured false row in ~186k published tokens is a floor, not a ceiling.Method — PR #13577 is the spec
Per behavioral claim (operator/API table rows, key names, behavioral sentences, code examples' asserted outputs):
content/docs/**follow-ups, not into ratcheted skill text) / NOT MEASURABLE (recorded with the reason — never silently skipped);Inventory and order
skills/objectuiBatching: one flight per package, ≤3 parallel, no file overlap. All objectstack
skills/**PRs: governed — draft, human merge.Members are filed one per package at dispatch time and link back here; this anchor tracks the roster. First calibration results re-size the schedule and get posted here.
Refs: PR #13577 (the class proof and the method precedent) · #13539 (the incident card) · #13597 (the separate PM-corpus audit — different corpus, parallel program).
Generated by Claude Code