Repository navigation
Five-layer test environment, and the twelve defects it found - #11
Merged
Merged
Conversation
Five layers over the skills in this repository, of which only the last costs money. TESTING.md is the reference; the short version: contract tests/skills/<name>/ vitest unit tests/template/unit/ vitest integration tests/template/integration/ vitest, real HTTP on loopback repo tests/repo/ vitest behaviour tests/skills/<name>/behaviour/ claude plugin eval The skill is the unit under test, so its tests live with it: one thin file per skill over a shared, family-aware contract in scripts/testkit/suite.ts. tests/ holds test cases and nothing else — the harness is in scripts/testkit/ behind the @testkit alias, so moving it never touches a case. repo.ts parses frontmatter with its own parser rather than importing anything this repository ships to users. A test that parses a skill with the code the skill is validated by cannot catch that parser being wrong. The contract enforces two conventions that had no mechanical check: every figure a skill states must be pinned by a citation, pointed at a script that re-derives it, or declared unpinnable naming where to look; and a description with same-family siblings must hand off to one by name, so the eval set can measure precision and not only recall. template/scripts/fetch_policy.py is the fetching contract the unit and integration layers bind: it names the script and this repository in the User-Agent and never impersonates a browser, reads robots.txt first (unreadable disallows everything, missing allows everything), honours crawl-delay with a one-request-per-second floor, stops on 403/429/503 rather than retrying, and states when a body came from a stored copy. There is no override flag. npm run test:prove applies 14 mutations one at a time and fails if the suite stays green for any of them. Two checks survived their first run and were strengthened: the fetching-policy check was satisfied by a header comment mentioning the policy rather than calling it. The suite is red at this commit. It reports seven defects in the tree it was pointed at, written up in ISSUES-FOUND.md; the commits that follow fix them one at a time. Closes #8 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-quotas.sh, fetch-review-requirements.py and both copies of
check-latest-version.sh called curl directly. They sent no User-Agent
naming the script or this repository, never read robots.txt, paced
nothing, and used --fail, which threw the status away so a 403, a 429
and a DNS failure were indistinguishable.
All four now go through fetch_policy.py, vendored beside them by
npm run sync:fetch-policy so each skill still works installed on its
own. tests/repo/coverage.test.ts fails if a copy drifts from the
template or if a script that reaches the network does not call it — the
check requires an actual invocation, because a header comment claiming
the policy is precisely what a bypass leaves behind.
Verified against the real hosts rather than only the fixture:
developers.google.com/robots.txt 200, disallows only /youtube/partner/
registry.npmjs.org/robots.txt 200 — and it is the npm package
*named* robots.txt, not a robots
file. No directive is found, so no
rules are imposed, which is what
RFC 9309 prescribes; the policy says
so in its reason rather than
reporting a successful parse of
nothing.
Each script's usage() prints a fixed line range of its own header, so
those ranges are updated for the lines added above them.
Closes #4
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apps-script-services stated 30 seconds (custom function runtime) and 6 minutes (script runtime) in bold with no citation, no live-check reference and no unpinnable declaration. The skill's own "Checking current quotas" section declares 9 KB, 500 KB and 6 hours unpinnable but never named these two. Both figures are correct. developers.google.com/apps-script/guides/ services/quotas returns "Custom function runtime 30 sec / execution" and "Script runtime 6 min / execution". They were unpinned, not wrong, so the fix points them at scripts/check-quotas.sh, which already covers both. apps-script-marketplace-publish documented "Drive app" as an --integration value. Google's page calls it "Google Drive app", so the documented command exits 3 with "No integration named "Drive app"" and the audit for a Drive integration never runs. Corrected, with a note that --list-integrations is the authority when the table has drifted. Found by running the shipped script against the live source, not by a test — worth remembering that the contract layer can only check that a figure is accounted for, never that it is right. Closes #2 Closes #3 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
None of the eleven descriptions named a same-family sibling. Each said what its skill covers and when to fire, but never where it stops. A description is a trigger, not a summary, and without a hand-off the eval set can only measure recall: a skill that fires on a neighbour's request scores perfectly and nothing says it was wrong. Each description now names the sibling that owns the adjacent request. The longest result is 686 characters, well inside the 1024-character cap. The four skills that vendor fetch_policy.py also announce it in their "Available files" section, since a file the model is never told about is dead weight in the bundle. Closes #1 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…is not ours No plugin entry in any of the three marketplace manifests carried a licence, nor did the three Gemini extension manifests. The marketplace entry is what a consumer reads at install time, which makes it the worst place for the licence to be absent. Every entry now carries Apache-2.0 and a test covers each one. .agents/skills/skill-creator/ is vendored from anthropics/skills and ships Anthropic's own LICENSE.txt, and nothing told a consumer that. THIRD-PARTY.md now records it, with what is deliberately not third-party alongside so the next reader does not have to work it out again. Upstream ships no NOTICE, so there is none to propagate — the LICENSE.txt beside the skill is the whole of it. The licence checks are pinned to the canonical text: LICENSE's terms block, and the vendored copy's, both hash to the SHA-256 of https://www.apache.org/licenses/LICENSE-2.0.txt with trailing whitespace stripped. The appendix and the copyright line sit outside the hash because every correct deployment of this licence edits those. Three manifests had also drifted apart: the bootgs description had lost "layered architecture enforcement" and apps-script had lost "UI (menus/sidebars/dialogs)", though both plugins ship those skills. Each manifest is one storefront's copy, and the stale one advertises skills it no longer describes. All copies now agree, and the invariant covers the Gemini manifests too. Closes #5 Closes #6 Closes #7 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…al set Each skill authors its eval set once in tests/skills/<name>/evals.ts, beside its contract test; gen-behaviour.ts turns that into the case tree claude plugin eval runs. The tree is committed so it is reviewable in a diff, and tests/repo/behaviour.test.ts fails when it has drifted from its source — a generated tree nobody can trust is a generated tree nobody reads. Two cases per skill: one for recall, and one that expects the skill to decline in favour of a named neighbour. The precision case asserts both halves mechanically — the neighbour's Skill fired, and this skill's did not (min 0, max 0) — so a skill that fires on everything cannot pass. Every tool_used: Skill grader is marked arm: with-only. The baseline arm has no skill to fire, so scoring it there marks the baseline down for the plugin's absence and flatters Δ. A repo test fails if one is left scored in both arms. Validated for free with --max-cost-usd 0: all 22 cases load and parse, 2 arms × 22 cases = 88 runs. No case has been run against a model, so no score is known. Also records ISSUES-FOUND.md, the seven defects the suite found in the tree it was pointed at, each traceable to a named failing test. Refs #8 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The check required the literal phrase "use when". A description opening "Use before anything goes public:" is a trigger — a better one than most in this repository — and was rejected as a summary. Found by running the contract against a skill from another library, which failed this check while getting the convention right. A check that fires on correct content is worse than no check, and this one would have rejected a well-written skill here just as readily. Now accepts use + a condition (when, before, after, whenever, while, if, during, any time). "Use for ..." stays rejected: that introduces a topic, which is what a summary does. Closes #9 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`files` was unset, so `npm pack` produced 187 files: the whole test suite, the harness, the vendored skill-creator, TESTING.md and ISSUES-FOUND.md — an internal list of this repository's own defects. `private: true` is what was preventing a publish, not intent. The allowlist brings the tarball to 42 files: the skills, the template, the plugin manifests, and the licence paperwork. Dropping .agents/ from the tarball also means the package no longer redistributes Anthropic's skill-creator; it is still published through git, where LICENSE.txt and THIRD-PARTY.md cover it. THIRD-PARTY.md told the reader to keep skills-lock.json's computedHash in sync by hand. That is not possible: the hash is written by the skills CLI and is not a SHA-256 of SKILL.md — neither the vendored copy nor upstream hashes to it. Replaced with the CLI instruction. Two claims that note makes are now verified rather than assumed: the vendored SKILL.md is byte-identical to current upstream (33168 bytes), and upstream ships no NOTICE at its root or beside the skill. Closes #10 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The conventions existed only in conversation. That has a cost this repository already paid: the "no session links" rule was broken 22 times in published history, and an agent working here later added the trailer again — because a harness instruction asked for it and there was no project instruction on disk to outrank it. Rule 3 now is that instruction. CLAUDE.md is a single @AGENTS.md import, so Claude Code reads the file itself rather than a second copy that drifts from it. Adapted rather than copied. MaksymStoianov/skills-private is private and dual-licensed, this one is public and wholly Apache-2.0, so rule 10 is the inverse of its counterpart and says so. Its rule about leaving existing session trailers alone is replaced by what actually happened here, kept as a warning rather than a precedent: the rewrite broke every clone, needed a force-push to two mirrors, and left the old commits reachable by SHA anyway. Four rules have no counterpart there, because the infrastructure is ours: the five-layer suite as the commit gate, the generated behaviour tree that must never be hand-edited, the vendored fetching policy that must be changed at the template and synced, and the figures and hand-off conventions the contract enforces. Every factual claim in it was checked against the tree rather than recalled — the three Gemini manifests, the ignore rules, the bootgs provider limitation, and that the two new files stay out of the npm tarball. Closes #12 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pull requests here were opened by hand with no title convention. create-pr detects the repository's own convention — commitlint or semantic-pull-request config, CONTRIBUTING, recent merged titles — rather than assuming one, and fills the body from the repository's template. Taken from the owner's public MaksymStoianov/skills, Apache-2.0, the same licence as this repository, so it carries no third-party obligation. Installed with the installer rather than copied, so skills-lock.json stays the single source of truth, as rule 1 requires. The installer leaves a mess, and it will do it again on the next add, so AGENTS.md now says what to clean up: a stray `agent/` directory at the root, and a `skills/<name>` symlink dropped into the directory for skills authored here — which the npm `files` allowlist would have shipped. Its Claude Code copy also fails under the Bash sandbox with EPERM, because `.claude/skills` is in the deny list, so that symlink is finished by hand. Closes #13 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The repository has a live Gitea mirror, which had to be caught up with a force-push yesterday, and nothing here could work with it: `gh` authenticates against github.com and cannot see a Gitea server at all. gitea-tea drives the official `tea` CLI for issues, pull requests, labels and releases, reading each repository's own issue templates and label set before drafting rather than assuming a shape. It is the pair to create-pr: their descriptions hand off to each other, so a request reaches whichever one can talk to the host in question. From the owner's public MaksymStoianov/skills, Apache-2.0, installed with the installer so skills-lock.json stays the source of truth. The installer left the same three things behind that AGENTS.md now documents — the stray `agent/` directory, a `skills/<name>` symlink in the directory for skills authored here, and an EPERM on the Claude Code copy under the Bash sandbox — and they were cleaned up the same way. Closes #14 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every file, commit message and document here is in English. Thirteen issues and one pull request were not, because the conversation that produced them was in Russian and nothing on disk said otherwise. On a public repository that is a real cost: the skills, the README and the commit log address an English-speaking audience, and the tracker explaining why those skills look the way they do did not. All fourteen have been rewritten; rule 12 is what stops it recurring. The same rule covers a second finding from a pre-publication check. AGENTS.md and two issue bodies named a private repository of the owner's by its full path and described its contents, and one named the Gitea mirror's internal hostname. Nothing there leaks a credential or anyone else's confidential material, but none of it was load-bearing either: rule 13 works without the name — a proprietary skill must not be copied into a public Apache-2.0 tree — and a remote is configured locally, so no issue needs to name a host. Both now say "a private repository". Closes #15 Closes #16 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ISSUES-FOUND.md was written when `gh` appeared to be broken and issues could not be opened. They can, and #1-#10 now carry all of it, so a static list of this repository's own defects sitting in a public tree is a stale duplicate of its tracker — which AGENTS.md rule 2 makes the authoritative record. The commits that mention the file are history and remain accurate as history; the one place that referenced it as a live path no longer does. Also drops `--experimental-strip-types` from the two npm scripts that carried it. Node 24 strips types without the flag, verified by running the script, and every other script here already omitted it. Closes #17 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suite existed and nothing ran it. A public repository with 322 tests had no checks on its pull requests, so the contract layer only caught what someone remembered to run locally — which is most of the point of having it gone. Runs `npm ci`, `tsc --noEmit`, `npm test`, the behaviour-tree drift check, and `npm run test:prove`. Python is installed because the shared fetching policy under test is Python, reached through scripts/testkit/bridge.py by the unit and integration layers. The behaviour layer is deliberately absent. It spends money and needs a model credential, so it stays a release gate run by hand, as AGENTS.md rule 9 says; CI verifies only that the committed case tree still matches the eval sets it was generated from. test:prove is included on purpose. A green suite is only evidence if red is reachable, and the cheapest place to discover a check has stopped firing is the pull request that broke it. Closes #18 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first run annotated all three actions: they target Node 20, which GitHub has deprecated and is already force-running on Node 24. Pinning to the current majors takes the warning away now rather than when the runners stop compensating. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Five layers over the skills in this repository, of which only the last costs money.
TESTING.mdis the reference.tests/skills/<name>/tests/template/unit/tests/template/integration/tests/repo/tests/skills/<name>/behaviour/claude plugin evalThe skill is the unit under test, so its tests live with it: one thin file per skill over a shared, family-aware contract.
tests/holds test cases and nothing else; the harness lives inscripts/testkit/behind the@testkitalias, so moving it never touches a case.The frontmatter parser is deliberately independent of anything this repository ships to its users. A test that parses a skill with the same code the skill is validated by cannot catch that parser being wrong — it agrees with itself.
State
npm test— 322 tests green,tsc --noEmitcleannpm run test:prove— 14 of 14 mutations caught--max-cost-usd 0(2 arms × 22 cases = 88 runs). No case has been run against a model, so no score is known.Two conventions that had no mechanical check
Every figure is pinned or declared unpinnable. Either a citation, a reference to a script that re-derives it from the source, or an explicit declaration naming where to look instead. A stated exception is auditable; silence looks identical to an oversight.
A description is a trigger, not a summary. Where a skill has same-family siblings, the description hands off to one by name, or the eval set can only measure recall.
The checks are proven, not assumed
npm run test:provebreaks the tree on purpose, one mutation at a time, and fails if the suite stays green. Two mutations survived the first run and both checks were strengthened: the "script uses the fetching policy" check was satisfied by a header comment mentioning the policy — exactly what a bypass leaves behind.Also in this branch
The working rules are now written down in
AGENTS.md, withCLAUDE.mdas a single import of it, and two skills are installed for the pull-request workflow:create-prfor GitHub andgitea-teafor the Gitea mirror.Closes
Closes #1, Closes #2, Closes #3, Closes #4, Closes #5, Closes #6, Closes #7, Closes #8, Closes #9, Closes #10, Closes #12, Closes #13, Closes #14, Closes #15, Closes #16
🤖 Generated with Claude Code