Skip to content

Five-layer test environment, and the twelve defects it found - #11

Merged
MaksymStoianov merged 15 commits into
mainfrom
test-environment
Sep 22, 2026
Merged

MaksymStoianov merged 15 commits into
mainfrom
test-environment

Conversation

@MaksymStoianov

@MaksymStoianov MaksymStoianov commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Five layers over the skills in this repository, of which only the last costs money. TESTING.md is the reference.

Layer Lives in Proves Runner
Contract tests/skills/<name>/ the skill is well-formed for its family, every figure is pinned or declared unpinnable, every cross-reference resolves vitest
Unit tests/template/unit/ the shared fetching module in isolation vitest
Integration tests/template/integration/ the fetching rules bind, over real HTTP against a loopback fixture vitest
Repo tests/repo/ invariants no single skill can see vitest
Behaviour tests/skills/<name>/behaviour/ the skill changes what the model does claude plugin eval

The skill is the unit under test, so its tests live with it: one thin file per skill over a shared, family-aware contract. tests/ holds test cases and nothing else; the harness lives in scripts/testkit/ behind the @testkit alias, so moving it never touches a case.

The frontmatter parser is deliberately independent of anything this repository ships to its users. A test that parses a skill with the same code the skill is validated by cannot catch that parser being wrong — it agrees with itself.

State

  • npm test — 322 tests green, tsc --noEmit clean
  • npm run test:prove — 14 of 14 mutations caught
  • Behaviour layer: 22 cases load and parse, validated for free with --max-cost-usd 0 (2 arms × 22 cases = 88 runs). No case has been run against a model, so no score is known.

Two conventions that had no mechanical check

Every figure is pinned or declared unpinnable. Either a citation, a reference to a script that re-derives it from the source, or an explicit declaration naming where to look instead. A stated exception is auditable; silence looks identical to an oversight.

A description is a trigger, not a summary. Where a skill has same-family siblings, the description hands off to one by name, or the eval set can only measure recall.

The checks are proven, not assumed

npm run test:prove breaks the tree on purpose, one mutation at a time, and fails if the suite stays green. Two mutations survived the first run and both checks were strengthened: the "script uses the fetching policy" check was satisfied by a header comment mentioning the policy — exactly what a bypass leaves behind.

Also in this branch

The working rules are now written down in AGENTS.md, with CLAUDE.md as a single import of it, and two skills are installed for the pull-request workflow: create-pr for GitHub and gitea-tea for the Gitea mirror.

Closes

Closes #1, Closes #2, Closes #3, Closes #4, Closes #5, Closes #6, Closes #7, Closes #8, Closes #9, Closes #10, Closes #12, Closes #13, Closes #14, Closes #15, Closes #16

🤖 Generated with Claude Code

MaksymStoianov and others added 11 commits September 21, 2026 20:10
Five layers over the skills in this repository, of which only the last
costs money. TESTING.md is the reference; the short version:

  contract     tests/skills/<name>/      vitest
  unit         tests/template/unit/      vitest
  integration  tests/template/integration/  vitest, real HTTP on loopback
  repo         tests/repo/               vitest
  behaviour    tests/skills/<name>/behaviour/  claude plugin eval

The skill is the unit under test, so its tests live with it: one thin
file per skill over a shared, family-aware contract in
scripts/testkit/suite.ts. tests/ holds test cases and nothing else — the
harness is in scripts/testkit/ behind the @testkit alias, so moving it
never touches a case.

repo.ts parses frontmatter with its own parser rather than importing
anything this repository ships to users. A test that parses a skill with
the code the skill is validated by cannot catch that parser being wrong.

The contract enforces two conventions that had no mechanical check:
every figure a skill states must be pinned by a citation, pointed at a
script that re-derives it, or declared unpinnable naming where to look;
and a description with same-family siblings must hand off to one by
name, so the eval set can measure precision and not only recall.

template/scripts/fetch_policy.py is the fetching contract the unit and
integration layers bind: it names the script and this repository in the
User-Agent and never impersonates a browser, reads robots.txt first
(unreadable disallows everything, missing allows everything), honours
crawl-delay with a one-request-per-second floor, stops on 403/429/503
rather than retrying, and states when a body came from a stored copy.
There is no override flag.

npm run test:prove applies 14 mutations one at a time and fails if the
suite stays green for any of them. Two checks survived their first run
and were strengthened: the fetching-policy check was satisfied by a
header comment mentioning the policy rather than calling it.

The suite is red at this commit. It reports seven defects in the tree it
was pointed at, written up in ISSUES-FOUND.md; the commits that follow
fix them one at a time.

Closes #8

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check-quotas.sh, fetch-review-requirements.py and both copies of
check-latest-version.sh called curl directly. They sent no User-Agent
naming the script or this repository, never read robots.txt, paced
nothing, and used --fail, which threw the status away so a 403, a 429
and a DNS failure were indistinguishable.

All four now go through fetch_policy.py, vendored beside them by
npm run sync:fetch-policy so each skill still works installed on its
own. tests/repo/coverage.test.ts fails if a copy drifts from the
template or if a script that reaches the network does not call it — the
check requires an actual invocation, because a header comment claiming
the policy is precisely what a bypass leaves behind.

Verified against the real hosts rather than only the fixture:

  developers.google.com/robots.txt  200, disallows only /youtube/partner/
  registry.npmjs.org/robots.txt     200 — and it is the npm package
                                    *named* robots.txt, not a robots
                                    file. No directive is found, so no
                                    rules are imposed, which is what
                                    RFC 9309 prescribes; the policy says
                                    so in its reason rather than
                                    reporting a successful parse of
                                    nothing.

Each script's usage() prints a fixed line range of its own header, so
those ranges are updated for the lines added above them.

Closes #4

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apps-script-services stated 30 seconds (custom function runtime) and
6 minutes (script runtime) in bold with no citation, no live-check
reference and no unpinnable declaration. The skill's own "Checking
current quotas" section declares 9 KB, 500 KB and 6 hours unpinnable but
never named these two.

Both figures are correct. developers.google.com/apps-script/guides/
services/quotas returns "Custom function runtime 30 sec / execution" and
"Script runtime 6 min / execution". They were unpinned, not wrong, so
the fix points them at scripts/check-quotas.sh, which already covers
both.

apps-script-marketplace-publish documented "Drive app" as an
--integration value. Google's page calls it "Google Drive app", so the
documented command exits 3 with "No integration named "Drive app"" and
the audit for a Drive integration never runs. Corrected, with a note
that --list-integrations is the authority when the table has drifted.

Found by running the shipped script against the live source, not by a
test — worth remembering that the contract layer can only check that a
figure is accounted for, never that it is right.

Closes #2
Closes #3

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
None of the eleven descriptions named a same-family sibling. Each said
what its skill covers and when to fire, but never where it stops.

A description is a trigger, not a summary, and without a hand-off the
eval set can only measure recall: a skill that fires on a neighbour's
request scores perfectly and nothing says it was wrong. Each description
now names the sibling that owns the adjacent request. The longest result
is 686 characters, well inside the 1024-character cap.

The four skills that vendor fetch_policy.py also announce it in their
"Available files" section, since a file the model is never told about is
dead weight in the bundle.

Closes #1

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…is not ours

No plugin entry in any of the three marketplace manifests carried a
licence, nor did the three Gemini extension manifests. The marketplace
entry is what a consumer reads at install time, which makes it the worst
place for the licence to be absent. Every entry now carries Apache-2.0
and a test covers each one.

.agents/skills/skill-creator/ is vendored from anthropics/skills and
ships Anthropic's own LICENSE.txt, and nothing told a consumer that.
THIRD-PARTY.md now records it, with what is deliberately not third-party
alongside so the next reader does not have to work it out again.
Upstream ships no NOTICE, so there is none to propagate — the
LICENSE.txt beside the skill is the whole of it.

The licence checks are pinned to the canonical text: LICENSE's terms
block, and the vendored copy's, both hash to the SHA-256 of
https://www.apache.org/licenses/LICENSE-2.0.txt with trailing whitespace
stripped. The appendix and the copyright line sit outside the hash
because every correct deployment of this licence edits those.

Three manifests had also drifted apart: the bootgs description had lost
"layered architecture enforcement" and apps-script had lost
"UI (menus/sidebars/dialogs)", though both plugins ship those skills.
Each manifest is one storefront's copy, and the stale one advertises
skills it no longer describes. All copies now agree, and the invariant
covers the Gemini manifests too.

Closes #5
Closes #6
Closes #7

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…al set

Each skill authors its eval set once in tests/skills/<name>/evals.ts,
beside its contract test; gen-behaviour.ts turns that into the case tree
claude plugin eval runs. The tree is committed so it is reviewable in a
diff, and tests/repo/behaviour.test.ts fails when it has drifted from
its source — a generated tree nobody can trust is a generated tree
nobody reads.

Two cases per skill: one for recall, and one that expects the skill to
decline in favour of a named neighbour. The precision case asserts both
halves mechanically — the neighbour's Skill fired, and this skill's did
not (min 0, max 0) — so a skill that fires on everything cannot pass.

Every tool_used: Skill grader is marked arm: with-only. The baseline arm
has no skill to fire, so scoring it there marks the baseline down for
the plugin's absence and flatters Δ. A repo test fails if one is left
scored in both arms.

Validated for free with --max-cost-usd 0: all 22 cases load and parse,
2 arms × 22 cases = 88 runs. No case has been run against a model, so
no score is known.

Also records ISSUES-FOUND.md, the seven defects the suite found in the
tree it was pointed at, each traceable to a named failing test.

Refs #8

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The check required the literal phrase "use when". A description opening
"Use before anything goes public:" is a trigger — a better one than most
in this repository — and was rejected as a summary.

Found by running the contract against a skill from another library,
which failed this check while getting the convention right. A check that
fires on correct content is worse than no check, and this one would have
rejected a well-written skill here just as readily.

Now accepts use + a condition (when, before, after, whenever, while, if,
during, any time). "Use for ..." stays rejected: that introduces a topic,
which is what a summary does.

Closes #9

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`files` was unset, so `npm pack` produced 187 files: the whole test
suite, the harness, the vendored skill-creator, TESTING.md and
ISSUES-FOUND.md — an internal list of this repository's own defects.
`private: true` is what was preventing a publish, not intent. The
allowlist brings the tarball to 42 files: the skills, the template, the
plugin manifests, and the licence paperwork.

Dropping .agents/ from the tarball also means the package no longer
redistributes Anthropic's skill-creator; it is still published through
git, where LICENSE.txt and THIRD-PARTY.md cover it.

THIRD-PARTY.md told the reader to keep skills-lock.json's computedHash
in sync by hand. That is not possible: the hash is written by the skills
CLI and is not a SHA-256 of SKILL.md — neither the vendored copy nor
upstream hashes to it. Replaced with the CLI instruction.

Two claims that note makes are now verified rather than assumed: the
vendored SKILL.md is byte-identical to current upstream (33168 bytes),
and upstream ships no NOTICE at its root or beside the skill.

Closes #10

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The conventions existed only in conversation. That has a cost this
repository already paid: the "no session links" rule was broken 22 times
in published history, and an agent working here later added the trailer
again — because a harness instruction asked for it and there was no
project instruction on disk to outrank it. Rule 3 now is that
instruction.

CLAUDE.md is a single @AGENTS.md import, so Claude Code reads the file
itself rather than a second copy that drifts from it.

Adapted rather than copied. MaksymStoianov/skills-private is private and
dual-licensed, this one is public and wholly Apache-2.0, so rule 10 is
the inverse of its counterpart and says so. Its rule about leaving
existing session trailers alone is replaced by what actually happened
here, kept as a warning rather than a precedent: the rewrite broke every
clone, needed a force-push to two mirrors, and left the old commits
reachable by SHA anyway.

Four rules have no counterpart there, because the infrastructure is
ours: the five-layer suite as the commit gate, the generated behaviour
tree that must never be hand-edited, the vendored fetching policy that
must be changed at the template and synced, and the figures and
hand-off conventions the contract enforces.

Every factual claim in it was checked against the tree rather than
recalled — the three Gemini manifests, the ignore rules, the bootgs
provider limitation, and that the two new files stay out of the npm
tarball.

Closes #12

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pull requests here were opened by hand with no title convention.
create-pr detects the repository's own convention — commitlint or
semantic-pull-request config, CONTRIBUTING, recent merged titles —
rather than assuming one, and fills the body from the repository's
template.

Taken from the owner's public MaksymStoianov/skills, Apache-2.0, the
same licence as this repository, so it carries no third-party
obligation. Installed with the installer rather than copied, so
skills-lock.json stays the single source of truth, as rule 1 requires.

The installer leaves a mess, and it will do it again on the next add,
so AGENTS.md now says what to clean up: a stray `agent/` directory at
the root, and a `skills/<name>` symlink dropped into the directory for
skills authored here — which the npm `files` allowlist would have
shipped. Its Claude Code copy also fails under the Bash sandbox with
EPERM, because `.claude/skills` is in the deny list, so that symlink is
finished by hand.

Closes #13

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The repository has a live Gitea mirror, which had to be caught up with a
force-push yesterday, and nothing here could work with it: `gh`
authenticates against github.com and cannot see a Gitea server at all.

gitea-tea drives the official `tea` CLI for issues, pull requests,
labels and releases, reading each repository's own issue templates and
label set before drafting rather than assuming a shape. It is the pair
to create-pr: their descriptions hand off to each other, so a request
reaches whichever one can talk to the host in question.

From the owner's public MaksymStoianov/skills, Apache-2.0, installed
with the installer so skills-lock.json stays the source of truth.

The installer left the same three things behind that AGENTS.md now
documents — the stray `agent/` directory, a `skills/<name>` symlink in
the directory for skills authored here, and an EPERM on the Claude Code
copy under the Bash sandbox — and they were cleaned up the same way.

Closes #14

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@MaksymStoianov MaksymStoianov changed the title Пятислойное тестовое окружение и исправление девяти дефектов, которые оно нашло Five-layer test environment, and the twelve defects it found Sep 22, 2026
MaksymStoianov and others added 2 commits September 22, 2026 13:24
Every file, commit message and document here is in English. Thirteen
issues and one pull request were not, because the conversation that
produced them was in Russian and nothing on disk said otherwise. On a
public repository that is a real cost: the skills, the README and the
commit log address an English-speaking audience, and the tracker
explaining why those skills look the way they do did not. All fourteen
have been rewritten; rule 12 is what stops it recurring.

The same rule covers a second finding from a pre-publication check.
AGENTS.md and two issue bodies named a private repository of the owner's
by its full path and described its contents, and one named the Gitea
mirror's internal hostname. Nothing there leaks a credential or anyone
else's confidential material, but none of it was load-bearing either:
rule 13 works without the name — a proprietary skill must not be copied
into a public Apache-2.0 tree — and a remote is configured locally, so
no issue needs to name a host. Both now say "a private repository".

Closes #15
Closes #16

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ISSUES-FOUND.md was written when `gh` appeared to be broken and issues
could not be opened. They can, and #1-#10 now carry all of it, so a
static list of this repository's own defects sitting in a public tree is
a stale duplicate of its tracker — which AGENTS.md rule 2 makes the
authoritative record. The commits that mention the file are history and
remain accurate as history; the one place that referenced it as a live
path no longer does.

Also drops `--experimental-strip-types` from the two npm scripts that
carried it. Node 24 strips types without the flag, verified by running
the script, and every other script here already omitted it.

Closes #17

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MaksymStoianov and others added 2 commits September 22, 2026 13:38
The suite existed and nothing ran it. A public repository with 322 tests
had no checks on its pull requests, so the contract layer only caught
what someone remembered to run locally — which is most of the point of
having it gone.

Runs `npm ci`, `tsc --noEmit`, `npm test`, the behaviour-tree drift
check, and `npm run test:prove`. Python is installed because the shared
fetching policy under test is Python, reached through
scripts/testkit/bridge.py by the unit and integration layers.

The behaviour layer is deliberately absent. It spends money and needs a
model credential, so it stays a release gate run by hand, as AGENTS.md
rule 9 says; CI verifies only that the committed case tree still matches
the eval sets it was generated from.

test:prove is included on purpose. A green suite is only evidence if red
is reachable, and the cheapest place to discover a check has stopped
firing is the pull request that broke it.

Closes #18

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first run annotated all three actions: they target Node 20, which
GitHub has deprecated and is already force-running on Node 24. Pinning
to the current majors takes the warning away now rather than when the
runners stop compensating.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@MaksymStoianov
MaksymStoianov merged commit 5e912ca into main Sep 22, 2026
1 check passed
@MaksymStoianov
MaksymStoianov deleted the test-environment branch September 22, 2026 11:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment