Thanks for considering a contribution. This document covers the PR process, per-file-type guidelines, and how to test changes locally.
- Project goals and non-goals
- Repo layout
- The manifest
- Threshold re-calibration
- PR process
- Per-file-type guidelines
- Local install test
- Hook test snippets
- Style conventions
Goals
- Spec-driven workflows on top of Claude Code: every non-trivial change starts with a spec.
- Stack-agnostic: works for .NET, Node, Python, Go, Rust, etc. No language assumptions in agents.
- Cross-platform: PowerShell (Windows) and bash (Unix/macOS) parity.
- Cost-aware: heavy reasoning on
sonnet, mechanical work onhaiku.
Non-goals
- Not an IDE plugin or VSCode extension. Lives in
~/.claude/. - Not a replacement for
gitworkflow tooling. - Not a code formatter or linter wrapper.
specwright/
commands/ # 14 slash commands (markdown with frontmatter)
agents/ # 6 subagent definitions (markdown with frontmatter)
hooks/
powershell/ # 3 PowerShell hooks
bash/ # 3 bash hooks (parity with PowerShell)
templates/ # 4 setup templates
specs/ # 6 spec templates
install/ # install.ps1 + install.sh + install/README.md
docs/ # architecture, usage, walkthrough, troubleshooting
examples/ # demo references
specwright.manifest.json is the canonical inventory contract. Check 7 of
scripts/validate.{ps1,sh} reads it and fails the build when a number published in the docs
disagrees with what is actually on disk. A discipline tool that misdescribes itself has no
standing to lecture anyone about specs.
The manifest stores no counts. It declares where assets live (areas, each with a glob or
an explicit files list) and where the docs make claims about them (docClaims); the numbers are
derived from disk at runtime. That is deliberate - a manifest holding hardcoded counts would be a
third place to update on every change and would reintroduce exactly the drift it exists to prevent.
What this means in practice:
- Adding a command, agent, skill, or template: add the file. Nothing else. The count follows.
- Publishing a number in the docs: add a
docClaimsentry -file, apatternwith exactly one capture group around the number, and theequalsquantity it must match. A number with no entry fails the build as an undeclared claim, so this is not optional. - Rewording a sentence that carries a number: update its
patterntoo. A pattern that matches nothing fails as a vacuous claim rather than passing quietly - otherwise a reword would turn the check into a no-op that still reports green. - Writing intentionally historical docs (superseded counts as a past-state record): put the
path in
historicalExclusions.docs/history/,docs/superpowers/andCHANGELOG.mdare already excluded. Never "fix" their numbers to match today's disk state. - Publishing the current released version: add a
versionClaimsentry (file+ apatternwith one capture group, noequals). It is checked against the newest dated## [x.y.z] - <date>heading inCHANGELOG.md, not against a manifest quantity - CHANGELOG is the single source of truth for "what version is released." UnlikedocClaims, there is no undeclared-claim scan for version strings yet: a version claim in a doc that is not listed here is not caught.
Two constraints on pattern: it must be valid in both POSIX ERE (bash [[ =~ ]]) and .NET
(PowerShell), so use [0-9] rather than \d and avoid lookarounds; and it is matched
case-sensitively on both platforms. This applies to versionClaims patterns too.
scripts/selftest-docs.{ps1,sh} proves Check 7 still bites, by corrupting a throwaway copy of the
repo and asserting the validator catches it. CI runs it on all three OSes.
tests/hooks/run-conformance.ps1 (single cross-platform pwsh script by design - it must run BOTH
hook implementations in one process, so a bash twin would itself be a drift risk) pipes every
golden fixture under tests/hooks/fixtures/ into the bash and PowerShell implementation of each
hook and fails if their normalized decisions diverge from each other or from the golden. Add a
fixture case whenever you add hook behavior; -SelfTest proves the harness still detects
divergence.
Check 7 needs jq on Unix and fails loudly without it. This is the opposite of the hook rule
below (hooks exit 0 silently when jq is missing so they never block a user on their own bugs) -
a validator that skipped itself for a missing tool would turn CI green while checking nothing.
Where Check 7 guards inventory, Check 8 guards the relationships between the prompt files:
which agent a command invokes, which skill an agent loads, which template a prompt reads, how many
hard gates a workflow declares. It is a script, not a prompt - scripts/contract-lint.{ps1,sh},
configured entirely from the manifest's contractLint subtree. Full rule catalogue and rationale:
docs/contract-lint.md.
Run it directly while iterating:
bash scripts/contract-lint.sh --root ..\scripts\contract-lint.ps1 -Root .Exit 0 means no BLOCK findings, 1 means at least one, and 2 means it could not run at all
(missing manifest, missing jq, or the registry parity guard tripped). Check 8 treats 2 as a
failure for the same reason Check 7 refuses to skip itself.
Suppressing a finding. Rarely, a violation is correct on purpose. Put a comment on the offending line or the line above it, naming the rule and giving a real reason:
<!-- contract-lint: allow CL305 - the option here buys a logged constitution exception rather than a way past the requirement -->
Three things constrain that escape hatch, and all three are enforced:
- The reason is mandatory. Under ten non-separator characters fails as CL900. "
- x" is not a reason. - The rule id must exist. A typo fails as CL901 rather than silently suppressing nothing.
- It must actually suppress something. A suppression that outlives the finding it was written for fails as CL902 - the same anti-rot posture as Check 7's vacuous-claim rule.
A suppression can never suppress CL900, CL901 or CL902; that would be a self-authorizing loophole.
Adding a rule means four edits, and skipping any one of them fails CI: a contractLint.rules
registry entry, a rule function in both implementations, a fixture case under
tests/contract-lint/ whose expected.json names the rule, and a row in docs/contract-lint.md.
Each edge of that square is guarded by a different mechanism - the linters' own registry parity
guard, and invariants C and D in tests/contract-lint/run-selftest.ps1.
tests/contract-lint/run-selftest.ps1 is the fixture suite. Like the hook conformance harness it is
a single pwsh script by design: it runs both implementations in one process, so parity is asserted
rather than inferred. -SelfTest swaps in a linter that reports nothing and asserts the harness
notices.
Every hardcoded threshold in this repo (Gate Complexity's tasks/layers/files limits,
retroStaleMinutes, debounceMinutes, maxLessons, metrics.maxSizeKb, the perf gate's noise
floor) started as an estimate, not a measurement - see docs/adr/0004-threshold-calibration.md.
Re-run the calibration pass every 20 closed specs, or at each minor release, whichever comes
first:
- Run
/sd:status --calibrationagainst the accumulated.specs/index.mdand.specs/_metrics/events.jsonl. - For each threshold, record the verdict - keep, change, or insufficient data - in a new ADR under
docs/adr/. "Insufficient data" is a legitimate, expected outcome at a thin corpus size; do not change a threshold without a stated measurement behind it. - Where a threshold's rationale in
templates/project-config.template.jsonis still a judgement call (no measured basis), leave its_..._usecaveat in place rather than removing it.
-
Open an issue first for anything larger than a typo or a small docs fix. State:
- What problem the change solves.
- Which files are affected.
- Whether it is a breaking change.
-
Branch from
mainwith a descriptive name:feat/<slug>for new commands / agents / hooks.fix/<slug>for bug fixes.docs/<slug>for docs-only changes.refactor/<slug>for internal restructuring.
-
Keep PRs focused. One workflow, one agent, one hook per PR. A 1500-line "improve everything" PR will be asked to split.
-
Run the validator before opening the PR:
scripts/validate.ps1(Windows) orscripts/validate.sh(Unix) runs every engine-invariant check at once (ASCII, hook-pair parity, model aliases, install-target counts, changelog gate, docs consistency). CI runs the same on Windows + Ubuntu. See also the Local install test for a manual install smoke test, and The manifest for what Check 7 enforces. -
Update the changelog. Add a line under
## [Unreleased]inCHANGELOG.md. -
Open the PR with:
- A short description.
- Screenshots or terminal output if behaviour changes.
- A note on whether docs were updated.
A command is a markdown file with YAML frontmatter that Claude Code reads when the user types /sd:<name>.
Required structure:
---
description: One-line summary shown in /help
argument-hint: <ID or slug>
---
# /sd:<name>
## Phase 0 - Bootstrap
- Read CLAUDE.md
- Read .specs/constitution.md
- Read .claude/project-config.json
- Detect state (resumable?)
## Phase 1 - <name>
... (with hard gates marked as Gate N)
## Rules
- Hard constraints that cannot be bypassed.Conventions:
- Phase 0 always bootstraps; do not skip.
- Hard gates use the explicit marker
Gate Nand prose "STOP. Wait for explicit user approval." - State machine documented at top of file (what happens on re-invocation).
- Subagent invocation uses the
sd-prefix, never bare names. - No stack-specific commands inside (no
dotnet test,npm test, etc.). Reference thecommands.testfield fromproject-config.json.
A subagent is a markdown file with YAML frontmatter consumed by the Task tool.
Required frontmatter:
---
name: sd-<role>
color: <color> # e.g. cyan, orange, purple, green, blue - used for display only
description: One-line summary used by routing.
model: sonnet # MUST be an alias: sonnet | haiku | opus | inherit
tools: Read, Grep, Glob, ... # MINIMAL allowlist
skills:
- sd-<shared-rule-pack> # any skill this agent's body references; see Skills below
---Critical:
model:MUST be an alias. Full IDs likeclaude-sonnet-4-7are not portable and may not even exist. The aliassonnetauto-resolves to the latest Sonnet.tools:should be the minimum set the agent needs. Read-only agents do not getWrite. Implementer does not getWebSearch.skills:must list every skill the agent body references (**skill-name**in prose). A rule used by multiple agents lives in oneSKILL.md, never copy-pasted into agent bodies.- Agent must read
CLAUDE.mdandconstitution.mdat runtime. No hardcoded stack assumptions (nocs,csproj,dotnet, etc. literal references unless they come from project config). - Every finding cites
file:line. No prose without citations.
Machine-readable input declarations. Under every heading that selects a distinct agent
behavior by field value (## Mode N: TASK = <mode>, ## Task type: , `### `TASK = <type>, ### WORKFLOW_TYPE = ``), add two lines immediately before the existing prose:
Inputs (required): SPEC, IMPACT
Inputs (optional): MODE, REPLAN_SCOPE
- Tokens are
UPPER_SNAKE, comma-separated, exactly the identifiers the calling command sets - never a paraphrase. - Write
noneexplicitly when a mode has no required (or no optional) inputs. Silence is not an assertion - an omitted line reads as "not yet documented," not as "empty." - The existing prose
Inputs: ...line (with parenthetical caveats, cross-references, etc.) stays below unchanged - the two new lines are additive, for tooling to grep, not a replacement for the explanatory prose. - When adding or changing an agent invocation in
commands/*.md, cross-check it against the target mode's declared inputs: a token the command passes that the mode doesn't declare, or a required token the mode declares that the command omits, is a real defect - fix the mismatch (add the missing token to the declaration or the invocation, whichever is actually correct) rather than leaving the two out of sync. - This cross-check is machine-enforced: Check 8's
CL100-CL104rules (docs/contract-lint.md) parse both sides and BLOCK on a mode mismatch or a missing required input. A legitimate mismatch the declaration can't express (an either/or required set, for example) gets a<!-- contract-lint: allow CLxxx - <reason> -->suppression, not a silent gap.
Hooks come in pairs. If you change hooks/powershell/foo.ps1, you also update hooks/bash/foo.sh. The repo CI runs scripts/validate.{ps1,sh}, which refuses PRs where the pair drifts (along with the other engine-invariant checks).
PowerShell (*.ps1) - critical encoding rule:
PowerShell 5.1 reads UTF-8 without BOM as Windows-1252. The em-dash - (U+2014) is byte sequence E2 80 94; PowerShell sees the final byte 0x94 as a curly closing quote, and your string terminates in the middle of nowhere. Result: cascading parse errors like Missing closing '}'.
Pure ASCII only in *.ps1 files. Substitute as follows:
| Forbidden | Use instead |
|---|---|
- (em-dash U+2014) |
- (ASCII hyphen-minus) |
-> (right arrow) |
-> |
> (triangular bullet) |
> |
OK: |
OK: or [OK] |
WARN: |
WARN: or [WARN] |
X |
X or [FAIL] |
i |
i or [INFO] |
Verify before committing:
grep -nP "[^\x00-\x7F]" hooks/powershell/*.ps1 install/install.ps1
# Empty output = OK. Any output = REJECT, fix it.Bash (*.sh):
- Shebang:
#!/usr/bin/env bash - Use
jqfor JSON; ifjqis missing, exit0silently (do not block the user). - Use
stat -c %Y(Linux) ANDstat -f %m(macOS) - detect and branch. - Set
chmod +xon commit (or rely on the installer to do it). - Test with
bash -n hooks/bash/foo.shto catch syntax errors before runtime.
Templates are filled by /sd:setup and by agents.
- Use
<<placeholder>>syntax for fields the user fills manually. - Keep templates short and scannable. A
CLAUDE.mdtemplate that ends up being 600 lines defeats the purpose. - Spec templates leave the cross-phase fields explicitly empty with a comment like
<!-- Filled by Phase 3 - do not pre-fill -->. This is intentional: workflows enforce sequencing through empty fields.
Before opening any PR that touches install, hooks, commands, or agents:
Windows (PowerShell):
# 1. From the repo root, dry-run first
.\install\install.ps1 -DryRun
# 2. Real install to a sandbox base path
.\install\install.ps1 -BasePath C:\temp\sd-test
# 3. Verify
Get-ChildItem C:\temp\sd-test\commands\sd\
Get-ChildItem C:\temp\sd-test\agents\sd\
Get-ChildItem C:\temp\sd-test\hooks\sd\
Get-ChildItem C:\temp\sd-test\templates\sd\
# 4. Cleanup
Remove-Item -Recurse -Force C:\temp\sd-testUnix/macOS (bash):
# 1. Dry-run
./install/install.sh --dry-run
# 2. Real install to a sandbox base path
./install/install.sh --base-path /tmp/sd-test
# 3. Verify
ls /tmp/sd-test/commands/sd/
ls /tmp/sd-test/agents/sd/
ls /tmp/sd-test/hooks/sd/
ls /tmp/sd-test/templates/sd/
# Verify hook executable bits
test -x /tmp/sd-test/hooks/sd/prompt-router.sh && echo "OK"
# 4. Cleanup
rm -rf /tmp/sd-testHooks read JSON from stdin. You can simulate Claude Code locally:
prompt-router (UserPromptSubmit):
echo '{"prompt":"fix bug INV-2501 in stock service","cwd":"/path/to/repo"}' \
| bash hooks/bash/prompt-router.sh'{"prompt":"fix bug INV-2501 in stock service","cwd":"C:\\path\\to\\repo"}' `
| powershell -File hooks\powershell\prompt-router.ps1spec-gate (PreToolUse):
echo '{"tool_name":"Edit","tool_input":{"file_path":"src/foo.cs"},"cwd":"/path/to/repo"}' \
| bash hooks/bash/spec-gate.shsubagent-retro (SubagentStop):
echo '{"cwd":"/path/to/repo","session_id":"test-session-001"}' \
| bash hooks/bash/subagent-retro.shExpected behaviour: every hook exits 0 and either prints a <context-router> / <retro-reminder> block to stdout, prints a warning to stderr, or stays silent.
- Markdown: ATX headers (
#,##), no trailing colons in headers, fenced code blocks with language hint, 100-char soft wrap. - PowerShell: PascalCase function names,
$camelCasevariables, explicitparam()block, pure ASCII. - Bash: lowercase function names,
snake_casevariables,set -euo pipefailat top of non-trivial scripts. - YAML frontmatter: keys in lowercase-with-hyphens (
argument-hint), values unquoted unless they contain special chars. - Commit messages: imperative mood, 50-char subject, optional body wrapped at 72.
Thanks for contributing. If anything in this doc is unclear, open an issue and we'll improve it.