Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 18 additions & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
"displayName": "Agentic Coding Loop",
"source": "./",
"strict": false,
"version": "0.8.0",
"version": "0.10.0",
"description": "Run the improvement loop in a repository: bootstrap the LOOP_STATE.md spine, run a falsifiable round against it, audit a claim, or send a learning upstream. The protocol is LOOP.md; every skill inlines it.",
"homepage": "https://github.com/max-friedman/agentic-coding-loop",
"repository": "https://github.com/max-friedman/agentic-coding-loop",
Expand All @@ -23,6 +23,23 @@
},
"category": "workflow",
"tags": ["loop", "planning", "long-running", "state", "governance"]
},
{
"name": "loop-ux-roast",
"displayName": "Agentic Coding Loop — UX Roast domain",
"source": "./domains/ux-roast",
"strict": false,
"version": "0.1.0",
"description": "Optional domain layer for the agentic coding loop, for user-facing products doing continuous UX-driven improvement: a coverage map keyed by surface, verification at scale for multi-complaint roast passes, and parallelized maker+checker fixes for disjoint findings. Additive to core §E, never a replacement; requires the core `loop` plugin.",
"homepage": "https://github.com/max-friedman/agentic-coding-loop",
"repository": "https://github.com/max-friedman/agentic-coding-loop",
"license": "MIT",
"author": {
"name": "Max Friedman",
"url": "https://github.com/max-friedman"
},
"category": "workflow",
"tags": ["loop", "ux", "domain", "roast", "product"]
}
]
}
71 changes: 71 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,77 @@ Versioning is [semantic](https://semver.org): MAJOR for a change to the protocol
that existing state files or rounds must adapt to, MINOR for new capability, PATCH
for wording and fixes.

## [0.10.0] — 2026-07-27

Optional domains: a project can layer project-shaped rules onto the core
protocol without the core protocol ever naming one.

**Added**
- **Domain discovery.** A Domains table in `llms.txt`, a domain-check step in §B
Bootstrap (step 2) and §0 Preconditions (step 1), and a `**Layers:**` field in
`templates/LOOP_STATE.template.md`. An agent adopting the loop reads one table
and decides whether a domain fits — no plugin install required just to
discover one, and no forcing a fit when none does.
- **`domains/ux-roast`** — the first domain, shipped as its own marketplace
plugin (`loop-ux-roast`, versioned independently of `loop`), additive to core
via a symlinked `LOOP.md` and its own `DOMAIN.md`: parallelized maker+checker
fixes for disjoint findings, and a coverage map keyed by user-facing surface.
Deliberately does not redefine roasting or verification — §E and its step 2
already cover that generically as of 0.9.0; the domain only adds what's still
genuinely project-shaped on top.

**Why**
- The review rubric's own posture (`docs/REVIEW_RUBRIC.md`, hard disqualifier 6)
is that project-specific workflow belongs in a project's own rules, not a
proposal to the shared protocol every project must read. A domain plugin is
the mechanism for exactly that gap: rules real enough to be worth sharing, but
shaped for one kind of project rather than every project running the loop.
Nothing here is a proposal to LOOP.md's required steps; a project that never
opts in behaves exactly as before.
- Explicitly excluded from the domain: an unattended auto-merge tier without a
separate review session. §D's current MUST (0.7.0) is that a round never
merges itself — that is a stop condition's cousin, not a project-specific
workflow choice, and disqualifier 1 rejects weakening a MUST "as an escape
hatch" even when the escape hatch is opt-in. A project that wants this stays
on it as a documented local divergence, not a shared plugin setting.

**Blast radius**
- Zero for a project that never fetches a domain — llms.txt gains one table, and
§B/§0 gain a check that resolves to "no match, proceed core-only" immediately
for any project without a fitting domain.

## [0.9.0] — 2026-07-27

A roast can cite something real and still be wrong about it. Now it has to survive
a second check before it reaches the queue.

**Added**
- **§E step 2 — Verify against ground truth.** Every complaint that survives
step 1's citation requirement is checked against application state, logs, or a
second run, and tagged `real`, `critic-mistake`, or `environment-artifact`. Only
`real` complaints shape the verdict (now step 3) or reach the queue.
- **`docs/PRINCIPLES.md` §11** — citing something is not the same as diagnosing it
correctly, with the anti-laundering guardrail: ground truth may correct or drop
an unreal complaint, never explain away one a real user would still hit.
- `templates/ROAST_LOG.template.md`'s complaint table gained a `verified` column.

**Why**
- §E's existing rule — every complaint must cite something a user could hit —
catches fabrication. It does not catch misattribution: a critic can genuinely
see a real screen and still guess wrong about why it looks wrong, or be looking
at a test-harness artifact no real user could reach. [Proposal
003](proposals/003-roast-findings-need-verification.md) evidenced a recurring
pattern of exactly this across many rounds of one project's history, predating
this repository's own §E.
- Verification is placed *before* the verdict, not after, so a complaint later
found to be an environment-artifact never shapes what the roast claims a real
user would experience.

**Blast radius**
- One additional step per roast, one additional table column. Near-zero cost for
a roaster that already has ground-truth access while roasting; this is the step
that catches the failure for a roaster deliberately kept blind to internals.

## [0.8.0] — 2026-07-28

An empty queue can now refill itself — **opt-in, and off by default**.
Expand Down
79 changes: 59 additions & 20 deletions LOOP.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,11 @@ That file is the memory. Read it first, write it last.
## 0. Preconditions

1. Read `docs/plans/LOOP_STATE.md` in full, including the bottom sections. The
standing invariants are at the bottom and are the most important part.
standing invariants are at the bottom and are the most important part. If
`## Current status` declares a `**Layers:**` beyond core, fetch that domain's
`DOMAIN.md` (see §B step 2) and apply it for the rest of this round — the
declaration was made once at bootstrap precisely so later rounds don't re-derive
it.
2. Read the project rules (`CLAUDE.md` or `AGENTS.md`) if present.
3. If `docs/plans/LOOP_STATE.md` does not exist, run **§B Bootstrap** instead of a
round, then stop.
Expand Down Expand Up @@ -216,23 +220,31 @@ Verdict is one of: `holds`, `fails`, `unmeasurable as stated`.
## §B Bootstrap (no state file yet)

1. If `docs/plans/LOOP_STATE.md` exists, STOP. Never overwrite it.
2. Read enough to fill it honestly: README and purpose docs; the exact gate command
2. Check for a domain fit before reading further. Fetch the Domains table in
`llms.txt` (`https://raw.githubusercontent.com/max-friedman/agentic-coding-loop/main/llms.txt`)
and compare its descriptions against the target project. A clear match → fetch
that domain's `DOMAIN.md` too and layer it in for the rest of bootstrap and every
round after. No clear match, or it's a stretch → proceed core-only. Record the
choice as a `**Layers:**` line in `## Current status` — never guess silently and
never force a fit. A domain is additive to the steps above; it never replaces one.
3. Read enough to fill it honestly: README and purpose docs; the exact gate command
(from `package.json`, `pyproject.toml`, `Makefile`, CI config); tests asserting
*properties* rather than behavior (candidate invariants); TODO/FIXME clusters and
known-weakness sections (candidate queue items).
3. Write `docs/plans/LOOP_STATE.md` with every section from step 6 above. Round 0.
4. Write `docs/plans/LOOP_STATE.md` with every section from step 6 above. Round 0.
Headline is often "unmeasured" — say so.
4. Add to the project rules file, at the very top:
5. Add to the project rules file, at the very top:

```markdown
**Working the improvement loop? Read [`docs/plans/LOOP_STATE.md`](docs/plans/LOOP_STATE.md)
first and write it last.** It holds the queue, the coverage map, the NEEDS-MAX list,
and the standing invariants. Context is lost between rounds; that file is not.
```

5. Write only rules you can justify from code you read. An invented rule is worse
6. Write only rules you can justify from code you read. An invented rule is worse
than no rule — it gets cited later as if load-bearing.
6. Report the drafted queue and invariants, and say which sections you guessed at.
7. Report the drafted queue and invariants, which domain (if any) was matched, and
say which sections you guessed at.

---

Expand Down Expand Up @@ -496,29 +508,56 @@ read only the docs a user would actually find.
Where the roast and the project's own docs disagree, the roast is the evidence —
the docs were written by whoever built the thing.

### 2. Write the verdict before the fixes
### 2. Verify against ground truth

One honest paragraph in the user's voice: what this is, whether it did the job,
and what you would say about it to someone considering it. Write it before
proposing a single improvement — a verdict written afterward is shaped to justify
the fixes already in mind.
Every complaint that survived step 1's citation requirement still gets checked
against what the blind pass didn't have — application state, logs, or a second
independent run — and tagged one of three ways:

- **Real.** The cause is correctly attributed, and a user could genuinely
encounter it.
- **Critic-mistake.** Something was seen, but the roast misattributed the cause.
- **Environment-artifact.** The cause is real but a user would never hit it — a
test-harness quirk, leftover state from a previous pass, tooling noise. Not
something this product does to a user; something this *check* did to itself.

Citing something you saw is not the same as diagnosing it correctly. A critic can
genuinely observe a real screen and still guess wrong about why it looks wrong —
step 1's bar catches fabrication, not misattribution.

**The guardrail this step must not become:** ground-truth access explaining away a
complaint a real user would still experience, just because the internal cause
differs from what the critic guessed. Only *environment-artifact* — a cause a user
could never reach — drops a complaint. A *critic-mistake* keeps the observation
and corrects the cause; it does not disappear.

Only **real** complaints proceed to step 3's verdict and step 5's queue.
Critic-mistakes and environment-artifacts are recorded in the roast log with their
disposition, so the same false alarm is not re-diagnosed from scratch next time.

### 3. Write the verdict before the fixes

One honest paragraph in the user's voice, built from the **real** complaints only:
what this is, whether it did the job, and what you would say about it to someone
considering it. Write it before proposing a single improvement — a verdict written
afterward is shaped to justify the fixes already in mind.

Be specific and unkind. "The quickstart is confusing" is not a complaint. "The
quickstart's first command fails because it assumes a config file that step 4
creates" is.

### 3. Deduplicate against the roast log
### 4. Deduplicate against the roast log

Read `docs/plans/ROAST_LOG.md`. A complaint already recorded there and deliberately
not queued MUST NOT be re-queued without **new** evidence — say what changed.

Without this, an indefinite loop cycles on the same three complaints forever and
reports motion.

### 4. Convert only what is falsifiable
### 5. Convert only what is falsifiable

Each surviving complaint faces one test: **can it be phrased as a question with a
measurement that could come back bad?**
Each surviving **real** complaint faces one test: **can it be phrased as a
question with a measurement that could come back bad?**

- Passes → queue it, in §2's form: the question, and what a negative result looks
like. It is now an ordinary queue item and the next round is an ordinary round.
Expand All @@ -528,7 +567,7 @@ measurement that could come back bad?**
A complaint that cannot be made falsifiable is not thereby wrong. It is not yet a
round.

### 5. Write the roast log
### 6. Write the roast log

Append to `docs/plans/ROAST_LOG.md` — create it from
[`templates/ROAST_LOG.template.md`](templates/ROAST_LOG.template.md) if absent.
Expand All @@ -539,17 +578,17 @@ edited afterward.
## Roast N — <date> — <one-line verdict>

**Ran as:** the journey actually performed — commands, entry points, what a user was assumed to want.
**Verdict:** the honest paragraph from step 2.
**Complaints:** table of complaint | evidence cited | falsifiable | disposition.
**Verdict:** the honest paragraph from step 3, built from real complaints only.
**Complaints:** table of complaint | evidence cited | verified (real / critic-mistake / environment-artifact) | falsifiable | disposition.
**Queued:** items added, each as a question.
**Noted, not queued:** complaints that could not be made falsifiable, with why.
**Noted, not queued:** complaints that could not be made falsifiable, or that verification found was a critic-mistake or environment-artifact, with why.
**New this roast:** how many complaints do not already appear above. Zero means stop.
```

Then update `## Queue — next rounds` in the state file and record the roast on the
current round's **Loop:** line. Never write `## Loop configuration`.

### 6. Decide whether to continue
### 7. Decide whether to continue

| condition | action |
|---|---|
Expand Down
6 changes: 5 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,7 +101,10 @@ force-pushes and deletions blocked.

Every skill inlines [`LOOP.md`](LOOP.md) at load time, so the protocol has exactly
one copy and no skill can drift from it. Other harnesses — any agent, any tool —
read `LOOP.md` directly; it is self-contained in a single fetch. See
read `LOOP.md` directly; it is self-contained in a single fetch. Then check the
**Domains** table in [`llms.txt`](llms.txt) — optional, additive layers (e.g. a
UX-roast domain for user-facing products) that a matching project should fetch
alongside `LOOP.md`, before bootstrapping. No match → proceed core-only. See
[`docs/ADOPTING.md`](docs/ADOPTING.md).

## The design decisions worth defending
Expand Down Expand Up @@ -138,6 +141,7 @@ The repo's credibility rests on this section being accurate rather than short.
|---|---|
| [`LOOP.md`](LOOP.md) | The protocol. Self-contained, one fetch, written to be executed. |
| [`skills/`](skills) | Six Claude Code skills; each inlines `LOOP.md`. |
| [`domains/`](domains) | Optional, additive layers for project-shaped rules — e.g. `ux-roast`. See `llms.txt`'s Domains table. |
| [`docs/REVIEW_RUBRIC.md`](docs/REVIEW_RUBRIC.md) | The standard a proposal must clear. Reject by default. |
| [`docs/PRINCIPLES.md`](docs/PRINCIPLES.md) | Each rule and the specific failure it prevents. |
| [`docs/CASE_STUDY.md`](docs/CASE_STUDY.md) | The rounds above, in full, including what they cost. |
Expand Down
26 changes: 23 additions & 3 deletions docs/ADOPTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,12 +53,31 @@ All six carry `when_to_use` triggers and are model-invocable — an agent picks
right one from intent ("run the loop", "does that claim still hold") without a human
typing a slash command.

**Updates** are pinned to an explicit `version`, so pushes to `main` do not reach
you until a release:
**Domain plugins are separate, optional installs** — layers on top of these same
six skills, for projects that fit a specific domain. Currently:

```
/plugin install loop-ux-roast@agentic-coding-loop
```

| skill | what it does |
|---|---|
| `loop-init-ux` | Bootstraps the state file with `Layers: core + ux-roast` declared, CUJ-framed queue, coverage map by surface. |
| `loop-roast-ux` | A roast round using core §E, specialized with verification-at-scale and parallelized maker+checker fixes for disjoint findings. |

Install this only if the target project is a user-facing product doing
continuous UX-driven improvement — see the Domains table in
[`../llms.txt`](../llms.txt) for the fit criteria. A domain's skills inline both
`LOOP.md` and its own `DOMAIN.md`, so core and domain stay separately versioned
and neither can drift from the other.

**Updates** are pinned to an explicit `version` per plugin, so pushes to `main` do
not reach you until a release:

```
/plugin marketplace update agentic-coding-loop
/plugin update loop
/plugin update loop-ux-roast
```

Pinning is deliberate. These instructions execute inside your repo; you choose when
Expand Down Expand Up @@ -124,10 +143,11 @@ claude plugin validate . --strict

`--strict` turns unrecognized-field warnings into errors, catching a misspelled
manifest key before it silently does nothing. To check discovery end to end, install
from a local path and confirm all six skills appear:
from a local path and confirm all six core skills (and any domain plugin's) appear:

```bash
claude plugin marketplace add /absolute/path/to/agentic-coding-loop
claude plugin install loop@agentic-coding-loop
claude plugin install loop-ux-roast@agentic-coding-loop
claude plugin list
```
32 changes: 31 additions & 1 deletion docs/PRINCIPLES.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Principles

Ten rules. Each one names a specific failure it prevents. None of them are
Eleven rules. Each one names a specific failure it prevents. None of them are
abstract — they were paid for.

Rules 9 and 10 arrived differently from the rest: they came from projects running
Expand Down Expand Up @@ -157,6 +157,36 @@ Prose that has already been ignored once will be ignored again.
a different step than proposed, on the strength of the submitter's own objection to
their primary placement.*

## 11. Citing something is not the same as diagnosing it correctly

A roast's one honesty rule — every complaint must cite something a user could
hit — catches fabrication. It does not catch misattribution: a critic can
genuinely see a real screen and still guess wrong about why it looks wrong, or be
looking at a test-harness artifact a real user could never reach.

The failure it prevents: **the confidently wrong complaint**. A project running an
informal predecessor of the roast round for months found a recurring minority of
complaints that passed the citation bar and were still wrong — a thing that looked
like two conflicting UI elements was two legitimate ones; a thing that looked like
data corruption was stale state left by a previous test pass; a thing that looked
like a dropped-input bug was an artifact of the harness driving the product, not
something a real user would ever hit. Every one satisfied "cite what you saw,"
because the critic genuinely did see it.

**The fix is a second check, not a stricter first one.** No citation requirement
distinguishes a correct diagnosis from a plausible-sounding wrong one; only
comparing the complaint against ground truth — state, logs, a second run — does.
The check may only correct or drop a complaint that turns out to be unreal. It may
never use internal knowledge to explain away a complaint a real user would still
experience — that would turn a truth check into a laundering step, which is the
one failure mode worse than not checking at all.

*Arrived as [proposal 003](../proposals/003-roast-findings-need-verification.md),
filed as evidence from a project's own history running that informal predecessor.
Accepted with the anti-laundering guardrail the submitter flagged against their own
proposal — the strongest objection to a proposal is sometimes the reason to keep it,
narrowed.*

---

## What these have in common
Expand Down
Loading
Loading