Skip to content

v3 roadmap: one instruction, one reviewer list, one owner of the wait #42

Description

@kristofferR

Tracking issue for the v3 work: turning the agent-facing review protocol from prose into commands, and scoping the queue to the reviewer that actually meters it.

Why

crq's protocol lived in documentation. llms.txt, SKILL.md and crq help loop each described the same rules — drain findings before starting a round, hold the head while a reviewer is pending, pick a delay — and drifted apart. Nine of fifteen merged PRs before this changed that prose rather than any decision.

Agents then improvised the parts crq did not own, and the improvisations were uniform. Across 16 real driving sessions, every single one wrapped the tool the same way:

set +e; crq loop OWNER/REPO N > /tmp/out.json; echo "CRQ_EXIT:$?"

An exit code smuggled out through an echoed sentinel into a temp file. Several runs were SIGTERM'd mid-wait or orphaned at a session teardown, and restarting was wrong both ways: it either re-fired a review (spending account quota) or hit the dedupe and reported a converged round whose findings were never collected.

Guiding invariant: anything in the docs that tells an agent when to wait, hold, push, or resolve is a bug in crq next.

The sizing rule

#39 took four review rounds and 52 findings, and every round after the first found regressions introduced by the previous round's fixes. The reviewers were right each time. The cause was size — 24 files, a new command, a reordered decision path and changed decline semantics all interacting, so each correction moved something else.

So, for everything below:

  1. One thesis per PR. If the description needs "and also", split it.
  2. Converge in ≤2 review rounds. A third round that finds a regression from the second means land what is verified and move the rest out.
  3. Prefer the visible failure. Replace the agent review protocol with one instruction and a killable wait #39's worst moment was suppressing a finding to break a deadlock. A repeated instruction is annoying; a swallowed finding is harmful.
  4. Verify before writing. Three of the first seven planned items turned out to be already fixed. Re-read scope against main before starting anything here.

Done

Released

crq 2.0.0 (00c377a) — everything above is on main. The state schema is now
v4: a binary that predates a schema refuses the payload rather than erasing it (#49),
so every host in a fleet has to be upgraded together. CRQ_STATE_REF keeps its
crq-state-v3 name — that string is where the state lives, not what version it is.

Still in review: fleet configuration in the state ref — #58. Scope, repos, exclude,
required bots and the timings move out of each host's environment, because per-host files
diverge silently: a repository was excluded on one machine and still reviewed by another,
and each host was behaving correctly according to what it could see.

Still open

Two things worth remembering about the work itself

Both cost several review rounds before being spotted:

  • When a claim keeps being wrong, consider deleting the claim. The queue's ordering was refined across four rounds — simulate it, fall back to Seq, drop positions while the slot is held — before it became clear that order past the front is not knowable at all, because slot release comes from acknowledgement or the in-flight timeout rather than pacing. Removing the prediction was less code than any of the refinements.
  • An exemption applied per-site will keep missing a site. "A co-only round is not governed by the queue" was fixed three times in three rounds — the account window, the ordering, the follower pass — before being made structural by partitioning those rounds out before any gate applies.

Ordering

The seven PRs in review are independent except for two stacks that have since
been flattened: #57 was rebased onto main after #51 merged, and #54 carries
#53's workspace. Merge order is otherwise free.

What is left after them has no coupling at all — the behaviour lab, per-bot
SeenActiveAt, declined re-review recovery, and the two small items split out
of #55 and #57 rather than widening either.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions