An agent skill for independent, unbiased review of completed work — code, PRs, diffs, files, whole projects, specs, or configs — using a clean-context subagent so the implementing agent's bias never leaks into the verdict.
The agent that did the work cannot reliably audit it. Same-context "please review what we just built" produces motivated-reasoning reviews: the model has already invested in the implementation and tends to defend it. This is the checkbox theater problem.
This skill solves it by delegating the review to a fresh subagent via Cursor's Task tool. The subagent has no implementation history, no prior justifications, and no anchoring bias. The parent agent's only job is to identify the review target, gather a comprehensive neutral context package, and pass it through.
| Target | Example trigger |
|---|---|
agent-changes (default) |
"review what you just did" |
diff |
"review HEAD vs main", "review last 3 commits" |
pr |
"review PR #482", a GitHub PR URL |
files |
"review src/auth/", "review these files: a.ts, b.ts" |
project |
"audit this codebase", "project health review" |
spec |
"review this spec at docs/sso.md" |
config |
"review CI config", "review my Dockerfile" |
flowchart LR
A[Identify target] --> B[Gather target-specific context]
B --> C[Build neutral brief]
C --> D[Launch Task subagent]
D --> E[Subagent reads CRITERIA.md and OUTPUT_FORMAT.md]
E --> F[Subagent investigates with read-only tools]
F --> G[Devil's-advocate self-pass]
G --> H[Return structured verdict]
H --> I[Present verbatim to user]
Key design choices:
- 13 review dimensions tagged with applicable targets (e.g.
Spec Qualityonly fires onspec,Architectureonly onproject/files) so a PR review doesn't inherit project-audit noise. - Evidence rule: every finding must cite
file:line, command output, or a diff hunk. No evidence → goes toUnverified claims. - Confidence scoring (0–100): borrowed from Anthropic's official code-review plugin. Threshold ≥80 to surface as a finding; lower findings are reported separately with a missing-evidence note.
- Devil's-advocate self-pass: the subagent challenges every finding (KEEP / WEAKEN / DROP) and runs a gap pass even when zero issues were found.
- Mechanical verdict: APPROVE / APPROVE_WITH_CHANGES / REQUEST_CHANGES / BLOCK is derived from severity counts, not from a feeling.
git clone https://github.com/koldovsky/skill-review-task.git ~/.cursor/skills/review-taskThat's it. Cursor auto-discovers skills in ~/.cursor/skills/. The skill is disable-model-invocation: true, so it activates only when the user explicitly asks for a review.
Symlink (or copy) into ~/.claude/skills/:
git clone https://github.com/koldovsky/skill-review-task.git ~/.claude/skills/review-taskThen either:
- Update the path references in
SKILL.mdfrom~/.cursor/skills/review-task/...to~/.claude/skills/review-task/..., or - Rely on the parent agent to substitute the absolute path when launching the subagent (the skill's prompt template explicitly accommodates this).
cd ~/.cursor/skills/review-task && git pullJust ask in plain English:
- "review what you just did"
- "review PR #482"
- "audit this codebase"
- "review the spec at docs/sso.md"
- "review my Dockerfile"
The skill will identify the target, gather context, delegate to a fresh Task subagent, and present a structured verdict.
Modes are orthogonal to target type. State the mode in your request.
| Mode | Effect |
|---|---|
comprehensive (default) |
All applicable dimensions |
security-only |
Only the Security dimension |
completeness-only |
Only Completeness vs Requirements |
quick |
Skip the gap pass; only surface findings with confidence ≥ 90 |
architecture-only |
Architecture & Maintainability + Project Health (best for project) |
Example: "review PR #482 in security-only mode"
The subagent emits a strict markdown structure (full schema in OUTPUT_FORMAT.md):
## Review target
<target type> — <one-line scope>
## Verdict
APPROVE | APPROVE_WITH_CHANGES | REQUEST_CHANGES | BLOCK
## Risk score
<1-10>
## Findings
### F1 — <title>
- Severity, Category, Confidence, Evidence, Why it matters, Recommended fix, DA verdict
## Unverified claims (confidence < 80)
## Gap pass (things I almost missed)
## Next steps| File | Purpose |
|---|---|
| SKILL.md | Entry point — parent-agent workflow, target identification, per-target context playbook, subagent prompt template |
| CRITERIA.md | 13 review dimensions, target applicability matrix, severity rubric, confidence scoring, devil's-advocate pass |
| OUTPUT_FORMAT.md | Rigid markdown schema for subagent output |
The skill is intentionally generic. To adapt it to your project's conventions, prefer runtime customization over forking:
- Maintain
AGENTS.md,.cursor/rules/, or equivalent in your project. The subagent reads these and applies them as part ofCode Quality & Conventions. - Document your project's verification commands (
npm test,pnpm lint, etc.) in your README orAGENTS.md— the parent agent picks them up automatically.
If you need project-specific dimensions, fork and extend CRITERIA.md. The dimension table at the top is the contract; add new rows and tag them with applicable targets.
- Agentic AI Code Review: From Confidently Wrong to Evidence-Based — the agentic-loop + terminal-tool pattern
- agent-review-orchestrator — devil's-advocate verification layer
- Anthropic's code-review plugin — confidence scoring rubric
- Checkbox Theater — why same-context self-review fails
MIT — see LICENSE.