Skip to content

PRD: Claude Code baseline for the Solo Dev Starter Suite #49

Description

@Jordak

Problem Statement

Agent Eval Lab now has a working Claude Code agent adapter and a completed Codex starter-suite baseline, but it does not yet have a fair Claude Code baseline. Without that baseline, the project cannot responsibly compare Claude Code against Codex or make evidence-scoped recommendations about Claude Code as an agent harness.

From the user's perspective, the core problem is: "I want to know how Claude Code behaves on the same realistic Solo Dev Starter Suite tasks that Codex already ran, while keeping the experiment cheap, quick, and honest. I do not want to tune the suite for Claude Code, overclaim from a few smokes, or jump to Codex-vs-Claude comparison before Claude Code has its own baseline evidence."

This PRD follows the completed Codex baseline planning anchor in #10. The Codex baseline should close as a completed baseline PRD; this Claude Code baseline becomes the next active planning anchor. A future comparison PRD can refer back to both #10 and this Claude Code baseline PRD after both baselines exist.

Existing landed work this PRD builds on:

  • e1a818c - Add Claude Code harness pilot.
  • 454476c - Add model identity event fixtures.
  • f0aecba - Refactor CLI into command-family package.

Solution

Create a Claude Code baseline for the existing 12-task Solo Dev Starter Suite.

The baseline should evaluate one explicit Claude Code agent harness configuration, prioritizing proof-of-concept cost and speed. The selected Claude Code model should be the cheapest/fastest model available to Jordan's installed Claude Code CLI and account, with the exact alias or model identifier verified before Phase 1 begins.

The baseline should match the existing Codex baseline configuration and trial semantics wherever Claude Code has a fair equivalent. It should preserve the same task prompts, graders, reference artifacts, task-local environments, human-review vocabulary, and five-fair-trials-per-task target. It should not tune prompts, graders, or permissions to help Claude Code succeed.

The baseline proceeds in phases:

  1. Readiness/config verification before Phase 1.
  2. Phase 1 full-suite shakedown: one sequential Claude Code smoke trial per starter-suite task.
  3. Phase 2 repeated collection: four additional fair trials per task after a fair smoke, including tasks whose fair smoke failed the grader.
  4. Evidence artifacts: selected Claude Code evidence set and generated capability evidence digest.
  5. Hand-authored Claude Code baseline report, scoped only to Claude Code under the evaluated conditions.

The PRD itself should not be treated as one large ready-for-agent task. After Jordan approves this PRD, split implementation into child issues for readiness/config verification, Phase 1 full-suite shakedown, Phase 2 evidence collection, and the hand-authored baseline report.

User Stories

  1. As a technical evaluator, I want a Claude Code baseline separate from the Codex baseline, so that each agent harness has its own fair evidence before comparison.
  2. As a technical evaluator, I want the Claude Code baseline to use the same Solo Dev Starter Suite as Codex, so that later comparison can reason over comparable task contracts.
  3. As a technical evaluator, I want the baseline to preserve existing task prompts and graders, so that Claude Code is not evaluated on a tuned variant of the suite.
  4. As a technical evaluator, I want one explicit Claude Code agent harness configuration, so that the baseline measures a reproducible setup rather than an implicit local default.
  5. As a technical evaluator, I want the model request recorded separately from runtime model identity, so that requested and actual model metadata do not collapse into one field.
  6. As a maintainer, I want the exact cheap Claude model alias or identifier verified before trials begin, so that Phase 1 does not depend on stale model-name assumptions.
  7. As a maintainer, I want the Claude Code preflight to run before Phase 1, so that auth, executable discovery, CLI version, and print-mode command shape are checked up front.
  8. As a maintainer, I want max-turn behavior verified before Phase 1, so that the baseline does not accidentally measure an arbitrary Claude Code turn cap.
  9. As a maintainer, I want the Claude Code configuration to stay as close as practical to the Codex baseline, so that differences are caused by the agent harness rather than avoidable configuration drift.
  10. As a maintainer, I want Claude Code trials to use non-interactive print mode, so that the baseline fits the same non-interactive evaluation harness shape as Codex.
  11. As a maintainer, I want Claude Code trials to avoid session persistence, so that each trial remains isolated and comparable.
  12. As a maintainer, I want no custom allowed/disallowed tool rules by default, so that the trial surface is not tuned per task unless a suite-level fairness or safety issue requires it.
  13. As a maintainer, I want any added tool policy to be suite-level rather than task-specific, so that Claude Code does not receive hidden task-by-task help.
  14. As a maintainer, I want Phase 1 to run one smoke trial per starter task, so that task-specific harness problems are discovered before repeated collection.
  15. As a maintainer, I want Phase 1 smoke trials to run sequentially with one job, so that failures are easy to interpret and do not mix with concurrency issues.
  16. As a maintainer, I want fair smoke trials to count toward the five-trial baseline, so that useful evidence is not discarded.
  17. As a maintainer, I want invalid smoke trials to be treated as diagnostics, so that setup, auth, eval-harness, operator, task-definition, or dependency failures do not pollute capability metrics.
  18. As a maintainer, I want invalid smoke trials fixed and rerun before repeated collection for that task, so that Phase 2 does not scale unfair conditions.
  19. As a technical evaluator, I want fair failed smoke trials to continue into repeated collection, so that one failure becomes part of the reliability picture rather than an early stop.
  20. As a technical evaluator, I want five fair trials per task, so that pass rate, pass@k, pass^k, variance, and resource behavior remain comparable to the Codex baseline shape.
  21. As a technical evaluator, I want the cheap/fast model used even if it performs worse, so that this PRD remains a proof-of-concept baseline rather than a performance-maximizing showcase.
  22. As a maintainer, I want the option to pause after Phase 1 if cost is too high, so that the trial target is not weakened preemptively.
  23. As a human reviewer, I want to review the first fair trial for each task, so that every task gets a calibrated human quality check without requiring review of all 60 selected trials.
  24. As a human reviewer, I want later trials reviewed selectively when they fail, look suspicious, are resource-heavy, or may need validity judgment, so that review time is spent where it changes the evidence.
  25. As a human reviewer, I want Claude Code patches labeled with existing human review labels, so that Claude evidence remains comparable to the Codex evidence vocabulary.
  26. As a model-quality engineer, I want deterministic grader outcomes separated from human review outcomes, so that passing tests do not hide messy, risky, or resource-heavy behavior.
  27. As a model-quality engineer, I want selected evidence to exclude only invalid trials with explicit reasons, so that capability claims remain fair and auditable.
  28. As a model-quality engineer, I want event-derived model, token, duration, and cost evidence captured when Claude Code exposes it, so that resource behavior can be interpreted alongside correctness.
  29. As a model-quality engineer, I want model and cost gaps surfaced honestly when unavailable, so that reports do not pretend to know runtime facts the harness cannot observe.
  30. As a solo developer, I want the Claude Code baseline report to explain what Claude Code does well, poorly, and inconsistently under the evaluated conditions, so that I can understand the practical tool-fit signal.
  31. As a solo developer, I want the report to be Markdown and AI-readable, so that another AI assistant can help interpret the evidence for my project.
  32. As a project maintainer, I want the generated capability evidence digest to remain separate from hand-authored interpretation, so that evidence and claims are not blurred.
  33. As a project maintainer, I want the Claude Code baseline report to avoid ranking Claude Code against Codex, so that comparison waits for a later comparison PRD.
  34. As a project maintainer, I want the future comparison PRD to reference both baseline PRDs, so that cross-harness claims are grounded in completed evidence sets.
  35. As an AFK coding agent, I want this PRD split into child implementation issues after approval, so that each implementation slice has a narrow scope and clear stop conditions.
  36. As an AFK coding agent, I want readiness work separated from evidence collection, so that adapter/configuration bugs are fixed before spending trial budget.
  37. As an AFK coding agent, I want Phase 1 separated from Phase 2, so that full-suite smoke evidence can be reviewed before repeated trials scale.
  38. As an AFK coding agent, I want report writing separated from trial execution, so that interpretation can wait until evidence selection and review are complete.

Implementation Decisions

  • This PRD is for a Claude Code baseline, not a Codex-vs-Claude comparison.
  • The baseline uses the existing 12-task Solo Dev Starter Suite.
  • The baseline evaluates one explicit Claude Code agent harness configuration.
  • The exact model alias or identifier is not hard-coded in this PRD. The readiness slice must verify the cheapest/fastest available Claude Code model supported by the installed CLI and Jordan's account.
  • The proof-of-concept baseline should prefer cost and speed over headline quality. A stronger Sonnet/Opus baseline can be a later PRD if this baseline earns the cost.
  • The Claude Code configuration should match the Codex baseline wherever the Claude Code CLI has a fair equivalent.
  • The current expected Claude Code configuration shape is print mode, stream JSON output, local edit permissions, no session persistence, and a 1800-second timeout unless readiness work reveals a fairness issue.
  • Codex used one non-interactive codex exec invocation per trial. Claude Code should mirror that as one non-interactive claude -p session per trial.
  • Claude Code max-turn behavior must be verified before Phase 1 because its --max-turns flag limits agentic turns and may produce an error when reached.
  • Do not add Claude-specific prompt changes.
  • Do not add Claude-specific grader changes unless Phase 1 exposes a fairness problem in the task contract or grader.
  • Do not add task-specific tool-policy tuning to help Claude Code pass.
  • Suite-level tool-policy changes are allowed only for documented fairness or safety reasons.
  • Phase 1 runs one sequential smoke trial for every starter-suite task.
  • Phase 1 uses one trial and one job per task.
  • Fair smoke trials count toward the five-fair-trials-per-task baseline.
  • Invalid smoke trials are diagnostics. They must be fixed or rerun before repeated collection for that task.
  • Fair failed smoke trials continue into Phase 2. Capability failure is evidence, not a stop condition.
  • Phase 2 runs four additional fair trials per task after a fair smoke, for a final selected evidence set of five fair trials per task.
  • Human review is required for the first fair trial for each task.
  • Later trials require human review when they fail, look suspicious, are unusually resource-heavy, have messy or risky patches, or need validity/exclusion judgment.
  • The selected evidence set must use existing trial validity and exclusion semantics.
  • The generated capability evidence digest is evidence, not final interpretation.
  • The hand-authored baseline report interprets Claude Code under the evaluated conditions and does not rank Claude Code against Codex.
  • The PRD should be approved before implementation issues are created.
  • Likely child implementation issues are readiness/config verification, Phase 1 full-suite shakedown, Phase 2 repeated trials and selected evidence set, and hand-authored Claude Code baseline report.
  • No ADR is needed for this sequencing choice because it is a product/evidence planning decision rather than a hard-to-reverse architecture decision.

Major modules or areas likely to be exercised:

  • The Claude Code agent adapter for command construction, event capture, model identity, resource usage, timeout behavior, and report metadata.
  • Shared model identity parsing for event-derived model names and requested-model fallback.
  • CLI agent-option construction for explicit Claude Code configuration.
  • Trial running and summarization for sequential Phase 1 smokes and repeated Phase 2 trials.
  • Outcome evidence loading and capability evidence digest generation for selected Claude Code evidence.
  • Human review metadata and trial validity/exclusion handling.

Potential deep-module opportunities:

  • If readiness exposes more Claude/Codex duplication, keep reusable terminal, reporting, resource, or model-identity behavior in shared modules and keep child adapters focused on adapter-specific command/event behavior.
  • If evidence-set selection becomes ad hoc during this baseline, prefer a small reusable evidence-set construction interface over hand-maintained lists.
  • If Claude configuration verification grows, prefer a focused readiness/preflight module over spreading model-alias, max-turn, and report-field checks across command handlers.

Testing Decisions

  • Good tests should exercise external behavior: CLI parsing, command construction, preflight results, result/report artifacts, event-derived model identity, resource usage, trial validity, evidence digest rows, and human review metadata.
  • Tests should not overfit to private helper internals unless the helper encodes an externally important invariant.
  • Existing Claude Code adapter tests are prior art for command construction, output format, permission mode, model request, max turns, no session persistence, runtime-accountability metadata, event parsing, and report rendering.
  • Existing model identity tests are prior art for captured event fixtures and fallback behavior when events do not expose a model.
  • Existing result-loading tests are prior art for backfilling model identity, edit size, and resource metrics from stored artifacts.
  • Existing CLI tests are prior art for agent-option parsing and command-family behavior.
  • Existing evidence/report tests are prior art for capability evidence digests and fair/excluded trial summaries.
  • Readiness/config verification should include tests or smoke checks that show the selected Claude Code configuration appears in result.json and report.md.
  • If a new Claude Code event shape is discovered, add a sanitized JSONL fixture based on the real event shape.
  • If max-turn behavior affects fairness, add behavior coverage around the chosen config or preflight/readiness check.
  • Phase 1 and Phase 2 trial execution should follow the repo rule: start new task, grader, environment, or agent-harness behavior with one trial and one job before repeated or parallel batches.
  • Human review and exclusion decisions should be auditable through existing review metadata rather than only prose notes.

Out of Scope

  • Comparing Claude Code against Codex in this PRD.
  • Ranking agent harnesses.
  • Changing the Solo Dev Starter Suite prompts to help Claude Code.
  • Changing graders because the cheap/fast model struggles, unless the grader or task contract is unfair.
  • Running a stronger Claude Code Sonnet/Opus baseline.
  • Running a large parallel batch before Phase 1 smoke evidence is fair.
  • Building a web dashboard.
  • Adding model-based graders.
  • Creating a permanent local planning document for this PRD outside GitHub Issues.
  • Closing PRD: Codex deep baseline and evidence-scoped capability reports #10 automatically as part of this PRD unless Jordan explicitly chooses to do that closure step.

Further Notes

  • This PRD should become the active planning anchor only after PRD: Codex deep baseline and evidence-scoped capability reports #10 is closed as the completed Codex baseline PRD.
  • A later Codex-vs-Claude comparison PRD should reference both PRD: Codex deep baseline and evidence-scoped capability reports #10 and this Claude Code baseline PRD.
  • Claude Code pilot evidence already exists in the repo documentation. It showed one excluded eval-harness error, one fair pass, and one fair failure. That is enough to justify a baseline PRD, but not enough to start with comparison.
  • The first implementation child issue should be readiness/config verification, not trial scaling.
  • The PRD issue itself should not be labeled ready-for-agent until Jordan approves the PRD and child implementation issues are created.

Activity

  1. Jordak commented on May 14, 2026

    @Jordak
    OwnerAuthor

    #10 is now closed as the completed Codex baseline, so this PRD is the active planning anchor for the Claude Code baseline.

    I split the PRD into child issues:

    The parent PRD should stay open until those slices produce the selected Claude Code evidence set, generated digest, baseline report, and closeout comment. #50 remains blocked until this baseline is complete.

  2. Jordak commented on May 16, 2026

    @Jordak
    OwnerAuthor

    Claude Code baseline closeout is now complete.

    Merged artifacts on main:

    Final selected evidence shape:

    • Agent harness: claude
    • Adapter/config: Claude Code CLI 2.1.139, acceptEdits, stream-json, 1800 second timeout, no explicit max turns
    • Model: claude-haiku-4-5-20251001, recovered from Claude Code event metadata
    • Suite: 12-task Solo Dev Starter Suite
    • Selected fair trials: 60 total, 5 fair trials per task
    • Deterministic grader result: 53/60 passes
    • Task-level result: 10 tasks passed 5/5, remotion-audio-context-autoplay-muted-001 passed 3/5, and click-should-strip-ansi-tests-001 passed 0/5
    • Recorded usage in the selected fair evidence set: $16.35

    The report keeps interpretation scoped to the evaluated Claude Code harness configuration and does not make Codex-vs-Claude comparison claims. #50 remains open as the later comparison PRD.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions