Skip to content

PRD: Codex and Claude Code starter-suite comparison #50

Description

@Jordak

Problem Statement

Agent Eval Lab is moving from a single-harness Codex baseline toward multiple agent harness baselines. The project already has a Codex starter-suite evidence set under #10 and a Claude Code baseline PRD under #49, but it does not yet have a dedicated plan for how to compare those baselines after both are complete.

From the user's perspective, the core problem is: "I want to compare Codex CLI and Claude Code in a way that helps technical evaluators and solo developers understand practical tool fit, but I do not want a premature leaderboard or a comparison that mixes unfair evidence, different task contracts, different trial counts, hidden model defaults, or unreviewed artifacts."

Without a separate comparison PRD, comparison work risks leaking into the Claude Code baseline, closing #10 with an unclear standard, or turning generated digest rows into overbroad claims. The project needs a future comparison plan that treats each baseline as a prerequisite and defines how cross-harness claims will be made, scoped, and caveated.

Prerequisites:

This PRD should not become active implementation work until #10 is closed and #49 has produced a selected Claude Code evidence set, generated digest, and hand-authored Claude Code baseline report.

Solution

Create a Codex-vs-Claude comparison report only after both agent harness baselines exist.

The comparison should use selected fair evidence sets from both baselines, align them by task, trial count, agent harness configuration, model identity when known, runtime conditions, deterministic grader outcomes, human review outcomes, and resource usage metrics. It should produce an evidence-scoped comparison report in Markdown that explains relative strengths, weaknesses, reliability, patch quality, and resource behavior under the evaluated conditions.

The comparison report should not replace either baseline report. It should sit on top of them:

  1. Codex baseline: completed evidence and interpretation for Codex CLI.
  2. Claude Code baseline: completed evidence and interpretation for Claude Code.
  3. Comparison report: cross-harness interpretation grounded in both baselines.

The report should avoid global "best model" or universal "best tool" claims. Its job is to help readers understand what these two agent harness configurations did on this suite, where evidence is strong, where it is weak, and which practical tool-fit questions remain unanswered.

User Stories

  1. As a technical evaluator, I want Codex and Claude Code compared only after each has a completed baseline, so that comparison claims are grounded in fair evidence.
  2. As a technical evaluator, I want the comparison to reference both baseline PRDs, so that readers can trace each claim back to its evidence program.
  3. As a technical evaluator, I want selected fair evidence sets from both baselines, so that invalid trials do not distort cross-harness conclusions.
  4. As a technical evaluator, I want task-by-task comparison across the same Solo Dev Starter Suite, so that each harness is compared on the same task contracts.
  5. As a technical evaluator, I want equal fair-trial counts per task where possible, so that pass rates and consistency metrics compare like with like.
  6. As a technical evaluator, I want deterministic grader outcomes compared separately from human review outcomes, so that correctness and patch quality are both visible.
  7. As a technical evaluator, I want pass rate, pass@k, and pass^k compared by task and aggregate group, so that "eventually succeeds" and "consistently succeeds" remain distinct.
  8. As a technical evaluator, I want human review labels compared across harnesses, so that messy successes, test gaps, tool misuse, over-editing, and resource inefficiency are visible.
  9. As a technical evaluator, I want exclusions compared and explained, so that harness or environment failures do not masquerade as capability gaps.
  10. As a technical evaluator, I want resource usage metrics compared when available, so that runtime, token budget, and cost are part of practical tool-fit analysis.
  11. As a technical evaluator, I want missing model, token, or cost data called out explicitly, so that the report does not pretend all metadata is equally complete.
  12. As a model-quality engineer, I want agent harness and underlying model kept separate, so that Codex CLI and Claude Code are not collapsed into ambiguous model-vs-model claims.
  13. As a model-quality engineer, I want model identity and model source reported for each evidence set, so that explicit model requests and event-derived runtime identity are not confused.
  14. As a model-quality engineer, I want comparison claims scoped to the evaluated agent harness configurations, so that changes in permissions, model defaults, or CLI versions do not get hidden.
  15. As a model-quality engineer, I want confidence caveats when one harness lacks comparable metadata, so that gaps in observability are not interpreted as performance differences.
  16. As a solo developer, I want to know which harness produced cleaner patches on this suite, so that I can choose a tool aligned with my tolerance for review burden.
  17. As a solo developer, I want to know which harness was more resource-efficient, so that I can factor time and cost into practical tool choice.
  18. As a solo developer, I want to know which tasks each harness handled reliably, so that I can map evidence to my own project type.
  19. As a solo developer, I want the report to avoid universal rankings, so that I do not mistake this starter-suite comparison for a general benchmark.
  20. As a solo developer, I want evidence-scoped recommendations, so that I can act on the report without overreading it.
  21. As a project maintainer, I want the comparison report generated from explicit evidence manifests, so that the selected input set is durable and auditable.
  22. As a project maintainer, I want report generation to reuse existing capability evidence digest machinery where appropriate, so that comparison does not fork the reporting model unnecessarily.
  23. As a project maintainer, I want the comparison to surface both aggregate and per-task views, so that a strong aggregate result does not hide task-specific brittleness.
  24. As a project maintainer, I want the comparison to identify evidence gaps, so that the next suite or baseline can target missing task types or metadata.
  25. As a project maintainer, I want the comparison to preserve public/private boundaries, so that public docs discuss evaluation evidence rather than private career or account context.
  26. As an AFK coding agent, I want the comparison work split into child issues after prerequisites are complete, so that implementation can proceed through narrow, verifiable slices.
  27. As an AFK coding agent, I want the comparison data preparation issue separate from report interpretation, so that artifact correctness can be verified before claims are written.
  28. As an AFK coding agent, I want clear stop conditions when evidence sets are not comparable, so that the agent does not force a misleading report.

Implementation Decisions

  • This PRD is blocked until PRD: Codex deep baseline and evidence-scoped capability reports #10 is closed and PRD: Claude Code baseline for the Solo Dev Starter Suite #49 has produced a completed Claude Code baseline.
  • The comparison report must reference both source PRDs and their final evidence artifacts.
  • The comparison should use selected fair evidence sets rather than scanning all local runs.
  • The comparison should use the same Solo Dev Starter Suite task contracts from the baselines.
  • The comparison should compare agent harness configurations, not just agent names.
  • Underlying model must remain a separate dimension from agent harness.
  • Explicit model requests, event-derived runtime model identity, and unknown model identity should remain distinguishable.
  • Invalid/excluded trials should be omitted from fair capability metrics but summarized as operational evidence.
  • Deterministic grader outcomes and human review outcomes should remain separate.
  • Human review labels should be compared both as primary labels and secondary labels where available.
  • Resource usage should be compared only when collected comparably enough to support a claim.
  • Missing cost or model data should be treated as an evidence gap, not silently rendered as equality.
  • The report should include task-by-task comparisons before aggregate recommendations.
  • The report should include aggregate summaries only after task-level evidence is inspectable.
  • The report should include evidence-scoped recommendations for solo developers, but not universal rankings.
  • The comparison should not reinterpret raw transcripts without linking to the underlying artifact or evidence digest.
  • If the existing reporting modules cannot express cross-harness comparison cleanly, prefer a small comparison-reporting module with a focused interface rather than duplicating digest code.
  • If evidence-set manifests need normalization before comparison, prefer a reusable evidence-set comparison input module over ad hoc string/path manipulation.
  • The comparison PRD should be split into child implementation issues after prerequisites are complete.

Likely child issue slices:

  • Verify prerequisite closure and collect final baseline artifact links.
  • Build or adapt comparison input selection from the Codex and Claude evidence sets.
  • Generate a comparison evidence appendix or digest that aligns trials by task, harness, and model.
  • Draft the hand-authored Codex-vs-Claude comparison report.
  • Review and publish evidence-scoped recommendations.

Potential deep-module opportunities:

  • A comparison input module that accepts explicit evidence-set manifests and produces normalized per-task/per-harness comparison rows.
  • A comparison summary module that computes pass rate, pass@k, pass^k, review-label counts, resource medians, and exclusion summaries across harnesses.
  • A report-rendering module that keeps generated comparison evidence separate from hand-authored interpretation.

Testing Decisions

  • Good tests should exercise external behavior: comparison rows, grouped metrics, exclusion handling, human-review label aggregation, resource usage summaries, and generated Markdown output.
  • Tests should use small fixture evidence sets that include at least two agent harnesses and at least two tasks.
  • Tests should cover valid trials, excluded trials, deterministic passes/failures, primary review labels, secondary review labels, and missing model/resource fields.
  • Tests should verify that invalid trials are excluded from fair pass metrics but still summarized as exclusions.
  • Tests should verify that model identity and agent harness remain separate grouping dimensions.
  • Tests should verify that missing model or cost fields render as explicit evidence gaps rather than misleading zeros.
  • Tests should avoid asserting private helper internals when output tables and structured comparison rows express the behavior.
  • Prior art includes capability evidence digest tests, evidence-set tests, result-loading backfill tests, and trials summary tests.
  • If a new comparison module is created, it should be testable through a small, stable interface that takes loaded evidence records and returns comparison summaries.

Out of Scope

  • Running the Claude Code baseline itself. That belongs to PRD: Claude Code baseline for the Solo Dev Starter Suite #49.
  • Reopening or rerunning the Codex baseline except to fix artifact integrity issues.
  • Claiming universal model superiority.
  • Creating a leaderboard.
  • Comparing other agent harnesses such as Cursor Agent.
  • Changing task prompts or graders for comparison convenience.
  • Mixing different suites or different task contracts into the main comparison.
  • Treating private account, billing, or career context as public report content.
  • Building a web dashboard.
  • Adding model-based graders.
  • Making the comparison PRD ready for implementation before the Claude Code baseline is complete.

Further Notes

Activity

  1. Jordak commented on May 16, 2026

    @Jordak
    OwnerAuthor

    Reviewed comparison report has shipped on main.

    Artifact:

    Commit: f9e4103

    The report explicitly scopes the comparison to the evaluated configurations: Codex CLI with recovered gpt-5.5 / xhigh metadata versus Claude Code with event-derived claude-haiku-4-5-20251001. It avoids treating the asymmetric baselines as a universal model ranking and calls out the remaining evidence gaps for lower-effort Codex, stronger Claude Code, generated comparison appendix work, and resource/cost comparability.

  2. Jordak commented on May 16, 2026

    @Jordak
    OwnerAuthor

    This was generated by AI during triage.

    Follow-up issues created from the next-step discussion:

    Naming decision: use comparison evidence digest, not appendix, for the generated multi-config evidence artifact.

    Execution shape: run the comparison evidence digest work first or alongside the new baselines; treat the Opus baseline as a smoke-gated human-budget issue; track low-effort Codex in one issue unless portable reasoning-effort plumbing grows large enough to split.

  3. Keesan12 commented on May 17, 2026

    @Keesan12

    The comparison report is useful, especially the explicit caveat that this is config-scoped evidence, not a universal ranking.

    One thing that’s helped us keep these evals honest is logging cost per accepted / verified outcome, not just model-level win/loss. A run that looks strong on raw completion can still be the worse operator experience if it burns retries before the verifier moves.

    If that’s interesting, I’m happy to share the small run-record fields we’ve been using in MartinLoop for that layer.

  4. Jordak commented on May 18, 2026

    @Jordak
    OwnerAuthor

    @Keesan12 Thanks for the suggestion! I created #64 to implement that (though I'll use "tokens" rather than "cost" since Codex apparently doesn't report dollar amounts).

    Thanks for the offer too! Feel free to link me to the eval structure of MartinLoop and I'll take a look! I like the name, btw. I assume named after Martin Prince?

  5. Keesan12 commented on May 18, 2026

    @Keesan12

    Appreciate it — and the scope you’re keeping on the comparison looks right.

    The cleanest MartinLoop reference points are probably:

    The fields that have mattered most for us are boring on purpose: task/run id, budget remaining, verifier status, attempt number, halt reason, and enough tool/patch context to explain why the next attempt was or was not admitted.

    And no, not Martin Prince — more “keep the loop from going in circles and leave a record.” If the shape ends up being useful, blunt feedback is welcome. And if MartinLoop looks relevant to your eval workflow, a star helps.

  6. Keesan12 commented on May 20, 2026

    @Keesan12

    A side-by-side starter suite gets much more valuable if it compares control-plane behavior as well as raw task performance.

    The useful rows are things like: what budget controls exist, what halt reasons you can inspect, how approval boundaries work, whether verifier state is explicit, and what kind of receipt you get after an interrupted run. Those are often the differences that determine whether a tool is workable in production, even when the benchmark scores look similar.

    That’s the comparison layer we keep coming back to on MartinLoop-style runs.

  7. Jordak commented on Jun 4, 2026

    @Jordak
    OwnerAuthor

    This was generated by AI during triage.

    Codex low-effort baseline evidence for #63 is ready in draft PR #87:

    Summary: 60/60 selected fair trials passed across the 12 starter-suite tasks using Codex CLI with gpt-5.5 and recovered runtime reasoning_effort=low. No trials were excluded. All selected trials have primary success_clean; 5 also have secondary resource_inefficient. Total recorded tokens are 18,660,131, and cost_usd remains unknown.

    Comparison interpretation should stay in #50 or a child issue under #50 rather than expanding #63 beyond baseline collection.

  8. added a commit that references this issue on Jun 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions