You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
PRD: Codex and Claude Code starter-suite comparison #50
Agent Eval Lab is moving from a single-harness Codex baseline toward multiple agent harness baselines. The project already has a Codex starter-suite evidence set under #10 and a Claude Code baseline PRD under #49, but it does not yet have a dedicated plan for how to compare those baselines after both are complete.
From the user's perspective, the core problem is: "I want to compare Codex CLI and Claude Code in a way that helps technical evaluators and solo developers understand practical tool fit, but I do not want a premature leaderboard or a comparison that mixes unfair evidence, different task contracts, different trial counts, hidden model defaults, or unreviewed artifacts."
Without a separate comparison PRD, comparison work risks leaking into the Claude Code baseline, closing #10 with an unclear standard, or turning generated digest rows into overbroad claims. The project needs a future comparison plan that treats each baseline as a prerequisite and defines how cross-harness claims will be made, scoped, and caveated.
This PRD should not become active implementation work until #10 is closed and #49 has produced a selected Claude Code evidence set, generated digest, and hand-authored Claude Code baseline report.
Solution
Create a Codex-vs-Claude comparison report only after both agent harness baselines exist.
The comparison should use selected fair evidence sets from both baselines, align them by task, trial count, agent harness configuration, model identity when known, runtime conditions, deterministic grader outcomes, human review outcomes, and resource usage metrics. It should produce an evidence-scoped comparison report in Markdown that explains relative strengths, weaknesses, reliability, patch quality, and resource behavior under the evaluated conditions.
The comparison report should not replace either baseline report. It should sit on top of them:
Codex baseline: completed evidence and interpretation for Codex CLI.
Claude Code baseline: completed evidence and interpretation for Claude Code.
Comparison report: cross-harness interpretation grounded in both baselines.
The report should avoid global "best model" or universal "best tool" claims. Its job is to help readers understand what these two agent harness configurations did on this suite, where evidence is strong, where it is weak, and which practical tool-fit questions remain unanswered.
User Stories
As a technical evaluator, I want Codex and Claude Code compared only after each has a completed baseline, so that comparison claims are grounded in fair evidence.
As a technical evaluator, I want the comparison to reference both baseline PRDs, so that readers can trace each claim back to its evidence program.
As a technical evaluator, I want selected fair evidence sets from both baselines, so that invalid trials do not distort cross-harness conclusions.
As a technical evaluator, I want task-by-task comparison across the same Solo Dev Starter Suite, so that each harness is compared on the same task contracts.
As a technical evaluator, I want equal fair-trial counts per task where possible, so that pass rates and consistency metrics compare like with like.
As a technical evaluator, I want deterministic grader outcomes compared separately from human review outcomes, so that correctness and patch quality are both visible.
As a technical evaluator, I want pass rate, pass@k, and pass^k compared by task and aggregate group, so that "eventually succeeds" and "consistently succeeds" remain distinct.
As a technical evaluator, I want human review labels compared across harnesses, so that messy successes, test gaps, tool misuse, over-editing, and resource inefficiency are visible.
As a technical evaluator, I want exclusions compared and explained, so that harness or environment failures do not masquerade as capability gaps.
As a technical evaluator, I want resource usage metrics compared when available, so that runtime, token budget, and cost are part of practical tool-fit analysis.
As a technical evaluator, I want missing model, token, or cost data called out explicitly, so that the report does not pretend all metadata is equally complete.
As a model-quality engineer, I want agent harness and underlying model kept separate, so that Codex CLI and Claude Code are not collapsed into ambiguous model-vs-model claims.
As a model-quality engineer, I want model identity and model source reported for each evidence set, so that explicit model requests and event-derived runtime identity are not confused.
As a model-quality engineer, I want comparison claims scoped to the evaluated agent harness configurations, so that changes in permissions, model defaults, or CLI versions do not get hidden.
As a model-quality engineer, I want confidence caveats when one harness lacks comparable metadata, so that gaps in observability are not interpreted as performance differences.
As a solo developer, I want to know which harness produced cleaner patches on this suite, so that I can choose a tool aligned with my tolerance for review burden.
As a solo developer, I want to know which harness was more resource-efficient, so that I can factor time and cost into practical tool choice.
As a solo developer, I want to know which tasks each harness handled reliably, so that I can map evidence to my own project type.
As a solo developer, I want the report to avoid universal rankings, so that I do not mistake this starter-suite comparison for a general benchmark.
As a solo developer, I want evidence-scoped recommendations, so that I can act on the report without overreading it.
As a project maintainer, I want the comparison report generated from explicit evidence manifests, so that the selected input set is durable and auditable.
As a project maintainer, I want report generation to reuse existing capability evidence digest machinery where appropriate, so that comparison does not fork the reporting model unnecessarily.
As a project maintainer, I want the comparison to surface both aggregate and per-task views, so that a strong aggregate result does not hide task-specific brittleness.
As a project maintainer, I want the comparison to identify evidence gaps, so that the next suite or baseline can target missing task types or metadata.
As a project maintainer, I want the comparison to preserve public/private boundaries, so that public docs discuss evaluation evidence rather than private career or account context.
As an AFK coding agent, I want the comparison work split into child issues after prerequisites are complete, so that implementation can proceed through narrow, verifiable slices.
As an AFK coding agent, I want the comparison data preparation issue separate from report interpretation, so that artifact correctness can be verified before claims are written.
As an AFK coding agent, I want clear stop conditions when evidence sets are not comparable, so that the agent does not force a misleading report.
The comparison report must reference both source PRDs and their final evidence artifacts.
The comparison should use selected fair evidence sets rather than scanning all local runs.
The comparison should use the same Solo Dev Starter Suite task contracts from the baselines.
The comparison should compare agent harness configurations, not just agent names.
Underlying model must remain a separate dimension from agent harness.
Explicit model requests, event-derived runtime model identity, and unknown model identity should remain distinguishable.
Invalid/excluded trials should be omitted from fair capability metrics but summarized as operational evidence.
Deterministic grader outcomes and human review outcomes should remain separate.
Human review labels should be compared both as primary labels and secondary labels where available.
Resource usage should be compared only when collected comparably enough to support a claim.
Missing cost or model data should be treated as an evidence gap, not silently rendered as equality.
The report should include task-by-task comparisons before aggregate recommendations.
The report should include aggregate summaries only after task-level evidence is inspectable.
The report should include evidence-scoped recommendations for solo developers, but not universal rankings.
The comparison should not reinterpret raw transcripts without linking to the underlying artifact or evidence digest.
If the existing reporting modules cannot express cross-harness comparison cleanly, prefer a small comparison-reporting module with a focused interface rather than duplicating digest code.
If evidence-set manifests need normalization before comparison, prefer a reusable evidence-set comparison input module over ad hoc string/path manipulation.
The comparison PRD should be split into child implementation issues after prerequisites are complete.
Likely child issue slices:
Verify prerequisite closure and collect final baseline artifact links.
Build or adapt comparison input selection from the Codex and Claude evidence sets.
Generate a comparison evidence appendix or digest that aligns trials by task, harness, and model.
Draft the hand-authored Codex-vs-Claude comparison report.
Review and publish evidence-scoped recommendations.
Potential deep-module opportunities:
A comparison input module that accepts explicit evidence-set manifests and produces normalized per-task/per-harness comparison rows.
A comparison summary module that computes pass rate, pass@k, pass^k, review-label counts, resource medians, and exclusion summaries across harnesses.
A report-rendering module that keeps generated comparison evidence separate from hand-authored interpretation.
Testing Decisions
Good tests should exercise external behavior: comparison rows, grouped metrics, exclusion handling, human-review label aggregation, resource usage summaries, and generated Markdown output.
Tests should use small fixture evidence sets that include at least two agent harnesses and at least two tasks.
Tests should cover valid trials, excluded trials, deterministic passes/failures, primary review labels, secondary review labels, and missing model/resource fields.
Tests should verify that invalid trials are excluded from fair pass metrics but still summarized as exclusions.
Tests should verify that model identity and agent harness remain separate grouping dimensions.
Tests should verify that missing model or cost fields render as explicit evidence gaps rather than misleading zeros.
Tests should avoid asserting private helper internals when output tables and structured comparison rows express the behavior.
Prior art includes capability evidence digest tests, evidence-set tests, result-loading backfill tests, and trials summary tests.
If a new comparison module is created, it should be testable through a small, stable interface that takes loaded evidence records and returns comparison summaries.
The report explicitly scopes the comparison to the evaluated configurations: Codex CLI with recovered gpt-5.5 / xhigh metadata versus Claude Code with event-derived claude-haiku-4-5-20251001. It avoids treating the asymmetric baselines as a universal model ranking and calls out the remaining evidence gaps for lower-effort Codex, stronger Claude Code, generated comparison appendix work, and resource/cost comparability.
Naming decision: use comparison evidence digest, not appendix, for the generated multi-config evidence artifact.
Execution shape: run the comparison evidence digest work first or alongside the new baselines; treat the Opus baseline as a smoke-gated human-budget issue; track low-effort Codex in one issue unless portable reasoning-effort plumbing grows large enough to split.
The comparison report is useful, especially the explicit caveat that this is config-scoped evidence, not a universal ranking.
One thing that’s helped us keep these evals honest is logging cost per accepted / verified outcome, not just model-level win/loss. A run that looks strong on raw completion can still be the worse operator experience if it burns retries before the verifier moves.
If that’s interesting, I’m happy to share the small run-record fields we’ve been using in MartinLoop for that layer.
@Keesan12 Thanks for the suggestion! I created #64 to implement that (though I'll use "tokens" rather than "cost" since Codex apparently doesn't report dollar amounts).
Thanks for the offer too! Feel free to link me to the eval structure of MartinLoop and I'll take a look! I like the name, btw. I assume named after Martin Prince?
The fields that have mattered most for us are boring on purpose: task/run id, budget remaining, verifier status, attempt number, halt reason, and enough tool/patch context to explain why the next attempt was or was not admitted.
And no, not Martin Prince — more “keep the loop from going in circles and leave a record.” If the shape ends up being useful, blunt feedback is welcome. And if MartinLoop looks relevant to your eval workflow, a star helps.
A side-by-side starter suite gets much more valuable if it compares control-plane behavior as well as raw task performance.
The useful rows are things like: what budget controls exist, what halt reasons you can inspect, how approval boundaries work, whether verifier state is explicit, and what kind of receipt you get after an interrupted run. Those are often the differences that determine whether a tool is workable in production, even when the benchmark scores look similar.
That’s the comparison layer we keep coming back to on MartinLoop-style runs.
Summary: 60/60 selected fair trials passed across the 12 starter-suite tasks using Codex CLI with gpt-5.5 and recovered runtime reasoning_effort=low. No trials were excluded. All selected trials have primary success_clean; 5 also have secondary resource_inefficient. Total recorded tokens are 18,660,131, and cost_usd remains unknown.
Comparison interpretation should stay in #50 or a child issue under #50 rather than expanding #63 beyond baseline collection.
Problem Statement
Agent Eval Lab is moving from a single-harness Codex baseline toward multiple agent harness baselines. The project already has a Codex starter-suite evidence set under #10 and a Claude Code baseline PRD under #49, but it does not yet have a dedicated plan for how to compare those baselines after both are complete.
From the user's perspective, the core problem is: "I want to compare Codex CLI and Claude Code in a way that helps technical evaluators and solo developers understand practical tool fit, but I do not want a premature leaderboard or a comparison that mixes unfair evidence, different task contracts, different trial counts, hidden model defaults, or unreviewed artifacts."
Without a separate comparison PRD, comparison work risks leaking into the Claude Code baseline, closing #10 with an unclear standard, or turning generated digest rows into overbroad claims. The project needs a future comparison plan that treats each baseline as a prerequisite and defines how cross-harness claims will be made, scoped, and caveated.
Prerequisites:
This PRD should not become active implementation work until #10 is closed and #49 has produced a selected Claude Code evidence set, generated digest, and hand-authored Claude Code baseline report.
Solution
Create a Codex-vs-Claude comparison report only after both agent harness baselines exist.
The comparison should use selected fair evidence sets from both baselines, align them by task, trial count, agent harness configuration, model identity when known, runtime conditions, deterministic grader outcomes, human review outcomes, and resource usage metrics. It should produce an evidence-scoped comparison report in Markdown that explains relative strengths, weaknesses, reliability, patch quality, and resource behavior under the evaluated conditions.
The comparison report should not replace either baseline report. It should sit on top of them:
The report should avoid global "best model" or universal "best tool" claims. Its job is to help readers understand what these two agent harness configurations did on this suite, where evidence is strong, where it is weak, and which practical tool-fit questions remain unanswered.
User Stories
Implementation Decisions
Likely child issue slices:
Potential deep-module opportunities:
Testing Decisions
Out of Scope
Further Notes
ready-for-agentyet.