You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Codex low-effort starter-suite baseline for config comparison #63
#50 - PRD: Codex and Claude Code starter-suite comparison
Current behavior / design
The project has a completed Codex starter-suite baseline using recovered runtime metadata gpt-5.5 / xhigh. The first comparison report calls out that this high-effort Codex result is asymmetric against the cheap Claude Haiku baseline.
Agent Eval Lab currently exposes --codex-model, --codex-profile, sandbox, approval policy, timeout, and Codex state DB options. It does not currently expose a first-class --codex-reasoning-effort option in the task/run CLI.
Why change
A low-effort Codex starter-suite baseline would show how much of the current 60/60 Codex result depends on xhigh reasoning effort. This is the fastest way to make the comparison matrix more useful for practical tool-fit decisions around cost, latency, reliability, and review burden.
One issue should be enough to track this because it is the same agent harness and same model family as the completed Codex baseline. If effort plumbing becomes larger than expected, split that implementation into a child issue before trial collection.
What to build
Collect a Codex low-effort starter-suite baseline for the same 12 tasks and five-fair-trials-per-task shape.
Preferred target configuration:
Agent harness: codex
Model: gpt-5.5
Reasoning effort: low
Sandbox: match the existing Codex baseline unless a documented fairness issue requires a change
Approval policy: match the existing Codex baseline unless a documented fairness issue requires a change
Timeout: 1800 seconds unless readiness reveals a fairness issue
Implementation phases:
Verify how to request and prove low reasoning effort portably.
If needed, add first-class Agent Eval Lab CLI/config support for Codex reasoning effort rather than relying on a private local profile name.
Run a small preflight or smoke check proving result metadata records model_name = gpt-5.5 and reasoning_effort = low through recovered runtime metadata or another trustworthy source.
Run the 12-task starter-suite baseline, preserving the same task prompts, graders, reference artifacts, and trial validity semantics as the high-effort Codex baseline.
Start with one trial and one job for any new effort/config plumbing before scaling repeated collection.
Select five fair trials per task, excluding only invalid trials with explicit reasons.
Generate a selected evidence-set manifest and capability evidence digest under an appropriate report bundle.
The selected low-effort Codex configuration is recorded in a readiness/config note or report-bundle note.
Agent Eval Lab can request gpt-5.5 with low reasoning effort through a portable mechanism, or the issue documents why a profile-based mechanism was used.
Stored or recovered result metadata distinguishes the low-effort baseline from the existing xhigh Codex baseline.
Each starter-suite task has five selected fair low-effort Codex trials.
Excluded trials, if any, have explicit validity and exclusion-reason metadata and are omitted from fair metrics.
A selected evidence-set manifest is committed under evidence-sets/.
A generated capability evidence digest is committed under an appropriate reports/codex-low-effort-.../ bundle.
Human review labels are recorded for failures, suspicious/resource-heavy trials, messy patches, and any trial requiring validity judgment.
Tests cover any new Codex reasoning-effort CLI/config plumbing.
The issue comment summarizes pass rate, pass@5, pass^5, runtime/resource caveats, and comparison implications versus the existing xhigh Codex baseline.
Stop conditions
Stop before repeated collection if the harness cannot prove the run used low effort.
Stop if the low-effort request is silently ignored or overwritten by local/global Codex config.
Stop if a new Agent Eval Lab config change fails tests.
Stop and ask Jordan before scaling if the first smoke path suggests unexpectedly high cost, quota pressure, or repeated environment invalidity.
Out of scope
Changing starter-suite tasks or graders for low-effort Codex.
Running a different Codex model unless Jordan explicitly chooses that as a separate comparison axis.
Selected evidence: 60 trials across the 12 starter-suite tasks.
Outcome: 60/60 passed; each task has five selected fair trials.
Exclusions: none.
Metadata proof: all selected trials record model_name=gpt-5.5, requested reasoning_effort=low, recovered runtime reasoning_effort=low, and run-surface metadata from local Codex state.
Review overlay: all 60 selected trials have primary success_clean; 5 also carry secondary resource_inefficient.
Resource caveat: total recorded tokens are 18,660,131; cost_usd is unknown for all selected trials, so dollar cost is not measured by the harness.
Scope note:
This issue should stop at the Codex low-effort baseline evidence. Cross-agent comparison, including how this low-effort Codex baseline should be interpreted against Claude Code evidence, belongs under parent issue #50. Report-link portability is tracked separately in #86.
Parent
#50 - PRD: Codex and Claude Code starter-suite comparison
Current behavior / design
The project has a completed Codex starter-suite baseline using recovered runtime metadata
gpt-5.5/xhigh. The first comparison report calls out that this high-effort Codex result is asymmetric against the cheap Claude Haiku baseline.Agent Eval Lab currently exposes
--codex-model,--codex-profile, sandbox, approval policy, timeout, and Codex state DB options. It does not currently expose a first-class--codex-reasoning-effortoption in the task/run CLI.Why change
A low-effort Codex starter-suite baseline would show how much of the current 60/60 Codex result depends on
xhighreasoning effort. This is the fastest way to make the comparison matrix more useful for practical tool-fit decisions around cost, latency, reliability, and review burden.One issue should be enough to track this because it is the same agent harness and same model family as the completed Codex baseline. If effort plumbing becomes larger than expected, split that implementation into a child issue before trial collection.
What to build
Collect a Codex low-effort starter-suite baseline for the same 12 tasks and five-fair-trials-per-task shape.
Preferred target configuration:
codexgpt-5.5low1800seconds unless readiness reveals a fairness issueImplementation phases:
lowreasoning effort portably.model_name = gpt-5.5andreasoning_effort = lowthrough recovered runtime metadata or another trustworthy source.Acceptance criteria
gpt-5.5withlowreasoning effort through a portable mechanism, or the issue documents why a profile-based mechanism was used.xhighCodex baseline.evidence-sets/.reports/codex-low-effort-.../bundle.xhighCodex baseline.Stop conditions
loweffort.Out of scope