Repository navigation
PRD: Codex deep baseline and evidence-scoped capability reports #10
Description
Activity
- addedready-for-agentFully specified and ready for an AFK agentFully specified and ready for an AFK agent
on May 8, 2026 Status Update - 2026-05-09
mainhas been fast-forwarded through the starter-suite expansion work. The Solo Dev Starter Suite now has 9 publishable task bundles, each with a pinned public repo/commit, generated task card, reference artifact, and passing committed reference verification artifact:2048-advanced-snake-params-001- Python simulation metadata bugfixclick-default-map-nargs-001- Python Click default_map regressionclick-help-option-refactor-001- Python Click production-code refactorclick-help-shadowed-option-001- Python Click help hint regressionclick-should-strip-ansi-tests-001- Python test-writing taskdatawrapper-mcp-docker-requirements-001- deterministic dependency/setup taskhttpx-verify-false-client-cert-001- ambiguous edge-case behavior/default taskreact-tabs-selected-focus-overlay-001- React/CSS visual UI task with human visual review noted in Issue Add React UI visual task to Solo Dev Starter Suite #7todomvc-toggle-all-checkbox-001- JavaScript frontend DOM/state task
Current validation from
main:python3 -m agentlab task validate tasks/starterpassed for 9 starter tasks.python3 .agents/skills/task-card/scripts/render_task_cards.py tasks --checkpassed.python3 -m unittest discoverpassed: 93 tests.- All committed
tasks/starter/*/reference-result.jsonartifacts reportstatus: passed/success: true.
Smoke evidence status:
- Issue comments for Add behavior-preserving refactor task to Solo Dev Starter Suite #4-Add ambiguous product behavior task to Solo Dev Starter Suite #9 report one fair passing Codex smoke trial for each new task.
- Local ignored
runs/artifacts in this checkout currently include the TodoMVC smoke run plus the earlier 2048 and Click repeated evidence; the other new smoke artifacts are referenced in issue comments but are not durable repo artifacts.
PRD progress:
- The previous "next high-leverage implementation" in this PRD, explicit invalid/excluded trial handling, is complete.
- The starter suite breadth milestone is now substantially complete: it covers Python, JavaScript, frontend DOM/state, visual UI, test-writing, setup/dependency, refactor, ambiguous behavior/defaults, and several realistic bug/regression tasks.
Recommended next milestone:
Run the Codex deep-baseline batch across the full starter suite:
- For each task that has only smoke evidence, confirm the single-trial path is still fair from current
main. - Run bounded repeated Codex trials to reach the PRD target of 5 fair trials per task.
- Human-review suspicious, failed, messy, or resource-heavy trials; exclude only invalid trials with explicit reasons.
- Create explicit evidence-set manifests for the report input set.
- Only after the 9-task evidence set is reviewed, draft the first evidence-scoped Codex capability report.
I recommend keeping this PRD open until the full repeated-trial evidence set and first hand-authored capability report are complete. The immediate ready-for-agent work is the repeated fair Codex batch and evidence-set preparation, not more task curation.
- added a commit that references this issue
on May 9, 2026 Starter-suite deep baseline collected
Completed the Issue #10 starter-suite evidence pass and pushed commit 679d626 to main.
Artifacts now on main:
- Evidence set: evidence-sets/codex-starter-suite-deep-baseline-2026-05-09.json
- Capability evidence digest: reports/codex-starter-suite-deep-baseline-2026-05-09-digest.md
- HTML overnight report: reports/codex-starter-suite-overnight-report-2026-05-09.html
- Narrative report: docs/codex-starter-suite-deep-baseline-report.md
Result:
- 9 starter tasks reached the target of 5 fair Codex trials each.
- Selected evidence set contains 45 fair trials.
- All 45 selected fair trials passed deterministic graders.
- Human review labels in the selected set: success_clean:41, resource_inefficient:3, success_messy:1; plus 4 additional successful Click help-option refactor trials carry secondary resource_inefficient labels.
Caveats:
- The Click help-option refactor passed consistently but was resource-heavy across all selected fair trials.
- Two React Tabs trials produced correct tiny style patches but took disproportionate wall-clock time.
- One Click ANSI test trial was valid but messy because it deleted an existing Jupyter-specific smoke test while adding the broader matrix.
- Excluded invalid artifacts remain outside the fair set: earlier click-default-map-nargs-001 setup errors and one click-help-shadowed-option-001 eval harness error.
Interpretation: this is a strong starter-suite Codex baseline, not yet a broad general capability report. Next PRD focus should be adding more varied tasks before making broader claims.
- added a commit that references this issue
on May 9, 2026 - removedready-for-agentFully specified and ready for an AFK agentFully specified and ready for an AFK agent
on May 12, 2026 Status Update - 2026-05-11
Main is clean after the starter-suite expansion and branch cleanup. The Solo Dev Starter Suite now has 12 publishable task bundles on origin/main:
- 2048-advanced-snake-params-001
- click-default-map-nargs-001
- click-help-option-refactor-001
- click-help-shadowed-option-001
- click-should-strip-ansi-tests-001
- datawrapper-mcp-docker-requirements-001
- httpx-verify-false-client-cert-001
- prettier-duplicate-dangling-comments-001
- react-tabs-selected-focus-overlay-001
- remotion-audio-context-autoplay-muted-001
- todomvc-toggle-all-checkbox-001
- vite-deno-workspace-root-001
Completed since the previous PRD status:
- The reporting integrity gap was fixed so secondary human review labels are surfaced in reports.
- Task-authoring conflicts were reduced by moving each task bundle into its own isolated test file/fixture surface.
- Task candidate tracking was moved to GitHub Issues rather than local candidate documents or suite README indexes.
- Three additional realistic JavaScript/TypeScript task bundles were added, verified, smoke-tested, merged to main, and their implementation issues closed only after origin/main contained the resolving commits.
- Obsolete merged codex branches and stale worktrees were cleaned up.
Current evidence state:
- The existing 2026-05-09 deep baseline covers the original 9 starter tasks with 45 selected fair Codex trials.
- The 3 newer tasks each have reference verification and one passing Codex smoke trial, but do not yet have the PRD target of 5 fair trials per task.
- Therefore the current evidence gap is bounded: collect and review repeated fair Codex trials for Prettier, Vite, and Remotion before updating the evidence set or digest.
Recommended next PRD focus:
- Run bounded repeated Codex trials for the 3 new starter tasks only, targeting 5 fair trials per task.
- Human-review suspicious, failed, messy, or resource-heavy trials before including them in any selected evidence set.
- Exclude only invalid trials with explicit reasons such as dependency_issue or eval_harness_error.
- Update or create an evidence set and capability evidence digest that reflects all 12 starter tasks.
- Defer broader capability-report claims until the suite has more breadth and the new 12-task evidence set has been reviewed.
The PRD issue is no longer labeled ready-for-agent because the next implementation slice should be discussed before another AFK agent picks it up.
- added a commit that references this issue
on May 12, 2026 Status Update - 2026-05-11
The 12-task starter-suite evidence set is now durable on main in commit 65f0ff9.
Added artifacts:
- evidence-sets/codex-starter-suite-12-task-baseline-2026-05-11.json
- reports/codex-starter-suite-12-task-baseline-2026-05-11-digest.md
Evidence shape:
- 60 selected fair Codex CLI trials total.
- 12 starter tasks, 5 fair trials per task.
- All 60 selected fair trials passed deterministic graders.
- The older pre-tightening Vite smoke run remains excluded from the selected evidence set and is marked invalid_task in local run metadata.
- Prettier and Remotion are included as passing but resource_inefficient across their selected trials.
Validation before commit:
- Generated digest from the explicit evidence set.
- Confirmed manifest has 60 unique trials and does not include the old pre-tightening Vite run.
- python3 -m unittest tests.test_evidence_sets tests.test_evidence tests.test_results passed.
- Pre-commit task-card and task validation checks passed.
Closing this as completed.
The Codex deep baseline now has the durable evidence and reader-facing report artifacts this PRD called for:
- 12-task Codex baseline report:
reports/codex-starter-suite-12-task-baseline-2026-05-11/report.md - Selected evidence set:
evidence-sets/codex-starter-suite-12-task-baseline-2026-05-11.json - Generated evidence digest:
reports/codex-starter-suite-12-task-baseline-2026-05-11/digest.md - Model attribution note:
reports/codex-starter-suite-12-task-baseline-2026-05-11/model-attribution.md
Summary: 60 selected fair Codex CLI trials, 12 starter-suite tasks, 5 trials per task, 60/60 deterministic grader passes, with model/effort recovered as
gpt-5.5/xhigh. The report also keeps the evidence scoped and calls out resource-inefficiency caveats rather than treating pass rate as the whole story.Follow-on work now belongs in the Claude Code baseline and later Codex-vs-Claude comparison PRDs.
- 12-task Codex baseline report:
Problem Statement
Agent Eval Lab has grown from a starter harness into the early shape of a credible coding-agent evaluation lab. It can run tasks, invoke Codex CLI, capture outcomes, verify reference artifacts, summarize repeated trials, and record human reviews. However, the project still needs a coherent next-phase product plan so work does not scatter across task curation, agent integration, reporting, environment setup, and reliability metrics.
From the user's perspective, the core problem is: "I want to understand, with evidence, what coding-agent harnesses do well, poorly, or inconsistently on realistic software engineering tasks, starting with Codex CLI. I want the output to be useful to technical evaluators and to solo developers deciding which AI coding tool fits their use case. I do not want the lab to overgeneralize beyond observed evidence."
The current state has also surfaced important operational lessons:
Solution
Build the next phase around a Codex deep baseline for the Solo Dev Starter Suite.
The lab should curate realistic task bundles, verify each task with a reference artifact, smoke-test each task with one Codex trial, then run bounded repeated Codex trials only after the single-trial path is fair. The output should be an evidence-scoped Markdown agent capability report that explains what Codex CLI did well, poorly, and inconsistently under the evaluated conditions.
The solution should continue to follow Anthropic-aligned terminology: tasks, trials, evaluation suites, agent harnesses, graders/assertions, traces/transcripts, outcomes, reference verification, and pass@k/pass^k. It should keep "agent harness" separate from "underlying model" and avoid global claims such as "model Y has capability Z" unless the evidence supports that scope.
The project should also strengthen the harness around invalid/excluded trials, task-local environments, outcome evidence, and report generation so future task batches are more trustworthy and easier to interpret.
User Stories
Implementation Decisions
Testing Decisions
Out of Scope
Further Notes