Repository navigation
[v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priorityagent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
on Jun 8, 2026 Agent context — #73 [v0.5.0] benchmark adapters — define OpenAI/Google model matrix
Status: FINAL (verified against merged
main@ad18406, PR #75 squash-merged 2026-07-08). Merge-readiness fixes now on main: snapshots are byte-for-byte copies via_write_snapshot; a non-empty--outdir is refused withBenchmarkValidationError(exit 2);yaml.YAMLErrorre-raises asBenchmarkValidationError.
Auditable copy: this content is mirrored as a comment on issue #73 (.planning/is gitignored — the issue comment is the copy pre-flight verifies and dispatch links).1. Roadmap excerpt (why this exists)
STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): "Public benchmark harness — all eligible target docs MCPs + no-MCP baseline; 50-question Python eval… Reproducible from a clean clone. Methodology disclosure mandatory."
Decision 5.8: token methodology = Claude tokenizer, measured after client-side rewrap, tokens AND latency reported.
Decision 5.17 (evidence ladder): no comparative claim in any public copy until the harness produces reproducible data.
Maintainer decision (PLAN.md Amendment 2026-07-08): Claude token counting uses the Anthropic count-tokens API confined to the maintainer-run live phase — your adapter contracts recordtoken_count_methodper matrix entry, but no adapter you write ever calls that API.2. Code touch-points (verify against merged main before starting)
benchmarks/runner.py— the runner core you extend:_EXECUTABLE_ADAPTERS = {"fake", "no-mcp-baseline", "no_mcp_baseline"}— the adapter dispatch set._execute_cell(cell)/_fake_provider_answer(cell)— per-cell execution; failure capture viaBenchmarkCellFailure(category, message)with categoriestool_failure | timeout | mcp_protocol_crash.Competitordataclass — manifest entries carryid,adapter, and free-formraw(e.g.provider,model,fake_answer,force_failure)._tool_model_key(competitor)— producesid:provider/modelkeys; your matrix must keep tool×model cells separate (no cross-family averaging; aggregates secondary and labeled).
benchmarks/__main__.py— argparse subparser tree (runsubcommand). Add adapter/matrix wiring conservatively; T4 will add a siblingreportmodule later — keep__main__.pychurn minimal.- Artifact layout the runner writes per run dir:
snapshots/,environment.json,planned-cells.json,run-summary.json, and per-celltranscripts|tokens|latency|scoring|failures/<competitor>/<qid>.json. Token records currently carryclient_wrapped_tokens: Noneplaceholders — your interfaces define where real values would go, mocks only. - New file:
docs/benchmarks/model-matrix.yml— ≥1 OpenAI-backed and ≥1 Google-backed family; each entry records provider, model ID, client/wrapper, token-count method, latency scope, headline-eligibility.
3. Test patterns to follow
tests/benchmarks/test_runner.py: module-level helpers_write_yaml/_corpus/_manifestbuild YAML fixtures undertmp_path; tests callrun_benchmark(BenchmarkConfig(...))directly and assert artifact paths/JSON content. Follow the same shape for adapter tests. New tests MUST live undertests/benchmarks/(souv run pytest tests/benchmarks -qexercises them). Mock OpenAI and mock Google adapters — zero network calls; the implementation must refuse live providers unless explicit env/config is present, and tests must cover that refusal.4. Known pitfalls
- Lint blind spot: CI and the canonical gate do NOT cover
benchmarks/— you must also runuv run ruff check benchmarks/anduv run pyright benchmarks/(see [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84). - Merged runner tests are §2 regression cover — additive changes only; if your change appears to require modifying/weakening any merged runner test, stop and comment per pipeline §2/§8.
- Token labeling: methodology tokenizer counts are labeled
Claude Tokens (Normalized Payload)— never confusable with OpenAI/Google billing tokens. - New runtime dependency = supervisor-review trigger. Expected answer: none needed (PyYAML already a runtime dep; mocks need no SDKs).
- PR conventions:
Refs #63(neverCloses #63),Closes #73only if all criteria met. Diff >500 lines →supervisor-reviewlabel + explanation section. - Sequencing: this task merges BEFORE [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74's reporting task (its report consumes your model-matrix artifact).
5. Decision log (worker fills in)
- …
Decision log (worker T3)
PR opened: #85 (
agent/73-benchmark-adapter-matrix→main).- Added
docs/benchmarks/model-matrix.yml(2 OpenAI + 2 Google model families) andbenchmarks/model_matrix.py(loader +tool_model_cells()cross-multiplication contract). - Added
benchmarks/adapters/(request/response/latency/token-metadata/failure-capture interface, mock OpenAI/Google adapters, live-provider refusal guardrail). No SDK or network library imported anywhere; live adapters are guarded stubs that always raiseLiveExecutionNotImplementedErroreven when guard env vars are set. - Did NOT wire the model matrix into
benchmarks/runner.py's_execute_celldispatch or touchbenchmarks/__main__.py— see PR [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85's "Why this approach" section for the reasoning (avoiding an unprescribed design guess against the merged, byte-assertedtests/benchmarks/test_runner.pyartifact shapes). Zero diff onrunner.py,__main__.py, andtest_runner.py, confirmed viagit diff ad18406 --name-only. - 8/9 acceptance criteria met; the "report design keeps per-model results visible" criterion is unchecked and deferred to [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 (report generation), so PR [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85 uses
Refs #63+ does not useCloses #73. - Full validation gate (7 commands) passes; output pasted in the PR body.
supervisor-reviewlabel added: diff is 1088 total lines / 582 source-only lines across the newbenchmarks/model_matrix.py+benchmarks/adapters/*.py, over the 500-line §7 threshold. No forbidden-territory file touched, no existing test modified, no new runtime dependency, nopyproject.tomlchange.
I am not stopping pending guidance — this is a normal §7 supervisor-review pause, not a blocker. Ready for Vision/maintainer review.
- Added
- added a commit that references this issue
on Jul 8, 2026 Closing: PR #85 merged (1f202b6) delivers 8/9 acceptance criteria. The remaining criterion — report design keeping per-model results visible with aggregates marked as aggregates — is #74's deliverable by design (methodology work-package sequencing). The deliberate runner-wiring deferral (composition of tool×model cells with the runner's tool×question artifact shapes) is tracked for the maintainer-filed WP3 product-runner-adapter issue or a scoped follow-up; see the decision-log comment above. Merge performed by the orchestrator under the maintainer's explicit delegation recorded in this session (PLAN.md Amendment 2026-07-08).
Context
Parent: #63. Methodology:
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md.Aymen explicitly wants OpenAI and Google model families included. This issue defines the model/client matrix and provider adapter contracts without letting model choice blur into a single fake tool-quality score.
Goal
Add the benchmark model matrix and provider adapter contracts for OpenAI and Google-backed runs, with tests that use mocks rather than paid/live calls.
Acceptance criteria
Claude Tokens (Normalized Payload)when using the methodology tokenizer, so they are not confused with OpenAI or Google API billing tokens.docs/benchmarks/model-matrix.ymldefines the model/client families to run, including at least one OpenAI-backed and one Google-backed model family.Scope boundaries
In scope:
Out of scope:
Forbidden-territory reminder
Do not modify MCP tool names, parameters, return shapes,
schema.sql,.github/workflows/,pyproject.tomlproject metadata,.planning/POSITIONING.md, the README hero section,LICENSE,SECURITY.md, or existing tests by weakening/deleting assertions.If a new third-party runtime dependency is proposed, trigger supervisor review and justify why stdlib/httpx is insufficient.
Validation commands
Run any new adapter tests directly as well.
PR template
Use
Refs #63, notCloses #63.The PR must include:
Recovery
If token counting cannot be made comparable across providers, stop and comment with the exact mismatch instead of averaging incompatible numbers.
Effort estimate
4-6 hours. This can be agent-ready for mocked adapter plumbing only; live benchmark execution remains Vision-controlled.