Skip to content

[v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73

Description

@ayhammouda

Context

Parent: #63. Methodology: docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md.

Aymen explicitly wants OpenAI and Google model families included. This issue defines the model/client matrix and provider adapter contracts without letting model choice blur into a single fake tool-quality score.

Goal

Add the benchmark model matrix and provider adapter contracts for OpenAI and Google-backed runs, with tests that use mocks rather than paid/live calls.

Acceptance criteria

  • The runner/model design cross-multiplies the evaluation matrix: every eligible tool is evaluated against every model in the matrix independently.
  • Tool scores are not averaged across different model families; any aggregate must remain secondary and clearly labeled.
  • Token metrics are labeled as Claude Tokens (Normalized Payload) when using the methodology tokenizer, so they are not confused with OpenAI or Google API billing tokens.
  • docs/benchmarks/model-matrix.yml defines the model/client families to run, including at least one OpenAI-backed and one Google-backed model family.
  • The matrix format records provider, model ID, client/wrapper, token-count method, latency scope, and whether the result is eligible for headline claims.
  • Provider adapter interfaces support request, response, latency, token metadata, and failure capture.
  • Tests cover mock OpenAI and mock Google adapter behavior without making network calls.
  • The report design keeps per-model results visible and marks any aggregate as aggregate.
  • The implementation refuses to run live providers unless explicit environment variables/config are present.

Scope boundaries

In scope:

  • Model matrix spec.
  • Provider adapter interfaces.
  • Mock adapter tests.
  • Live-call guardrails.

Out of scope:

  • Running the full benchmark.
  • Spending API credits.
  • Adding README claims.
  • Choosing final headline wording.

Forbidden-territory reminder

Do not modify MCP tool names, parameters, return shapes, schema.sql, .github/workflows/, pyproject.toml project metadata, .planning/POSITIONING.md, the README hero section, LICENSE, SECURITY.md, or existing tests by weakening/deleting assertions.

If a new third-party runtime dependency is proposed, trigger supervisor review and justify why stdlib/httpx is insufficient.

Validation commands

uv run ruff check src/ tests/
uv run pyright src/
uv run pytest --tb=short -q
uv run python-docs-mcp-server doctor

Run any new adapter tests directly as well.

PR template

Use Refs #63, not Closes #63.

The PR must include:

  • Model matrix summary.
  • Explanation of live-call guardrails.
  • Test output.
  • Confirmation that no paid/live provider calls were made.

Recovery

If token counting cannot be made comparable across providers, stop and comment with the exact mismatch instead of averaging incompatible numbers.

Effort estimate

4-6 hours. This can be agent-ready for mocked adapter plumbing only; live benchmark execution remains Vision-controlled.

Activity

  1. added
    enhancementNew feature or request
    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
    on Jun 8, 2026
  2. added this to the v0.5.0 milestone on Jul 8, 2026
  3. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Agent context — #73 [v0.5.0] benchmark adapters — define OpenAI/Google model matrix

    Status: FINAL (verified against merged main @ ad18406, PR #75 squash-merged 2026-07-08). Merge-readiness fixes now on main: snapshots are byte-for-byte copies via _write_snapshot; a non-empty --out dir is refused with BenchmarkValidationError (exit 2); yaml.YAMLError re-raises as BenchmarkValidationError.
    Auditable copy: this content is mirrored as a comment on issue #73 (.planning/ is gitignored — the issue comment is the copy pre-flight verifies and dispatch links).

    1. Roadmap excerpt (why this exists)

    STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): "Public benchmark harness — all eligible target docs MCPs + no-MCP baseline; 50-question Python eval… Reproducible from a clean clone. Methodology disclosure mandatory."
    Decision 5.8: token methodology = Claude tokenizer, measured after client-side rewrap, tokens AND latency reported.
    Decision 5.17 (evidence ladder): no comparative claim in any public copy until the harness produces reproducible data.
    Maintainer decision (PLAN.md Amendment 2026-07-08): Claude token counting uses the Anthropic count-tokens API confined to the maintainer-run live phase — your adapter contracts record token_count_method per matrix entry, but no adapter you write ever calls that API.

    2. Code touch-points (verify against merged main before starting)

    • benchmarks/runner.py — the runner core you extend:
      • _EXECUTABLE_ADAPTERS = {"fake", "no-mcp-baseline", "no_mcp_baseline"} — the adapter dispatch set.
      • _execute_cell(cell) / _fake_provider_answer(cell) — per-cell execution; failure capture via BenchmarkCellFailure(category, message) with categories tool_failure | timeout | mcp_protocol_crash.
      • Competitor dataclass — manifest entries carry id, adapter, and free-form raw (e.g. provider, model, fake_answer, force_failure).
      • _tool_model_key(competitor) — produces id:provider/model keys; your matrix must keep tool×model cells separate (no cross-family averaging; aggregates secondary and labeled).
    • benchmarks/__main__.py — argparse subparser tree (run subcommand). Add adapter/matrix wiring conservatively; T4 will add a sibling report module later — keep __main__.py churn minimal.
    • Artifact layout the runner writes per run dir: snapshots/, environment.json, planned-cells.json, run-summary.json, and per-cell transcripts|tokens|latency|scoring|failures/<competitor>/<qid>.json. Token records currently carry client_wrapped_tokens: None placeholders — your interfaces define where real values would go, mocks only.
    • New file: docs/benchmarks/model-matrix.yml — ≥1 OpenAI-backed and ≥1 Google-backed family; each entry records provider, model ID, client/wrapper, token-count method, latency scope, headline-eligibility.

    3. Test patterns to follow

    tests/benchmarks/test_runner.py: module-level helpers _write_yaml/_corpus/_manifest build YAML fixtures under tmp_path; tests call run_benchmark(BenchmarkConfig(...)) directly and assert artifact paths/JSON content. Follow the same shape for adapter tests. New tests MUST live under tests/benchmarks/ (so uv run pytest tests/benchmarks -q exercises them). Mock OpenAI and mock Google adapters — zero network calls; the implementation must refuse live providers unless explicit env/config is present, and tests must cover that refusal.

    4. Known pitfalls

    • Lint blind spot: CI and the canonical gate do NOT cover benchmarks/ — you must also run uv run ruff check benchmarks/ and uv run pyright benchmarks/ (see [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84).
    • Merged runner tests are §2 regression cover — additive changes only; if your change appears to require modifying/weakening any merged runner test, stop and comment per pipeline §2/§8.
    • Token labeling: methodology tokenizer counts are labeled Claude Tokens (Normalized Payload) — never confusable with OpenAI/Google billing tokens.
    • New runtime dependency = supervisor-review trigger. Expected answer: none needed (PyYAML already a runtime dep; mocks need no SDKs).
    • PR conventions: Refs #63 (never Closes #63), Closes #73 only if all criteria met. Diff >500 lines → supervisor-review label + explanation section.
    • Sequencing: this task merges BEFORE [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74's reporting task (its report consumes your model-matrix artifact).

    5. Decision log (worker fills in)

    • …
  4. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Decision log (worker T3)

    PR opened: #85 (agent/73-benchmark-adapter-matrix → main).

    • Added docs/benchmarks/model-matrix.yml (2 OpenAI + 2 Google model families) and benchmarks/model_matrix.py (loader + tool_model_cells() cross-multiplication contract).
    • Added benchmarks/adapters/ (request/response/latency/token-metadata/failure-capture interface, mock OpenAI/Google adapters, live-provider refusal guardrail). No SDK or network library imported anywhere; live adapters are guarded stubs that always raise LiveExecutionNotImplementedError even when guard env vars are set.
    • Did NOT wire the model matrix into benchmarks/runner.py's _execute_cell dispatch or touch benchmarks/__main__.py — see PR [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85's "Why this approach" section for the reasoning (avoiding an unprescribed design guess against the merged, byte-asserted tests/benchmarks/test_runner.py artifact shapes). Zero diff on runner.py, __main__.py, and test_runner.py, confirmed via git diff ad18406 --name-only.
    • 8/9 acceptance criteria met; the "report design keeps per-model results visible" criterion is unchecked and deferred to [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 (report generation), so PR [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85 uses Refs #63 + does not use Closes #73.
    • Full validation gate (7 commands) passes; output pasted in the PR body.
    • supervisor-review label added: diff is 1088 total lines / 582 source-only lines across the new benchmarks/model_matrix.py + benchmarks/adapters/*.py, over the 500-line §7 threshold. No forbidden-territory file touched, no existing test modified, no new runtime dependency, no pyproject.toml change.

    I am not stopping pending guidance — this is a normal §7 supervisor-review pause, not a blocker. Ready for Vision/maintainer review.

  5. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Closing: PR #85 merged (1f202b6) delivers 8/9 acceptance criteria. The remaining criterion — report design keeping per-model results visible with aggregates marked as aggregates — is #74's deliverable by design (methodology work-package sequencing). The deliberate runner-wiring deferral (composition of tool×model cells with the runner's tool×question artifact shapes) is tracked for the maintainer-filed WP3 product-runner-adapter issue or a scoped follow-up; see the decision-log comment above. Merge performed by the orchestrator under the maintainer's explicit delegation recorded in this session (PLAN.md Amendment 2026-07-08).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestpriority:P2Medium priority

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions