Skip to content

[v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86

Description

@ayhammouda

Context

Parent: #63. Methodology: docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md (work package 3). Runner: benchmarks/runner.py (merged via #75). Model matrix + provider adapters: benchmarks/model_matrix.py, benchmarks/adapters/ (merged via #85).

Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, resolve the one design decision below, create the .planning/agent-context/ file (decision 5.14, mirrored as an issue comment), and apply agent-ready before any dispatch.

Goal

Add the benchmark adapter that runs python-docs-mcp-server itself as a system under test: offline stdio invocation of the local server, per-question docs retrieval, response capture into the runner's artifact shapes.

Design decision required from maintainer (blocks agent-ready)

Cell composition semantics for tool×model. Recommended: no runner core change — the competitor manifest enumerates one entry per tool×model pairing (the runner's _tool_model_key already renders id:provider/model keys), and benchmarks/model_matrix.py gains a validator that checks manifest entries against matrix families. Alternative (rejected by default): change the runner's cell shape to competitor×model×question, which would rework merged artifact layouts and tests. Confirm or override before labeling agent-ready.

Acceptance criteria (draft — Vision edits before agent-ready)

  • A python-docs-mcp adapter in benchmarks/adapters/ launches the local server over stdio (offline, no network — principle 2.2 preserved at runtime), issues the documented retrieval flow (search_docs → get_docs) per corpus question, and records raw tool payloads into the transcript artifact.
  • Retrieval requires a built local index; the adapter fails with a recorded tool_failure (not a crash) when the index is missing, and the dry-run path validates without needing one.
  • Manifest↔matrix validation: a manifest entry whose provider/model pair is absent from docs/benchmarks/model-matrix.yml fails validation with a clean BenchmarkValidationError.
  • Tests (under tests/benchmarks/, additive only) cover the adapter against a fake/stubbed MCP transport — no real server spawn required in unit tests; one optional integration test may spawn the real server if the local index exists, skipped otherwise.
  • No README/benchmark claims; Refs #63, never Closes #63.

Scope boundaries

In scope: the product's own tool adapter, manifest/matrix validation, mocked tests. Out of scope: LLM/provider calls (adapters from #85 handle that in the live phase), competitor MCP adapters (separate issue), scoring, reporting, README.

Forbidden-territory reminder

Standard §2 list applies; additionally: merged tests in tests/benchmarks/ are regression cover — additive only. MCP tool names/parameters/return shapes of the server itself must not change to accommodate the benchmark.

Validation commands

uv run ruff check src/ tests/
uv run ruff check benchmarks/
uv run pyright src/
uv run pyright benchmarks/
uv run pytest --tb=short -q
uv run pytest tests/benchmarks -q
uv run python-docs-mcp-server doctor

Recovery

If stdio invocation or artifact composition needs anything the merged runner does not support, stop and comment here rather than modifying benchmarks/runner.py beyond the issue's prescribed scope.

Effort estimate

4–8 hours.

Refs #63.

Activity

  1. added this to the v0.5.0 milestone on Jul 8, 2026
  2. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Maintainer decision — cell composition CONFIRMED (2026-07-08)

    The embedded design decision in this issue's body is confirmed as recommended, with one code-verified precision (PLAN.md orchestrator note 1, Codex-reviewed):

    • The competitor manifest enumerates one entry per tool×model pairing; the runner's _tool_model_key (benchmarks/runner.py:311-318) already renders id:provider/model keys.
    • benchmarks/model_matrix.py gains a manifest↔matrix validator: a manifest entry whose provider/model pair is absent from docs/benchmarks/model-matrix.yml fails with a clean BenchmarkValidationError.
    • No cell-shape or artifact-layout change — cells stay competitor×question; the rejected alternative (competitor×model×question) would rework merged, test-pinned artifact layouts and report.py join keys.
    • Precision: "no runner core change" does not mean zero runner edits — the adapter must be dispatchable, so a minimal dispatch-registry change at the _execute_cell seam (replacing the hardcoded _EXECUTABLE_ADAPTERS allowlist check, runner.py:23, 288-292) is in scope. Cell construction (runner.py:85-89) is untouched.

    Pre-flight amendments to the acceptance criteria (maintainer-approved, sign-off mirrored on #63):

    • Added criterion (D5): flip docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md line Status: Draft to Status: Accepted (blessed 2026-07-08; see PLAN-2026-07-08-benchmark-plumbing.md Amendment 1) — one line, rides this PR.
    • Named merged-test sanction (D6a): this issue MAY update the known-gap disclosure strings at report.py:26-32, report.py:455-460, report.py:664-667, the stale docstrings at runner.py:3-4 and model_matrix.py:8-15, and the exact tests/benchmarks/test_report.py assertions that pin the known-gap paragraph — updating a pinned string to the new sanctioned text is in scope. Any other merged-test change remains stop-and-comment.

    Refs #63

  3. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Agent context — #86 [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio)

    Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
    Auditable copy: this content is mirrored as a comment on issue #86 (.planning/ is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).

    1. Roadmap excerpt (why this exists)

    • STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness — all target docs MCPs + no-MCP baseline, 50-question eval, reproducible, methodology disclosure mandatory. This issue makes the product itself a system under test.
    • Decision 5.17: no comparative/benchmark claim enters README/PyPI/launch copy before reproducible data exists. Your PR makes zero claims.
    • Cell composition CONFIRMED (maintainer comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86, 2026-07-08): the competitor manifest enumerates one entry per tool×model pairing; benchmarks/model_matrix.py gains a manifest↔matrix validator (manifest provider/model pair absent from docs/benchmarks/model-matrix.yml → clean BenchmarkValidationError); NO cell-shape or artifact-layout change — cells stay competitor×question. "No runner core change" does not mean zero runner edits: a minimal dispatch-registry change at the _execute_cell seam is in scope; cell construction is untouched.
    • Amendment D5 rides this PR: flip the docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md Status line (currently **Status:** Draft methodology for issue [#63], lines 3-4) to Status: Accepted (blessed 2026-07-08; see PLAN-2026-07-08-benchmark-plumbing.md Amendment 1) — one line, nothing else in that file.

    2. Code touch-points (verify against merged main before starting)

    • benchmarks/runner.py:23 — _EXECUTABLE_ADAPTERS = {"fake", "no-mcp-baseline", "no_mcp_baseline"}; the allowlist check lives in _fake_provider_answer (runner.py:287-300, gate at 288-292). Replace this hardcoded gate with a minimal dispatch registry so the product adapter is dispatchable from _execute_cell (runner.py:218-284).
    • benchmarks/runner.py:85-89 — cell construction (competitor×question) — UNTOUCHED.
    • benchmarks/runner.py:311-318 — _tool_model_key already renders id:provider/model; do not change it — the manifest entry carries the pairing.
    • benchmarks/adapters/base.py:78-110 — ProviderAdapter contract: implement _generate(); generate() (89-105) wraps latency measurement + failure capture (BenchmarkCellFailure, categories tool_failure | timeout | mcp_protocol_crash).
    • NEW file under benchmarks/adapters/ — the stdio product adapter: spawn the local server over stdio (offline, no network — principle 2.2 at runtime), run the documented search_docs → get_docs flow per corpus question, record raw tool payloads into the transcript artifact. Missing local index → recorded tool_failure, never a crash; dry-run path validates indexless; timeouts map to category "timeout".
    • benchmarks/model_matrix.py — load_model_matrix (line 73), tool_model_cells (line 175); add the manifest↔matrix validator here. The stale deferral docstring at lines 8-15 ("wiring ... is deferred to a follow-up issue") gets updated — D6a-sanctioned.
    • benchmarks/runner.py:357-376 — _write_cell_artifacts defines the per-cell artifact shapes (transcripts|tokens|latency|scoring|failures/<competitor>/<qid>.json) your results must fit; raw tool payloads go inside the transcript record, no new artifact directories.
    • D6a-sanctioned stale-text updates (ONLY these): benchmarks/report.py:26-31 module-docstring known-gap paragraph; report.py:453-460 "Known gap (inherited, not fixed by this report)" block; report.py:663-667 "No cross-model aggregation applies to this run..." block; runner.py:3-4 stale docstring; model_matrix.py:8-15 docstring; and the exact tests/benchmarks/test_report.py assertions pinning those strings — test_report.py:145 ("Model / Client Matrix") and test_report.py:538 ("No cross-model aggregation applies to this run"). Updating a pinned string to the new sanctioned text is in scope. Any other merged-test change: stop and comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86.

    3. Test patterns to follow

    • Selective subprocess.run monkeypatch that fakes only the matching call and delegates the rest to the real function: tests/benchmarks/test_runner.py:140-157.
    • CLI exercised as a real subprocess via sys.executable -m benchmarks: test_runner.py:353-372.
    • Hand-authored minimal-but-valid run-dir factory: tests/benchmarks/test_report.py:232-295 (_minimal_valid_run_dir).
    • Unit tests drive the adapter against a fake/stubbed MCP transport ONLY — no real server spawn. One optional integration test may spawn the real server when the local index exists, skipped otherwise. All new tests under tests/benchmarks/ (additive).

    4. Known pitfalls

    • Merged tests in tests/benchmarks/ are regression cover: additive-only EXCEPT the named D6a strings in §2.
    • Lint blind spot until [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84 lands: CI does not run lint over benchmarks/ — additionally run uv run ruff check benchmarks/ and uv run pyright benchmarks/ yourself.
    • Diff >500 lines → apply the supervisor-review label + an explanation section in the PR body.
    • pyproject.toml and uv.lock are untouchable this run — the adapter must work with existing dependencies.
    • Refs #63, never Closes #63; Closes #86 only if every acceptance criterion is met.
    • The server's MCP tool surface is a CLIENT dependency here — never modify the server's tool names, parameters, or return shapes to accommodate the benchmark.
    • If stdio invocation or artifact composition needs anything the merged runner does not support beyond the prescribed _execute_cell seam, stop and comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 (issue recovery clause) instead of widening runner scope.
    • Full gate before merge decision: uv run ruff check src/ tests/ benchmarks/, uv run pyright src/ benchmarks/, uv run pytest --tb=short -q, uv run pytest tests/benchmarks -q, uv run python-docs-mcp-server doctor.
    • Branch: agent/86-product-runner-adapter.

    5. Decision log (worker fills in)

    • …
  4. added
    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
    on Jul 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestpriority:P2Medium priority

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions