Repository navigation
[v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86
Copy link
Copy link
Closed
Labels
agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
Milestone
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
on Jul 8, 2026 Maintainer decision — cell composition CONFIRMED (2026-07-08)
The embedded design decision in this issue's body is confirmed as recommended, with one code-verified precision (PLAN.md orchestrator note 1, Codex-reviewed):
- The competitor manifest enumerates one entry per tool×model pairing; the runner's
_tool_model_key(benchmarks/runner.py:311-318) already rendersid:provider/modelkeys. benchmarks/model_matrix.pygains a manifest↔matrix validator: a manifest entry whose provider/model pair is absent fromdocs/benchmarks/model-matrix.ymlfails with a cleanBenchmarkValidationError.- No cell-shape or artifact-layout change — cells stay competitor×question; the rejected alternative (competitor×model×question) would rework merged, test-pinned artifact layouts and
report.pyjoin keys. - Precision: "no runner core change" does not mean zero runner edits — the adapter must be dispatchable, so a minimal dispatch-registry change at the
_execute_cellseam (replacing the hardcoded_EXECUTABLE_ADAPTERSallowlist check,runner.py:23,288-292) is in scope. Cell construction (runner.py:85-89) is untouched.
Pre-flight amendments to the acceptance criteria (maintainer-approved, sign-off mirrored on #63):
- Added criterion (D5): flip
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.mdlineStatus: DrafttoStatus: Accepted (blessed 2026-07-08; see PLAN-2026-07-08-benchmark-plumbing.md Amendment 1)— one line, rides this PR. - Named merged-test sanction (D6a): this issue MAY update the known-gap disclosure strings at
report.py:26-32,report.py:455-460,report.py:664-667, the stale docstrings atrunner.py:3-4andmodel_matrix.py:8-15, and the exacttests/benchmarks/test_report.pyassertions that pin the known-gap paragraph — updating a pinned string to the new sanctioned text is in scope. Any other merged-test change remains stop-and-comment.
Refs #63
- The competitor manifest enumerates one entry per tool×model pairing; the runner's
Agent context — #86 [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio)
Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
Auditable copy: this content is mirrored as a comment on issue #86 (.planning/is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).1. Roadmap excerpt (why this exists)
- STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness — all target docs MCPs + no-MCP baseline, 50-question eval, reproducible, methodology disclosure mandatory. This issue makes the product itself a system under test.
- Decision 5.17: no comparative/benchmark claim enters README/PyPI/launch copy before reproducible data exists. Your PR makes zero claims.
- Cell composition CONFIRMED (maintainer comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86, 2026-07-08): the competitor manifest enumerates one entry per tool×model pairing;
benchmarks/model_matrix.pygains a manifest↔matrix validator (manifest provider/model pair absent fromdocs/benchmarks/model-matrix.yml→ cleanBenchmarkValidationError); NO cell-shape or artifact-layout change — cells stay competitor×question. "No runner core change" does not mean zero runner edits: a minimal dispatch-registry change at the_execute_cellseam is in scope; cell construction is untouched. - Amendment D5 rides this PR: flip the
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.mdStatus line (currently**Status:** Draft methodology for issue [#63], lines 3-4) toStatus: Accepted (blessed 2026-07-08; see PLAN-2026-07-08-benchmark-plumbing.md Amendment 1)— one line, nothing else in that file.
2. Code touch-points (verify against merged main before starting)
benchmarks/runner.py:23—_EXECUTABLE_ADAPTERS = {"fake", "no-mcp-baseline", "no_mcp_baseline"}; the allowlist check lives in_fake_provider_answer(runner.py:287-300, gate at 288-292). Replace this hardcoded gate with a minimal dispatch registry so the product adapter is dispatchable from_execute_cell(runner.py:218-284).benchmarks/runner.py:85-89— cell construction (competitor×question) — UNTOUCHED.benchmarks/runner.py:311-318—_tool_model_keyalready rendersid:provider/model; do not change it — the manifest entry carries the pairing.benchmarks/adapters/base.py:78-110—ProviderAdaptercontract: implement_generate();generate()(89-105) wraps latency measurement + failure capture (BenchmarkCellFailure, categoriestool_failure | timeout | mcp_protocol_crash).- NEW file under
benchmarks/adapters/— the stdio product adapter: spawn the local server over stdio (offline, no network — principle 2.2 at runtime), run the documentedsearch_docs→get_docsflow per corpus question, record raw tool payloads into the transcript artifact. Missing local index → recordedtool_failure, never a crash; dry-run path validates indexless; timeouts map to category"timeout". benchmarks/model_matrix.py—load_model_matrix(line 73),tool_model_cells(line 175); add the manifest↔matrix validator here. The stale deferral docstring at lines 8-15 ("wiring ... is deferred to a follow-up issue") gets updated — D6a-sanctioned.benchmarks/runner.py:357-376—_write_cell_artifactsdefines the per-cell artifact shapes (transcripts|tokens|latency|scoring|failures/<competitor>/<qid>.json) your results must fit; raw tool payloads go inside the transcript record, no new artifact directories.- D6a-sanctioned stale-text updates (ONLY these):
benchmarks/report.py:26-31module-docstring known-gap paragraph;report.py:453-460"Known gap (inherited, not fixed by this report)" block;report.py:663-667"No cross-model aggregation applies to this run..." block;runner.py:3-4stale docstring;model_matrix.py:8-15docstring; and the exacttests/benchmarks/test_report.pyassertions pinning those strings —test_report.py:145("Model / Client Matrix") andtest_report.py:538("No cross-model aggregation applies to this run"). Updating a pinned string to the new sanctioned text is in scope. Any other merged-test change: stop and comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86.
3. Test patterns to follow
- Selective
subprocess.runmonkeypatch that fakes only the matching call and delegates the rest to the real function:tests/benchmarks/test_runner.py:140-157. - CLI exercised as a real subprocess via
sys.executable -m benchmarks:test_runner.py:353-372. - Hand-authored minimal-but-valid run-dir factory:
tests/benchmarks/test_report.py:232-295(_minimal_valid_run_dir). - Unit tests drive the adapter against a fake/stubbed MCP transport ONLY — no real server spawn. One optional integration test may spawn the real server when the local index exists, skipped otherwise. All new tests under
tests/benchmarks/(additive).
4. Known pitfalls
- Merged tests in
tests/benchmarks/are regression cover: additive-only EXCEPT the named D6a strings in §2. - Lint blind spot until [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84 lands: CI does not run lint over
benchmarks/— additionally runuv run ruff check benchmarks/anduv run pyright benchmarks/yourself. - Diff >500 lines → apply the
supervisor-reviewlabel + an explanation section in the PR body. pyproject.tomlanduv.lockare untouchable this run — the adapter must work with existing dependencies.Refs #63, neverCloses #63;Closes #86only if every acceptance criterion is met.- The server's MCP tool surface is a CLIENT dependency here — never modify the server's tool names, parameters, or return shapes to accommodate the benchmark.
- If stdio invocation or artifact composition needs anything the merged runner does not support beyond the prescribed
_execute_cellseam, stop and comment on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 (issue recovery clause) instead of widening runner scope. - Full gate before merge decision:
uv run ruff check src/ tests/ benchmarks/,uv run pyright src/ benchmarks/,uv run pytest --tb=short -q,uv run pytest tests/benchmarks -q,uv run python-docs-mcp-server doctor. - Branch:
agent/86-product-runner-adapter.
5. Decision log (worker fills in)
- …
- addedagent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
on Jul 8, 2026
Metadata
Metadata
Assignees
Labels
agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
Context
Parent: #63. Methodology:
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md(work package 3). Runner:benchmarks/runner.py(merged via #75). Model matrix + provider adapters:benchmarks/model_matrix.py,benchmarks/adapters/(merged via #85).Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, resolve the one design decision below, create the
.planning/agent-context/file (decision 5.14, mirrored as an issue comment), and applyagent-readybefore any dispatch.Goal
Add the benchmark adapter that runs
python-docs-mcp-serveritself as a system under test: offline stdio invocation of the local server, per-question docs retrieval, response capture into the runner's artifact shapes.Design decision required from maintainer (blocks agent-ready)
Cell composition semantics for tool×model. Recommended: no runner core change — the competitor manifest enumerates one entry per tool×model pairing (the runner's
_tool_model_keyalready rendersid:provider/modelkeys), andbenchmarks/model_matrix.pygains a validator that checks manifest entries against matrix families. Alternative (rejected by default): change the runner's cell shape to competitor×model×question, which would rework merged artifact layouts and tests. Confirm or override before labeling agent-ready.Acceptance criteria (draft — Vision edits before agent-ready)
python-docs-mcpadapter inbenchmarks/adapters/launches the local server over stdio (offline, no network — principle 2.2 preserved at runtime), issues the documented retrieval flow (search_docs→get_docs) per corpus question, and records raw tool payloads into the transcript artifact.tool_failure(not a crash) when the index is missing, and the dry-run path validates without needing one.docs/benchmarks/model-matrix.ymlfails validation with a cleanBenchmarkValidationError.tests/benchmarks/, additive only) cover the adapter against a fake/stubbed MCP transport — no real server spawn required in unit tests; one optional integration test may spawn the real server if the local index exists, skipped otherwise.Refs #63, neverCloses #63.Scope boundaries
In scope: the product's own tool adapter, manifest/matrix validation, mocked tests. Out of scope: LLM/provider calls (adapters from #85 handle that in the live phase), competitor MCP adapters (separate issue), scoring, reporting, README.
Forbidden-territory reminder
Standard §2 list applies; additionally: merged tests in
tests/benchmarks/are regression cover — additive only. MCP tool names/parameters/return shapes of the server itself must not change to accommodate the benchmark.Validation commands
Recovery
If stdio invocation or artifact composition needs anything the merged runner does not support, stop and comment here rather than modifying
benchmarks/runner.pybeyond the issue's prescribed scope.Effort estimate
4–8 hours.
Refs #63.