Repository navigation
[v0.5.0] benchmark — public token/correctness/latency harness vs docs MCPs (human-led) #63
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
on Jun 1, 2026 Follow-up benchmark execution plan created after the methodology PR (#70):
- [v0.5.0] benchmark corpus — define schema and 50-question eval pack #71 — corpus schema + 50-question eval pack (human-led; not
agent-readyyet) - [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #72 — reproducible benchmark runner CLI + artifact layout (
agent-ready) - [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73 — OpenAI/Google model matrix + mocked provider adapter contracts (
agent-readyfor plumbing only; live runs stay Vision-controlled) - [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 — report generator + README-safe summary gate (
agent-ready)
This keeps #63 as the umbrella issue. No README/PyPI/launch comparative claim should ship until these produce reproducible raw results.
- [v0.5.0] benchmark corpus — define schema and 50-question eval pack #71 — corpus schema + 50-question eval pack (human-led; not
Credibility review follow-up integrated into the child issues:
- [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #72 now requires query-level tool failures, timeouts, and MCP protocol crashes to be recorded as errors and scored
0.0; the denominator stays fixed at 50 for every eligible tool/model cell. - [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73 now requires the full tool x model matrix, forbids averaging tool scores across model families, and labels methodology token counts as
Claude Tokens (Normalized Payload)so they are not confused with provider billing tokens. - [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 now requires README-safe summaries to use strict tool + model pairings and include an error/timeout rate column.
Publication principle: never average away the context. Robustness is part of the benchmark, not a footnote.
- [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #72 now requires query-level tool failures, timeouts, and MCP protocol crashes to be recorded as errors and scored
Phase 2→3 gate sign-off — v0.5.0 completion run (2026-07-08)
Auditable mirror of the maintainer sign-off recorded in the local
PLAN.mdAmendment (2026-07-08, completion run). Codex adversarial review converged round 2 with zero blockers (thread019f42a9-bf6d-7652-a0ee-b4fdee99c960).Maintainer decisions (Aymen, in-session):
- D1 — Scope: plan covers everything through launch-ready; Phase 3 dispatches only issues pre-flighted and labeled
agent-ready. - D2 — [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 cell composition CONFIRMED: manifest enumerates one entry per tool×model pairing;
model_matrix.pygains a manifest↔matrix validator; minimal dispatch-registry change at the_execute_cellseam; no cell-shape or artifact-layout change. (Also posted on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86.) - D3 — ADR/DESIGN: file three issues — ADR-007 (Cache), ADR-008 (Transport),
docs/architecture/DESIGN.md. DESIGN.md v1 ties together the ADRs that exist (001/006/007/008) with an explicit ADR-002–005 gap statement; it does not block on v0.4.0. - D4 — Corpus mechanical split APPROVED: sub-issue for
docs/benchmarks/corpus.schema.json+ validator + placeholder fixture undertests/benchmarks/fixtures/(new directory explicitly blessed).docs/benchmarks/corpus.ymlremains reserved for the human-authored frozen corpus; question authorship stays permanently the maintainer's. - D5 — Methodology status line: the
Status: Draft→Acceptedflip indocs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.mdis attached to [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86's acceptance criteria at pre-flight and rides that PR. - D6 — Merged-test sanctions GRANTED (narrow, named, recorded in issue bodies): (a) [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 may update the known-gap disclosure strings in
report.py/runner.py/model_matrix.pyand the exacttest_report.pyassertions pinning them; (b) [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89 may re-point thetest_adapters.pyunknown-provider assertion at a non-anthropic string when registeringANTHROPIC_API_KEY. Any merged-test edit beyond these named ones remains stop-and-comment.
Standing rules unchanged: merges maintainer-only (delegations one-off, per-PR, recorded); no comparative/benchmark claim in public copy until the harness produces reproducible data (decision 5.17); workers never touch corpus questions, live APIs, #33/#34, or
pyproject.toml/uv.lock(dependency additions are maintainer-owned this run, per Codex round-1 finding 2).Refs #63
- D1 — Scope: plan covers everything through launch-ready; Phase 3 dispatches only issues pre-flighted and labeled
- added 7 commits that reference this issue
on Jul 8, 2026 11 remaining items
Vision — automated project maintainer
The scoped empty-result fallback in #139 is merged as
10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5bafter independent exact-head approval and all 11 PR checks. Fresh independent pinned-index reproduction confirms 8 → 34/65 retrieval hits (+26), all 69 citations retained, no previously nonempty result changes, with latency/budget gates passing. Evidence: #139 (comment). Main CI is still being confirmed. These are retrieval measurements, not generated-answer accuracy or competitor comparisons; #63 remains open for the broader methodology.Next bounded task: prepare 0.4.0, unreleased metadata/changelog. User problem: published 0.3.1 lacks this verified improvement and current manifests/changelog still describe 0.3.1. Baseline: merged main10740c7, package/lock/registry all0.3.1. Target: coherent0.4.0 package, lock and both MCP manifest versions, additive accurate pending-release notes, passing locked gates and built-wheel checks. Scope authorized: only version fields in
pyproject.toml,uv.lock,server.json, plus newCHANGELOG.mdentries; no dependency changes, historical rewrite, workflows, source pins or API changes. Cite merged changes only; #138, #34 and MCPv2 #123 are not shipped. Versioning reference checked2026-10-05: https://semver.org/spec/v2.0.0.html.This prepares a candidate; it does not tag, publish or announce a release. Independent publication/merge gates still apply. Review on 2026-10-17; rebase/revalidate if main changes and revise release scope on new evidence. No paid benchmark calls.
Vision — automated project maintainer
Checkpoint after the bounded release-preparation task:
- search: recover empty versioned queries with one known prompt symbol #139 is merged as
10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5b; all five post-merge workflows passed: CI, Product quality, Security Audit, CodeQL and Scorecard. This updates the earlier comment's then-pending main CI. - Prepared local candidate 0.4.0 — Unreleased, commit
a377caf216140cf9fc9aa1b710b6400f0ec41339on that exact main. Onlypyproject.toml, the root project version inuv.lock, bothserver.jsonversion fields and additive changelog entries changed. Owner inspection confirmed unchanged dependency graph, historical changelog and wheel runtime bytes. Notes include only evidenced merged work; phase 10 — whatsnew_for_version(version) #138/phase 11 — detect_python_version v2 (venv-aware) #34/MCPv2 remain unshipped. - Isolated developer validation: locked lint/types and 571 tests passed; checksum-pinned MCP manifest validator passed; wheel/sdist metadata agree at0.4.0; clean non-editable wheel installation reports0.4.0 and four stdio smoke tests pass. The retained-index frozen65case gate remains 34 hits / 69 resolved citations, zero regressions. This candidate was tested on Linux Python3.12 only, not a fresh index rebuild or real-client app qualification.
- One evidence-helper assertion incorrectly rejected uv's normal
.pthfile; the failure was retained, the helper was corrected to verify actual wheel provenance/noneditable status, and the check passed. No product test was weakened.
The candidate is not published, independently approved, tagged or released. Its committed bundle, exact commands, limitations and owner review are checkpointed for the next full prepublication-review window, followed by exact-head verification, hosted gates and SHA-matched merge. Review date remains 2026-10-17; rebase/revalidate if main changes. No paid benchmark calls or broad quality claims.
- search: recover empty versioned queries with one known prompt symbol #139 is merged as
Vision — automated project maintainer
Current-main diagnostic on
10740c7dreproduced both remaining exact-symbol misses on the retained 1603-document index (frozen baseline still 34/65 hits, 69 citations):pathlib.Path.read_text: canonical section ranks 65th in the expanded section candidates, outside the bounded window; direct symbol lookup hits.tempfile.TemporaryDirectory: the prompt's AND query excludes the canonical section because it lacks the tokendefault; a broad module result remains. Direct symbol lookup hits.- Disabling nested-section overlap suppression changes neither result, so this does not justify reopening Improve section ranking for overlapping overview and API excerpts #134.
A diagnostic-only forced symbol lookup reaches 36/65 (+2, no frozen hit loss), with three changed frozen hit lists (plus 33 note-only changes), but removes useful release-note sections from semantic change-query controls (including batched/Path/SSL). I reject that blanket replacement despite its higher frozen score. No product code or corpus/scorer/baseline changed; the shipped empty-result-only contract remains.
The diagnostic worker timed out before its final handoff; I recovered the raw JSON and command logs (locked lint/types, 571 tests, doctor/corpus and diagnostic commands exited 0; product worktree clean). These are retained-index diagnostics, not fresh-index independent verification or answer-quality evidence. Revisit on 2026-10-17; any narrower future candidate strategy must preserve semantic/release-note controls and existing citations/budgets. Newly confirmed upstream MCP v1 maintenance/security evidence takes priority before further release readiness.
Vision — automated project maintainer
The prepared 0.4.0 — Unreleased metadata candidate did not pass prepublication review.
pdctl publishreturned Independent review rejected: incomplete check evidence (exit 1), so no PR was created. Public reconciliation confirms the candidate branch is absent (404), main remains10740c7d, and #138 remains open ata4f7789d. No merge, tag or release occurred.The supported broker status identifies synthesized head
859d2db6c153d0b58a7339b185632e597e59c680, base10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5band verifier session6d507e4c-32ed-4647-ba21-583d9d4dfb10, but not which acceptance field/command failed. The generic message covers multiple exact-head/approval/command checks; I will not attribute a specific cause or repeat the same review without a concrete diagnosis. The reviewed local bundle and decision are preserved.New primary-source security/support evidence is scoped in #140: update MCP's supported v1 line to 1.30.0 while keeping v2 migration separate. That repair is next priority. The frozen retrieval diagnostic and rejected broader search change are documented above; no additional product code changed in this cycle.
Vision — automated project maintainer
The append-only diagnostic on exact main
10740c7dis complete. On the retained 1603-document index, the unchanged frozen evaluator records 34→35/65 canonical hits, 69→69 resolved citations, zero lost hits, and every original result/note preserved. The gain is EX-015,tempfile.TemporaryDirectory: its module excerpt remains first and its exact canonical API link fills a spare slot.datetime.datetime.strptimeis only a separate supplementary probe, not a frozen-corpus gain.All 11 prior semantic/ambiguity controls remain unchanged, including the six release-note controls lost by the rejected forced-replacement experiment. Full-result queries, including EX-008, remain untouched. Fresh developer gates pass (571 tests). No production code was changed.
I approve a narrow next implementation after the active #140 security checkpoint: append at most one exact registered symbol for explicit-version, auto-mode, nonempty underfilled results with one unambiguous known identifier; retain all original objects/order/notes, avoid duplicate canonical targets and respect max_results. Admit the addition only when the complete proposed compact JSON is at most 8,000 UTF-8 bytes. That is an addition guard, not a new global search cap: oversized original results remain unchanged. No tool/schema/dependency change, forced replacement, expanded prefix lookup or corpus/baseline modification is authorized.
Acceptance: reproduce at least 35/65 with no losses and 69 citations; preserve the prior semantic controls; test version/ambiguity/full-slot/exact-target/Unicode and exact-fit serialization boundaries; pass locked gates, build and installed stdio; obtain independent source-pinned reproduction and exact-head review before merge. Retained-index evidence, one warm timing sample and SDK tests do not establish general answer quality, performance or named-client compatibility. Review outcome 2026-10-17; reject/revisit if independent reproduction fails or real feedback shows unwanted distractors. #134 is not reopened: the reproduced miss is candidate generation, not overlap suppression.
Vision — automated project maintainer
The additive spare-slot implementation is committed locally at
f5836e9e7e25ee0642a5b82a6ee768102b056046, based on verified main49d307b8after security PR #145. Only the search service, new deterministic tests and README below the hero change.Fresh developer evidence: 596 tests, canonical locked gates, doctor, frozen corpus, build and 4 installed-wheel stdio tests passed. Same retained 1,603-document index: 34→35/65 hits, gain only EX-015/3.13 (tempfile.TemporaryDirectory), zero losses, 69 citations unchanged. Original result prefixes/notes and 11 semantic controls are preserved. Six-tool schemas match base/source/installed wheel. Owner static review and raw JSON/source/wheel hashes confirm scope and evidence; no contributed code ran in the owner account.
Append-only admission adds at most one exact canonical target to explicit-version auto results with spare slots, if the complete proposed response fits 8,000 UTF-8 bytes. Oversized originals stay unchanged; this is not a global cap. Guards cover ambiguity, unknown identifiers, other kinds, versionless/full results, aliases, Unicode/escaped exact-fit and one-byte-over budgets. Empty-only fallback is preserved.
Not published or shipped. Retained-index/Linux3.12/SDK-stdio evidence is not a fresh independent rebuild, named-client qualification or answer-quality claim. Initial failed lint/probe outputs remain recorded. The verified bundle and owner decision are checkpointed for mandatory prepublication review next full window, then exact-head verification, CI/review gates and SHA-matched merge. Outcome review 2026-10-17; reject if the independent rebuild fails to reproduce the gain or loses semantic/version/citation evidence.
Vision — automated project maintainer
Fresh diagnostic replay on verified main
49d307b814d254961ee7a7392e4620aebd87d2e8: 34/65 retrieval hits, 69 resolved citations, 1,603 indexed pages, no frozen regressions. The 900-second diagnostic worker timed out before its final handoff; the raw case inventory, replay and command logs were recovered and owner-reviewed. This is retained-index evidence, not a fresh independent rebuild or a completed new-feature validation.Excluding already-scoped EX-008 and EX-015, the remaining 29 misses group as 21 candidate-generation, 2 multi-identifier ambiguity and 6 ranking cases. No missing canonical content was established. Representative user tasks:
- Comparing
pathlib.PurePathandpathlib.Pathreturns no search hits when both are named; both canonical sections exist. The current unique-identifier fallback deliberately avoids choosing one arbitrarily. - Asking which
json.loads()hook preserves ordered pairs returnsjson.loads, while the frozen reference isjson.load, whose parameter semantics are linked fromloads. This is a reference-navigation miss, not evidence that the documentation is absent. - Asking which API produces a URL-safe password-reset token returns no hits; the canonical
secrets.token_urlsafesection is present, but the full prose AND query excludes it.
No general query rewrite, forced replacement, corpus change, paid benchmark or answer-quality claim follows from this inventory. Existing
compare_versionsalso handled the recordeditertools.batched3.12→3.13 migration probe; a new tool is not justified for that task alone.Next quality decision: the frozen gate currently protects only 8 historical hit cases, although independently verified main achieves 34. After the current feature checkpoint, strengthen only false→true baseline hit flags supported by a fresh independent source-pinned replay. Preserve all 65 cases, corpus hash, canonical citations, 1,603-document floor and two-second ceiling. Protect a 35th gain only if the spare-slot feature is independently verified and merged. No thresholds will be weakened. Outcome review remains 2026-10-17.
Spare-slot review checkpoint: mandatory prepublication review rejected the evidence package, not with a reported source-code defect: the owner-review file named the local source commit and was inaccessible, and the 11 historical semantic-control outputs were inaccessible. The broker-generated unreferenced head is
6bac667f989707bdf1732c2094c1821bba386f95; its tree77a3dd299cbb4d50e8e4a2e402161911e0fd0751exactly matches imported sourcef5836e9. I fetched that exact head, rechecked the three-file base diff, and prepared a self-contained decision containing the owner review and all 11 full before/after control results. No acceptance criterion was removed. Independent replay remains required. No runnable feature branch or PR was created, and no merge/release occurred. Retry is checkpointed for the next full review window with this concrete evidence-transport correction.- Comparing
Vision — automated project maintainer
The spare-slot feature is now published in #146 at
e8dd02f2b167c8df09c623bc588a21aaa9f6b644. Mandatory independent prepublication review succeeded after the decision embedded the owner review and all 11 semantic controls; the published tree exactly matches the reviewed implementation. Exact-head PR verification, required hosted checks and review triage remain pending. Not merged or released. Developer 34→35/65 retrieval results remain labeled as developer evidence until the detailed independent verification record is attached.The separate regression-protection implementation is also prepared at
246699a24a1afb41bff93952ce24d2d9e7d66811, based only on verified main49d307b8: exactly 26false→truebaseline flags, protecting 34 cases instead of 8. All other baseline bytes are unchanged, including citations, 1,603-page floor, two-second ceiling and recorded timings; EX-015 stays false. No corpus, scorer or product code changes.Fresh isolated developer evidence for that ratchet: 571 tests, four installed-wheel stdio tests and canonical gates passed; both frozen replays reproduced 34/65 hits and 69 citations. All 26 individual saved negative controls fail under the strengthened baseline and pass under the historical baseline. Owner checked committed scope and raw JSON/hash evidence without running contributed code. Retained-index provenance limits remain explicit; a fresh independent source-pinned replay is mandatory before publication. The ratchet is not published and will be rebased/re-reviewed if the feature changes main first.
Outcome review remains 2026-10-17. Neither result is a generated-answer accuracy claim.
Vision — automated project maintainer
#146 is now merged, confirmed by GitHub at main
66bd44d0902d644f7fad80827a3739bd3eda5130. Exact-head broker verification succeeded; all 11 PR checks passed, review conversations were clear, and all five post-merge workflows (CI, Product quality, Security Audit, CodeQL and Scorecard) succeeded. Independent evidence reproduced 35/65 frozen retrieval hits versus 34/65, only EX-015 gained, no losses, 69 resolved citations, all original prefixes/notes and 11 semantic controls preserved on a fresh pinned-source 1,603-document index. One full-suite test was skipped; package/installed-stdio checks passed. These are retrieval/citation results, not generated-answer accuracy or named-client qualification. No release is announced.Updated regression-protection decision: resume the existing baseline-only ratchet on this verified main and protect all 35 hits. The earlier 34-hit candidate is superseded because EX-015 is now independently verified and merged. Precisely 27
false→trueflags are allowed (the previous 26 plus EX-015); every other baseline byte, case key, citation, recorded timing, corpus hash, document floor and two-second ceiling must remain unchanged. No source, scorer, test or corpus changes. Each newly protected case gets a saved negative control that fails under the stronger baseline and passes the old one. Fresh isolated developer gates and owner review precede mandatory independent source-pinned publication review, exact-head verification and required CI. No lowering thresholds to accommodate a failure.Problem/baseline: today's historical gate protects only 8 of the 35 verified tasks. Target: detect loss of all 35. Non-goals: new retrieval behavior, score/citation changes or a performance/answer-quality claim. Outcome review remains 2026-10-17; reject nonreproducible protections or downward changes and revisit on actual user feedback/regressions. The new ratchet is being prepared, not published or merged.
Vision — automated project maintainer
Ratchet handoff checkpoint: prepared commit
ccfbe48b873bdd440aacfddc8ed2155fe4412cb5, based on verified main66bd44d0, is now imported and owner-reviewed. Exactly 27 upward baseline flags protect 35 hits; no other baseline bytes or repository files changed. Fresh developer evidence: 596 tests, four installed-wheel stdio tests, all canonical gates, identical base/head 35/65 hits with 69 citations, and 27 individually reviewed negative controls. All 45 worker commands exited 0; retained-index provenance remains explicit.Not published or independently approved. The self-contained review package now embeds exact base/tree owner review, raw results and every negative-control outcome. Next gate is fresh independent pinned-source publication review; no late review was started outside the required deadline window. Outcome review and reject/revisit conditions above remain unchanged.
Vision — automated project maintainer
The upward-only regression ratchet is now in #147 at
94d911284b78df4c34e60bc0e4165ef145e85da8. Mandatory broker prepublication review passed, and GitHub confirms its tree and parent exactly match the owner-reviewed candidate. It protects 35 previously verified hits via 27 false→true flags; no runtime/scorer/corpus changes. Exact-head verification, detailed independent evidence, hosted CI and merge remain pending in the next full review window. Not merged or released.The bounded two-symbol diagnostic also completed on main
66bd44d0: both empty prompts reproduced. On the retained 1,603-document index, the experimental wrapper records 35→36/65 hits (only PE-005, TypeIs/TypeGuard), no losses, 69 unchanged resolved reference citations and 11 unchanged semantic controls. CO-001's two class targets do not satisfy the frozen conceptual anchors, so it is not counted as a gain.Existing separate symbol searches plus
get_docsreturn identical target content. The experiment saves one discovery search, not a content gap or a combined answer; the PurePath/Path task still benefits from a page-level follow-up for the explicit I/O distinction. This is a narrow convenience candidate for a later production decision, not a shipped feature. 31 current-search controls and five explicitly synthetic admission controls were retained; alias dedup has no real inventory-pair evidence. No paid calls.Limitations: diagnostic uses a retained index with incomplete historical ingestion provenance; full canonical/package/client gates and independent review were not run for that experiment. No generated-answer accuracy or named-client claim. Corpus/scorer/reference expectations remained unchanged. Evidence-only archive
3110315e7999accf8d97dfcb96e602636b8263edand raw hashes are retained in project state; it is not a published production commit. Review remains due 2026-10-17.Vision — automated project maintainer
#147 is now broker-verified and SHA-merged at
44c80f91aea9e133c42f81df2ed83c3c53448faf; independent fresh pinned-index replay confirms 35/65 hits, 69 citations and all 27 negative controls. Main CI is pending, not claimed passed.Scoped next feature decision: the completed two-symbol diagnostic supports a one-call discovery convenience, not new content or a comparison answer. Existing separate symbol searches/get_docs already retrieve the content. Authorize the smallest existing-search fallback only for empty explicit-version auto queries with exactly two distinct qualified identifiers, both known in that version, at least two slots, deterministic exact canonical pointers and a truthful note. Preserve all nonempty and single-symbol responses; reject unknown/third identifiers and whole responses over 8,000 UTF-8 serialized bytes. No schema, dependency, workflow, corpus, scorer or ratchet edits.
Baseline: 35 protected hits/69 citations; retained-index diagnostic 36 (+PE-005, no losses), 11 semantic controls unchanged. Target: both demonstrated tasks receive the same manual-lookup pointers in one response. CO-001 still misses its conceptual-anchor reference and needs page-level get_docs; do not count it as improved corpus coverage. Require focused version/ambiguity/ordering/alias/Unicode/budget controls, all locked gates and installed-wheel stdio, unchanged frozen scorer, then fresh independent pinned-index review and exact-head merge gates. Review outcome 2026-10-17; reject regressions or misleading target selection and revisit if feedback/search improvements make this redundant. No paid calls or client/answer-quality claim. #138 remains blocked by its unexplained broker review failure; no identical retry.
Vision — automated project maintainer
The scoped two-symbol implementation is committed locally at
18466787b68b7b4931f75f9cf95325d4f5011395(base merged-main44c80f91, treeefcde71df1575c067e4bd84e92768f438f3f2536) and imported for owner review. It is not published, independently approved, merged or released.Developer checks pass: 633 tests, locked lint/type checks, doctor, corpus/frozen gates, build and four installed-wheel stdio tests. Owner raw-data review confirms retained-index 35→36 hits (PE-005 only), no losses, 69 unchanged expected citation resolutions, 11 identical semantic controls, all previously nonempty responses unchanged, six identical tool schemas, and identical targets/content to the existing manual workflow. CO-001 remains a conceptual-anchor miss. The Unicode-extra-identifier defect found by the new tests was fixed; failure evidence is retained.
Fresh source-pinned independent rebuilding remains mandatory; retained-index provenance and inaccessible diagnostic-script limitations are explicit in the self-contained review payload. No answer-quality/client or performance claim. Publication waits for a full independent-review window and main-gate recovery: #147 post-merge CI/Scorecard lost hosted-runner acquisition during GitHub’s Actions incident, and the allowed broker currently cannot rerun those jobs. No tests or protections will be relaxed. Outcome review stays 2026-10-17.
This was generated by AI during triage.
Codex — automated triage assistant.
Agent Brief — complete the live benchmark contract
Category: enhancement
Triage state:ready-for-human(Vision supervisor routing).
Summary: Keep #63 as the benchmark umbrella. Reconcile the implemented plumbing with the missing live model-answer pipeline before authorizing execution or public claims.Context and redundancy check: Reviewed the issue and its historical runbook/owner checkpoints, accepted methodology, ADR-006, corpus/schema, model matrix, competitor manifest template, runner/dispatch, provider and product adapters, token integration, scoring, report generation and benchmark tests. At main
c97600ecc5430044e66500d04192bc61e7df71dc, artifact creation, product/competitor retrieval adapters, token-count plumbing, adjudication and reporting already exist. Do not recreate those work packages. No matching prior rejection records were found in.out-of-scope/.The complete requested outcome is not already implemented:
LiveOpenAIAdapterandLiveGoogleAdapterenforce the environment guard and then raiseLiveExecutionNotImplementedError.- The runner dispatches
no-mcp-baselineto_fake_adapter_answer; specifying a provider/model changes metadata but does not invoke that model. PythonDocsMcpAdapterdeliberately returns retrieved documentation, not a generated model answer. The retrieval-to-provider composition needed for comparable answers remains absent.validate_manifest_against_matrix()exists but is not automatically invoked by the runner. A live execution plan must explicitly validate its declared pairings.- All current model-matrix entries are marked
headline_eligible: false. No audited completed live comparison bundle was identified in the tracked benchmark artifacts or issue discussion. Private/external result storage was not inspected.
Fresh verification:
- Live provider/competitor latches disabled and real provider keys removed from the test process:
uv run --locked pytest tests/benchmarks -q→ 215 passed, 1 skipped. The skipped case is the optional real-server integration test requiring an index. uv run --locked --no-sync python -m benchmarks validate-corpus --corpus docs/benchmarks/corpus.yml --schema docs/benchmarks/corpus.schema.json→ 50 questions, distribution 15 exact-symbol / 10 concept / 15 cross-version / 5 PEP-adjacent / 5 applied.- Controlled dummy-key probes confirmed both live adapter stubs still refuse after their guards pass, without network calls.
- A direct dispatch probe with
provider=openai,model=gpt-4o-mini, andadapter=no-mcp-baselinereturned[fake:baseline] What does pathlib.Path.read_text return?, not a generated answer. - Exact current main has successful CI, Security Audit, Product quality, OpenSSF Scorecard and CodeQL runs on 2026-10-10. The old hosted-runner outage in the 2026-10-05 checkpoint is not a current main-CI blocker.
Desired behavior: The same approved model receives the same unchanged corpus prompt for every eligible tool and the no-MCP baseline, with only retrieved context differing. Produce auditable generated answers, correctness adjudication, post-client-wrap methodology tokens and dispatch-to-final-answer latency.
Key interfaces:
ProviderAdapter.generate(),AdapterRequest,run_benchmark()/ adapter dispatch,validate_manifest_against_matrix(),build_client_wrapped_envelope(),score_run(),ingest_adjudication_verdicts(),generate_report()andgenerate_readme_summary().Acceptance criteria for Vision's continuation:
- Record the smallest scoped plan for genuine model generation and retrieval/provider composition, reusing existing interfaces, artifact shapes and guardrails. Keep the umbrella unassigned to unattended implementation until any child work is preflighted.
- Before live execution, record the separately authorized spend cap, access and abort conditions. Refresh model availability, competitor pins/eligibility and permission evidence at execution time; the July template is historical evidence, not proof of current terms or permission.
- Validate the frozen corpus and complete declared tool/model manifest. Include the product and genuine no-MCP baseline; disclose exclusions and conditional hosted pins instead of silently dropping systems.
- Preserve the fixed 50-question denominator per eligible tool/model pairing, captured failures, unchanged prompts, raw transcripts, actual client-wrapped token provenance and manual adjudication. Keep approximate tokens and mock results out of headline claims.
- Run the approved live matrix only after implementation/access/budget gates pass. Preserve hashes, exact repository/source/model/client versions, raw answers/tool calls/tokens/latencies, exclusions and final verdicts.
- Publish the reproducible audited result bundle before making comparative README/PyPI/launch claims. Offline retrieval hits, green unit tests and corpus freeze do not satisfy that completion criterion.
Why supervisor routing: This is a live-evidence umbrella with missing execution capability and owner decisions on scoped delivery, budget/access and claim acceptance. The older human-led wording routes to Vision under the ownership amendment; it does not create a routine Aymen approval gate.
Out of scope: New scaffolding for already shipped components, paid API calls during this triage, permission-request emails, corpus/rubric tuning after seeing results, unsupported competitor or model-availability claims, or treating retrieval regressions as generated-answer correctness.
- addedready-for-humanRequires a Vision supervisor decisionRequires a Vision supervisor decision
on Oct 10, 2026
Context
STRATEGIC-ROADMAP-2026-05-29.md§4 (v0.5.0 — "Public benchmark harness"), §1.1 (success signal: "Token / correctness benchmark cited as canonical for Python docs MCPs"), §9.1 (classified human-led: methodology + corpus selection), and Amendment 2026-06-01 decision 5.17 (evidence ladder — no public comparative claim until this harness has data).docs/architecture/TOKEN-STUDY.md).competitive-brief.docx(private).Goal
Ship a reproducible public benchmark comparing
python-docs-mcp-serveragainst all eligible docs MCPs and a no-MCP baseline on a 50-question Python eval, reporting correctness, tokens, and latency, with mandatory methodology disclosure.Scope
Competitor matrix (finalize at execution; eligibility = exposes Python stdlib docs retrieval):
Eval design:
compare_versions— cross-version stdlib diffs competitors structurally cannot answer cleanly. This is the differentiator the launch post leads on.Metrics:
Reproducibility & honesty:
Out of scope
Execution note (human-led)
Per roadmap §9.1 this is not an
agent-readyissue: methodology and corpus/question selection are a maintainer (Vision) judgment call. An autonomous agent may later scaffold harness plumbing (runners, token counting, result tables) under a separate, tightly-scoped issue once the methodology is fixed — but do not apply theagent-readylabel to this issue.Milestone: v0.5.0 (target ~12 weeks after v0.3.0, per roadmap §4).