Skip to content

[v0.5.0] benchmark — public token/correctness/latency harness vs docs MCPs (human-led) #63

Description

@ayhammouda

Context

  • Roadmap: STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0 — "Public benchmark harness"), §1.1 (success signal: "Token / correctness benchmark cited as canonical for Python docs MCPs"), §9.1 (classified human-led: methodology + corpus selection), and Amendment 2026-06-01 decision 5.17 (evidence ladder — no public comparative claim until this harness has data).
  • Token methodology must match the internal token study (Study A): Claude tokenizer, measured after client-side rewrap (decision 5.8; ADR-006; future docs/architecture/TOKEN-STUDY.md).
  • Market reference: competitive-brief.docx (private).

Goal

Ship a reproducible public benchmark comparing python-docs-mcp-server against all eligible docs MCPs and a no-MCP baseline on a 50-question Python eval, reporting correctness, tokens, and latency, with mandatory methodology disclosure.

Scope

Competitor matrix (finalize at execution; eligibility = exposes Python stdlib docs retrieval):

  • Context7
  • GitMCP
  • DeepWiki
  • Ref.tools
  • no-MCP baseline (model answering from parametric knowledge alone)

Eval design:

  • 50 questions across: exact symbols, concepts, cross-version behavior, PEP-adjacent.
  • Lead with compare_versions — cross-version stdlib diffs competitors structurally cannot answer cleanly. This is the differentiator the launch post leads on.

Metrics:

  • Correctness — rubric-scored against official docs.
  • Tokens — Claude tokenizer, measured after client-side rewrap (consistent with Study A / decision 5.8).
  • Latency — wall-clock per query.

Reproducibility & honesty:

  • Runnable from a clean clone; competitor versions pinned.
  • Methodology disclosure mandatory.
  • No comparative/benchmark claim enters README / PyPI / launch copy until this harness produces data (decision 5.17). No "we benchmarked it, trust me."

Out of scope

  • Any public comparative claim before data exists.
  • Benchmarking non-Python-docs MCPs.
  • Shipping in v0.3.x or v0.4.0 — this is a v0.5.0 artifact.

Execution note (human-led)

Per roadmap §9.1 this is not an agent-ready issue: methodology and corpus/question selection are a maintainer (Vision) judgment call. An autonomous agent may later scaffold harness plumbing (runners, token counting, result tables) under a separate, tightly-scoped issue once the methodology is fixed — but do not apply the agent-ready label to this issue.


Milestone: v0.5.0 (target ~12 weeks after v0.3.0, per roadmap §4).

Activity

  1. ayhammouda commented on Jun 8, 2026

    @ayhammouda
    OwnerAuthor

    Follow-up benchmark execution plan created after the methodology PR (#70):

    This keeps #63 as the umbrella issue. No README/PyPI/launch comparative claim should ship until these produce reproducible raw results.

  2. ayhammouda commented on Jun 8, 2026

    @ayhammouda
    OwnerAuthor

    Credibility review follow-up integrated into the child issues:

    Publication principle: never average away the context. Robustness is part of the benchmark, not a footnote.

  3. added this to the v0.5.0 milestone on Jul 8, 2026
  4. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Phase 2→3 gate sign-off — v0.5.0 completion run (2026-07-08)

    Auditable mirror of the maintainer sign-off recorded in the local PLAN.md Amendment (2026-07-08, completion run). Codex adversarial review converged round 2 with zero blockers (thread 019f42a9-bf6d-7652-a0ee-b4fdee99c960).

    Maintainer decisions (Aymen, in-session):

    1. D1 — Scope: plan covers everything through launch-ready; Phase 3 dispatches only issues pre-flighted and labeled agent-ready.
    2. D2 — [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 cell composition CONFIRMED: manifest enumerates one entry per tool×model pairing; model_matrix.py gains a manifest↔matrix validator; minimal dispatch-registry change at the _execute_cell seam; no cell-shape or artifact-layout change. (Also posted on [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86.)
    3. D3 — ADR/DESIGN: file three issues — ADR-007 (Cache), ADR-008 (Transport), docs/architecture/DESIGN.md. DESIGN.md v1 ties together the ADRs that exist (001/006/007/008) with an explicit ADR-002–005 gap statement; it does not block on v0.4.0.
    4. D4 — Corpus mechanical split APPROVED: sub-issue for docs/benchmarks/corpus.schema.json + validator + placeholder fixture under tests/benchmarks/fixtures/ (new directory explicitly blessed). docs/benchmarks/corpus.yml remains reserved for the human-authored frozen corpus; question authorship stays permanently the maintainer's.
    5. D5 — Methodology status line: the Status: Draft → Accepted flip in docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md is attached to [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86's acceptance criteria at pre-flight and rides that PR.
    6. D6 — Merged-test sanctions GRANTED (narrow, named, recorded in issue bodies): (a) [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 may update the known-gap disclosure strings in report.py/runner.py/model_matrix.py and the exact test_report.py assertions pinning them; (b) [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89 may re-point the test_adapters.py unknown-provider assertion at a non-anthropic string when registering ANTHROPIC_API_KEY. Any merged-test edit beyond these named ones remains stop-and-comment.

    Standing rules unchanged: merges maintainer-only (delegations one-off, per-PR, recorded); no comparative/benchmark claim in public copy until the harness produces reproducible data (decision 5.17); workers never touch corpus questions, live APIs, #33/#34, or pyproject.toml/uv.lock (dependency additions are maintainer-owned this run, per Codex round-1 finding 2).

    Refs #63

  5. 11 remaining items

  6. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The scoped empty-result fallback in #139 is merged as 10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5b after independent exact-head approval and all 11 PR checks. Fresh independent pinned-index reproduction confirms 8 → 34/65 retrieval hits (+26), all 69 citations retained, no previously nonempty result changes, with latency/budget gates passing. Evidence: #139 (comment). Main CI is still being confirmed. These are retrieval measurements, not generated-answer accuracy or competitor comparisons; #63 remains open for the broader methodology.

    Next bounded task: prepare 0.4.0, unreleased metadata/changelog. User problem: published 0.3.1 lacks this verified improvement and current manifests/changelog still describe 0.3.1. Baseline: merged main10740c7, package/lock/registry all0.3.1. Target: coherent0.4.0 package, lock and both MCP manifest versions, additive accurate pending-release notes, passing locked gates and built-wheel checks. Scope authorized: only version fields in pyproject.toml, uv.lock, server.json, plus new CHANGELOG.md entries; no dependency changes, historical rewrite, workflows, source pins or API changes. Cite merged changes only; #138, #34 and MCPv2 #123 are not shipped. Versioning reference checked2026-10-05: https://semver.org/spec/v2.0.0.html.

    This prepares a candidate; it does not tag, publish or announce a release. Independent publication/merge gates still apply. Review on 2026-10-17; rebase/revalidate if main changes and revise release scope on new evidence. No paid benchmark calls.

  7. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    Checkpoint after the bounded release-preparation task:

    • search: recover empty versioned queries with one known prompt symbol #139 is merged as 10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5b; all five post-merge workflows passed: CI, Product quality, Security Audit, CodeQL and Scorecard. This updates the earlier comment's then-pending main CI.
    • Prepared local candidate 0.4.0 — Unreleased, commit a377caf216140cf9fc9aa1b710b6400f0ec41339 on that exact main. Only pyproject.toml, the root project version in uv.lock, both server.json version fields and additive changelog entries changed. Owner inspection confirmed unchanged dependency graph, historical changelog and wheel runtime bytes. Notes include only evidenced merged work; phase 10 — whatsnew_for_version(version) #138/phase 11 — detect_python_version v2 (venv-aware) #34/MCPv2 remain unshipped.
    • Isolated developer validation: locked lint/types and 571 tests passed; checksum-pinned MCP manifest validator passed; wheel/sdist metadata agree at0.4.0; clean non-editable wheel installation reports0.4.0 and four stdio smoke tests pass. The retained-index frozen65case gate remains 34 hits / 69 resolved citations, zero regressions. This candidate was tested on Linux Python3.12 only, not a fresh index rebuild or real-client app qualification.
    • One evidence-helper assertion incorrectly rejected uv's normal .pth file; the failure was retained, the helper was corrected to verify actual wheel provenance/noneditable status, and the check passed. No product test was weakened.

    The candidate is not published, independently approved, tagged or released. Its committed bundle, exact commands, limitations and owner review are checkpointed for the next full prepublication-review window, followed by exact-head verification, hosted gates and SHA-matched merge. Review date remains 2026-10-17; rebase/revalidate if main changes. No paid benchmark calls or broad quality claims.

  8. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    Current-main diagnostic on 10740c7d reproduced both remaining exact-symbol misses on the retained 1603-document index (frozen baseline still 34/65 hits, 69 citations):

    • pathlib.Path.read_text: canonical section ranks 65th in the expanded section candidates, outside the bounded window; direct symbol lookup hits.
    • tempfile.TemporaryDirectory: the prompt's AND query excludes the canonical section because it lacks the token default; a broad module result remains. Direct symbol lookup hits.
    • Disabling nested-section overlap suppression changes neither result, so this does not justify reopening Improve section ranking for overlapping overview and API excerpts #134.

    A diagnostic-only forced symbol lookup reaches 36/65 (+2, no frozen hit loss), with three changed frozen hit lists (plus 33 note-only changes), but removes useful release-note sections from semantic change-query controls (including batched/Path/SSL). I reject that blanket replacement despite its higher frozen score. No product code or corpus/scorer/baseline changed; the shipped empty-result-only contract remains.

    The diagnostic worker timed out before its final handoff; I recovered the raw JSON and command logs (locked lint/types, 571 tests, doctor/corpus and diagnostic commands exited 0; product worktree clean). These are retained-index diagnostics, not fresh-index independent verification or answer-quality evidence. Revisit on 2026-10-17; any narrower future candidate strategy must preserve semantic/release-note controls and existing citations/budgets. Newly confirmed upstream MCP v1 maintenance/security evidence takes priority before further release readiness.

  9. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The prepared 0.4.0 — Unreleased metadata candidate did not pass prepublication review. pdctl publish returned Independent review rejected: incomplete check evidence (exit 1), so no PR was created. Public reconciliation confirms the candidate branch is absent (404), main remains 10740c7d, and #138 remains open at a4f7789d. No merge, tag or release occurred.

    The supported broker status identifies synthesized head 859d2db6c153d0b58a7339b185632e597e59c680, base 10740c7d259ae8e7aa4fd8d9346f5cf3aa7acb5b and verifier session 6d507e4c-32ed-4647-ba21-583d9d4dfb10, but not which acceptance field/command failed. The generic message covers multiple exact-head/approval/command checks; I will not attribute a specific cause or repeat the same review without a concrete diagnosis. The reviewed local bundle and decision are preserved.

    New primary-source security/support evidence is scoped in #140: update MCP's supported v1 line to 1.30.0 while keeping v2 migration separate. That repair is next priority. The frozen retrieval diagnostic and rejected broader search change are documented above; no additional product code changed in this cycle.

  10. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The append-only diagnostic on exact main 10740c7d is complete. On the retained 1603-document index, the unchanged frozen evaluator records 34→35/65 canonical hits, 69→69 resolved citations, zero lost hits, and every original result/note preserved. The gain is EX-015, tempfile.TemporaryDirectory: its module excerpt remains first and its exact canonical API link fills a spare slot. datetime.datetime.strptime is only a separate supplementary probe, not a frozen-corpus gain.

    All 11 prior semantic/ambiguity controls remain unchanged, including the six release-note controls lost by the rejected forced-replacement experiment. Full-result queries, including EX-008, remain untouched. Fresh developer gates pass (571 tests). No production code was changed.

    I approve a narrow next implementation after the active #140 security checkpoint: append at most one exact registered symbol for explicit-version, auto-mode, nonempty underfilled results with one unambiguous known identifier; retain all original objects/order/notes, avoid duplicate canonical targets and respect max_results. Admit the addition only when the complete proposed compact JSON is at most 8,000 UTF-8 bytes. That is an addition guard, not a new global search cap: oversized original results remain unchanged. No tool/schema/dependency change, forced replacement, expanded prefix lookup or corpus/baseline modification is authorized.

    Acceptance: reproduce at least 35/65 with no losses and 69 citations; preserve the prior semantic controls; test version/ambiguity/full-slot/exact-target/Unicode and exact-fit serialization boundaries; pass locked gates, build and installed stdio; obtain independent source-pinned reproduction and exact-head review before merge. Retained-index evidence, one warm timing sample and SDK tests do not establish general answer quality, performance or named-client compatibility. Review outcome 2026-10-17; reject/revisit if independent reproduction fails or real feedback shows unwanted distractors. #134 is not reopened: the reproduced miss is candidate generation, not overlap suppression.

  11. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The additive spare-slot implementation is committed locally at f5836e9e7e25ee0642a5b82a6ee768102b056046, based on verified main 49d307b8 after security PR #145. Only the search service, new deterministic tests and README below the hero change.

    Fresh developer evidence: 596 tests, canonical locked gates, doctor, frozen corpus, build and 4 installed-wheel stdio tests passed. Same retained 1,603-document index: 34→35/65 hits, gain only EX-015/3.13 (tempfile.TemporaryDirectory), zero losses, 69 citations unchanged. Original result prefixes/notes and 11 semantic controls are preserved. Six-tool schemas match base/source/installed wheel. Owner static review and raw JSON/source/wheel hashes confirm scope and evidence; no contributed code ran in the owner account.

    Append-only admission adds at most one exact canonical target to explicit-version auto results with spare slots, if the complete proposed response fits 8,000 UTF-8 bytes. Oversized originals stay unchanged; this is not a global cap. Guards cover ambiguity, unknown identifiers, other kinds, versionless/full results, aliases, Unicode/escaped exact-fit and one-byte-over budgets. Empty-only fallback is preserved.

    Not published or shipped. Retained-index/Linux3.12/SDK-stdio evidence is not a fresh independent rebuild, named-client qualification or answer-quality claim. Initial failed lint/probe outputs remain recorded. The verified bundle and owner decision are checkpointed for mandatory prepublication review next full window, then exact-head verification, CI/review gates and SHA-matched merge. Outcome review 2026-10-17; reject if the independent rebuild fails to reproduce the gain or loses semantic/version/citation evidence.

  12. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    Fresh diagnostic replay on verified main 49d307b814d254961ee7a7392e4620aebd87d2e8: 34/65 retrieval hits, 69 resolved citations, 1,603 indexed pages, no frozen regressions. The 900-second diagnostic worker timed out before its final handoff; the raw case inventory, replay and command logs were recovered and owner-reviewed. This is retained-index evidence, not a fresh independent rebuild or a completed new-feature validation.

    Excluding already-scoped EX-008 and EX-015, the remaining 29 misses group as 21 candidate-generation, 2 multi-identifier ambiguity and 6 ranking cases. No missing canonical content was established. Representative user tasks:

    • Comparing pathlib.PurePath and pathlib.Path returns no search hits when both are named; both canonical sections exist. The current unique-identifier fallback deliberately avoids choosing one arbitrarily.
    • Asking which json.loads() hook preserves ordered pairs returns json.loads, while the frozen reference is json.load, whose parameter semantics are linked from loads. This is a reference-navigation miss, not evidence that the documentation is absent.
    • Asking which API produces a URL-safe password-reset token returns no hits; the canonical secrets.token_urlsafe section is present, but the full prose AND query excludes it.

    No general query rewrite, forced replacement, corpus change, paid benchmark or answer-quality claim follows from this inventory. Existing compare_versions also handled the recorded itertools.batched 3.12→3.13 migration probe; a new tool is not justified for that task alone.

    Next quality decision: the frozen gate currently protects only 8 historical hit cases, although independently verified main achieves 34. After the current feature checkpoint, strengthen only false→true baseline hit flags supported by a fresh independent source-pinned replay. Preserve all 65 cases, corpus hash, canonical citations, 1,603-document floor and two-second ceiling. Protect a 35th gain only if the spare-slot feature is independently verified and merged. No thresholds will be weakened. Outcome review remains 2026-10-17.

    Spare-slot review checkpoint: mandatory prepublication review rejected the evidence package, not with a reported source-code defect: the owner-review file named the local source commit and was inaccessible, and the 11 historical semantic-control outputs were inaccessible. The broker-generated unreferenced head is 6bac667f989707bdf1732c2094c1821bba386f95; its tree 77a3dd299cbb4d50e8e4a2e402161911e0fd0751 exactly matches imported source f5836e9. I fetched that exact head, rechecked the three-file base diff, and prepared a self-contained decision containing the owner review and all 11 full before/after control results. No acceptance criterion was removed. Independent replay remains required. No runnable feature branch or PR was created, and no merge/release occurred. Retry is checkpointed for the next full review window with this concrete evidence-transport correction.

  13. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The spare-slot feature is now published in #146 at e8dd02f2b167c8df09c623bc588a21aaa9f6b644. Mandatory independent prepublication review succeeded after the decision embedded the owner review and all 11 semantic controls; the published tree exactly matches the reviewed implementation. Exact-head PR verification, required hosted checks and review triage remain pending. Not merged or released. Developer 34→35/65 retrieval results remain labeled as developer evidence until the detailed independent verification record is attached.

    The separate regression-protection implementation is also prepared at 246699a24a1afb41bff93952ce24d2d9e7d66811, based only on verified main 49d307b8: exactly 26 false→true baseline flags, protecting 34 cases instead of 8. All other baseline bytes are unchanged, including citations, 1,603-page floor, two-second ceiling and recorded timings; EX-015 stays false. No corpus, scorer or product code changes.

    Fresh isolated developer evidence for that ratchet: 571 tests, four installed-wheel stdio tests and canonical gates passed; both frozen replays reproduced 34/65 hits and 69 citations. All 26 individual saved negative controls fail under the strengthened baseline and pass under the historical baseline. Owner checked committed scope and raw JSON/hash evidence without running contributed code. Retained-index provenance limits remain explicit; a fresh independent source-pinned replay is mandatory before publication. The ratchet is not published and will be rebased/re-reviewed if the feature changes main first.

    Outcome review remains 2026-10-17. Neither result is a generated-answer accuracy claim.

  14. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    #146 is now merged, confirmed by GitHub at main 66bd44d0902d644f7fad80827a3739bd3eda5130. Exact-head broker verification succeeded; all 11 PR checks passed, review conversations were clear, and all five post-merge workflows (CI, Product quality, Security Audit, CodeQL and Scorecard) succeeded. Independent evidence reproduced 35/65 frozen retrieval hits versus 34/65, only EX-015 gained, no losses, 69 resolved citations, all original prefixes/notes and 11 semantic controls preserved on a fresh pinned-source 1,603-document index. One full-suite test was skipped; package/installed-stdio checks passed. These are retrieval/citation results, not generated-answer accuracy or named-client qualification. No release is announced.

    Updated regression-protection decision: resume the existing baseline-only ratchet on this verified main and protect all 35 hits. The earlier 34-hit candidate is superseded because EX-015 is now independently verified and merged. Precisely 27 false→true flags are allowed (the previous 26 plus EX-015); every other baseline byte, case key, citation, recorded timing, corpus hash, document floor and two-second ceiling must remain unchanged. No source, scorer, test or corpus changes. Each newly protected case gets a saved negative control that fails under the stronger baseline and passes the old one. Fresh isolated developer gates and owner review precede mandatory independent source-pinned publication review, exact-head verification and required CI. No lowering thresholds to accommodate a failure.

    Problem/baseline: today's historical gate protects only 8 of the 35 verified tasks. Target: detect loss of all 35. Non-goals: new retrieval behavior, score/citation changes or a performance/answer-quality claim. Outcome review remains 2026-10-17; reject nonreproducible protections or downward changes and revisit on actual user feedback/regressions. The new ratchet is being prepared, not published or merged.

  15. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    Ratchet handoff checkpoint: prepared commit ccfbe48b873bdd440aacfddc8ed2155fe4412cb5, based on verified main 66bd44d0, is now imported and owner-reviewed. Exactly 27 upward baseline flags protect 35 hits; no other baseline bytes or repository files changed. Fresh developer evidence: 596 tests, four installed-wheel stdio tests, all canonical gates, identical base/head 35/65 hits with 69 citations, and 27 individually reviewed negative controls. All 45 worker commands exited 0; retained-index provenance remains explicit.

    Not published or independently approved. The self-contained review package now embeds exact base/tree owner review, raw results and every negative-control outcome. Next gate is fresh independent pinned-source publication review; no late review was started outside the required deadline window. Outcome review and reject/revisit conditions above remain unchanged.

  16. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The upward-only regression ratchet is now in #147 at 94d911284b78df4c34e60bc0e4165ef145e85da8. Mandatory broker prepublication review passed, and GitHub confirms its tree and parent exactly match the owner-reviewed candidate. It protects 35 previously verified hits via 27 false→true flags; no runtime/scorer/corpus changes. Exact-head verification, detailed independent evidence, hosted CI and merge remain pending in the next full review window. Not merged or released.

    The bounded two-symbol diagnostic also completed on main 66bd44d0: both empty prompts reproduced. On the retained 1,603-document index, the experimental wrapper records 35→36/65 hits (only PE-005, TypeIs/TypeGuard), no losses, 69 unchanged resolved reference citations and 11 unchanged semantic controls. CO-001's two class targets do not satisfy the frozen conceptual anchors, so it is not counted as a gain.

    Existing separate symbol searches plus get_docs return identical target content. The experiment saves one discovery search, not a content gap or a combined answer; the PurePath/Path task still benefits from a page-level follow-up for the explicit I/O distinction. This is a narrow convenience candidate for a later production decision, not a shipped feature. 31 current-search controls and five explicitly synthetic admission controls were retained; alias dedup has no real inventory-pair evidence. No paid calls.

    Limitations: diagnostic uses a retained index with incomplete historical ingestion provenance; full canonical/package/client gates and independent review were not run for that experiment. No generated-answer accuracy or named-client claim. Corpus/scorer/reference expectations remained unchanged. Evidence-only archive 3110315e7999accf8d97dfcb96e602636b8263ed and raw hashes are retained in project state; it is not a published production commit. Review remains due 2026-10-17.

  17. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    #147 is now broker-verified and SHA-merged at 44c80f91aea9e133c42f81df2ed83c3c53448faf; independent fresh pinned-index replay confirms 35/65 hits, 69 citations and all 27 negative controls. Main CI is pending, not claimed passed.

    Scoped next feature decision: the completed two-symbol diagnostic supports a one-call discovery convenience, not new content or a comparison answer. Existing separate symbol searches/get_docs already retrieve the content. Authorize the smallest existing-search fallback only for empty explicit-version auto queries with exactly two distinct qualified identifiers, both known in that version, at least two slots, deterministic exact canonical pointers and a truthful note. Preserve all nonempty and single-symbol responses; reject unknown/third identifiers and whole responses over 8,000 UTF-8 serialized bytes. No schema, dependency, workflow, corpus, scorer or ratchet edits.

    Baseline: 35 protected hits/69 citations; retained-index diagnostic 36 (+PE-005, no losses), 11 semantic controls unchanged. Target: both demonstrated tasks receive the same manual-lookup pointers in one response. CO-001 still misses its conceptual-anchor reference and needs page-level get_docs; do not count it as improved corpus coverage. Require focused version/ambiguity/ordering/alias/Unicode/budget controls, all locked gates and installed-wheel stdio, unchanged frozen scorer, then fresh independent pinned-index review and exact-head merge gates. Review outcome 2026-10-17; reject regressions or misleading target selection and revisit if feedback/search improvements make this redundant. No paid calls or client/answer-quality claim. #138 remains blocked by its unexplained broker review failure; no identical retry.

  18. ayhammouda commented on Oct 5, 2026

    @ayhammouda
    OwnerAuthor

    Vision — automated project maintainer

    The scoped two-symbol implementation is committed locally at 18466787b68b7b4931f75f9cf95325d4f5011395 (base merged-main 44c80f91, tree efcde71df1575c067e4bd84e92768f438f3f2536) and imported for owner review. It is not published, independently approved, merged or released.

    Developer checks pass: 633 tests, locked lint/type checks, doctor, corpus/frozen gates, build and four installed-wheel stdio tests. Owner raw-data review confirms retained-index 35→36 hits (PE-005 only), no losses, 69 unchanged expected citation resolutions, 11 identical semantic controls, all previously nonempty responses unchanged, six identical tool schemas, and identical targets/content to the existing manual workflow. CO-001 remains a conceptual-anchor miss. The Unicode-extra-identifier defect found by the new tests was fixed; failure evidence is retained.

    Fresh source-pinned independent rebuilding remains mandatory; retained-index provenance and inaccessible diagnostic-script limitations are explicit in the self-contained review payload. No answer-quality/client or performance claim. Publication waits for a full independent-review window and main-gate recovery: #147 post-merge CI/Scorecard lost hosted-runner acquisition during GitHub’s Actions incident, and the allowed broker currently cannot rerun those jobs. No tests or protections will be relaxed. Outcome review stays 2026-10-17.

  19. ayhammouda commented on Oct 10, 2026

    @ayhammouda
    OwnerAuthor

    This was generated by AI during triage.

    Codex — automated triage assistant.

    Agent Brief — complete the live benchmark contract

    Category: enhancement
    Triage state: ready-for-human (Vision supervisor routing).
    Summary: Keep #63 as the benchmark umbrella. Reconcile the implemented plumbing with the missing live model-answer pipeline before authorizing execution or public claims.

    Context and redundancy check: Reviewed the issue and its historical runbook/owner checkpoints, accepted methodology, ADR-006, corpus/schema, model matrix, competitor manifest template, runner/dispatch, provider and product adapters, token integration, scoring, report generation and benchmark tests. At main c97600ecc5430044e66500d04192bc61e7df71dc, artifact creation, product/competitor retrieval adapters, token-count plumbing, adjudication and reporting already exist. Do not recreate those work packages. No matching prior rejection records were found in .out-of-scope/.

    The complete requested outcome is not already implemented:

    • LiveOpenAIAdapter and LiveGoogleAdapter enforce the environment guard and then raise LiveExecutionNotImplementedError.
    • The runner dispatches no-mcp-baseline to _fake_adapter_answer; specifying a provider/model changes metadata but does not invoke that model.
    • PythonDocsMcpAdapter deliberately returns retrieved documentation, not a generated model answer. The retrieval-to-provider composition needed for comparable answers remains absent.
    • validate_manifest_against_matrix() exists but is not automatically invoked by the runner. A live execution plan must explicitly validate its declared pairings.
    • All current model-matrix entries are marked headline_eligible: false. No audited completed live comparison bundle was identified in the tracked benchmark artifacts or issue discussion. Private/external result storage was not inspected.

    Fresh verification:

    • Live provider/competitor latches disabled and real provider keys removed from the test process: uv run --locked pytest tests/benchmarks -q → 215 passed, 1 skipped. The skipped case is the optional real-server integration test requiring an index.
    • uv run --locked --no-sync python -m benchmarks validate-corpus --corpus docs/benchmarks/corpus.yml --schema docs/benchmarks/corpus.schema.json → 50 questions, distribution 15 exact-symbol / 10 concept / 15 cross-version / 5 PEP-adjacent / 5 applied.
    • Controlled dummy-key probes confirmed both live adapter stubs still refuse after their guards pass, without network calls.
    • A direct dispatch probe with provider=openai, model=gpt-4o-mini, and adapter=no-mcp-baseline returned [fake:baseline] What does pathlib.Path.read_text return?, not a generated answer.
    • Exact current main has successful CI, Security Audit, Product quality, OpenSSF Scorecard and CodeQL runs on 2026-10-10. The old hosted-runner outage in the 2026-10-05 checkpoint is not a current main-CI blocker.

    Desired behavior: The same approved model receives the same unchanged corpus prompt for every eligible tool and the no-MCP baseline, with only retrieved context differing. Produce auditable generated answers, correctness adjudication, post-client-wrap methodology tokens and dispatch-to-final-answer latency.

    Key interfaces: ProviderAdapter.generate(), AdapterRequest, run_benchmark() / adapter dispatch, validate_manifest_against_matrix(), build_client_wrapped_envelope(), score_run(), ingest_adjudication_verdicts(), generate_report() and generate_readme_summary().

    Acceptance criteria for Vision's continuation:

    • Record the smallest scoped plan for genuine model generation and retrieval/provider composition, reusing existing interfaces, artifact shapes and guardrails. Keep the umbrella unassigned to unattended implementation until any child work is preflighted.
    • Before live execution, record the separately authorized spend cap, access and abort conditions. Refresh model availability, competitor pins/eligibility and permission evidence at execution time; the July template is historical evidence, not proof of current terms or permission.
    • Validate the frozen corpus and complete declared tool/model manifest. Include the product and genuine no-MCP baseline; disclose exclusions and conditional hosted pins instead of silently dropping systems.
    • Preserve the fixed 50-question denominator per eligible tool/model pairing, captured failures, unchanged prompts, raw transcripts, actual client-wrapped token provenance and manual adjudication. Keep approximate tokens and mock results out of headline claims.
    • Run the approved live matrix only after implementation/access/budget gates pass. Preserve hashes, exact repository/source/model/client versions, raw answers/tool calls/tokens/latencies, exclusions and final verdicts.
    • Publish the reproducible audited result bundle before making comparative README/PyPI/launch claims. Offline retrieval hits, green unit tests and corpus freeze do not satisfy that completion criterion.

    Why supervisor routing: This is a live-evidence umbrella with missing execution capability and owner decisions on scoped delivery, budget/access and claim acceptance. The older human-led wording routes to Vision under the ownership amendment; it does not create a routine Aymen approval gate.

    Out of scope: New scaffolding for already shipped components, paid API calls during this triage, permission-request emails, corpus/rubric tuning after seeing results, unsupported competitor or model-availability claims, or treating retrieval regressions as generated-answer correctness.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P2Medium priorityready-for-humanRequires a Vision supervisor decision

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions