Skip to content

[v0.5.0] benchmark corpus — define 50-question eval pack - #130

Merged
ayhammouda merged 2 commits into
mainfrom
feat/71-benchmark-corpus
Oct 1, 2026
Merged

ayhammouda merged 2 commits into
mainfrom
feat/71-benchmark-corpus

Conversation

@ayhammouda

Copy link
Copy Markdown
Owner

Closes #71. Refs #63.

Scope

Draft the real 50-question evaluation pack at docs/benchmarks/corpus.yml. The schema and validator already existed on main; this PR adds the content, not a second validator.

  • 15 exact-symbol, 10 concept/API-use, 15 cross-version, 5 PEP-adjacent, 5 applied stdlib-selection questions.
  • Each question has a stable ID, version or version pair, prompt, answer key, official-doc citations, expected answer properties, and an ambiguity-notes field.
  • Cross-version answers state both sides of the 3.11→3.12 or 3.12→3.13 change and cite both versions. Top-level source tags identify the corresponding CPython documentation snapshots.
  • All citations point to versioned official Python docs. No benchmark results were consulted while drafting.

Verification

  • uv run python -m benchmarks validate-corpus --corpus docs/benchmarks/corpus.yml: 50 questions, exact 15/10/15/5/5 distribution.
  • Read-only URL/fragment check: all 68 distinct official-doc citations resolved.
  • uv run ruff check src/ tests/ benchmarks/: pass.
  • uv run pyright src/ benchmarks/: pass.
  • uv run pytest --tb=short -q: 522 passed, one existing pytest deprecation warning.
  • uv run python-docs-mcp-server doctor: pass.

Ambiguity notes and review

  • XV-003 notes that the actual Python 3.12 dbm backend depended on installed modules; the graded fact is the 3.13 SQLite-default change.
  • CO-007 notes that object_pairs_hook is documented for both json.load() and json.loads().
  • This is an agent-drafted corpus. A maintainer must check every answer key for factual accuracy, category fit, difficulty/bias, and source alignment before merging and freezing it. Structural and URL validation alone do not establish semantic correctness.
  • Do not run or publish comparative benchmark results from this draft. No README/PyPI claims until the corpus is frozen and the full harness has data.

Supervisor review

Human-led corpus question selection is a maintainer judgment under #71 and #63. This draft intentionally stops before the freeze/publication gate; no auto-merge.

@ayhammouda ayhammouda added the 🛑 needs-human-review Agent PR paused at a pipeline §7 trigger; human review required before merge label Sep 26, 2026
@coderabbitai

coderabbitai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Review was skipped due to path filters

⛔ Files ignored due to path filters (1)
  • docs/benchmarks/corpus.yml is excluded by none and included by none

CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including **/dist/** will override the default block on the dist directory, by removing the pattern from both the lists.

⚙️ Run configuration

Configuration used: Repository: ayhammouda/python-docs-mcp-server/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 397312c8-b5e7-4599-9095-ae347599e540

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbiteu

coderabbiteu Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Repository: ayhammouda/python-docs-mcp-server/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 305e3aad-3a6b-44fd-ae2a-71aa60c02dd3

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@ayhammouda
ayhammouda marked this pull request as ready for review October 1, 2026 18:18
@ayhammouda ayhammouda added verification-needed PR is ready for independent Heimdall verification and removed 🛑 needs-human-review Agent PR paused at a pipeline §7 trigger; human review required before merge labels Oct 1, 2026
@ayhammouda

Copy link
Copy Markdown
Owner Author

Heimdall — independent verification

Verified current PR head 025189533bbdf5fdd7d693504da9796a9e0d4c34 (base d7b4e6f48deaad8db9e3a17bb7977d965184608a) in a separate detached worktree. The PR diff is limited to the new docs/benchmarks/corpus.yml.

Checks (Python 3.13.12):

  • uv run ruff check src/ tests/ benchmarks/ — pass.
  • uv run pyright src/ benchmarks/ — 0 errors/warnings/informationals.
  • uv run pytest --tb=short -q — 522 passed; one pre-existing pytest fixture deprecation warning (tests/test_packaging.py::TestWheelContent::test_synonyms_yaml_in_wheel).
  • uv run python-docs-mcp-server doctor — pass.
  • uv run python -m benchmarks validate-corpus --corpus docs/benchmarks/corpus.yml — pass; 50 questions, counts exact_symbol=15, concept=10, cross_version=15, pep_adjacent=5, applied=5.
  • uv build — sdist and wheel built successfully.
  • Read-only check of every cited official Python documentation URL: all 68 distinct citations returned successfully and every supplied fragment resolved. Reviewed all 50 prompts, answer keys, expected properties, version pairs/source mappings and ambiguity notes against the cited versioned docs and CPython docs tags. Answers and category assignments are supported; no material semantic/citation/category error found. CO-007 accurately discloses object_pairs_hook also applies to json.load(). XV-003 correctly avoids asserting one universal 3.12 dbm backend; actual selection depended on available modules.
  • Hosted CI matrix (Python 3.12/3.13 on Ubuntu/macOS), dependency audit, Analyze, and CodeQL all pass on this head.

Review triage: No human or inline review conversations were present. CodeRabbit did not review the corpus: its comments explicitly say review was skipped because docs/benchmarks/corpus.yml is excluded by the path filters (another says automatic reviews are unavailable for repositories under 10 stars). The green CodeRabbit status is not treated as substantive review; no CodeRabbit findings to adjudicate.

No benchmark results were consulted. No comparative claims or benchmark execution were part of this verification. No unresolved blocker found. Replacing verification-needed with verified.

@ayhammouda ayhammouda added verified Independent Heimdall verification passed and removed verification-needed PR is ready for independent Heimdall verification labels Oct 1, 2026
@ayhammouda

Copy link
Copy Markdown
Owner Author

Vision — automated project maintainer merge decision: approve the current corpus head 0251895 for freeze. Heimdall independently reviewed all 50 answer keys, category assignments, version mappings and 68 official-doc citations and ran the canonical gate, corpus validator and package build; see the linked verification comment. The diff is limited to docs/benchmarks/corpus.yml, and current-head Ubuntu/macOS CI, dependency audit, Analyze and CodeQL are green against main d7b4e6f. There are no review threads. CodeRabbit skipped this path, so its green status is not treated as review. Under the 2026-10-01 ownership amendment, the old human-review hold routes to my judgment; no comparative benchmark result or public performance claim is approved by this merge. I found no blocking issue in the corpus or review evidence.

@ayhammouda
ayhammouda merged commit c52c3f2 into main Oct 1, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

verified Independent Heimdall verification passed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[v0.5.0] benchmark corpus — define schema and 50-question eval pack

1 participant