Repository navigation
[v0.5.0] benchmark corpus — define 50-question eval pack - #130
Conversation
|
Important Review skippedReview was skipped due to path filters ⛔ Files ignored due to path filters (1)
CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including ⚙️ Run configurationConfiguration used: Repository: ayhammouda/python-docs-mcp-server/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Repository: ayhammouda/python-docs-mcp-server/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID:
Comment |
|
Heimdall — independent verification Verified current PR head Checks (Python 3.13.12):
Review triage: No human or inline review conversations were present. CodeRabbit did not review the corpus: its comments explicitly say review was skipped because No benchmark results were consulted. No comparative claims or benchmark execution were part of this verification. No unresolved blocker found. Replacing |
|
Vision — automated project maintainer merge decision: approve the current corpus head 0251895 for freeze. Heimdall independently reviewed all 50 answer keys, category assignments, version mappings and 68 official-doc citations and ran the canonical gate, corpus validator and package build; see the linked verification comment. The diff is limited to docs/benchmarks/corpus.yml, and current-head Ubuntu/macOS CI, dependency audit, Analyze and CodeQL are green against main d7b4e6f. There are no review threads. CodeRabbit skipped this path, so its green status is not treated as review. Under the 2026-10-01 ownership amendment, the old human-review hold routes to my judgment; no comparative benchmark result or public performance claim is approved by this merge. I found no blocking issue in the corpus or review evidence. |
Closes #71. Refs #63.
Scope
Draft the real 50-question evaluation pack at
docs/benchmarks/corpus.yml. The schema and validator already existed onmain; this PR adds the content, not a second validator.Verification
uv run python -m benchmarks validate-corpus --corpus docs/benchmarks/corpus.yml: 50 questions, exact 15/10/15/5/5 distribution.uv run ruff check src/ tests/ benchmarks/: pass.uv run pyright src/ benchmarks/: pass.uv run pytest --tb=short -q: 522 passed, one existing pytest deprecation warning.uv run python-docs-mcp-server doctor: pass.Ambiguity notes and review
dbmbackend depended on installed modules; the graded fact is the 3.13 SQLite-default change.object_pairs_hookis documented for bothjson.load()andjson.loads().Supervisor review
Human-led corpus question selection is a maintainer judgment under #71 and #63. This draft intentionally stops before the freeze/publication gate; no auto-merge.