Skip to content

[v0.5.0] benchmark scoring — correctness scorer with manual-adjudication hooks #88

Description

@ayhammouda

Context

Parent: #63. Methodology: docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md (work package 5; "Correctness Scoring" section defines the 1.0/0.5/0.0 rubric). Runner scoring records (merged #75) already carry score: None + requires_manual_scoring: True placeholders and included_in_correctness_denominator.

Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, create the .planning/agent-context/ file (decision 5.14, mirrored as an issue comment), and apply agent-ready before any dispatch. Actual scoring of real answers stays human-adjudicated — this issue is the plumbing only.

Goal

Add the correctness scorer: rubric application against the corpus answer keys, with manual-adjudication hooks, writing completed scoring records the report generator (#74) consumes.

Acceptance criteria (draft — Vision edits before agent-ready)

  • Scorer reads a run's scoring placeholders + the corpus answer keys and produces per-cell records with score ∈ {1.0, 0.5, 0.0}, per-category rollups, and the methodology's "correct-but-ungrounded" flag (answers lacking evidence from supplied docs are marked and reported separately).
  • Manual-adjudication hook: cells the automatic pass cannot decide are emitted to an adjudication queue file; a CLI subcommand ingests human verdicts back and re-emits final records. Automatic scoring never overwrites a human verdict.
  • Failed cells (already scored 0.0 by the runner) pass through untouched; the denominator stays the full corpus count per tool×model cell.
  • Tests (additive, under tests/benchmarks/) use fixture answer keys and transcripts only — no LLM-as-judge, no network, no real corpus questions authored.
  • No README/benchmark claims; Refs #63, never Closes #63.

Scope boundaries

In scope: scorer, adjudication queue/ingest, tests on fixtures. Out of scope: authoring corpus questions or answer keys (#71, human-led), LLM-based auto-scoring (not in methodology), running the benchmark, reporting layout (#74), README.

Forbidden-territory reminder

Standard §2 list applies; merged tests additive-only.

Validation commands

uv run ruff check src/ tests/
uv run ruff check benchmarks/
uv run pyright src/
uv run pyright benchmarks/
uv run pytest --tb=short -q
uv run pytest tests/benchmarks -q
uv run python-docs-mcp-server doctor

Recovery

If the corpus schema (#71) lacks fields the rubric needs (expected answer properties, required citations), stop and comment here so #71 can be corrected — do not invent schema fields.

Effort estimate

4–6 hours.

Refs #63.

Activity

  1. added this to the v0.5.0 milestone on Jul 8, 2026
  2. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Agent context — #88 [v0.5.0] benchmark scoring — correctness scorer with manual-adjudication hooks

    Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
    Auditable copy: this content is mirrored as a comment on issue #88 (.planning/ is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).

    1. Roadmap excerpt (why this exists)

    • STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness deliverable; scoring is methodology work package 5.
    • docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md "Correctness Scoring" (lines 119–136): one score per answer from {1.0, 0.5, 0.0}; public reporting needs BOTH mean and per-category correctness; an answer that looks correct but lacks evidence from the supplied docs is marked in the raw results and discussed separately (the "correct-but-ungrounded" flag).
    • [v0.5.0] benchmark — public token/correctness/latency harness vs docs MCPs (human-led) #63 decisions (2026-06-08): failed cells score 0.0 and stay in the denominator; the denominator is fixed at the full corpus count per tool×model cell. The merged runner already enforces both — the scorer INHERITS these decisions, it never re-litigates them.
    • Actual scoring of real answers stays human-adjudicated: this issue is plumbing only.

    2. Code touch-points (verify against merged main before starting)

    • benchmarks/runner.py:321-354 _scoring_record: failed cells → status: "failed", score: 0.0, requires_manual_scoring: False (337–347); succeeded cells → status: "placeholder", score: None, requires_manual_scoring: True (348–354). Every record carries included_in_correctness_denominator: True, denominator_unit: "corpus_query", tool_model_key, error_category.
    • benchmarks/runner.py:363 (_write_cell_artifacts): scoring artifacts live at <run-dir>/scoring/<competitor_id>/<corpus_id>.json; transcripts (your grounding-evidence input) at transcripts/<competitor_id>/<corpus_id>.json.
    • benchmarks/runner.py:139,143 run-summary: correctness_denominator_cells and failed_cells_included_in_correctness_denominator: True — completed records must stay consistent with these; never shrink a denominator.
    • Downstream consumer benchmarks/report.py:98-100,217-220: reads score, requires_manual_scoring, included_in_correctness_denominator per record; mean uses non-None scores (line 359), pending count uses requires_manual_scoring (line 360). A completed record therefore sets a numeric score and requires_manual_scoring: False so [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74's report picks it up unchanged.
    • Corpus schema is NOT on main. git ls-files docs/benchmarks/ shows only PUBLIC-BENCHMARK-METHODOLOGY.md and model-matrix.yml — no corpus.schema.json ([v0.5.0] benchmark corpus — mechanical slice: schema, validator, placeholder fixture (split from #71) #94 unmerged at verification time). Re-check at start; if still absent or missing rubric fields, the §4 recovery clause fires immediately.
    • New code: scorer module under benchmarks/ + additive python -m benchmarks subcommand(s) — one for scoring, one for adjudication-queue ingest. Follow the argparse subparsers pattern in benchmarks/__main__.py (parser build lines 17–62, dispatch in main() 65–111); do not disturb the existing run/report subcommands.
    • Score vocabulary is closed: {1.0, 0.5, 0.0}. Human verdicts ALWAYS win — ingest must never let an automatic pass overwrite an ingested human verdict; failed cells (runner-scored 0.0) pass through untouched.
    • tool_model_key format (benchmarks/runner.py:311-318 _tool_model_key): id:provider/model when both set, id:model with model only, else bare id. Preserve it verbatim on completed records — per-model reporting keys on it.

    3. Test patterns to follow

    • Model: tests/benchmarks/test_report.py:232-295 _minimal_valid_run_dir — a hand-authored run-dir factory under tmp_path (run-summary.json, environment.json, snapshots/, per-cell scoring/latency/tokens JSON). Build the scorer's fixtures the same way, adding fixture answer keys and fixture transcripts.
    • Fixture answer keys + transcripts ONLY — no LLM-as-judge, no network, no real corpus questions authored.
    • Pinned test file: tests/benchmarks/test_scoring.py (so uv run pytest tests/benchmarks -q exercises it).
    • Cover: rubric application for each score value, per-category rollups, correct-but-ungrounded flag, adjudication-queue emission, human-verdict ingest + never-overwrite guarantee, failed-cell passthrough, denominator invariance.

    4. Known pitfalls

    5. Decision log (worker fills in)

    • …
  3. added
    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
    on Jul 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestpriority:P2Medium priority

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions