Repository navigation
[v0.5.0] benchmark scoring — correctness scorer with manual-adjudication hooks #88
Copy link
Copy link
Closed
Labels
agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
Milestone
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
on Jul 8, 2026 Agent context — #88 [v0.5.0] benchmark scoring — correctness scorer with manual-adjudication hooks
Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
Auditable copy: this content is mirrored as a comment on issue #88 (.planning/is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).1. Roadmap excerpt (why this exists)
- STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness deliverable; scoring is methodology work package 5.
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md"Correctness Scoring" (lines 119–136): one score per answer from{1.0, 0.5, 0.0}; public reporting needs BOTH mean and per-category correctness; an answer that looks correct but lacks evidence from the supplied docs is marked in the raw results and discussed separately (the "correct-but-ungrounded" flag).- [v0.5.0] benchmark — public token/correctness/latency harness vs docs MCPs (human-led) #63 decisions (2026-06-08): failed cells score 0.0 and stay in the denominator; the denominator is fixed at the full corpus count per tool×model cell. The merged runner already enforces both — the scorer INHERITS these decisions, it never re-litigates them.
- Actual scoring of real answers stays human-adjudicated: this issue is plumbing only.
2. Code touch-points (verify against merged main before starting)
benchmarks/runner.py:321-354_scoring_record: failed cells →status: "failed",score: 0.0,requires_manual_scoring: False(337–347); succeeded cells →status: "placeholder",score: None,requires_manual_scoring: True(348–354). Every record carriesincluded_in_correctness_denominator: True,denominator_unit: "corpus_query",tool_model_key,error_category.benchmarks/runner.py:363(_write_cell_artifacts): scoring artifacts live at<run-dir>/scoring/<competitor_id>/<corpus_id>.json; transcripts (your grounding-evidence input) attranscripts/<competitor_id>/<corpus_id>.json.benchmarks/runner.py:139,143run-summary:correctness_denominator_cellsandfailed_cells_included_in_correctness_denominator: True— completed records must stay consistent with these; never shrink a denominator.- Downstream consumer
benchmarks/report.py:98-100,217-220: readsscore,requires_manual_scoring,included_in_correctness_denominatorper record; mean uses non-None scores (line 359), pending count usesrequires_manual_scoring(line 360). A completed record therefore sets a numericscoreandrequires_manual_scoring: Falseso [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74's report picks it up unchanged. - Corpus schema is NOT on main.
git ls-files docs/benchmarks/shows onlyPUBLIC-BENCHMARK-METHODOLOGY.mdandmodel-matrix.yml— nocorpus.schema.json([v0.5.0] benchmark corpus — mechanical slice: schema, validator, placeholder fixture (split from #71) #94 unmerged at verification time). Re-check at start; if still absent or missing rubric fields, the §4 recovery clause fires immediately. - New code: scorer module under
benchmarks/+ additivepython -m benchmarkssubcommand(s) — one for scoring, one for adjudication-queue ingest. Follow the argparse subparsers pattern inbenchmarks/__main__.py(parser build lines 17–62, dispatch inmain()65–111); do not disturb the existingrun/reportsubcommands. - Score vocabulary is closed:
{1.0, 0.5, 0.0}. Human verdicts ALWAYS win — ingest must never let an automatic pass overwrite an ingested human verdict; failed cells (runner-scored 0.0) pass through untouched. tool_model_keyformat (benchmarks/runner.py:311-318_tool_model_key):id:provider/modelwhen both set,id:modelwith model only, else bareid. Preserve it verbatim on completed records — per-model reporting keys on it.
3. Test patterns to follow
- Model:
tests/benchmarks/test_report.py:232-295_minimal_valid_run_dir— a hand-authored run-dir factory undertmp_path(run-summary.json,environment.json,snapshots/, per-cell scoring/latency/tokens JSON). Build the scorer's fixtures the same way, adding fixture answer keys and fixture transcripts. - Fixture answer keys + transcripts ONLY — no LLM-as-judge, no network, no real corpus questions authored.
- Pinned test file:
tests/benchmarks/test_scoring.py(souv run pytest tests/benchmarks -qexercises it). - Cover: rubric application for each score value, per-category rollups, correct-but-ungrounded flag, adjudication-queue emission, human-verdict ingest + never-overwrite guarantee, failed-cell passthrough, denominator invariance.
4. Known pitfalls
- Recovery clause (verbatim from [v0.5.0] benchmark scoring — correctness scorer with manual-adjudication hooks #88): "If the corpus schema ([v0.5.0] benchmark corpus — define schema and 50-question eval pack #71) lacks fields the rubric needs (expected answer properties, required citations), stop and comment here so [v0.5.0] benchmark corpus — define schema and 50-question eval pack #71 can be corrected — do not invent schema fields."
- Merged tests are additive-only regression cover; NO sanction to modify any merged test exists for this issue.
- Never author corpus questions or answer keys ([v0.5.0] benchmark corpus — define schema and 50-question eval pack #71 is human-led); tests use synthetic fixtures only.
pyproject.tomlanduv.lockare untouchable — the scorer needs no new dependency.- No README edits, no benchmark claims anywhere.
Refs #63, neverCloses #63;Closes #88only if all criteria met. - Lint blind spot: run
uv run ruff check benchmarks/anduv run pyright benchmarks/in addition to thesrc//tests/gates (full command list in the issue). - Branch:
agent/88-correctness-scorer. - Sequencing: [v0.5.0] benchmark corpus — mechanical slice: schema, validator, placeholder fixture (split from #71) #94/[v0.5.0] benchmark corpus — define schema and 50-question eval pack #71 must land the corpus schema before real scoring can be exercised; merge latest main into your branch and re-run the full validation gate before any merge decision.
5. Decision log (worker fills in)
- …
- addedagent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
on Jul 8, 2026
Metadata
Metadata
Assignees
Labels
agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
Context
Parent: #63. Methodology:
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md(work package 5; "Correctness Scoring" section defines the 1.0/0.5/0.0 rubric). Runner scoring records (merged #75) already carryscore: None+requires_manual_scoring: Trueplaceholders andincluded_in_correctness_denominator.Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, create the
.planning/agent-context/file (decision 5.14, mirrored as an issue comment), and applyagent-readybefore any dispatch. Actual scoring of real answers stays human-adjudicated — this issue is the plumbing only.Goal
Add the correctness scorer: rubric application against the corpus answer keys, with manual-adjudication hooks, writing completed scoring records the report generator (#74) consumes.
Acceptance criteria (draft — Vision edits before agent-ready)
score ∈ {1.0, 0.5, 0.0}, per-category rollups, and the methodology's "correct-but-ungrounded" flag (answers lacking evidence from supplied docs are marked and reported separately).tests/benchmarks/) use fixture answer keys and transcripts only — no LLM-as-judge, no network, no real corpus questions authored.Refs #63, neverCloses #63.Scope boundaries
In scope: scorer, adjudication queue/ingest, tests on fixtures. Out of scope: authoring corpus questions or answer keys (#71, human-led), LLM-based auto-scoring (not in methodology), running the benchmark, reporting layout (#74), README.
Forbidden-territory reminder
Standard §2 list applies; merged tests additive-only.
Validation commands
Recovery
If the corpus schema (#71) lacks fields the rubric needs (expected answer properties, required citations), stop and comment here so #71 can be corrected — do not invent schema fields.
Effort estimate
4–6 hours.
Refs #63.