Repository navigation
[v0.5.0] benchmark reporting — generate raw report and README-safe summary #74
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priorityagent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
on Jun 8, 2026 Agent context — #74 [v0.5.0] benchmark reporting — generate raw report and README-safe summary
Status: FINAL (verified against merged
main@1f202b6; PR #75 mergedad18406, PR #85 merged1f202b6, both 2026-07-08).
Auditable copy: this content is mirrored as a comment on issue #74 (.planning/is gitignored — the issue comment is the copy pre-flight verifies and dispatch links).1. Roadmap excerpt (why this exists)
STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness with mandatory methodology disclosure. Decision 5.17 (evidence ladder): no comparative/benchmark claim enters README/PyPI/launch copy until reproducible data exists — your generator is the enforcement point: it produces a README-safe template, never a README edit, and refuses to render when required metadata/corpus-hash/raw files are missing.
2. Code touch-points (verify against merged main before starting)
- Input: the runner's per-run artifact tree —
run-summary.json,environment.json,planned-cells.json,snapshots/(byte-copies of corpus + manifest after the [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #75 merge-readiness fixes — corpus hash = hash of the snapshot bytes, which must equal the frozen input file's), and per-celltranscripts|tokens|latency|scoring|failures/<competitor>/<qid>.json. - Key fields: scoring records carry
included_in_correctness_denominator,error_category,score(0.0 for failures,None+requires_manual_scoringfor placeholders); latency records carrylatency_ms,status,error_category; token records carryclient_wrapped_tokens/raw_payload_tokens(may beNoneplaceholders — report must surface that honestly, not as zero);tool_model_key= strict tool+model pairing (id:provider/model) — README-safe tables use these pairings, never tool-only rows. docs/benchmarks/model-matrix.yml(from [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73/PR [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85, on main): entries record provider, model_id, client, token_count_method, latency_scope, headline_eligible. Loader:benchmarks/model_matrix.py(validates schema, exposestool_model_cells()for tool×model cross-multiplication andMETHODOLOGY_TOKEN_LABEL = "Claude Tokens (Normalized Payload)"). The report must include the model/client matrix, keep per-model results visible, mark any aggregate as aggregate, and never average across model families — this is [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73's one deferred acceptance criterion, now YOURS to satisfy.- Known gap you inherit but do not fix: the matrix is NOT yet wired into
benchmarks/runner.py's_execute_celldispatch (deliberate [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73 deferral, tracked for the WP3 issue). Your generator consumes run artifacts + the matrix file as they exist. If the report needs a per-model artifact field the runner doesn't emit, that is exactly issue [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74's recovery clause: STOP and comment on [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 naming the missing fields — do not invent fields and do not wire the runner yourself. - Output:
docs/benchmarks/results/<run-id>/REPORT.mdgenerator + separately generated README-safe summary block (compact tables + links to methodology and raw bundle) + error/timeout-rate column so unstable tools can't hide behind surviving-query correctness. - New module:
benchmarks/report.pybehind a thin subcommand — keepbenchmarks/__main__.pychurn minimal (shared surface with [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73's changes).
3. Test patterns to follow
tests/benchmarks/test_runner.py:_write_yaml-style helpers build fixtures undertmp_path; assert file paths and parsed JSON content. For reporting: generate a fake run-artifact tree undertmp_path, run the generator, assert REPORT.md content, the refusal path (missing metadata/corpus-hash/raw files → clean validation error), failed-competitor disclosure, and aggregate/per-model separation. New tests MUST live undertests/benchmarks/(souv run pytest tests/benchmarks -qexercises them).4. Known pitfalls
- Fixture outputs go to
tmp_pathortests/benchmarks/fixtures/— NEVER commit anything underdocs/benchmarks/results/unless unmistakably labeled SYNTHETIC/FIXTURE DATA in the title and every table. The PR-template's "example generated report from fixture data" belongs in the PR body, not committed to the public docs tree. - README is not touched at all in this task — the README-safe block is a generated template file; publication is a maintainer act gated on real data (decision 5.17).
- Lint blind spot: also run
uv run ruff check benchmarks/anduv run pyright benchmarks/(see [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84). - Merged benchmark tests are §2 regression cover — additive only; stop and comment per §2/§8 if a merged test seems to need weakening.
- Missing result fields: if the report needs fields the runner doesn't emit, STOP and comment on [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 naming the missing fields (issue's recovery clause) — do not invent fields.
- PR conventions:
Refs #63(neverCloses #63),Closes #74only if all criteria met. Diff >500 lines →supervisor-reviewlabel + explanation section. - Sequencing: dispatch/merge strictly after [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73 — your report consumes its model-matrix artifact; merge main into your branch and re-run the full gate before your merge decision.
5. Decision log (worker fills in)
- …
- Input: the runner's per-run artifact tree —
- added 5 commits that reference this issue
on Jul 8, 2026
Context
Parent: #63. Methodology:
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md.The benchmark must produce public-facing output that is credible, boringly auditable, and impossible to confuse with hand-picked marketing numbers.
Goal
Add benchmark reporting that converts raw run artifacts into a full report plus a README-safe summary block gated on reproducible data.
Acceptance criteria
python-docs-mcp-server + gpt-4o, instead of tool-only rows.docs/benchmarks/results/<run-id>/REPORT.md.Scope boundaries
In scope:
Out of scope:
Forbidden-territory reminder
Do not modify MCP tool names, parameters, return shapes,
schema.sql,.github/workflows/,pyproject.tomlproject metadata,.planning/POSITIONING.md, the README hero section,LICENSE,SECURITY.md, or existing tests by weakening/deleting assertions.README edits, if any, must be below the install/tooling sections and must not touch the hero.
Validation commands
Run any new reporting tests directly as well.
PR template
Use
Refs #63, notCloses #63.The PR must include:
Recovery
If report generation needs result fields not specified by the runner, stop and comment with the missing fields so the runner issue can be corrected.
Effort estimate
4-6 hours.