Skip to content

[v0.5.0] benchmark reporting — generate raw report and README-safe summary #74

Description

@ayhammouda

Context

Parent: #63. Methodology: docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md.

The benchmark must produce public-facing output that is credible, boringly auditable, and impossible to confuse with hand-picked marketing numbers.

Goal

Add benchmark reporting that converts raw run artifacts into a full report plus a README-safe summary block gated on reproducible data.

Acceptance criteria

  • README-safe summaries present strict tool + model pairings, for example python-docs-mcp-server + gpt-4o, instead of tool-only rows.
  • Generated tables include an error/timeout rate column so unstable tools cannot hide behind surviving-query correctness.
  • A report generator reads raw benchmark artifacts and writes docs/benchmarks/results/<run-id>/REPORT.md.
  • The report includes methodology link, corpus hash, repo commit, model/client matrix, competitor manifest, correctness by category, token counts after client rewrap, latency median/p95, failures/exclusions, and environment metadata.
  • A README summary template is generated separately and includes only compact tables plus links to the methodology and raw result bundle.
  • The generator refuses to produce a README-ready block if required metadata, corpus hash, or raw result files are missing.
  • Tests cover report generation, missing metadata failure, failed competitor disclosure, and aggregate/per-model separation.
  • README itself is not updated with benchmark claims until real data exists.

Scope boundaries

In scope:

  • Report generator.
  • README-safe summary template generation.
  • Tests for gating and disclosure rules.

Out of scope:

  • Running providers.
  • Scoring answer correctness manually.
  • Editing README with final benchmark results.

Forbidden-territory reminder

Do not modify MCP tool names, parameters, return shapes, schema.sql, .github/workflows/, pyproject.toml project metadata, .planning/POSITIONING.md, the README hero section, LICENSE, SECURITY.md, or existing tests by weakening/deleting assertions.

README edits, if any, must be below the install/tooling sections and must not touch the hero.

Validation commands

uv run ruff check src/ tests/
uv run pyright src/
uv run pytest --tb=short -q
uv run python-docs-mcp-server doctor

Run any new reporting tests directly as well.

PR template

Use Refs #63, not Closes #63.

The PR must include:

  • Example generated report from fixture data.
  • Test output.
  • Confirmation that README claim publication is still gated on real benchmark data.

Recovery

If report generation needs result fields not specified by the runner, stop and comment with the missing fields so the runner issue can be corrected.

Effort estimate

4-6 hours.

Activity

  1. added
    enhancementNew feature or request
    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
    on Jun 8, 2026
  2. added this to the v0.5.0 milestone on Jul 8, 2026
  3. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Agent context — #74 [v0.5.0] benchmark reporting — generate raw report and README-safe summary

    Status: FINAL (verified against merged main @ 1f202b6; PR #75 merged ad18406, PR #85 merged 1f202b6, both 2026-07-08).
    Auditable copy: this content is mirrored as a comment on issue #74 (.planning/ is gitignored — the issue comment is the copy pre-flight verifies and dispatch links).

    1. Roadmap excerpt (why this exists)

    STRATEGIC-ROADMAP-2026-05-29.md §4 (v0.5.0): public benchmark harness with mandatory methodology disclosure. Decision 5.17 (evidence ladder): no comparative/benchmark claim enters README/PyPI/launch copy until reproducible data exists — your generator is the enforcement point: it produces a README-safe template, never a README edit, and refuses to render when required metadata/corpus-hash/raw files are missing.

    2. Code touch-points (verify against merged main before starting)

    3. Test patterns to follow

    tests/benchmarks/test_runner.py: _write_yaml-style helpers build fixtures under tmp_path; assert file paths and parsed JSON content. For reporting: generate a fake run-artifact tree under tmp_path, run the generator, assert REPORT.md content, the refusal path (missing metadata/corpus-hash/raw files → clean validation error), failed-competitor disclosure, and aggregate/per-model separation. New tests MUST live under tests/benchmarks/ (so uv run pytest tests/benchmarks -q exercises them).

    4. Known pitfalls

    • Fixture outputs go to tmp_path or tests/benchmarks/fixtures/ — NEVER commit anything under docs/benchmarks/results/ unless unmistakably labeled SYNTHETIC/FIXTURE DATA in the title and every table. The PR-template's "example generated report from fixture data" belongs in the PR body, not committed to the public docs tree.
    • README is not touched at all in this task — the README-safe block is a generated template file; publication is a maintainer act gated on real data (decision 5.17).
    • Lint blind spot: also run uv run ruff check benchmarks/ and uv run pyright benchmarks/ (see [v0.5.0] ci — lint and type-check the benchmarks/ package (maintainer-only) #84).
    • Merged benchmark tests are §2 regression cover — additive only; stop and comment per §2/§8 if a merged test seems to need weakening.
    • Missing result fields: if the report needs fields the runner doesn't emit, STOP and comment on [v0.5.0] benchmark reporting — generate raw report and README-safe summary #74 naming the missing fields (issue's recovery clause) — do not invent fields.
    • PR conventions: Refs #63 (never Closes #63), Closes #74 only if all criteria met. Diff >500 lines → supervisor-review label + explanation section.
    • Sequencing: dispatch/merge strictly after [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #73 — your report consumes its model-matrix artifact; merge main into your branch and re-run the full gate before your merge decision.

    5. Decision log (worker fills in)

    • …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestpriority:P2Medium priority

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions