Skip to content

Make the ledger's merge provenance real - #280

Merged
drevendev merged 1 commit into
masterfrom
policy/ledger-provenance
Sep 8, 2026
Merged

drevendev merged 1 commit into
masterfrom
policy/ledger-provenance

Conversation

@drevendev

Copy link
Copy Markdown
Owner

Closes #279

The mechanical half of #278, which stays open for the judgement half.

Goal

Fill the provenance field nobody was responsible for filling, stop rows losing their
evidence to an unquoted comma, and refuse to release a milestone on provenance that does
not exist.

Evidence

Found by the external QA voice in #278 and verified independently.

MERGE_COMMIT blank on 11 of 23 rows, 8 of them IMPLEMENTED. The AUTHOR cannot
write it — the row lands inside its own pull request and the squash commit does not exist
until that pull request merges — so AUTHOR_RUNBOOK.md called it optional, while
release_tag.py, added the day before, consumes it as authoritative provenance.

The cost is on the digest: coverage_digest hashes (REQ_ID, MERGE_COMMIT) pairs, so
with blanks a milestone repaired at new commits hashes identically to the one released
before and the patch tag that exists for exactly that case is never cut.

10 of 23 rows parsed into more than six fields. An unquoted comma in EVIDENCE
splits the row; csv.DictReader files the surplus under None, validate() never
looked, and every consumer read only the text before it. REQ-CORE-005 was rendering
120 characters of a 620-character cell.

Scope — every changed path

  • docs/spec/implementation_status.csv — 11 rows gain their real merge commit; 10 rows
    regain the evidence that was being truncated; written through a real CSV writer.
  • docs/spec/IMPLEMENTATION_STATUS.md — regenerated.
  • scripts/implementation_status.py — validate() rejects a row that parses into more
    than the six declared fields.
  • scripts/backfill_merge_commits.py — new. Asks the forge where each row's merged pull
    request landed. Never rewrites a recorded commit, never touches an unmerged pull
    request, never changes STATUS.
  • scripts/release_tag.py — missing_provenance; a complete milestone with a blank
    commit warns and is not tagged.
  • .github/workflows/release-tag.yml — backfill runs before the tagger; git identity
    hoisted so both steps share it.
  • scripts/tests/test_backfill_merge_commits.py — new, 12 tests.
  • docs/zendev/AUTHOR_RUNBOOK.md — the field is filled after the merge, not optional;
    quote EVIDENCE that holds a comma.

Non-goals

No STATUS is touched. REQ-CORE-006 stays PARTIAL here even though #278 argues
convincingly that merged #239 completed it. Deciding a requirement is satisfied is the
ACCEPTOR's judgement against acceptance criteria, not an operator's while repairing data.

Acceptance criteria

  • No row naming a merged pull request lacks its commit; a test asserts it of this
    repository's own ledger.
  • --check rejects an over-wide row, naming it and why.
  • Evidence survives a parse round-trip.
  • The tagger warns and declines rather than hashing an empty string.
  • Backfill runs before the tagger and cannot loop: a run that changes nothing pushes
    nothing.

Verification

python -m unittest discover -s scripts/tests        →  Ran 426 tests … OK
python scripts/implementation_status.py --check     →  matches 23 ledger rows
python scripts/release_tag.py --dry-run             →  no milestone is newly complete

Before and after on one row: REQ-CORE-005 evidence parsed as 120 characters, now 620;
overflow rows 10 → 0; blank merge commits 11 → 0.

🤖 Generated with Claude Code

Found by the external QA voice in #278, verified independently, and true in a way I did
not expect: the release tagger I added the day before consumes MERGE_COMMIT as
authoritative provenance, while AUTHOR_RUNBOOK calls that field optional. Nothing filled
it, because nothing could — the row lands inside the AUTHOR's own pull request and the
squash commit does not exist until that pull request merges. Eleven of twenty-three rows
were blank, eight of them IMPLEMENTED. Optional and nobody's job are the same thing.

The cost lands on the digest: coverage_digest hashes (REQ_ID, MERGE_COMMIT) pairs, so
with blanks a milestone repaired at new commits hashes identically to the one released
before, and the patch tag that exists for exactly that case is never cut. So
backfill_merge_commits asks the forge where each row's merged pull request landed and
writes it down — no judgement, the mapping is a fact GitHub already holds — and the
tagger now refuses a milestone whose provenance is still blank rather than hashing an
empty string.

The second defect was quieter. Ten rows parsed into more than the six declared fields
because EVIDENCE held an unquoted comma; DictReader files the surplus under the key None,
validate() never looked, and every consumer read only the text before that comma.
REQ-CORE-005 was rendering 120 characters of a 620-character cell. The check now names
such a row and the ledger is rewritten through a real CSV writer, so nothing this repo
writes can create one again.

No STATUS is touched. REQ-CORE-006 stays PARTIAL here even though #278 argues
convincingly that merged #239 completed it: deciding a requirement is satisfied is the
ACCEPTOR's call against acceptance criteria, not an operator's while repairing data.
That half of #278 stays open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@drevendev
drevendev merged commit 55090ab into master Sep 8, 2026
6 checks passed
@drevendev
drevendev deleted the policy/ledger-provenance branch September 8, 2026 00:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The ledger records provenance nobody fills, and loses evidence to an unquoted comma

1 participant