Skip to content

Add the audit, which is every quality claim in executable form - #36

Merged
tamnd merged 1 commit into
mainfrom
audit
Aug 17, 2026
Merged

tamnd merged 1 commit into
mainfrom
audit

Conversation

@tamnd

@tamnd tamnd commented Aug 17, 2026

Copy link
Copy Markdown
Owner

Forty-one checks over the committed corpus, in six groups: structure (S01-S08), placeholders (P01-P08), glossary (G01-G06), language (L01-L08), availability (A01-A05) and hygiene (H01-H06).

Nothing here calls a model, opens a socket or reads a key. Every hard check is computable from the catalogs plus the upstream pin plus the glossary and nothing else, which is what makes it fast enough to run on every push and what makes a green run worth believing without knowing which route happened to answer that day.

Hard checks stop the build, soft ones are reported and do not. The glossary group is soft until M7. L05, L06 and L08 are soft on purpose: each has real exceptions, a heading that is genuinely an instruction or a sentence Vietnamese splits in two, and a hard rule that a correct translation breaks is a rule somebody switches off.

L04

Worth reading before the rest. The check started as "no English sentence copied through", re-deriving the classifier's verdict from the msgid and looking for English function words in whatever came back. A no-op is defined as a string with no run of two letters left once the markup is gone, so a function word cannot survive in one, and the check was incapable of producing a finding. Confirmed against the real content repo: zero hits over all 77 839 distinct msgids, by construction rather than by the corpus being clean.

Re-scoping it to literal blocks was measured and rejected: 162 candidates in the real corpus, every one of them a license text (PSF, BeOpen, zlib, W3C, WIDE, MT19937), all deliberately untranslated.

It now compares what the corpus recorded in its passthrough= provenance comment against what the classifier says today. That is the one signal that actually drifts, and only prose is reported: a string that was a no-op and is now a literal block still reads correctly, a string that is now prose is one the corpus is showing a Vietnamese reader in English.

What moved out of the library

Three things are shared rather than redefined, because a second definition is a definition that can drift from the first.

  • Entry.line, the line a block started on, excluded from equality because where a string sits in a file is not part of what the string is.
  • invariants.SHORT and invariants.edges, public now, because the audit asks the same two questions of a committed catalog.
  • Pin.from_yaml and Pin.read, which read back what as_yaml writes and refuse anything else. S01 exists to catch a pin that has drifted from the corpus, so a reader that shrugged at a malformed file would be checking nothing at exactly the moment it mattered.

What it finds on the real corpus

Run against the content repo as it stands:

check findings what it is
S08 548 header drift, all files
A03 388 memory holds a machine translation the catalog never got
H04 382 file ends in more than one newline, verified with xxd
L07 306 written by gpt-5-6-mini though all three routes ask for gpt-5
S03 228 upstream drift
P07 48 doctest with a translated comment
P04 46 model added a leading space
S02 28 msgid edited here
H03 0 no key-shaped string tracked

The 306 is the one to act on and gets its own issue. All three routes request gpt-5 and 306 committed entries came back from something else.

Tests

174 new tests across seven files, one per group plus the model. Nine CLI tests for exit codes, --only, --skip, --lang, --json, --report and --fail-soft.

make check locally: 1273 passed, 97.17% coverage against an 85% floor, make secrets clean.

Forty-one checks over the committed corpus, in six groups: structure,
placeholders, glossary, language, availability, hygiene. Nothing here
calls a model, opens a socket or reads a key, which is what makes it
fast enough for every push and what makes a green run worth believing
without knowing which route happened to answer that day.

Hard checks stop the build, soft ones are reported and do not. The
glossary group is soft until M7, and L05, L06 and L08 are soft on
purpose: each has real exceptions, a heading that is genuinely an
instruction or a sentence Vietnamese splits in two, and a hard rule that
a correct translation breaks is not a rule.

L04 is the one worth reading. It started as "no English sentence copied
through", re-deriving the classifier's verdict from the msgid and
looking for English function words in whatever came back. A no-op is
defined as a string with no run of two letters left once the markup is
gone, so a function word cannot survive in one, and the check was
incapable of a finding: zero over all 77 839 distinct strings in the
corpus, by construction rather than by the corpus being clean. It now
compares what the corpus recorded as a passthrough against what the
classifier says today, which is the one thing that actually drifts.

Three things moved out of the library to be shared rather than
redefined. Entry carries the line it started on, excluded from equality
because where a string sits in a file is not part of what the string is.
invariants.SHORT and invariants.edges are public now, because the audit
asks the same two questions of a committed catalog and a second floor
would mean a string one rule calls ordinary and the other calls English.
Pin.from_yaml reads back what as_yaml writes, narrowly, because S01
exists to catch a pin that has drifted and a reader that shrugged at a
malformed file would be checking nothing at the moment it mattered.
@tamnd
tamnd merged commit 9e84542 into main Aug 17, 2026
@tamnd
tamnd deleted the audit branch August 17, 2026 21:46
@tamnd tamnd mentioned this pull request Aug 17, 2026
7 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant