Skip to content

Write the memory onto the catalogs, with precedence and a byte check - #29

Merged
tamnd merged 1 commit into
mainfrom
m5-apply
Aug 17, 2026
Merged

tamnd merged 1 commit into
mainfrom
m5-apply

Conversation

@tamnd

@tamnd tamnd commented Aug 17, 2026

Copy link
Copy Markdown
Owner

Third box of #11. apply is the step that turns the memory into the content repo.

What it does

Given a memory, a set of upstream catalogs and a stamp, it produces the exact bytes of every file. No model call, no clock read. That is what makes --check possible: CI renders the same inputs and compares, so an edit made by hand on the content repo fails the build by name rather than being reverted by the next run and mentioned in a log.

Structure comes from upstream and translations come from the memory. Upstream owns which strings exist, in what order, with which source references and extracted comments. The memory owns what they say in Vietnamese. Keeping those apart is what lets one memory serve several version branches, since the 85 192 strings 3.14 and 3.15 share are the same segments in a different arrangement.

Precedence

Enforced at the moment of writing, not by a check that runs afterwards. An entry already in the target that is translated and not fuzzy is a string a person signed off on, and it is returned whole, with the physical lines it was read from, so it produces zero diff bytes.

Spec 01 §3 says apply refuses to run rather than overwrite a person's work. Refusing the entry is the same guarantee and a workable one. A reviewer unfuzzying a string is the normal case, and refusing the whole run would leave the pipeline failing until the memory caught up with the catalog. What matters is that no machine string ever lands on top of a human one, and none does.

Provenance

Everything written carries #, fuzzy and one comment line:

# pydocvi: model=gpt-5.1 prompt=a3f91c2e glossary=v3 batch=b7e1 run=2026-08-16T14:02Z

Five fields, all from the segment, none from the run doing the writing. They describe when and how the string was translated, not when it was last copied into a file. A field that fell back to the applying run would say something untrue and would report the whole corpus as changed the day after it was written. --check renders each file with the revision date that file already carries, for the same reason. Both of those were bugs I had until I diffed rendered output against disk.

A passthrough gets a shorter line naming what it is instead: # pydocvi: passthrough=version_marker.

Two decisions about untouched entries

An untranslated upstream entry is passed through with its own lines. with_msgstr drops raw, and render_field reproduces 86.31% of this corpus's wrapping, so re-rendering tens of thousands of fields that nobody has translated would reflow them all to say exactly what they already said.

An upstream entry that carries a fuzzy translation is blanked instead. Fuzzy means gettext is not confident the string still matches its source, sync already refuses to take one into the memory as ground truth, and carrying it into this repo would launder it into something that looks like our work.

Writing

Atomic, and only the files that changed. Leaving 548 mtimes alone is what makes git status after a nine-hour run a report of the run rather than a list of every file in the repo.

CLI

pydocvi apply [--branch 3.15] [--file library/os.po] [--dry-run] [--check]

--check exits 1 and names the files that differ. --dry-run prints the counts and writes nothing.

Known consequence

X-Generator is part of the header apply owns, so bumping the tool version makes --check report every file as changed until a re-apply. That is correct, since the generator did change, but it means a release lands with a re-apply in the same PR.

Tests

33 new: 26 in tests/test_apply.py covering precedence, what upstream owns, the header, planning and writing, and --check; 7 in tests/test_cli.py for the four endings of the command plus the usage error. apply.py is at 100% line and branch coverage. Suite is 945 passing, 98.55% overall.

apply takes a memory, a set of upstream catalogs and a stamp, and produces
the exact bytes of every file in the content repo. It calls no model and
reads no clock, which is what makes --check possible: CI renders the same
inputs and compares, so a hand edit on the content repo fails the build by
name instead of being overwritten on the next run and mentioned in a log.

Structure comes from upstream and translations come from the memory.
Upstream owns which strings exist, in what order, with which source
references; the memory owns what they say. Keeping those apart is what lets
one memory serve several version branches, since the 85 192 strings 3.14 and
3.15 share are the same segments in a different arrangement.

Precedence is enforced where it matters, at the moment of writing. An entry
already in the target that is translated and not fuzzy is returned whole,
with the physical lines it was read from, so a reviewed string produces zero
diff bytes and no machine translation can land on top of it. The spec says
apply refuses to run in that case. Refusing the entry is the same guarantee
and a workable one: a reviewer unfuzzying a string is the normal case, and
refusing the run would leave the pipeline failing until the memory caught up
with the catalog.

Everything written carries a fuzzy flag and one provenance line naming the
model, prompt, glossary, batch and run. Every field comes from the segment
and none from the applying run. A field that fell back to the current run
would say something untrue and would make --check report the whole corpus as
changed the day after it was written, which is a check that fails so
reliably it stops being read. --check renders each file with the revision
date that file already carries, for the same reason.

An untranslated upstream entry keeps its own lines rather than being
re-rendered, because render_field reproduces 86.31% of this corpus's
wrapping and rewriting tens of thousands of untouched fields to say what
they already said is a diff nobody can read. An upstream entry that is
fuzzy is blanked instead: gettext is not confident the string still matches
its source, sync will not take it as ground truth, and carrying it here
would launder it into something that looks like our work.

Writes are atomic and only touch the files that changed, so git status after
a nine-hour run is a report of the run rather than a list of every file in
the repo.
@tamnd
tamnd merged commit 99fcb51 into main Aug 17, 2026
7 checks passed
@tamnd
tamnd deleted the m5-apply branch August 17, 2026 11:24
@tamnd tamnd mentioned this pull request Aug 17, 2026
7 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant