Skip to content

docs(ai): record measured retrain gate status, not authoring-time status - #1350

Merged
lusoris merged 3 commits into
masterfrom
docs/retrain-gate-status-1246
Sep 6, 2026
Merged

lusoris merged 3 commits into
masterfrom
docs/retrain-gate-status-1246

Conversation

@lusoris

@lusoris lusoris commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Summary

The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no longer hold — G3 in particular blamed "PR #1307 & fix/cambi-cuda-context unmerged", and both merged on 2026-09-05/06. That is the table the retrain is supposed to be gated on, so a stale FAIL is as costly as a wrong PASS.

Each gate was re-run, not reasoned about:

Gate Was Now Why
G1 FAIL FAIL 12 epics open, now listed by number, with the caveat that the bodies are snapshots and several items have already shipped
G2 FAIL PASS Zero failing runs on master HEAD. The release-please failure it cited is gone — ADR-1171 made the missing release-bot credential warn-not-error on push
G3 FAIL PASS #1307, #1312, #1324 all on master; container rebuilt from cd52f2670, default model verified on all four backends
G4 FAIL FAIL Still #1302 — now with the precise reason
G5 FAIL FAIL Maintainer sign-off, unchanged

G3 was checked more broadly than the runbook asked. §3.2 specifies CUDA only. A model reaches the GPU twins through the model, not through --feature, so a CUDA-only check would not have covered SYCL or HIP at all. All four were run: CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each exiting 0 — evidence and container digest.

G4's blocker is now specific. master's extract_k150k_features.py has no --vmaf-model flag (grep -c returns 0), which is exactly what §4.2's teacher_model assertion needs. #1302 supplies it, and #1302 has no failing check — its only red mark is the aggregator's draft guard (Draft PRs must not satisfy Required Checks Aggregator), and its ADR-0108 validator passes six of six. It needs promotion, not repair.

Type

  • docs — documentation only

Checklist

  • Commits follow Conventional Commits.
  • make format && make lint green locally — pre-commit clean on the touched files.
  • Unit tests pass — no code touched; this is the retrain runbook's gate table.
  • SIMD/GPU code path — none touched. The GPU numbers quoted were measured, not changed.
  • Feature extractor twins — none touched.
  • New C/C++ source — none added.
  • Breaking change — no.
  • ADR — none: this records measurements. No decision is made that another engineer could have made differently.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated — no state delta: gate status is epic bookkeeping, not a tracked defect.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: trivial. Five verification commands re-run and their results recorded.
  • Decision matrix — no alternatives: only-one-way fix. A gate is either passing or not; the table either says so accurately or it does not.
  • AGENTS.md invariant note — no rebase-sensitive invariants: the one thing a future editor must not do (carry a status forward without re-running it) is stated in docs/rebase-notes.md and inline in the table's own note.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/changed/retrain-gate-status-1246.md.
  • Rebase note — docs/retrain-gate-status-1246 — measured retrain gate status (2026-09-06).

Reproducer

# G2 — zero failing runs on master HEAD
gh run list --branch master --limit 20 --json conclusion,name \
  --jq '[.[]|select(.conclusion=="failure")]'

# G4 — the flag §4.2 needs is absent from master
git show origin/master:ai/scripts/extract_k150k_features.py | grep -c 'vmaf.model'   # 0

# ...and #1302 supplies it, with no failing check of its own
gh pr checks 1302 --json name,state --jq '[.[]|select(.state=="FAILURE")]'

Known follow-ups

  • G1 needs an audit, not a count. Twelve open epics overstate the work: the bodies are snapshots and several items have already shipped. Each should be checked against the code before being treated as outstanding — the same drift this PR just corrected in G3.
  • G4 unblocks the moment feat(ai): distil from the ADR-1168 default model with per-row teacher provenance #1302 is promoted, after which §4 can run and the gate can be re-measured.

@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 6, 2026
@lusoris

lusoris commented Sep 6, 2026

Copy link
Copy Markdown
Contributor Author

Extended: this PR now also fixes a path that would have aborted the retrain.

The runbook pointed --scores at .corpus/konvid-150k/scores.csv, which does not exist — KoNViD-150k splits its scores into k150ka_scores.csv / k150kb_scores.csv. The wrong path appeared twice: in the §4 five-clip smoke and in the §5.1 multi-day production extraction. extract_k150k_features.py validates it and exits error: scores CSV not found, so the smoke would have died on its first line and the ~105–110 h K150K run at its very start.

Verified against the workstation corpus:

k150ka_scores.csv   154,746 rows   <- exists, and the script's own default
scores.csv          absent
clips/              153,841 files  <- superset of k150ka_extracted/ (152,265)

Corrected to k150ka_scores.csv. --clips-dir is left as clips/ because lookup is by video_name and either directory resolves; both paths are now in a verify-first command so this is caught before the long run rather than hours into it.

lusoris and others added 3 commits September 6, 2026 14:53
The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no
longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context
unmerged" -- both merged on 2026-09-05/06.

Re-ran each gate rather than reasoning about it:

  G2  PASS  zero failing runs on master HEAD. The release-please failure the
            row was written for is gone; ADR-1171 made the missing release-bot
            credential warn-not-error on push and the workflow reports success.
  G3  PASS  #1307, #1312 and #1324 all on master; container rebuilt from
            cd52f26 and the default model verified on ALL FOUR backends --
            CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each
            exiting 0. The runbook only asked for CUDA; the others were checked
            because a model reaches the GPU twins through the model, not
            through --feature, so a CUDA-only check would not have covered them.
  G1  FAIL  12 epics open, listed by number, with the caveat that the epic
            bodies are snapshots and several of their items have already
            shipped -- the count overstates the work.
  G4  FAIL  still blocked on #1302, and now says why precisely: master's
            extract_k150k_features.py has no --vmaf-model flag (grep returns
            0), which is what §4.2's teacher_model assertion needs. #1302 has
            no failing check -- its only red mark is the aggregator's draft
            guard, and its ADR-0108 validator passes six of six. It needs
            promotion, not repair.

Adds a note that the table is a measurement and each row must be re-run rather
than carried forward, since stale-status drift is what it just corrected.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion run

The runbook pointed --scores at .corpus/konvid-150k/scores.csv. That file does
not exist. KoNViD-150k splits its scores by part, and the corpus on this
workstation holds k150ka_scores.csv and k150kb_scores.csv (plus the matching
*_votes.csv and a manifest.csv).

The wrong path appeared TWICE: in the section 4 five-clip smoke and in the
section 5.1 multi-day production extraction. extract_k150k_features.py
validates the file and exits with 'error: scores CSV not found', so each would
have aborted on its first line -- the smoke immediately, and the ~105-110 hour
K150K run at its very start.

Corrected to k150ka_scores.csv, which is also the script's own argparse default
and the path in its module docstring example.

--clips-dir is deliberately left as clips/: a real directory of 153,841 files
and a superset of the script's default k150ka_extracted/ (152,265). Lookup is by
video_name, so either resolves. Both paths are now listed in a verify-first
command so an operator checks them before committing to the long run.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ran the section 4 five-clip smoke end to end in the container. As written it
failed on its first line, then its second, then on every clip. Each fix below
comes from an actual run, not from reading.

1. --scores named scores.csv, which does not exist. KoNViD-150k splits its
   scores by part; the corpus holds k150ka_scores.csv (154,746 rows) and
   k150kb_scores.csv. k150ka_scores.csv is the script's own default.

2. --cpu-vmaf-bin was missing entirely. It is required, and its default
   /build/vmaf/core/build-cpu/tools/vmaf does not exist in the container, so the
   run aborts with 'error: cpu-vmaf-bin not found'. /usr/local/bin/vmaf serves
   both roles.

3. --clips-dir named clips/, which is unusable from inside the container. Its
   153,841 entries are symlinks to HOST absolute paths under
   /home/kilian/dev/vmaf/.workingdir2/konvid-150k/, which do not resolve in the
   container mount. Every clip failed with 'ffprobe ... returned non-zero exit
   status 1' -- which reads like corrupt media and is really a dangling link.
   The real files are in k150ka_extracted/ (152,265 files, ffprobe reports
   960x540), again the script's default.

All three appeared in BOTH the section 4 smoke and the section 5.1 multi-day
production extraction, so the ~105-110 hour K150K run would have aborted at its
very start.

With them fixed the pipeline runs clean: ok=5 fail=0 at 1.26 clip/s,
status complete, schema k150k-feature-extraction-manifest-v1, 5 parquet rows.

Of section 4.2's assertions, schema, status and stats.ok already pass on master.
Three do not, and all three come from #1302: teacher_model in the manifest, the
teacher_model parquet column, and adm3_mean (grep -c adm3 returns 0 on master's
extractor and 3 on #1302's). G4 is blocked on that PR alone -- corpus, binary
and pipeline are all verified working.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the docs/retrain-gate-status-1246 branch from 43846d6 to e77ad73 Compare September 6, 2026 12:54
@lusoris
lusoris marked this pull request as ready for review September 6, 2026 13:04
@lusoris
lusoris merged commit e91ab82 into master Sep 6, 2026
108 of 109 checks passed
@lusoris
lusoris deleted the docs/retrain-gate-status-1246 branch September 6, 2026 13:26
@lusoris lusoris added the type:docs Documentation updates label Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:docs Documentation updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant