Repository navigation
docs(ai): record measured retrain gate status, not authoring-time status - #1350
Conversation
|
Extended: this PR now also fixes a path that would have aborted the retrain. The runbook pointed Verified against the workstation corpus: Corrected to |
The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context unmerged" -- both merged on 2026-09-05/06. Re-ran each gate rather than reasoning about it: G2 PASS zero failing runs on master HEAD. The release-please failure the row was written for is gone; ADR-1171 made the missing release-bot credential warn-not-error on push and the workflow reports success. G3 PASS #1307, #1312 and #1324 all on master; container rebuilt from cd52f26 and the default model verified on ALL FOUR backends -- CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each exiting 0. The runbook only asked for CUDA; the others were checked because a model reaches the GPU twins through the model, not through --feature, so a CUDA-only check would not have covered them. G1 FAIL 12 epics open, listed by number, with the caveat that the epic bodies are snapshots and several of their items have already shipped -- the count overstates the work. G4 FAIL still blocked on #1302, and now says why precisely: master's extract_k150k_features.py has no --vmaf-model flag (grep returns 0), which is what §4.2's teacher_model assertion needs. #1302 has no failing check -- its only red mark is the aggregator's draft guard, and its ADR-0108 validator passes six of six. It needs promotion, not repair. Adds a note that the table is a measurement and each row must be re-run rather than carried forward, since stale-status drift is what it just corrected. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion run The runbook pointed --scores at .corpus/konvid-150k/scores.csv. That file does not exist. KoNViD-150k splits its scores by part, and the corpus on this workstation holds k150ka_scores.csv and k150kb_scores.csv (plus the matching *_votes.csv and a manifest.csv). The wrong path appeared TWICE: in the section 4 five-clip smoke and in the section 5.1 multi-day production extraction. extract_k150k_features.py validates the file and exits with 'error: scores CSV not found', so each would have aborted on its first line -- the smoke immediately, and the ~105-110 hour K150K run at its very start. Corrected to k150ka_scores.csv, which is also the script's own argparse default and the path in its module docstring example. --clips-dir is deliberately left as clips/: a real directory of 153,841 files and a superset of the script's default k150ka_extracted/ (152,265). Lookup is by video_name, so either resolves. Both paths are now listed in a verify-first command so an operator checks them before committing to the long run. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ran the section 4 five-clip smoke end to end in the container. As written it failed on its first line, then its second, then on every clip. Each fix below comes from an actual run, not from reading. 1. --scores named scores.csv, which does not exist. KoNViD-150k splits its scores by part; the corpus holds k150ka_scores.csv (154,746 rows) and k150kb_scores.csv. k150ka_scores.csv is the script's own default. 2. --cpu-vmaf-bin was missing entirely. It is required, and its default /build/vmaf/core/build-cpu/tools/vmaf does not exist in the container, so the run aborts with 'error: cpu-vmaf-bin not found'. /usr/local/bin/vmaf serves both roles. 3. --clips-dir named clips/, which is unusable from inside the container. Its 153,841 entries are symlinks to HOST absolute paths under /home/kilian/dev/vmaf/.workingdir2/konvid-150k/, which do not resolve in the container mount. Every clip failed with 'ffprobe ... returned non-zero exit status 1' -- which reads like corrupt media and is really a dangling link. The real files are in k150ka_extracted/ (152,265 files, ffprobe reports 960x540), again the script's default. All three appeared in BOTH the section 4 smoke and the section 5.1 multi-day production extraction, so the ~105-110 hour K150K run would have aborted at its very start. With them fixed the pipeline runs clean: ok=5 fail=0 at 1.26 clip/s, status complete, schema k150k-feature-extraction-manifest-v1, 5 parquet rows. Of section 4.2's assertions, schema, status and stats.ok already pass on master. Three do not, and all three come from #1302: teacher_model in the manifest, the teacher_model parquet column, and adm3_mean (grep -c adm3 returns 0 on master's extractor and 3 on #1302's). G4 is blocked on that PR alone -- corpus, binary and pipeline are all verified working. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
43846d6 to
e77ad73
Compare
Summary
The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no longer hold — G3 in particular blamed "PR #1307 & fix/cambi-cuda-context unmerged", and both merged on 2026-09-05/06. That is the table the retrain is supposed to be gated on, so a stale FAIL is as costly as a wrong PASS.
Each gate was re-run, not reasoned about:
masterHEAD. The release-please failure it cited is gone — ADR-1171 made the missing release-bot credential warn-not-error on pushcd52f2670, default model verified on all four backendsG3 was checked more broadly than the runbook asked. §3.2 specifies CUDA only. A model reaches the GPU twins through the model, not through
--feature, so a CUDA-only check would not have covered SYCL or HIP at all. All four were run: CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each exiting 0 — evidence and container digest.G4's blocker is now specific.
master'sextract_k150k_features.pyhas no--vmaf-modelflag (grep -creturns 0), which is exactly what §4.2'steacher_modelassertion needs. #1302 supplies it, and #1302 has no failing check — its only red mark is the aggregator's draft guard (Draft PRs must not satisfy Required Checks Aggregator), and its ADR-0108 validator passes six of six. It needs promotion, not repair.Type
docs— documentation onlyChecklist
make format && make lintgreen locally —pre-commitclean on the touched files.Bug-status hygiene (ADR-0165)
docs/state.mdupdated — no state delta: gate status is epic bookkeeping, not a tracked defect.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score.Deep-dive deliverables (ADR-0108)
AGENTS.mdinvariant note — no rebase-sensitive invariants: the one thing a future editor must not do (carry a status forward without re-running it) is stated indocs/rebase-notes.mdand inline in the table's own note.changelog.d/changed/retrain-gate-status-1246.md.docs/retrain-gate-status-1246 — measured retrain gate status (2026-09-06).Reproducer
Known follow-ups