Repository navigation
fix(sycl): make the default model run on Intel Arc (cambi/speed twins, fp64-free speed extractors) - #1307
Merged
Conversation
11 of 14 tasks
lusoris
force-pushed
the
fix/sycl-v1-model-crash
branch
from
September 5, 2026 16:05
7bc7587 to
2c16d2d
Compare
19 of 26 tasks
…laims; add the option-table sync invariant Verifier corrections to the agent's ADR-1179, research digest, docs/backends/sycl/overview.md and docs/state.md: the default model on SYCL (Arc A380) is identical at %.6f on the 1080p checkerboard pairs and drifts 2.66e-3 on cambi / 2.0e-3 on vmaf for the 576x324 src01 pair; the Recently-closed row T-SYCL-V1-MODEL-SEGFAULT-2026-09-04 was missing (only the _Updated line existed). core/src/feature/sycl/AGENTS.md gains the per-feature option-table sync invariant (every option the CPU twin declares that a model can set must exist in the SYCL twin's table, or the model's feature name diverges).
lusoris
force-pushed
the
fix/sycl-v1-model-crash
branch
from
September 5, 2026 21:15
2c16d2d to
15c3a07
Compare
lusoris
marked this pull request as ready for review
September 5, 2026 21:15
lusoris
added a commit
that referenced
this pull request
Sep 5, 2026
Document the complete operational procedure for the one-shot tiny-AI model retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246). Records maintainer decisions D1-D6 (2026-09-04): - D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6 - D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock) - D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only - D4: drop and count <216 px geometry refusals; never a second teacher - D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline - D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes - Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate docs/state.md: no bug row (operator procedure documentation only). docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
…195) CLAUDE.md rule 15 says to rebuild the container when "its image predates the last master sync". That is stated in terms of time, and time cannot answer the question. A build run against a checkout that is behind master produces an image newer than every commit in the repository and missing exactly the work it was rebuilt for. That happened today. The container was rebuilt specifically to pick up the GPU default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run against them. The build succeeded, the image was the newest thing on disk, and it contained none of the three: the build context was 28 commits behind origin/master. It was caught only because a test file added by one of those PRs was missing. Had the smoke run instead, it would have reported green numbers for code that was not in the image, and those numbers would have been cited as a retrain gate. ADR-1102's marker answers "did a container build this?". It cannot answer "which code was in that container?", and that second question is the one that was wrong. dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref, source_repo) from a VMAFX_SOURCE_REV build argument supplied by dev/docker-compose.yml. It is written in the LAST stage, deliberately away from the ADR-1102 marker in the first: the first stage is reused by every rebuild, so a marker there would report the revision of whichever build first populated the layer cache -- authoritative and stale, which is worse than absent. scripts/dev/check-container-source.sh answers it in both directions. --pre-build refuses a checkout that is behind the reference and lists the commits under baked-in paths the image would be missing; --image reads the marker out of an existing image and reports current, stale (naming what is missing), or unverifiable. A build that never received the argument records `unknown`, and `unknown` is reported as "cannot verify", not as a pass: an image that cannot say what it holds is not evidence. dev/scripts/container-build.sh makes the correct path the easy one -- check, build with the verified revision, re-verify the result. --allow-behind exists for local experiments and says so loudly. Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions, hermetic apart from one Docker case that skips when no daemon is reachable. The stale-context case reproduces today's shape; ahead-of-master is deliberately allowed, since a feature branch legitimately leads master. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246). Records maintainer decisions D1-D6 (2026-09-04): - D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6 - D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock) - D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only - D4: drop and count <216 px geometry refusals; never a second teacher - D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline - D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes - Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate docs/state.md: no bug row (operator procedure documentation only). docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
… (ADR-1192) Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5. All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter path against a container-built current-master libvmaf, but every backend's pooled score has drifted from testdata/netflix_benchmark_results.json (recorded by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the 576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU merges #1307/#1312/#1324 — shows the same drift, so none of it comes from today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no throughput baseline is recorded, because the run also reproduced two pre-existing GPU defects: - vmaf --threads N aborts on every GPU backend (exit 234, "context could not be synchronized"); without --threads both CUDA and SYCL score correctly and are bit-stable over 10 runs. bench_all.sh hard-codes --threads 1. - The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs on master and 8 of 40 on 5a08030 — inside binomial noise of each other. Harness fixes in the same change: - bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit as "backend likely unavailable", which is how a hard abort passed for a missing device for months. It now captures stderr per row and prints FAIL with the exit code and the real last line. Its flag sets also drop --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend. - benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench host, so the SYCL rows failed outright). Both are environment overrides now, VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern. No golden assertions touched; no snapshot regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
…195) CLAUDE.md rule 15 says to rebuild the container when "its image predates the last master sync". That is stated in terms of time, and time cannot answer the question. A build run against a checkout that is behind master produces an image newer than every commit in the repository and missing exactly the work it was rebuilt for. That happened today. The container was rebuilt specifically to pick up the GPU default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run against them. The build succeeded, the image was the newest thing on disk, and it contained none of the three: the build context was 28 commits behind origin/master. It was caught only because a test file added by one of those PRs was missing. Had the smoke run instead, it would have reported green numbers for code that was not in the image, and those numbers would have been cited as a retrain gate. ADR-1102's marker answers "did a container build this?". It cannot answer "which code was in that container?", and that second question is the one that was wrong. dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref, source_repo) from a VMAFX_SOURCE_REV build argument supplied by dev/docker-compose.yml. It is written in the LAST stage, deliberately away from the ADR-1102 marker in the first: the first stage is reused by every rebuild, so a marker there would report the revision of whichever build first populated the layer cache -- authoritative and stale, which is worse than absent. scripts/dev/check-container-source.sh answers it in both directions. --pre-build refuses a checkout that is behind the reference and lists the commits under baked-in paths the image would be missing; --image reads the marker out of an existing image and reports current, stale (naming what is missing), or unverifiable. A build that never received the argument records `unknown`, and `unknown` is reported as "cannot verify", not as a pass: an image that cannot say what it holds is not evidence. dev/scripts/container-build.sh makes the correct path the easy one -- check, build with the verified revision, re-verify the result. --allow-behind exists for local experiments and says so loudly. Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions, hermetic apart from one Docker case that skips when no daemon is reachable. The stale-context case reproduces today's shape; ahead-of-master is deliberately allowed, since a feature branch legitimately leads master. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
…195) CLAUDE.md rule 15 says to rebuild the container when "its image predates the last master sync". That is stated in terms of time, and time cannot answer the question. A build run against a checkout that is behind master produces an image newer than every commit in the repository and missing exactly the work it was rebuilt for. That happened today. The container was rebuilt specifically to pick up the GPU default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run against them. The build succeeded, the image was the newest thing on disk, and it contained none of the three: the build context was 28 commits behind origin/master. It was caught only because a test file added by one of those PRs was missing. Had the smoke run instead, it would have reported green numbers for code that was not in the image, and those numbers would have been cited as a retrain gate. ADR-1102's marker answers "did a container build this?". It cannot answer "which code was in that container?", and that second question is the one that was wrong. dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref, source_repo) from a VMAFX_SOURCE_REV build argument supplied by dev/docker-compose.yml. It is written in the LAST stage, deliberately away from the ADR-1102 marker in the first: the first stage is reused by every rebuild, so a marker there would report the revision of whichever build first populated the layer cache -- authoritative and stale, which is worse than absent. scripts/dev/check-container-source.sh answers it in both directions. --pre-build refuses a checkout that is behind the reference and lists the commits under baked-in paths the image would be missing; --image reads the marker out of an existing image and reports current, stale (naming what is missing), or unverifiable. A build that never received the argument records `unknown`, and `unknown` is reported as "cannot verify", not as a pass: an image that cannot say what it holds is not evidence. dev/scripts/container-build.sh makes the correct path the easy one -- check, build with the verified revision, re-verify the result. --allow-behind exists for local experiments and says so loudly. Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions, hermetic apart from one Docker case that skips when no daemon is reachable. The stale-context case reproduces today's shape; ahead-of-master is deliberately allowed, since a feature branch legitimately leads master. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
…195) (#1337) * ci(dev): make the container say which source it was built from (ADR-1195) CLAUDE.md rule 15 says to rebuild the container when "its image predates the last master sync". That is stated in terms of time, and time cannot answer the question. A build run against a checkout that is behind master produces an image newer than every commit in the repository and missing exactly the work it was rebuilt for. That happened today. The container was rebuilt specifically to pick up the GPU default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run against them. The build succeeded, the image was the newest thing on disk, and it contained none of the three: the build context was 28 commits behind origin/master. It was caught only because a test file added by one of those PRs was missing. Had the smoke run instead, it would have reported green numbers for code that was not in the image, and those numbers would have been cited as a retrain gate. ADR-1102's marker answers "did a container build this?". It cannot answer "which code was in that container?", and that second question is the one that was wrong. dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref, source_repo) from a VMAFX_SOURCE_REV build argument supplied by dev/docker-compose.yml. It is written in the LAST stage, deliberately away from the ADR-1102 marker in the first: the first stage is reused by every rebuild, so a marker there would report the revision of whichever build first populated the layer cache -- authoritative and stale, which is worse than absent. scripts/dev/check-container-source.sh answers it in both directions. --pre-build refuses a checkout that is behind the reference and lists the commits under baked-in paths the image would be missing; --image reads the marker out of an existing image and reports current, stale (naming what is missing), or unverifiable. A build that never received the argument records `unknown`, and `unknown` is reported as "cannot verify", not as a pass: an image that cannot say what it holds is not evidence. dev/scripts/container-build.sh makes the correct path the easy one -- check, build with the verified revision, re-verify the result. --allow-behind exists for local experiments and says so loudly. Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions, hermetic apart from one Docker case that skips when no daemon is reachable. The stale-context case reproduces today's shape; ahead-of-master is deliberately allowed, since a feature branch legitimately leads master. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(state): cite PR #1337 in the container-staleness row (ADR-0165 touch gate) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(state): drop the duplicate rows a keep-both rebase created Each dropped row restates one origin/master already carries; master is the authoritative record. Verified with scripts/ci/check-state-md-rows.sh. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
… (ADR-1192) Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5. All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter path against a container-built current-master libvmaf, but every backend's pooled score has drifted from testdata/netflix_benchmark_results.json (recorded by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the 576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU merges #1307/#1312/#1324 — shows the same drift, so none of it comes from today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no throughput baseline is recorded, because the run also reproduced two pre-existing GPU defects: - vmaf --threads N aborts on every GPU backend (exit 234, "context could not be synchronized"); without --threads both CUDA and SYCL score correctly and are bit-stable over 10 runs. bench_all.sh hard-codes --threads 1. - The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs on master and 8 of 40 on 5a08030 — inside binomial noise of each other. Harness fixes in the same change: - bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit as "backend likely unavailable", which is how a hard abort passed for a missing device for months. It now captures stderr per row and prints FAIL with the exit code and the real last line. Its flag sets also drop --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend. - benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench host, so the SYCL rows failed outright). Both are environment overrides now, VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern. No golden assertions touched; no snapshot regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
… (ADR-1192) (#1334) * fix(testdata): make the Netflix benchmark harness honest and portable (ADR-1192) Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5. All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter path against a container-built current-master libvmaf, but every backend's pooled score has drifted from testdata/netflix_benchmark_results.json (recorded by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the 576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU merges #1307/#1312/#1324 — shows the same drift, so none of it comes from today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no throughput baseline is recorded, because the run also reproduced two pre-existing GPU defects: - vmaf --threads N aborts on every GPU backend (exit 234, "context could not be synchronized"); without --threads both CUDA and SYCL score correctly and are bit-stable over 10 runs. bench_all.sh hard-codes --threads 1. - The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs on master and 8 of 40 on 5a08030 — inside binomial noise of each other. Harness fixes in the same change: - bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit as "backend likely unavailable", which is how a hard abort passed for a missing device for months. It now captures stderr per row and prints FAIL with the exit code and the real last line. Its flag sets also drop --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend. - benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench host, so the SYCL rows failed outright). Both are environment overrides now, VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern. No golden assertions touched; no snapshot regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(state): drop the duplicate rows a keep-both rebase created Each dropped row restates one origin/master already carries; master is the authoritative record. Verified with scripts/ci/check-state-md-rows.sh. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246). Records maintainer decisions D1-D6 (2026-09-04): - D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6 - D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock) - D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only - D4: drop and count <216 px geometry refusals; never a second teacher - D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline - D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes - Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate docs/state.md: no bug row (operator procedure documentation only). docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246). Records maintainer decisions D1-D6 (2026-09-04): - D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6 - D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock) - D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only - D4: drop and count <216 px geometry refusals; never a second teacher - D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline - D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes - Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate docs/state.md: no bug row (operator procedure documentation only). docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
… (#1313) * docs(ai): operator runbook for the v1.0.16 teacher retrain (epic #1246) Document the complete operational procedure for the one-shot tiny-AI model retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246). Records maintainer decisions D1-D6 (2026-09-04): - D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6 - D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock) - D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only - D4: drop and count <216 px geometry refusals; never a second teacher - D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline - D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes - Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate docs/state.md: no bug row (operator procedure documentation only). docs/rebase-notes.md: no rebase impact (fork-added docs). * docs(rebase-notes): drop the duplicate blank line MD012 flags --------- Co-authored-by: Lusoris <lusoris@pm.me>
7 of 17 tasks
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context unmerged" -- both merged on 2026-09-05/06. Re-ran each gate rather than reasoning about it: G2 PASS zero failing runs on master HEAD. The release-please failure the row was written for is gone; ADR-1171 made the missing release-bot credential warn-not-error on push and the workflow reports success. G3 PASS #1307, #1312 and #1324 all on master; container rebuilt from cd52f26 and the default model verified on ALL FOUR backends -- CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each exiting 0. The runbook only asked for CUDA; the others were checked because a model reaches the GPU twins through the model, not through --feature, so a CUDA-only check would not have covered them. G1 FAIL 12 epics open, listed by number, with the caveat that the epic bodies are snapshots and several of their items have already shipped -- the count overstates the work. G4 FAIL still blocked on #1302, and now says why precisely: master's extract_k150k_features.py has no --vmaf-model flag (grep returns 0), which is what §4.2's teacher_model assertion needs. #1302 has no failing check -- its only red mark is the aggregator's draft guard, and its ADR-0108 validator passes six of six. It needs promotion, not repair. Adds a note that the table is a measurement and each row must be re-run rather than carried forward, since stale-status drift is what it just corrected. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 6, 2026
…tus (#1350) * docs(ai): record measured retrain gate status, not authoring-time status The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context unmerged" -- both merged on 2026-09-05/06. Re-ran each gate rather than reasoning about it: G2 PASS zero failing runs on master HEAD. The release-please failure the row was written for is gone; ADR-1171 made the missing release-bot credential warn-not-error on push and the workflow reports success. G3 PASS #1307, #1312 and #1324 all on master; container rebuilt from cd52f26 and the default model verified on ALL FOUR backends -- CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each exiting 0. The runbook only asked for CUDA; the others were checked because a model reaches the GPU twins through the model, not through --feature, so a CUDA-only check would not have covered them. G1 FAIL 12 epics open, listed by number, with the caveat that the epic bodies are snapshots and several of their items have already shipped -- the count overstates the work. G4 FAIL still blocked on #1302, and now says why precisely: master's extract_k150k_features.py has no --vmaf-model flag (grep returns 0), which is what §4.2's teacher_model assertion needs. #1302 has no failing check -- its only red mark is the aggregator's draft guard, and its ADR-0108 validator passes six of six. It needs promotion, not repair. Adds a note that the table is a measurement and each row must be re-run rather than carried forward, since stale-status drift is what it just corrected. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(ai): fix the K150K scores path in both the smoke and the production run The runbook pointed --scores at .corpus/konvid-150k/scores.csv. That file does not exist. KoNViD-150k splits its scores by part, and the corpus on this workstation holds k150ka_scores.csv and k150kb_scores.csv (plus the matching *_votes.csv and a manifest.csv). The wrong path appeared TWICE: in the section 4 five-clip smoke and in the section 5.1 multi-day production extraction. extract_k150k_features.py validates the file and exits with 'error: scores CSV not found', so each would have aborted on its first line -- the smoke immediately, and the ~105-110 hour K150K run at its very start. Corrected to k150ka_scores.csv, which is also the script's own argparse default and the path in its module docstring example. --clips-dir is deliberately left as clips/: a real directory of 153,841 files and a superset of the script's default k150ka_extracted/ (152,265). Lookup is by video_name, so either resolves. Both paths are now listed in a verify-first command so an operator checks them before committing to the long run. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(ai): correct four K150K command errors found by running the smoke Ran the section 4 five-clip smoke end to end in the container. As written it failed on its first line, then its second, then on every clip. Each fix below comes from an actual run, not from reading. 1. --scores named scores.csv, which does not exist. KoNViD-150k splits its scores by part; the corpus holds k150ka_scores.csv (154,746 rows) and k150kb_scores.csv. k150ka_scores.csv is the script's own default. 2. --cpu-vmaf-bin was missing entirely. It is required, and its default /build/vmaf/core/build-cpu/tools/vmaf does not exist in the container, so the run aborts with 'error: cpu-vmaf-bin not found'. /usr/local/bin/vmaf serves both roles. 3. --clips-dir named clips/, which is unusable from inside the container. Its 153,841 entries are symlinks to HOST absolute paths under /home/kilian/dev/vmaf/.workingdir2/konvid-150k/, which do not resolve in the container mount. Every clip failed with 'ffprobe ... returned non-zero exit status 1' -- which reads like corrupt media and is really a dangling link. The real files are in k150ka_extracted/ (152,265 files, ffprobe reports 960x540), again the script's default. All three appeared in BOTH the section 4 smoke and the section 5.1 multi-day production extraction, so the ~105-110 hour K150K run would have aborted at its very start. With them fixed the pipeline runs clean: ok=5 fail=0 at 1.26 clip/s, status complete, schema k150k-feature-extraction-manifest-v1, 5 parquet rows. Of section 4.2's assertions, schema, status and stats.ok already pass on master. Three do not, and all three come from #1302: teacher_model in the manifest, the teacher_model parquet column, and adm3_mean (grep -c adm3 returns 0 on master's extractor and 3 on #1302's). G4 is blocked on that PR alone -- corpus, binary and pipeline are all verified working. no digest needed: trivial Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
vmaf --backend syclwith the default modelvmaf_v1.0.16_3d0h(ADR-1169) crashed on Intel Arc A380 right afterSYCL: using device, while--model version=vmaf_v0.6.1passed — a 1.0.0 blocker, since only a model reaches the SYCL twins (--feature <name>resolves to the CPU extractor by plain name match). Three fixes in the twins:integer_cambi_sycl.cppnow uses the sharedvmaf_cambi_init_tvi_and_vlt()helper exported fromcambi.cinstead of a private TVI/VLT bisection, sizes the c-values histogram byMAX(num_bins, v_band_size), gains thecambi_high_res_speedup(hrs) option with the CPU's threshold / window / scale-0 decimation semantics (its absence made the feature namecambi_cmxv_17_vlt_0.06instead of the model'scambi_hrs_1080_cmxv_17_vlt_0.06, so prediction failed with-EAGAIN), and serialises the option dictionary before the geometry defaults;speed_chroma_sycl.cpp/speed_temporal_sycl.cppdrop theirdoubleaccumulators andsycl::local_accessor<double>(the Arc A-series has noaspect::fp64; ADR-0220).core/src/meson.buildalso propagates-fp-model=preciseto thex86_avx2/x86_avx512static libs under icx. Verified by the reviewer on the Arc A380 with the branch build: the 576x324 src01 pair and both 1080p checkerboard pairs exit 0 with avmafkey; pooledvmafSYCL vs CPU (--no_sycl --no_cuda) 82.814061 vs 82.816062 (delta 2.0e-3), 45.315104 vs 45.315104, 0.000000 vs 0.000000;integer_adm3andinteger_motion3identical at six decimals on all three pairs;speed_chroma_uvdelta 2e-6. Two parity observations are opened indocs/state.mdrather than hidden: pooledcambiat 576x324 is 0.262341 (SYCL) vs 0.259678 (CPU), delta 2.66e-3 — above the 1e-3 bound this fix was asked to meet — andinteger_motion2on the 3-frame checkerboards is 12.554712 vs 12.000000 (a feature the model does not consume;integer_motion_sycl.cppis untouched). Unverified: the agent's log carried no backtrace, so the per-defect crash attribution is the fixing agent's, and the icxfloat_motiondrift claim behind the meson change was not reproduced without the flag. Netflix golden gate: 271 passed / 12 skipped / 0 failed (CPU-forced, 274 s) on the rebased head SYCL unit tests on Arc A380: test_integer_cambi_sycl 4/4, test_sycl_cambi_parity 2/2, test_sycl_speed_chroma_parity 2/2, test_sycl_speed_temporal_parity 2/2; python/test/sycl_default_model_test.py 1/1Type
fix— SYCL twins crashed / failed prediction under the default model on Intel Arc.Checklist
make format && make lintis green locally (pre-commit on every touched file).test_integer_cambi_sycl4/4,test_sycl_cambi_parity2/2,test_sycl_speed_chroma_parity2/2,test_sycl_speed_temporal_parity2/2 (Arc A380, branch build),python/test/sycl_default_model_test.py1/1; Netflix golden gate 271 passed / 12 skipped / 0 failed (274 s, CPU-forced).docs/backends/sycl/overview.md(fixed twins, model-only twin selection, measured cambi drift),docs/adr/1179-sycl-v1-model-crash-fix.md,docs/research/2026-09-05-sycl-v1-model-crash-digest.md.cambi_sycl,speed_chroma_sycl,speed_temporal_sycltwins changed; SIMD: only the icx-fp-model=preciseflag on the AVX2/AVX-512 static libs (no source change); new C sources: none (newpython/test/sycl_default_model_test.py); breaking change: none; ADR: ADR-1179.Bug-status hygiene (ADR-0165)
docs/state.md—T-SYCL-V1-MODEL-SEGFAULT-2026-09-04added to Recently closed;T-SYCL-CAMBI-PARITY-DRIFT-2026-09-05andT-SYCL-MOTION2-CHECKERBOARD-DRIFT-2026-09-05added to Open bugs.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Deep-dive deliverables (ADR-0108)
docs/research/2026-09-05-sycl-v1-model-crash-digest.mddocs/adr/1179-sycl-v1-model-crash-fix.md§ Alternatives consideredAGENTS.mdinvariant note —core/src/feature/sycl/AGENTS.md§ Per-feature option-table sync invariantchangelog.d/fixed/sycl-v1-model-crash.mddocs/rebase-notes.mdentryReproducer
🤖 Generated with Claude Code