Skip to content

fix(sycl): make the default model run on Intel Arc (cambi/speed twins, fp64-free speed extractors) - #1307

Merged
lusoris merged 4 commits into
masterfrom
fix/sycl-v1-model-crash
Sep 5, 2026
Merged

lusoris merged 4 commits into
masterfrom
fix/sycl-v1-model-crash

Conversation

@lusoris

@lusoris lusoris commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

vmaf --backend sycl with the default model vmaf_v1.0.16_3d0h (ADR-1169) crashed on Intel Arc A380 right after SYCL: using device, while --model version=vmaf_v0.6.1 passed — a 1.0.0 blocker, since only a model reaches the SYCL twins (--feature <name> resolves to the CPU extractor by plain name match). Three fixes in the twins: integer_cambi_sycl.cpp now uses the shared vmaf_cambi_init_tvi_and_vlt() helper exported from cambi.c instead of a private TVI/VLT bisection, sizes the c-values histogram by MAX(num_bins, v_band_size), gains the cambi_high_res_speedup (hrs) option with the CPU's threshold / window / scale-0 decimation semantics (its absence made the feature name cambi_cmxv_17_vlt_0.06 instead of the model's cambi_hrs_1080_cmxv_17_vlt_0.06, so prediction failed with -EAGAIN), and serialises the option dictionary before the geometry defaults; speed_chroma_sycl.cpp / speed_temporal_sycl.cpp drop their double accumulators and sycl::local_accessor<double> (the Arc A-series has no aspect::fp64; ADR-0220). core/src/meson.build also propagates -fp-model=precise to the x86_avx2 / x86_avx512 static libs under icx. Verified by the reviewer on the Arc A380 with the branch build: the 576x324 src01 pair and both 1080p checkerboard pairs exit 0 with a vmaf key; pooled vmaf SYCL vs CPU (--no_sycl --no_cuda) 82.814061 vs 82.816062 (delta 2.0e-3), 45.315104 vs 45.315104, 0.000000 vs 0.000000; integer_adm3 and integer_motion3 identical at six decimals on all three pairs; speed_chroma_uv delta 2e-6. Two parity observations are opened in docs/state.md rather than hidden: pooled cambi at 576x324 is 0.262341 (SYCL) vs 0.259678 (CPU), delta 2.66e-3 — above the 1e-3 bound this fix was asked to meet — and integer_motion2 on the 3-frame checkerboards is 12.554712 vs 12.000000 (a feature the model does not consume; integer_motion_sycl.cpp is untouched). Unverified: the agent's log carried no backtrace, so the per-defect crash attribution is the fixing agent's, and the icx float_motion drift claim behind the meson change was not reproduced without the flag. Netflix golden gate: 271 passed / 12 skipped / 0 failed (CPU-forced, 274 s) on the rebased head SYCL unit tests on Arc A380: test_integer_cambi_sycl 4/4, test_sycl_cambi_parity 2/2, test_sycl_speed_chroma_parity 2/2, test_sycl_speed_temporal_parity 2/2; python/test/sycl_default_model_test.py 1/1

Type

  • fix — SYCL twins crashed / failed prediction under the default model on Intel Arc.

Checklist

  • Commits follow Conventional Commits.
  • make format && make lint is green locally (pre-commit on every touched file).
  • Unit tests: test_integer_cambi_sycl 4/4, test_sycl_cambi_parity 2/2, test_sycl_speed_chroma_parity 2/2, test_sycl_speed_temporal_parity 2/2 (Arc A380, branch build), python/test/sycl_default_model_test.py 1/1; Netflix golden gate 271 passed / 12 skipped / 0 failed (274 s, CPU-forced).
  • Docs in the same PR: docs/backends/sycl/overview.md (fixed twins, model-only twin selection, measured cambi drift), docs/adr/1179-sycl-v1-model-crash-fix.md, docs/research/2026-09-05-sycl-v1-model-crash-digest.md.
  • SIMD/GPU, twins, new C sources, breaking change, ADR — GPU: SYCL cambi_sycl, speed_chroma_sycl, speed_temporal_sycl twins changed; SIMD: only the icx -fp-model=precise flag on the AVX2/AVX-512 static libs (no source change); new C sources: none (new python/test/sycl_default_model_test.py); breaking change: none; ADR: ADR-1179.

Bug-status hygiene (ADR-0165)

  • docs/state.md — T-SYCL-V1-MODEL-SEGFAULT-2026-09-04 added to Recently closed; T-SYCL-CAMBI-PARITY-DRIFT-2026-09-05 and T-SYCL-MOTION2-CHECKERBOARD-DRIFT-2026-09-05 added to Open bugs.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/2026-09-05-sycl-v1-model-crash-digest.md
  • Decision matrix — docs/adr/1179-sycl-v1-model-crash-fix.md § Alternatives considered
  • AGENTS.md invariant note — core/src/feature/sycl/AGENTS.md § Per-feature option-table sync invariant
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/fixed/sycl-v1-model-crash.md
  • Rebase note — docs/rebase-notes.md entry

Reproducer

source /opt/intel/oneapi/setvars.sh
Y=python/test/resource/yuv
core/build/tools/vmaf -r $Y/src01_hrc00_576x324.yuv -d $Y/src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 \
  --backend sycl --model version=vmaf_v1.0.16_3d0h --json -o sycl.json
# before: exit 139 right after "SYCL: using device: Intel(R) Arc(TM) A380 Graphics"
# after:  exit 0; pooled_metrics.vmaf.mean = 82.814061 (CPU --no_sycl --no_cuda: 82.816062)
core/build/tools/vmaf -r $Y/checkerboard_1920_1080_10_3_0_0.yuv -d $Y/checkerboard_1920_1080_10_3_1_0.yuv -w 1920 -h 1080 -p 420 -b 8 \
  --backend sycl --model version=vmaf_v1.0.16_3d0h --json -o cb1.json
# after: exit 0; pooled_metrics.vmaf.mean = 45.315104 on both backends
PYTHONPATH=$PWD/python python3 -m pytest python/test/sycl_default_model_test.py -v
# PYTEST_LINE
make test-netflix-golden
# GOLDEN_SHORT

🤖 Generated with Claude Code

…laims; add the option-table sync invariant

Verifier corrections to the agent's ADR-1179, research digest, docs/backends/sycl/overview.md
and docs/state.md: the default model on SYCL (Arc A380) is identical at %.6f on the 1080p
checkerboard pairs and drifts 2.66e-3 on cambi / 2.0e-3 on vmaf for the 576x324 src01 pair;
the Recently-closed row T-SYCL-V1-MODEL-SEGFAULT-2026-09-04 was missing (only the _Updated
line existed). core/src/feature/sycl/AGENTS.md gains the per-feature option-table sync
invariant (every option the CPU twin declares that a model can set must exist in the SYCL
twin's table, or the model's feature name diverges).
@lusoris
lusoris force-pushed the fix/sycl-v1-model-crash branch from 2c16d2d to 15c3a07 Compare September 5, 2026 21:15
@lusoris
lusoris marked this pull request as ready for review September 5, 2026 21:15
@lusoris
lusoris merged commit 2c7dc01 into master Sep 5, 2026
128 of 131 checks passed
@lusoris
lusoris deleted the fix/sycl-v1-model-crash branch September 5, 2026 21:41
lusoris added a commit that referenced this pull request Sep 5, 2026
Document the complete operational procedure for the one-shot tiny-AI model
retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246).

Records maintainer decisions D1-D6 (2026-09-04):
- D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6
- D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock)
- D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only
- D4: drop and count <216 px geometry refusals; never a second teacher
- D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline
- D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes
- Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate

docs/state.md: no bug row (operator procedure documentation only).
docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris added a commit that referenced this pull request Sep 6, 2026
…195)

CLAUDE.md rule 15 says to rebuild the container when "its image predates the
last master sync". That is stated in terms of time, and time cannot answer the
question. A build run against a checkout that is behind master produces an
image newer than every commit in the repository and missing exactly the work
it was rebuilt for.

That happened today. The container was rebuilt specifically to pick up the GPU
default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run
against them. The build succeeded, the image was the newest thing on disk, and
it contained none of the three: the build context was 28 commits behind
origin/master. It was caught only because a test file added by one of those PRs
was missing. Had the smoke run instead, it would have reported green numbers
for code that was not in the image, and those numbers would have been cited as
a retrain gate.

ADR-1102's marker answers "did a container build this?". It cannot answer
"which code was in that container?", and that second question is the one that
was wrong.

dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref,
source_repo) from a VMAFX_SOURCE_REV build argument supplied by
dev/docker-compose.yml. It is written in the LAST stage, deliberately away from
the ADR-1102 marker in the first: the first stage is reused by every rebuild,
so a marker there would report the revision of whichever build first populated
the layer cache -- authoritative and stale, which is worse than absent.

scripts/dev/check-container-source.sh answers it in both directions.
--pre-build refuses a checkout that is behind the reference and lists the
commits under baked-in paths the image would be missing; --image reads the
marker out of an existing image and reports current, stale (naming what is
missing), or unverifiable. A build that never received the argument records
`unknown`, and `unknown` is reported as "cannot verify", not as a pass: an
image that cannot say what it holds is not evidence.

dev/scripts/container-build.sh makes the correct path the easy one -- check,
build with the verified revision, re-verify the result. --allow-behind exists
for local experiments and says so loudly.

Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions,
hermetic apart from one Docker case that skips when no daemon is reachable.
The stale-context case reproduces today's shape; ahead-of-master is
deliberately allowed, since a feature branch legitimately leads master.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model
retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246).

Records maintainer decisions D1-D6 (2026-09-04):
- D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6
- D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock)
- D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only
- D4: drop and count <216 px geometry refusals; never a second teacher
- D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline
- D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes
- Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate

docs/state.md: no bug row (operator procedure documentation only).
docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris added a commit that referenced this pull request Sep 6, 2026
… (ADR-1192)

Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5.
All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter
path against a container-built current-master libvmaf, but every backend's
pooled score has drifted from testdata/netflix_benchmark_results.json (recorded
by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the
576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU
merges #1307/#1312/#1324 — shows the same drift, so none of it comes from
today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no
throughput baseline is recorded, because the run also reproduced two
pre-existing GPU defects:

- vmaf --threads N aborts on every GPU backend (exit 234, "context could not be
  synchronized"); without --threads both CUDA and SYCL score correctly and are
  bit-stable over 10 runs. bench_all.sh hard-codes --threads 1.
- The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs
  on master and 8 of 40 on 5a08030 — inside binomial noise of each other.

Harness fixes in the same change:

- bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit
  as "backend likely unavailable", which is how a hard abort passed for a
  missing device for months. It now captures stderr per row and prints FAIL
  with the exit code and the real last line. Its flag sets also drop
  --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend.
- benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and
  /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench
  host, so the SYCL rows failed outright). Both are environment overrides now,
  VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern.

No golden assertions touched; no snapshot regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
…195)

CLAUDE.md rule 15 says to rebuild the container when "its image predates the
last master sync". That is stated in terms of time, and time cannot answer the
question. A build run against a checkout that is behind master produces an
image newer than every commit in the repository and missing exactly the work
it was rebuilt for.

That happened today. The container was rebuilt specifically to pick up the GPU
default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run
against them. The build succeeded, the image was the newest thing on disk, and
it contained none of the three: the build context was 28 commits behind
origin/master. It was caught only because a test file added by one of those PRs
was missing. Had the smoke run instead, it would have reported green numbers
for code that was not in the image, and those numbers would have been cited as
a retrain gate.

ADR-1102's marker answers "did a container build this?". It cannot answer
"which code was in that container?", and that second question is the one that
was wrong.

dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref,
source_repo) from a VMAFX_SOURCE_REV build argument supplied by
dev/docker-compose.yml. It is written in the LAST stage, deliberately away from
the ADR-1102 marker in the first: the first stage is reused by every rebuild,
so a marker there would report the revision of whichever build first populated
the layer cache -- authoritative and stale, which is worse than absent.

scripts/dev/check-container-source.sh answers it in both directions.
--pre-build refuses a checkout that is behind the reference and lists the
commits under baked-in paths the image would be missing; --image reads the
marker out of an existing image and reports current, stale (naming what is
missing), or unverifiable. A build that never received the argument records
`unknown`, and `unknown` is reported as "cannot verify", not as a pass: an
image that cannot say what it holds is not evidence.

dev/scripts/container-build.sh makes the correct path the easy one -- check,
build with the verified revision, re-verify the result. --allow-behind exists
for local experiments and says so loudly.

Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions,
hermetic apart from one Docker case that skips when no daemon is reachable.
The stale-context case reproduces today's shape; ahead-of-master is
deliberately allowed, since a feature branch legitimately leads master.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
…195)

CLAUDE.md rule 15 says to rebuild the container when "its image predates the
last master sync". That is stated in terms of time, and time cannot answer the
question. A build run against a checkout that is behind master produces an
image newer than every commit in the repository and missing exactly the work
it was rebuilt for.

That happened today. The container was rebuilt specifically to pick up the GPU
default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run
against them. The build succeeded, the image was the newest thing on disk, and
it contained none of the three: the build context was 28 commits behind
origin/master. It was caught only because a test file added by one of those PRs
was missing. Had the smoke run instead, it would have reported green numbers
for code that was not in the image, and those numbers would have been cited as
a retrain gate.

ADR-1102's marker answers "did a container build this?". It cannot answer
"which code was in that container?", and that second question is the one that
was wrong.

dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref,
source_repo) from a VMAFX_SOURCE_REV build argument supplied by
dev/docker-compose.yml. It is written in the LAST stage, deliberately away from
the ADR-1102 marker in the first: the first stage is reused by every rebuild,
so a marker there would report the revision of whichever build first populated
the layer cache -- authoritative and stale, which is worse than absent.

scripts/dev/check-container-source.sh answers it in both directions.
--pre-build refuses a checkout that is behind the reference and lists the
commits under baked-in paths the image would be missing; --image reads the
marker out of an existing image and reports current, stale (naming what is
missing), or unverifiable. A build that never received the argument records
`unknown`, and `unknown` is reported as "cannot verify", not as a pass: an
image that cannot say what it holds is not evidence.

dev/scripts/container-build.sh makes the correct path the easy one -- check,
build with the verified revision, re-verify the result. --allow-behind exists
for local experiments and says so loudly.

Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions,
hermetic apart from one Docker case that skips when no daemon is reachable.
The stale-context case reproduces today's shape; ahead-of-master is
deliberately allowed, since a feature branch legitimately leads master.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
…195) (#1337)

* ci(dev): make the container say which source it was built from (ADR-1195)

CLAUDE.md rule 15 says to rebuild the container when "its image predates the
last master sync". That is stated in terms of time, and time cannot answer the
question. A build run against a checkout that is behind master produces an
image newer than every commit in the repository and missing exactly the work
it was rebuilt for.

That happened today. The container was rebuilt specifically to pick up the GPU
default-model fixes (#1307, #1312, #1324) so the epic #1246 GPU smoke could run
against them. The build succeeded, the image was the newest thing on disk, and
it contained none of the three: the build context was 28 commits behind
origin/master. It was caught only because a test file added by one of those PRs
was missing. Had the smoke run instead, it would have reported green numbers
for code that was not in the image, and those numbers would have been cited as
a retrain gate.

ADR-1102's marker answers "did a container build this?". It cannot answer
"which code was in that container?", and that second question is the one that
was wrong.

dev/Containerfile now records /etc/vmafx-dev-source (source_rev, source_ref,
source_repo) from a VMAFX_SOURCE_REV build argument supplied by
dev/docker-compose.yml. It is written in the LAST stage, deliberately away from
the ADR-1102 marker in the first: the first stage is reused by every rebuild,
so a marker there would report the revision of whichever build first populated
the layer cache -- authoritative and stale, which is worse than absent.

scripts/dev/check-container-source.sh answers it in both directions.
--pre-build refuses a checkout that is behind the reference and lists the
commits under baked-in paths the image would be missing; --image reads the
marker out of an existing image and reports current, stale (naming what is
missing), or unverifiable. A build that never received the argument records
`unknown`, and `unknown` is reported as "cannot verify", not as a pass: an
image that cannot say what it holds is not evidence.

dev/scripts/container-build.sh makes the correct path the easy one -- check,
build with the verified revision, re-verify the result. --allow-behind exists
for local experiments and says so loudly.

Verified by scripts/ci/tests/test-check-container-source.sh: 8 assertions,
hermetic apart from one Docker case that skips when no daemon is reachable.
The stale-context case reproduces today's shape; ahead-of-master is
deliberately allowed, since a feature branch legitimately leads master.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(state): cite PR #1337 in the container-staleness row (ADR-0165 touch gate)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(state): drop the duplicate rows a keep-both rebase created

Each dropped row restates one origin/master already carries; master is the
authoritative record. Verified with scripts/ci/check-state-md-rows.sh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
… (ADR-1192)

Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5.
All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter
path against a container-built current-master libvmaf, but every backend's
pooled score has drifted from testdata/netflix_benchmark_results.json (recorded
by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the
576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU
merges #1307/#1312/#1324 — shows the same drift, so none of it comes from
today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no
throughput baseline is recorded, because the run also reproduced two
pre-existing GPU defects:

- vmaf --threads N aborts on every GPU backend (exit 234, "context could not be
  synchronized"); without --threads both CUDA and SYCL score correctly and are
  bit-stable over 10 runs. bench_all.sh hard-codes --threads 1.
- The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs
  on master and 8 of 40 on 5a08030 — inside binomial noise of each other.

Harness fixes in the same change:

- bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit
  as "backend likely unavailable", which is how a hard abort passed for a
  missing device for months. It now captures stderr per row and prints FAIL
  with the exit code and the real last line. Its flag sets also drop
  --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend.
- benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and
  /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench
  host, so the SYCL rows failed outright). Both are environment overrides now,
  VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern.

No golden assertions touched; no snapshot regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
… (ADR-1192) (#1334)

* fix(testdata): make the Netflix benchmark harness honest and portable (ADR-1192)

Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5.
All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter
path against a container-built current-master libvmaf, but every backend's
pooled score has drifted from testdata/netflix_benchmark_results.json (recorded
by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the
576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU
merges #1307/#1312/#1324 — shows the same drift, so none of it comes from
today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no
throughput baseline is recorded, because the run also reproduced two
pre-existing GPU defects:

- vmaf --threads N aborts on every GPU backend (exit 234, "context could not be
  synchronized"); without --threads both CUDA and SYCL score correctly and are
  bit-stable over 10 runs. bench_all.sh hard-codes --threads 1.
- The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs
  on master and 8 of 40 on 5a08030 — inside binomial noise of each other.

Harness fixes in the same change:

- bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit
  as "backend likely unavailable", which is how a hard abort passed for a
  missing device for months. It now captures stderr per row and prints FAIL
  with the exit code and the real last line. Its flag sets also drop
  --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend.
- benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and
  /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench
  host, so the SYCL rows failed outright). Both are environment overrides now,
  VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern.

No golden assertions touched; no snapshot regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(state): drop the duplicate rows a keep-both rebase created

Each dropped row restates one origin/master already carries; master is the
authoritative record. Verified with scripts/ci/check-state-md-rows.sh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model
retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246).

Records maintainer decisions D1-D6 (2026-09-04):
- D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6
- D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock)
- D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only
- D4: drop and count <216 px geometry refusals; never a second teacher
- D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline
- D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes
- Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate

docs/state.md: no bug row (operator procedure documentation only).
docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris added a commit that referenced this pull request Sep 6, 2026
Document the complete operational procedure for the one-shot tiny-AI model
retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246).

Records maintainer decisions D1-D6 (2026-09-04):
- D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6
- D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock)
- D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only
- D4: drop and count <216 px geometry refusals; never a second teacher
- D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline
- D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes
- Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate

docs/state.md: no bug row (operator procedure documentation only).
docs/rebase-notes.md: no rebase impact (fork-added docs).
lusoris added a commit that referenced this pull request Sep 6, 2026
… (#1313)

* docs(ai): operator runbook for the v1.0.16 teacher retrain (epic #1246)

Document the complete operational procedure for the one-shot tiny-AI model
retraining pass against the vmaf_v1.0.16_3d0h teacher (epic #1246).

Records maintainer decisions D1-D6 (2026-09-04):
- D1: raw union extraction (FULL_FEATURES + adm3); student locked to canonical-6
- D2: all corpora (Netflix, CHUG, BVI-DVC, UGC, K150K; ~130 h total wall-clock)
- D3: single vmaf_v1.0.16_3d0h teacher across all rows; HDR rows train MOS head only
- D4: drop and count <216 px geometry refusals; never a second teacher
- D5: int8 static PTQ or QAT in QDQ format is the shipped format; fp32 is baseline
- D6: CPU + CUDA extraction; SYCL lane excluded pending #1307 and drift fixes
- Ensemble parked under ADR-1105 until real-corpus LOSO passes production gate

docs/state.md: no bug row (operator procedure documentation only).
docs/rebase-notes.md: no rebase impact (fork-added docs).

* docs(rebase-notes): drop the duplicate blank line MD012 flags

---------

Co-authored-by: Lusoris <lusoris@pm.me>
lusoris added a commit that referenced this pull request Sep 6, 2026
The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no
longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context
unmerged" -- both merged on 2026-09-05/06.

Re-ran each gate rather than reasoning about it:

  G2  PASS  zero failing runs on master HEAD. The release-please failure the
            row was written for is gone; ADR-1171 made the missing release-bot
            credential warn-not-error on push and the workflow reports success.
  G3  PASS  #1307, #1312 and #1324 all on master; container rebuilt from
            cd52f26 and the default model verified on ALL FOUR backends --
            CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each
            exiting 0. The runbook only asked for CUDA; the others were checked
            because a model reaches the GPU twins through the model, not
            through --feature, so a CUDA-only check would not have covered them.
  G1  FAIL  12 epics open, listed by number, with the caveat that the epic
            bodies are snapshots and several of their items have already
            shipped -- the count overstates the work.
  G4  FAIL  still blocked on #1302, and now says why precisely: master's
            extract_k150k_features.py has no --vmaf-model flag (grep returns
            0), which is what §4.2's teacher_model assertion needs. #1302 has
            no failing check -- its only red mark is the aggregator's draft
            guard, and its ADR-0108 validator passes six of six. It needs
            promotion, not repair.

Adds a note that the table is a measurement and each row must be re-run rather
than carried forward, since stale-status drift is what it just corrected.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
…tus (#1350)

* docs(ai): record measured retrain gate status, not authoring-time status

The epic #1246 gate table listed G2 and G3 as FAIL against conditions that no
longer hold. G3 in particular blamed "PR #1307 & fix/cambi-cuda-context
unmerged" -- both merged on 2026-09-05/06.

Re-ran each gate rather than reasoning about it:

  G2  PASS  zero failing runs on master HEAD. The release-please failure the
            row was written for is gone; ADR-1171 made the missing release-bot
            credential warn-not-error on push and the workflow reports success.
  G3  PASS  #1307, #1312 and #1324 all on master; container rebuilt from
            cd52f26 and the default model verified on ALL FOUR backends --
            CPU 82.816062, CUDA 82.814062, SYCL 82.814061, HIP 82.816061, each
            exiting 0. The runbook only asked for CUDA; the others were checked
            because a model reaches the GPU twins through the model, not
            through --feature, so a CUDA-only check would not have covered them.
  G1  FAIL  12 epics open, listed by number, with the caveat that the epic
            bodies are snapshots and several of their items have already
            shipped -- the count overstates the work.
  G4  FAIL  still blocked on #1302, and now says why precisely: master's
            extract_k150k_features.py has no --vmaf-model flag (grep returns
            0), which is what §4.2's teacher_model assertion needs. #1302 has
            no failing check -- its only red mark is the aggregator's draft
            guard, and its ADR-0108 validator passes six of six. It needs
            promotion, not repair.

Adds a note that the table is a measurement and each row must be re-run rather
than carried forward, since stale-status drift is what it just corrected.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(ai): fix the K150K scores path in both the smoke and the production run

The runbook pointed --scores at .corpus/konvid-150k/scores.csv. That file does
not exist. KoNViD-150k splits its scores by part, and the corpus on this
workstation holds k150ka_scores.csv and k150kb_scores.csv (plus the matching
*_votes.csv and a manifest.csv).

The wrong path appeared TWICE: in the section 4 five-clip smoke and in the
section 5.1 multi-day production extraction. extract_k150k_features.py
validates the file and exits with 'error: scores CSV not found', so each would
have aborted on its first line -- the smoke immediately, and the ~105-110 hour
K150K run at its very start.

Corrected to k150ka_scores.csv, which is also the script's own argparse default
and the path in its module docstring example.

--clips-dir is deliberately left as clips/: a real directory of 153,841 files
and a superset of the script's default k150ka_extracted/ (152,265). Lookup is by
video_name, so either resolves. Both paths are now listed in a verify-first
command so an operator checks them before committing to the long run.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(ai): correct four K150K command errors found by running the smoke

Ran the section 4 five-clip smoke end to end in the container. As written it
failed on its first line, then its second, then on every clip. Each fix below
comes from an actual run, not from reading.

1. --scores named scores.csv, which does not exist. KoNViD-150k splits its
   scores by part; the corpus holds k150ka_scores.csv (154,746 rows) and
   k150kb_scores.csv. k150ka_scores.csv is the script's own default.

2. --cpu-vmaf-bin was missing entirely. It is required, and its default
   /build/vmaf/core/build-cpu/tools/vmaf does not exist in the container, so the
   run aborts with 'error: cpu-vmaf-bin not found'. /usr/local/bin/vmaf serves
   both roles.

3. --clips-dir named clips/, which is unusable from inside the container. Its
   153,841 entries are symlinks to HOST absolute paths under
   /home/kilian/dev/vmaf/.workingdir2/konvid-150k/, which do not resolve in the
   container mount. Every clip failed with 'ffprobe ... returned non-zero exit
   status 1' -- which reads like corrupt media and is really a dangling link.
   The real files are in k150ka_extracted/ (152,265 files, ffprobe reports
   960x540), again the script's default.

All three appeared in BOTH the section 4 smoke and the section 5.1 multi-day
production extraction, so the ~105-110 hour K150K run would have aborted at its
very start.

With them fixed the pipeline runs clean: ok=5 fail=0 at 1.26 clip/s,
status complete, schema k150k-feature-extraction-manifest-v1, 5 parquet rows.

Of section 4.2's assertions, schema, status and stats.ok already pass on master.
Three do not, and all three come from #1302: teacher_model in the manifest, the
teacher_model parquet column, and adm3_mean (grep -c adm3 returns 0 on master's
extractor and 3 on #1302's). G4 is blocked on that PR alone -- corpus, binary
and pipeline are all verified working.

no digest needed: trivial

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@lusoris lusoris added the type:bug Something isn't working label Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant