Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
712c77d
port: reconcile with upstream Netflix/vmaf September 2026 (ADM/VIF SI…
lusoris Sep 18, 2026
5ad4972
fix(adm): guard the zero-bit shift rounding in the AVX2 and AVX-512 p…
lusoris Sep 18, 2026
c666cc5
fix(gpu): score integer ADM correctly on frames 17 to 32 pixels wide
lusoris Sep 18, 2026
859f624
fix(ci): make the cuda and hip tidy lanes measure their kernel files
lusoris Sep 18, 2026
01141d0
refactor(sycl): clear the integer ADM file's lint debt without suppre…
lusoris Sep 18, 2026
7a2ae71
docs(adm): record the GPU tiny-frame fix, the HIP test flags and the …
lusoris Sep 18, 2026
f17ec42
test(cuda): clean the ADM parity test and report a missing device as …
lusoris Sep 18, 2026
ab42221
refactor(hip): clear the clang-tidy findings in the integer ADM CM ke…
lusoris Sep 18, 2026
93c0f4a
refactor(hip): clear the clang-tidy findings in the integer ADM host …
lusoris Sep 18, 2026
6b8a2f6
fix(hip): fail integer ADM init when the feature-name dictionary cann…
lusoris Sep 18, 2026
1e4ba3d
refactor(cuda): bring the integer ADM CM kernels to zero clang-tidy f…
lusoris Sep 18, 2026
5873e59
refactor(cuda): bring the integer ADM CUDA host glue to zero clang-ti…
lusoris Sep 18, 2026
38c250e
fix(cuda): release integer ADM device state when init fails, and use …
lusoris Sep 18, 2026
d5add29
perf(cuda): stop allocating two integer ADM buffers no kernel reads
lusoris Sep 18, 2026
f74f0d3
fix(build): rebuild CUDA and HIP kernels when a header they include c…
lusoris Sep 18, 2026
0845350
fix(sycl): wrap the scale-0 integer ADM intermediates to 16 bits like…
lusoris Sep 18, 2026
03ecaaf
test(gpu): compare the integer ADM twins with scalar CPU on full-rang…
lusoris Sep 18, 2026
b185335
fix(sycl): round the integer ADM scale 1-3 filter terms like the CPU
lusoris Sep 18, 2026
16dbf5f
docs(sycl): record the SYCL integer ADM wrap and rounding fixes
lusoris Sep 18, 2026
6835538
fix(ci): find the CUDA and HIP kernels that now carry a depfile
lusoris Sep 18, 2026
fc28ddc
docs(rebase): record the SYCL integer ADM narrowing and rounding inva…
lusoris Sep 18, 2026
9226ea0
docs(adm): note the SYCL twin's full-range and rounding fixes on the …
lusoris Sep 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/lint-and-format.yml
Original file line number Diff line number Diff line change
Expand Up @@ -323,6 +323,12 @@ jobs:
with:
fetch-depth: 0
persist-credentials: false
- name: Self-test the ratchet tooling
# Runs on every PR, not only C ones: the tooling itself lives under
# scripts/ci/, and the GPU lanes run it only on workstations.
run: |
python3 -B -m unittest discover -s scripts/ci/tests -p 'test_tidy_ratchet.py'
python3 -B -m unittest discover -s scripts/ci/tests -p 'test_gen_gpu_compile_commands.py'
- name: Plan CI impact (ADR-1140)
id: impact
env:
Expand Down
101 changes: 101 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11451,6 +11451,17 @@ core internal headers (`framesync.h`, `thread_pool.h`, `picture_pool.h`,
projection).


- **macOS on Apple silicon now computes integer ADM like every other
platform** (ADR-1257). Since August 2026 the Apple AArch64 build routed the
first column of every ADM DWT2 row through a compatibility wrapper that
reproduced a historical three-tap result, to keep a macOS-only expected
score in the Python quality tests. Upstream Netflix/vmaf fixed the same
dropped tap in `cba9343ed`, so the wrapper is removed. macOS scores now match
Linux AArch64 and x86. On the Netflix akiyo fixture, VMAF moves from the
recorded macOS value 88.030322 to about 88.030459, the value Linux AArch64
and x86 produce.


- **docs(adm)**: replace open `//TODO: if we integrate adm_p_norm` marker in
`integer_adm.c` with a formal DEFERRED block citing ADR-0481; documents why
`adm_p_norm` remains hardcoded at 3.0 and what must happen before it can be
Expand Down Expand Up @@ -13069,6 +13080,13 @@ values in divergent-branch kernels. The 25 `.cu` files with nested `if`-divergen
all candidates for silent score corruption under the 13.2 toolchain.


- **The CUDA integer ADM extractor allocates about 50 MB less device memory
at 1080p.** Two scratch buffers sized to the frame (`tmp_accum`,
3 x w x h x 8 bytes, and `tmp_accum_h`) were allocated at init and passed to
kernels that never read them, a leftover from an earlier reduction scheme.
They are gone; scores are unchanged.


- **chore(deps): Bump pinned CUDA version from 13.2.0 to 13.3.0** across Dockerfile, `dev/Containerfile`, and CI workflow files (`build.yml`, `libvmaf-build-matrix.yml`).


Expand Down Expand Up @@ -19540,6 +19558,17 @@ integer reformulation of it. CPU scores are unchanged (the compiled
`integer_adm.c` is byte-identical to before).


- **AVX-512 `integer_adm_scale0` now matches the scalar path on frames 17 to
32 pixels wide.** At those widths one of scale 0's right shifts is by zero
bits, and the AVX2 and AVX-512 paths computed its rounding term by
converting infinity to an integer, which is undefined behaviour. The
AVX-512 build turned it into `0xFFFFFFFF` and scored scale 0 up to 0.01
away from scalar; the AVX2 build happened to produce the correct 0. Both now
use the scalar path's guarded helper, so the scalar, AVX2, AVX-512 and NEON
paths give identical scores at every such width. Scalar and AVX2 scores do
not change, and neither does any frame wider than 32 pixels.


- CUDA & HIP backends: fixed two GPU numerical defects in integer ADM contrast masking kernels (`adm_cm.cu` and `adm_cm.hip`):
(1) Fixed border row selection at `i == 0 && top <= 0` (scale 3 height <= 14 px) by replacing running pointer offsets with explicit absolute indexing `{row_top, row_bot, col_l, col_r}` and evaluating `csf_a` at the row 0 center instead of row 2.
(2) Fixed distributed rounding shift by enforcing row-level accumulation across columns in 64-bit precision before applying `(row_total + add_shift_inner_accum) >> shift_inner_accum` once per row, eliminating warp-distributed (CUDA) and thread-distributed (HIP) rounding bias drift on wide frames (>= 1920 px).
Expand Down Expand Up @@ -19603,6 +19632,19 @@ their defaults in `init_fex_metal` rather than being user-configurable via
table-column-style violations in the extractor overview and DISTS-Sq options tables.


- **`integer_adm_scale3` is now correct and reproducible on frames 17 to 32
pixels high or wide.** For those sizes the fourth ADM DWT level has only two
output samples per row. The mirror-index table then replaced its first entry
with index -1, so scale 3 read one row before its input band and one int32
before its scratch allocation, which AddressSanitizer reports as a heap
overflow. The score depended on whatever those bytes held and could change
between runs on identical input. The table now keeps the symmetric mirror
for every size. The ADM working buffers are also zero-initialised, as
upstream Netflix/vmaf does since `1786bd961`, so a future out-of-region
read yields a reproducible value. Frames of 33 pixels or more in both
dimensions score exactly as before.


- Make integer ADM AVX2/AVX-512 temporary scopes and read-only aliases analyzer-clean while preserving dispatch signatures and numerical expression order.


Expand Down Expand Up @@ -23383,6 +23425,17 @@ is addressed.
the correct key.


- **The CUDA and HIP `integer_adm` twins now score frames 17 to 32 pixels wide
correctly, and every GPU twin refuses frames below 17x17.** At those widths
one of scale 0's right shifts is by zero bits. The CUDA and HIP host code
set its rounding term to 2^31 instead of 0, and their scale-0 kernels read
one column and one row past the band at the right and bottom edges. Scale 0
came out up to 0.2 away from the CPU, and 32x32 frames scored NaN. The CUDA,
HIP and SYCL twins also accepted frames smaller than 17x17, which the CPU
and Metal extractors refuse; they now fail with `-EINVAL` and an error that
names the extractor. Scores for frames wider than 32 pixels do not change.


- **CUDA and SYCL CAMBI scores drifted from the CPU reference on real
content.** Both GPU twins mis-mirrored two host-side stages of
`cambi.c`. (1) The spatial-mask kernel clamped (replicated) out-of-image
Expand Down Expand Up @@ -23472,6 +23525,13 @@ is addressed.
uses `-f docker/Dockerfile.node --target <stage>` directly.


- **Editing a header now rebuilds the CUDA and HIP kernels that include it.**
The nvcc and hipcc build steps did not record header dependencies, so an
incremental build after a header change could link host code against
kernels compiled for an older struct layout, which crashed or, worse,
computed wrong results. Clean builds were never affected.


- **GPU `motion_v2` twins emit `motion3_v2_score` (SYCL / HIP / Metal).** The
SYCL, HIP, and Metal `motion_v2` twins now emit
`VMAF_integer_feature_motion3_v2_score` and accept the
Expand Down Expand Up @@ -27165,6 +27225,19 @@ backtick code spans.
ADR-1066).


- **The SYCL `integer_adm` twin now scores full-range content like the CPU.**
The CPU keeps its scale-0 intermediate values in 16 bits, so very large
values wrap around there, and the SYCL twin kept them in 32 or 64 bits.
On noisy content, such as two frames of independent random noise,
`integer_adm_scale0` came out 2.1e-4 away from the CPU, over the 1e-4
cross-backend tolerance, and up to 4.4e-3 away with larger CSF weights.
The SYCL twin now wraps at the same points. It also rounds two scale 1-3
terms the way the CPU, CUDA and HIP do (the long-standing Netflix#955
quirk). That moves SYCL's scale 1-3 and `integer_adm2` scores on ordinary
content by less than 1e-6, toward the CPU. The remaining differences are
below 3e-7.



- **The SYCL integer-ADM bit-scan proves its own shift bound.** Both
`get_best15_from32` equivalents in `core/src/feature/sycl/integer_adm_sycl.cpp`
Expand Down Expand Up @@ -27732,6 +27805,17 @@ ADR-0513.
which also removes the duplication ADR-1208 exists to prevent.


- **`make tidy-ratchet LANE=cuda`, `LANE=hip` and `LANE=sycl` run again, and
measure the CUDA and HIP kernel files.** `tidy-ratchet.py` handed the lanes'
compiler flags to clang-tidy as clang-tidy options, and resolved the build
directory and the SYCL wrapper relative to each translation unit's
directory, so every GPU lane failed before measuring anything. The `.cu` and
`.hip` files were also missing from `compile_commands.json`, because meson
builds them through custom targets; the new
`scripts/ci/gen-gpu-compile-commands.py` adds them, and the make target runs
it first. The Tidy Ratchet job now tests this tooling on every pull request.


- CI: the required `Tidy Changed` and `Tidy Ratchet` gates failed on every C/C++
pull request. `core/meson.build`'s `b_lto_threads=4` (ADR-1172) renders as
GCC's `-flto=4`, which the tidy lane's gcc build writes into
Expand Down Expand Up @@ -27884,6 +27968,23 @@ legs. (ADR-0603, triggered by Renovate PR #1402)
reports `x86 feature: SHSTK`.


- **Integer ADM SIMD kernels no longer store outside their output bands**
(ports of upstream Netflix/vmaf `03b5562c5` and `ea012e387`).
`adm_decouple_avx2` computed its 8-wide tail bound from column 0 although its
loop starts at the left border. Its last store therefore ran up to seven
samples past the decouple region, and for band widths 32 and 40 past the end
of the row. No consumer reads those samples, so x86 scores are unchanged.
`adm_dwt2_8_neon` had no horizontal tail and stored one sample past every
band row. On the last row that sample overwrote the first element of the
next band in the shared ADM buffer with a value computed from over-read
scratch. On AArch64 this changed `integer_adm_scale0` for frame widths 24,
32 and 40 at heights below 50, differently from run to run. Both loops now
stop at a bound that keeps every store inside the band, and a scalar tail
handles the remainder. Every NEON-dispatched width is bit-exact with the
scalar reference, and scores for larger frames, including every Netflix
golden fixture, are unchanged.


- **Upstream issue harvest (ADR-1166): nine stale Netflix/vmaf reports
verified against the fork and fixed, eight more recorded.** The fork
diverged far enough from upstream that an open report there is neither
Expand Down
7 changes: 7 additions & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -177,13 +177,20 @@ TIDY_RATCHET_EXTRA_cuda := --extra-arg=--cuda-host-only --extra-arg=-nocudalib
TIDY_RATCHET_EXTRA_hip := --extra-arg=-x --extra-arg=hip \
--extra-arg=-D__HIP_PLATFORM_AMD__=1 --extra-arg=-I/opt/rocm/include
TIDY_RATCHET_EXTRA_sycl := --clang-tidy scripts/ci/clang-tidy-sycl.sh
# nvcc, hipcc and icpx compile through meson custom targets, which leave their
# translation units out of compile_commands.json; add them before measuring.
TIDY_RATCHET_PREP_cuda := python3 scripts/ci/gen-gpu-compile-commands.py $(TIDY_RATCHET_BUILD_DIR)
TIDY_RATCHET_PREP_hip := python3 scripts/ci/gen-gpu-compile-commands.py $(TIDY_RATCHET_BUILD_DIR)
TIDY_RATCHET_PREP_sycl := python3 scripts/ci/gen-sycl-compile-commands.py $(TIDY_RATCHET_BUILD_DIR)
tidy-ratchet:
$(call require-tool,clang-tidy,install clang-tools)
$(TIDY_RATCHET_PREP_$(LANE))
python3 scripts/ci/tidy-ratchet.py --lane $(LANE) \
--build-dir $(TIDY_RATCHET_BUILD_DIR) $(TIDY_RATCHET_EXTRA_$(LANE)) $(TIDY_RATCHET_ARGS)

tidy-ratchet-write:
$(call require-tool,clang-tidy,install clang-tools)
$(TIDY_RATCHET_PREP_$(LANE))
python3 scripts/ci/tidy-ratchet.py --lane $(LANE) --write \
--build-dir $(TIDY_RATCHET_BUILD_DIR) $(TIDY_RATCHET_EXTRA_$(LANE)) $(TIDY_RATCHET_ARGS)

Expand Down
9 changes: 9 additions & 0 deletions changelog.d/changed/adm-darwin-neon-dwt2-four-tap.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **macOS on Apple silicon now computes integer ADM like every other
platform** (ADR-1257). Since August 2026 the Apple AArch64 build routed the
first column of every ADM DWT2 row through a compatibility wrapper that
reproduced a historical three-tap result, to keep a macOS-only expected
score in the Python quality tests. Upstream Netflix/vmaf fixed the same
dropped tap in `cba9343ed`, so the wrapper is removed. macOS scores now match
Linux AArch64 and x86. On the Netflix akiyo fixture, VMAF moves from the
recorded macOS value 88.030322 to about 88.030459, the value Linux AArch64
and x86 produce.
5 changes: 5 additions & 0 deletions changelog.d/changed/cuda-adm-unused-accum-buffers.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
- **The CUDA integer ADM extractor allocates about 50 MB less device memory
at 1080p.** Two scratch buffers sized to the frame (`tmp_accum`,
3 x w x h x 8 bytes, and `tmp_accum_h`) were allocated at init and passed to
kernels that never read them, a leftover from an earlier reduction scheme.
They are gone; scores are unchanged.
9 changes: 9 additions & 0 deletions changelog.d/fixed/adm-avx512-tiny-width-rounding.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **AVX-512 `integer_adm_scale0` now matches the scalar path on frames 17 to
32 pixels wide.** At those widths one of scale 0's right shifts is by zero
bits, and the AVX2 and AVX-512 paths computed its rounding term by
converting infinity to an integer, which is undefined behaviour. The
AVX-512 build turned it into `0xFFFFFFFF` and scored scale 0 up to 0.01
away from scalar; the AVX2 build happened to produce the correct 0. Both now
use the scalar path's guarded helper, so the scalar, AVX2, AVX-512 and NEON
paths give identical scores at every such width. Scalar and AVX2 scores do
not change, and neither does any frame wider than 32 pixels.
11 changes: 11 additions & 0 deletions changelog.d/fixed/adm-scale3-tiny-frame-oob-read.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
- **`integer_adm_scale3` is now correct and reproducible on frames 17 to 32
pixels high or wide.** For those sizes the fourth ADM DWT level has only two
output samples per row. The mirror-index table then replaced its first entry
with index -1, so scale 3 read one row before its input band and one int32
before its scratch allocation, which AddressSanitizer reports as a heap
overflow. The score depended on whatever those bytes held and could change
between runs on identical input. The table now keeps the symmetric mirror
for every size. The ADM working buffers are also zero-initialised, as
upstream Netflix/vmaf does since `1786bd961`, so a future out-of-region
read yields a reproducible value. Frames of 33 pixels or more in both
dimensions score exactly as before.
9 changes: 9 additions & 0 deletions changelog.d/fixed/gpu-adm-tiny-frames.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **The CUDA and HIP `integer_adm` twins now score frames 17 to 32 pixels wide
correctly, and every GPU twin refuses frames below 17x17.** At those widths
one of scale 0's right shifts is by zero bits. The CUDA and HIP host code
set its rounding term to 2^31 instead of 0, and their scale-0 kernels read
one column and one row past the band at the right and bottom edges. Scale 0
came out up to 0.2 away from the CPU, and 32x32 frames scored NaN. The CUDA,
HIP and SYCL twins also accepted frames smaller than 17x17, which the CPU
and Metal extractors refuse; they now fail with `-EINVAL` and an error that
names the extractor. Scores for frames wider than 32 pixels do not change.
5 changes: 5 additions & 0 deletions changelog.d/fixed/gpu-kernel-header-deps.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
- **Editing a header now rebuilds the CUDA and HIP kernels that include it.**
The nvcc and hipcc build steps did not record header dependencies, so an
incremental build after a header change could link host code against
kernels compiled for an older struct layout, which crashed or, worse,
computed wrong results. Clean builds were never affected.
11 changes: 11 additions & 0 deletions changelog.d/fixed/sycl-adm-int16.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
- **The SYCL `integer_adm` twin now scores full-range content like the CPU.**
The CPU keeps its scale-0 intermediate values in 16 bits, so very large
values wrap around there, and the SYCL twin kept them in 32 or 64 bits.
On noisy content, such as two frames of independent random noise,
`integer_adm_scale0` came out 2.1e-4 away from the CPU, over the 1e-4
cross-backend tolerance, and up to 4.4e-3 away with larger CSF weights.
The SYCL twin now wraps at the same points. It also rounds two scale 1-3
terms the way the CPU, CUDA and HIP do (the long-standing Netflix#955
quirk). That moves SYCL's scale 1-3 and `integer_adm2` scores on ordinary
content by less than 1e-6, toward the CPU. The remaining differences are
below 3e-7.
9 changes: 9 additions & 0 deletions changelog.d/fixed/tidy-gpu-lanes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **`make tidy-ratchet LANE=cuda`, `LANE=hip` and `LANE=sycl` run again, and
measure the CUDA and HIP kernel files.** `tidy-ratchet.py` handed the lanes'
compiler flags to clang-tidy as clang-tidy options, and resolved the build
directory and the SYCL wrapper relative to each translation unit's
directory, so every GPU lane failed before measuring anything. The `.cu` and
`.hip` files were also missing from `compile_commands.json`, because meson
builds them through custom targets; the new
`scripts/ci/gen-gpu-compile-commands.py` adds them, and the make target runs
it first. The Tidy Ratchet job now tests this tooling on every pull request.
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
- **Integer ADM SIMD kernels no longer store outside their output bands**
(ports of upstream Netflix/vmaf `03b5562c5` and `ea012e387`).
`adm_decouple_avx2` computed its 8-wide tail bound from column 0 although its
loop starts at the left border. Its last store therefore ran up to seven
samples past the decouple region, and for band widths 32 and 40 past the end
of the row. No consumer reads those samples, so x86 scores are unchanged.
`adm_dwt2_8_neon` had no horizontal tail and stored one sample past every
band row. On the last row that sample overwrote the first element of the
next band in the shared ADM buffer with a value computed from over-read
scratch. On AArch64 this changed `integer_adm_scale0` for frame widths 24,
32 and 40 at heights below 50, differently from run to run. Both loops now
stop at a bound that keeps every store inside the band, and a scalar tail
handles the remainder. Every NEON-dispatched width is bit-exact with the
scalar reference, and scores for larger frames, including every Netflix
golden fixture, are unchanged.
12 changes: 12 additions & 0 deletions core/src/feature/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -433,6 +433,18 @@ feature/
[ADR-0155](../../../docs/adr/0155-adm-i4-rounding-deferred-netflix-955.md)
and [rebase-notes 0048](../../../docs/rebase-notes.md).

- **`integer_adm.c` DWT mirror table for tiny extents** (fork-only fix,
[Research-2063](../../../docs/research/2063-upstream-sync-2026-09-adm-vif-simd.md)):
`dwt2_src_indices_1d()` starts its mirrored tail at
`(n_half > 2u) ? n_half - 2u : 1u` and bounds the first loop with
`i + 2 < n_half`. Upstream's `n_half - 2` restarts the tail at 0 when
`n_half == 2` (scale 3 for any frame dimension from 17 to 32), replaces the
`{1, 0, 1, 2}` mirror with `{-1, 0, 1, 2}` and reads index -1 before the
band and before the `tmp_ref` allocation. `init_buffers()` also zeroes
`data_buf` (upstream `1786bd961`), but only as defence in depth: the
zeroing does not make upstream's bound safe. Guarded by
`test_integer_adm_tiny_frames`, which the ASan lane aborts on the old bound.

- **`integer_adm` GPU row-level rounding invariant** (fork-local, ADR-1167):
In integer ADM contrast masking kernels (`cuda/integer_adm/adm_cm.cu` and
`hip/integer_adm/adm_cm.hip`), the inner accumulation rounding shift
Expand Down
Loading
Loading