Skip to content

cuda: use the CPU's reflect-101 edge mirror in the motion kernels - #1644

Open
lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/cuda-motion-mirror
Open

lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/cuda-motion-mirror

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown

The CUDA motion kernels mirror an out-of-range index with sup - (out - sup + 1), which is 2 * sup - idx - 1. The CPU path has used reflect-101, 2 * size - idx - 2, since a44e5e6, and the CUDA copy was not updated. integer_motion2 and VMAF from the CUDA extractor therefore differ from the CPU wherever the edge samples carry detail (the two checkerboard test pairs are unaffected). This addresses the first root cause of #1562. The second one, the rounding order of blurring each frame before subtracting (a residual of about 1e-5 per frame), is not part of this change.

Cause

mirror() in libvmaf/src/feature/cuda/integer_motion/motion_score.cu is called by both calculate_motion_score_kernel_8bpc and calculate_motion_score_kernel_16bpc for every filter tap that leaves the picture. For idx >= sup it returns sup - (out - sup + 1). a44e5e6 changed the same expression in the CPU path (libvmaf/src/feature/integer_motion.c) from height - (i_tap - height + 1) to height - (i_tap - height + 2), which is the mirror() there today. The two outermost rows and columns of every frame are therefore filtered with the wrong neighbours on the GPU. The effect is a border effect and shrinks with the frame width.

Reproducer

Master 8e7a1ac4e, release build with -Denable_cuda=true, RTX 4090, CUDA 13.4. Netflix src01 pair (48 frames, 8-bit) and its yuv420p10le version (3 frames), vmaf_v0.6.1:

vmaf -r src01_hrc00_576x324.yuv -d src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 \
    --model version=vmaf_v0.6.1 -q --json --gpumask 4294967295 -o cpu.json   # CPU
vmaf ... --gpumask 0 -o gpu.json                                              # CUDA

Largest absolute per-frame difference between the two, and the pooled VMAF mean (CPU / CUDA):

input integer_motion2 master this branch vmaf master this branch pooled VMAF master this branch
8-bit, 48 frames 2.62e-3 1.3e-5 2.85e-3 1.4e-5 76.667831 / 76.668905 76.667831 / 76.66783
10-bit, 3 frames 1.35e-3 3e-6 1.51e-3 3e-6 82.564227 / 82.56523 82.564227 / 82.564225

The CLI prints six digits. What remains on this branch is the blur-then-subtract rounding of the second root cause.

Fix

The offset in mirror() changes from + 1 to + 2, as in integer_motion.c.

Tests

test_cuda_motion (new, libvmaf/test/, needs a CUDA device) runs motion on the CPU and motion_cuda on four 96x64 frames of noise at 8 and 10 bits and requires VMAF_integer_feature_motion2_score to agree within 1e-3 on every frame. On unpatched master the 8-bit case fails with a difference of 5.0e-2; on this branch the differences are 1.7e-5 (8-bit) and 1.9e-5 (10-bit).

Validation

x86-64 Linux, GCC 16.2.1, nvcc 13.4, RTX 4090, release build on master 8e7a1ac4e with -Denable_cuda=true.

Rebased on master 9e48141b (2026-10-02). On that base, with -Denable_cuda=true -Denable_float=true -Denable_checkasm=true, meson test gives 28 of 29 (the failure is test_cuda_pic_preallocation, which also fails on unpatched 9e48141b, 27 of 28 there), and the CPU -Db_sanitize=address,undefined -Db_lto=false build gives 22 of 25 on master and here (test_predict and test_pic_preallocation abort under LeakSanitizer and checkasm aborts on a heap-buffer-overflow in adm_dwt2_16 (integer_adm.c:2603); all three also abort on unpatched 9e48141b). The CLI output of the three Netflix pairs (--gpumask 0, vmaf_v0.6.1) against unpatched 9e48141b: src01 changes as described above (pooled VMAF 76.668905 to 76.66783, largest per-frame change 2.852e-3), both checkerboard pairs are identical (35.068667 and 7.985899). The sweeps and the other measurements below were taken on 8e7a1ac4e and not repeated; the files and x86 code paths they depend on are unchanged since (the one upstream change in between is an arm64-only ADM kernel).

  • meson test: 27 of 28 pass on master, 28 of 29 here (the new test). The failure on both is test_cuda_pic_preallocation, which segfaults in the host-pinned case on unpatched master; this change does not touch it.
  • CUDA CLI output (--gpumask 0, vmaf_v0.6.1) on the three Netflix pairs against master. The src01 pair changes by design: pooled VMAF 76.668905 to 76.66783 (the CPU gives 76.667831), largest per-frame change 2.9e-3. Both checkerboard pairs are identical (35.068667 and 7.985899). CPU output cannot change: no CPU file is touched, and no Python golden assertion changes.
  • Not run: an ASan/UBSan build, 12 and 16-bit input.

Overlap: #1582 edits the same function (if (sup == 1) return 0; for tiny frames); git merge-file on motion_score.cu merges the two without conflict, and #1582 itself currently conflicts with master in integer_motion.c. #1614, #1553, #1613 and #1619 add tests to libvmaf/test/meson.build, #1583 and #1612 edit integer_motion_cuda.c: git merge-tree against each of them reports no conflict.

The workflow run on this PR needs a maintainer's approval.

@lusoris
lusoris force-pushed the fix/cuda-motion-mirror branch from b4d093f to 72f88b5 Compare October 2, 2026 05:23
mirror() in motion_score.cu returns sup - (out - sup + 1) for an index past the last sample, which is 2 * sup - idx - 1. a44e5e6 changed the same expression in the CPU motion path from height - (i_tap - height + 1) to height - (i_tap - height + 2), which is 2 * size - idx - 2 (mirror() in integer_motion.c), and the CUDA copy was not updated. Both kernels, 8bpc and 16bpc, call it, so the two outermost rows and columns of every frame are filtered with the wrong neighbours.

On the Netflix 576x324 pair (48 frames, vmaf_v0.6.1) the largest per-frame difference between --gpumask 4294967295 and --gpumask 0 is 2.62e-3 in integer_motion2 and 2.85e-3 in vmaf, and the pooled VMAF differs by 1.07e-3. With the offset changed from 1 to 2 it is 1.3e-5 and 1.4e-5, pooled 1e-6. The 10-bit version of the pair (3 frames) goes from 1.35e-3 to 3e-6. The delta scales with 1/width, as expected from a border effect. What remains is the rounding of the blur-then-subtract order of the CUDA kernel against the CPU's single convolution of the frame difference; that is a separate change.

test_cuda_motion compares integer_motion2 of motion_cuda and of the CPU extractor on four frames of noise at 8 and 10 bits and requires agreement within 1e-3. On unpatched master the 8-bit case differs by 5e-2 and fails; with this change both cases pass with a difference of 2e-5. The test needs a CUDA device.

CUDA scores on the Netflix src01 pair change by design (pooled VMAF 76.668905 to 76.66783, the CPU gives 76.667831); the two checkerboard pairs are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the fix/cuda-motion-mirror branch from 72f88b5 to b435b6c Compare October 2, 2026 18:45
4KVCD added a commit to 4KVCD/libvmaf-fast that referenced this pull request Oct 5, 2026
…AF alone and with NEG, VMAF v1; RTX 5090, Intel iGPU, CPU), and every claim checked against the code, git and the runs

Problem
The landing page of 14a1322 had wrong or unsupported statements:
- Its speed table came from older runs: libvmaf's CUDA code only from host
  memory (46 fps at 4K), no route from GPU memory, VMAF v1 on another film
  (3840x1608), and app-level numbers from VideoMetricsLab.
- "Moving a 4K frame pair through system memory costs the CPU more than the
  GPU half of VMAF v1 saves": not so (the hybrid from system memory is 1.7x
  the CPU's speed at 4K with fewer cores busy).
- "It is also faster than the CUDA code it ports": only for VMAF and NEG
  together, or from system memory. For VMAF alone from GPU memory,
  libvmaf's CUDA code is faster (465 against 382 fps at 4K).
- "integer arithmetic only": the shaders use float for estimates the result
  does not depend on (ADM's angle pre-check, VIF's division), and double
  with the optional NATIVE_F64.
- The AMD fault was described as a read "at a negative index"; it is the
  table's entries for values below -1, read at their usual place.
- The self-test was called the engine's; it is the Python bindings' probe().
- VMAF v1's 71-case matrix was said to be identical "on all three GPUs"; the
  Radeon 780M ran 48 frames of film, not the matrix.
- "libvmaf runs the two feature sets separately": it runs motion once and
  VIF and ADM twice (fex_ctx_vector.c merges extractors with equal options).
- libvmaf's CUDA code "within about 0.001" of the CPU: that was before Netflix#1644;
  now only motion differs, by at most 0.000029 on the benchmark video.
- Also: "fully on the GPU" (libvmaf predicts on the CPU); the five-frame
  window and moving average belong to the HFR models only; Netflix#1612's leak and
  Netflix#1652's hang were missing; the GPU half of VMAF v1 "alone at 655 fps" and
  "2.9 cores busy instead of 7.7" could not be reproduced here (dropped); the
  whole-film check ran before the AMD fix to one shader.
- Of the ten shaders, only common.slang carried the Netflix and NVIDIA
  copyright notice the page says the engine keeps.
- Source comments repeated "no float or double on the GPU" and pointed to
  fast/README.md for the pull requests, which the root README lists.

Change
- fast/tests/bench_readme.py (new): the page's tables. Frames decoded into
  memory first, then each route scores them in a loop: libvmaf on the CPU at
  4-24 threads; libvmaf's CUDA code from pinned host pictures and from
  device pictures filled on the GPU; Vulkan from system memory and from GPU
  memory (its buffers imported into CUDA); VMAF v1 with each GPU. Speed and
  the process's CPU time per route; scores compared within each kind (GPU
  routes of VMAF v0.6.1 against libvmaf's CUDA code). --vmaf-only for VMAF
  without NEG.
- README.md rewritten from the measurements and the checks:
  - Speed: one table for VMAF and VMAF + NEG at 4K and 1080p (CPU, CUDA and
    Vulkan from system and GPU memory, Intel iGPU), one for VMAF v1 with
    cores busy, and what they show, including where CUDA is faster.
  - How it was done: what is reproduced exactly and how (per-pixel float and
    double, tables, rounding of partial sums), where float is still used,
    the Intel and AMD driver faults as found, the self-tests as the
    bindings' probe(), VMAF v1's work split measured on the benchmark video
    (ADM3 + motion3 68% at 4K, 63% at 1080p), the route from GPU memory.
  - The pull request table with complete fixes; Netflix#1477 merged upstream on
    2026-10-02; which seven changed since; Netflix#1562's second cause.
  - How it is checked: the two matrices' actual cases, real video, the whole
    film, and exactly what was run on the Radeon 780M.
- fast/vulkan/shaders/*.slang: the libvmaf notice (Netflix 2016-2023, NVIDIA
  2021, BSD+Patent) on the nine that lacked it; motion_v1 and vif_filter
  name the CPU files they come from. common.slang's note on arithmetic says
  where float and double are used.
- fast/vulkan/vmaf_vulkan.cpp, common.slang, build_libvmaf_cuda.ps1: comments
  point to README.md for the pull requests.
- fast/python/vmaf_fast: vulkan.py and v1.py docstrings no longer say
  "integer arithmetic only", "several times faster" or native/vmaf_vulkan;
  v1.py drops app numbers from VideoMetricsLab.
- fast/README.md: bench_readme.py under Testing.

Effect
No change to any library: vmaf_vulkan.dll rebuilt from this tree is byte
for byte the release's (SHA-256 47b95eeb...), so v3.2.0-fast.1 stays
current. The landing page states only what was measured or read here.

Verification (RTX 5090, Core Ultra 9 285K and its Intel GPU)
- bench_readme.py, HoneyBee 3840x2160 10-bit against its x265 encode, 48
  frames in memory: 480 pairs at 4K, 960 at 1920x1080, with and without
  --vmaf-only. Repeat runs within about 1%. libvmaf's CUDA code from
  pinned host pictures with 0, 2, 4 and 8 threads: 58-65 fps at 4K.
- compare_vmaf_vulkan.py --matrix (45) and compare_vmaf_v1.py --matrix (71),
  each with --device 0 and 1: ALL IDENTICAL. diagnose_vmaf_vulkan.py: every
  sum and buffer is the reference's, on both.
- Real video, 48 frames of HoneyBee at 4K 10-bit and 1080p 8-bit, both GPUs:
  VMAF, NEG and VMAF v1 IDENTICAL.
- libvmaf CUDA against CPU, feature by feature on the benchmark video: VIF
  and ADM (VMAF and NEG) bit-identical; motion2 at most 2.87e-5 (4K) and
  4.53e-6 (1080p).
- VMAF v1 per feature, libvmaf on one thread, 96 frames: ADM3 55.2 ms,
  motion3 6.0, CAMBI 12.4, SpEED 16.0 at 4K.
- The whole Beekeeper film again (VideoMetricsLab's
  scripts/verify_vmaf_vulkan_movie.py, the release's DLLs, RTX 5090):
  151,919 of 151,919 frames, all 11 features bit-identical to CUDA, VMAF
  and NEG identical (means 94.927743 and 90.299552). The Intel run of
  2026-10-04 (earlier build) is cited as such.
- The pull requests' states, authors and head commits read with gh on
  2026-10-05; upstream's tree has no GPU code but CUDA.

Limits
- One PC. The Radeon 780M results are the other PC's reports, not re-run.
- Cores busy are compared at different speeds; libvmaf's CPU time per frame
  grows with its thread count, so they are not a per-frame cost.
- libvmaf's CUDA code from device pictures crashed in this benchmark with
  libvmaf's thread pool on (threads > 0); its table row is with threads = 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant