Repository navigation
Conversation
lusoris
force-pushed
the
fix/cuda-motion-mirror
branch
from
October 2, 2026 05:23
b4d093f to
72f88b5
Compare
mirror() in motion_score.cu returns sup - (out - sup + 1) for an index past the last sample, which is 2 * sup - idx - 1. a44e5e6 changed the same expression in the CPU motion path from height - (i_tap - height + 1) to height - (i_tap - height + 2), which is 2 * size - idx - 2 (mirror() in integer_motion.c), and the CUDA copy was not updated. Both kernels, 8bpc and 16bpc, call it, so the two outermost rows and columns of every frame are filtered with the wrong neighbours. On the Netflix 576x324 pair (48 frames, vmaf_v0.6.1) the largest per-frame difference between --gpumask 4294967295 and --gpumask 0 is 2.62e-3 in integer_motion2 and 2.85e-3 in vmaf, and the pooled VMAF differs by 1.07e-3. With the offset changed from 1 to 2 it is 1.3e-5 and 1.4e-5, pooled 1e-6. The 10-bit version of the pair (3 frames) goes from 1.35e-3 to 3e-6. The delta scales with 1/width, as expected from a border effect. What remains is the rounding of the blur-then-subtract order of the CUDA kernel against the CPU's single convolution of the frame difference; that is a separate change. test_cuda_motion compares integer_motion2 of motion_cuda and of the CPU extractor on four frames of noise at 8 and 10 bits and requires agreement within 1e-3. On unpatched master the 8-bit case differs by 5e-2 and fails; with this change both cases pass with a difference of 2e-5. The test needs a CUDA device. CUDA scores on the Netflix src01 pair change by design (pooled VMAF 76.668905 to 76.66783, the CPU gives 76.667831); the two checkerboard pairs are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
force-pushed
the
fix/cuda-motion-mirror
branch
from
October 2, 2026 18:45
72f88b5 to
b435b6c
Compare
4KVCD
added a commit
to 4KVCD/libvmaf-fast
that referenced
this pull request
Oct 5, 2026
…AF alone and with NEG, VMAF v1; RTX 5090, Intel iGPU, CPU), and every claim checked against the code, git and the runs Problem The landing page of 14a1322 had wrong or unsupported statements: - Its speed table came from older runs: libvmaf's CUDA code only from host memory (46 fps at 4K), no route from GPU memory, VMAF v1 on another film (3840x1608), and app-level numbers from VideoMetricsLab. - "Moving a 4K frame pair through system memory costs the CPU more than the GPU half of VMAF v1 saves": not so (the hybrid from system memory is 1.7x the CPU's speed at 4K with fewer cores busy). - "It is also faster than the CUDA code it ports": only for VMAF and NEG together, or from system memory. For VMAF alone from GPU memory, libvmaf's CUDA code is faster (465 against 382 fps at 4K). - "integer arithmetic only": the shaders use float for estimates the result does not depend on (ADM's angle pre-check, VIF's division), and double with the optional NATIVE_F64. - The AMD fault was described as a read "at a negative index"; it is the table's entries for values below -1, read at their usual place. - The self-test was called the engine's; it is the Python bindings' probe(). - VMAF v1's 71-case matrix was said to be identical "on all three GPUs"; the Radeon 780M ran 48 frames of film, not the matrix. - "libvmaf runs the two feature sets separately": it runs motion once and VIF and ADM twice (fex_ctx_vector.c merges extractors with equal options). - libvmaf's CUDA code "within about 0.001" of the CPU: that was before Netflix#1644; now only motion differs, by at most 0.000029 on the benchmark video. - Also: "fully on the GPU" (libvmaf predicts on the CPU); the five-frame window and moving average belong to the HFR models only; Netflix#1612's leak and Netflix#1652's hang were missing; the GPU half of VMAF v1 "alone at 655 fps" and "2.9 cores busy instead of 7.7" could not be reproduced here (dropped); the whole-film check ran before the AMD fix to one shader. - Of the ten shaders, only common.slang carried the Netflix and NVIDIA copyright notice the page says the engine keeps. - Source comments repeated "no float or double on the GPU" and pointed to fast/README.md for the pull requests, which the root README lists. Change - fast/tests/bench_readme.py (new): the page's tables. Frames decoded into memory first, then each route scores them in a loop: libvmaf on the CPU at 4-24 threads; libvmaf's CUDA code from pinned host pictures and from device pictures filled on the GPU; Vulkan from system memory and from GPU memory (its buffers imported into CUDA); VMAF v1 with each GPU. Speed and the process's CPU time per route; scores compared within each kind (GPU routes of VMAF v0.6.1 against libvmaf's CUDA code). --vmaf-only for VMAF without NEG. - README.md rewritten from the measurements and the checks: - Speed: one table for VMAF and VMAF + NEG at 4K and 1080p (CPU, CUDA and Vulkan from system and GPU memory, Intel iGPU), one for VMAF v1 with cores busy, and what they show, including where CUDA is faster. - How it was done: what is reproduced exactly and how (per-pixel float and double, tables, rounding of partial sums), where float is still used, the Intel and AMD driver faults as found, the self-tests as the bindings' probe(), VMAF v1's work split measured on the benchmark video (ADM3 + motion3 68% at 4K, 63% at 1080p), the route from GPU memory. - The pull request table with complete fixes; Netflix#1477 merged upstream on 2026-10-02; which seven changed since; Netflix#1562's second cause. - How it is checked: the two matrices' actual cases, real video, the whole film, and exactly what was run on the Radeon 780M. - fast/vulkan/shaders/*.slang: the libvmaf notice (Netflix 2016-2023, NVIDIA 2021, BSD+Patent) on the nine that lacked it; motion_v1 and vif_filter name the CPU files they come from. common.slang's note on arithmetic says where float and double are used. - fast/vulkan/vmaf_vulkan.cpp, common.slang, build_libvmaf_cuda.ps1: comments point to README.md for the pull requests. - fast/python/vmaf_fast: vulkan.py and v1.py docstrings no longer say "integer arithmetic only", "several times faster" or native/vmaf_vulkan; v1.py drops app numbers from VideoMetricsLab. - fast/README.md: bench_readme.py under Testing. Effect No change to any library: vmaf_vulkan.dll rebuilt from this tree is byte for byte the release's (SHA-256 47b95eeb...), so v3.2.0-fast.1 stays current. The landing page states only what was measured or read here. Verification (RTX 5090, Core Ultra 9 285K and its Intel GPU) - bench_readme.py, HoneyBee 3840x2160 10-bit against its x265 encode, 48 frames in memory: 480 pairs at 4K, 960 at 1920x1080, with and without --vmaf-only. Repeat runs within about 1%. libvmaf's CUDA code from pinned host pictures with 0, 2, 4 and 8 threads: 58-65 fps at 4K. - compare_vmaf_vulkan.py --matrix (45) and compare_vmaf_v1.py --matrix (71), each with --device 0 and 1: ALL IDENTICAL. diagnose_vmaf_vulkan.py: every sum and buffer is the reference's, on both. - Real video, 48 frames of HoneyBee at 4K 10-bit and 1080p 8-bit, both GPUs: VMAF, NEG and VMAF v1 IDENTICAL. - libvmaf CUDA against CPU, feature by feature on the benchmark video: VIF and ADM (VMAF and NEG) bit-identical; motion2 at most 2.87e-5 (4K) and 4.53e-6 (1080p). - VMAF v1 per feature, libvmaf on one thread, 96 frames: ADM3 55.2 ms, motion3 6.0, CAMBI 12.4, SpEED 16.0 at 4K. - The whole Beekeeper film again (VideoMetricsLab's scripts/verify_vmaf_vulkan_movie.py, the release's DLLs, RTX 5090): 151,919 of 151,919 frames, all 11 features bit-identical to CUDA, VMAF and NEG identical (means 94.927743 and 90.299552). The Intel run of 2026-10-04 (earlier build) is cited as such. - The pull requests' states, authors and head commits read with gh on 2026-10-05; upstream's tree has no GPU code but CUDA. Limits - One PC. The Radeon 780M results are the other PC's reports, not re-run. - Cores busy are compared at different speeds; libvmaf's CPU time per frame grows with its thread count, so they are not a per-frame cost. - libvmaf's CUDA code from device pictures crashed in this benchmark with libvmaf's thread pool on (threads > 0); its table row is with threads = 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The CUDA motion kernels mirror an out-of-range index with
sup - (out - sup + 1), which is2 * sup - idx - 1. The CPU path has used reflect-101,2 * size - idx - 2, since a44e5e6, and the CUDA copy was not updated.integer_motion2and VMAF from the CUDA extractor therefore differ from the CPU wherever the edge samples carry detail (the two checkerboard test pairs are unaffected). This addresses the first root cause of #1562. The second one, the rounding order of blurring each frame before subtracting (a residual of about 1e-5 per frame), is not part of this change.Cause
mirror()inlibvmaf/src/feature/cuda/integer_motion/motion_score.cuis called by bothcalculate_motion_score_kernel_8bpcandcalculate_motion_score_kernel_16bpcfor every filter tap that leaves the picture. Foridx >= supit returnssup - (out - sup + 1). a44e5e6 changed the same expression in the CPU path (libvmaf/src/feature/integer_motion.c) fromheight - (i_tap - height + 1)toheight - (i_tap - height + 2), which is themirror()there today. The two outermost rows and columns of every frame are therefore filtered with the wrong neighbours on the GPU. The effect is a border effect and shrinks with the frame width.Reproducer
Master
8e7a1ac4e, release build with-Denable_cuda=true, RTX 4090, CUDA 13.4. Netflixsrc01pair (48 frames, 8-bit) and itsyuv420p10leversion (3 frames),vmaf_v0.6.1:vmaf -r src01_hrc00_576x324.yuv -d src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 \ --model version=vmaf_v0.6.1 -q --json --gpumask 4294967295 -o cpu.json # CPU vmaf ... --gpumask 0 -o gpu.json # CUDALargest absolute per-frame difference between the two, and the pooled VMAF mean (CPU / CUDA):
integer_motion2mastervmafmasterThe CLI prints six digits. What remains on this branch is the blur-then-subtract rounding of the second root cause.
Fix
The offset in
mirror()changes from+ 1to+ 2, as ininteger_motion.c.Tests
test_cuda_motion(new,libvmaf/test/, needs a CUDA device) runsmotionon the CPU andmotion_cudaon four 96x64 frames of noise at 8 and 10 bits and requiresVMAF_integer_feature_motion2_scoreto agree within 1e-3 on every frame. On unpatched master the 8-bit case fails with a difference of 5.0e-2; on this branch the differences are 1.7e-5 (8-bit) and 1.9e-5 (10-bit).Validation
x86-64 Linux, GCC 16.2.1, nvcc 13.4, RTX 4090, release build on master
8e7a1ac4ewith-Denable_cuda=true.Rebased on master
9e48141b(2026-10-02). On that base, with-Denable_cuda=true -Denable_float=true -Denable_checkasm=true,meson testgives 28 of 29 (the failure istest_cuda_pic_preallocation, which also fails on unpatched9e48141b, 27 of 28 there), and the CPU-Db_sanitize=address,undefined -Db_lto=falsebuild gives 22 of 25 on master and here (test_predictandtest_pic_preallocationabort under LeakSanitizer andcheckasmaborts on a heap-buffer-overflow inadm_dwt2_16(integer_adm.c:2603); all three also abort on unpatched9e48141b). The CLI output of the three Netflix pairs (--gpumask 0,vmaf_v0.6.1) against unpatched9e48141b:src01changes as described above (pooled VMAF 76.668905 to 76.66783, largest per-frame change 2.852e-3), both checkerboard pairs are identical (35.068667 and 7.985899). The sweeps and the other measurements below were taken on8e7a1ac4eand not repeated; the files and x86 code paths they depend on are unchanged since (the one upstream change in between is an arm64-only ADM kernel).meson test: 27 of 28 pass on master, 28 of 29 here (the new test). The failure on both istest_cuda_pic_preallocation, which segfaults in the host-pinned case on unpatched master; this change does not touch it.--gpumask 0,vmaf_v0.6.1) on the three Netflix pairs against master. Thesrc01pair changes by design: pooled VMAF 76.668905 to 76.66783 (the CPU gives 76.667831), largest per-frame change 2.9e-3. Both checkerboard pairs are identical (35.068667 and 7.985899). CPU output cannot change: no CPU file is touched, and no Python golden assertion changes.Overlap: #1582 edits the same function (
if (sup == 1) return 0;for tiny frames);git merge-fileonmotion_score.cumerges the two without conflict, and #1582 itself currently conflicts with master ininteger_motion.c. #1614, #1553, #1613 and #1619 add tests tolibvmaf/test/meson.build, #1583 and #1612 editinteger_motion_cuda.c:git merge-treeagainst each of them reports no conflict.The workflow run on this PR needs a maintainer's approval.