Skip to content

cuda/vif: reset accumulators before scale 0 - #1614

Open
BardieJoensen wants to merge 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-vif-reset-order
Open

BardieJoensen wants to merge 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-vif-reset-order

Conversation

@BardieJoensen

Copy link
Copy Markdown
Contributor

CUDA VIF resets its accumulators on a different stream from the scale-0 kernel. The reset can erase that kernel's results and produce NaN scores.

Move the reset onto the picture stream. The existing event keeps scales 1–3 ordered. Add a regression that delays the reset and checks all four scales across repeated 8/10-bit frames.

The regression failed before the fix and passed afterwards. CUDA initcheck/memcheck and all 21 CPU tests passed. The CUDA suite passed 23/24 tests; the existing pinned-picture test still crashes.

@lusoris

lusoris commented Oct 1, 2026

Copy link
Copy Markdown

Tested on master 8e7a1ac4e + this PR (merges without conflicts), RTX 4090, CUDA 13.4.92, driver 615.71.09, gcc 16.2.1, release build with -Denable_cuda=true.

The regression test fails without the fix and passes with it. test_cuda_vif_order built against master (the PR's test file and meson.build hunk only): bpc=8 frame=2 scale=0 VIF=-nan expected=1, fail, VIF kernels did not wait for their accumulator reset. With the PR: pass. meson test with the PR: 26 ok, 1 fail (test_cuda_pic_preallocation, SIGSEGV); master has 25 ok and the same failure, the extra pass is the new test.

The same race without the test, from the public API: four instances in one process, one thread each, own VmafCudaState each, 48 frames of src01 576x324, vmaf_v0.6.1, every feature compared with a single-instance run after the flush, 12 runs per build:

build runs with wrong vif_scale0 runs with wrong vif_scale1..3 runs with wrong motion2
master 9 of 12 (1 to 7 of 192 values) 0 12 of 12
this PR 0 of 12 0 12 of 12
this PR + #1583 0 of 12 0 0 of 12

So scales 1 to 3, which are launched on the extractor's own stream behind the reset, were never affected, as the description says. The motion2 column is the separate motion SAD reset in the second commit of #1583, which this PR does not touch (git merge-tree of the two is clean). git merge-tree against #1563 conflicts in test/meson.build only.

4KVCD added a commit to 4KVCD/libvmaf-fast that referenced this pull request Oct 7, 2026
… are preallocated and reused, not allocated, page-locked and zeroed for every frame (VMAF from system memory 64 -> 345-370 fps at 4K)

Problem
libvmaf's CUDA code fed host pictures through its own API
(vmaf_cuda_preallocate_pictures with the HOST or HOST_PINNED method, then
vmaf_cuda_fetch_preallocated_picture) scored 4K 10-bit VMAF at about 60 fps
on an RTX 5090, against about 465 fps for the same code fed device
pictures and 310 for Vulkan fed from system memory.

Cause
Nothing was preallocated for the host methods ("//TODO: preallocate host
pics" in vmaf_cuda_fetch_preallocated_picture). Every fetch allocated a new
picture: for HOST_PINNED, vmaf_cuda_picture_alloc_pinned, which page-locks
a new buffer (cuMemHostAlloc, 25 MB at 4K 10-bit, a synchronization point
for the device) and zeroes it; the release freed it again (cuMemFreeHost).
For HOST, vmaf_picture_alloc and a zeroing memset. Timed per 4K pair: the
two fetches 8.95 ms, filling the luma 1.63 ms, vmaf_read_pictures 5.20 ms.
No upstream issue or pull request covers it (searched Netflix/vmaf).

Change
- picture_pool: optional allocate/free callbacks and cookie in
  VmafPicturePoolConfig (default: vmaf_picture_alloc and aligned_free), the
  buffer type its pictures report, and vmaf_picture_pool_try_fetch, which
  returns -EAGAIN instead of waiting when no picture is free. The error
  path of vmaf_picture_pool_init now frees the pictures already allocated
  (it called vmaf_picture_unref on pictures whose ref it had closed, which
  freed nothing).
- vmaf_cuda_preallocate_pictures, HOST and HOST_PINNED: a host picture pool
  beside the ring buffer, 2 * n_threads + 6 pictures (a pair in flight, the
  two reference pictures kept for PREV_REF extractors, the thread pool's
  batches); HOST_PINNED's are page-locked
  (vmaf_cuda_picture_pool_alloc_pinned / _free_pinned in picture_cuda.c).
- vmaf_cuda_fetch_preallocated_picture takes a picture from that pool; when
  all are in use it allocates one as before, so a caller holding many
  pictures never waits. vmaf_close closes the pool before the ring buffer
  and the CUDA context.
- vmaf_read_pictures: an upload from page-locked memory continues after
  cuMemcpy2DAsync returns, and a pooled picture is reused as soon as it is
  released, so for HOST_PINNED pictures it now waits for both uploads (the
  device pictures' ready events) before any host picture can be released.
  Pageable pictures need no wait: the call returns once the data is
  staged.
- test_cuda_pic_preallocation: two tests. The same 12 differing 320x240
  frames scored with HOST, HOST_PINNED and freshly allocated pictures must
  give identical scores, and HOST/HOST_PINNED pictures must come back
  holding an earlier frame (reused, not newly zeroed). 20 HOST_PINNED
  pictures taken before any is released must all be handed out (pool of 6,
  then fallback) and vmaf_close must return.

Effect (4K 10-bit HoneyBee, frames in memory, RTX 5090, two runs each,
release libvmaf.dll and patched alternating)
- Pinned host pictures: VMAF 64 -> 345-370 fps; VMAF + NEG 60 -> 280-283.
  As fast as a program uploading into libvmaf's device pictures through
  its own reused page-locked buffer (276 against 278 fps VMAF + NEG at 4K,
  705 against 697 at 1080p).
- libvmaf's ordinary CPU picture pool (vmaf_preallocate_pictures, as
  VideoMetricsLab's CUDA scorer uses): unchanged, 174-178 / 147-167 fps
  before, 175-176 / 159-167 after.
- Scores identical on every frame.

Verification
- meson test (CUDA build, tests and tools on, MSVC): 31 of 32 pass before
  and after, plus the two new tests. test_vmaf_cuda_gpumask fails with exit
  127 before and after (a shell test that cannot run on Windows), and
  test_cuda_vif_order (from Netflix#1614) does not build on Windows (pthread.h
  and nanosleep); both unrelated.
- The new reuse test fails with the pool disabled ("HOST pictures were not
  reused") and passes with it.
- fast/tests/bench_readme.py: per-frame VMAF + NEG scores of libvmaf's
  pinned pictures, the program's upload and GPU-decoded frames identical,
  4K and 1080p; compare_vmaf_vulkan.py --matrix (45) and compare_vmaf_v1.py
  --matrix (71) on the RTX 5090: ALL IDENTICAL.

Limits
- A pool of 6 pictures (n_threads = 0) holds 6 page-locked pictures for
  the context's lifetime: 150 MB at 4K 10-bit, allocated by
  vmaf_cuda_preallocate_pictures.
- Pictures from the pool keep their previous contents; callers fill every
  plane they use, as with vmaf_fetch_preallocated_picture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants