Repository navigation
cuda/vif: reset accumulators before scale 0 - #1614
BardieJoensen wants to merge 1 commit into
Conversation
|
Tested on master The regression test fails without the fix and passes with it. The same race without the test, from the public API: four instances in one process, one thread each, own
So scales 1 to 3, which are launched on the extractor's own stream behind the reset, were never affected, as the description says. The |
… are preallocated and reused, not allocated, page-locked and zeroed for every frame (VMAF from system memory 64 -> 345-370 fps at 4K)
Problem
libvmaf's CUDA code fed host pictures through its own API
(vmaf_cuda_preallocate_pictures with the HOST or HOST_PINNED method, then
vmaf_cuda_fetch_preallocated_picture) scored 4K 10-bit VMAF at about 60 fps
on an RTX 5090, against about 465 fps for the same code fed device
pictures and 310 for Vulkan fed from system memory.
Cause
Nothing was preallocated for the host methods ("//TODO: preallocate host
pics" in vmaf_cuda_fetch_preallocated_picture). Every fetch allocated a new
picture: for HOST_PINNED, vmaf_cuda_picture_alloc_pinned, which page-locks
a new buffer (cuMemHostAlloc, 25 MB at 4K 10-bit, a synchronization point
for the device) and zeroes it; the release freed it again (cuMemFreeHost).
For HOST, vmaf_picture_alloc and a zeroing memset. Timed per 4K pair: the
two fetches 8.95 ms, filling the luma 1.63 ms, vmaf_read_pictures 5.20 ms.
No upstream issue or pull request covers it (searched Netflix/vmaf).
Change
- picture_pool: optional allocate/free callbacks and cookie in
VmafPicturePoolConfig (default: vmaf_picture_alloc and aligned_free), the
buffer type its pictures report, and vmaf_picture_pool_try_fetch, which
returns -EAGAIN instead of waiting when no picture is free. The error
path of vmaf_picture_pool_init now frees the pictures already allocated
(it called vmaf_picture_unref on pictures whose ref it had closed, which
freed nothing).
- vmaf_cuda_preallocate_pictures, HOST and HOST_PINNED: a host picture pool
beside the ring buffer, 2 * n_threads + 6 pictures (a pair in flight, the
two reference pictures kept for PREV_REF extractors, the thread pool's
batches); HOST_PINNED's are page-locked
(vmaf_cuda_picture_pool_alloc_pinned / _free_pinned in picture_cuda.c).
- vmaf_cuda_fetch_preallocated_picture takes a picture from that pool; when
all are in use it allocates one as before, so a caller holding many
pictures never waits. vmaf_close closes the pool before the ring buffer
and the CUDA context.
- vmaf_read_pictures: an upload from page-locked memory continues after
cuMemcpy2DAsync returns, and a pooled picture is reused as soon as it is
released, so for HOST_PINNED pictures it now waits for both uploads (the
device pictures' ready events) before any host picture can be released.
Pageable pictures need no wait: the call returns once the data is
staged.
- test_cuda_pic_preallocation: two tests. The same 12 differing 320x240
frames scored with HOST, HOST_PINNED and freshly allocated pictures must
give identical scores, and HOST/HOST_PINNED pictures must come back
holding an earlier frame (reused, not newly zeroed). 20 HOST_PINNED
pictures taken before any is released must all be handed out (pool of 6,
then fallback) and vmaf_close must return.
Effect (4K 10-bit HoneyBee, frames in memory, RTX 5090, two runs each,
release libvmaf.dll and patched alternating)
- Pinned host pictures: VMAF 64 -> 345-370 fps; VMAF + NEG 60 -> 280-283.
As fast as a program uploading into libvmaf's device pictures through
its own reused page-locked buffer (276 against 278 fps VMAF + NEG at 4K,
705 against 697 at 1080p).
- libvmaf's ordinary CPU picture pool (vmaf_preallocate_pictures, as
VideoMetricsLab's CUDA scorer uses): unchanged, 174-178 / 147-167 fps
before, 175-176 / 159-167 after.
- Scores identical on every frame.
Verification
- meson test (CUDA build, tests and tools on, MSVC): 31 of 32 pass before
and after, plus the two new tests. test_vmaf_cuda_gpumask fails with exit
127 before and after (a shell test that cannot run on Windows), and
test_cuda_vif_order (from Netflix#1614) does not build on Windows (pthread.h
and nanosleep); both unrelated.
- The new reuse test fails with the pool disabled ("HOST pictures were not
reused") and passes with it.
- fast/tests/bench_readme.py: per-frame VMAF + NEG scores of libvmaf's
pinned pictures, the program's upload and GPU-decoded frames identical,
4K and 1080p; compare_vmaf_vulkan.py --matrix (45) and compare_vmaf_v1.py
--matrix (71) on the RTX 5090: ALL IDENTICAL.
Limits
- A pool of 6 pictures (n_threads = 0) holds 6 page-locked pictures for
the context's lifetime: 150 MB at 4K 10-bit, allocated by
vmaf_cuda_preallocate_pictures.
- Pictures from the pool keep their previous contents; callers fill every
plane they use, as with vmaf_fetch_preallocated_picture.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CUDA VIF resets its accumulators on a different stream from the scale-0 kernel. The reset can erase that kernel's results and produce NaN scores.
Move the reset onto the picture stream. The existing event keeps scales 1–3 ordered. Add a regression that delays the reset and checks all four scales across repeated 8/10-bit frames.
The regression failed before the fix and passed afterwards. CUDA initcheck/memcheck and all 21 CPU tests passed. The CUDA suite passed 23/24 tests; the existing pinned-picture test still crashes.