Repository navigation
Conversation
Running `vmaf --gpumask 0 --threads N` aborts at flush with "context could not be synchronized" (exit 234), preceded by a flood of "feature ... cannot be overwritten at index N" warnings. The read dispatch in vmaf_read_pictures already notes that "multithreading for GPU does not yield performance benefits / disabled for now", but it is not actually disabled: when a thread pool is present the function returns early via threaded_read_pictures_batch, bypassing the HAVE_CUDA device- picture cleanup below it (the cuEventRecord(finished) and the unref of ref_device/dist_device), and feature extractors flush from worker threads that lack the bound CUDA context. Tear the thread pool down in vmaf_cuda_import_state so CUDA runs take the single-threaded path, matching the existing comment's intent. GPU feature threading yields no throughput benefit. Reproduced on master with an RTX 5090 / CUDA 13.2: before, exit 234; after, completes with scores identical to the single-threaded path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
Re-verified against the v3.2.0 release (
The branch is based on |
|
Tested on master The abort reproduces on master and is gone with the PR. The cause on this machine is not the one in the description. The line before the error is a single That is the cost of the approach here. Destroying the pool in
Overlap: it merges textually with #1583, #1553 and #1613, but #1613 and #1583 keep CPU threading on CUDA runs, which this PR turns off. |
Summary
Running
vmaf --gpumask 0 --threads Naborts at flush withcontext could not be synchronized(exit 234), preceded by a flood offeature ... cannot be overwritten at index Nwarnings.Cause
The read dispatch in
vmaf_read_picturescarries the note "multithreading forGPU does not yield performance benefits / disabled for now", but GPU feature
threading is not actually disabled. When a thread pool is present,
vmaf_read_picturesreturns early viathreaded_read_pictures_batch, which:HAVE_CUDAdevice-picture cleanup below it — thecuEventRecord(finished)and thevmaf_picture_unrefofref_device/dist_device; andcontext.
The CUDA context is therefore left unsynchronized and
flush_contextfails.Fix
Tear the thread pool down in
vmaf_cuda_import_stateso CUDA runs take thesingle-threaded path, matching the existing comment's intent. GPU feature
threading yields no throughput benefit (the GPU pipeline, not host
feature-extraction parallelism, is the bound).
Validation
RTX 5090, CUDA 13.2,
vmaf_v0.6.1, 1080p y4m:--gpumask 0(single-thread)--gpumask 0 --threads 4context could not be synchronized)Re-verified on the v3.2.0 release (master
ac9467ff)The bug still reproduces on current
master, and this fix still appliescleanly —
vmaf_cuda_import_stateis unchanged since this PR was opened, so thepatch cherry-picks without conflict. Built pure upstream
masterwith CUDA(
-Denable_cuda=true -Denable_float=true -Denable_nvcc=true) on an RTX 5090 /CUDA 13.3, 854p y4m:
--gpumask 0 --threads 085.194703)--gpumask 0 --threads 485.194703(bit-identical to single-thread)--threads 4(CPU, no CUDA import)