Skip to content

cuda: flush temporal extractors once with a thread pool - #1553

Open
BardieJoensen wants to merge 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-fix-threaded-double-flush
Open

BardieJoensen wants to merge 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-fix-threaded-double-flush

Conversation

@BardieJoensen

@BardieJoensen BardieJoensen commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

With CPU threads enabled, motion_cuda is flushed by both the temporal loop and the CUDA loop. The second flush appends the final score again and fails with -EINVAL.

This change skips CUDA extractors in the temporal loop and adds a regression test covering 0/1/4 threads, 1/2/12 frames, 8/10-bit input and mixed CPU/CUDA extraction. The test checks every final score. It failed before the fix and passed afterward.

Related: #1538. This fix preserves CPU threading.

Skip CUDA extractors in the temporal flush loop to avoid duplicate final
scores. Add a regression across thread counts and short clips.
@BardieJoensen
BardieJoensen force-pushed the cuda-fix-threaded-double-flush branch from 64c7845 to 4b2fe89 Compare September 23, 2026 18:13
@BardieJoensen BardieJoensen changed the title cuda: don't flush CUDA feature extractors twice when a thread pool is present cuda: flush temporal extractors once with a thread pool Sep 23, 2026
@lusoris

lusoris commented Oct 1, 2026

Copy link
Copy Markdown

Tested on master 8e7a1ac4e + this PR (merges without conflicts), RTX 4090, CUDA 13.4.92, driver 615.71.09, gcc 16.2.1, release build with -Denable_cuda=true.

The regression test fails without the fix and passes with it. test_cuda_motion_flush built against master (the PR's test file and meson.build hunk only, libvmaf.c untouched): libvmaf ERROR context could not be synchronized, threads=1 bpc=8 frames=2 cuda=1 mixed=0 flush=-22, fail, motion flush failed (possible duplicate final score). With the PR: pass. meson test with the PR: 26 ok, 1 fail (test_cuda_pic_preallocation, SIGSEGV); master has 25 ok and the same failure, the extra pass is the new test.

Through the CLI: on master, vmaf --gpumask 0 --threads 4 on src01 576x324 (48 frames, vmaf_v0.6.1) exits 234 with feature "VMAF_integer_feature_motion2_score" cannot be overwritten at index 47 followed by context could not be synchronized. With this PR --threads 0 and --threads 4 both exit 0, and all 12 metrics of all 48 frames are identical to the master --threads 0 run (max abs difference 0; the CLI prints 6 digits). The thread pool is kept, so CPU extractors still run threaded.

Overlap: the same guard is in the first commit of #1583 (git merge-tree conflicts in libvmaf.c because the comment differs) and in a hunk of #1563 (conflict only in test/meson.build with this PR). #1538 avoids the abort differently, by destroying the thread pool on CUDA import.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants