Repository navigation
fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host - #2282
Merged
Merged
Conversation
…wn stub headers (#2281) * chore(tidy): measure the MATLAB MEX sources in the cpu lane against own stub headers The ten MEX sources of compat/python-vmaf/matlab/ include mex.h and matrix.h from the MATLAB SDK, which no runner has, so no lane measured them and they sat in the exception list. Self-authored stubs under scripts/ci/lint-stubs/matlab/ declare only what the ten files use; gen-mex-compile-commands.py adds their entries to the lane's compilation database (cpu lane and changed-files job), and nothing builds or links the stubs. The ten exceptions and the workflow exclusion are removed. The first measurement found about 230 findings. The mechanical ones are fixed (braces, isolated declarations, static, const, unused parameters, widening, missing includes, dead stores) together with one defect: ical_std.c passed the data pointer of a matrix to mxDestroyArray(). Thirteen function-size findings stay in tidy-baseline-cpu.json: nothing here runs MATLAB, so a refactor of those functions needs a test first. Closes T-TIDY-MATLAB-MEX-UNMEASURED-2026-09-22. * test(ci): type the generated-module handle of the MEX compile-command test * docs: regenerate the indexes and the citation map after rebasing
…DR numbers (#2275) * docs(adr): remove four unfilled template blocks and file four cited ADR numbers ADR-0643, ADR-0665, ADR-0666 and ADR-0673 began with an unfilled allocator template (title `<fill in title>`) in front of their real text; the block is removed and a status update says so. ADR-0228, ADR-0636, ADR-0867 and ADR-0979 are cited by other ADRs and had no file; each gets a short record written from the commits and ADRs that name it. * docs: regenerate the indexes and the citation map after rebasing
…nd-tripping every plane through the host (#2282) * fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host float_ms_ssim_cuda copied every scored plane of both input pictures to pinned host memory, waited for the copy, converted it with picture_copy() on the host and uploaded the floats again: per plane and frame two plane-sized copies to the host, two uploads and two host waits, for pictures that were already on the device. Level 0 of the pyramids is now picture_copy() on the device (ms_ssim_picture_to_float: the same samples, division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference picture's stream after the distorted picture's ready event, and the private stream waits behind it. The pinned staging buffers are gone. test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts every copy with a host side while they are scored. On the old code, 8-bit 4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18 plane copies with a host side; now 0 and 0, and what goes to the host is exactly the per-window term planes the host adds in raster order. The device-free contract test refuses the host staging. Scores are unchanged: every output equals the CPU's bits, and the exact-twin matrix keeps 36 of 36 float_ms_ssim cells equal.
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
…nd-tripping every plane through the host float_ms_ssim_cuda copied every scored plane of both input pictures to pinned host memory, waited for the copy, converted it with picture_copy() on the host and uploaded the floats again: per plane and frame two plane-sized copies to the host, two uploads and two host waits, for pictures that were already on the device. Level 0 of the pyramids is now picture_copy() on the device (ms_ssim_picture_to_float: the same samples, division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference picture's stream after the distorted picture's ready event, and the private stream waits behind it. The pinned staging buffers are gone. test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts every copy with a host side while they are scored. On the old code, 8-bit 4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18 plane copies with a host side; now 0 and 0, and what goes to the host is exactly the per-window term planes the host adds in raster order. The device-free contract test refuses the host staging. Scores are unchanged: every output equals the CPU's bits, and the exact-twin matrix keeps 36 of 36 float_ms_ssim cells equal. This is #2282 on master, carried by the RC4 WP3 CUDA lane until the stack is restacked onto a master that has it; the restack drops this commit and takes master's side of both files.
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
…s (RC4 WP3, ADR-2023) The VMAFx API scores frames that already live on a CUDA device without a copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. - Devices: CUDA devices by index (retained primary context) or from the caller's context and stream (external[0], external[1]); count, info and describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC. vmafx_context_use_device imports the device into the context's engine, and features registered afterwards run on their CUDA twins. - Import: device pointers are bound where they are when each plane starts 8-byte aligned with a pitch that is a multiple of 8 (the twins' vector loads); other layouts are refused naming the field, or copied on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (import_convert.cu); CUDA arrays and GL textures are read out on the device. No import path copies through the host, and every host copy site calls vmafx_count_host_copy(). - Fences: a CUDA_EVENT acquire fence is waited on by the device's library stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context; the new VmafxFrameImport release callback (release, user; ABI 0.1.4) lets a producer make its stream wait on the release event without a host stall. The ADR-1199 barrier stays only for pictures that carry no fence ordering. - Pools: CUDA frame pools hand out device frames; a returned frame is reused only after the device readers of its previous use finished. integer_vif_cuda read both pictures with the pitch it computed at init, so an imported plane with another pitch was read from the wrong rows; scale 0 now reads each picture with its own pitch. Master cannot hand the engine a CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's host round trip is fixed in the commit before this one (#2282 on master). Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4): test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0 differing, 6028 imports and 0 host copies. A planted skipped acquire wait gives 15 bad frames of 16 under device load and the real wait 0; a planted early release gives 15 bad canaries, the real release 0. One import scored by two contexts equals each context's own run and is released only after the second context finished. The exact-twin matrix keeps 48 of 48 vif and float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3 skipped.
Merged
15 of 18 tasks
lusoris
force-pushed
the
fix/cuda-ms-ssim-device-level0
branch
from
October 6, 2026 12:54
3de937a to
afd35e8
Compare
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
…s (RC4 WP3, ADR-2023) The VMAFx API scores frames that already live on a CUDA device without a copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. - Devices: CUDA devices by index (retained primary context) or from the caller's context and stream (external[0], external[1]); count, info and describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC. vmafx_context_use_device imports the device into the context's engine, and features registered afterwards run on their CUDA twins. - Import: device pointers are bound where they are when each plane starts 8-byte aligned with a pitch that is a multiple of 8 (the twins' vector loads); other layouts are refused naming the field, or copied on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (import_convert.cu); CUDA arrays and GL textures are read out on the device. No import path copies through the host, and every host copy site calls vmafx_count_host_copy(). - Fences: a CUDA_EVENT acquire fence is waited on by the device's library stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context; the new VmafxFrameImport release callback (release, user; ABI 0.1.4) lets a producer make its stream wait on the release event without a host stall. The ADR-1199 barrier stays only for pictures that carry no fence ordering. - Pools: CUDA frame pools hand out device frames; a returned frame is reused only after the device readers of its previous use finished. integer_vif_cuda read both pictures with the pitch it computed at init, so an imported plane with another pitch was read from the wrong rows; scale 0 now reads each picture with its own pitch. Master cannot hand the engine a CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's host round trip is fixed in the commit before this one (#2282 on master). Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4): test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0 differing, 6028 imports and 0 host copies. A planted skipped acquire wait gives 15 bad frames of 16 under device load and the real wait 0; a planted early release gives 15 bad canaries, the real release 0. One import scored by two contexts equals each context's own run and is released only after the second context finished. The exact-twin matrix keeps 48 of 48 vif and float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3 skipped. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
…s (RC4 WP3, ADR-2023) (#2277) * feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023) The VMAFx API scores frames that already live on a CUDA device without a copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. - Devices: CUDA devices by index (retained primary context) or from the caller's context and stream (external[0], external[1]); count, info and describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC. vmafx_context_use_device imports the device into the context's engine, and features registered afterwards run on their CUDA twins. - Import: device pointers are bound where they are when each plane starts 8-byte aligned with a pitch that is a multiple of 8 (the twins' vector loads); other layouts are refused naming the field, or copied on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (import_convert.cu); CUDA arrays and GL textures are read out on the device. No import path copies through the host, and every host copy site calls vmafx_count_host_copy(). - Fences: a CUDA_EVENT acquire fence is waited on by the device's library stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context; the new VmafxFrameImport release callback (release, user; ABI 0.1.4) lets a producer make its stream wait on the release event without a host stall. The ADR-1199 barrier stays only for pictures that carry no fence ordering. - Pools: CUDA frame pools hand out device frames; a returned frame is reused only after the device readers of its previous use finished. integer_vif_cuda read both pictures with the pitch it computed at init, so an imported plane with another pitch was read from the wrong rows; scale 0 now reads each picture with its own pitch. Master cannot hand the engine a CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's host round trip is fixed in the commit before this one (#2282 on master). Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4): test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0 differing, 6028 imports and 0 host copies. A planted skipped acquire wait gives 15 bad frames of 16 under device load and the real wait 0; a planted early release gives 15 bad canaries, the real release 0. One import scored by two contexts equals each context's own run and is released only after the second context finished. The exact-twin matrix keeps 48 of 48 vif and float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3 skipped. * docs(api): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197) * docs(agents): write the vif_cuda pitch note in the internal register (praetor caveman lint) Signed-off-by: Lusoris <lusoris@proton.me>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
float_ms_ssim_cudano longer sends its input pictures through the host: level 0 of its pyramids is nowpicture_copy()on the device, and the scores are unchanged bit for bit. Before, the twin copied every scored plane of both pictures to pinned host memory, waited for the copy (cuStreamSynchronize()), converted it withpicture_copy()on the host and uploaded the floats again. That is two plane-sized copies to the host, two uploads and two host waits per plane per frame, for pictures that were already on the device (what the FFmpeglibvmaf_cudafilter hands libvmaf). Found by a vendor-profiler trace of the RC4 WP3 CUDA lane (#2277). This is a master defect, so it lands on master first; #2277 carries the same change as already landed.core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu: new kernelms_ssim_picture_to_float. It ispicture_copy()sample for sample: one byte per sample at 8 bits, and at 10, 12 and 16 bits a 16-bit sample divided by 4, 16 or 256 (a power of two, so the quotient is exact). It takes the destination asfloat *, so it needs no integer-to-pointer cast.core/src/feature/cuda/integer_ms_ssim_cuda.c:ms_ssim_stage_inputs()launches it for both pictures on the reference picture's stream, after the distorted picture's ready event and the previous frame'slc.finished. The private stream waits onlc.submitbehind it. The pinned staging buffers (h_input_uint, per-planeh_ref/h_cmp) and thepicture_copy.hinclude are gone.Type
fix— bug fixChecklist
make format && make lintis green locally. The commit hooks pass, among them clang-format, semgrep, source ADR citations, tidy coverage, generated-index freshness, the HISS audit and assertion density.-Denable_cuda=true -Denable_dnn=disabled -Db_lto=false, release):--suite fast --no-suite gpu --num-processes 4: 374 OK, 0 fail.fast+gpuCUDA tests under the device lock: 63 OK, 0 fail.test_cuda_parity_gate_default_run: OK./cross-backend-diffand the worst ULP is ≤ 2.scripts/ci/exact_twin_matrix.py --backends cuda --features float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chromagives 36 cells (8, 10, 12, 16 bit × 4:2:0, 4:2:2, 4:4:4), all 36 equal to the CPU (0 ULP).integer_ms_ssim_sycl.cpp,integer_ms_ssim_hip.c: "pictures arrive as CPU VmafPictures"), runpicture_copy()on the host and upload the floats once. That is one upload and no copy back. Device-resident input on those backends belongs to their RC4 WP3 lanes..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (EUPL-1.2, fork-authored).!orBREAKING CHANGE:and the migration path is documented below. Not a breaking change: no API, option or output changes.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt. Not applicable: no ADR (a bug fix; the arithmetic contract stays ADR-1403 / ADR-1465).tidy: cuda
core/src/feature/cuda/integer_ms_ssim_cuda.c,core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu,core/test/test_cuda_float_ms_ssim_host_traffic.c0 findings, 0 uncited NOLINT. Measured in the dev container with clang-tidy 22.1.8 (scripts/dev/tidy-lane.sh --only ... cuda).memcmpof doubles (now a bit comparison), and the nesting depth of the fill loop. All are fixed.scripts/ci/tidy-baseline-cuda.jsongains the new test inmeasured_sourcesthroughtidy-lane.sh --write --only(generated, not hand-edited).check-tidy-coveragefailed on it before that.Also passing:
scripts/dev/preflight.sh --stage msvcism,core/test/test_win32_pthread_shim_contract.py,scripts/ci/assertion-density.shandpraetorctl audit.Bug-status hygiene (ADR-0165)
docs/state.mdupdated with the closed row T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Golden gate: 280 passed, 3 skipped. Run as
GOLDEN_NINJA_JOBS=4 make test-netflix-golden, withcore/build-goldenbuilt with gcc and pytest from the repository venv.Cross-backend numerical results
Scores are unchanged. Measured on an RTX 4090 (sm_89), driver 615.71.09, CUDA 13.4:
float_ms_ssim,_lcsand_chromacells equal to the CPU.test_cuda_float_ms_ssim_parity(==),_order,_parity_largeandtest_cuda_exact_twinspass.Failing first
test_cuda_float_ms_ssim_host_traffic(new,fast+gpu) feeds device pictures from the pool, as the FFmpeg filter does. It wraps the copy entries of the CUDA state's driver table and counts every copy with a host side while the frames are scored. Results:enable_chromae671d628eenable_chromaOn master the test fails with
float_ms_ssim_cuda uploads while scoring. With the fix, what goes to the host is exactly the per-window term planes, which the host adds in raster order (ADR-1465).The device-free contract
test_cuda_float_ms_ssim_exact_contract.pygains level-0 checks. Run against master's two sources it reports 8 failures (no device conversion, noms_ssim_picture_to_float,CU_MEMORYTYPE_HOSTin the host file). Four planted cases now run every time: host staging (four constructs), an unconverted plane, another divisor, and a kernel without the division.Vendor trace (Nsight Systems 2026.3.2,
--trace=cuda) of the new test:cuStreamSynchronizecalls.Performance (if
perforfeat)No timing claim (RC7). Per plane and frame there are now no copies to the host, no uploads and no host waits.
docs/metrics/ms-ssim.mdnotes that its CUDA timings were measured before this change.Deep-dive deliverables (ADR-0108)
picture_copy()on the device, in its own kernel. Folding the conversion into the decimate kernel would change a kernel the exact-twin contract pins, for no gain in copies.AGENTS.mdinvariant note —core/src/feature/cuda/AGENTS.d/ms-ssim.md, new section "Level 0 is converted on the device" (never bring the host staging back; guards). The per-plane buffer list is updated.changelog.d/changed/cuda-ms-ssim-device-level0.md.docs/rebase-notes.md, "float_ms_ssim_cudabuilds level 0 on the device".User documentation:
docs/metrics/ms-ssim.mdHistory (2026-10-06).Reproducer
Known follow-ups
ms_ssim_score.cuandinteger_ms_ssim_cuda.c.