Skip to content

fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host - #2282

Merged
lusoris merged 3 commits into
masterfrom
fix/cuda-ms-ssim-device-level0
Oct 6, 2026
Merged

lusoris merged 3 commits into
masterfrom
fix/cuda-ms-ssim-device-level0

Conversation

@lusoris

@lusoris lusoris commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Summary

float_ms_ssim_cuda no longer sends its input pictures through the host: level 0 of its pyramids is now picture_copy() on the device, and the scores are unchanged bit for bit. Before, the twin copied every scored plane of both pictures to pinned host memory, waited for the copy (cuStreamSynchronize()), converted it with picture_copy() on the host and uploaded the floats again. That is two plane-sized copies to the host, two uploads and two host waits per plane per frame, for pictures that were already on the device (what the FFmpeg libvmaf_cuda filter hands libvmaf). Found by a vendor-profiler trace of the RC4 WP3 CUDA lane (#2277). This is a master defect, so it lands on master first; #2277 carries the same change as already landed.

  • core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu: new kernel ms_ssim_picture_to_float. It is picture_copy() sample for sample: one byte per sample at 8 bits, and at 10, 12 and 16 bits a 16-bit sample divided by 4, 16 or 256 (a power of two, so the quotient is exact). It takes the destination as float *, so it needs no integer-to-pointer cast.
  • core/src/feature/cuda/integer_ms_ssim_cuda.c: ms_ssim_stage_inputs() launches it for both pictures on the reference picture's stream, after the distorted picture's ready event and the previous frame's lc.finished. The private stream waits on lc.submit behind it. The pinned staging buffers (h_input_uint, per-plane h_ref / h_cmp) and the picture_copy.h include are gone.

Type

  • fix — bug fix

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. The commit hooks pass, among them clang-format, semgrep, source ADR citations, tidy coverage, generated-index freshness, the HISS audit and assertion density.
  • Unit tests pass. On a CUDA build (-Denable_cuda=true -Denable_dnn=disabled -Db_lto=false, release):
    • --suite fast --no-suite gpu --num-processes 4: 374 OK, 0 fail.
    • The 63 fast + gpu CUDA tests under the device lock: 63 OK, 0 fail.
    • test_cuda_parity_gate_default_run: OK.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. scripts/ci/exact_twin_matrix.py --backends cuda --features float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chroma gives 36 cells (8, 10, 12, 16 bit × 4:2:0, 4:2:2, 4:4:4), all 36 equal to the CPU (0 ULP).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. The SYCL and HIP twins do not have this defect. They receive host pictures (integer_ms_ssim_sycl.cpp, integer_ms_ssim_hip.c: "pictures arrive as CPU VmafPictures"), run picture_copy() on the host and upload the floats once. That is one upload and no copy back. Device-resident input on those backends belongs to their RC4 WP3 lanes.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (EUPL-1.2, fork-authored).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not a breaking change: no API, option or output changes.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt. Not applicable: no ADR (a bug fix; the arithmetic contract stays ADR-1403 / ADR-1465).

tidy: cuda core/src/feature/cuda/integer_ms_ssim_cuda.c, core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu, core/test/test_cuda_float_ms_ssim_host_traffic.c 0 findings, 0 uncited NOLINT. Measured in the dev container with clang-tidy 22.1.8 (scripts/dev/tidy-lane.sh --only ... cuda).

  • The first run of the new test found 3 findings: braces, a memcmp of doubles (now a bit comparison), and the nesting depth of the fill loop. All are fixed.
  • scripts/ci/tidy-baseline-cuda.json gains the new test in measured_sources through tidy-lane.sh --write --only (generated, not hand-edited). check-tidy-coverage failed on it before that.

Also passing: scripts/dev/preflight.sh --stage msvcism, core/test/test_win32_pthread_shim_contract.py, scripts/ci/assertion-density.sh and praetorctl audit.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated with the closed row T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below. Not applicable.

Golden gate: 280 passed, 3 skipped. Run as GOLDEN_NINJA_JOBS=4 make test-netflix-golden, with core/build-golden built with gcc and pytest from the repository venv.

Cross-backend numerical results

Scores are unchanged. Measured on an RTX 4090 (sm_89), driver 615.71.09, CUDA 13.4:

  • The exact-twin matrix gives 36 of 36 float_ms_ssim, _lcs and _chroma cells equal to the CPU.
  • test_cuda_float_ms_ssim_parity (==), _order, _parity_large and test_cuda_exact_twins pass.
  • The new test compares every output of every frame with the CPU's bits.

Failing first

test_cuda_float_ms_ssim_host_traffic (new, fast + gpu) feeds device pictures from the pool, as the FFmpeg filter does. It wraps the copy entries of the CUDA state's driver table and counts every copy with a host side while the frames are scored. Results:

Case (3 frames) Code Bytes to the device Plane copies with a host side Bytes to the host Term planes
8-bit 4:4:4, enable_chroma master e671d628e 3538944 18 11197296 10312560
8-bit 4:4:4, enable_chroma this PR 0 0 10312560 10312560
10-bit 4:2:0, luma this PR 0 0 3437520 3437520

On master the test fails with float_ms_ssim_cuda uploads while scoring. With the fix, what goes to the host is exactly the per-window term planes, which the host adds in raster order (ADR-1465).

The device-free contract test_cuda_float_ms_ssim_exact_contract.py gains level-0 checks. Run against master's two sources it reports 8 failures (no device conversion, no ms_ssim_picture_to_float, CU_MEMORYTYPE_HOST in the host file). Four planted cases now run every time: host staging (four constructs), an unconverted plane, another divisor, and a kernel without the division.

Vendor trace (Nsight Systems 2026.3.2, --trace=cuda) of the new test:

  • Before (master code, the 8-bit 4:4:4 case): 36 host-to-device copies (4.42 MB, of which 3.54 MB are the staging uploads), 153 device-to-host copies (11.20 MB, of which 0.88 MB are the 18 plane copies), 35 cuStreamSynchronize calls.
  • After (both cases): host to device only the test's own picture fills (1.77 MB); device to host only the term planes (13.75 MB = 10312560 + 3437520 bytes).

Performance (if perf or feat)

No timing claim (RC7). Per plane and frame there are now no copies to the host, no uploads and no host waits. docs/metrics/ms-ssim.md notes that its CUDA timings were measured before this change.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: one defect, measured by the test and the trace above.
  • Decision matrix — no alternatives: only-one-way fix. picture_copy() on the device, in its own kernel. Folding the conversion into the decimate kernel would change a kernel the exact-twin contract pins, for no gain in copies.
  • AGENTS.md invariant note — core/src/feature/cuda/AGENTS.d/ms-ssim.md, new section "Level 0 is converted on the device" (never bring the host staging back; guards). The per-plane buffer list is updated.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/changed/cuda-ms-ssim-device-level0.md.
  • Rebase note — docs/rebase-notes.md, "float_ms_ssim_cuda builds level 0 on the device".

User documentation: docs/metrics/ms-ssim.md History (2026-10-06).

Reproducer

meson setup build-cuda core -Denable_cuda=true -Denable_dnn=disabled -Db_lto=false
nice -n 10 ninja -C build-cuda -j6
flock ~/.cache/vmafx-locks/cuda-4090.lock timeout 300 build-cuda/test/test_cuda_float_ms_ssim_host_traffic
python3 core/test/test_cuda_float_ms_ssim_exact_contract.py
python3 scripts/ci/exact_twin_matrix.py --vmaf-binary build-cuda/tools/vmaf --backends cuda \
  --features float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chroma   # takes the lock itself

Known follow-ups

@lusoris lusoris added type:bug Something isn't working cuda labels Oct 6, 2026
…wn stub headers (#2281)

* chore(tidy): measure the MATLAB MEX sources in the cpu lane against own stub headers

The ten MEX sources of compat/python-vmaf/matlab/ include mex.h and matrix.h
from the MATLAB SDK, which no runner has, so no lane measured them and they sat
in the exception list. Self-authored stubs under scripts/ci/lint-stubs/matlab/
declare only what the ten files use; gen-mex-compile-commands.py adds their
entries to the lane's compilation database (cpu lane and changed-files job),
and nothing builds or links the stubs. The ten exceptions and the workflow
exclusion are removed.

The first measurement found about 230 findings. The mechanical ones are fixed
(braces, isolated declarations, static, const, unused parameters, widening,
missing includes, dead stores) together with one defect: ical_std.c passed the
data pointer of a matrix to mxDestroyArray(). Thirteen function-size findings
stay in tidy-baseline-cpu.json: nothing here runs MATLAB, so a refactor of those
functions needs a test first.

Closes T-TIDY-MATLAB-MEX-UNMEASURED-2026-09-22.

* test(ci): type the generated-module handle of the MEX compile-command test

* docs: regenerate the indexes and the citation map after rebasing
…DR numbers (#2275)

* docs(adr): remove four unfilled template blocks and file four cited ADR numbers

ADR-0643, ADR-0665, ADR-0666 and ADR-0673 began with an unfilled allocator
template (title `<fill in title>`) in front of their real text; the block is
removed and a status update says so. ADR-0228, ADR-0636,
ADR-0867 and ADR-0979 are cited by other ADRs and had no file; each gets a
short record written from the commits and ADRs that name it.

* docs: regenerate the indexes and the citation map after rebasing
…nd-tripping every plane through the host (#2282)

* fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host

float_ms_ssim_cuda copied every scored plane of both input pictures to
pinned host memory, waited for the copy, converted it with picture_copy()
on the host and uploaded the floats again: per plane and frame two
plane-sized copies to the host, two uploads and two host waits, for
pictures that were already on the device. Level 0 of the pyramids is now
picture_copy() on the device (ms_ssim_picture_to_float: the same samples,
division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference
picture's stream after the distorted picture's ready event, and the
private stream waits behind it. The pinned staging buffers are gone.

test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts
every copy with a host side while they are scored. On the old code, 8-bit
4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18
plane copies with a host side; now 0 and 0, and what goes to the host is
exactly the per-window term planes the host adds in raster order. The
device-free contract test refuses the host staging. Scores are unchanged:
every output equals the CPU's bits, and the exact-twin matrix keeps 36 of
36 float_ms_ssim cells equal.
lusoris added a commit that referenced this pull request Oct 6, 2026
…nd-tripping every plane through the host

float_ms_ssim_cuda copied every scored plane of both input pictures to
pinned host memory, waited for the copy, converted it with picture_copy()
on the host and uploaded the floats again: per plane and frame two
plane-sized copies to the host, two uploads and two host waits, for
pictures that were already on the device. Level 0 of the pyramids is now
picture_copy() on the device (ms_ssim_picture_to_float: the same samples,
division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference
picture's stream after the distorted picture's ready event, and the
private stream waits behind it. The pinned staging buffers are gone.

test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts
every copy with a host side while they are scored. On the old code, 8-bit
4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18
plane copies with a host side; now 0 and 0, and what goes to the host is
exactly the per-window term planes the host adds in raster order. The
device-free contract test refuses the host staging. Scores are unchanged:
every output equals the CPU's bits, and the exact-twin matrix keeps 36 of
36 float_ms_ssim cells equal.

This is #2282 on master, carried by the RC4 WP3 CUDA lane until the stack
is restacked onto a master that has it; the restack drops this commit and
takes master's side of both files.
lusoris added a commit that referenced this pull request Oct 6, 2026
…s (RC4 WP3, ADR-2023)

The VMAFx API scores frames that already live on a CUDA device without a
copy through the host, and imported frames score bit for bit as the same
frames uploaded from the host for every CUDA twin declared exact.

- Devices: CUDA devices by index (retained primary context) or from the
  caller's context and stream (external[0], external[1]); count, info and
  describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and
  GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC.
  vmafx_context_use_device imports the device into the context's engine, and
  features registered afterwards run on their CUDA twins.
- Import: device pointers are bound where they are when each plane starts
  8-byte aligned with a pitch that is a multiple of 8 (the twins' vector
  loads); other layouts are refused naming the field, or copied on the
  device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on
  the device (import_convert.cu); CUDA arrays and GL textures are read out
  on the device. No import path copies through the host, and every host copy
  site calls vmafx_count_host_copy().
- Fences: a CUDA_EVENT acquire fence is waited on by the device's library
  stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL
  textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after
  the last reader in every context; the new VmafxFrameImport release callback
  (release, user; ABI 0.1.4) lets a producer make its stream wait on the
  release event without a host stall. The ADR-1199 barrier stays only for
  pictures that carry no fence ordering.
- Pools: CUDA frame pools hand out device frames; a returned frame is reused
  only after the device readers of its previous use finished.

integer_vif_cuda read both pictures with the pitch it computed at init, so
an imported plane with another pitch was read from the wrong rows; scale 0
now reads each picture with its own pitch. Master cannot hand the engine a
CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's
host round trip is fixed in the commit before this one (#2282 on master).

Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4):
test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both
checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0
differing, 6028 imports and 0 host copies. A planted skipped acquire wait
gives 15 bad frames of 16 under device load and the real wait 0; a planted
early release gives 15 bad canaries, the real release 0. One import scored
by two contexts equals each context's own run and is released only after
the second context finished. The exact-twin matrix keeps 48 of 48 vif and
float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3
skipped.
@lusoris
lusoris force-pushed the fix/cuda-ms-ssim-device-level0 branch from 3de937a to afd35e8 Compare October 6, 2026 12:54
@lusoris
lusoris merged commit afd35e8 into master Oct 6, 2026
5 of 90 checks passed
@lusoris
lusoris deleted the fix/cuda-ms-ssim-device-level0 branch October 6, 2026 12:54
lusoris added a commit that referenced this pull request Oct 8, 2026
…s (RC4 WP3, ADR-2023)

The VMAFx API scores frames that already live on a CUDA device without a
copy through the host, and imported frames score bit for bit as the same
frames uploaded from the host for every CUDA twin declared exact.

- Devices: CUDA devices by index (retained primary context) or from the
  caller's context and stream (external[0], external[1]); count, info and
  describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and
  GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC.
  vmafx_context_use_device imports the device into the context's engine, and
  features registered afterwards run on their CUDA twins.
- Import: device pointers are bound where they are when each plane starts
  8-byte aligned with a pitch that is a multiple of 8 (the twins' vector
  loads); other layouts are refused naming the field, or copied on the
  device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on
  the device (import_convert.cu); CUDA arrays and GL textures are read out
  on the device. No import path copies through the host, and every host copy
  site calls vmafx_count_host_copy().
- Fences: a CUDA_EVENT acquire fence is waited on by the device's library
  stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL
  textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after
  the last reader in every context; the new VmafxFrameImport release callback
  (release, user; ABI 0.1.4) lets a producer make its stream wait on the
  release event without a host stall. The ADR-1199 barrier stays only for
  pictures that carry no fence ordering.
- Pools: CUDA frame pools hand out device frames; a returned frame is reused
  only after the device readers of its previous use finished.

integer_vif_cuda read both pictures with the pitch it computed at init, so
an imported plane with another pitch was read from the wrong rows; scale 0
now reads each picture with its own pitch. Master cannot hand the engine a
CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's
host round trip is fixed in the commit before this one (#2282 on master).

Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4):
test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both
checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0
differing, 6028 imports and 0 host copies. A planted skipped acquire wait
gives 15 bad frames of 16 under device load and the real wait 0; a planted
early release gives 15 bad canaries, the real release 0. One import scored
by two contexts equals each context's own run and is released only after
the second context finished. The exact-twin matrix keeps 48 of 48 vif and
float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3
skipped.

Signed-off-by: Lusoris <lusoris@proton.me>
lusoris added a commit that referenced this pull request Oct 8, 2026
…s (RC4 WP3, ADR-2023) (#2277)

* feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023)

The VMAFx API scores frames that already live on a CUDA device without a
copy through the host, and imported frames score bit for bit as the same
frames uploaded from the host for every CUDA twin declared exact.

- Devices: CUDA devices by index (retained primary context) or from the
  caller's context and stream (external[0], external[1]); count, info and
  describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and
  GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC.
  vmafx_context_use_device imports the device into the context's engine, and
  features registered afterwards run on their CUDA twins.
- Import: device pointers are bound where they are when each plane starts
  8-byte aligned with a pitch that is a multiple of 8 (the twins' vector
  loads); other layouts are refused naming the field, or copied on the
  device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on
  the device (import_convert.cu); CUDA arrays and GL textures are read out
  on the device. No import path copies through the host, and every host copy
  site calls vmafx_count_host_copy().
- Fences: a CUDA_EVENT acquire fence is waited on by the device's library
  stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL
  textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after
  the last reader in every context; the new VmafxFrameImport release callback
  (release, user; ABI 0.1.4) lets a producer make its stream wait on the
  release event without a host stall. The ADR-1199 barrier stays only for
  pictures that carry no fence ordering.
- Pools: CUDA frame pools hand out device frames; a returned frame is reused
  only after the device readers of its previous use finished.

integer_vif_cuda read both pictures with the pitch it computed at init, so
an imported plane with another pitch was read from the wrong rows; scale 0
now reads each picture with its own pitch. Master cannot hand the engine a
CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's
host round trip is fixed in the commit before this one (#2282 on master).

Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4):
test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both
checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0
differing, 6028 imports and 0 host copies. A planted skipped acquire wait
gives 15 bad frames of 16 under device load and the real wait 0; a planted
early release gives 15 bad canaries, the real release 0. One import scored
by two contexts equals each context's own run and is released only after
the second context finished. The exact-twin matrix keeps 48 of 48 vif and
float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3
skipped.

* docs(api): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197)

* docs(agents): write the vif_cuda pitch note in the internal register (praetor caveman lint)

Signed-off-by: Lusoris <lusoris@proton.me>

This branch was successfully deployed

1 active deployment
github-pages — afd35e89 Deployed Oct 6, 2026 by lusoris via deploy #5187
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant