Skip to content

refactor(cuda): bring the CUDA kernels and their headers to the lint and HISS standard (ADR-1142) - #2109

Merged
lusoris merged 5 commits into
masterfrom
refactor/tidy-zero-cuda-1
Oct 6, 2026
Merged

lusoris merged 5 commits into
masterfrom
refactor/tidy-zero-cuda-1

Conversation

@lusoris

@lusoris lusoris commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Brings the cuda clang-tidy lane to zero findings (581 before, 0 after; measured in the dev container, clang-tidy 22.1.8, scripts/dev/tidy-lane.sh --write cuda). The 62 files are the CUDA kernels (.cu, .cuh), their CUDA-only headers, headers shared with the C hosts, two CUDA tests and the shared picture.h DEVICE_CODE branch. scripts/ci/tidy-baseline-cuda.json is written by the tool, not by hand. Uncited NOLINT: 0.

  • Device pointers are rebuilt from a CUdeviceptr with VMAF_CUDA_DPTR(T, address) (core/src/cuda/cuda_device_ptr.cuh).
  • Helpers and device structs live in anonymous namespaces; __global__ kernels stay extern "C".
  • No designated initializers in device code (nvcc MSVC host frontend, preflight.sh --stage msvcism passes): aggregates are filled field by field.
  • Long functions are split (vif_statistic_calculation into vif_sigmas, vif_gain, vif_accumulate_log, vif_statistic_pixel; SpEED, CAMBI, ADM, SSIM, MS-SSIM, PSNR-HVS helpers).
  • Headers shared with C carry the ADR-1138 / ADR-1470 NOLINTBEGIN brackets, as the cpu lane did.
  • Source-text contract tests follow the new spelling; their assertions are unchanged.

Behaviour check

  • Twin scores at --precision max, 20 extractors on --backend cuda: Netflix 576x324 pair (48 frames), both 1080p checkerboard pairs, a 10-bit 480x270 pair; 0 of 3127 values differ from master, pooled metrics equal.
  • All 264 libvmaf tests with a cuda / gpu / vif / adm / ssim / psnr / motion / cambi / ciede / moment / speed name pass on the RTX 4090 under the lock, including every test_cuda_*_parity and *_large.
  • sm_89 SASS before/after: identical for 15 of 21 kernel TUs; the rest differ by commutative operand order, register names, unsigned compares on non-negative indices, or (filter1d.cu) the split of the VIF statistic; fp64 and conversion instruction counts of filter1d are identical, LDG.E.CONSTANT counts equal everywhere.
  • 4K wall times (bbb 200 frames, 18 features, 2 reps) within noise; the machine was loaded by other lanes.
  • Netflix golden gate: not run; no CPU extractor code changed (comments in integer_vif.h, ssimulacra2_eotf_lut.h, NOLINT brackets only).

NOLINTs added

  • core/src/cuda/cuda_device_ptr.cuh: performance-no-int-to-ptr on the VMAF_CUDA_DPTR macro (ADR-1142). A CUdeviceptr is an integer by the C ABI of the kernel argument structs; an inline function in place of the macro turns 134 of speed_score.cu's 378 LDG.E.CONSTANT loads into plain loads (measured, sm_89, nvcc 13.4).
  • ADR-1138 / ADR-1470 brackets (modernize-use-using, performance-enum-size) in C headers: ssimulacra2_cuda.h, integer_adm_cuda.h, integer_psnr_hvs_cuda.h, integer_cambi_cuda.h, speed_cuda_params.h, integer_vif_cuda.h, integer_vif.h, float_vif_gpu_common.h, float_vif_device.h, float_adm_gpu_common.h, float_adm_device.h, ciede_device.h, libvmaf_cuda.h.
  • ssimulacra2_eotf_lut.h (and its generator): modernize-use-std-numbers around the table; entries are EOTF samples, not constants.
  • integer_vif.h: cert-dcl03-c,misc-static-assert on a runtime assert() (ADR-1142 forbids weakening it).

Checklist (ADR-0108)

  • Research digest: no digest needed: lint refactor, no behaviour change
  • Decision matrix: no alternatives: only-one-way fix (macro vs function measured above)
  • AGENTS.md invariant note: core/src/cuda/AGENTS.md section "Device pointers and device-code lint"
  • Reproducer / smoke-test command: scripts/dev/tidy-lane.sh cuda exits 0; meson test -C build --suite fast+gpu on a CUDA build
  • CHANGELOG fragment: changelog.d/changed/tidy-zero-cuda-lane.md
  • Rebase notes: docs/rebase-notes.md entry added
  • docs/state.md: no state delta: lint refactor, no bug found or closed
  • Netflix golden data: no assertAlmostEqual touched

… frame (#2254)

* fix(cli): say "problem scoring picture" when libvmaf fails to score a frame

vmaf_read_pictures() failing was printed as "problem reading pictures", the
wording of the input read failure ("problem while reading pictures", exit
102), so an extractor refusing a frame or its options read like a broken
input file.

The CLI now prints "problem scoring picture N: libvmaf returned E" after the
libvmaf message that names the extractor. The exit status and the read-failure
wording are unchanged. The exit-code table of docs/usage/cli.md lists the
case.

Closes T-CLI-EXTRACTOR-ERROR-PROPAGATION-2026-10-05.
* fix(test): keep the vmaf-tune tests off the vmaf on PATH

About 40 vmaf-tune tests ran whatever vmaf the host had: the backend probe
defaults to the bare name vmaf, so even with encode and score faked the CLI
tests started it (42 starts measured with a logging stub first on PATH).
pkg/fast's integration test looked vmaf up on PATH as well.

A suite-wide fixture now gives a probe with neither runner nor path a
CPU-only build; a test that exercises the probe passes a runner or the path
of a vmaf of its own. The Go test takes the binary from internal/vmaftest and
the parity test hands it over as VMAF_BIN instead of editing PATH.

Closes T-VMAFTUNE-TESTS-PROBE-PATH-VMAF-2026-10-05 and
T-GO-FAST-INTEGRATION-PATH-VMAF-2026-10-05.
#2252)

* fix(mcp): report backends from vmaf --list-backends, not the help text

Both MCP servers took a GPU backend as built in when vmaf --help listed its
--no_<name> flag. The CLI prints all four on every build, so list_backends
reported CUDA, SYCL, HIP and Metal on a CPU-only build and the backend
allowlist admitted them; the run then failed with exit 100.

The probes now read vmaf --list-backends (ADR-1874) and take a backend as
available only when its row is usable. The Go server goes through
pkg/scorebackend.Detect; the Python server parses the same document. A vmaf
that cannot print the report is treated as CPU-only, with a logged reason.

Closes T-MCP-BACKENDS-FROM-HELP-TEXT-2026-10-05.
…#2249)

* fix(sycl): stop the fused vif_sycl scales overwriting their own input

With vif_fused=true, one launch per scale reads the scale's input from
d_rd_ref / d_rd_dis and writes the next scale's downsampled input into the
same planes. Work-groups of one launch do not wait for each other, so a
work-group could overwrite samples another had not loaded yet. From
1920x1080 up, scales 1 to 3 differed from the CPU vif on every frame
(BBB 3840x2160: 0 of 200 frames identical, up to 4.89e-4). The separate
passes (the default) read and write in two kernels and were exact.

The fused scales now alternate between two pairs of planes
(vif_rd_output()): scale 1 writes a second pair, allocated in fused mode
only at scale 1's output size (4.1 MB at 3840x2160); scale 2 writes the
first pair again after scale 1 has read it, the queue being in order.
close_fex_sycl() frees the buffers through vif_free_buffers().

test_sycl_vif_parity (and _large) compares the CPU, the separate passes
and the fused passes with == at the fixture size, 1920x1080, 1919x1079 and
3840x2160 at 8 and 10 bits. On master the four cases from 1920x1080 up
fail on every frame of scales 1-3; with the fix the test passed 10 of 10
runs on an Arc A380, and BBB 3840x2160, the Netflix pair and both 1080p
checkerboards are identical to the CPU in both modes.
…and HISS standard (ADR-1142) (#2109)

* refactor(cuda): bring the CUDA kernels and their headers to the lint and HISS standard (ADR-1142)
@lusoris
lusoris force-pushed the refactor/tidy-zero-cuda-1 branch from bc3c0a0 to fd8b8c9 Compare October 6, 2026 10:43
@lusoris
lusoris merged commit fd8b8c9 into master Oct 6, 2026
6 of 19 checks passed
@lusoris
lusoris deleted the refactor/tidy-zero-cuda-1 branch October 6, 2026 10:43

This branch was successfully deployed

No deployments
github-pages — fd8b8c93 Deployed Oct 6, 2026 by lusoris via deploy #5143
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:refactor Internal refactor

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant