Skip to content

fix(cuda): build every kernel without FMA contraction; make float_ms_ssim_cuda the CPU's arithmetic - #1695

Merged
lusoris merged 2 commits into
masterfrom
fix/cuda-fmad-off-every-kernel
Oct 1, 2026
Merged

lusoris merged 2 commits into
masterfrom
fix/cuda-fmad-off-every-kernel

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Every CUDA kernel is now built without FMA contraction, through one flag list instead of a per-kernel table, and float_ms_ssim_cuda is bit-identical to the CPU extractor. No twin is measurably slower at 3840x2160.

nvcc fuses a * b + c into one FMA by default. Six of the 21 fatbins were built with --fmad=false and fifteen were not, while the CPU build and, since ADR-1367, the SYCL twins never fuse. This closes T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29 and sets the CUDA floating-point policy in ADR-1403.

What changed

  • One flag list. cuda_device_strict_fp_args in core/src/meson.build: -Xcompiler=<host strict FP> plus --fmad=false under nvcc, -ffp-contract=off under clang's CUDA driver. Every fatbin's command takes it. cuda_cu_extra_flags is empty and may not carry a floating-point flag.
  • Pinned. core/test/test_strict_fp_compiler_args.py executes the policy for both compilers and rejects a per-kernel FP flag, --fmad=true, a fatbin command without the list and a second definition.
  • Explicit FMA where the reference fuses. ms_ssim_decimate.c fuses every tap on purpose, so the ms_ssim_score.cu decimate now writes __fmaf_rn(). ssimulacra2_device.cu and speed_score.cu already spelled theirs.
  • float_ms_ssim_cuda follows the CPU's arithmetic. The flag alone made this twin 2x to 3x worse at 4K, which exposed four differences from the reference. They are fixed here: fused decimate taps, window sums as an exact fp32 pair standing for iqa_convolve()'s fp64 sum, the CPU's mixed fp32 / fp64 operand types for l / c / s, and fp32 per-scale means and constants on the host.
  • The clang CUDA path configures again. -Denable_nvcc=false stopped at Unknown variable name "nvcc_ccbin_flags".

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. (clang-format on the touched files; the clang-tidy ratchet on the touched translation units, CUDA lane: counts equal the baseline.)
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. (Fast suite on a CUDA build, RTX 4090.)
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. (Per-twin table below; float_ms_ssim goes to 0.)
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. (The SYCL, HIP and Metal float_ms_ssim twins are listed there.)
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md). (None added.)
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. (Not a breaking change.)
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29 moved to Recently closed; T-CUDA-FLOAT-MS-SSIM-NOT-CPU-ARITHMETIC-2026-10-01 and T-CUDA-CLANG-PATH-UNCONFIGURABLE-2026-10-01 found and closed; three RC3 rows opened (Known follow-ups).

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. (None changes: no CPU code is touched.)

Cross-backend numerical results

RTX 4090 (sm_89, driver 615.71.09), nvcc 13.4.92, gcc 16.2.1, release without LTO, master bd061d95a against this change, --precision max. Fixtures: Netflix 576x324 pair (48 frames), both 1920x1080 checkerboard pairs (3 frames each), BBB 3840x2160 (50 frames). Full tables: docs/research/1403-cuda-strict-fp-every-kernel.md.

Which kernels change. Fifteen of 21 fatbins keep their machine code on all six shipped architectures: the six that already had the flag (psnr_hvs_score joined them with ADR-1397), and nine in which nvcc had nothing to fuse (adm_*, cambi_score, moment_score, motion_v2_score, psnr_score, and ssim_score, which spells every rounding as an intrinsic since #1677). Six change: filter1d, float_psnr_score, float_motion_score, float_vif_score, ciede_score, ms_ssim_score.

Bit-identical to the CPU on every frame of every fixture

before and after   vif  motion  motion_v2  psnr  psnr_hvs  float_psnr  float_moment  float_ssim  cambi  speed_temporal
new                float_ms_ssim   0/104 frames -> 104/104
                     Netflix 6.9e-8, checkerboards 2.1e-6 and 4.4e-6, 4K 5.8e-7 -> 0

Same value as before on every frame (not bit-identical to the CPU, unchanged by this PR): adm, float_adm, ssim, ssimulacra2, speed_chroma.

Moved (max abs diff against the CPU, before -> after)

twin          Netflix 576x324        BBB 3840x2160          note
ciede         1.14e-5 -> 1.14e-5     1.40e-6 -> 1.49e-6     4K mean 1.17e-6 -> 7.7e-7
float_vif     2.72e-5 -> 3.81e-5     5.75e-6 -> 7.03e-6     tolerance 5e-5; same numbers as SYCL before/after ADR-1367
float_motion  3.01e-6 -> 3.12e-6     2.37e-5 -> 2.36e-5     checkerboards 1.35e-4 -> 1.36e-4

Every twin is inside its ADR-0214 tolerance on the Netflix pair, the fixture the gate runs: cross_backend_parity_gate.py --backends cpu cuda passes all 18 cells before and after. float_motion on the checkerboards is outside its tolerance before and after alike, for a reason the flag does not touch: the CPU keeps one fp32 running sum. A host replay shows it (the fp32 running sum reproduces the CPU score bit for bit; the exact sum of the same terms is 1.35e-4 lower and the twin is within 6e-7 of it).

float_ms_ssim_cuda, step by step (max abs diff)

step                                       Netflix   checker 1px  checker 10px  BBB 4K
before                                     6.89e-8   2.14e-6      4.42e-6       5.82e-7
--fmad=false only                          5.53e-8   1.05e-6      2.97e-6       1.23e-6
+ decimate taps as __fmaf_rn()             6.80e-8   1.57e-6      3.16e-6       1.13e-6
+ window sums in fp64                      2.73e-8   3.19e-8      3.45e-8       3.49e-8
+ reference operand types, fp32 means      0         0            0             0
window sums as an fp32 pair, not fp64      0         0            0             0

The 15 enable_lcs outputs were up to 1.3e-6 off and are identical too.

clang CUDA path (clang 22.1.8, same tree): 18 of 19 twins give the nvcc build's value on every output of every frame (3 952 outputs); ciede differs (device math functions). float_ms_ssim_cuda is bit-identical to the CPU under both compilers; that needed __fsqrt_rn(), because clang compiles sqrtf() to sqrt.approx.f32.

Performance (if perf or feat)

3840x2160, BBB, (t(N) - t(2)) / (N - 2) ms per frame, alternating pairs on a shared host (load average 9 to 16). Median before and after, then the median and quartiles of the paired difference.

15 pairs, N = 102
twin            code       before  after   paired difference
vif             changed     3.42    3.44   -0.07 (-0.30 .. +0.21)
float_psnr      changed     2.62    2.61   -0.07 (-1.03 .. +0.13)
float_motion    changed     2.73    3.03   +0.21 (-0.24 .. +0.44)
float_vif       changed     2.53    2.65   +0.12 (-0.02 .. +0.22)
ciede           changed     2.42    2.46   +0.04 (-0.05 .. +0.59)
float_ms_ssim   changed    11.26   10.87   -0.12 (-1.34 .. +0.62)
psnr_hvs        identical  12.18   12.09   -0.08 (-0.16 .. +0.10)
psnr            identical   2.84    3.12   +0.07 (-0.12 .. +0.30)
adm             identical   3.50    3.47   +0.04 (-0.11 .. +0.14)
speed_chroma    identical   2.38    2.44   +0.04 (-0.08 .. +0.16)

21 pairs, N = 200 (the two largest again)
float_motion    changed     2.82    2.89   +0.10 (-0.22 .. +0.27)
float_vif       changed     2.43    2.56    0.00 (-0.04 .. +0.05)
psnr            identical   2.61    2.38   -0.06 (-0.35 .. +0.07)
float_moment    identical   2.53    2.45   -0.03 (-0.11 .. +0.24)

No difference is distinguishable from the controls that run identical machine code. The other thirteen twins run the same code as before. One thing was measurable and was avoided: plain fp64 accumulators for the float_ms_ssim window sums give the same scores and cost 3.4 ms per 4K frame (10.7 -> 14.1); the fp32 pair costs nothing.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1403-cuda-strict-fp-every-kernel.md.
  • Decision matrix — docs/adr/1403-cuda-strict-fp-every-kernel.md, ## Alternatives considered.
  • AGENTS.md invariant note — core/src/cuda/AGENTS.md (one FP list, never per-kernel), core/src/feature/cuda/AGENTS.md (the policy and the five things that keep float_ms_ssim_cuda exact), docs/development/rebase-sensitive-invariants.md.
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/changed/cuda-strict-fp-every-kernel.md.
  • Rebase note — docs/rebase-notes.md, ADR-1403.

Reproducer

meson setup build-cuda core -Denable_cuda=true -Denable_sycl=false --buildtype=release
ninja -C build-cuda
# No device needed: the policy for nvcc and clang, and the float_ms_ssim arithmetic.
python3 -m pytest core/test/test_strict_fp_compiler_args.py core/test/test_cuda_kernel_source_contract.py \
  core/test/test_cuda_device_resident_contract.py
# RTX 4090: 16 outputs of 3 frames compared with ==; all 48 differ on master.
build-cuda/test/test_cuda_float_ms_ssim_parity
# Every frame of the Netflix pair and of BBB 4K identical to --backend cpu.
python3 scripts/dev/speed_gpu_parity.py --backend cuda --vmaf "$PWD/build-cuda/tools/vmaf" \
  --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb --feature float_ms_ssim --no-timing
# FMAs left in a kernel.
cuobjdump -sass -arch sm_89 build-cuda/src/float_vif_score.fatbin | grep -c FFMA

Known follow-ups

  • T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01: float_motion_cuda is 1.36e-4 from the CPU on the 1080p checkerboards, above the 5e-5 tolerance, because the CPU keeps one fp32 running sum. Fix: reduce in the CPU's order on the device.
  • T-CUDA-FLOAT-VIF-RESIDUAL-2026-10-01: float_vif_cuda matches on no frame and uses 3.8e-5 of its 5e-5; not isolated (device log2f, order of sums).
  • T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01: the SYCL, HIP and Metal float_ms_ssim twins keep the arithmetic the CUDA twin had (read from source, not run).
  • HIP got the same policy in fix(hip): compile every HIP kernel with contraction off through one shared flag list #1694 (ADR-1407).
  • Kernel-only timing needs a quiet host or an in-process harness; vmaf_bench --gpu-profile is SYCL-only.
  • Bases: parity, fatbin comparison and the gate are on bd061d95a (after feat(cuda): run float_ssim on the device at every scale, equal to the CPU #1677 rewrote ssim_score.cu); the timings are from bd17c352f, where the six changed fatbins are the same; the float_ms_ssim step table and the fp64 cost are from 2c3acf1c9.
  • T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01 has a fix stacked on this branch (fix/cuda-float-motion-cpu-float-sum).

@lusoris
lusoris force-pushed the fix/cuda-fmad-off-every-kernel branch 2 times, most recently from 08ce781 to 271932f Compare October 1, 2026 12:02
@lusoris
lusoris force-pushed the fix/cuda-fmad-off-every-kernel branch 2 times, most recently from b69128a to 832e5b1 Compare October 1, 2026 12:44
…ssim_cuda the CPU's arithmetic

nvcc fuses a * b + c into one FMA by default. Six of the 21 CUDA fatbins
were built with --fmad=false through a per-kernel table and fifteen were
not, while the CPU build and, since ADR-1367, the SYCL twins never fuse.
ADR-1403 makes this one policy:

- core/src/meson.build: every fatbin takes cuda_device_strict_fp_args,
  defined once (--fmad=false with the host strict args under nvcc,
  -ffp-contract=off under clang CUDA). cuda_cu_extra_flags is empty and
  may not carry a floating-point flag.
- test_strict_fp_compiler_args.py executes the policy for both compilers
  and rejects a per-kernel FP flag, --fmad=true, a fatbin command without
  the list and a second definition.
- A kernel whose reference fuses on purpose writes __fmaf_rn(): the
  ms_ssim_score.cu decimate now does (ms_ssim_decimate.c uses
  vmaf_fmaf_exact()).

float_ms_ssim_cuda got 2x to 3x worse at 4K with the flag alone, which
exposed four differences from the CPU extractor. Its kernels now follow
the reference operation for operation: fused decimate taps, window sums
as an exact fp32 pair standing for iqa_convolve()'s fp64 sum, the CPU's
fp32 denominators and quotient for l / c / s, and fp32 constants and
per-scale means on the host.

The clang CUDA path (-Denable_nvcc=false) configures again: two flag
lists were assigned in the nvcc branch only. Built with clang, 18 of the
19 twins give the nvcc build's value on every frame.

Measured on an RTX 4090 against --backend cpu at --precision max
(Netflix 576x324 pair, both 1080p checkerboard pairs, BBB 3840x2160):

- Fifteen fatbins keep their machine code on all six architectures;
  six change.
- float_ms_ssim: 0 of 104 frames identical before (up to 4.4e-6 off),
  104 of 104 after, enable_lcs outputs included.
- adm, vif, motion, motion_v2, psnr, psnr_hvs, float_psnr, float_moment,
  float_ssim, cambi, float_adm, ssim, ssimulacra2, speed_chroma and
  speed_temporal produce the same values as before.
- ciede, float_vif and float_motion move in their last digits, each
  where it was against its ADR-0214 tolerance. The parity gate passes
  all 18 cells before and after.
- No twin is measurably slower at 3840x2160.

Closes T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29. Opens
T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01,
T-CUDA-FLOAT-VIF-RESIDUAL-2026-10-01 and
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
@lusoris
lusoris force-pushed the fix/cuda-fmad-off-every-kernel branch from 832e5b1 to 2789a33 Compare October 1, 2026 12:58
@lusoris
lusoris merged commit 691ca56 into master Oct 1, 2026
63 of 69 checks passed
@lusoris
lusoris deleted the fix/cuda-fmad-off-every-kernel branch October 1, 2026 12:59
@github-actions github-actions Bot added the type:bug Something isn't working label Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant