Repository navigation
fix(cuda): build every kernel without FMA contraction; make float_ms_ssim_cuda the CPU's arithmetic - #1695
Merged
Conversation
9 of 26 tasks
lusoris
force-pushed
the
fix/cuda-fmad-off-every-kernel
branch
2 times, most recently
from
October 1, 2026 12:02
08ce781 to
271932f
Compare
16 of 26 tasks
lusoris
force-pushed
the
fix/cuda-fmad-off-every-kernel
branch
2 times, most recently
from
October 1, 2026 12:44
b69128a to
832e5b1
Compare
…ssim_cuda the CPU's arithmetic nvcc fuses a * b + c into one FMA by default. Six of the 21 CUDA fatbins were built with --fmad=false through a per-kernel table and fifteen were not, while the CPU build and, since ADR-1367, the SYCL twins never fuse. ADR-1403 makes this one policy: - core/src/meson.build: every fatbin takes cuda_device_strict_fp_args, defined once (--fmad=false with the host strict args under nvcc, -ffp-contract=off under clang CUDA). cuda_cu_extra_flags is empty and may not carry a floating-point flag. - test_strict_fp_compiler_args.py executes the policy for both compilers and rejects a per-kernel FP flag, --fmad=true, a fatbin command without the list and a second definition. - A kernel whose reference fuses on purpose writes __fmaf_rn(): the ms_ssim_score.cu decimate now does (ms_ssim_decimate.c uses vmaf_fmaf_exact()). float_ms_ssim_cuda got 2x to 3x worse at 4K with the flag alone, which exposed four differences from the CPU extractor. Its kernels now follow the reference operation for operation: fused decimate taps, window sums as an exact fp32 pair standing for iqa_convolve()'s fp64 sum, the CPU's fp32 denominators and quotient for l / c / s, and fp32 constants and per-scale means on the host. The clang CUDA path (-Denable_nvcc=false) configures again: two flag lists were assigned in the nvcc branch only. Built with clang, 18 of the 19 twins give the nvcc build's value on every frame. Measured on an RTX 4090 against --backend cpu at --precision max (Netflix 576x324 pair, both 1080p checkerboard pairs, BBB 3840x2160): - Fifteen fatbins keep their machine code on all six architectures; six change. - float_ms_ssim: 0 of 104 frames identical before (up to 4.4e-6 off), 104 of 104 after, enable_lcs outputs included. - adm, vif, motion, motion_v2, psnr, psnr_hvs, float_psnr, float_moment, float_ssim, cambi, float_adm, ssim, ssimulacra2, speed_chroma and speed_temporal produce the same values as before. - ciede, float_vif and float_motion move in their last digits, each where it was against its ADR-0214 tolerance. The parity gate passes all 18 cells before and after. - No twin is measurably slower at 3840x2160. Closes T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29. Opens T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01, T-CUDA-FLOAT-VIF-RESIDUAL-2026-10-01 and T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
lusoris
force-pushed
the
fix/cuda-fmad-off-every-kernel
branch
from
October 1, 2026 12:58
832e5b1 to
2789a33
Compare
This was referenced Oct 1, 2026
3 of 8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every CUDA kernel is now built without FMA contraction, through one flag list instead of a per-kernel table, and
float_ms_ssim_cudais bit-identical to the CPU extractor. No twin is measurably slower at 3840x2160.nvcc fuses
a * b + cinto one FMA by default. Six of the 21 fatbins were built with--fmad=falseand fifteen were not, while the CPU build and, since ADR-1367, the SYCL twins never fuse. This closesT-CUDA-FP-CONTRACT-DEFAULT-2026-09-29and sets the CUDA floating-point policy in ADR-1403.What changed
cuda_device_strict_fp_argsincore/src/meson.build:-Xcompiler=<host strict FP>plus--fmad=falseunder nvcc,-ffp-contract=offunder clang's CUDA driver. Every fatbin's command takes it.cuda_cu_extra_flagsis empty and may not carry a floating-point flag.core/test/test_strict_fp_compiler_args.pyexecutes the policy for both compilers and rejects a per-kernel FP flag,--fmad=true, a fatbin command without the list and a second definition.ms_ssim_decimate.cfuses every tap on purpose, so thems_ssim_score.cudecimate now writes__fmaf_rn().ssimulacra2_device.cuandspeed_score.cualready spelled theirs.float_ms_ssim_cudafollows the CPU's arithmetic. The flag alone made this twin 2x to 3x worse at 4K, which exposed four differences from the reference. They are fixed here: fused decimate taps, window sums as an exact fp32 pair standing foriqa_convolve()'s fp64 sum, the CPU's mixed fp32 / fp64 operand types forl/c/s, and fp32 per-scale means and constants on the host.-Denable_nvcc=falsestopped atUnknown variable name "nvcc_ccbin_flags".Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. (clang-format on the touched files; the clang-tidy ratchet on the touched translation units, CUDA lane: counts equal the baseline.)python3 scripts/ci/run_meson_test.py -- -C build. (Fast suite on a CUDA build, RTX 4090.)/cross-backend-diffand the worst ULP is ≤ 2. (Per-twin table below;float_ms_ssimgoes to 0.)float_ms_ssimtwins are listed there.).c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md). (None added.)!orBREAKING CHANGE:and the migration path is documented below. (Not a breaking change.)docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29moved to Recently closed;T-CUDA-FLOAT-MS-SSIM-NOT-CPU-ARITHMETIC-2026-10-01andT-CUDA-CLANG-PATH-UNCONFIGURABLE-2026-10-01found and closed; three RC3 rows opened (Known follow-ups).Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
RTX 4090 (sm_89, driver 615.71.09), nvcc 13.4.92, gcc 16.2.1, release without LTO,
masterbd061d95aagainst this change,--precision max. Fixtures: Netflix 576x324 pair (48 frames), both 1920x1080 checkerboard pairs (3 frames each), BBB 3840x2160 (50 frames). Full tables:docs/research/1403-cuda-strict-fp-every-kernel.md.Which kernels change. Fifteen of 21 fatbins keep their machine code on all six shipped architectures: the six that already had the flag (
psnr_hvs_scorejoined them with ADR-1397), and nine in which nvcc had nothing to fuse (adm_*,cambi_score,moment_score,motion_v2_score,psnr_score, andssim_score, which spells every rounding as an intrinsic since #1677). Six change:filter1d,float_psnr_score,float_motion_score,float_vif_score,ciede_score,ms_ssim_score.Bit-identical to the CPU on every frame of every fixture
Same value as before on every frame (not bit-identical to the CPU, unchanged by this PR):
adm,float_adm,ssim,ssimulacra2,speed_chroma.Moved (max abs diff against the CPU, before -> after)
Every twin is inside its ADR-0214 tolerance on the Netflix pair, the fixture the gate runs:
cross_backend_parity_gate.py --backends cpu cudapasses all 18 cells before and after.float_motionon the checkerboards is outside its tolerance before and after alike, for a reason the flag does not touch: the CPU keeps one fp32 running sum. A host replay shows it (the fp32 running sum reproduces the CPU score bit for bit; the exact sum of the same terms is 1.35e-4 lower and the twin is within 6e-7 of it).float_ms_ssim_cuda, step by step (max abs diff)The 15
enable_lcsoutputs were up to 1.3e-6 off and are identical too.clang CUDA path (clang 22.1.8, same tree): 18 of 19 twins give the nvcc build's value on every output of every frame (3 952 outputs);
ciedediffers (device math functions).float_ms_ssim_cudais bit-identical to the CPU under both compilers; that needed__fsqrt_rn(), because clang compilessqrtf()tosqrt.approx.f32.Performance (if
perforfeat)3840x2160, BBB,
(t(N) - t(2)) / (N - 2)ms per frame, alternating pairs on a shared host (load average 9 to 16). Median before and after, then the median and quartiles of the paired difference.No difference is distinguishable from the controls that run identical machine code. The other thirteen twins run the same code as before. One thing was measurable and was avoided: plain fp64 accumulators for the
float_ms_ssimwindow sums give the same scores and cost 3.4 ms per 4K frame (10.7 -> 14.1); the fp32 pair costs nothing.Deep-dive deliverables (ADR-0108)
docs/research/1403-cuda-strict-fp-every-kernel.md.docs/adr/1403-cuda-strict-fp-every-kernel.md,## Alternatives considered.AGENTS.mdinvariant note —core/src/cuda/AGENTS.md(one FP list, never per-kernel),core/src/feature/cuda/AGENTS.md(the policy and the five things that keepfloat_ms_ssim_cudaexact),docs/development/rebase-sensitive-invariants.md.changelog.d/changed/cuda-strict-fp-every-kernel.md.docs/rebase-notes.md, ADR-1403.Reproducer
Known follow-ups
T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01:float_motion_cudais 1.36e-4 from the CPU on the 1080p checkerboards, above the 5e-5 tolerance, because the CPU keeps one fp32 running sum. Fix: reduce in the CPU's order on the device.T-CUDA-FLOAT-VIF-RESIDUAL-2026-10-01:float_vif_cudamatches on no frame and uses 3.8e-5 of its 5e-5; not isolated (devicelog2f, order of sums).T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01: the SYCL, HIP and Metalfloat_ms_ssimtwins keep the arithmetic the CUDA twin had (read from source, not run).vmaf_bench --gpu-profileis SYCL-only.bd061d95a(after feat(cuda): run float_ssim on the device at every scale, equal to the CPU #1677 rewrotessim_score.cu); the timings are frombd17c352f, where the six changed fatbins are the same; thefloat_ms_ssimstep table and the fp64 cost are from2c3acf1c9.T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01has a fix stacked on this branch (fix/cuda-float-motion-cpu-float-sum).