Repository navigation
fix(cuda): reproduce the CPU's running float sum in psnr_hvs_cuda - #1666
Merged
Merged
Conversation
lusoris
force-pushed
the
fix/psnr-hvs-twins-cpu-float-sum
branch
from
October 1, 2026 07:27
4aac3fb to
f2b637e
Compare
lusoris
force-pushed
the
fix/psnr-hvs-twins-cpu-float-sum
branch
from
October 1, 2026 09:18
f2b637e to
9c300cd
Compare
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before, it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the ADR-1361 parity tolerance (3.34e-3 dB). calc_psnrhvs() adds every masked coefficient error of a plane into one running float, so its result depends on the order of the additions. The twin summed the 64 terms of a block on the device and the blocks on the host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation; a double accumulator on the CPU was ruled out because it moves a Netflix golden value past places=4. - The kernel stores the 64 terms of every block, computed in the CPU's arithmetic: masking table and threshold in double, integer coefficient difference, fatbin built with --fmad=false. - psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms the combined score and the dB value with the CPU's expressions. It is built with the strict floating-point arguments of the scalar reference. - The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0 at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361. - test_cuda_psnr_hvs_parity asserts equality on all four outputs, including two 3840x2160 cases that fail on master by 1.6e-2 dB. test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are device-free. Measured on an RTX 4090 at --precision max: every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores are unchanged. The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of 2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09), and 256 bytes of readback per block. Tuning is tracked as T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins are unchanged until #1658 and #1657 land (T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01). Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris
force-pushed
the
fix/psnr-hvs-twins-cpu-float-sum
branch
from
October 1, 2026 09:37
9c300cd to
bd51dfb
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
psnr_hvs_cudanow returns the CPU extractor's scores bit for bit. On master it is up to 1.66e-2 dB from--backend cpuat 3840x2160, beyond the ADR-1361 tolerance (3.34e-3 dB), although it is the more accurate of the two.The cause is the reference:
calc_psnrhvs()adds every masked coefficient error of a plane into one runningfloat(10.8 million terms for 3840x2160 luma), so its value depends on the order of the additions, and the twin summed per block. Per maintainer decision (2026-10-01, ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation instead of widening the tolerance; adoubleaccumulator on the CPU was ruled out because it moves a Netflix golden value pastplaces=4. The kernel now stores the 64 terms of every block in the CPU's arithmetic and the host adds them in the CPU's order. ClosesT-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.This costs throughput (numbers below), which the decision accepts: correctness first, tuning afterwards.
Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. The pre-commit hook set passed on the whole diff; the touched C files (integer_psnr_hvs_cuda.c,psnr_hvs_score.c,test_cuda_psnr_hvs_parity.c,test_psnr_hvs_score.c) measure 0 clang-tidy findings on the CUDA lane, which is their baseline, so no tidy baseline changes.python3 scripts/ci/run_meson_test.py -- -C build --suite=faston a CPU + CUDA build (RTX 4090), includingtest_cuda_psnr_hvs_parity,test_cuda_psnr_hvs_parity_large,test_psnr_hvs_scoreandtest_psnr_hvs_twin_exact_sum_contract./cross-backend-diffand the worst ULP is ≤ 2. The worst difference is 0: every output of every frame is bit-identical to the CPU (table below)..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30moved to Recently closed;T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01,T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01andT-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01opened (RC3).Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.No CPU extractor code changes.
--backend cpuoutput of this branch equals master's on all 210 frames of the fixtures below.Cross-backend numerical results
RTX 4090 (driver 615.71.09, CUDA 13.4),
--precision max,psnr_hvson--backend cpuagainstpsnr_hvs_cuda. Largest absolute difference in dB over the frames; "before" is masterc7f28317f, "after" this branch (re-measured on headf2b637ed4, based on master5593a70ae).The 10-bit 1920x1080 and 3840x2160 fixtures are ffmpeg conversions of the 8-bit ones (the BBB 3840x2160 distorted side with
noise=alls=6:allf=t); BBB 1920x1080 is a bicubic downscale. The parity gate on all 200 frames of BBB 3840x2160 reports a maximum difference of 0 with tolerance 0.Each kernel change is needed. With the CPU-order sum in place, restoring one of the three remaining differences breaks identity on a few frames per fixture: nvcc contraction (up to 8.5e-7 dB), a float square root in the masking threshold (5.0e-7), a float masking table (7.4e-7). Details in Research-1397.
Performance (if
perforfeat)Slower, as expected. ms per frame on the RTX 4090, median of three runs of (t(N) − t(2)) / (N − 2), N = 960 at 576x324 (the Netflix pair concatenated 20 times), 102 at 1920x1080 and 3840x2160; other sessions shared the CPU (load average 5 to 11), the GPU was held under the device lock.
Of the 3840x2160 figure, 6.4 ms is the host sum (16.2 million additions in one dependency chain); the rest is the readback, 64.8 MB per frame instead of 1.0 MB. 15 to 32 % of a plane's terms are nonzero, so compacting them is the first tuning step;
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01tracks it.Deep-dive deliverables (ADR-0108)
docs/research/1397-psnr-hvs-twins-cpu-float-sum.md.## Alternatives considered.AGENTS.mdinvariant note —core/src/feature/cuda/AGENTS.md(the exact-sum contract of the kernel and host),core/src/feature/AGENTS.md(psnr_hvs_score.c),scripts/ci/AGENTS.md(EXACT_TWINS), and the index entry indocs/development/rebase-sensitive-invariants.md.changelog.d/fixed/cuda-psnr-hvs-cpu-float-sum.md.docs/rebase-notes.md, "ADR-1397 — psnr_hvs_cuda returns the CPU's scores bit for bit".Reproducer
ninja -C build test/test_cuda_psnr_hvs_parity test/test_psnr_hvs_score ./build/test/test_cuda_psnr_hvs_parity && ./build/test/test_psnr_hvs_score python3 core/test/test_psnr_hvs_twin_exact_sum_contract.py python3 scripts/ci/cross_backend_parity_gate.py --vmaf-binary build/tools/vmaf \ --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --features psnr_hvs --backends cpu cudaOn master the first command fails (
test_psnr_hvs_cpu_cuda_identical: 3.8e-6 dB at 256x144; the 3840x2160 cases differ by 1.63e-2 and 1.66e-2 dB) and the gate reportsmax_abs_diff=8.370e-05 FAILunder the new contract.Known follow-ups
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01: the exact twin is slower than sixteen CPU threads at every size measured.T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01:psnr_hvs_hipis unchanged here. perf(hip): upload native samples and convert on device in psnr_hvs #1658 rewrites that twin and was open when this was written; the same change goes on top of it and is verified on the gfx1036.T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01:psnr_hvs_syclis unchanged here. perf(sycl): make psnr_hvs kernel scratch-free on Arc A380 #1657 rewrites its kernel and was open when this was written. The SYCL kernel is fp64-free (ADR-0220), so the threshold's double product and square root need fp32-pair arithmetic there. Until then its gate cells keep the ADR-1361 tolerance.docs/metrics/psnr-hvs.mdanddocs/metrics/features.mdlistedenable_chromaas defaultfalseor absent,psnr_hvsas the luma score and the input as 8-bit only.Breaking changes / migration
None in the API.
psnr_hvs_cudascores change in their last digits (by up to 8.4e-5 dB at 576x324 and 1.7e-2 dB at 3840x2160) to the values--backend cpureturns.