Repository navigation
fix(cuda): add the float_motion SAD in the CPU's order so the twin is bit-identical - #1699
Merged
Merged
Conversation
lusoris
force-pushed
the
fix/cuda-float-motion-cpu-float-sum
branch
3 times, most recently
from
October 1, 2026 13:08
8cf21b7 to
1e52ad6
Compare
lusoris
marked this pull request as ready for review
October 1, 2026 13:08
… bit-identical float_motion.c adds the absolute differences of a row into one float, the row sums into a second float, and divides by the pixel count in float. Those running sums round at every step, so the score depends on the order of the additions. float_motion_cuda summed each 16x16 block on the device and the blocks in double on the host: closer to the exact sum, and up to 1.36e-4 from the CPU on 1920x1080 checkerboards, above the 5e-5 cross-backend tolerance (ADR-0214). ADR-1409: - float_motion_score.cu: the blur kernels write the blurred plane only. A new float_motion_row_sad kernel runs one thread per row and adds |cur[j] - prev[j]| left to right into one fp32 accumulator; the readback is one float per row. - float_motion_sad.h (new, backend neutral): the tail of compute_motion_simd(), one float over the rows and a float division. float_motion_cuda.c keeps no sum of its own. - EXACT_TWINS lists float_motion: cuda, so the parity gate compares the cell with tolerance 0 at --precision max. The blur was already the CPU's once every fatbin was built without FMA contraction (ADR-1403); this change depends on that flag. Measured on an RTX 4090 against --backend cpu at --precision max, motion / motion2 / motion3, before and after: - Netflix 576x324, 48 frames: 3.1e-6 -> identical on every frame - 1920x1080 checkerboards, 3 frames each: 1.36e-4 -> identical - BBB 3840x2160, 200 frames: 2.4e-5 -> identical - also identical: 1280x720 10-bit, a 573x163 4:4:4 crop, and the Netflix pair with the fps-weight, cap and blend options set Time per BBB 3840x2160 frame, 25 paired 200-frame runs: 3.00 ms before, 2.98 after (paired difference median -0.12 ms, quartiles -0.57 to +0.25; the untouched float_psnr_cuda read 2.84 and 2.80). Tests: test_cuda_float_motion_parity and its 960x540 variant compare with == at 8, 10 and 12 bits and fail on the old twin by 2.0e-5 and 6.2e-5; test_float_motion_sad pins the host helper without a device; test_cuda_kernel_source_contract.py has five planted regressions. docs/state.md: T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01 closed; T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01 opened for the SYCL, HIP and Metal twins, which still sum per block.
lusoris
force-pushed
the
fix/cuda-float-motion-cpu-float-sum
branch
from
October 1, 2026 13:34
1e52ad6 to
636c057
Compare
This was referenced Oct 1, 2026
3 of 8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
float_motion_cudanow returns the CPUfloat_motionscores bit for bit:motion,motion2andmotion3are identical on every frame measured, with no measurable change in time.Builds on #1695 (every CUDA fatbin without FMA contraction, landed as
691ca5677): the blur is the CPU's only with that flag.Closes
T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01, which #1695 found. The CPU extractor adds the absolute differences of a row into onefloat, the row sums into a secondfloat, and divides infloat. Those running sums round at every step, so the score depends on the order of the additions. The twin summed each 16x16 block on the device and the blocks indoubleon the host, which is closer to the exact mean and up to 1.36e-4 from the CPU on 1920x1080 checkerboards, above the 5e-5 tolerance of ADR-0214. The gate runs the Netflix pair only (3.1e-6), so it passed.What changed
core/src/feature/cuda/float_motion/float_motion_score.cu: the two blur kernels write the blurred plane only. A newfloat_motion_row_sadkernel runs one thread per row and adds|cur[j] - prev[j]|left to right into one fp32 accumulator; the readback is one float per row.core/src/feature/float_motion_sad.h(new, backend neutral): the tail ofcompute_motion_simd(), onefloatover the rows and afloatdivision.float_motion_cuda.ckeeps no sum of its own.scripts/ci/cross_backend_calibration.py:EXACT_TWINSlistsfloat_motion:cuda, so the parity gate compares that cell with tolerance 0 at--precision max.T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01, opened here).Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. (clang-format and clang-tidy on the touched files: 0 warnings in them; the commit hooks.)python3 scripts/ci/run_meson_test.py -- -C build. (--suite=faston a CUDA build, RTX 4090: 264 of 264 on the current tip.)/cross-backend-diffand the worst ULP is ≤ 2. (0; table below.).c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. (Not a breaking change.)docs/adr/_index_fragments/<NNNN-slug>.md.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-CUDA-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01moved to Recently closed with the evidence;T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01opened for the other three twins.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
RTX 4090, nvcc 13.4,
--precision max,float_motionon--backend cpuagainstfloat_motion_cuda. Frames identical and largest absolute difference overmotion,motion2andmotion3. "Before" ismaster691ca5677(#1695).The one frame that matched before is frame 0, whose
motionandmotion2are 0 by definition.scripts/ci/cross_backend_parity_gate.py --features float_motion --backends cpu cudareportstol=0.0e+00 (exact:ADR-1397) max_abs_diff=0.000e+00 OKon all four fixtures (BBB: all 200 frames).Performance (if
perforfeat)Not a performance change; measured because the row kernel uses one thread per row.
vmaftool, BBB 3840x2160, 200 frames, 25 alternating before/after pairs, host load 12 to 14:Deep-dive deliverables (ADR-0108)
docs/state.mdrow.## Alternatives considered.AGENTS.mdinvariant note —core/src/feature/cuda/AGENTS.md(one thread per row, plain loop, host helper, no other reduction),scripts/ci/AGENTS.md(EXACT_TWINS), anddocs/development/rebase-sensitive-invariants.md.changelog.d/fixed/cuda-float-motion-cpu-float-sum.md.docs/rebase-notes.md, ADR-1409.Reproducer
Known follow-ups
float_motion_sycl,float_motion_hip,float_motion_metalstill sum per block:T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01. The host helper is backend neutral; each twin needs the row kernel; the blur is already built without contraction on SYCL (ADR-1367) and HIP (fix(hip): compile every HIP kernel with contraction off through one shared flag list #1694).float_motioncell comparesmotionandmotion2;motion3is covered bytest_cuda_float_motion_parity, because the SYCL twin does not emit it yet.