Repository navigation
fix(sycl): make float_ms_ssim_sycl the CPU's arithmetic so its per-scale means are bit-identical - #1709
Merged
Conversation
lusoris
force-pushed
the
fix/sycl-float-ms-ssim-cpu-arithmetic
branch
2 times, most recently
from
October 1, 2026 16:38
b6e8582 to
651e31f
Compare
…ale means are bit-identical ADR-1403 found four differences between the GPU float_ms_ssim twins and the CPU extractor and fixed the CUDA twin. The SYCL twin had all four: the decimate added sample * tap in two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums where iqa_convolve() adds fp32 products in fp64; l / c / s were fp32 quotients where the CPU divides fp64 numerators by fp32 denominators; the host combined unrounded means without fabs() on l and c. On an Arc A380 it matched the CPU on none of 104 frames (up to 2.98e-6). ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01. A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl already had: - sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients as exact fp32 pairs, int64 fixed-point terms, the exact host sum). Both SSIM twins include it. - integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and fabs() of all three on the host. - EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl. Measured on an Arc A380 (xe, Level Zero) at --precision max against a GCC build's --backend cpu, score, before and after: - Netflix 576x324, 48 frames: 6.9e-8 -> identical - 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical - BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame 1.1e-16 off through the host pow() of the icx build - enable_lcs (16 outputs) and enable_chroma outputs identical on every frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit Against the CPU extractor of its own icx binary the twin needs #1706 on an AVX-512 host (that extractor's ssim_avx512.c tails were contracted). Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15 paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and 1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms. No kernel uses scratch memory. Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18 outputs of 3 frames with == (53 and 54 of 54 differ on the old code); test_sycl_kernel_source_contract.py has seven planted regressions. docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed; T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP and Metal twins.
lusoris
force-pushed
the
fix/sycl-float-ms-ssim-cpu-arithmetic
branch
from
October 1, 2026 16:52
651e31f to
061d0e1
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
float_ms_ssim_syclnow computes the CPU extractor's arithmetic. On an Arc A380 every per-scale mean of every measured frame is the CPU's bit for bit, withenable_lcsandenable_chroma. Before, the twin matched the CPU on none of 104 frames (up to 2.98e-6). It costs time: 42.6 ms per 3840x2160 frame against 31.4.This is the SYCL part of
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01, which #1695 opened when it fixed the CUDA twin. The SYCL twin had the same four differences: the decimate addedsample * tapin two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums whereiqa_convolve()adds fp32 products indouble;l/c/swere fp32 quotients where the CPU dividesdoublenumerators by fp32 denominators; the host combined unrounded means withoutfabs()onlandc.A SYCL kernel may not use
double(ADR-0220), so the CUDA fix does not carry over as written.float_ssim_syclalready had the same arithmetic without fp64, and this PR shares it instead of copying it.What changed
core/src/feature/sycl/sycl_ssim_terms.h(new): the per-pixel SSIM arithmetic, moved out ofinteger_ssim_sycl.cppunchanged: window sums and thel/cquotients as exact fp32 pairs, each term to int64 units of 2^-52, the exact host sum. Both SSIM twins include it.core/src/feature/sycl/integer_ms_ssim_sycl.cpp:sycl::fma()per decimate tap, pair window sums,ssim_terms()forl/c/s, int64 group partials, fp32 means andfabs()of all three on the host.scripts/ci/cross_backend_calibration.py:EXACT_TWINSlistsfloat_ms_ssimandfloat_ms_ssim_lcs:sycl.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. (clang-format on the touched files; the SYCL clang-tidy wrapper oninteger_ms_ssim_sycl.cpp,integer_ssim_sycl.cpp, the new header and the test: 0 warnings in them, nine older ones in the test file removed; the commit hooks.)python3 scripts/ci/run_meson_test.py -- -C build. (--suite=faston a GCC CPU build: 208 of 208. On the icx SYCL build:--suite syclon the Arc A380 58 of 60, and 207 of 208 without a device. The three failures fail onmastertoo:test_sycl_float_adm_parityand its large variant becausefloat_adm_syclstill uses scratch memory on this GPU (ADR-1395), andtest_integer_adm_simdon icx builds since fix(adm): keep the scale-0 masking centre tap in int32 (ADR-1402) #1700, fixed by fix(build): build every x86 SIMD library without FP contraction #1706.)/cross-backend-diffand the worst ULP is ≤ 2. (0 on the device; tables below.).c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. (Not a breaking change.)docs/adr/_index_fragments/<NNNN-slug>.md.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01added under Recently closed with the evidence;T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01now names Metal only and stays open (HIP landed in fix(hip): make float_ms_ssim_hip the CPU's arithmetic so the twin is bit-identical #1710).Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
Arc A380 (dg2-g11, xe driver, Level Zero, icx 2026.0),
--precision max,float_ms_ssimon--backend cpuof a GCC build againstfloat_ms_ssim_sycl. Frames identical and largest absolute difference of the score. "Before" ismasterdff31c445.The device arithmetic matches on every frame: all per-scale means are identical. The three places that are not zero are host math calls. The twin's host tail calls the
pow()andlog10()of its own icx build, which links Intel'slibimf, and the reference here is a GCC build (T-ICX-LIBIMF-HOST-MATH-2026-10-01in #1706).Against the CPU extractor of the same icx binary, which is what the gate compares,
scripts/ci/cross_backend_parity_gate.py --features float_ms_ssim float_ms_ssim_lcs --backends cpu sycl:The failures are the CPU side, not the twin: on this AVX-512 host the icx build contracted the scalar tails of
ssim_avx512.c, so the CPU extractor itself was one fp32 unit off in a per-scale mean on 4 of 104 frames. #1706 fixes that; with its two flag changes applied to this tree the cell is 0 on all four fixtures. The twin equals the scalar CPU path of its own binary on every frame either way.Bit-identical in practice, not by construction: the CPU forms
landcindoubleand adds them into a runningdoublein raster order; the twin carries them as pairs good to about 2^-46 and adds them exactly. The fp32 rounding of each per-scale mean absorbs the difference.test_sycl_kernel_scratchon the A380 audits 110 kernels; none of this PR's uses scratch memory and the ratchet list is unchanged.Performance (if
perforfeat)Not a performance change, and it costs time.
vmaftool on the Arc A380, alternating before/after runs, medians, host load 10 to 20:ms per frame. A cheaper tap accumulation (one two-sum per tap instead of a full pair addition) gave the same values and 42.2 ms, so the taps are not where the time goes, and the shared arithmetic was left as
float_ssim_syclhad it. Both passes read every tap from device memory (22 and 55 samples per pixel), as before.Deep-dive deliverables (ADR-0108)
docs/state.mdrow.## Alternatives considered.AGENTS.mdinvariant note —core/src/feature/sycl/AGENTS.md(the four rules, the shared header),scripts/ci/AGENTS.md(EXACT_TWINS), anddocs/development/rebase-sensitive-invariants.md.changelog.d/fixed/sycl-float-ms-ssim-cpu-arithmetic.md.docs/rebase-notes.md, ADR-1414.Reproducer
Known follow-ups
float_ms_ssim_metalkeeps the old arithmetic:T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01(integer_ms_ssim_hiplanded in fix(hip): make float_ms_ssim_hip the CPU's arithmetic so the twin is bit-identical #1710).float_ssim_sycldoes in its horizontal pass, is tuning left for after the exactness work.float_ms_ssim_cudais bit-identical too (ADR-1403) and not listed inEXACT_TWINS; this PR does not touch that entry.