Repository navigation
fix(build): build every x86 SIMD library without FP contraction - #1706
Merged
Merged
Conversation
The x86 SIMD kernels are bit-exact twins of scalar references that live in baseline libraries, where no fused multiply-add exists. Most kernels finish the last n % lanes elements in plain C, and the two general libraries (x86_avx2, x86_avx512) were built with vmaf_fp_model_args only. Under icx that is -fp-model=precise, which implies -ffp-contract=on, so an icx build (every SYCL build is one) rounded those tails once where the reference rounds twice. In ssim_avx512.c that moved a score. Measured on an AVX-512 host (9950X3D, icx 2026.0) at --precision max: the CPU float_ms_ssim of an icx build differed from a GCC build and from its own scalar path by one fp32 unit in a per-scale mean, 7.7e-9 to 1.4e-8 in the score, on 4 of 104 frames (Netflix 576x324 pair, 1080p checkerboards, BBB 3840x2160). adm, vif and speed files were contracted too (42, 2 and 43 to 47 FMA instructions against 0, 0 and 3 with GCC), with no score difference measured. ADR-1415: - core/src/meson.build: both general libraries take vmaf_strict_fp_args, like the nine carve-out libraries. - test_strict_fp_compiler_args.py: both are strict targets. - test_ssim_x86_simd (new): AVX2 and AVX-512 precompute, variance and accumulate against transcriptions of the scalar reference, bit for bit, at element counts with and without a tail. Fails on the unfixed icx build at ssim_variance_avx512 n=1. - test_integer_adm_simd takes _simd_strict_fp_args: it compiles the scalar ADM kernels into its own translation unit and failed on icx builds since #1700 (the compiler's default fast model). GCC: the disassembly of all 28 objects of the two libraries is identical with and without the flag, so GCC builds compute what they did. icx: only explicit fmadd intrinsics and fmaf() calls remain; fast suite 207 of 207 (device suites left out), simd suite 21 of 21. GCC build against icx build, --backend cpu, 19 features on four fixtures: 15 bit-identical; psnr, psnr_hvs, ciede and speed_chroma still differ (at most 7.1e-15, 7.1e-15, 5.7e-12, 1.2e-6) because icx links Intel's math library. docs/state.md: T-ICX-SSIM-AVX512-FP-CONTRACT-2026-10-01, T-ICX-X86-SIMD-GENERAL-LIBS-CONTRACT-2026-10-01 and T-ICX-INTEGER-ADM-SIMD-TEST-FP-MODEL-2026-10-01 found and closed; T-ICX-LIBIMF-HOST-MATH-2026-10-01 opened.
lusoris
force-pushed
the
fix/icx-ssim-avx512-fp-contract
branch
from
October 1, 2026 14:50
7fb94c1 to
8d2e0c8
Compare
This was referenced Oct 1, 2026
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
…ale means are bit-identical ADR-1403 found four differences between the GPU float_ms_ssim twins and the CPU extractor and fixed the CUDA twin. The SYCL twin had all four: the decimate added sample * tap in two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums where iqa_convolve() adds fp32 products in fp64; l / c / s were fp32 quotients where the CPU divides fp64 numerators by fp32 denominators; the host combined unrounded means without fabs() on l and c. On an Arc A380 it matched the CPU on none of 104 frames (up to 2.98e-6). ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01. A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl already had: - sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients as exact fp32 pairs, int64 fixed-point terms, the exact host sum). Both SSIM twins include it. - integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and fabs() of all three on the host. - EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl. Measured on an Arc A380 (xe, Level Zero) at --precision max against a GCC build's --backend cpu, score, before and after: - Netflix 576x324, 48 frames: 6.9e-8 -> identical - 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical - BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame 1.1e-16 off through the host pow() of the icx build - enable_lcs (16 outputs) and enable_chroma outputs identical on every frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit Against the CPU extractor of its own icx binary the twin needs #1706 on an AVX-512 host (that extractor's ssim_avx512.c tails were contracted). Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15 paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and 1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms. No kernel uses scratch memory. Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18 outputs of 3 frames with == (53 and 54 of 54 differ on the old code); test_sycl_kernel_source_contract.py has seven planted regressions. docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed; T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP and Metal twins.
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
…ale means are bit-identical ADR-1403 found four differences between the GPU float_ms_ssim twins and the CPU extractor and fixed the CUDA twin. The SYCL twin had all four: the decimate added sample * tap in two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums where iqa_convolve() adds fp32 products in fp64; l / c / s were fp32 quotients where the CPU divides fp64 numerators by fp32 denominators; the host combined unrounded means without fabs() on l and c. On an Arc A380 it matched the CPU on none of 104 frames (up to 2.98e-6). ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01. A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl already had: - sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients as exact fp32 pairs, int64 fixed-point terms, the exact host sum). Both SSIM twins include it. - integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and fabs() of all three on the host. - EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl. Measured on an Arc A380 (xe, Level Zero) at --precision max against a GCC build's --backend cpu, score, before and after: - Netflix 576x324, 48 frames: 6.9e-8 -> identical - 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical - BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame 1.1e-16 off through the host pow() of the icx build - enable_lcs (16 outputs) and enable_chroma outputs identical on every frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit Against the CPU extractor of its own icx binary the twin needs #1706 on an AVX-512 host (that extractor's ssim_avx512.c tails were contracted). Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15 paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and 1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms. No kernel uses scratch memory. Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18 outputs of 3 frames with == (53 and 54 of 54 differ on the old code); test_sycl_kernel_source_contract.py has seven planted regressions. docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed; T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP and Metal twins.
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
…ale means are bit-identical ADR-1403 found four differences between the GPU float_ms_ssim twins and the CPU extractor and fixed the CUDA twin. The SYCL twin had all four: the decimate added sample * tap in two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums where iqa_convolve() adds fp32 products in fp64; l / c / s were fp32 quotients where the CPU divides fp64 numerators by fp32 denominators; the host combined unrounded means without fabs() on l and c. On an Arc A380 it matched the CPU on none of 104 frames (up to 2.98e-6). ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01. A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl already had: - sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients as exact fp32 pairs, int64 fixed-point terms, the exact host sum). Both SSIM twins include it. - integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and fabs() of all three on the host. - EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl. Measured on an Arc A380 (xe, Level Zero) at --precision max against a GCC build's --backend cpu, score, before and after: - Netflix 576x324, 48 frames: 6.9e-8 -> identical - 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical - BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame 1.1e-16 off through the host pow() of the icx build - enable_lcs (16 outputs) and enable_chroma outputs identical on every frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit Against the CPU extractor of its own icx binary the twin needs #1706 on an AVX-512 host (that extractor's ssim_avx512.c tails were contracted). Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15 paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and 1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms. No kernel uses scratch memory. Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18 outputs of 3 frames with == (53 and 54 of 54 differ on the old code); test_sycl_kernel_source_contract.py has seven planted regressions. docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed; T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP and Metal twins.
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
…ale means are bit-identical (#1709) * fix(sycl): make float_ms_ssim_sycl the CPU's arithmetic so its per-scale means are bit-identical ADR-1403 found four differences between the GPU float_ms_ssim twins and the CPU extractor and fixed the CUDA twin. The SYCL twin had all four: the decimate added sample * tap in two roundings where the reference fuses each tap; the Gaussian window sums were fp32 running sums where iqa_convolve() adds fp32 products in fp64; l / c / s were fp32 quotients where the CPU divides fp64 numerators by fp32 denominators; the host combined unrounded means without fabs() on l and c. On an Arc A380 it matched the CPU on none of 104 frames (up to 2.98e-6). ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01. A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl already had: - sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients as exact fp32 pairs, int64 fixed-point terms, the exact host sum). Both SSIM twins include it. - integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and fabs() of all three on the host. - EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl. Measured on an Arc A380 (xe, Level Zero) at --precision max against a GCC build's --backend cpu, score, before and after: - Netflix 576x324, 48 frames: 6.9e-8 -> identical - 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical - BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame 1.1e-16 off through the host pow() of the icx build - enable_lcs (16 outputs) and enable_chroma outputs identical on every frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit Against the CPU extractor of its own icx binary the twin needs #1706 on an AVX-512 host (that extractor's ssim_avx512.c tails were contracted). Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15 paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and 1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms. No kernel uses scratch memory. Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18 outputs of 3 frames with == (53 and 54 of 54 differ on the old code); test_sycl_kernel_source_contract.py has seven planted regressions. docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed; T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP and Metal twins.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every x86 SIMD library is now built without FP contraction. On an AVX-512 host, the CPU
float_ms_ssimof an Intel-compiler build then gives the same values as a GCC build and as its own scalar path. Every SYCL build is an icx build, so this is the CPU extractor a SYCL twin gets compared with.The x86 SIMD kernels are bit-exact twins of scalar references that live in baseline libraries, where no fused multiply-add exists. Most kernels finish the last
n % laneselements in plain C. The two general libraries (x86_avx2,x86_avx512) were built withvmaf_fp_model_argsonly; under icx that is-fp-model=precise, which implies-ffp-contract=on, so those tails were rounded once where the reference rounds twice. Nine files had been given their own strict library one at a time;ssim_avx512.cwas missed when its AVX2 sibling got one.Found while making
float_ms_ssim_syclexact: the twin matched a GCC CPU build on every frame and the CPU extractor of its own icx build on 100 of 104.What changed
core/src/meson.build:x86_avx2_static_libandx86_avx512_static_libtakevmaf_strict_fp_args. No source file changes.core/test/test_strict_fp_compiler_args.py: both libraries are listed inSTRICT_TARGETS.core/test/test_ssim_x86_simd.c(new, fast and simd suites): AVX2 and AVX-512ssim_precompute,ssim_varianceandssim_accumulateagainst transcriptions of the scalar reference, bit for bit, at 24 element counts with and without a tail. The x86 sibling oftest_ssim_neon.core/test/meson.build:test_integer_adm_simdtakes_simd_strict_fp_args. It compiles the scalar ADM kernels into its own translation unit and has failed on icx builds since fix(adm): keep the scale-0 masking centre tap in int32 (ADR-1402) #1700 (adm_cm_avx2 12x10 dense case 1 aim=1 p_norm=3: 0x1.06965p+3, scalar 0x1.06964ep+3), because that unit was built with icx's default fast model.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. (clang-format on the new test; the commit hooks.)python3 scripts/ci/run_meson_test.py -- -C build. (--suite=faston a GCC build: 207 of 207; on an icx build without the device suites: 207 of 207;--suite=simdon the icx build: 21 of 21, 20 of 21 onmaster408dcaad5.)/cross-backend-diffand the worst ULP is ≤ 2. (GCC build against icx build, CPU only; tables below.).c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. (Not a breaking change.)docs/adr/_index_fragments/<NNNN-slug>.md.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-ICX-SSIM-AVX512-FP-CONTRACT-2026-10-01,T-ICX-X86-SIMD-GENERAL-LIBS-CONTRACT-2026-10-01andT-ICX-INTEGER-ADM-SIMD-TEST-FP-MODEL-2026-10-01added under Recently closed with the evidence;T-ICX-LIBIMF-HOST-MATH-2026-10-01opened.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
ryzen-4090-arc(9950X3D, AVX-512), icx 2026.0 against GCC,--backend cpu --precision max.float_ms_ssim=enable_lcs=true, frames on which any of the 16 outputs differs between the two builds, and the largest difference in the score:Each differing frame has one per-scale mean (
float_ms_ssim_s_scale4orfloat_ms_ssim_c_scale3) off by one fp32 unit, 5.96e-8. Before the change the builds agreed with--cpumask 48(AVX-512 off) and with--cpumask 4294967295(scalar), and the GCC build agreed with its own scalar path, which isolates the AVX-512 object of the icx build.Fused multiply-add instructions per object (
objdump -d | grep -c 'vfn\?m\(add\|sub\)'), icx before, icx after, GCC (before and after):GCC: the disassembly of all 28 objects of the two libraries is identical with and without the flag.
All 19 features that have a GPU twin, GCC build against icx build after the change, on the Netflix pair, both checkerboard pairs and 20 frames of BBB 3840x2160: 15 are bit-identical (
vif,adm,motion,motion_v2,ssim,float_ssim,float_ms_ssim,float_psnr,float_moment,float_motion,float_vif,float_adm,cambi,ssimulacra2,speed_temporal).psnrandpsnr_hvs(7.1e-15),ciede(5.7e-12) andspeed_chroma(1.2e-6) differ with and without SIMD: icx links Intel'slibimfahead of glibc'slibm.Deep-dive deliverables (ADR-0108)
docs/state.mdrows.## Alternatives considered.AGENTS.mdinvariant note —core/src/feature/x86/AGENTS.md, "Strict FP for every x86 SIMD library" and the SSIM accumulate twin group.changelog.d/fixed/x86-simd-libraries-strict-fp.md.docs/rebase-notes.md, ADR-1415.Reproducer
Known follow-ups
T-ICX-LIBIMF-HOST-MATH-2026-10-01(opened here):psnr,psnr_hvs,ciedeandspeed_chromaof an icx build differ from a GCC build through the math library. The parity gate runs one binary and does not see it.float_ms_ssim_syclitself (the SYCL part ofT-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01) follows in its own PR; its exact gate cell compares the twin with the CPU extractor of the same icx binary and needs this fix on an AVX-512 host.