Skip to content

fix(build): build every x86 SIMD library without FP contraction - #1706

Merged
lusoris merged 2 commits into
masterfrom
fix/icx-ssim-avx512-fp-contract
Oct 1, 2026
Merged

lusoris merged 2 commits into
masterfrom
fix/icx-ssim-avx512-fp-contract

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Every x86 SIMD library is now built without FP contraction. On an AVX-512 host, the CPU float_ms_ssim of an Intel-compiler build then gives the same values as a GCC build and as its own scalar path. Every SYCL build is an icx build, so this is the CPU extractor a SYCL twin gets compared with.

The x86 SIMD kernels are bit-exact twins of scalar references that live in baseline libraries, where no fused multiply-add exists. Most kernels finish the last n % lanes elements in plain C. The two general libraries (x86_avx2, x86_avx512) were built with vmaf_fp_model_args only; under icx that is -fp-model=precise, which implies -ffp-contract=on, so those tails were rounded once where the reference rounds twice. Nine files had been given their own strict library one at a time; ssim_avx512.c was missed when its AVX2 sibling got one.

Found while making float_ms_ssim_sycl exact: the twin matched a GCC CPU build on every frame and the CPU extractor of its own icx build on 100 of 104.

What changed

  • core/src/meson.build: x86_avx2_static_lib and x86_avx512_static_lib take vmaf_strict_fp_args. No source file changes.
  • core/test/test_strict_fp_compiler_args.py: both libraries are listed in STRICT_TARGETS.
  • core/test/test_ssim_x86_simd.c (new, fast and simd suites): AVX2 and AVX-512 ssim_precompute, ssim_variance and ssim_accumulate against transcriptions of the scalar reference, bit for bit, at 24 element counts with and without a tail. The x86 sibling of test_ssim_neon.
  • core/test/meson.build: test_integer_adm_simd takes _simd_strict_fp_args. It compiles the scalar ADM kernels into its own translation unit and has failed on icx builds since fix(adm): keep the scale-0 masking centre tap in int32 (ADR-1402) #1700 (adm_cm_avx2 12x10 dense case 1 aim=1 p_norm=3: 0x1.06965p+3, scalar 0x1.06964ep+3), because that unit was built with icx's default fast model.
  • ADR-1415 records the policy.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. (clang-format on the new test; the commit hooks.)
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. (--suite=fast on a GCC build: 207 of 207; on an icx build without the device suites: 207 of 207; --suite=simd on the icx build: 21 of 21, 20 of 21 on master 408dcaad5.)
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. (GCC build against icx build, CPU only; tables below.)
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. (Build flags only; no kernel source changes.)
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. (Not a breaking change.)
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-ICX-SSIM-AVX512-FP-CONTRACT-2026-10-01, T-ICX-X86-SIMD-GENERAL-LIBS-CONTRACT-2026-10-01 and T-ICX-INTEGER-ADM-SIMD-TEST-FP-MODEL-2026-10-01 added under Recently closed with the evidence; T-ICX-LIBIMF-HOST-MATH-2026-10-01 opened.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. (None changes: the gate runs a GCC or clang build, and GCC's objects are identical with and without the flag.)

Cross-backend numerical results

ryzen-4090-arc (9950X3D, AVX-512), icx 2026.0 against GCC, --backend cpu --precision max.

float_ms_ssim=enable_lcs=true, frames on which any of the 16 outputs differs between the two builds, and the largest difference in the score:

fixture                         frames  before             after
Netflix 576x324                 48      1,  7.7e-9         0
checkerboard 1 px, 1920x1080    3       0                  0
checkerboard 10 px, 1920x1080   3       1,  7.9e-9         0
BBB 3840x2160                   50      2,  1.4e-8         0

Each differing frame has one per-scale mean (float_ms_ssim_s_scale4 or float_ms_ssim_c_scale3) off by one fp32 unit, 5.96e-8. Before the change the builds agreed with --cpumask 48 (AVX-512 off) and with --cpumask 4294967295 (scalar), and the GCC build agreed with its own scalar path, which isolates the AVX-512 object of the icx build.

Fused multiply-add instructions per object (objdump -d | grep -c 'vfn\?m\(add\|sub\)'), icx before, icx after, GCC (before and after):

object                         icx before  icx after  GCC
ssim_avx512.c                  24          0          0
adm_avx2.c, adm_avx512.c       42          0          0
vif_avx2.c, vif_statistic_avx2.c  2        0          0
speed_avx2.c                   43          21         3     explicit fmaf() calls
speed_avx512.c                 47          21         3     explicit fmaf() calls
common/convolution_avx512.c    11          11         90    explicit fmadd intrinsics
ms_ssim_decimate_avx512.c      54          54         109   explicit fmadd intrinsics

GCC: the disassembly of all 28 objects of the two libraries is identical with and without the flag.

All 19 features that have a GPU twin, GCC build against icx build after the change, on the Netflix pair, both checkerboard pairs and 20 frames of BBB 3840x2160: 15 are bit-identical (vif, adm, motion, motion_v2, ssim, float_ssim, float_ms_ssim, float_psnr, float_moment, float_motion, float_vif, float_adm, cambi, ssimulacra2, speed_temporal). psnr and psnr_hvs (7.1e-15), ciede (5.7e-12) and speed_chroma (1.2e-6) differ with and without SIMD: icx links Intel's libimf ahead of glibc's libm.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: the cause is a compiler flag, and the measurements are in ADR-1415 and the docs/state.md rows.
  • Decision matrix — ADR-1415 ## Alternatives considered.
  • AGENTS.md invariant note — core/src/feature/x86/AGENTS.md, "Strict FP for every x86 SIMD library" and the SSIM accumulate twin group.
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/fixed/x86-simd-libraries-strict-fp.md.
  • Rebase note — docs/rebase-notes.md, ADR-1415.

Reproducer

source /opt/intel/oneapi/setvars.sh
CC=icx CXX=icpx meson setup build-icx core -Denable_sycl=true -Denable_cuda=false --buildtype=release
ninja -C build-icx
# Fails on master at ssim_variance_avx512 n=1; needs an AVX-512 host.
build-icx/test/test_ssim_x86_simd
# Fails on master since #1700 (icx only).
build-icx/test/test_integer_adm_simd
# 24 on master, 0 here.
objdump -d "$(find build-icx/src -name '*ssim_avx512.c.o')" | grep -c 'vfn\?m\(add\|sub\)'
python3 core/test/test_strict_fp_compiler_args.py

Known follow-ups

  • T-ICX-LIBIMF-HOST-MATH-2026-10-01 (opened here): psnr, psnr_hvs, ciede and speed_chroma of an icx build differ from a GCC build through the math library. The parity gate runs one binary and does not see it.
  • The nine carve-out libraries now carry the same flags as the general ones and can be folded into them in a cleanup.
  • float_ms_ssim_sycl itself (the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01) follows in its own PR; its exact gate cell compares the twin with the CPU extractor of the same icx binary and needs this fix on an AVX-512 host.

The x86 SIMD kernels are bit-exact twins of scalar references that live
in baseline libraries, where no fused multiply-add exists. Most kernels
finish the last n % lanes elements in plain C, and the two general
libraries (x86_avx2, x86_avx512) were built with vmaf_fp_model_args
only. Under icx that is -fp-model=precise, which implies
-ffp-contract=on, so an icx build (every SYCL build is one) rounded
those tails once where the reference rounds twice.

In ssim_avx512.c that moved a score. Measured on an AVX-512 host
(9950X3D, icx 2026.0) at --precision max: the CPU float_ms_ssim of an
icx build differed from a GCC build and from its own scalar path by
one fp32 unit in a per-scale mean, 7.7e-9 to 1.4e-8 in the score, on 4
of 104 frames (Netflix 576x324 pair, 1080p checkerboards, BBB
3840x2160). adm, vif and speed files were contracted too (42, 2 and 43
to 47 FMA instructions against 0, 0 and 3 with GCC), with no score
difference measured.

ADR-1415:

- core/src/meson.build: both general libraries take
  vmaf_strict_fp_args, like the nine carve-out libraries.
- test_strict_fp_compiler_args.py: both are strict targets.
- test_ssim_x86_simd (new): AVX2 and AVX-512 precompute, variance and
  accumulate against transcriptions of the scalar reference, bit for
  bit, at element counts with and without a tail. Fails on the unfixed
  icx build at ssim_variance_avx512 n=1.
- test_integer_adm_simd takes _simd_strict_fp_args: it compiles the
  scalar ADM kernels into its own translation unit and failed on icx
  builds since #1700 (the compiler's default fast model).

GCC: the disassembly of all 28 objects of the two libraries is
identical with and without the flag, so GCC builds compute what they
did. icx: only explicit fmadd intrinsics and fmaf() calls remain; fast
suite 207 of 207 (device suites left out), simd suite 21 of 21. GCC
build against icx build, --backend cpu, 19 features on four fixtures:
15 bit-identical; psnr, psnr_hvs, ciede and speed_chroma still differ
(at most 7.1e-15, 7.1e-15, 5.7e-12, 1.2e-6) because icx links Intel's
math library.

docs/state.md: T-ICX-SSIM-AVX512-FP-CONTRACT-2026-10-01,
T-ICX-X86-SIMD-GENERAL-LIBS-CONTRACT-2026-10-01 and
T-ICX-INTEGER-ADM-SIMD-TEST-FP-MODEL-2026-10-01 found and closed;
T-ICX-LIBIMF-HOST-MATH-2026-10-01 opened.
@lusoris
lusoris force-pushed the fix/icx-ssim-avx512-fp-contract branch from 7fb94c1 to 8d2e0c8 Compare October 1, 2026 14:50
@lusoris
lusoris merged commit fb52219 into master Oct 1, 2026
66 of 76 checks passed
@lusoris
lusoris deleted the fix/icx-ssim-avx512-fp-contract branch October 1, 2026 14:50
lusoris added a commit that referenced this pull request Oct 1, 2026
…ale means are bit-identical

ADR-1403 found four differences between the GPU float_ms_ssim twins and
the CPU extractor and fixed the CUDA twin. The SYCL twin had all four:
the decimate added sample * tap in two roundings where the reference
fuses each tap; the Gaussian window sums were fp32 running sums where
iqa_convolve() adds fp32 products in fp64; l / c / s were fp32
quotients where the CPU divides fp64 numerators by fp32 denominators;
the host combined unrounded means without fabs() on l and c. On an Arc
A380 it matched the CPU on none of 104 frames (up to 2.98e-6).

ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl
already had:

- sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of
  integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients
  as exact fp32 pairs, int64 fixed-point terms, the exact host sum).
  Both SSIM twins include it.
- integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window
  sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and
  fabs() of all three on the host.
- EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl.

Measured on an Arc A380 (xe, Level Zero) at --precision max against a
GCC build's --backend cpu, score, before and after:

- Netflix 576x324, 48 frames: 6.9e-8 -> identical
- 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical
- BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame
  1.1e-16 off through the host pow() of the icx build
- enable_lcs (16 outputs) and enable_chroma outputs identical on every
  frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit

Against the CPU extractor of its own icx binary the twin needs #1706 on
an AVX-512 host (that extractor's ssim_avx512.c tails were contracted).

Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15
paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and
1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms.
No kernel uses scratch memory.

Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18
outputs of 3 frames with == (53 and 54 of 54 differ on the old code);
test_sycl_kernel_source_contract.py has seven planted regressions.

docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed;
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP
and Metal twins.
lusoris added a commit that referenced this pull request Oct 1, 2026
…ale means are bit-identical

ADR-1403 found four differences between the GPU float_ms_ssim twins and
the CPU extractor and fixed the CUDA twin. The SYCL twin had all four:
the decimate added sample * tap in two roundings where the reference
fuses each tap; the Gaussian window sums were fp32 running sums where
iqa_convolve() adds fp32 products in fp64; l / c / s were fp32
quotients where the CPU divides fp64 numerators by fp32 denominators;
the host combined unrounded means without fabs() on l and c. On an Arc
A380 it matched the CPU on none of 104 frames (up to 2.98e-6).

ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl
already had:

- sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of
  integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients
  as exact fp32 pairs, int64 fixed-point terms, the exact host sum).
  Both SSIM twins include it.
- integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window
  sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and
  fabs() of all three on the host.
- EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl.

Measured on an Arc A380 (xe, Level Zero) at --precision max against a
GCC build's --backend cpu, score, before and after:

- Netflix 576x324, 48 frames: 6.9e-8 -> identical
- 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical
- BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame
  1.1e-16 off through the host pow() of the icx build
- enable_lcs (16 outputs) and enable_chroma outputs identical on every
  frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit

Against the CPU extractor of its own icx binary the twin needs #1706 on
an AVX-512 host (that extractor's ssim_avx512.c tails were contracted).

Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15
paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and
1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms.
No kernel uses scratch memory.

Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18
outputs of 3 frames with == (53 and 54 of 54 differ on the old code);
test_sycl_kernel_source_contract.py has seven planted regressions.

docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed;
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP
and Metal twins.
lusoris added a commit that referenced this pull request Oct 1, 2026
…ale means are bit-identical

ADR-1403 found four differences between the GPU float_ms_ssim twins and
the CPU extractor and fixed the CUDA twin. The SYCL twin had all four:
the decimate added sample * tap in two roundings where the reference
fuses each tap; the Gaussian window sums were fp32 running sums where
iqa_convolve() adds fp32 products in fp64; l / c / s were fp32
quotients where the CPU divides fp64 numerators by fp32 denominators;
the host combined unrounded means without fabs() on l and c. On an Arc
A380 it matched the CPU on none of 104 frames (up to 2.98e-6).

ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl
already had:

- sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of
  integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients
  as exact fp32 pairs, int64 fixed-point terms, the exact host sum).
  Both SSIM twins include it.
- integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window
  sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and
  fabs() of all three on the host.
- EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl.

Measured on an Arc A380 (xe, Level Zero) at --precision max against a
GCC build's --backend cpu, score, before and after:

- Netflix 576x324, 48 frames: 6.9e-8 -> identical
- 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical
- BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame
  1.1e-16 off through the host pow() of the icx build
- enable_lcs (16 outputs) and enable_chroma outputs identical on every
  frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit

Against the CPU extractor of its own icx binary the twin needs #1706 on
an AVX-512 host (that extractor's ssim_avx512.c tails were contracted).

Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15
paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and
1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms.
No kernel uses scratch memory.

Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18
outputs of 3 frames with == (53 and 54 of 54 differ on the old code);
test_sycl_kernel_source_contract.py has seven planted regressions.

docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed;
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP
and Metal twins.
lusoris added a commit that referenced this pull request Oct 1, 2026
…ale means are bit-identical (#1709)

* fix(sycl): make float_ms_ssim_sycl the CPU's arithmetic so its per-scale means are bit-identical

ADR-1403 found four differences between the GPU float_ms_ssim twins and
the CPU extractor and fixed the CUDA twin. The SYCL twin had all four:
the decimate added sample * tap in two roundings where the reference
fuses each tap; the Gaussian window sums were fp32 running sums where
iqa_convolve() adds fp32 products in fp64; l / c / s were fp32
quotients where the CPU divides fp64 numerators by fp32 denominators;
the host combined unrounded means without fabs() on l and c. On an Arc
A380 it matched the CPU on none of 104 frames (up to 2.98e-6).

ADR-1414, the SYCL part of T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01.
A SYCL kernel may not use fp64, so the fix reuses what float_ssim_sycl
already had:

- sycl_ssim_terms.h (new): the per-pixel SSIM arithmetic moved out of
  integer_ssim_sycl.cpp unchanged (window sums and the l / c quotients
  as exact fp32 pairs, int64 fixed-point terms, the exact host sum).
  Both SSIM twins include it.
- integer_ms_ssim_sycl.cpp: sycl::fma() per decimate tap, pair window
  sums, ssim_terms() for l / c / s, int64 group partials, fp32 means and
  fabs() of all three on the host.
- EXACT_TWINS lists float_ms_ssim and float_ms_ssim_lcs: sycl.

Measured on an Arc A380 (xe, Level Zero) at --precision max against a
GCC build's --backend cpu, score, before and after:

- Netflix 576x324, 48 frames: 6.9e-8 -> identical
- 1920x1080 checkerboards, 3 frames each: 1.06e-6, 2.98e-6 -> identical
- BBB 3840x2160: 1.23e-6 -> identical on 199 of 200 frames; one frame
  1.1e-16 off through the host pow() of the icx build
- enable_lcs (16 outputs) and enable_chroma outputs identical on every
  frame; also identical at 10, 12 and 16 bits and as 4:2:2 10-bit

Against the CPU extractor of its own icx binary the twin needs #1706 on
an AVX-512 host (that extractor's ssim_avx512.c tails were contracted).

Time per frame, medians: 3840x2160 31.4 ms before, 42.6 after (15
paired 100-frame runs; the CPU extractor 117 ms); 576x324 0.84 and
1.20 ms. float_ssim_sycl, whose helpers only moved: 22.9 and 23.0 ms.
No kernel uses scratch memory.

Tests: test_sycl_ms_ssim_parity and its 960x540 variant compare 18
outputs of 3 frames with == (53 and 54 of 54 differ on the old code);
test_sycl_kernel_source_contract.py has seven planted regressions.

docs/state.md: T-SYCL-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 closed;
T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 stays open for the HIP
and Metal twins.
@github-actions github-actions Bot added the type:bug Something isn't working label Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant