Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -1341,6 +1341,26 @@
([SSIM](docs/metrics/ssim.md#options)).


- **`float_ms_ssim_sycl` computes the CPU's arithmetic.** The SYCL twin
added the decimate taps in two roundings where the CPU fuses them, kept the
Gaussian window sums and the luminance, contrast and structure terms in
fp32 where the CPU uses `double`, and combined unrounded per-scale means.
Measured on an Arc A380 it matched the CPU on none of 104 frames (6.9e-8 on
the Netflix 576x324 pair, up to 2.98e-6 on 1080p checkerboards, 1.23e-6 at
3840x2160). It now follows the reference operation for operation, with the
CPU's `double` values carried as exact pairs of floats and the frame sums
in 64-bit fixed point
([ADR-1414](docs/adr/1414-sycl-float-ms-ssim-cpu-arithmetic.md), after
ADR-1403 for CUDA). Every per-scale mean of every measured frame equals
the CPU's, `enable_lcs` and `enable_chroma` outputs included; the score
equals a GCC build's on 253 of 254 frames, the other differing by 1.1e-16
through the host `pow()` of an Intel-compiler build. The parity gate
compares the twin with tolerance 0. A 3840x2160 frame takes 42.6 ms on the
A380 against 31.4 before. Stored `float_ms_ssim_sycl` outputs change in
their low digits by at most the differences above. The HIP and Metal twins
keep the old arithmetic.


- **`float_ssim_sycl` combined formula residual eliminated.**
Arithmetic alignment from PR #1645 (`core/src/feature/sycl/integer_ssim_sycl.cpp`,
evaluating exact per-pixel $l \cdot c \cdot s$ in fp32 pairs, fixed-point work-group
Expand Down
18 changes: 18 additions & 0 deletions changelog.d/fixed/sycl-float-ms-ssim-cpu-arithmetic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
- **`float_ms_ssim_sycl` computes the CPU's arithmetic.** The SYCL twin
added the decimate taps in two roundings where the CPU fuses them, kept the
Gaussian window sums and the luminance, contrast and structure terms in
fp32 where the CPU uses `double`, and combined unrounded per-scale means.
Measured on an Arc A380 it matched the CPU on none of 104 frames (6.9e-8 on
the Netflix 576x324 pair, up to 2.98e-6 on 1080p checkerboards, 1.23e-6 at
3840x2160). It now follows the reference operation for operation, with the
CPU's `double` values carried as exact pairs of floats and the frame sums
in 64-bit fixed point
([ADR-1414](docs/adr/1414-sycl-float-ms-ssim-cpu-arithmetic.md), after
ADR-1403 for CUDA). Every per-scale mean of every measured frame equals
the CPU's, `enable_lcs` and `enable_chroma` outputs included; the score
equals a GCC build's on 253 of 254 frames, the other differing by 1.1e-16
through the host `pow()` of an Intel-compiler build. The parity gate
compares the twin with tolerance 0. A 3840x2160 frame takes 42.6 ms on the
A380 against 31.4 before. Stored `float_ms_ssim_sycl` outputs change in
their low digits by at most the differences above. The HIP and Metal twins
keep the old arithmetic.
50 changes: 41 additions & 9 deletions core/src/feature/sycl/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -281,7 +281,10 @@ HIP / Metal motion twins listed in Twin-update table above — same PR.
`enable_lcs` = separate kernel `launch_vert_combine_lcs` (default path
keeps one reduction), `float_ssim_lcs()` = `iqa/ssim_tools.c` L/C/S in
fp32 (clamped variances, flat-window covariance clamp, C3 = C2 / 2),
partials `[l | c | s]`. **On rebase**: do not fold products back into
partials `[l | c | s]`. Since ADR-1414 the per-pixel helpers
(`ssim_terms`, `add_*_tap`, `term_fixed`, `FixedSum`) live in
`sycl_ssim_terms.h`, shared with the MS-SSIM twin.
**On rebase**: do not fold products back into
expressions (icpx contracts `a * b + c`, ADR-1358) or restore the
left-to-right four-term variance sum; the identical-frame cases in
`test_sycl_twin_option_parity` fail on either.
Expand Down Expand Up @@ -644,13 +647,42 @@ HIP / Metal motion twins listed in Twin-update table above — same PR.

- **`integer_ms_ssim_sycl.cpp` waits once per frame (ADR-1363).** Each
(plane, scale) owns the span `partial_offset[plane][scale]` of
`d_partials` / `h_partials` (`[l x groups][c x groups][s x groups]`);
`submit()` enqueues the pyramid and every scale's `enqueue_scale_lcs` plus one
copy of the whole buffer, and `collect()` waits once and sums each span in
group order (`sum_scale_lcs`). Output is bit-identical to the per-scale
readback it replaced. **On rebase**: do not share one partials buffer across
scales again (the reuse is what forced the per-scale wait); the horizontal
workspace may be shared because the queue is in order.
`d_partials` / `h_partials` (`[l x groups][c x groups][s x groups]`, int64
since ADR-1414); `submit()` enqueues the pyramid and every scale's
`enqueue_scale_lcs` plus one copy of the whole buffer, and `collect()`
waits once and sums each span (`sum_scale_lcs`). **On rebase**: do not
share one partials buffer across scales again (the reuse is what forced
the per-scale wait); the horizontal workspace may be shared because the
queue is in order.

- **`integer_ms_ssim_sycl.cpp` = CPU arithmetic, type for type (ADR-1414).**
Four things, each one a regression if undone:
(1) `decimate_pixel()`: every tap `sycl::fma(sample, tap, acc)`, rows
first, then the nine row sums = `ms_ssim_decimate.c` (`vmaf_fmaf_exact`).
Plain `acc += sample * tap` is NOT contracted in this TU (ADR-1367) ->
1e-6 off. (2) Window sums: `add_horizontal_tap` / `add_vertical_tap` /
`round_moments` from `sycl_ssim_terms.h` = `iqa_convolve()`'s fp32
products summed in fp64, as fp32 pairs, one rounding per pass. (3)
`ssim_terms()`: fp32 variances + clamp, fp32 denominators, l and c as
pairs (CPU fp64 quotients), s = `div_rn` fp32 quotient, C3 = C2 / 2.0f.
(4) Sums: `term_fixed()` -> int64 units of 2^-52, `reduce_over_group`
exact, host `FixedSum`, mean `(double)(float)(sum / pixels)` (=
`iqa_ssim()` float return), `combine_ms_ssim()` with `fabs()` on l, c, s.
`sycl_ssim_terms.h` is shared with `float_ssim_sycl`
(`integer_ssim_sycl.cpp`): ONE copy of the arithmetic, no private
`ssim_terms` / `add_*_tap` / `term_fixed` in either TU (contract test).
Header holds host-only `double` (`FixedSum::value`); never call it from a
kernel (ADR-0220). Exact in practice, not by construction: pairs ~2^-46
vs CPU fp64, exact sum vs CPU running fp64 sum; fp32 mean rounding absorbs
it. Arc A380: every per-scale mean identical on Netflix pair (48),
checkerboards (3 + 3), BBB 4K (50), chroma, 10 / 12 / 16 bit. Score vs a
GCC build: 1 of 254 frames 1.1e-16 off = host `pow()` of the icx build
(libimf), not the device. Same-binary CPU on AVX-512 host needs #1706.
Scratch-free (ADR-1395). `EXACT_TWINS`: `float_ms_ssim`,
`float_ms_ssim_lcs`: `sycl`. Guards: `test_sycl_ms_ssim_parity` (+
`_large`; `==` on 18 outputs x 3 frames, runs FIRST in the binary so its
scalar-CPU cpumask precedes the process-wide SSIM dispatch install),
`test_sycl_kernel_source_contract.py` (seven planted regressions).

## icpx-aware clang-tidy

Expand Down Expand Up @@ -716,7 +748,7 @@ ADR-0884 / ADR-0946 backlog must update in same PR.
| `integer_ciede_sycl.cpp` | `ciede.c` | `test_sycl_ciede_parity.c` | ADR-0884 (round 2) |
| `integer_ssim_sycl.cpp` | `integer_ssim.c` | `test_sycl_ssim_parity.c` | ADR-0884 (round 2) |
| `integer_ssim_sycl.cpp` (`float_ssim_sycl`) | `float_ssim.c` + `ssim.c` | `test_sycl_float_ssim_parity.c` (+ `_large`) | ADR-1370 |
| `integer_ms_ssim_sycl.cpp` | `ms_ssim.c` | `test_sycl_ms_ssim_parity.c` | ADR-0884 (round 2) |
| `integer_ms_ssim_sycl.cpp` | `ms_ssim.c` | `test_sycl_ms_ssim_parity.c` (+ `_large`; bit-exact, 18 outputs x 3 frames) | ADR-0884 (round 2), ADR-1414 |
| `integer_motion_v2_sycl.cpp` | `integer_motion_v2.c` | `test_sycl_motion_v2_parity.c` | ADR-0884 (round 2) |
| `float_psnr_sycl.cpp` | `float_psnr.c` | `test_sycl_float_psnr_parity.c` | ADR-0946 (round 3) |
| `float_adm_sycl.cpp` | `float_adm.c` | `test_sycl_float_adm_parity.c` | ADR-0946 (round 3) |
Expand Down
Loading
Loading