Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -1169,6 +1169,30 @@
`float_ssim_l/c/s` means the same way; linear scores move by less than 6e-8.


- **SYCL and HIP `psnr_hvs` return the CPU extractor's scores bit for bit**
(`T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01`, `T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01`,
[ADR-1401](docs/adr/1401-psnr-hvs-sycl-hip-exact-twins.md)). Like the CUDA
twin (ADR-1397), `psnr_hvs_sycl` and `psnr_hvs_hip` now store the 64 terms of
every block in the CPU's arithmetic and the host adds them in the CPU's order.
Before, they summed each block on the device and were up to 1.7e-2 dB from
`--backend cpu` at 3840x2160, beyond the parity tolerance. `psnr_hvs`,
`psnr_hvs_y`, `psnr_hvs_cb` and `psnr_hvs_cr` are identical to the CPU at
`--precision max` from 576x324 to 3840x2160 and at 8 to 12 bits, measured on
an Arc A380 and a gfx1036, and the parity gate compares every `psnr_hvs` cell
with tolerance 0. The SYCL kernel has no fp64: it takes the CPU's `double`
masking threshold from a new integer square root of the exact product
(`sqrt_prod_rn()` in `sycl_exact_fp.h`), and it stays free of scratch memory.
The scores of both twins therefore change in their last digits (by up to
1.7e-2 dB at 3840x2160). Both are slower for it: a 3840x2160 frame takes
35.9 ms instead of 22.4 ms on an Arc A380 and about 38 ms instead of 18 ms on
a gfx1036, and the term buffer needs 65 MB per 3840x2160 frame; tuning is
tracked as `T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01`. Compare a
twin with the CPU extractor of the same `vmaf` binary: the dB value uses the
host's `log10`, which differs by one unit in the last place between an `icx`
and a gcc build. See
[the psnr_hvs page](docs/metrics/psnr-hvs.md#agreement-with-the-cpu-extractor).


- **SYCL: kernels stay out of scratch memory, and a wrong-result driver is
reported.** On an Arc A380 under the Linux xe driver, a SYCL kernel that keeps
a private array in memory or spills registers returns wrong values, with no
Expand Down
22 changes: 22 additions & 0 deletions changelog.d/fixed/sycl-hip-psnr-hvs-cpu-float-sum.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
- **SYCL and HIP `psnr_hvs` return the CPU extractor's scores bit for bit**
(`T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01`, `T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01`,
[ADR-1401](docs/adr/1401-psnr-hvs-sycl-hip-exact-twins.md)). Like the CUDA
twin (ADR-1397), `psnr_hvs_sycl` and `psnr_hvs_hip` now store the 64 terms of
every block in the CPU's arithmetic and the host adds them in the CPU's order.
Before, they summed each block on the device and were up to 1.7e-2 dB from
`--backend cpu` at 3840x2160, beyond the parity tolerance. `psnr_hvs`,
`psnr_hvs_y`, `psnr_hvs_cb` and `psnr_hvs_cr` are identical to the CPU at
`--precision max` from 576x324 to 3840x2160 and at 8 to 12 bits, measured on
an Arc A380 and a gfx1036, and the parity gate compares every `psnr_hvs` cell
with tolerance 0. The SYCL kernel has no fp64: it takes the CPU's `double`
masking threshold from a new integer square root of the exact product
(`sqrt_prod_rn()` in `sycl_exact_fp.h`), and it stays free of scratch memory.
The scores of both twins therefore change in their last digits (by up to
1.7e-2 dB at 3840x2160). Both are slower for it: a 3840x2160 frame takes
35.9 ms instead of 22.4 ms on an Arc A380 and about 38 ms instead of 18 ms on
a gfx1036, and the term buffer needs 65 MB per 3840x2160 frame; tuning is
tracked as `T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01`. Compare a
twin with the CPU extractor of the same `vmaf` binary: the dB value uses the
host's `log10`, which differs by one unit in the last place between an `icx`
and a gcc build. See
[the psnr_hvs page](docs/metrics/psnr-hvs.md#agreement-with-the-cpu-extractor).
7 changes: 6 additions & 1 deletion core/src/feature/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -672,7 +672,12 @@ feature/
(`x + 0.0f == x`); nothing else. **On rebase**: upstream change to the
tail of `calc_psnrhvs()` or to `extract()` in
`third_party/xiph/psnr_hvs.c` -> same change here + every twin that
calls it (`cuda/integer_psnr_hvs_cuda.c`).
calls it (`cuda/integer_psnr_hvs_cuda.c`,
`sycl/integer_psnr_hvs_sycl.cpp`, `hip/integer_psnr_hvs_hip.c`;
ADR-1401). `log10` behind the dB value = host libm: an icx-built
binary (libimf) and a gcc-built one (glibc) differ by one ulp on some
frames, CPU extractor and twins alike, so compare twin and CPU from
one binary.

- **`psnr_hvs` AVX2 DCT bit-exactness** (fork-local, ADR-0159):
[`x86/psnr_hvs_avx2.c`](x86/psnr_hvs_avx2.c) vectorizes
Expand Down
4 changes: 3 additions & 1 deletion core/src/feature/cuda/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -262,7 +262,9 @@ HIP / Metal motion twins listed in Twin-update table below — same PR.
- no sum of terms in kernel or host TU (per-block partials round
differently: 1e-2 dB at 3840x2160).
Guards: `test_cuda_psnr_hvs_parity{,_large}` (device, `==` on all four
outputs, 3840x2160 included), `test_psnr_hvs_twin_exact_sum_contract.py`
outputs, 3840x2160 included; cases live in
`core/test/psnr_hvs_twin_parity.h`, shared with the SYCL and HIP twins,
ADR-1401), `test_psnr_hvs_twin_exact_sum_contract.py`
and `test_psnr_hvs_score` (device-free). Gate cell = tolerance 0
(`EXACT_TWINS`, `scripts/ci/cross_backend_calibration.py`). Upstream
change to `calc_psnrhvs()` arithmetic or order -> mirror it here in
Expand Down
29 changes: 29 additions & 0 deletions core/src/feature/hip/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -229,6 +229,35 @@ same list). gfx1036: hipcc default fails it with 126577 fused multiply-adds;
`-fno-hip-fp32-correctly-rounded-divide-sqrt` with 303181 divisions and
158765 square roots off.

## `psnr_hvs_hip` = CPU scores bit for bit (ADR-1397, ADR-1401)

Kernel (`integer_psnr_hvs/psnr_hvs_score.hip`) stores the 64 terms
`calc_psnrhvs()` sums per block (`hvs_store_terms`, row-major, blocks in plane
then raster order); `psnr_hvs_plane_scores()` hands each plane to
`vmaf_psnr_hvs_plane_score()` (`../psnr_hvs_score.c`: one running `float`, CPU
order), combined score + dB via the same file. Same contract as the CUDA twin.
Load-bearing, each one breaks bit-identity on its own:

- masking table = `(csf * 0.3885746225901003)^2` in `double`, stored `float`
(`hvs_mask_value`, constexpr -> static data);
- threshold = `sqrt((double)energy * ratio) / 32`, one rounding to `float`
(`hvs_threshold`); each work-item forms its own, the reference item takes
the larger after the barrier;
- `-ffp-contract=off` + `-fhip-fp32-correctly-rounded-divide-sqrt` on the
HSACO (`hip_cu_extra_flags`);
- coefficient error = integer `abs()` cast to `float`;
- no sum of terms in kernel or host TU (per-block partials round differently:
1e-2 dB at 3840x2160).

Guards: `test_hip_psnr_hvs_parity{,_large}` (device, `==` on all four outputs,
3840x2160 included), `test_psnr_hvs_twin_exact_sum_contract.py` and
`test_psnr_hvs_score` (device-free). Gate cell = tolerance 0 (`EXACT_TWINS`,
`scripts/ci/cross_backend_calibration.py`). Upstream change to
`calc_psnrhvs()` arithmetic or order -> mirror it here in the same PR.
Readback = 256 bytes per block (65 MB per 3840x2160 4:2:0 frame, device +
pinned); tuning tracked as T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01,
must stay bit-exact.

## Per-thread atomicAdd replaces CUDA per-warp `__shfl_down_sync` reduce (ADR-0539)

CUDA twin's `cuda_helper.cuh::warp_reduce` hard-codes
Expand Down
Loading
Loading