Skip to content

fix(hip): compile every HIP kernel with contraction off through one shared flag list - #1694

Merged
lusoris merged 2 commits into
masterfrom
fix/hip-fp-contract-off
Oct 1, 2026
Merged

lusoris merged 2 commits into
masterfrom
fix/hip-fp-contract-off

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes T-HIP-FP-CONTRACT-DEFAULT-2026-09-29.

hipcc compiles device code with -ffp-contract=fast, so a * b + c becomes one fused multiply-add. The CPU reference build does not contract. core/src/meson.build turned contraction off for three kernels only, through a per-kernel table (hip_cu_extra_flags): ssimulacra2_blur, integer_ssim_score and speed_pipeline. The other eighteen kernels rounded differently from the extractors they mirror.

Every kernel now gets one list, defined once between the VMAF HIP strict FP policy markers (ADR-1407, the HIP port of ADR-1367):

hip_strict_fp_args = ['-ffp-contract=off', '-fhip-fp32-correctly-rounded-divide-sqrt']

The table is gone, so a kernel cannot be exempted and a new kernel needs no flag entry. The second flag is hipcc's default, pinned.

The CUDA policy ADR was not on master when this was written, so this PR carries the HIP ADR; the CUDA row is T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29.

Pinned by two tests

  • test_hip_fp_arith_contract (device): a probe kernel built with the list computes a * b + c, a / b and sqrtf(|a|) for 1 048 576 random and 2 197 boundary operand triples; the results are compared with correctly rounded host values. Its host side is test_sycl_fp_arith_contract.c built for the HIP probe, so both backends share one set of operands and references. On the gfx1036:
Probe compiled with a * b + c differing a / b sqrtf
hip_strict_fp_args 0 0 0
hipcc defaults 126 577 0 0
-ffp-contract=off -fno-hip-fp32-correctly-rounded-divide-sqrt 0 303 181 158 765
  • test_hip_strict_fp_policy.py (device-free): the list holds both flags and is defined once, every HSACO compile gets it, no per-kernel table exists, no kernel source turns contraction back on, the probe is built with the same list. Seven planted regressions.

Parity per twin on ryzen-4090-arc (gfx1036, ROCm 7.2.4), before -> after

Max abs diff against --backend cpu at --precision max (bit-identical frames in brackets). 576x324 is the Netflix pair (48 frames), 4K is BBB 3840x2160 (22 frames).

Twin 576x324 4K
float_adm_hip 2.50e-5 -> 2.53e-6 6.04e-6 -> 1.28e-5
float_vif_hip 2.72e-5 -> 3.82e-5 5.74e-6 -> 7.02e-6
float_motion_hip 3.01e-6 -> 3.12e-6 (1/48) 2.37e-5 -> 2.36e-5
float_ssim_hip 1.79e-7 -> 1.19e-7 (0/48 -> 7/48) scale=1: 7.21e-6 -> 4.83e-6
integer_ms_ssim_hip 6.89e-8 -> 5.53e-8 5.82e-7 -> 1.22e-6
ciede_hip 1.133e-5 -> 1.134e-5 1.65e-6 -> 1.43e-6
psnr_hvs_hip 8.37e-5 -> 8.37e-5 1.10e-2 -> 1.10e-2
vif_hip 5.36e-7, output unchanged 2.98e-7, output unchanged
integer_ssim_hip 2.32e-14, output unchanged 5.58e-13, output unchanged
speed_chroma_hip 1.19e-6 (47/48), output unchanged 1.43e-6 (20/22), output unchanged
adm_hip, motion_hip, motion_v2_hip, psnr_hip, float_psnr_hip, float_moment_hip, cambi_hip, ssimulacra2_hip, speed_temporal_hip 0 (48/48), output unchanged 0 (22/22), output unchanged
  • Twelve twins produce bit-identical output before and after. Seven change, each inside its ADR-0214 tolerance.
  • float_vif_hip's worst frame gets worse while its mean stays (3.07e-6 -> 3.18e-6 at 576x324, 1.52e-6 -> 1.39e-6 at 4K). SYCL measured the same under ADR-1367 (2.71e-5 -> 3.81e-5): the residual is not in the operations the list controls.
  • cross_backend_parity_gate.py --backends cpu hip on the Netflix pair: the 17 runnable cells pass before and after. The motion cell aborts on a metric name in both (T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29).

Cost at 4K, ms per frame, (t(22) - t(2)) / 20, before -> after

Load average 12 to 26 from other jobs, so the two builds were run interleaved:

Twin Before After Reps
float_psnr_hip 4.49 4.44 7
psnr_hip 8.62 8.44 7
motion_hip 12.92 13.08 7
motion_v2_hip 13.98 13.81 7
float_motion_hip 20.38 19.97 7
adm_hip 76.76 72.96 5
float_adm_hip 77.95 78.75 5
float_vif_hip 82.94 86.49 5
float_ssim_hip (scale=1) 93.60 84.76 5
vif_hip 141.17 139.15 5
integer_ms_ssim_hip 164.61 157.06 5

Not interleaved, median of 3: speed_chroma 5.46 -> 5.50, float_moment 8.76 -> 8.00, speed_temporal 14.26 -> 14.60, psnr_hvs 18.26 -> 18.60, ciede 73.60 -> 74.68, cambi 99.49 -> 98.79, integer_ssim 136.9 -> 139.1, ssimulacra2 3765 -> 3649.

float_vif_hip is 4.3% slower in every interleaved sample. Nothing else leaves its run-to-run spread, and nothing reaches ADR-1367's 10% bar for an exemption.

Device defect seen during the measurement

Five runs had one or two wrong frames (vif_hip, float_moment_hip, adm_hip), on both builds. Repeats were clean; an interleaved repeat of adm_hip at 4K gave 15 of 15 clean runs with the list and 14 of 15 without. That is T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01. The tables are from clean runs.

Tests

python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast --suite=gpu: 261 OK on the rebased tip.

Type

  • fix — bug fix

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1407-hip-strict-fp-every-kernel.md.
  • Decision matrix — ## Alternatives considered in ADR-1407.
  • AGENTS.md invariant note — core/src/feature/hip/AGENTS.md, "Every kernel is IEEE-strict: hip_strict_fp_args".
  • Reproducer / smoke-test command — under "Reproducer" below.
  • CHANGELOG fragment — changelog.d/changed/hip-strict-fp-every-kernel.md.
  • Rebase note — docs/rebase-notes.md, "ADR-1407 — every HIP kernel compiles with hip_strict_fp_args".

Reproducer

On an AMD host (ryzen-4090-arc, gfx1036):

meson setup build-hip core -Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036 -Denable_cuda=false -Denable_sycl=false --buildtype=release -Db_lto=false
ninja -C build-hip
python3 scripts/ci/run_meson_test.py -- -C build-hip test_hip_fp_arith_contract test_hip_strict_fp_policy test_hip_device_resident_contract

Device-free: python3 core/test/test_hip_strict_fp_policy.py.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-HIP-FP-CONTRACT-DEFAULT-2026-09-29 moved to Recently closed with the per-twin numbers.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

Known follow-ups

…hared flag list

hipcc compiles device code with -ffp-contract=fast, so a * b + c becomes one
fused multiply-add. The CPU reference build does not contract, and
core/src/meson.build turned contraction off for three kernels only, through
a per-kernel table (hip_cu_extra_flags): ssimulacra2_blur, integer_ssim_score
and speed_pipeline. The other eighteen kernels rounded differently from the
extractors they mirror (T-HIP-FP-CONTRACT-DEFAULT-2026-09-29).

Every kernel now gets one list, hip_strict_fp_args = ['-ffp-contract=off',
'-fhip-fp32-correctly-rounded-divide-sqrt'], defined once between the VMAF
HIP strict FP policy markers; the table is gone (ADR-1407, the HIP port of
ADR-1367).

Two tests pin it. test_hip_fp_arith_contract runs a probe kernel built with
the list over 1 048 576 random and 2 197 boundary a * b + c, a / b and sqrtf
operands and compares them with correctly rounded host values; its host side
is test_sycl_fp_arith_contract.c built for the HIP probe. On a gfx1036 it
passes with the list, fails with hipcc's defaults (126 577 contracted
multiply-adds) and fails with approximate division (303 181 divisions,
158 765 square roots). test_hip_strict_fp_policy.py checks the build files
without a device, with seven planted regressions.

Measured on the gfx1036 against --backend cpu at --precision max, before ->
after, Netflix 576x324: float_adm 2.50e-5 -> 2.53e-6, float_ssim 1.79e-7 ->
1.19e-7, float_vif 2.72e-5 -> 3.82e-5 (mean 3.07e-6 -> 3.18e-6), float_motion
3.01e-6 -> 3.12e-6; twelve twins produce bit-identical output; the parity
gate's 17 runnable HIP cells pass before and after. At 3840x2160 float_vif
is 4.3% slower in interleaved runs and no other twin leaves its run-to-run
spread. The fast and gpu suites pass on the HIP build (259 OK).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant