Repository navigation
fix(hip): compile every HIP kernel with contraction off through one shared flag list - #1694
Merged
Merged
Conversation
15 of 26 tasks
lusoris
force-pushed
the
fix/hip-fp-contract-off
branch
from
October 1, 2026 11:20
e58bd08 to
d3e907a
Compare
…hared flag list hipcc compiles device code with -ffp-contract=fast, so a * b + c becomes one fused multiply-add. The CPU reference build does not contract, and core/src/meson.build turned contraction off for three kernels only, through a per-kernel table (hip_cu_extra_flags): ssimulacra2_blur, integer_ssim_score and speed_pipeline. The other eighteen kernels rounded differently from the extractors they mirror (T-HIP-FP-CONTRACT-DEFAULT-2026-09-29). Every kernel now gets one list, hip_strict_fp_args = ['-ffp-contract=off', '-fhip-fp32-correctly-rounded-divide-sqrt'], defined once between the VMAF HIP strict FP policy markers; the table is gone (ADR-1407, the HIP port of ADR-1367). Two tests pin it. test_hip_fp_arith_contract runs a probe kernel built with the list over 1 048 576 random and 2 197 boundary a * b + c, a / b and sqrtf operands and compares them with correctly rounded host values; its host side is test_sycl_fp_arith_contract.c built for the HIP probe. On a gfx1036 it passes with the list, fails with hipcc's defaults (126 577 contracted multiply-adds) and fails with approximate division (303 181 divisions, 158 765 square roots). test_hip_strict_fp_policy.py checks the build files without a device, with seven planted regressions. Measured on the gfx1036 against --backend cpu at --precision max, before -> after, Netflix 576x324: float_adm 2.50e-5 -> 2.53e-6, float_ssim 1.79e-7 -> 1.19e-7, float_vif 2.72e-5 -> 3.82e-5 (mean 3.07e-6 -> 3.18e-6), float_motion 3.01e-6 -> 3.12e-6; twelve twins produce bit-identical output; the parity gate's 17 runnable HIP cells pass before and after. At 3840x2160 float_vif is 4.3% slower in interleaved runs and no other twin leaves its run-to-run spread. The fast and gpu suites pass on the HIP build (259 OK).
lusoris
force-pushed
the
fix/hip-fp-contract-off
branch
from
October 1, 2026 12:05
d3e907a to
b943e49
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes
T-HIP-FP-CONTRACT-DEFAULT-2026-09-29.hipcc compiles device code with
-ffp-contract=fast, soa * b + cbecomes one fused multiply-add. The CPU reference build does not contract.core/src/meson.buildturned contraction off for three kernels only, through a per-kernel table (hip_cu_extra_flags):ssimulacra2_blur,integer_ssim_scoreandspeed_pipeline. The other eighteen kernels rounded differently from the extractors they mirror.Every kernel now gets one list, defined once between the
VMAF HIP strict FP policymarkers (ADR-1407, the HIP port of ADR-1367):The table is gone, so a kernel cannot be exempted and a new kernel needs no flag entry. The second flag is hipcc's default, pinned.
The CUDA policy ADR was not on
masterwhen this was written, so this PR carries the HIP ADR; the CUDA row isT-CUDA-FP-CONTRACT-DEFAULT-2026-09-29.Pinned by two tests
test_hip_fp_arith_contract(device): a probe kernel built with the list computesa * b + c,a / bandsqrtf(|a|)for 1 048 576 random and 2 197 boundary operand triples; the results are compared with correctly rounded host values. Its host side istest_sycl_fp_arith_contract.cbuilt for the HIP probe, so both backends share one set of operands and references. On the gfx1036:a * b + cdifferinga / bsqrtfhip_strict_fp_args-ffp-contract=off -fno-hip-fp32-correctly-rounded-divide-sqrttest_hip_strict_fp_policy.py(device-free): the list holds both flags and is defined once, every HSACO compile gets it, no per-kernel table exists, no kernel source turns contraction back on, the probe is built with the same list. Seven planted regressions.Parity per twin on
ryzen-4090-arc(gfx1036, ROCm 7.2.4), before -> afterMax abs diff against
--backend cpuat--precision max(bit-identical frames in brackets). 576x324 is the Netflix pair (48 frames), 4K is BBB 3840x2160 (22 frames).float_adm_hipfloat_vif_hipfloat_motion_hipfloat_ssim_hipscale=1: 7.21e-6 -> 4.83e-6integer_ms_ssim_hipciede_hippsnr_hvs_hipvif_hipinteger_ssim_hipspeed_chroma_hipadm_hip,motion_hip,motion_v2_hip,psnr_hip,float_psnr_hip,float_moment_hip,cambi_hip,ssimulacra2_hip,speed_temporal_hipfloat_vif_hip's worst frame gets worse while its mean stays (3.07e-6 -> 3.18e-6 at 576x324, 1.52e-6 -> 1.39e-6 at 4K). SYCL measured the same under ADR-1367 (2.71e-5 -> 3.81e-5): the residual is not in the operations the list controls.cross_backend_parity_gate.py --backends cpu hipon the Netflix pair: the 17 runnable cells pass before and after. Themotioncell aborts on a metric name in both (T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29).Cost at 4K, ms per frame,
(t(22) - t(2)) / 20, before -> afterLoad average 12 to 26 from other jobs, so the two builds were run interleaved:
float_psnr_hippsnr_hipmotion_hipmotion_v2_hipfloat_motion_hipadm_hipfloat_adm_hipfloat_vif_hipfloat_ssim_hip(scale=1)vif_hipinteger_ms_ssim_hipNot interleaved, median of 3:
speed_chroma5.46 -> 5.50,float_moment8.76 -> 8.00,speed_temporal14.26 -> 14.60,psnr_hvs18.26 -> 18.60,ciede73.60 -> 74.68,cambi99.49 -> 98.79,integer_ssim136.9 -> 139.1,ssimulacra23765 -> 3649.float_vif_hipis 4.3% slower in every interleaved sample. Nothing else leaves its run-to-run spread, and nothing reaches ADR-1367's 10% bar for an exemption.Device defect seen during the measurement
Five runs had one or two wrong frames (
vif_hip,float_moment_hip,adm_hip), on both builds. Repeats were clean; an interleaved repeat ofadm_hipat 4K gave 15 of 15 clean runs with the list and 14 of 15 without. That isT-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01. The tables are from clean runs.Tests
python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast --suite=gpu: 261 OK on the rebased tip.Type
fix— bug fixDeep-dive deliverables (ADR-0108)
docs/research/1407-hip-strict-fp-every-kernel.md.## Alternatives consideredin ADR-1407.AGENTS.mdinvariant note —core/src/feature/hip/AGENTS.md, "Every kernel is IEEE-strict:hip_strict_fp_args".changelog.d/changed/hip-strict-fp-every-kernel.md.docs/rebase-notes.md, "ADR-1407 — every HIP kernel compiles with hip_strict_fp_args".Reproducer
On an AMD host (
ryzen-4090-arc, gfx1036):Device-free:
python3 core/test/test_hip_strict_fp_policy.py.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-HIP-FP-CONTRACT-DEFAULT-2026-09-29moved to Recently closed with the per-twin numbers.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Known follow-ups
float_vif_hipkeeps a 3e-6 mean offset and a worst frame at 3.8e-5 with strict arithmetic; not isolated here.float_motion_hip) and perf(hip): decimate float_ssim on the device so 1080p and 4K run on the twin #1685 (float_ssim_hip) were measured without this list; their numbers will move within the ranges above once both land.