Repository navigation
perf(hip): run cambi and SpEED entirely on the device - #1661
Merged
Merged
Conversation
17 of 26 tasks
lusoris
force-pushed
the
fix/hip-rc3-parity
branch
2 times, most recently
from
October 1, 2026 06:56
3139555 to
b6d0877
Compare
lusoris
force-pushed
the
perf/hip-rc3-device-resident
branch
from
October 1, 2026 07:06
47401e5 to
68d9204
Compare
cambi_hip, speed_chroma_hip and speed_temporal_hip no longer run any stage of the CPU extractors on the host. Each frame is one staged upload (vmaf_hip_picture_upload_staged()), the whole pipeline on the extractor's stream and one small readback; collect() is the only host wait. This ports the SYCL designs of ADR-1357 (CAMBI) and ADR-1358 (SpEED) to HIP as ADR-1378 and ADR-1384. CAMBI keeps cambi.c's sliding column histograms and sums the top-K c-values exactly, so it is bit-identical to the CPU wherever the CPU's own double sum is exact, and it now applies cambi.c's reciprocal-table window guard. The window, mask index, resize tables, contrast weights, top-K mean and the guard are new shared helpers in cambi.c, which the CPU init and the SYCL twin call too. SpEED reproduces speed.c in fp32 operation for operation. On HIP that is a per-TU build flag (-ffp-contract=off, correctly rounded divide and sqrt), not per-operation intrinsics: HIP's __fmul_rn/__fadd_rn/__fdiv_rn are the plain contracting operators and __fsqrt_rn the native approximation. The init-time configure is one shared routine, speed_internal_gpu_configure(), used by the SYCL twins as well. Both kernels keep their per-work-item arithmetic in one header that a host test replays against the CPU extractors in the fast suite (test_hip_cambi_device_math, test_hip_speed_device_math), with planted regressions; test_hip_device_resident_contract.py pins one upload, one readback and one wait per frame. Built for gfx90a, gfx1030, gfx1036 and gfx1100; not yet run on an AMD device (verify commands in docs/state.md). Depends on #1636 (vmaf_hip_picture_upload_staged(), vmaf_hip_rc_to_errno()).
The CAMBI and SpEED checks repeated the same host-stage, frame-path and readback logic; one helper each removes the repeats and keeps ruff's branch limit.
ADR-1264 has a build without hipcc report -ENOSYS and nothing else; the speed_temporal_hip ran their host configure first, so a bad window or format returned -EINVAL there. They now return -ENOSYS before any check, and the window-guard test checks only cambi.c's half on a scaffold build. The contract test pins the order, with a planted regression per family.
described the CUDA, HIP and Metal twins as running that walk on the host. With cambi_hip device-resident that no longer holds for HIP (nor for CUDA since ADR-1379): only the Metal twin calls vmaf_cambi_calculate_c_values(). The feature notes, the metric pages and the changelog entry say so, and the state row records that cambi_hip equals the CPU on the gfx1036 on narrow frames as well.
The rebase onto master took master's map for the conflicted file, which still listed the deleted core/src/feature/hip/speed/speed_score.hip under ADR-0567 and lacked this branch's new sites.
lusoris
force-pushed
the
perf/hip-rc3-device-resident
branch
from
October 1, 2026 07:18
68d9204 to
806e4d1
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29 and T-HIP-SPEED-HOST-RESIDUAL-2026-09-29.
Built on #1636 (
fix/hip-rc3-parity), which is merged; the base is nowmaster.Ports the SYCL device-resident architectures of ADR-1357 (CAMBI) and ADR-1358 (SpEED) to the HIP backend as ADR-1378 and ADR-1384.
cambi_hip,speed_chroma_hip, andspeed_temporal_hipno longer run any stage of the CPU extractors on the host or perform per-frame host-device round trips.Each frame is one staged upload (
vmaf_hip_picture_upload_staged()), the entire pipeline executes on the extractor's stream, followed by one compact readback (CambiHipResults88 bytes,SpeedGpuFrameResult24/40 bytes), withcollect()being the only host wait.Parity & Validation on
ryzen-4090-arc(gfx1036, ROCm 7.2.4)log2f(LD_PRELOAD=/tmp/crlog2f.so),speed_gpu_parity.py --backend hipexits 0, 100% bit-identical to CPU on all channels and frames.log2fmisrounding causes max diff 1.43e-6 on 1-2 frames ofspeed_chroma_v.speed_chroma576x324: 1.41 -> 0.70 ms/frame (CPU16: 0.07)speed_chroma3840x2160: 22.81 -> 5.55 ms/frame (CPU16: 5.37) — 4.1x speedupspeed_temporal576x324: 1.25 -> 0.74 ms/frame (CPU16: 0.40)speed_temporal3840x2160: 46.54 -> 15.15 ms/frame (CPU16: 20.21) — 3.1x speedupRe-run after the rebase (2026-10-01)
Rebased onto master
591d53449(after #1636, #1637, #1630 and the CAMBI short-frame fix #1642). On the gfx1036, same host and ROCm:python3 scripts/dev/speed_gpu_parity.py --backend hip --feature cambi --no-timing: 48/48 and 50/50 frames identical (max diff 0.0), as above.cambi_hipagainst the CPUcambiwith fix(cambi): keep the c-values walks inside short frames #1642's clipped window: identical on banded 1920x64, 1920x128, 1920x160, 1920x176, 2560x240 and 3840x128 clips, and on narrow 64x1920 and 96x1080 vertical ramps. This is the short- and narrow-frame check ADR-1393 asks of the device c-values kernel.LD_PRELOAD=/tmp/crlog2f.so: every output identical on both clips, exit 0. On stock glibc:speed_chroma_v47/48 and 49/50 frames identical (max 1.43e-6),speed_chroma_u48/48 and 49/50 (9.5e-7),speed_chroma_uv47/48 and 48/50 (9.5e-7),speed_temporalidentical.speed_temporal_hipequals the CPU at 1920x128, 1920x160 and 3840x128;speed_chroma_hipequals it at 1920x160 and, like the CPUspeed_chroma, refuses 1920x64, 1920x128 and 3840x128 with-EINVAL(exit 234), whose chroma planes are below the 80-pixel minimum.python3 scripts/ci/run_meson_test.py -- -C build-hip --suite gpu --num-processes 1: 52 OK, noMemory access fault; the fast suite passes on the HIP build (250 OK) and on a CPU-only build (194 OK, 1 skipped). One of three full device runs failedtest_hip_upload_raceonvif_hip(4 of 32 scores, the four scales of one frame). That is the gfx1036 command loss recorded asT-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01, in an extractor this PR does not touch; the test then passed 20 of 20 times on its own, on this build and on the fix(hip): bring motion_hip and the HIP option twins onto the CPU's arithmetic #1636 build.The timings above were not measured again; the rebase changes no HIP source of this PR. Two commits were added: one aligns the CAMBI notes #1642 wrote (it described the HIP twin as running the c-values walk on the host) with the device-resident twin, and one regenerates the ADR citation map.
Deep-dive deliverables (ADR-0108)
docs/research/1378-hip-cambi-speed-device-resident.md.## Alternatives consideredin ADR-1378 and ADR-1384.AGENTS.mdinvariant note —core/src/feature/hip/AGENTS.md(device-resident CAMBI and SpEED pipelines) anddocs/development/rebase-sensitive-invariants.md;core/src/feature/AGENTS.mdnow says which twins still call the host c-values walk.changelog.d/changed/perf-hip-cambi-device-resident.md,changelog.d/changed/perf-hip-speed-device-resident.md.docs/rebase-notes.md, "ADR-1378 / ADR-1384 — HIP CAMBI and SpEED run entirely on the device".Reproducer
On an AMD host (
ryzen-4090-arc, gfx1036):Device-free:
python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast(test_hip_cambi_device_math,test_hip_speed_device_math,test_hip_device_resident_contract).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29andT-HIP-SPEED-HOST-RESIDUAL-2026-09-29closed with the measured gfx1036 results.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.