Skip to content

perf(hip): run cambi and SpEED entirely on the device - #1661

Merged
lusoris merged 8 commits into
masterfrom
perf/hip-rc3-device-resident
Oct 1, 2026
Merged

lusoris merged 8 commits into
masterfrom
perf/hip-rc3-device-resident

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29 and T-HIP-SPEED-HOST-RESIDUAL-2026-09-29.
Built on #1636 (fix/hip-rc3-parity), which is merged; the base is now master.

Ports the SYCL device-resident architectures of ADR-1357 (CAMBI) and ADR-1358 (SpEED) to the HIP backend as ADR-1378 and ADR-1384. cambi_hip, speed_chroma_hip, and speed_temporal_hip no longer run any stage of the CPU extractors on the host or perform per-frame host-device round trips.

Each frame is one staged upload (vmaf_hip_picture_upload_staged()), the entire pipeline executes on the extractor's stream, followed by one compact readback (CambiHipResults 88 bytes, SpeedGpuFrameResult 24/40 bytes), with collect() being the only host wait.

Parity & Validation on ryzen-4090-arc (gfx1036, ROCm 7.2.4)

  • CAMBI:
    • Netflix 576x324 pair: 48/48 frames identical to CPU (max diff 0.0).
    • BBB 3840x2160 (50 frames): 50/50 frames identical to CPU (max diff 0.0, bound 2.2e-15).
    • Short-frame domain (1920x64, 1920x128, 1920x160, 3840x128): runs clean and bit-identical to CPU reference.
    • Timings (ms/frame, median of 3 of (t(22) - t(2)) / 20, before -> after, CPU at 16 threads):
      • 576x324: 1.71 -> 1.80 ms/frame (CPU16: 0.14)
      • 3840x2160: 91.21 -> 130.55 ms/frame (CPU16: 10.48)
  • SpEED:
    • With exact log2f (LD_PRELOAD=/tmp/crlog2f.so), speed_gpu_parity.py --backend hip exits 0, 100% bit-identical to CPU on all channels and frames.
    • On stock glibc 2.44, glibc log2f misrounding causes max diff 1.43e-6 on 1-2 frames of speed_chroma_v.
    • Short-frame domain: 1920x64 cleanly rejected with -234 (< 80 min dim); 1920x128/160 and 3840x128 run clean and stay in bounds.
    • Timings (ms/frame, median of 3, before -> after, CPU at 16 threads):
      • speed_chroma 576x324: 1.41 -> 0.70 ms/frame (CPU16: 0.07)
      • speed_chroma 3840x2160: 22.81 -> 5.55 ms/frame (CPU16: 5.37) — 4.1x speedup
      • speed_temporal 576x324: 1.25 -> 0.74 ms/frame (CPU16: 0.40)
      • speed_temporal 3840x2160: 46.54 -> 15.15 ms/frame (CPU16: 20.21) — 3.1x speedup

Re-run after the rebase (2026-10-01)

Rebased onto master 591d53449 (after #1636, #1637, #1630 and the CAMBI short-frame fix #1642). On the gfx1036, same host and ROCm:

  • python3 scripts/dev/speed_gpu_parity.py --backend hip --feature cambi --no-timing: 48/48 and 50/50 frames identical (max diff 0.0), as above.
  • cambi_hip against the CPU cambi with fix(cambi): keep the c-values walks inside short frames #1642's clipped window: identical on banded 1920x64, 1920x128, 1920x160, 1920x176, 2560x240 and 3840x128 clips, and on narrow 64x1920 and 96x1080 vertical ramps. This is the short- and narrow-frame check ADR-1393 asks of the device c-values kernel.
  • SpEED with LD_PRELOAD=/tmp/crlog2f.so: every output identical on both clips, exit 0. On stock glibc: speed_chroma_v 47/48 and 49/50 frames identical (max 1.43e-6), speed_chroma_u 48/48 and 49/50 (9.5e-7), speed_chroma_uv 47/48 and 48/50 (9.5e-7), speed_temporal identical.
  • Short 4:2:0 frames: speed_temporal_hip equals the CPU at 1920x128, 1920x160 and 3840x128; speed_chroma_hip equals it at 1920x160 and, like the CPU speed_chroma, refuses 1920x64, 1920x128 and 3840x128 with -EINVAL (exit 234), whose chroma planes are below the 80-pixel minimum.
  • python3 scripts/ci/run_meson_test.py -- -C build-hip --suite gpu --num-processes 1: 52 OK, no Memory access fault; the fast suite passes on the HIP build (250 OK) and on a CPU-only build (194 OK, 1 skipped). One of three full device runs failed test_hip_upload_race on vif_hip (4 of 32 scores, the four scales of one frame). That is the gfx1036 command loss recorded as T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01, in an extractor this PR does not touch; the test then passed 20 of 20 times on its own, on this build and on the fix(hip): bring motion_hip and the HIP option twins onto the CPU's arithmetic #1636 build.

The timings above were not measured again; the rebase changes no HIP source of this PR. Two commits were added: one aligns the CAMBI notes #1642 wrote (it described the HIP twin as running the c-values walk on the host) with the device-resident twin, and one regenerates the ADR citation map.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1378-hip-cambi-speed-device-resident.md.
  • Decision matrix — ## Alternatives considered in ADR-1378 and ADR-1384.
  • AGENTS.md invariant note — core/src/feature/hip/AGENTS.md (device-resident CAMBI and SpEED pipelines) and docs/development/rebase-sensitive-invariants.md; core/src/feature/AGENTS.md now says which twins still call the host c-values walk.
  • Reproducer / smoke-test command — under "Reproducer" below.
  • CHANGELOG fragment — changelog.d/changed/perf-hip-cambi-device-resident.md, changelog.d/changed/perf-hip-speed-device-resident.md.
  • Rebase note — docs/rebase-notes.md, "ADR-1378 / ADR-1384 — HIP CAMBI and SpEED run entirely on the device".

Reproducer

On an AMD host (ryzen-4090-arc, gfx1036):

meson setup build-hip core -Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036 -Db_lto=false
ninja -C build-hip
python3 scripts/ci/run_meson_test.py -- -C build-hip --suite gpu --num-processes 1
python3 scripts/dev/speed_gpu_parity.py --backend hip --vmaf $PWD/build-hip/tools/vmaf --feature cambi --no-timing
python3 scripts/dev/speed_gpu_parity.py --backend hip --vmaf $PWD/build-hip/tools/vmaf --no-timing

Device-free: python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast (test_hip_cambi_device_math, test_hip_speed_device_math, test_hip_device_resident_contract).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29 and T-HIP-SPEED-HOST-RESIDUAL-2026-09-29 closed with the measured gfx1036 results.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

@github-actions github-actions Bot added the type:perf Performance improvement label Oct 1, 2026
@lusoris
lusoris force-pushed the fix/hip-rc3-parity branch 2 times, most recently from 3139555 to b6d0877 Compare October 1, 2026 06:56
Base automatically changed from fix/hip-rc3-parity to master October 1, 2026 06:56
@lusoris
lusoris force-pushed the perf/hip-rc3-device-resident branch from 47401e5 to 68d9204 Compare October 1, 2026 07:06
lusoris and others added 8 commits October 1, 2026 09:12
cambi_hip, speed_chroma_hip and speed_temporal_hip no longer run any stage
of the CPU extractors on the host. Each frame is one staged upload
(vmaf_hip_picture_upload_staged()), the whole pipeline on the extractor's
stream and one small readback; collect() is the only host wait. This ports
the SYCL designs of ADR-1357 (CAMBI) and ADR-1358 (SpEED) to HIP as
ADR-1378 and ADR-1384.

CAMBI keeps cambi.c's sliding column histograms and sums the top-K
c-values exactly, so it is bit-identical to the CPU wherever the CPU's own
double sum is exact, and it now applies cambi.c's reciprocal-table window
guard. The window, mask index, resize tables, contrast weights, top-K mean
and the guard are new shared helpers in cambi.c, which the CPU init and the
SYCL twin call too.

SpEED reproduces speed.c in fp32 operation for operation. On HIP that is a
per-TU build flag (-ffp-contract=off, correctly rounded divide and sqrt),
not per-operation intrinsics: HIP's __fmul_rn/__fadd_rn/__fdiv_rn are the
plain contracting operators and __fsqrt_rn the native approximation. The
init-time configure is one shared routine, speed_internal_gpu_configure(),
used by the SYCL twins as well.

Both kernels keep their per-work-item arithmetic in one header that a host
test replays against the CPU extractors in the fast suite
(test_hip_cambi_device_math, test_hip_speed_device_math), with planted
regressions; test_hip_device_resident_contract.py pins one upload, one
readback and one wait per frame. Built for gfx90a, gfx1030, gfx1036 and
gfx1100; not yet run on an AMD device (verify commands in docs/state.md).

Depends on #1636 (vmaf_hip_picture_upload_staged(), vmaf_hip_rc_to_errno()).
The CAMBI and SpEED checks repeated the same host-stage, frame-path and
readback logic; one helper each removes the repeats and keeps ruff's
branch limit.
ADR-1264 has a build without hipcc report -ENOSYS and nothing else; the
speed_temporal_hip ran their host configure first, so a bad window or
format returned -EINVAL there. They now return -ENOSYS before any check,
and the window-guard test checks only cambi.c's half on a scaffold build.
The contract test pins the order, with a planted regression per family.
described the CUDA, HIP and Metal twins as running that walk on the host.
With cambi_hip device-resident that no longer holds for HIP (nor for CUDA
since ADR-1379): only the Metal twin calls
vmaf_cambi_calculate_c_values(). The feature notes, the metric pages and
the changelog entry say so, and the state row records that cambi_hip
equals the CPU on the gfx1036 on narrow frames as well.
The rebase onto master took master's map for the conflicted file, which
still listed the deleted core/src/feature/hip/speed/speed_score.hip under
ADR-0567 and lacked this branch's new sites.
@lusoris
lusoris force-pushed the perf/hip-rc3-device-resident branch from 68d9204 to 806e4d1 Compare October 1, 2026 07:18
@lusoris
lusoris merged commit 5593a70 into master Oct 1, 2026
73 of 76 checks passed
@lusoris
lusoris deleted the perf/hip-rc3-device-resident branch October 1, 2026 07:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant