Skip to content

perf(hip): upload each plane of a frame once and let every twin read it - #1701

Merged
lusoris merged 2 commits into
masterfrom
perf/hip-shared-plane-uploads
Oct 1, 2026
Merged

lusoris merged 2 commits into
masterfrom
perf/hip-shared-plane-uploads

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29 and T-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19.

The HIP backend is host-picture only, and every twin copied the planes it reads into device buffers of its own, each with the wait that keeps a pageable picture from being refilled under the copy. With thirteen twins in one process a 4:2:0 frame pair was uploaded as 31 planes where it has 6, and vmaf_float_v0.6.1 had lost 21% of its throughput to the per-twin wait.

A VmafContext now owns the device copy of the frame (VmafHipSharedFrame, core/src/hip/shared_frame.{c,h}, ADR-1408):

  • vmaf_read_pictures() announces the frame's pictures before the dispatch loop and ends the frame after it.
  • A twin asks for its planes with vmaf_hip_plane_source_acquire[_luma](). The first request for a plane uploads it with the waiting vmaf_hip_picture_upload(), together with the planes the twins asked for in the frame before; every later request gets the same device pointer and waits for nothing. A frame is one upload call and one host wait.
  • Frames alternate between two slots. A twin's hold on a slot ends at its next acquire; an upload into a slot a frame-skipping twin still holds (n_subsample) waits for the device first.
  • The pictures are read only between begin and end, so the pageable-upload race (T-HIP-PAGEABLE-UPLOAD-RACE-2026-09-18) stays closed. Without a shared frame (the extractor API used directly) the same call uploads into the twin's own buffers.

Adopted: psnr_hip, float_psnr_hip, float_moment_hip, ciede_hip, integer_ssim_hip, float_ssim_hip, vif_hip, float_vif_hip, adm_hip, float_adm_hip, motion_hip, motion_v2_hip, float_motion_hip. The motion twins keep the previous frame with a device-to-device copy behind the SAD instead of pinned staging and a two-plane ping-pong.

The shared frame belongs to the context, not to the imported VmafHipState (the SYCL shape): one state can serve more than one context, and nothing in the public API changes.

Output on ryzen-4090-arc (gfx1036, ROCm 7.2.4), --precision max

Every metric of every frame is bit-identical before and after (origin/master 7dc4526 against the change). "Thirteen twins" names adm_hip, float_vif_hip and float_moment_hip, because --backend hip does not select the first two for adm and float_vif.

Run Fixture Frames Metric series Identical
Thirteen twins in one process Netflix 576x324 48 40 all
Thirteen twins Netflix 576x324, 10 bits 3 40 all
Thirteen twins 1080p checkerboard 1 px / 10 px 3 / 3 40 each all
Thirteen twins BBB 3840x2160 50 40 all
Thirteen twins, --subsample 2 Netflix / BBB 4K 12 / 12 40 each all
--model version=vmaf_v0.6.1 Netflix / checkerboard / BBB 4K 48 / 3 / 50 15 each all
--model version=vmaf_float_v0.6.1 Netflix / checkerboard / BBB 4K 48 / 3 / 50 15 each all
psnr + psnr_hvs + motion_v2 (the row's command) Netflix / BBB 4K 48 / 50 10 each all

390 series, max abs diff 0 in each. The twins' distance from the CPU is therefore unchanged.

Throughput, ms per frame

Steady state inside one process (the CLI's frames-per-second line at frame 11 and at the last frame), median of interleaved runs of the two builds, load average 6 to 35 from other jobs.

Run 1080p before 1080p after 4K before 4K after
--model version=vmaf_float_v0.6.1 57.30 46.86 294.32 226.43
--model version=vmaf_v0.6.1 (default model) 37.36 37.37 about 155 or 230 about 155 or 230
thirteen twins 182.71 184.20 731.17 734.96
adm_hip + vif + motion 52.87 52.51 227.75 213.55
vif + motion 41.83 38.03 224.35 222.51
psnr + psnr_hvs + motion_v2 8.00 7.97 31.74 32.27
psnr + float_psnr + float_moment_hip 4.95 4.41 17.16 15.11
psnr + motion_v2 4.21 4.14 18.49 17.96
motion 2.71 2.71 12.39 11.04
psnr 1.63 1.67 7.14 7.35
  • vmaf_float_v0.6.1 gets back what the per-twin wait cost it: 17.5 to 21.3 frames per second at 1080p (seven pairs; six of seven samples after between 45.8 and 47.6 ms, every sample before between 54.8 and 65.5), and the change is the faster one in each of five pairs at 4K. float_motion_hip no longer waits behind float_adm_hip's kernels, so the host reaches the frame's CPU float_vif while they run.
  • The default model is unchanged (nine pairs at 1080p). At 4K both builds run at about 155 or about 230 ms depending on what else loads the host (samples before: 193 / 149 / 161 / 230 / 233, after: 152 / 155 / 229 / 232 / 234).
  • Runs entirely on the device are bound by their kernels and do not move.
  • The row's command is unchanged because psnr_hvs_hip is most of it and still stages its own planes (fix(sycl): return the CPU's psnr_hvs scores bit for bit on SYCL and HIP #1689 / fix(sycl): score 4:0:0 input in psnr_hvs_sycl and psnr_hvs_hip #1692 own that file).

Two other ways to bring a shared plane to the device were measured and rejected (Research-1408): pinned host planes the kernels read in place (psnr 7.06 to 9.70 ms per 4K frame, psnr + motion_v2 19.15 to 22.00) and pinned staging with a device copy (13.24 for motion_v2, where the waiting upload gives 10.86). On this iGPU the runtime's upload of a pageable picture is cheaper than a host copy of it.

Tests

  • test_hip_shared_frame (new, fast, no device): shared_frame.c against stubs of the runtime. One upload per plane and frame, the planes of the frame before taken along, the picture read before an acquire returns and never after the frame ended, slot alternation, the wait for a frame-skipping twin, private planes for everything the shared frame does not serve, a failed upload is not remembered.
  • test_hip_shared_frame_contract.py (new, fast): the sources, with 17 planted regressions (a twin with its own upload, the frame ended before dispatch, the upload without the wait, an unfenced slot, ...).
  • test_hip_upload_race (extended, device): every twin in one context against the CPU with and without n_subsample, and both pictures refilled the moment the frame ended. With the wait taken out of the shared upload it fails on every run: 7 of 12 extractors off in the pooled check, 8 of 272 scores changed in the refill check.
  • python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast --suite=gpu on the pushed tip: 265 OK, 1 failed (on 24ac5bb). The failure is test_cuda_parity_gate_default_run, which fails on every build without CUDA on master and is fixed by fix(ci): skip the CUDA parity-gate default run on a build without CUDA #1698.

Type

  • perf — performance

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1408-hip-shared-frame-planes.md.
  • Decision matrix — ## Alternatives considered in ADR-1408.
  • AGENTS.md invariant note — core/src/feature/hip/AGENTS.md ("Picture uploads") and core/src/hip/AGENTS.md.
  • Reproducer / smoke-test command — under "Reproducer" below.
  • CHANGELOG fragment — changelog.d/changed/hip-shared-frame-planes.md.
  • Rebase note — docs/rebase-notes.md, "ADR-1408 — a VmafContext uploads each frame plane once for all HIP twins".

Reproducer

On an AMD host (ryzen-4090-arc, gfx1036):

meson setup build-hip core -Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036 -Denable_cuda=false -Denable_sycl=false --buildtype=release -Db_lto=false
ninja -C build-hip
python3 scripts/ci/run_meson_test.py -- -C build-hip test_hip_shared_frame test_hip_shared_frame_contract test_hip_upload_race
build-hip/tools/vmaf -r ref.yuv -d dis.yuv -w 1920 -h 1080 -p 420 -b 8 --backend hip --model version=vmaf_float_v0.6.1 --json -o out.json --precision max

Device-free: python3 core/test/test_hip_shared_frame_contract.py and build-hip/test/test_hip_shared_frame.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29 and T-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19 moved to Recently closed with the gfx1036 results; T-HIP-SHARED-FRAME-REMAINING-TWINS-2026-10-01 opened; T-HIP-SHARED-UPLOAD-DISCRETE-GPU-2026-10-01 deferred.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

Known follow-ups

  • psnr_hvs_hip, ssimulacra2_hip, float_ms_ssim_hip, cambi_hip, speed_chroma_hip and speed_temporal_hip still stage their own planes (T-HIP-SHARED-FRAME-REMAINING-TWINS-2026-10-01).
  • A discrete AMD GPU is unmeasured (T-HIP-SHARED-UPLOAD-DISCRETE-GPU-2026-10-01).
  • float_vif_hip is not selected for float_vif (enable_float_vif_hip_autodispatch is off), so vmaf_float_v0.6.1 on HIP computes VIF on the CPU.

The HIP backend is host-picture only, and every twin copied the planes it
reads into device buffers of its own, each with the wait that keeps a
pageable picture from being refilled under the copy. Thirteen twins in one
process uploaded a 4:2:0 frame pair as 31 planes where it has 6, and the
host blocked once per twin and frame.

A VmafContext now owns the device copy of the frame its HIP twins read
(VmafHipSharedFrame, core/src/hip/shared_frame.c, ADR-1408).
vmaf_read_pictures() announces the frame's pictures before the dispatch
loop and ends the frame after it. A twin asks for its planes with
vmaf_hip_plane_source_acquire(): the first request for a plane uploads it
with the waiting vmaf_hip_picture_upload(), together with the planes the
twins asked for in the frame before, and every later request in the frame
gets the same device pointer: one upload call and one host wait per frame.
Frames alternate between two slots; an upload into a slot a frame-skipping
twin still holds waits for the device first. The pictures are read only
between begin and end, so the pageable-upload race stays closed. Without a
shared frame the same call uploads into the twin's own buffers.

psnr, float_psnr, float_moment, ciede, integer_ssim, float_ssim, vif,
float_vif, adm, float_adm, motion, motion_v2 and float_motion read the
shared planes. The motion twins keep the previous frame with a
device-to-device copy behind the SAD instead of pinned staging and a
ping-pong.

On a gfx1036 every metric of every frame is bit-identical before and after
(thirteen twins, both shipped models, 576x324 to 3840x2160, 8 and 10 bits,
with and without frame subsampling). vmaf_float_v0.6.1 goes from 57.3 to
46.9 ms per 1080p frame and from 294 to 226 ms per 4K frame; vmaf_v0.6.1
and runs whose time is all device kernels are unchanged.

Closes T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29 and
T-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19.
@lusoris
lusoris force-pushed the perf/hip-shared-plane-uploads branch from a7a3ae4 to 05fae1e Compare October 1, 2026 14:09
@lusoris
lusoris merged commit 6283ee5 into master Oct 1, 2026
68 of 75 checks passed
@lusoris
lusoris deleted the perf/hip-shared-plane-uploads branch October 1, 2026 14:09
@github-actions github-actions Bot added the type:perf Performance improvement label Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant