Repository navigation
perf(hip): upload each plane of a frame once and let every twin read it - #1701
Merged
Merged
Conversation
The HIP backend is host-picture only, and every twin copied the planes it reads into device buffers of its own, each with the wait that keeps a pageable picture from being refilled under the copy. Thirteen twins in one process uploaded a 4:2:0 frame pair as 31 planes where it has 6, and the host blocked once per twin and frame. A VmafContext now owns the device copy of the frame its HIP twins read (VmafHipSharedFrame, core/src/hip/shared_frame.c, ADR-1408). vmaf_read_pictures() announces the frame's pictures before the dispatch loop and ends the frame after it. A twin asks for its planes with vmaf_hip_plane_source_acquire(): the first request for a plane uploads it with the waiting vmaf_hip_picture_upload(), together with the planes the twins asked for in the frame before, and every later request in the frame gets the same device pointer: one upload call and one host wait per frame. Frames alternate between two slots; an upload into a slot a frame-skipping twin still holds waits for the device first. The pictures are read only between begin and end, so the pageable-upload race stays closed. Without a shared frame the same call uploads into the twin's own buffers. psnr, float_psnr, float_moment, ciede, integer_ssim, float_ssim, vif, float_vif, adm, float_adm, motion, motion_v2 and float_motion read the shared planes. The motion twins keep the previous frame with a device-to-device copy behind the SAD instead of pinned staging and a ping-pong. On a gfx1036 every metric of every frame is bit-identical before and after (thirteen twins, both shipped models, 576x324 to 3840x2160, 8 and 10 bits, with and without frame subsampling). vmaf_float_v0.6.1 goes from 57.3 to 46.9 ms per 1080p frame and from 294 to 226 ms per 4K frame; vmaf_v0.6.1 and runs whose time is all device kernels are unchanged. Closes T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29 and T-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19.
lusoris
force-pushed
the
perf/hip-shared-plane-uploads
branch
from
October 1, 2026 14:09
a7a3ae4 to
05fae1e
Compare
This was referenced Oct 1, 2026
3 of 6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes
T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29andT-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19.The HIP backend is host-picture only, and every twin copied the planes it reads into device buffers of its own, each with the wait that keeps a pageable picture from being refilled under the copy. With thirteen twins in one process a 4:2:0 frame pair was uploaded as 31 planes where it has 6, and
vmaf_float_v0.6.1had lost 21% of its throughput to the per-twin wait.A
VmafContextnow owns the device copy of the frame (VmafHipSharedFrame,core/src/hip/shared_frame.{c,h}, ADR-1408):vmaf_read_pictures()announces the frame's pictures before the dispatch loop and ends the frame after it.vmaf_hip_plane_source_acquire[_luma](). The first request for a plane uploads it with the waitingvmaf_hip_picture_upload(), together with the planes the twins asked for in the frame before; every later request gets the same device pointer and waits for nothing. A frame is one upload call and one host wait.n_subsample) waits for the device first.T-HIP-PAGEABLE-UPLOAD-RACE-2026-09-18) stays closed. Without a shared frame (the extractor API used directly) the same call uploads into the twin's own buffers.Adopted:
psnr_hip,float_psnr_hip,float_moment_hip,ciede_hip,integer_ssim_hip,float_ssim_hip,vif_hip,float_vif_hip,adm_hip,float_adm_hip,motion_hip,motion_v2_hip,float_motion_hip. The motion twins keep the previous frame with a device-to-device copy behind the SAD instead of pinned staging and a two-plane ping-pong.The shared frame belongs to the context, not to the imported
VmafHipState(the SYCL shape): one state can serve more than one context, and nothing in the public API changes.Output on
ryzen-4090-arc(gfx1036, ROCm 7.2.4),--precision maxEvery metric of every frame is bit-identical before and after (
origin/master7dc4526 against the change). "Thirteen twins" namesadm_hip,float_vif_hipandfloat_moment_hip, because--backend hipdoes not select the first two foradmandfloat_vif.--subsample 2--model version=vmaf_v0.6.1--model version=vmaf_float_v0.6.1psnr+psnr_hvs+motion_v2(the row's command)390 series, max abs diff 0 in each. The twins' distance from the CPU is therefore unchanged.
Throughput, ms per frame
Steady state inside one process (the CLI's frames-per-second line at frame 11 and at the last frame), median of interleaved runs of the two builds, load average 6 to 35 from other jobs.
--model version=vmaf_float_v0.6.1--model version=vmaf_v0.6.1(default model)adm_hip+vif+motionvif+motionpsnr+psnr_hvs+motion_v2psnr+float_psnr+float_moment_hippsnr+motion_v2motionpsnrvmaf_float_v0.6.1gets back what the per-twin wait cost it: 17.5 to 21.3 frames per second at 1080p (seven pairs; six of seven samples after between 45.8 and 47.6 ms, every sample before between 54.8 and 65.5), and the change is the faster one in each of five pairs at 4K.float_motion_hipno longer waits behindfloat_adm_hip's kernels, so the host reaches the frame's CPUfloat_vifwhile they run.psnr_hvs_hipis most of it and still stages its own planes (fix(sycl): return the CPU's psnr_hvs scores bit for bit on SYCL and HIP #1689 / fix(sycl): score 4:0:0 input in psnr_hvs_sycl and psnr_hvs_hip #1692 own that file).Two other ways to bring a shared plane to the device were measured and rejected (Research-1408): pinned host planes the kernels read in place (
psnr7.06 to 9.70 ms per 4K frame,psnr+motion_v219.15 to 22.00) and pinned staging with a device copy (13.24 formotion_v2, where the waiting upload gives 10.86). On this iGPU the runtime's upload of a pageable picture is cheaper than a host copy of it.Tests
test_hip_shared_frame(new,fast, no device):shared_frame.cagainst stubs of the runtime. One upload per plane and frame, the planes of the frame before taken along, the picture read before an acquire returns and never after the frame ended, slot alternation, the wait for a frame-skipping twin, private planes for everything the shared frame does not serve, a failed upload is not remembered.test_hip_shared_frame_contract.py(new,fast): the sources, with 17 planted regressions (a twin with its own upload, the frame ended before dispatch, the upload without the wait, an unfenced slot, ...).test_hip_upload_race(extended, device): every twin in one context against the CPU with and withoutn_subsample, and both pictures refilled the moment the frame ended. With the wait taken out of the shared upload it fails on every run: 7 of 12 extractors off in the pooled check, 8 of 272 scores changed in the refill check.python3 scripts/ci/run_meson_test.py -- -C build-hip --suite=fast --suite=gpuon the pushed tip: 265 OK, 1 failed (on 24ac5bb). The failure istest_cuda_parity_gate_default_run, which fails on every build without CUDA on master and is fixed by fix(ci): skip the CUDA parity-gate default run on a build without CUDA #1698.Type
perf— performanceDeep-dive deliverables (ADR-0108)
docs/research/1408-hip-shared-frame-planes.md.## Alternatives consideredin ADR-1408.AGENTS.mdinvariant note —core/src/feature/hip/AGENTS.md("Picture uploads") andcore/src/hip/AGENTS.md.changelog.d/changed/hip-shared-frame-planes.md.docs/rebase-notes.md, "ADR-1408 — a VmafContext uploads each frame plane once for all HIP twins".Reproducer
On an AMD host (
ryzen-4090-arc, gfx1036):Device-free:
python3 core/test/test_hip_shared_frame_contract.pyandbuild-hip/test/test_hip_shared_frame.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29andT-HIP-UPLOAD-WAIT-THROUGHPUT-2026-09-19moved to Recently closed with the gfx1036 results;T-HIP-SHARED-FRAME-REMAINING-TWINS-2026-10-01opened;T-HIP-SHARED-UPLOAD-DISCRETE-GPU-2026-10-01deferred.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Known follow-ups
psnr_hvs_hip,ssimulacra2_hip,float_ms_ssim_hip,cambi_hip,speed_chroma_hipandspeed_temporal_hipstill stage their own planes (T-HIP-SHARED-FRAME-REMAINING-TWINS-2026-10-01).T-HIP-SHARED-UPLOAD-DISCRETE-GPU-2026-10-01).float_vif_hipis not selected forfloat_vif(enable_float_vif_hip_autodispatchis off), sovmaf_float_v0.6.1on HIP computes VIF on the CPU.