Repository navigation
feat(api): import HIP device frames, dma-bufs and GL textures with event and sync_file fences (RC4 WP3, ADR-2092) - #2341
Merged
Merged
Conversation
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
…++ 2026.1.1 The exit evidence of the SYCL lane was first measured with the host's oneAPI 2026.0, while build-config.env pins 2026.1. It is repeated in throwaway containers of the dev image (icpx 2026.1.1, Level Zero GPU driver 26.35.39758.10) with the Arc A380 passed through, and every result holds: 188 + 48 bit-exact cells with 0 differing values, the acquire, release and canary tests, the recorded reads, two contexts, dma-buf and GL imports, 130 scratch-free kernels and the AOT compile for 19 targets. The barrier against depends_on() timing and the three DPC++ runtime findings hold on 2026.1.1 too; the research note lists them by version, with the 2026.0 numbers as footnotes. Its claim that a 2D copy is one command per row is withdrawn: a Level Zero trace on 2026.1.1 shows one kernel launch, and 2026.0 could not be traced. The research note moves from 2159 to 2160: the HIP lane (#2341) also took 2159, and no allocator exists for research ids, so 2160 is the highest id on master, every origin branch and every sibling worktree plus one. ADR-2091 is accepted as designed (maintainer popup 2026-10-06).
This was referenced Oct 6, 2026
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
Conflicts resolved per hunk: - core/api/vmafx.toml, core/src/AGENTS.d/vmafx-device-frames.md: the HIP lane's text with integration's ABI number (the release callback and the GL texture memory are ABI 0.1.7 here, 0.1.4 on the lane). - core/test/test_device_target_header_dependencies.py: the lane's counts (CUDA 22, HIP 21) and its kernel-source roots (feature_src_dir, cuda_dir, src_dir). - docs/development/rebase-sensitive-invariants.md: both entries kept. - docs/state.md: the resolver; T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU stays closed (fixed on the WP1 branch, merged here). - Generated files took integration's side and were regenerated (vmafx-api.py --write, make docs-fragments-write, the ADR citation registry, where the lane's rename of core/test/vmafx_cuda_cells.h to vmafx_device_cells.h is carried). Fixes for the integration base: the HIP import sources join libvmafx_sources (the WP6 split, ADR-2094), and hip/import_device.c leaves the engine with vmafx_engine_leave(context, previous). CPU fast suite 387 passed, 0 failed.
This was referenced Oct 6, 2026
Draft
lusoris
force-pushed
the
rc4/api-wp3-cuda
branch
2 times, most recently
from
October 8, 2026 12:08
468f5cc to
c595358
Compare
lusoris
force-pushed
the
rc4/api-wp3-hip
branch
from
October 8, 2026 17:43
20a3f4d to
228b5af
Compare
lusoris
marked this pull request as ready for review
October 8, 2026 17:43
lusoris
force-pushed
the
rc4/api-wp3-hip
branch
from
October 9, 2026 09:34
228b5af to
142a898
Compare
…ent and sync_file fences (RC4 WP3, ADR-2092) (#2341) * feat(api): import HIP device frames, dma-bufs and GL textures with event and sync_file fences (RC4 WP3, ADR-2092) The VMAFx API now scores frames that already live on a HIP device without a copy through the host. A HIP device is opened by index or from the caller's stream; imports take device pointers at any offset and pitch, dma-bufs as external memory (the buffer's own size, released behind an event), HIP arrays, and OpenGL textures from a GLX context on the device's GPU. NV12, P010 and P016 are planarised by the conversion kernels the CUDA lane wrote, now shared by both backends. Every frame of a device is read on one library stream. The HIP twins were written for host pictures, so their upload becomes a device-to-device copy on that stream, and the reading twin's stream and the null stream wait for it; psnr_hvs_hip, ssimulacra2_hip and float_ms_ssim_hip no longer stage a device frame on the host. HIP_EVENT acquire fences are waits on the library stream; SYNC_FILE and GL_SYNC fences are checked on the host and waited for by the import rule, because ROCm 7.2.4 aborts on external semaphores. HOST and HIP_EVENT release fences are recorded where the last reference drops, through the release-event table now shared with CUDA. Each import submits its work with hipStreamQuery(), which the gfx1036 needs to keep later copies from overtaking it. enable_float_vif_hip_autodispatch now defaults to true, so float_vif can be scored on imported HIP frames and --backend hip runs float_vif_hip. Imported frames equal host uploads for every exact HIP twin on the Netflix pair, both checkerboards, Sparks 10-bit and 4K (354 cells, 16098 values); a skipped acquire wait and an early release fail under device load; the host-copy counter stays 0. * docs(api): move the HIP rebase note to a fragment and write the HIP pages in the internal register Under render at landing (ADR-2197) a pull request carries no rendered file: the rebase note of the HIP lane moves to docs/rebase-notes.d/vmafx-device-frames-hip.md and the ADR tag pages take master's text; the landing render writes them. The two HIP AGENTS pages this pull request edits drop their articles for the praetor caveman lint (article density at most 2 per 100 words). * ci(tidy): measure the HIP device-frame import's units in the cpu and cuda lanes The units this pull request adds or touches were measured on the rebased tree in the dev container (scripts/dev/tidy-lane.sh --write --only ..., clang-tidy 22.1.8), 0 findings and 0 uncited NOLINT in each lane: cpu: libvmaf.c, core/src/vmafx/{device,device_context,fence,frame_import, frame_import_admit,frame_pool,release_events,submit}.c and test_vmafx_fence_kinds.c (gl_sync.c and sync_file.c stay read by the cuda and hip lanes only: #2342 folds them into sync_object.c); cuda: the CUDA import sources and tests with the shared units (19). The hip lane's record of its 30 units is unchanged. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-wp3-hip
branch
from
October 9, 2026 10:42
142a898 to
e0f59be
Compare
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…lane Signed-off-by: Lusoris <lusoris@proton.me> #2341 added both files to every build, but its scoped cpu write did not list them in measured_sources, so the drift guard of the lane selectors read them as files only the cuda and hip lanes measure. A scoped cpu write in the dev container (scripts/dev/tidy-lane.sh --write --only, clang-tidy 22.1.8) measures both with no finding and adds them; no allowance changes.
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…lane Signed-off-by: Lusoris <lusoris@proton.me> #2341 added both files to every build, but its scoped cpu write did not list them in measured_sources, so the drift guard of the lane selectors read them as files only the cuda and hip lanes measure. A scoped cpu write in the dev container (scripts/dev/tidy-lane.sh --write --only, clang-tidy 22.1.8) measures both with no finding and adds them; no allowance changes.
12 of 18 tasks
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…lane Signed-off-by: Lusoris <lusoris@proton.me> #2341 added both files to every build, but its scoped cpu write did not list them in measured_sources, so the drift guard of the lane selectors read them as files only the cuda and hip lanes measure. A scoped cpu write in the dev container (scripts/dev/tidy-lane.sh --write --only, clang-tidy 22.1.8), run at this tip after the rebase took master's baseline, measures both with no finding and adds them; no allowance changes.
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…lane Signed-off-by: Lusoris <lusoris@proton.me> #2341 added both files to every build, but its scoped cpu write did not list them in measured_sources, so the drift guard of the lane selectors read them as files only the cuda and hip lanes measure. A scoped cpu write in the dev container (scripts/dev/tidy-lane.sh --write --only, clang-tidy 22.1.8), run at this tip after the rebase took master's baseline, measures both with no finding and adds them; no allowance changes.
14 of 15 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RC4 work package 3, HIP lane: the VMAFx API scores frames that already live on a HIP device (device pointers, dma-bufs, HIP arrays, and OpenGL textures where the runtime can read them) with no copy through the host. On the project's pinned ROCm 10.1.0, imported frames score bit for bit as the same frames uploaded from the host for every HIP twin declared exact. The PR is stacked on the CUDA lane #2277 (
rc4/api-wp3-cuda, head15bd091e2) and reuses its GL kinds, release callback and release-event table. The lane's choices are recorded in ADR-2092, now Accepted. The measurements behind them are in Research-2159.All evidence below comes from the pinned toolchain:
build-config.envROCM_BUILDERon master,rocm/dev-ubuntu-26.04:10.1.0-full@sha256:4f5ed1bf…(HIP runtime 7.16.26385). It ran in that image with the gfx1036 passed through. The lane was first built on the host's ROCm 7.2.4; those numbers stay as footnotes.What it adds:
vmafx_device_createwithVMAFX_BACKEND_HIPtakes either a device index (the library creates a stream) or the caller'shipStream_tinexternal[0]. HIP has no context object, soexternal[1]must be 0.vmafx_device_count,infoanddescribereport DEVICE_POINTER, DEVICE_ARRAY, GL_TEXTURE and DMABUF memory, and NONE, HOST, HIP_EVENT, GL_SYNC and SYNC_FILE fences.vmafx_context_use_device, features registered on the context run on their HIP twins.VMAF_PICTURE_BUFFER_TYPE_HIP_DEVICEpicture that carries the library stream.psnr_hvs_hip,ssimulacra2_hipandfloat_ms_ssim_hiplevel 0 used to stage planes on the host; they now copy or convert on the device.core/src/vmafx/import_convert_kernels.h, built by both nvcc and hipcc.VMAFX_E_RANGE.HIP_EVENTacquire fence is a wait on the library stream.SYNC_FILEandGL_SYNCacquire fences are checked on the host. If one is unsignalled the import isVMAFX_E_BUSY, and the D8 rule waits and retries.HOSTandHIP_EVENTrelease fences are signalled after the last reader, through the shared release-event table (core/src/vmafx/release_events.c).SYNC_FILErelease fence is refused withVMAFX_E_NOTSUP. ROCm 10.1 has no sync_file semaphore type: it refuses a DRM syncobj as an opaque descriptor (hipErrorNotSupported) and as a timeline descriptor (hipErrorInvalidValue). ROCm 7.2.4 aborts the process on the opaque descriptor.hipMemcpy2DFromArray(Async),hipMemcpyParam2DAsyncandhipMemcpy3DAsyncreturnhipErrorInvalidValue.VMAFX_E_NOTSUP, namingdesc.memory, the texture's extent and the runtime version. ROCm 7.2.4 reads these textures.float_vif_hipregistered by default.enable_float_vif_hip_autodispatchnow defaults totrue, per the maintainer popup of 2026-10-06 ("Default on (Recommended)").The ABI stays at
0.1.4.--abi-check --against-ref origin/rc4/api-generation-prototypereportsdefinition is an append-only successor of origin/rc4/api-generation-prototype (141 additions), and the same check againstorigin/rc4/api-wp3-cudareports(0 additions).python3 scripts/codegen/vmafx-api.py --check:38 generated files match core/api/vmafx.toml.Found on the way (rows in
docs/state.md; none of them is on master, so no master PR was opened):test_device_target_header_dependencies.pycounted onlyfeature_src_dirkernels. It failed in a CUDA build of the CUDA lane (22 fatbins against 21) and never checked the import kernels' headers. Fixed in commit 2 (T-DEVICE-HEADER-TEST-IMPORT-KERNELS-UNCOUNTED-2026-10-06).test_vmafx_api_generatorcannot see a planted wrong version node and fails. The checker's guess thatnmprints no versions is wrong there. Recorded but not changed here, because those are WP1 files (T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06).build-config.envonrc4/api-wp3-cudais behind master; the restack picks up 10.1.0.Landing (Q-083)
Lands bottom-up after #2287 (WP4) and #2290 (incremental motion window), one API PR at a time. The PR's whole diff (old base
15bd091e2, the #2277 head it was written on, to20a3f4d4d) was squashed into one commit, put on the restacked chain on 2026-10-07 and is now rebased onto masterb037dffb8, after #2640 took #2290's change out of #2615's squash and #2290 re-landed as its own commit (e584c79a6, maintainer decision Q-307).ABI check against master:
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/master->definition is an append-only successor of origin/master (0 additions). This PR changes the documentation of existing entries only, soabi_versionstays 0.1.8 and no entry is renumbered.Conflicts, resolved per hunk when the PR was first put on the chain:
core/api/vmafx.toml: this PR's doc text with the parent's ABI numbers.core/src/picture.h: this PR's comment onVMAF_PICTURE_BUFFER_TYPE_HIP_DEVICE(master had only renumbered the ADR in the old one).core/test/test_device_target_header_dependencies.py: the HIP count 21 and thesrc_dirroot of this PR on top of the CUDA lane's count 22.docs/api/vmafx/index.md,docs/development/rebase-sensitive-invariants.md: both sides kept (window scores of feat(api): add VMAFx window scores, the window clock and the in-flight bound (RC4 WP4, ADR-2074) #2287 and the CUDA / HIP imports; both invariant entries).docs/state.md: the resolver, thenT-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01keeps master's label and gains this PR's ROCm 10.1 sightings;T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06stays closed, as the generator PR closed it.Rebased onto #2290 and master since: the generated
core/src/AGENTS.mdindex (regenerated);core/src/vmafx/fence.c, where #2637 (c67a5240b) made a host fence wait on a condition variable: master's include of<pthread.h>and its waiting paragraph are kept, together with this PR's lane paragraph, which git merged;docs/state.mdby the resolver; the clang-tidy records, which took master's side and were measured again (below). The rebase note is the fragmentdocs/rebase-notes.d/vmafx-device-frames-hip.md(ADR-2197); the rendered files (CHANGELOG.md, ADR indexes,docs/rebase-notes.md) are master's.Fixed while landing (the stack moved under the PR):
core/src/meson.buildappended the HIP import sources tolibvmaf_sources, which the WP6 split renamed tolibvmafx_sources, so a HIP build did not configure.core/src/hip/import_device.ccalledvmafx_engine_leave(previous); since WP4 it takes the context.core/src/vmafx/import_convert_kernels.h(shared by CUDA and HIP) had an anonymous namespace and__global__definitions in a header: 4 clang-tidy findings in the hip lane. The header now holdsstatic __device__ __forceinline__bodies andimport_convert.cu/import_convert.hipdefine theextern "C"entry points, asfeature/float_moment_sum_gpu.hdoes.core/test/test_hip_shared_frame_contract.py: master's ruff (PLR2004) refused a literal 2; it is a named constant now.core/src/AGENTS.d/vmafx-device-frames.mdnamed the release callback "ABI 0.1.4"; the definition says 0.1.7.ci(tidy): the cpu lane's baseline records 10 of this PR's units on top of feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) #2290's (the earlier record fell away in the restack); the cuda lane's records its 19.Rebased onto master
b037dffb8after #2290 landed: the feature and docs commits applied unchanged (git range-diff:=); the cpu clang-tidy baseline, which #2290, #2646, #2647 and #2643 had moved, took master's side and the 10 cpu units were measured again (0 findings); the cuda baseline is this PR's record, master has not changed it since.No integration-branch commit is carried.
Local gate (on
142a89892, masterb037dffb8)-Db_lto=false,-j4, warnings as errors): build 0 warnings;--suite=fast428 OK, 2 failed:test_vmafx_window_liveandtest_vmafx_window_cli, the wall-clock budget under host load (T-WINDOW-LIVE-LATENCY-UNDER-LOAD-2026-10-09, moved to the serialtimingsuite by fix(test): check the live window budget alone in a timing suite and its pacing on a virtual clock #2664); both pass on a re-run through the runner and 3 of 3 direct runs.test_gpu_picture_pool_uafon its own withMALLOC_PERTURB_=0: OK; codegen tests 207 passed;make test-netflix-golden GOLDEN_NINJA_JOBS=4280 passed, 3 skipped;preflight.sh --stage msvcismpass; affected suites: tooling 2648 passed, 6 skipped.vmafx-api.py --check: 80 generated files match;--abi-check --against-ref origin/master: 0 additions.vmafx-hip-lane:rocm10.1.0-egl(the pinnedrocm/dev-ubuntu-26.04:10.1.0-full@sha256:4f5ed1bf…, HIP 7.16.26385, plus the EGL / GLES / X11 development libraries; ROCm unchanged), gfx1036, every device runflock hip-gfx1036.lock timeout 290. Build (-Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036,-Dwerror=true, fatal link warnings): 0 warnings.test_vmafx_import_hip16/16;test_vmafx_import_hip_fence8/8;test_vmafx_import_hip_glskipped with the ROCm 10.1 refusal, as designed (feat(hip): import GL textures through EGL dma-buf export so the pinned ROCm 10.1 reads them (ADR-2132) #2360 lands the EGL route);test_vmafx_fence_kinds5/5;test_hip_shared_frame9/9; with feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) #2290's advance callback on master underneath:test_hip_motion_five_frame_window,test_hip_motion_parity,test_hip_motion_v2_paritypass.test_vmafx_import_hip_bitexact, one run per clip (0 to 3 and the 4K pair): 354 cells, 16098 values, 0 cells differing, 0 attempts repeated, 9042 imports, 6028 conversions, 0 host copies.scripts/dev/tidy-lane.sh --write --only ...), measured on this tree (cpu again after the rebase ontob037dffb8): cpucore/src/libvmaf.c,core/src/vmafx/{device,device_context,fence,frame_import,frame_import_admit,frame_pool,release_events,submit}.c,core/test/test_vmafx_fence_kinds.c(10 TUs) 0 findings, recorded in the cpu baseline (gl_sync.candsync_file.c, 0 findings as well, stay read by the cuda and hip lanes only: feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342 folds them intosync_object.c); cuda the CUDA import sources and tests with the shared units (19 TUs) 0 findings; hip the 30 units of the HIP import, the HIP twins this PR touches and the shared units, 0 findings (record unchanged).docs/state.mdtouch and rows, silent revert against master,praetorctl audit): pass.The earlier device evidence on this PR (2026-10-07, before the restacks) is in the sections below and in
T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01.Type
feat— new featureChecklist
make format && make lintis green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, FFmpeg patch stack, HISS audit, assertion density).-Denable_hip=true -Denable_hipcc=true -Dhip_gfx_targets=gfx1036): fast suite 451 OK, 1 skipped, and 1 failed (test_vmafx_api_generator, the WP1 row above). The fast suite includes the 72fast+gpuHIP twin tests and the lane's device-free contracts. The lane programs pass as listed below.--suite=fast: 366 OK.test_vmafx_import_cuda12/12,_fence6/6,_gl2/2,_bitexact236 cells with 0 differing.test_hip_exact_twin_matrixandtest_hip_v1_models_no_fallbackpass./cross-backend-diffand the worst ULP is ≤ 2. — Every exact HIP cell compares imported frames with host uploads at 0 ULP (test_vmafx_import_hip_bitexact). The twins' arithmetic is unchanged..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (EUPL-1.2, fork-authored).!orBREAKING CHANGE:and the migration path is documented below. — not a breaking change: no ABI change.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— ADR-2092, number claimed withscripts/adr/next-free.sh --claim; indexes regenerated.tidy: hip
core/src/hip/{import_device,import_dmabuf,import_fence,import_frame,import_gl,picture_hip,shared_frame}.c,core/src/feature/hip/{float_vif_hip,integer_ms_ssim_hip,integer_psnr_hvs_hip,ssimulacra2_hip}.c,core/src/libvmaf.c,core/src/vmafx/{device,device_context,fence,frame_import,frame_import_admit,frame_pool,gl_sync,release_events,submit,sync_file}.c,core/test/test_hip_shared_frame.c,core/test/test_vmafx_fence_kinds.c,core/test/test_vmafx_import_hip{,_bitexact,_fence,_gl}.c(28 TUs; through themvmafx_hip{,_internal}.h,vmafx_hip_test_util.handvmafx_device_cells.hare covered) 0 findings, 0 uncited NOLINT. After the ROCm 10.1 changes,import_frame.cand the four lane tests were re-measured: 0 findings. cpu: the touchedcore/src/vmafx/*.c,core/src/libvmaf.candtest_vmafx_fence_kinds.c(12 TUs) 0 findings. cuda:core/src/cuda/{import_device,import_fence,import_frame,import_gl}.c,core/test/test_vmafx_import_cuda{,_bitexact,_fence,_gl}.cand the sharedcore/src/vmafx/*.c(19 TUs) 0 findings. All runs used the dev container, clang-tidy 22.1.8 andscripts/dev/tidy-lane.sh --only ... <lane>. The first hip run found 18 findings and the first cuda run 1. All were fixed without NOLINT.scripts/dev/preflight.sh --stage msvcism: pass.core/test/test_win32_pthread_shim_contract.py: pass.scripts/ci/assertion-density.sh: pass.praetorctl audit: governance gates passed, with no HISS finding in a touched file.Bug-status hygiene (ADR-0165)
docs/state.mdupdated.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Golden gate (
GOLDEN_NINJA_JOBS=4 make test-netflix-golden): 280 passed, 3 skipped (core/build-goldenbuilt with gcc).Cross-backend numerical results
Measured on
ryzen-4090-arc's AMD gfx1036 (integrated GPU, Linux 7.2.9-1-cachyos), inside the pinned ROCm 10.1.0 image. The image had/dev/kfdand/dev/dripassed through, the render and video groups added, and the X display mounted. Every device run wasflock ~/.cache/vmafx-locks/hip-gfx1036.lock timeout 300 docker run … timeout 290 <test>. Other lanes loaded the host throughout (load average 13 to 48).test_vmafx_import_hip_bitexact: every cell ofscripts/ci/exact_twins.d/*.hipin three layouts — planar with odd offsets and pitches, semi-planar, and semi-planar in dma-bufs (through the D8 rule, with an exported sync_file)test_vmafx_import_hip_fenceHIP_EVENT: 0 bad of 16 with the wait for psnr and vif, 16 of 16 withVMAFX_TEST_SKIP_ACQUIRE_WAIT.SYNC_FILE(a dma-buf written through GBM, with its exported sync_file as acquire fence): 32 pending at import, 32 waited for, 0 bad. With the check skipped, 32 imports took an unsignalled fence.test_release_canaryVMAFX_TEST_EARLY_RELEASE, 16 of 16VMAFX_TEST_FORCE_HOST_COPY, 3test_one_import_two_contexts(vmaf_v1.0.16_3d0handvmaf_v0.6.1)test_vmafx_import_hipVMAFX_E_BUSY.test_vmafx_import_hip_glthe HIP runtime (version 71626385) maps texture 0 (array 640x360) but refuses to read it: hipErrorInvalidValue.2Vendor profiler.
rocprofv3 --memory-copy-trace --hip-runtime-trace --kernel-trace, run in the image on the checkerboard import sessions (VMAFX_TEST_IMPORT_ONLY=1 VMAFX_TEST_CLIPS=1, 432 imports):__amd_rocclr_copyBufferRect*, the device-to-device copies), 288 NV12 conversions and 90float_ms_ssim_hiplevel-0 conversions. It ran no memory copy at all — none with a host side.rocprofv3at finalisation (ring_buffer.cpp:106 mmap failed with errno 22). Their HIP API log (AMD_LOG_LEVEL=3, 6624 imports) shows 7056 device-to-device copies on the library stream and no other copy there.Re-tests requested on 10.1:
hipErrorNotSupported; timeline:hipErrorInvalidValuehipGLGetDevices()without a context or under another vendor's GLX context returnshipErrorInvalidValue, and a later GLX call workshipErrorInvalidValue; row copy or texture read faults the GPU), with Mesa 26.0.8 and 26.2.4VMAFX_E_NOTSUPnaming the runtimehipStreamQuery()after importGates shown failing on planted defects (measured on this branch)
VMAFX_TEST_SKIP_ACQUIRE_WAIT)test_vmafx_import_hip_fencetest_vmafx_import_hip_fenceVMAFX_TEST_EARLY_RELEASE)test_vmafx_import_hip_fenceVMAFX_TEST_FORCE_HOST_COPY)test_vmafx_import_hiphipStreamQuery()after the import removedtest_vmafx_import_hiptest_vmafx_import_hip_glVMAFX_E_DEVICE; the contract test refuses the sourcetest_vmafx_import_hip_bitexactfloat_vif_hipwithout the HIP flagtest_vmafx_import_hip_bitexacttest_vmafx_import_hiptest_release_fence_kinds: not signalled after the releasetest_vmafx_import_hiptest_release_callbackfails (the 5 s wait runs out)test_vmafx_import_hiptest_no_frame_poolsfailstest_hip_shared_frame(device-free fakes)test_vmafx_import_hip_contract.pytest_vmafx_import_cuda_cells_contract.pyvmafx/import_convert_kernels.hmissing from a shared-header listtest_device_target_header_dependencies.pyThe eleven source regressions the contract test refuses:
psnr_hvs_hip,ssimulacra2_hipandfloat_ms_ssim_hiplevel 0Performance (if
perforfeat)No timing claim (RC8). An import costs one device-to-device copy per plane into the twins' buffers, on the library stream. The shared frame makes that copy once per frame for every twin that reads the plane. Reading producer memory in place is the RC8 row
T-HIP-IMPORT-TWIN-DEVICE-COPY-2026-10-06.Deep-dive deliverables (ADR-0108)
docs/research/2159-vmafx-hip-device-frames.md: runtime behaviour on ROCm 10.1 (with 7.2.4 for comparison) — external memory and semaphores, events,hipFree, pools, GBM dma-bufs, unsubmitted work, HIP-GL setup and read-out, dropped commands — plus the exit evidence.docs/adr/2092-vmafx-hip-device-frames.md## Alternatives considered.AGENTS.mdinvariant note — newcore/src/hip/AGENTS.d/vmafx-device-frames.md;core/src/hip/AGENTS.d/host-picture-staging.md,core/src/feature/hip/AGENTS.d/picture-upload-sync.mdandcore/src/AGENTS.d/vmafx-device-frames.mdupdated;docs/development/rebase-sensitive-invariants.mdentry "VMAFx HIP frames are read on one stream per device".changelog.d/added/api-hip-device-frames.md,changelog.d/changed/float-vif-hip-autodispatch-default.md,changelog.d/changed/hip-twins-device-level0.md.docs/rebase-notes.d/vmafx-device-frames-hip.md(ADR-2197), "VMAFx device frames on HIP (RC4 WP3 HIP lane)".User documentation:
docs/api/vmafx/index.mdgains a "HIP devices" section (descriptors, memory kinds, the ROCm 10.1 GL refusal, fences, a dma-buf example, ordering, profiling).docs/backends/hip/overview.md,uploads.md,build-flags.md,vif.mdandfeatures.mdare updated to match.Reproducer
The image is the pinned
ROCM_BUILDER, plus meson, nasm, zimg, xxd, gbm, drm, X11 and Mesa GLX for the build and the tests:The fixtures are
python/test/resource/yuvandtestdata/bbb(4K). The dma-buf cases need libgbm and the device's render node, found through sysfs when/dev/dri/by-pathis missing, as in a container. The GL test needs an X display and a GLX driver for the HIP device's GPU, and setsDRI_PRIMEitself.Known follow-ups
build-config.envto ROCm 10.1.0. Until thencheck-silent-revert.pyruns againstorigin/rc4/api-wp3-cuda.T-HIP-ROCM10-GL-TEXTURE-READ-2026-10-06).T-HIP-VMAFX-NO-FRAME-POOLS-2026-10-06).T-HIP-IMPORT-TWIN-DEVICE-COPY-2026-10-06).T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06) belongs to the WP1 owner.hipStreamQuery()workaround stays until a ROCm or driver update passes 8 runs without it.Footnotes
On the host's ROCm 7.2.4 the same cells had 0 differing in every attempt, with 2 cells repeated once (1 value each, the platform defect). ↩
On ROCm 7.2.4: 42 values, 0 differing. ↩