Repository navigation
feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) - #2375
Merged
Merged
Conversation
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
…152) Brings the Vulkan import lane (#2375) onto the integration branch. The lane's first commit (02aa3a7, its own merge of the SYCL and HIP lanes) is not taken: both lanes are here already. Resolved per hunk: - core/api/vmafx.toml: the lane's twelve additions are ABI 0.1.10 here (0.1.5 on the lane, ADR-1897); generated files regenerated. - cuda/import_frame.c: the Vulkan check runs first; a packed or MSB layout is refused for an OPTIMAL Vulkan image as for CUDA arrays. - hip/import_frame.c: the GL dma-buf export (ADR-2132) and the Vulkan dma-buf path are both kept; the lane's VmafxHipGl registration struct is dropped, since the GL interop it served is gone here. - vmafx/internal.h: pooling and window declarations and the Vulkan declarations are both kept. - SYNC_FILE fences: vmafx_fence_destroy() closes the descriptor (ADR-2091 item 6), as the lane resolves it; the API index, the device-frames agents page, the rebase note and the SYCL sync_file test follow. - test_vmafx_import_vulkan_api: designated initialisers for the layout, which gained the packed fields. (cherry picked from commit cff8500)
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
The integration branch took the Vulkan import lane (#2375) at ABI 0.1.10, the number WP9's spec functions had. Per ADR-1897 the spec functions move to ABI 0.1.11 (definition, changelog fragment, ADR-2125, rebase note). Both branches added a Research-2162; the WP9 digest is Research-2163. The merge put integration's new rebase bullet under the WP9 section; it is back in the integration section. Generated files: one side, regenerated.
lusoris
force-pushed
the
rc4/api-wp3-cuda
branch
2 times, most recently
from
October 8, 2026 12:08
468f5cc to
c595358
Compare
This was referenced Oct 8, 2026
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…-2091 item 6 decides The WP3 lanes met with two rules for a SYNC_FILE fence: the SYCL lane (ADR-2091 item 6) closes the descriptor in vmafx_fence_destroy(), the HIP lane refused the call. rc4/integration kept the refusal when it merged the lanes and restored item 6 only with the Vulkan import (#2375). This PR carries item 6 itself, so no landed state contradicts its own ADR: vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor in every build (refused on Windows, which has no sync_files) for every lane, and still refuses a GL_SYNC fence (the producer's, glDeleteSync()). Test: test_vmafx_fence_kinds test_sync_file_states checks that destroy closes the descriptor; it fails on the refusal. test_vmafx_import_sycl's sync_file case destroys its acquire fence instead of closing it itself. Signed-off-by: Lusoris <lusoris@proton.me>
17 of 19 tasks
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
…GL textures (RC4 WP3, ADR-2091) (#2342) * feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) The VMAFx API now scores frames that already live on a SYCL device without a copy through the host, and they score bit for bit as the same frames uploaded from the host for every SYCL twin declared exact. - SYCL devices are the Level Zero GPUs, opened by index or in the context of the caller's queue; each has one in-order library queue with immediate command lists, the precaution of draft PR #2217 (evaluated: its i915 defect did not reproduce on xe). - USM planes of any address and pitch and linear dma-bufs are bound; Intel Y-tiled and Tile4 dma-bufs are de-tiled on the device with the address math the VA-surface import now shares (core/src/sycl/detile.h); NV12 / P010 / P016 are planarised on the device. GL textures are imported through an EGL dma-buf export. - Every SYCL twin that stages its own planes, the chroma twins included, copies a device picture on the device through vmaf_sycl_picture_read_plane(), which records the read on the frame. - Acquire and release are joins (an empty kernel with depends_on()), not barriers: a barrier on the event of a command behind a host task waits on the host (Research-2159). SYCL_EVENT acquire fences are waited on by the device; HOST, SYNC_FILE, GL_SYNC and a dma-buf's implicit write fences are checked on the host for the import rule's retry. HOST and SYCL_EVENT release fences and the release callback follow the last reader in every context. - The sync_file, dma-buf fence and GL sync helpers move to core/src/vmafx/sync_object.c, shared with the CUDA lane. - Windows shared textures and SYNC_FILE release fences are refused, tracked as T-SYCL-VMAFX-IMPORT-REMAINDER-2026-10-06. Landing (Q-083): the branch's diff (15bd091..5dafa1c) squashed onto (a082db7): the HIP lane's vmafx/gl_sync.c and sync_file.c folded into sync_object.c; SYNC_FILE fences borrowed (vmafx_fence_destroy() refuses them); dmabuf_import.cpp keeps master's driver_fd() (#2377) inside vmaf_sycl_dmabuf_import_queue(). Carried from rc4/integration: 44cbb8c (the SYCL EGL export folded onto vmafx/egl_export.c, VmafxEglTiled; the engine_leave signature) without its d3d11 stub, which the WP6 generator fix replaces, and the test of 869ec15 (an import closes only its own descriptors), whose library half master's #2377 already holds, and the SYCL parts of docs/api/vmafx/index.md that the integration merge dropped (restored there by 3b9a414), written for CUDA, SYCL and HIP. The rebase note is a fragment (ADR-2197). vmafx_sycl_producer.cpp is brought to 0 clang-tidy findings in the sycl lane, and core/src/compat/crt_portable.h's C spelling vmaf_tmpfile_portable(void), which the sycl lane reports through sycl/common.cpp, carries the ADR-1138 NOLINT that x86/avx512_warm_up.h has. * fix(api): destroy a SYNC_FILE fence by closing its descriptor, as ADR-2091 item 6 decides The WP3 lanes met with two rules for a SYNC_FILE fence: the SYCL lane (ADR-2091 item 6) closes the descriptor in vmafx_fence_destroy(), the HIP lane refused the call. rc4/integration kept the refusal when it merged the lanes and restored item 6 only with the Vulkan import (#2375). This PR carries item 6 itself, so no landed state contradicts its own ADR: vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor in every build (refused on Windows, which has no sync_files) for every lane, and still refuses a GL_SYNC fence (the producer's, glDeleteSync()). Test: test_vmafx_fence_kinds test_sync_file_states checks that destroy closes the descriptor; it fails on the refusal. test_vmafx_import_sycl's sync_file case destroys its acquire fence instead of closing it itself. * fix(test): key the one-EGL-export contract by POSIX paths so it holds on Windows test_vmafx_one_egl_export_contract.py keyed the library sources by str(path.relative_to(SRC)), which is "vmafx\sync_object.c" on Windows, so test_sync_objects_defined_once compared it with "vmafx/sync_object.c" and failed there. The keys are now relative_to(SRC).as_posix(), as the other contract tests do (#1850). scripts/dev/preflight.sh --stage msvcism, which refuses that pattern, failed on the branch and passes now. * ci(tidy): measure the SYCL device-frame import's units in the cpu, sycl, cuda and hip lanes On master after #2367: every lane's baseline is master's, and the units this pull request adds or changes were measured again in the dev container (scripts/dev/tidy-lane.sh --write --only ..., clang-tidy 22.1.8), 0 findings and 0 uncited NOLINT in each lane: - cpu: libvmaf.c, the shared core/src/vmafx/ units (device, device_context, egl_export, fence, frame_import, frame_import_admit, frame_pool, submit, sync_object) and test_vmafx_fence_kinds.c; - sycl: the shared units, the SYCL import sources and runtime, the SYCL feature twins this pull request touches and the import tests; - cuda: the shared units and core/src/cuda/import_gl.c; - hip: the shared units, core/src/hip/import_frame.c and import_gl.c. core/src/vmafx/gl_sync.c and sync_file.c, folded into sync_object.c, leave the cuda and hip lanes: a scoped write drops the entries of deleted files since #2636, so the whole-lane measurements the earlier record needed are gone. This replaces the pull request's two earlier tidy commits. * fix(build): compile the SYCL import tests' producer with the SYCL argument list Master's test_sycl_math_constants_contract (#2654) requires every icpx compile command to take sycl_common_args or sycl_feature_tail_args, which carry the Windows math-constant define. The vmafx_sycl_producer.cpp command of this pull request spelled out the parts of sycl_common_args without that define and failed the contract on the rebased tree. It now uses sycl_common_args itself; the contract test passes again. Signed-off-by: Lusoris <lusoris@proton.me>
14 of 15 tasks
lusoris
marked this pull request as ready for review
October 10, 2026 01:18
lusoris
force-pushed
the
rc4/api-wp3-vulkan-import
branch
from
October 10, 2026 01:25
8505d76 to
9c460c4
Compare
12 of 19 tasks
lusoris
added a commit
that referenced
this pull request
Oct 10, 2026
…an tests build and run on CPU in every lane (#2714) * fix(dev): put the Vulkan loader and lavapipe in the dev image so Vulkan tests build and run on CPU in every lane The dev image had no Vulkan loader, so meson never found dependency('vulkan') in it. The Vulkan frame-import tests of #2375 (ADR-2152) were not built there, and the hosted cuda, hip and sycl clang-tidy lanes, which run in the image's pinned digest, failed closed with "not measured" on the five test files #2375 records. Per operator decision ci-config-19 (ADR-3137): the build-deps stage installs libvulkan1, libvulkan-dev, mesa-vulkan-drivers (lavapipe, the CPU Vulkan driver) and vulkan-tools from the archive of the digest-pinned DEV_BASE, like the stage's other distribution packages, and the build fails when vulkaninfo --summary lists no lavapipe. One image, no tidy exception; real-GPU Vulkan runs stay a separate signal. The hosted lanes pick it up when the tidy-lane action's pinned digest moves to the published build of this change; #2375 moves it and re-measures its baselines. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-wp3-vulkan-import
branch
from
October 10, 2026 06:30
9c460c4 to
4b8610d
Compare
…152) (#2375) * feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) VMAFX_MEMORY_VULKAN and VMAFX_FENCE_VULKAN_SEMAPHORE with the handle type, tiling, flags and PCI location of the producer; VmafxFrameImport.acquire_more and VmafxDeviceInfo.pci; vmafx_frame_signal_on_release(). CUDA imports opaque memory as arrays or buffers and waits on and signals the producer's timeline on its stream; SYCL and HIP read LINEAR and DRM-modifier frames through their dma-buf paths. Memory of another GPU, multi-plane images, Windows handles, OPTIMAL tiling off CUDA and Vulkan semaphores off CUDA are refused by name. Import only (Q-011): the library links no Vulkan loader and makes no Vulkan call. A SYNC_FILE fence is destroyed by closing its descriptor (ADR-2091 item 6), which the SYCL lane (#2342) now lands; this lane adds the VULKAN_SEMAPHORE refusals of vmafx_fence_wait() / vmafx_fence_destroy() next to it. Landing (Q-083): the branch merges the SYCL and HIP lanes, so its tip is not squashed. The change is rc4/integration's squashed variant (798c42c, 541a91a, bae72c1, a21d20c: the lane's cff8500, 3decc47, 43b2231, c101d50 on the merged lanes) plus the lane's last commit 8505d76 (the release canary scores each frame in two contexts), cherry-picked onto #2368's landing head. The twelve definition additions are ABI 0.1.10 (0.1.5 on the lane, ADR-1897). ADR-2152 is Accepted (maintainer, Q-089). The rebase note is a fragment (ADR-2197); the Vulkan agent page is in the caveman register. The touched units are measured in the cpu, cuda, hip and sycl lanes (dev image with the Vulkan loader and headers): 0 findings. cuda/import_vulkan.c included <unistd.h> and called dup() / close() unconditionally, which the Windows MSVC+CUDA build cannot compile; it now uses VMAF_DUP (new in compat/crt_portable.h, _dup() on Windows) and VMAF_CLOSE, as the other C runtime calls of the tree do. * test(vulkan): judge the fence test's skipped wait only where no import waited for its writer The acquire arm of test_vmafx_import_vulkan_fence demanded bad frames from the planted skipped wait whenever any import returned before its write had finished. On the gfx1036 under ROCm 10.1 the import normally waits for the writer; one landing run of the HIP lane had 2 of 16 imports ahead, both frames still read finished (0 bad), and the arm failed although the library behaved correctly: a copy queued behind an early import can run after the write has landed. The arm now demands bad frames on CUDA, whose import never waits, and wherever every import ran ahead of its write (SYCL); it demands none where no import ran ahead, and prints a partial count without judging it. The real arm is unchanged. Not reproduced on demand (one run in 79 had imports ahead). With the new rule: gfx1036 40 of 40, RTX 4090 10 of 10 and Arc A380 10 of 10 (skipped wait 8 bad, 16 imports ahead, every run). * ci(tidy): measure the translation units of the Vulkan frame import in the cpu, cuda, hip and sycl lanes The tidy coverage rule requires every tracked translation unit to be read by a lane. The units this pull request adds or touches were measured on the rebased tree in the dev image with the Vulkan loader and headers (scripts/dev/tidy-lane.sh --image vmaf-dev-mcp:local-vulkan --write --only ..., clang-tidy 22.1.8): cpu 8, cuda 16, hip 15 and sycl 15 translation units, 0 findings and 0 uncited NOLINT in each lane. The earlier record of these units fell away when the branch was rebased onto the restacked SYCL lanes. * ci(impact): route the Vulkan import tests to the cuda, hip and sycl tidy lanes Master's test_ci_impact (TidyLaneRouting, from #2704) requires every file a GPU tidy lane measures and the cpu lane does not to match that lane's tidy_<lane> patterns in .github/ci-impact.json. This pull request's Vulkan import tests (core/test/test_vmafx_import_vulkan*.c and vmafx_vulkan_producer.c) are measured by the cuda, hip and sycl lanes only and matched none, so the test failed on the rebased tree. core/test/*vulkan* joins the three lanes' patterns; the test passes. * test(vulkan): skip with the reason and run no case when the lane has no GPU Vulkan device On a CPU-only lane (a hosted runner whose only Vulkan device is lavapipe) the Vulkan import tests exit 77, as maintainer decision Q-346 settles: the import reads a GPU's Vulkan memory, which lavapipe does not have, and the real-GPU runs are the device signal. The skip now says so, and it never reads like a pass: - vk_open() names the reason on stderr: "skipped: no CUDA device, so no GPU Vulkan memory to import (lavapipe is a CPU device; the producer needs the importing GPU's memory)", or that no GPU Vulkan device sits on the lane device; it names the lane (CUDA, SYCL, HIP). - test_vmafx_import_vulkan, _bitexact and _fence run no case without the GPU, so no case prints "pass" for work it did not do; the harness reports "0 tests run, skipped" and exits 77. The FFmpeg test names its two other reasons (inputs not set, no Vulkan decode on the GPU) the same way. With the GPU the tests run as before (CUDA lane: import 5/5, fence 2/2). * ci(tidy): run the hosted tidy lanes in the dev image that carries the Vulkan loader #2714 (ADR-3137, operator decision ci-config-19) put the Vulkan loader, its headers and lavapipe into the dev image. dev-container-publish.yml published that build of master 9d3734f as ghcr.io/vmafx/vmafx-dev-mcp:sha-9d3734f0602a3114849d8edce4136d9c8e7c850a, digest sha256:29b13030cacce60cec4209f645c6772feda5b2d936fa511fc807eddccf12722c (read from the registry with docker buildx imagetools inspect). .github/actions/tidy-lane pins it, as its description asks of the pull request that re-measures the baselines after a dev/Containerfile change. Measured in that image (scripts/dev/tidy-lane.sh --image <digest> --write --only ...): the cuda (16), hip (15) and sycl (15) units of this pull request, the five Vulkan test files among them, 0 findings, 0 uncited NOLINT, no compile failure; the records equal the ones already committed. * test(vulkan): say when the FFmpeg test's inputs are not readable With VMAFX_TEST_VULKAN_REF / _DIST set to files that do not exist the skip said "no Vulkan decode of the inputs on this GPU", which blamed the device. It now says "skipped: the inputs are not readable (<ref>, <dist>)" and keeps the decode reason for inputs that exist but do not decode. The cuda, hip and sycl tidy lanes measured the file in the published dev image: 0 findings, records unchanged. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-wp3-vulkan-import
branch
from
October 10, 2026 11:08
4b8610d to
db110bd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RC4 work package 3, Vulkan import lane: frames a Vulkan producer exported (FFmpeg's Vulkan decode and filters, libplacebo, a GStreamer Vulkan element) are imported into CUDA, SYCL and HIP devices with no copy through the host, and score bit for bit as the same frames uploaded from the host on every exact cell of each backend. Import only, per maintainer decision Q-011: the library links no Vulkan loader and makes no Vulkan call. The decision is ADR-2152 (Proposed); the measurements behind it are Research-2162.
Base. The branch starts at
rc4/api-wp3-common(027aebc56) and first merges the SYCL lane (#2342,rc4/api-wp3-sycl,5dafa1cbd) and the HIP lane (#2341,rc4/api-wp3-hip,20a3f4d4d), because the Vulkan import runs on both. Both lanes are stacked on the CUDA lane (#2277), so this PR's base isrc4/api-wp3-cuda, as theirs is: the diff holds the two lanes, the merge and the Vulkan commits; review the Vulkan work commit by commit after the merge (cff8500bband later). The merge commit (02aa3a7d2) keeps one copy of each shared helper:core/src/vmafx/sync_object.{c,h}(the HIP lane'ssync_file.candgl_sync.cfolded in) andrelease_events.c; the resolution is described indocs/rebase-notes.md. Correction to the merge commit's message: it saysvmafx_fence_destroy()refuses aSYNC_FILEfence. That is wrong for this branch:vmafx_fence_destroy()closes aSYNC_FILEdescriptor, the SYCL lane's decision (ADR-2091 item 6), which the HIP lane's refusal gave way to incff8500bb;test_vmafx_fence_kindschecks that the descriptor is closed, and the API page and the changelog say so.What it adds (ABI 0.1.5, append-only):
VMAFX_MEMORY_VULKANwithVmafxFrameImport.vulkan_handle_type(OPAQUE_FD,DMA_BUF; the Windows types declared and refused),vulkan_tiling(OPTIMAL,LINEAR,DRM_FORMAT_MODIFIER),vulkan_flags(VMAFX_VULKAN_DEDICATED) andvulkan_pci.VMAFX_FENCE_VULKAN_SEMAPHORE: an exported timeline semaphore and a value.acquire_more[2]: the further acquire fences of a frame written under several timelines (one image per plane, asAVVkFrame.sem[i]).VmafxDeviceInfo.pci: each device's PCI location, so a producer picks the same GPU.vmafx_frame_signal_on_release(): the device signals the producer's timeline behind the frame's last reader in every context.core/src/cuda/import_vulkan.c): opaque-fd memory of OPTIMAL images as CUDA arrays (NV12 / P010 / P016 planarised on the device; a planar OPTIMAL frame needsVMAFX_IMPORT_ALLOW_COPY, as CUDA arrays do), of LINEAR images and buffers as device pointers; the acquire timelines waited on the library stream; the release signalled there.vmafx_import_vulkan_as_dmabuf()), with async_fileof the producer's write as the acquire fence and a HOST release fence.desc.vulkan_pci), a plane of a multi-plane image (desc.plane[1].plane_index), Windows handles, OPTIMAL tiling on SYCL and HIP,DMA_BUFand DRM-modifier tiling on CUDA (the CUDA driver imports dma-bufs on Tegra only), a modifier the device does not read, Vulkan semaphores on SYCL and HIP (ROCm 10.1 imports none; the SYCL lane does not wait on one yet).vmafx_fence_wait()/vmafx_fence_destroy()refuse aVULKAN_SEMAPHOREfence.python3 scripts/codegen/vmafx-api.py --check: 38 generated files match.--abi-check --against-ref origin/rc4/api-generation-prototype:definition is an append-only successor of origin/rc4/api-generation-prototype (153 additions); againstorigin/rc4/api-wp3-common: 22 additions.Landing (Q-083)
On master (2026-10-10)
Rebased onto master
9d3734f06, after #2368 (f6513a6de) and #2714 (c96b8204e, the dev image with the Vulkan loader, ADR-3137) landed. Only this PR's commits are carried; every commit of #2368, #2342, #2367, #2360, #2341 and #2290 is gone, among them4613fc64fwith the leaked identity. Every commit is by and signed off aslusoris@proton.me.docs/state.md: the RC4 label row conflicted; it is master's list with this PR's fourT-VULKAN-*rows added where the PR put them..github/actions/tidy-lanepins the published build of fix(dev): put the Vulkan loader and lavapipe in the dev image so Vulkan tests build and run on CPU in every lane #2714,ghcr.io/vmafx/vmafx-dev-mcp:sha-9d3734f0602a3114849d8edce4136d9c8e7c850a=sha256:29b13030cacce60cec4209f645c6772feda5b2d936fa511fc807eddccf12722c(read from the registry withdocker buildx imagetools inspect). Measured in that image (scripts/dev/tidy-lane.sh --image <digest> --write --only ...): cuda 16, hip 15, sycl 15 units of this PR, the five Vulkan test files among them, 0 findings, 0 uncited NOLINT, no compile failure; cpu 8 units (measured on the earlier restack, unchanged). No exception.test_ci_impactTidyLaneRouting(ci(tidy): defer changed files to the lane that measures them and lint C-only SYCL headers through their includers (Q-341, Q-342) #2704) wanted the Vulkan tests in the cuda, hip and sycl patterns of.github/ci-impact.json;core/test/*vulkan*joins them.abi_version0.1.9 -> 0.1.10.Local gate (on
b34c7c5c3, master9d3734f06; the last commit4b8610dd7changes only the FFmpeg test's skip message: rebuilt, checked on CUDA with unreadable inputs, and its file measured again in the three GPU lanes, 0 findings)-Db_lto=false,-j4, one link at a time, warnings as errors): build 0 warnings;--suite=fast441 OK, 0 failed;test_gpu_picture_pool_uaf(MALLOC_PERTURB_=0) OK; codegen tests 209 passed;make test-netflix-golden280 passed, 3 skipped;preflight.sh --stage msvcismpass; affected suites: tooling 2748 passed, 0 failed, 6 skipped.vmafx-api.py --check: 80 generated files match.--suite=fast514 OK, 0 failed;test_vmafx_import_vulkan_cuda5/5,_fence2/2 (0 repeated arms),test_vmafx_import_vulkan_cuda_bitexactNetflix and bbb pass; regressiontest_vmafx_import_cuda12/12,_fence6/6,_gl2/2,_bitexactpass,test_vmafx_fence_kinds5/5. Without the GPU (CUDA_VISIBLE_DEVICES=on the same build):test_vmafx_import_vulkan_cuda,_bitexactand_fenceeach print the "no CUDA device" reason above and "0 tests run, skipped", exit 77.test_vmafx_import_vulkan_sycl5/5,_fence2/2,test_vmafx_import_vulkan_sycl_bitexactNetflix and bbb pass; regressiontest_vmafx_import_sycl17/17,_fence6/6,_gl2/2, bitexact Netflix and bbb pass,test_sycl_kernel_scratchpass,test_vmafx_fence_kinds5/5.vmafx-hip-lane:rocm10.1.0-vulkan(ROCm 10.1.0), gfx1036: build 0 warnings;test_vmafx_import_vulkan_hip5/5,_fence2/2,test_vmafx_import_vulkan_hip_bitexactNetflix and bbb pass; regressiontest_vmafx_import_hip16/16,_fence8/8,test_vmafx_fence_kinds5/5,test_hip_shared_frame9/9.*_ffmpegtests (FFmpeg-decoded Vulkan frames) skip here: their FFmpeg n9.0.2 + Vulkan inputs were removed by a disk clean-up (CUDA and SYCL now say "the inputs are not readable"; the HIP build does not register the test without FFmpeg's Vulkan). Their device evidence is the earlier run recorded below.docs/state.mdtouch and rows, silent revert,praetorctl audit): pass.git merge-tree --write-tree origin/master HEAD: clean.The sections below record how the PR was first put on the stack and the evidence measured there.
Lands bottom-up after #2368 (SYCL 4:2:2 / 4:4:4), the last WP3 lane, one API PR at a time. The branch first merges the SYCL lane (#2342) and the HIP lane (#2341), which landed before it, so its tip is not squashed onto its base (that would bring both lanes again). What is cherry-picked onto the landing head of #2368 is the change
rc4/integrationcarries for this PR, where both lanes were already in:798c42c5d(the lane'scff8500bb),541a91a01(3decc47da),bae72c148(43b2231f3),a21d20ce1(c101d5030), and the lane's last commit8505d76a7(the release canary scores each frame in two contexts). Compared file by file with the lane's own change (02aa3a7d2..8505d76a7, after its merge commit): the same 79 files; the differences are the ABI number, the merge's single copies the lanes landed with, and what the lanes' landing already changed (the HIP GL export, the 4:2:2 / 4:4:4 layouts, and the SYNC_FILE destroy of ADR-2091 item 6, which #2342 now carries).ADR-2152 was Proposed on the branch; the maintainer accepted it as written (ledger
Q-089, 2026-10-07): this PR flips it (status line, deciders, a## Referencesentry citing Q-089, the index fragment). Its frozen body names ABI 0.1.5, the number free when the lane branched.ABI check against the parent:
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref land-2368->definition is an append-only successor of land-2368 (18 additions).abi_versiongoes from 0.1.9 to 0.1.10 in nodeVMAFX_0.1; the new entries (VMAFX_MEMORY_VULKAN,VMAFX_FENCE_VULKAN_SEMAPHORE,VmafxVulkanHandleType,VmafxVulkanTiling,VmafxVulkanFlags,VmafxDeviceInfo.pci, theVmafxFrameImportfieldsacquire_more,vulkan_handle_type,vulkan_tiling,vulkan_flags,vulkan_pci,vmafx_frame_signal_on_release()) say "Added in ABI 0.1.10"; no entry of the parents is renumbered.Conflicts, resolved per hunk:
core/src/AGENTS.d/vmafx-device-frames.md: the parents' page, unchanged here; its fence rule (destroy closes a SYNC_FILE fence's descriptor, ADR-2091 item 6; a GL_SYNC fence stays refused) lands with feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342.core/src/hip/AGENTS.d/vmafx-device-frames.md,docs/backends/sycl/zero-copy.md: the parents' wording plus this PR's addition (onesync_object.ccopy shared with SYCL; "Vulkan memory" in the SYCL ingestion row).docs/state.md: the resolver; the RC4 label row keeps the parents' description and gains this PR's four Vulkan rows; the RC8 and deferred label rows gainT-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06andT-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06.Changed while landing:
core/src/cuda/import_vulkan.cincluded<unistd.h>and calleddup()/close()unconditionally; the required Windows MSVC+CUDA build has neither (andscripts/dev/preflight.sh --stage msvcismdoes not look for it). It now callsVMAF_DUP()(added tocore/src/compat/crt_portable.h:_dup()on Windows,dup()elsewhere) andVMAF_CLOSE(), the tree's spelling for the C runtime calls Windows renames. The CUDA build and every CUDA test below were run again after the change (same numbers); the SYCL and HIP builds compile neither file's changed code.vmafx_fence_destroy()for it,test_vmafx_fence_kinds,test_vmafx_import_sycl's sync_file case or the API page's fence paragraph. Its changelog fragment drops the sentence, and its rebase note says the rule landed with feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342. After the restack onto the new parents the code of this PR is the gated head's (f0c73a33d): the tree diff between the two heads is four fragment files.docs/rebase-notes.d/vmafx-vulkan-frame-import.md;core/AGENTS.d/vmafx-vulkan-frames.mdis in the caveman register (praetorctl caveman check --kind=context: pass;caveman_keep_check.py: 0 missing).Local gate on the landing head (CPU,
-Db_lto=false,-j4, warnings as errors; host Vulkan loader and headers 1.4.363, so the Vulkan tests are built): build 0 warnings;--suite=fast420 OK, 0 failed (test_vmafx_import_vulkan_apiincluded;test_gpu_picture_pool_uafwithMALLOC_PERTURB_=0, #2547: 1 OK); codegen tests 124 passed;make test-netflix-golden GOLDEN_NINJA_JOBS=4280 passed, 3 skipped;preflight.sh --stage msvcismpass; affected suites: mcp 590 passed, tooling 2382 passed (6 skipped), 0 failed. Platform check: the library makes no Vulkan call;vmafx/frame_import_vulkan.ckeeps its dma-buf translation behind__linux__and refuses the Windows handle types by name; the Vulkan tests are built only wheredependency('vulkan')is found.Train gates against the parent head (
BASE_SHA=land-2368faffc033d, the stack's scripts with ADR-2197):deliverables-check.sh,state-md-touch-check.sh,check-state-md-rows.sh,check-silent-revert.py(clean) andpraetorctl audit(engine7458a220) all rc 0.Device runs, each under its lock with
timeout 290inside (the bit-exactness programs in theirVMAFX_TEST_ONLY=netflixandbbbhalves to fit it); the FFmpeg tests decode the Vulkan lane's two 576x324 H.264 inputs.VMAF_DUPchange: build 0 warnings;--suite=fast493 OK, 0 failed.test_vmafx_import_vulkan_cuda_bitexact376 + 96 = 472 cells, 21464 values, 0 differing, 0 repeated attempts (12056 imports, 6028 conversions);_fenceacquire: skipped wait 8 bad (16 imports ahead of their write), with the wait 0 bad of 8; release: planted early 4 bad, real 0 of 8;_ffmpeg48 frames, 23 cells, 3936 values, 0 differing (the decoder's own two-plane image refused namingplane_index);test_vmafx_import_vulkan_cuda5 of 5. Lane regression:test_vmafx_import_cuda12 of 12,_fence6 of 6,_gl2 of 2,_bitexact236 cells, 10732 values, 0 differing;test_vmafx_fence_kinds5 of 5.vmaf-dev-mcp:local-vulkan(the pinned dev image plus the Vulkan loader and headers 1.4.341; ANV from Mesa 26.0.8, Vulkan 1.4.335;ANV_DEBUG=video-decode), Arc A380 at 0000:03:00.0 (xe); build undersycl-build.lock, 0 compiler warnings (two ocloc RetryManager notes for the untouchedSs2SsimUnitsKernelandTermKernel).test_vmafx_import_vulkan_sycl_bitexact376 + 96 = 472 cells, 21464 values, 0 differing, 0 repeated attempts (12056 imports, 9042 conversions);_fenceacquire: skipped wait 8 bad (16 ahead), with the wait 0 bad of 8; release: planted early 8 bad, real 0 of 8;_ffmpeg48 frames, 23 cells, 3936 values, 0 differing;test_vmafx_import_vulkan_sycl5 of 5. Lane regression:test_vmafx_import_sycl17 of 17,_fence6 of 6,_gl2 of 2,_bitexact188 + 48 cells, 9712 + 1020 values, 0 differing;test_vmafx_fence_kinds5 of 5;test_sycl_kernel_scratch132 kernels, 0 use scratch memory;VMAF_SYCL_AOT_JOBS=8 meson test --suite sycl-aot: OK, 36 SYCL translation units for 19 targets, 220.8 s.vmafx-hip-lane:rocm10.1.0-vulkan(the pinned image plus the Vulkan loader and headers 1.4.341; RADV from Mesa 26.0.8), gfx1036 at 0000:7d:00.0: build 0 warnings.test_vmafx_import_vulkan_hip_bitexact376 + 96 = 472 cells, 21464 values, 0 differing in every attempt, 2 attempts repeated once (the gfx1036's dropped dispatches,T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01);_ffmpeg48 frames, 23 cells, 3936 values, 0 differing;test_vmafx_import_vulkan_hip5 of 5;_fence5 of 6 runs pass (acquire: 0 bad with and without the wait, 0 imports ahead, asT-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06describes; release: planted early 8 bad, real 0 of 8). The first run failed its acquire arm: 2 imports returned ahead of their write and still read finished frames (0 bad), and the test asserts that a skipped wait is seen whenever an import is ahead; see Known follow-ups. Lane regression:test_vmafx_import_hip16 of 16,_fence8 of 8,test_vmafx_fence_kinds5 of 5,test_hip_shared_frame9 of 9. FFmpeg for the Vulkan decode tests: the Vulkan lane's own build (FFmpeg n9.0.2 libraries, 63.1.102) through-Dpkg_config_pathin the SYCL and HIP builds, as the lane built them.tidy (dev image with the Vulkan loader and headers, clang-tidy 22.1.8,
scripts/dev/tidy-lane.sh --image vmaf-dev-mcp:local-vulkan --jobs 4 --write --only ...): cpucore/src/vmafx/{device,fence,frame_import,frame_import_vulkan,sync_object}.c,core/test/{test_vmafx_abi_layout,test_vmafx_fence_kinds,test_vmafx_import_vulkan_api}.c(8 TUs) 0 findings; cuda the five sharedvmafx/*.c,core/src/cuda/import_{device,fence,frame,vulkan}.c,core/test/test_vmafx_import_cuda{,_bitexact}.cand the five Vulkan test sources (16 TUs) 0 findings, andimport_vulkan.cagain after theVMAF_DUPchange: 0 (cpulibvmaf.canddict.cppthrough the changedcrt_portable.h: 0); hip the five shared,core/src/hip/import_{device,fence,frame}.c,core/test/test_vmafx_import_hip{,_bitexact}.cand the five Vulkan test sources (15 TUs) 0 findings; sycl the five shared,core/src/sycl/import_{device,fence,frame}.c,core/src/sycl/vmafx_sycl_rt.cpp,core/test/test_vmafx_import_sycl.cand the five Vulkan test sources (15 TUs) 0 findings.check-tidy-coverage.py: every tracked translation unit is read or excepted.Type
feat— new featuresycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, HISS audit and evidence replay with the branch's pinned engine0af07a733e65).--suite fast --no-suite gpu: 377 OK, 0 fail, 1 skipped (test_vmafx_api_abi_append_only, 77: its reference commit predatescore/api/vmafx.toml); every Vulkan and lane device program per backend in the table below./cross-backend-diffand the worst ULP is ≤ 2. — no extractor changed; every exact cell of every backend is compared value for value against host uploads below: 0 differing..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (EUPL-1.2, fork-authored).!orBREAKING CHANGE:and the migration path is documented below. — not a breaking change: append-only.docs/adr/_index_fragments/<NNNN-slug>.md— ADR-2152 (claimed withscripts/adr/next-free.sh --claim), fragment added; the index, by-tag and title pages are rendered at landing (ADR-2197).tidy (dev container
vmaf-dev-mcp:localplus the Vulkan headers and loader, clang-tidy 22.1.8,scripts/dev/tidy-lane.sh --only ... <lane>,--jobs 4): cpu lane, 8 units (core/src/vmafx/{device,fence,frame_import,frame_import_vulkan,sync_object}.c,core/test/{test_vmafx_abi_layout,test_vmafx_fence_kinds,test_vmafx_import_vulkan_api}.c): 0. cuda lane, 11 units (core/src/cuda/import_{device,fence,frame,vulkan}.c,core/test/test_vmafx_import_cuda{,_bitexact}.cand the five Vulkan test sources): 0. hip lane, 11 units (core/src/hip/import_{device,fence,frame,gl}.c,core/test/test_vmafx_import_hip{,_bitexact}.c, the Vulkan test sources): 0. sycl lane, 10 units (core/src/sycl/import_{device,fence,frame}.c,vmafx_sycl_rt.cpp,core/test/test_vmafx_import_sycl.c, the Vulkan test sources): 0 in this PR's files, 0 uncited NOLINT; the 7 findings left are incore/include/libvmaf/picture.h, untouched here and suppressed on master by #2186 (this stack's base predates it), as #2342 reports. The first runs found 12 in this PR's files (analyzer array bounds and a NULL path in the CUDA release-signal registration,sscanfin three PCI parsers, an increment in a condition, two misplacedconst, an enum zero-initialised, an implicit widening, a test function's size); all fixed, the three parsers became one (vmafx_parse_pci_bus_id()), and the changed units re-measured at 0.scripts/dev/preflight.sh --stage msvcism: pass. Pre-push gate (lefthook, the branch's praetor engine): pass, MkDocs strict build included; its first run caught two findings, fixed inc101d5030(two fork-added functions of 20+ lines without an assert; a computed test source name the stale-reference check could not resolve).Bug-status hygiene (ADR-0165)
docs/state.mdupdated — opened T-VULKAN-IMPORT-SYCL-HOST-FENCES-2026-10-06, T-VULKAN-IMPORT-OPTIMAL-CUDA-ONLY-2026-10-06, T-VULKAN-IMPORT-WIN32-HANDLES-UNTESTED-2026-10-06, T-VULKAN-DECODER-FRAMES-NEED-PRODUCER-COPY-2026-10-06 (RC4), T-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06 (RC8 tuning) and T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06 (deferred, the platform); a Vulkan-semaphore sighting onT-HIP-ROCM-NO-SYNC-FILE-SEMAPHORE-2026-10-06. Each is in the first-release classification table.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests, except by porting Netflix's own updated assertion verbatim from upstream (value and places as upstream has them, measured against the fork's CPU build first; ADR-1828).Golden gate (
make test-netflix-golden GOLDEN_NINJA_JOBS=4,core/build-goldenbuilt with gcc; the pytest step with the repository venv,VENV=, because the worktree venv has no pytest): 280 passed, 3 skipped.Cross-backend numerical results
Toolchains (binding, the brief's pins): CUDA 13.4 on the host (driver API 13040, NVIDIA 615.71.09), RTX 4090 at 0000:06:00.0, Vulkan 1.4.351 on the device, host loader and headers 1.4.363. SYCL DPC++ 2026.1.1 in a throwaway container of
vmaf-dev-mcp:localwith the Vulkan loader and headers added (1.4.341), Level Zero GPU driver 26.35.39758.10, loader 1.34.0, ANV from the image's Mesa 26.0.8 (Vulkan 1.4.335), Arc A380 at 0000:03:00.0 (xe). HIP ROCm 10.1.0 (HIP 71626385) invmafx-hip-lane:rocm10.1.0with the Vulkan loader and headers added (1.4.341), RADV from the image's Mesa 26.0.8 (Vulkan 1.4.335), gfx1036 at 0000:7d:00.0. The producers in the SYCL and HIP containers use the image's Mesa 26.0.8, not the host's 26.2.4. FFmpeg n9.0.2 (libavutil 61.1.102): the host's for CUDA, a build of the same tag in the ROCm image for SYCL and HIP. Every device run under its lock with the time limit inside it (flock <lock> timeout <N> ...).test_vmafx_import_vulkan_<lane>_bitexactVMAFX_TEST_SKIP_ACQUIRE_WAIT) gives bad frames; with the wait 0 of Ntest_vmafx_import_vulkan_<lane>_fencetest_acquire_orderT-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06)VMAFX_TEST_EARLY_RELEASE); each frame scored by two contexts, the second submitted after the first returnedtest_release_canary, 5 runs per lanevmafx_frame_signal_on_release()on the producer's timeline)test_vmafx_import_vulkan_<lane>test_scores_every_layouttest_refusals_named,test_other_gpu_refused,test_vmafx_import_vulkan_api(11 tests on the CPU)desc.vulkan_pci; Windows handle; multi-planesignal_on_release; another GPU (RADV)-hwaccel vulkanequivalent (AV_HWDEVICE_TYPE_VULKAN+ H.264 decode), device copy into per-plane images, importedtest_vmafx_import_vulkan_<lane>_ffmpegplane_indexANV_DEBUG=video-decode); decoder memory not exportable on ANVtest_vmafx_import_<lane>,_fence,test_vmafx_fence_kindsVendor profiler traces, each of an import-only run (
VMAFX_TEST_IMPORT_ONLY=1, the host sessions skipped):--trace=cuda, 472 cells, 12056 imports: the library stream ran 15070 array-to-device and 4680 device-to-device copies and no host-to-device or device-to-host copy; elsewhere 120 host-to-device copies of at most 64 KiB (the twins' tables) and the twins' result and term readbacks on their own streams (17060, the exact twins' raster-order term planes of ADR-1424 / ADR-1464 included).rocprofv3 --memory-copy-trace --kernel-trace(ROCm 10.1.0), 472 cells: no memory copy on the library stream (its kernels:__amd_rocclr_copyBufferRectAligned, the import's de-interleave and shift kernels,ms_ssim_picture_to_float); elsewhere 40 host-to-device copies onfloat_adm's stream and 4364 readbacks of the twins.sycl-trace --ur.call(DPC++ 2026.1.1), the 4K half (96 cells, 1152 imports): the imported memory is never the source or destination of a copy and is read by 852 kernel pointer arguments; host to device 28 copies, 1648784 bytes in all (largest 262148, tables); the other copies are the twins' readbacks and device-to-device copies of their own buffers.Gates shown failing on planted defects (measured on this branch)
VMAFX_TEST_SKIP_ACQUIRE_WAIT)test_vmafx_import_vulkan_<lane>_fenceVMAFX_TEST_EARLY_RELEASE)8505d76a7) SYCL saw it on 1 or 2 of 8: a SYCL submit returns only after its readsacquireonly,acquire_moreignored (mutation ofvmafx_cuda_wait_acquires())test_vmafx_import_vulkan_cuda_fenceacquire_more); the release arm also fails ("the early release is seen")test_vmafx_import_vulkan_apitest_pci_bus_id_refusedfailsPerformance (if
perforfeat)No timing claim (RC7). A CUDA import imports and maps the exported objects anew per frame (30140 memory imports for 12056 imports in the trace); a cache keyed by the exported object is the RC8 row
T-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06.Deep-dive deliverables (ADR-0108)
docs/research/2162-vmafx-vulkan-frame-import.md: what each runtime does with opaque memory, DRM-modifier dma-bufs, OPTIMAL images, multi-plane images, timeline andSYNC_FDsemaphores; FFmpeg's decoder output and pool export; GStreamer 1.28's allocator; the traces.docs/adr/2152-vmafx-vulkan-frame-import.md## Alternatives considered(lane-specific routes, a Vulkan device in the library, DMABUF only, OPTIMAL on HIP, Level Zero semaphores, accepting the decoder's two-plane image).AGENTS.mdinvariant note —core/AGENTS.d/vmafx-vulkan-frames.md(new page;core/AGENTS.mdregenerated): import only, the shared checks, the CUDA interop, PCI per backend, the tests;core/src/hip/AGENTS.d/vmafx-device-frames.mdpoints tosync_object.cafter the merge.changelog.d/added/api-vulkan-frame-import.md.docs/rebase-notes.d/vmafx-vulkan-frame-import.md(ADR-2197): "VMAFx Vulkan frame import (RC4 WP3 Vulkan lane)" (the lanes' single copies it builds on, the SYNC_FILE destroy that feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342 carries, the API, the new files).User documentation:
docs/api/vmafx/index.mdgains "Vulkan frames" (what the producer does, the descriptor table, what each backend reads, the refusals, anAVVkFrameexample for CUDA, the release rule, FFmpeg decoder and GStreamer notes, how to check for host copies);docs/backends/cuda/overview.md,docs/backends/sycl/zero-copy.mdanddocs/backends/hip/uploads.mdpoint to it; the generated reference pages cover the new kinds, fields and function.Reproducer
Known follow-ups
test_vmafx_import_vulkan_hip_fence: on the landing head one of six runs failed it. Two imports returned ahead of their write (on this platform the import normally waits for the writer,T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06) and still read finished frames (0 bad), and the arm asserts that a skipped wait is seen whenever an import is ahead. The next five runs had 0 imports ahead and passed. The arm's rule was wrong, not the library: an import ahead of its write can still read the finished frame. Fixed in its own pull request on top of this one (fix/vulkan-hip-fence-test-early-import,2dd86ed8c,T-VULKAN-FENCE-TEST-PARTIAL-EARLY-IMPORT-2026-10-08): the skipped wait must be seen on CUDA and where every import ran ahead, none may read early where no import did, and a partial count is printed, not judged.WP3-vulkan-1): thevmafxFFmpeg filter takesAV_PIX_FMT_VULKANframes with the device copy into per-plane images; the GStreamer element allocates exportable images itself (1.28's allocator memory cannot be exported). The request names libplacebo's output frames too (AVVkFrames of FFmpeg's pool on the same device, the decode → tone map → score route); WP9 covers them with the filter, no test here.zeDeviceImportExternalSemaphoreExt()imports a Vulkan timeline on the A380): RC4 follow-up rowT-VULKAN-IMPORT-SYCL-HOST-FENCES-2026-10-06, not in this PR.T-VULKAN-IMPORT-OPTIMAL-CUDA-ONLY-2026-10-06; the import blocking on the producer's write:T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06.T-VULKAN-IMPORT-WIN32-HANDLES-UNTESTED-2026-10-06.docs/rebase-notes.md). Its base is behindorigin/master; master's praetor pin is newer (04cc813ff054) than the branch's (0af07a733e65), and the hooks here ran with the branch's engine.