Skip to content

feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) - #2375

Merged
lusoris merged 1 commit into
masterfrom
rc4/api-wp3-vulkan-import
Oct 10, 2026
Merged

lusoris merged 1 commit into
masterfrom
rc4/api-wp3-vulkan-import

Conversation

@lusoris

@lusoris lusoris commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

RC4 work package 3, Vulkan import lane: frames a Vulkan producer exported (FFmpeg's Vulkan decode and filters, libplacebo, a GStreamer Vulkan element) are imported into CUDA, SYCL and HIP devices with no copy through the host, and score bit for bit as the same frames uploaded from the host on every exact cell of each backend. Import only, per maintainer decision Q-011: the library links no Vulkan loader and makes no Vulkan call. The decision is ADR-2152 (Proposed); the measurements behind it are Research-2162.

Base. The branch starts at rc4/api-wp3-common (027aebc56) and first merges the SYCL lane (#2342, rc4/api-wp3-sycl, 5dafa1cbd) and the HIP lane (#2341, rc4/api-wp3-hip, 20a3f4d4d), because the Vulkan import runs on both. Both lanes are stacked on the CUDA lane (#2277), so this PR's base is rc4/api-wp3-cuda, as theirs is: the diff holds the two lanes, the merge and the Vulkan commits; review the Vulkan work commit by commit after the merge (cff8500bb and later). The merge commit (02aa3a7d2) keeps one copy of each shared helper: core/src/vmafx/sync_object.{c,h} (the HIP lane's sync_file.c and gl_sync.c folded in) and release_events.c; the resolution is described in docs/rebase-notes.md. Correction to the merge commit's message: it says vmafx_fence_destroy() refuses a SYNC_FILE fence. That is wrong for this branch: vmafx_fence_destroy() closes a SYNC_FILE descriptor, the SYCL lane's decision (ADR-2091 item 6), which the HIP lane's refusal gave way to in cff8500bb; test_vmafx_fence_kinds checks that the descriptor is closed, and the API page and the changelog say so.

What it adds (ABI 0.1.5, append-only):

  • VMAFX_MEMORY_VULKAN with VmafxFrameImport.vulkan_handle_type (OPAQUE_FD, DMA_BUF; the Windows types declared and refused), vulkan_tiling (OPTIMAL, LINEAR, DRM_FORMAT_MODIFIER), vulkan_flags (VMAFX_VULKAN_DEDICATED) and vulkan_pci. VMAFX_FENCE_VULKAN_SEMAPHORE: an exported timeline semaphore and a value. acquire_more[2]: the further acquire fences of a frame written under several timelines (one image per plane, as AVVkFrame.sem[i]). VmafxDeviceInfo.pci: each device's PCI location, so a producer picks the same GPU. vmafx_frame_signal_on_release(): the device signals the producer's timeline behind the frame's last reader in every context.
  • CUDA (core/src/cuda/import_vulkan.c): opaque-fd memory of OPTIMAL images as CUDA arrays (NV12 / P010 / P016 planarised on the device; a planar OPTIMAL frame needs VMAFX_IMPORT_ALLOW_COPY, as CUDA arrays do), of LINEAR images and buffers as device pointers; the acquire timelines waited on the library stream; the release signalled there.
  • SYCL and HIP: a LINEAR or DRM-modifier frame whose memory is a dma-buf (opaque descriptors are dma-bufs on the Mesa drivers) goes through the lanes' dma-buf paths (vmafx_import_vulkan_as_dmabuf()), with a sync_file of the producer's write as the acquire fence and a HOST release fence.
  • Refused by name: memory of another GPU (desc.vulkan_pci), a plane of a multi-plane image (desc.plane[1].plane_index), Windows handles, OPTIMAL tiling on SYCL and HIP, DMA_BUF and DRM-modifier tiling on CUDA (the CUDA driver imports dma-bufs on Tegra only), a modifier the device does not read, Vulkan semaphores on SYCL and HIP (ROCm 10.1 imports none; the SYCL lane does not wait on one yet). vmafx_fence_wait() / vmafx_fence_destroy() refuse a VULKAN_SEMAPHORE fence.

python3 scripts/codegen/vmafx-api.py --check: 38 generated files match. --abi-check --against-ref origin/rc4/api-generation-prototype: definition is an append-only successor of origin/rc4/api-generation-prototype (153 additions); against origin/rc4/api-wp3-common: 22 additions.

Landing (Q-083)

On master (2026-10-10)

Rebased onto master 9d3734f06, after #2368 (f6513a6de) and #2714 (c96b8204e, the dev image with the Vulkan loader, ADR-3137) landed. Only this PR's commits are carried; every commit of #2368, #2342, #2367, #2360, #2341 and #2290 is gone, among them 4613fc64f with the leaked identity. Every commit is by and signed off as lusoris@proton.me.

  • docs/state.md: the RC4 label row conflicted; it is master's list with this PR's four T-VULKAN-* rows added where the PR put them.
  • The hosted tidy lanes now build and measure the Vulkan files (operator decision ci-config-19, Q-345). .github/actions/tidy-lane pins the published build of fix(dev): put the Vulkan loader and lavapipe in the dev image so Vulkan tests build and run on CPU in every lane #2714, ghcr.io/vmafx/vmafx-dev-mcp:sha-9d3734f0602a3114849d8edce4136d9c8e7c850a = sha256:29b13030cacce60cec4209f645c6772feda5b2d936fa511fc807eddccf12722c (read from the registry with docker buildx imagetools inspect). Measured in that image (scripts/dev/tidy-lane.sh --image <digest> --write --only ...): cuda 16, hip 15, sycl 15 units of this PR, the five Vulkan test files among them, 0 findings, 0 uncited NOLINT, no compile failure; cpu 8 units (measured on the earlier restack, unchanged). No exception.
  • CPU-only lanes skip with the reason (maintainer decision Q-346): without the lane's GPU the Vulkan import tests run no case and exit 77, and say why, for example "skipped: no CUDA device, so no GPU Vulkan memory to import (lavapipe is a CPU device; the producer needs the importing GPU's memory)"; the harness prints "0 tests run, skipped", so a skip never reads like a pass. The FFmpeg test names its own reasons (inputs not set, not readable, no Vulkan decode).
  • Fixed while landing: master's test_ci_impact TidyLaneRouting (ci(tidy): defer changed files to the lane that measures them and lint C-only SYCL headers through their includers (Q-341, Q-342) #2704) wanted the Vulkan tests in the cuda, hip and sycl patterns of .github/ci-impact.json; core/test/*vulkan* joins them.
  • ABI check against master: 18 additions, abi_version 0.1.9 -> 0.1.10.

Local gate (on b34c7c5c3, master 9d3734f06; the last commit 4b8610dd7 changes only the FFmpeg test's skip message: rebuilt, checked on CUDA with unreadable inputs, and its file measured again in the three GPU lanes, 0 findings)

  • CPU (-Db_lto=false, -j4, one link at a time, warnings as errors): build 0 warnings; --suite=fast 441 OK, 0 failed; test_gpu_picture_pool_uaf (MALLOC_PERTURB_=0) OK; codegen tests 209 passed; make test-netflix-golden 280 passed, 3 skipped; preflight.sh --stage msvcism pass; affected suites: tooling 2748 passed, 0 failed, 6 skipped. vmafx-api.py --check: 80 generated files match.
  • CUDA 13.4, RTX 4090, driver 615.78.08: build 0 warnings; --suite=fast 514 OK, 0 failed; test_vmafx_import_vulkan_cuda 5/5, _fence 2/2 (0 repeated arms), test_vmafx_import_vulkan_cuda_bitexact Netflix and bbb pass; regression test_vmafx_import_cuda 12/12, _fence 6/6, _gl 2/2, _bitexact pass, test_vmafx_fence_kinds 5/5. Without the GPU (CUDA_VISIBLE_DEVICES= on the same build): test_vmafx_import_vulkan_cuda, _bitexact and _fence each print the "no CUDA device" reason above and "0 tests run, skipped", exit 77.
  • SYCL, DPC++ 2026.1.1, Arc A380: build 0 warnings; test_vmafx_import_vulkan_sycl 5/5, _fence 2/2, test_vmafx_import_vulkan_sycl_bitexact Netflix and bbb pass; regression test_vmafx_import_sycl 17/17, _fence 6/6, _gl 2/2, bitexact Netflix and bbb pass, test_sycl_kernel_scratch pass, test_vmafx_fence_kinds 5/5.
  • HIP, pinned vmafx-hip-lane:rocm10.1.0-vulkan (ROCm 10.1.0), gfx1036: build 0 warnings; test_vmafx_import_vulkan_hip 5/5, _fence 2/2, test_vmafx_import_vulkan_hip_bitexact Netflix and bbb pass; regression test_vmafx_import_hip 16/16, _fence 8/8, test_vmafx_fence_kinds 5/5, test_hip_shared_frame 9/9.
  • The three *_ffmpeg tests (FFmpeg-decoded Vulkan frames) skip here: their FFmpeg n9.0.2 + Vulkan inputs were removed by a disk clean-up (CUDA and SYCL now say "the inputs are not readable"; the HIP build does not register the test without FFmpeg's Vulkan). Their device evidence is the earlier run recorded below.
  • Train gates against master (deliverables, docs/state.md touch and rows, silent revert, praetorctl audit): pass. git merge-tree --write-tree origin/master HEAD: clean.

The sections below record how the PR was first put on the stack and the evidence measured there.

Lands bottom-up after #2368 (SYCL 4:2:2 / 4:4:4), the last WP3 lane, one API PR at a time. The branch first merges the SYCL lane (#2342) and the HIP lane (#2341), which landed before it, so its tip is not squashed onto its base (that would bring both lanes again). What is cherry-picked onto the landing head of #2368 is the change rc4/integration carries for this PR, where both lanes were already in: 798c42c5d (the lane's cff8500bb), 541a91a01 (3decc47da), bae72c148 (43b2231f3), a21d20ce1 (c101d5030), and the lane's last commit 8505d76a7 (the release canary scores each frame in two contexts). Compared file by file with the lane's own change (02aa3a7d2..8505d76a7, after its merge commit): the same 79 files; the differences are the ABI number, the merge's single copies the lanes landed with, and what the lanes' landing already changed (the HIP GL export, the 4:2:2 / 4:4:4 layouts, and the SYNC_FILE destroy of ADR-2091 item 6, which #2342 now carries).

ADR-2152 was Proposed on the branch; the maintainer accepted it as written (ledger Q-089, 2026-10-07): this PR flips it (status line, deciders, a ## References entry citing Q-089, the index fragment). Its frozen body names ABI 0.1.5, the number free when the lane branched.

ABI check against the parent: python3 scripts/codegen/vmafx-api.py --abi-check --against-ref land-2368 -> definition is an append-only successor of land-2368 (18 additions). abi_version goes from 0.1.9 to 0.1.10 in node VMAFX_0.1; the new entries (VMAFX_MEMORY_VULKAN, VMAFX_FENCE_VULKAN_SEMAPHORE, VmafxVulkanHandleType, VmafxVulkanTiling, VmafxVulkanFlags, VmafxDeviceInfo.pci, the VmafxFrameImport fields acquire_more, vulkan_handle_type, vulkan_tiling, vulkan_flags, vulkan_pci, vmafx_frame_signal_on_release()) say "Added in ABI 0.1.10"; no entry of the parents is renumbered.

Conflicts, resolved per hunk:

  • core/src/AGENTS.d/vmafx-device-frames.md: the parents' page, unchanged here; its fence rule (destroy closes a SYNC_FILE fence's descriptor, ADR-2091 item 6; a GL_SYNC fence stays refused) lands with feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342.
  • core/src/hip/AGENTS.d/vmafx-device-frames.md, docs/backends/sycl/zero-copy.md: the parents' wording plus this PR's addition (one sync_object.c copy shared with SYCL; "Vulkan memory" in the SYCL ingestion row).
  • docs/state.md: the resolver; the RC4 label row keeps the parents' description and gains this PR's four Vulkan rows; the RC8 and deferred label rows gain T-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06 and T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06.
  • Generated files (API headers, symbol list, Python binding, ABI layout test, reference pages, AGENTS indexes) regenerated; the files ADR-2197 renders at landing and the derived citation map (ADR-2200) are the parent's.

Changed while landing:

  • core/src/cuda/import_vulkan.c included <unistd.h> and called dup() / close() unconditionally; the required Windows MSVC+CUDA build has neither (and scripts/dev/preflight.sh --stage msvcism does not look for it). It now calls VMAF_DUP() (added to core/src/compat/crt_portable.h: _dup() on Windows, dup() elsewhere) and VMAF_CLOSE(), the tree's spelling for the C runtime calls Windows renames. The CUDA build and every CUDA test below were run again after the change (same numbers); the SYCL and HIP builds compile neither file's changed code.
  • The SYNC_FILE destroy rule of ADR-2091 item 6 (destroy closes the descriptor, every lane) moved into feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342 by maintainer decision on 2026-10-08, so master never contradicts it; this PR no longer changes vmafx_fence_destroy() for it, test_vmafx_fence_kinds, test_vmafx_import_sycl's sync_file case or the API page's fence paragraph. Its changelog fragment drops the sentence, and its rebase note says the rule landed with feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342. After the restack onto the new parents the code of this PR is the gated head's (f0c73a33d): the tree diff between the two heads is four fragment files.
  • The rebase note is the fragment docs/rebase-notes.d/vmafx-vulkan-frame-import.md; core/AGENTS.d/vmafx-vulkan-frames.md is in the caveman register (praetorctl caveman check --kind=context: pass; caveman_keep_check.py: 0 missing).

Local gate on the landing head (CPU, -Db_lto=false, -j4, warnings as errors; host Vulkan loader and headers 1.4.363, so the Vulkan tests are built): build 0 warnings; --suite=fast 420 OK, 0 failed (test_vmafx_import_vulkan_api included; test_gpu_picture_pool_uaf with MALLOC_PERTURB_=0, #2547: 1 OK); codegen tests 124 passed; make test-netflix-golden GOLDEN_NINJA_JOBS=4 280 passed, 3 skipped; preflight.sh --stage msvcism pass; affected suites: mcp 590 passed, tooling 2382 passed (6 skipped), 0 failed. Platform check: the library makes no Vulkan call; vmafx/frame_import_vulkan.c keeps its dma-buf translation behind __linux__ and refuses the Windows handle types by name; the Vulkan tests are built only where dependency('vulkan') is found.

Train gates against the parent head (BASE_SHA = land-2368 faffc033d, the stack's scripts with ADR-2197): deliverables-check.sh, state-md-touch-check.sh, check-state-md-rows.sh, check-silent-revert.py (clean) and praetorctl audit (engine 7458a220) all rc 0.

Device runs, each under its lock with timeout 290 inside (the bit-exactness programs in their VMAFX_TEST_ONLY=netflix and bbb halves to fit it); the FFmpeg tests decode the Vulkan lane's two 576x324 H.264 inputs.

  • CUDA 13.4 on the host, RTX 4090 at 0000:06:00.0 (NVIDIA 615.71.09, Vulkan 1.4.351), on the head with the VMAF_DUP change: build 0 warnings; --suite=fast 493 OK, 0 failed. test_vmafx_import_vulkan_cuda_bitexact 376 + 96 = 472 cells, 21464 values, 0 differing, 0 repeated attempts (12056 imports, 6028 conversions); _fence acquire: skipped wait 8 bad (16 imports ahead of their write), with the wait 0 bad of 8; release: planted early 4 bad, real 0 of 8; _ffmpeg 48 frames, 23 cells, 3936 values, 0 differing (the decoder's own two-plane image refused naming plane_index); test_vmafx_import_vulkan_cuda 5 of 5. Lane regression: test_vmafx_import_cuda 12 of 12, _fence 6 of 6, _gl 2 of 2, _bitexact 236 cells, 10732 values, 0 differing; test_vmafx_fence_kinds 5 of 5.
  • SYCL: DPC++ 2026.1.1 in vmaf-dev-mcp:local-vulkan (the pinned dev image plus the Vulkan loader and headers 1.4.341; ANV from Mesa 26.0.8, Vulkan 1.4.335; ANV_DEBUG=video-decode), Arc A380 at 0000:03:00.0 (xe); build under sycl-build.lock, 0 compiler warnings (two ocloc RetryManager notes for the untouched Ss2SsimUnitsKernel and TermKernel). test_vmafx_import_vulkan_sycl_bitexact 376 + 96 = 472 cells, 21464 values, 0 differing, 0 repeated attempts (12056 imports, 9042 conversions); _fence acquire: skipped wait 8 bad (16 ahead), with the wait 0 bad of 8; release: planted early 8 bad, real 0 of 8; _ffmpeg 48 frames, 23 cells, 3936 values, 0 differing; test_vmafx_import_vulkan_sycl 5 of 5. Lane regression: test_vmafx_import_sycl 17 of 17, _fence 6 of 6, _gl 2 of 2, _bitexact 188 + 48 cells, 9712 + 1020 values, 0 differing; test_vmafx_fence_kinds 5 of 5; test_sycl_kernel_scratch 132 kernels, 0 use scratch memory; VMAF_SYCL_AOT_JOBS=8 meson test --suite sycl-aot: OK, 36 SYCL translation units for 19 targets, 220.8 s.
  • HIP: ROCm 10.1.0 (HIP 7.16.26385) in vmafx-hip-lane:rocm10.1.0-vulkan (the pinned image plus the Vulkan loader and headers 1.4.341; RADV from Mesa 26.0.8), gfx1036 at 0000:7d:00.0: build 0 warnings. test_vmafx_import_vulkan_hip_bitexact 376 + 96 = 472 cells, 21464 values, 0 differing in every attempt, 2 attempts repeated once (the gfx1036's dropped dispatches, T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01); _ffmpeg 48 frames, 23 cells, 3936 values, 0 differing; test_vmafx_import_vulkan_hip 5 of 5; _fence 5 of 6 runs pass (acquire: 0 bad with and without the wait, 0 imports ahead, as T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06 describes; release: planted early 8 bad, real 0 of 8). The first run failed its acquire arm: 2 imports returned ahead of their write and still read finished frames (0 bad), and the test asserts that a skipped wait is seen whenever an import is ahead; see Known follow-ups. Lane regression: test_vmafx_import_hip 16 of 16, _fence 8 of 8, test_vmafx_fence_kinds 5 of 5, test_hip_shared_frame 9 of 9. FFmpeg for the Vulkan decode tests: the Vulkan lane's own build (FFmpeg n9.0.2 libraries, 63.1.102) through -Dpkg_config_path in the SYCL and HIP builds, as the lane built them.

tidy (dev image with the Vulkan loader and headers, clang-tidy 22.1.8, scripts/dev/tidy-lane.sh --image vmaf-dev-mcp:local-vulkan --jobs 4 --write --only ...): cpu core/src/vmafx/{device,fence,frame_import,frame_import_vulkan,sync_object}.c, core/test/{test_vmafx_abi_layout,test_vmafx_fence_kinds,test_vmafx_import_vulkan_api}.c (8 TUs) 0 findings; cuda the five shared vmafx/*.c, core/src/cuda/import_{device,fence,frame,vulkan}.c, core/test/test_vmafx_import_cuda{,_bitexact}.c and the five Vulkan test sources (16 TUs) 0 findings, and import_vulkan.c again after the VMAF_DUP change: 0 (cpu libvmaf.c and dict.cpp through the changed crt_portable.h: 0); hip the five shared, core/src/hip/import_{device,fence,frame}.c, core/test/test_vmafx_import_hip{,_bitexact}.c and the five Vulkan test sources (15 TUs) 0 findings; sycl the five shared, core/src/sycl/import_{device,fence,frame}.c, core/src/sycl/vmafx_sycl_rt.cpp, core/test/test_vmafx_import_sycl.c and the five Vulkan test sources (15 TUs) 0 findings. check-tidy-coverage.py: every tracked translation unit is read or excepted.

Type

  • feat — new feature
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, HISS audit and evidence replay with the branch's pinned engine 0af07a733e65).
  • Unit tests pass: CUDA build, --suite fast --no-suite gpu: 377 OK, 0 fail, 1 skipped (test_vmafx_api_abi_append_only, 77: its reference commit predates core/api/vmafx.toml); every Vulkan and lane device program per backend in the table below.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. — no extractor changed; every exact cell of every backend is compared value for value against host uploads below: 0 differing.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. — no extractor touched.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (EUPL-1.2, fork-authored).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. — not a breaking change: append-only.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md — ADR-2152 (claimed with scripts/adr/next-free.sh --claim), fragment added; the index, by-tag and title pages are rendered at landing (ADR-2197).

tidy (dev container vmaf-dev-mcp:local plus the Vulkan headers and loader, clang-tidy 22.1.8, scripts/dev/tidy-lane.sh --only ... <lane>, --jobs 4): cpu lane, 8 units (core/src/vmafx/{device,fence,frame_import,frame_import_vulkan,sync_object}.c, core/test/{test_vmafx_abi_layout,test_vmafx_fence_kinds,test_vmafx_import_vulkan_api}.c): 0. cuda lane, 11 units (core/src/cuda/import_{device,fence,frame,vulkan}.c, core/test/test_vmafx_import_cuda{,_bitexact}.c and the five Vulkan test sources): 0. hip lane, 11 units (core/src/hip/import_{device,fence,frame,gl}.c, core/test/test_vmafx_import_hip{,_bitexact}.c, the Vulkan test sources): 0. sycl lane, 10 units (core/src/sycl/import_{device,fence,frame}.c, vmafx_sycl_rt.cpp, core/test/test_vmafx_import_sycl.c, the Vulkan test sources): 0 in this PR's files, 0 uncited NOLINT; the 7 findings left are in core/include/libvmaf/picture.h, untouched here and suppressed on master by #2186 (this stack's base predates it), as #2342 reports. The first runs found 12 in this PR's files (analyzer array bounds and a NULL path in the CUDA release-signal registration, sscanf in three PCI parsers, an increment in a condition, two misplaced const, an enum zero-initialised, an implicit widening, a test function's size); all fixed, the three parsers became one (vmafx_parse_pci_bus_id()), and the changed units re-measured at 0.

scripts/dev/preflight.sh --stage msvcism: pass. Pre-push gate (lefthook, the branch's praetor engine): pass, MkDocs strict build included; its first run caught two findings, fixed in c101d5030 (two fork-added functions of 20+ lines without an assert; a computed test source name the stale-reference check could not resolve).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated — opened T-VULKAN-IMPORT-SYCL-HOST-FENCES-2026-10-06, T-VULKAN-IMPORT-OPTIMAL-CUDA-ONLY-2026-10-06, T-VULKAN-IMPORT-WIN32-HANDLES-UNTESTED-2026-10-06, T-VULKAN-DECODER-FRAMES-NEED-PRODUCER-COPY-2026-10-06 (RC4), T-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06 (RC8 tuning) and T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06 (deferred, the platform); a Vulkan-semaphore sighting on T-HIP-ROCM-NO-SYNC-FILE-SEMAPHORE-2026-10-06. Each is in the first-release classification table.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests, except by porting Netflix's own updated assertion verbatim from upstream (value and places as upstream has them, measured against the fork's CPU build first; ADR-1828).
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. — not applicable.

Golden gate (make test-netflix-golden GOLDEN_NINJA_JOBS=4, core/build-golden built with gcc; the pytest step with the repository venv, VENV=, because the worktree venv has no pytest): 280 passed, 3 skipped.

Cross-backend numerical results

Toolchains (binding, the brief's pins): CUDA 13.4 on the host (driver API 13040, NVIDIA 615.71.09), RTX 4090 at 0000:06:00.0, Vulkan 1.4.351 on the device, host loader and headers 1.4.363. SYCL DPC++ 2026.1.1 in a throwaway container of vmaf-dev-mcp:local with the Vulkan loader and headers added (1.4.341), Level Zero GPU driver 26.35.39758.10, loader 1.34.0, ANV from the image's Mesa 26.0.8 (Vulkan 1.4.335), Arc A380 at 0000:03:00.0 (xe). HIP ROCm 10.1.0 (HIP 71626385) in vmafx-hip-lane:rocm10.1.0 with the Vulkan loader and headers added (1.4.341), RADV from the image's Mesa 26.0.8 (Vulkan 1.4.335), gfx1036 at 0000:7d:00.0. The producers in the SYCL and HIP containers use the image's Mesa 26.0.8, not the host's 26.2.4. FFmpeg n9.0.2 (libavutil 61.1.102): the host's for CUDA, a build of the same tag in the ROCm image for SYCL and HIP. Every device run under its lock with the time limit inside it (flock <lock> timeout <N> ...).

Exit evidence (ADR-1829) Test CUDA (RTX 4090) SYCL (Arc A380) HIP (gfx1036)
Imported == host-uploaded, every exact cell, 576x324 pair, both 1080p checkerboards, the 10-bit Sparks pair, 4K bbb; each clip in each of the lane's layouts as planar and as NV12 / P010 test_vmafx_import_vulkan_<lane>_bitexact 472 cells, 21464 values, 0 differing (12056 imports, 6028 conversions); layouts OPTIMAL image + buffer 472 cells, 21464 values, 0 differing (12056 imports, 9042 conversions); layouts buffer + Tile4 DRM image 472 cells, 21464 values, 0 differing (12056 imports, 6028 conversions; 0 repeated attempts in this run, up to 4 allowed for the gfx1036's dropped dispatches); layouts buffer + linear DRM image
Acquire wait skipped (planted VMAFX_TEST_SKIP_ACQUIRE_WAIT) gives bad frames; with the wait 0 of N test_vmafx_import_vulkan_<lane>_fence test_acquire_order skipped: 8 bad (16 imports ahead of their write); with the wait: 0 bad of 8 skipped: 8 bad (16 ahead); with the wait: 0 bad of 8 0 bad of 8 with and without the wait: the HIP runtime's dma-buf import blocks until the producer's write is done (0 imports ahead; the test asserts that none read an unfinished frame; T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06)
Release canary with a planted early release (VMAFX_TEST_EARLY_RELEASE); each frame scored by two contexts, the second submitted after the first returned same, test_release_canary, 5 runs per lane early: 4 or 5 bad of 8 (bound: at least 1; CUDA copies the frame once, at the import); real: 0 of 8 in 5 of 5 (release = vmafx_frame_signal_on_release() on the producer's timeline) early: 8 of 8 in 5 of 5 (asserted: all 8); real: 0 of 8 in 5 of 5 (HOST release fence) early: 8 of 8 in 5 of 5 (asserted: all 8); real: 0 of 8 in 5 of 5 (HOST release fence)
Host-copy counter 0 every Vulkan test 0 0 0
One import scored by two contexts test_vmafx_import_vulkan_<lane> test_scores_every_layout each layout and pixel format: both contexts equal the host session same same
Refusals named: format / modifier / cross-device / unsupported semaphore test_refusals_named, test_other_gpu_refused, test_vmafx_import_vulkan_api (11 tests on the CPU) DRM tiling, dma-buf handle, planar OPTIMAL without ALLOW_COPY; another GPU (A380) by desc.vulkan_pci; Windows handle; multi-plane OPTIMAL tiling, AMD modifier, Vulkan semaphore acquire, further acquires, signal_on_release; another GPU (RADV) the same as SYCL; another GPU (A380)
Real decode: FFmpeg n9.0.2 -hwaccel vulkan equivalent (AV_HWDEVICE_TYPE_VULKAN + H.264 decode), device copy into per-plane images, imported test_vmafx_import_vulkan_<lane>_ffmpeg 48 frames, 23 cells, 3936 values, 0 differing; the decoder's own frame (one 2-plane image) refused naming plane_index 48 frames, 3936 values, 0 differing (ANV_DEBUG=video-decode); decoder memory not exportable on ANV 48 frames, 3936 values, 0 differing; decoder memory not exportable on RADV
Lane tests unchanged by the merge and the Vulkan code test_vmafx_import_<lane>, _fence, test_vmafx_fence_kinds 12 / 6 / 5 passed 16 / 6 / 5 passed 16 / 8 / 5 passed

Vendor profiler traces, each of an import-only run (VMAFX_TEST_IMPORT_ONLY=1, the host sessions skipped):

  • CUDA, Nsight Systems 2026.3.2 --trace=cuda, 472 cells, 12056 imports: the library stream ran 15070 array-to-device and 4680 device-to-device copies and no host-to-device or device-to-host copy; elsewhere 120 host-to-device copies of at most 64 KiB (the twins' tables) and the twins' result and term readbacks on their own streams (17060, the exact twins' raster-order term planes of ADR-1424 / ADR-1464 included).
  • HIP, rocprofv3 --memory-copy-trace --kernel-trace (ROCm 10.1.0), 472 cells: no memory copy on the library stream (its kernels: __amd_rocclr_copyBufferRectAligned, the import's de-interleave and shift kernels, ms_ssim_picture_to_float); elsewhere 40 host-to-device copies on float_adm's stream and 4364 readbacks of the twins.
  • SYCL, sycl-trace --ur.call (DPC++ 2026.1.1), the 4K half (96 cells, 1152 imports): the imported memory is never the source or destination of a copy and is read by 852 kernel pointer arguments; host to device 28 copies, 1648784 bytes in all (largest 262148, tables); the other copies are the twins' readbacks and device-to-device copies of their own buffers.

Gates shown failing on planted defects (measured on this branch)

Planted defect Test Result
acquire wait skipped (VMAFX_TEST_SKIP_ACQUIRE_WAIT) test_vmafx_import_vulkan_<lane>_fence CUDA 8 bad, SYCL 8 bad, asserted every run; HIP: the platform orders the import (above)
release when the submit returns (VMAFX_TEST_EARLY_RELEASE) same SYCL 8 of 8 and HIP 8 of 8 in 5 of 5 runs each, all 8 asserted; CUDA 4 or 5 of 8, at least 1 asserted. With one context (before 8505d76a7) SYCL saw it on 1 or 2 of 8: a SYCL submit returns only after its reads
CUDA waits on acquire only, acquire_more ignored (mutation of vmafx_cuda_wait_acquires()) test_vmafx_import_vulkan_cuda_fence the real acquire arm: 8 bad of 8, fails (the reference's timeline is handed over in acquire_more); the release arm also fails ("the early release is seen")
the PCI parser's guard against signs and blanks removed test_vmafx_import_vulkan_api test_pci_bus_id_refused fails

Performance (if perf or feat)

No timing claim (RC7). A CUDA import imports and maps the exported objects anew per frame (30140 memory imports for 12056 imports in the trace); a cache keyed by the exported object is the RC8 row T-CUDA-VULKAN-IMPORT-PER-FRAME-INTEROP-2026-10-06.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/2162-vmafx-vulkan-frame-import.md: what each runtime does with opaque memory, DRM-modifier dma-bufs, OPTIMAL images, multi-plane images, timeline and SYNC_FD semaphores; FFmpeg's decoder output and pool export; GStreamer 1.28's allocator; the traces.
  • Decision matrix — docs/adr/2152-vmafx-vulkan-frame-import.md ## Alternatives considered (lane-specific routes, a Vulkan device in the library, DMABUF only, OPTIMAL on HIP, Level Zero semaphores, accepting the decoder's two-plane image).
  • AGENTS.md invariant note — core/AGENTS.d/vmafx-vulkan-frames.md (new page; core/AGENTS.md regenerated): import only, the shared checks, the CUDA interop, PCI per backend, the tests; core/src/hip/AGENTS.d/vmafx-device-frames.md points to sync_object.c after the merge.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/added/api-vulkan-frame-import.md.
  • Rebase note — docs/rebase-notes.d/vmafx-vulkan-frame-import.md (ADR-2197): "VMAFx Vulkan frame import (RC4 WP3 Vulkan lane)" (the lanes' single copies it builds on, the SYNC_FILE destroy that feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) #2342 carries, the API, the new files).

User documentation: docs/api/vmafx/index.md gains "Vulkan frames" (what the producer does, the descriptor table, what each backend reads, the refusals, an AVVkFrame example for CUDA, the release rule, FFmpeg decoder and GStreamer notes, how to check for host copies); docs/backends/cuda/overview.md, docs/backends/sycl/zero-copy.md and docs/backends/hip/uploads.md point to it; the generated reference pages cover the new kinds, fields and function.

Reproducer

python3 scripts/codegen/vmafx-api.py --check
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/rc4/api-generation-prototype
# CUDA on the host (Vulkan loader + headers installed)
meson setup build-cuda core -Denable_cuda=true && nice -n 10 ninja -C build-cuda -j4
# fixtures: python/test/resource/yuv and testdata/bbb; FFmpeg inputs: two H.264 files of one size
export VMAFX_TEST_VULKAN_REF=ref.mp4 VMAFX_TEST_VULKAN_DIST=dist.mp4
for t in test_vmafx_import_vulkan_api test_vmafx_import_vulkan_cuda test_vmafx_import_vulkan_cuda_fence \
         test_vmafx_import_vulkan_cuda_ffmpeg test_vmafx_import_vulkan_cuda_bitexact; do
  flock ~/.cache/vmafx-locks/cuda-4090.lock timeout 1700 build-cuda/test/$t
done
# SYCL: the same tests (…_sycl…) in the dev image with /dev/dri, ONEAPI_DEVICE_SELECTOR=level_zero:gpu
#       (ANV_DEBUG=video-decode for the FFmpeg test); HIP: (…_hip…) in vmafx-hip-lane:rocm10.1.0
#       with /dev/kfd and /dev/dri; each under its device lock.
# vendor traces: VMAFX_TEST_IMPORT_ONLY=1 with nsys profile --trace=cuda / sycl-trace --ur.call /
#       rocprofv3 --memory-copy-trace --kernel-trace on the _bitexact program

Known follow-ups

  • HIP acquire arm of test_vmafx_import_vulkan_hip_fence: on the landing head one of six runs failed it. Two imports returned ahead of their write (on this platform the import normally waits for the writer, T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06) and still read finished frames (0 bad), and the arm asserts that a skipped wait is seen whenever an import is ahead. The next five runs had 0 imports ahead and passed. The arm's rule was wrong, not the library: an import ahead of its write can still read the finished frame. Fixed in its own pull request on top of this one (fix/vulkan-hip-fence-test-early-import, 2dd86ed8c, T-VULKAN-FENCE-TEST-PARTIAL-EARLY-IMPORT-2026-10-08): the skipped wait must be seen on CUDA and where every import ran ahead, none may read early where no import did, and a partial count is printed, not judged.
  • Coordination with WP9 (request WP3-vulkan-1): the vmafx FFmpeg filter takes AV_PIX_FMT_VULKAN frames with the device copy into per-plane images; the GStreamer element allocates exportable images itself (1.28's allocator memory cannot be exported). The request names libplacebo's output frames too (AVVkFrames of FFmpeg's pool on the same device, the decode → tone map → score route); WP9 covers them with the filter, no test here.
  • SYCL device-side Vulkan semaphores through Level Zero (zeDeviceImportExternalSemaphoreExt() imports a Vulkan timeline on the A380): RC4 follow-up row T-VULKAN-IMPORT-SYCL-HOST-FENCES-2026-10-06, not in this PR.
  • HIP OPTIMAL images (one of five probe runs read a P010 chroma plane wrong; one size measured): T-VULKAN-IMPORT-OPTIMAL-CUDA-ONLY-2026-10-06; the import blocking on the producer's write: T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER-2026-10-06.
  • Windows handles: T-VULKAN-IMPORT-WIN32-HANDLES-UNTESTED-2026-10-06.
  • CUDA release canary margin: the planted early release is seen on 4 or 5 of 8 frames (bound: at least 1). CUDA reads the producer's memory once, in the import copy, and only frames whose copy is still held behind the delay timeline when the submit returns show it; SYCL and HIP readers copy per context and show it on 8 of 8.
  • Restack: this branch carries the SYCL and HIP lanes; once they land, the restack drops the merge and keeps its single-copy resolution (docs/rebase-notes.md). Its base is behind origin/master; master's praetor pin is newer (04cc813ff054) than the branch's (0af07a733e65), and the hooks here ran with the branch's engine.

@lusoris lusoris added the rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag label Oct 6, 2026
@github-actions github-actions Bot added the type:feature New feature or request label Oct 6, 2026
lusoris added a commit that referenced this pull request Oct 6, 2026
…152)

Brings the Vulkan import lane (#2375) onto the integration branch. The
lane's first commit (02aa3a7, its own merge of the SYCL and HIP lanes)
is not taken: both lanes are here already.

Resolved per hunk:

- core/api/vmafx.toml: the lane's twelve additions are ABI 0.1.10 here
  (0.1.5 on the lane, ADR-1897); generated files regenerated.
- cuda/import_frame.c: the Vulkan check runs first; a packed or MSB layout
  is refused for an OPTIMAL Vulkan image as for CUDA arrays.
- hip/import_frame.c: the GL dma-buf export (ADR-2132) and the Vulkan
  dma-buf path are both kept; the lane's VmafxHipGl registration struct is
  dropped, since the GL interop it served is gone here.
- vmafx/internal.h: pooling and window declarations and the Vulkan
  declarations are both kept.
- SYNC_FILE fences: vmafx_fence_destroy() closes the descriptor (ADR-2091
  item 6), as the lane resolves it; the API index, the device-frames
  agents page, the rebase note and the SYCL sync_file test follow.
- test_vmafx_import_vulkan_api: designated initialisers for the layout,
  which gained the packed fields.

(cherry picked from commit cff8500)
@lusoris lusoris mentioned this pull request Oct 6, 2026
9 of 19 tasks
lusoris added a commit that referenced this pull request Oct 6, 2026
The integration branch took the Vulkan import lane (#2375) at ABI 0.1.10,
the number WP9's spec functions had. Per ADR-1897 the spec functions move
to ABI 0.1.11 (definition, changelog fragment, ADR-2125, rebase note).
Both branches added a Research-2162; the WP9 digest is Research-2163.
The merge put integration's new rebase bullet under the WP9 section; it
is back in the integration section. Generated files: one side,
regenerated.
@lusoris
lusoris force-pushed the rc4/api-wp3-cuda branch 2 times, most recently from 468f5cc to c595358 Compare October 8, 2026 12:08
Base automatically changed from rc4/api-wp3-cuda to master October 8, 2026 12:11
lusoris added a commit that referenced this pull request Oct 9, 2026
…-2091 item 6 decides

The WP3 lanes met with two rules for a SYNC_FILE fence: the SYCL lane
(ADR-2091 item 6) closes the descriptor in vmafx_fence_destroy(), the HIP
lane refused the call. rc4/integration kept the refusal when it merged
the lanes and restored item 6 only with the Vulkan import (#2375). This
PR carries item 6 itself, so no landed state contradicts its own ADR:
vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor in every
build (refused on Windows, which has no sync_files) for every lane, and
still refuses a GL_SYNC fence (the producer's, glDeleteSync()).

Test: test_vmafx_fence_kinds test_sync_file_states checks that destroy
closes the descriptor; it fails on the refusal. test_vmafx_import_sycl's
sync_file case destroys its acquire fence instead of closing it itself.

Signed-off-by: Lusoris <lusoris@proton.me>
lusoris added a commit that referenced this pull request Oct 9, 2026
…GL textures (RC4 WP3, ADR-2091) (#2342)

* feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091)

The VMAFx API now scores frames that already live on a SYCL device
without a copy through the host, and they score bit for bit as the
same frames uploaded from the host for every SYCL twin declared exact.

- SYCL devices are the Level Zero GPUs, opened by index or in the
  context of the caller's queue; each has one in-order library queue
  with immediate command lists, the precaution of draft PR #2217
  (evaluated: its i915 defect did not reproduce on xe).
- USM planes of any address and pitch and linear dma-bufs are bound;
  Intel Y-tiled and Tile4 dma-bufs are de-tiled on the device with the
  address math the VA-surface import now shares (core/src/sycl/detile.h);
  NV12 / P010 / P016 are planarised on the device. GL textures are
  imported through an EGL dma-buf export.
- Every SYCL twin that stages its own planes, the chroma twins
  included, copies a device picture on the device through
  vmaf_sycl_picture_read_plane(), which records the read on the frame.
- Acquire and release are joins (an empty kernel with depends_on()),
  not barriers: a barrier on the event of a command behind a host task
  waits on the host (Research-2159). SYCL_EVENT acquire fences are
  waited on by the device; HOST, SYNC_FILE, GL_SYNC and a dma-buf's
  implicit write fences are checked on the host for the import rule's
  retry. HOST and SYCL_EVENT release fences and the release callback
  follow the last reader in every context.
- The sync_file, dma-buf fence and GL sync helpers move to
  core/src/vmafx/sync_object.c, shared with the CUDA lane.
- Windows shared textures and SYNC_FILE release fences are refused,
  tracked as T-SYCL-VMAFX-IMPORT-REMAINDER-2026-10-06.

Landing (Q-083): the branch's diff (15bd091..5dafa1c) squashed onto
(a082db7): the HIP lane's vmafx/gl_sync.c and sync_file.c folded into
sync_object.c; SYNC_FILE fences borrowed (vmafx_fence_destroy() refuses
them); dmabuf_import.cpp keeps master's driver_fd() (#2377) inside
vmaf_sycl_dmabuf_import_queue(). Carried from rc4/integration: 44cbb8c
(the SYCL EGL export folded onto vmafx/egl_export.c, VmafxEglTiled; the
engine_leave signature) without its d3d11 stub, which the WP6 generator
fix replaces, and the test of 869ec15 (an import closes only its own
descriptors), whose library half master's #2377 already holds, and the
SYCL parts of docs/api/vmafx/index.md that the integration merge dropped
(restored there by 3b9a414), written for CUDA, SYCL and HIP. The
rebase note is a fragment (ADR-2197). vmafx_sycl_producer.cpp is brought
to 0 clang-tidy findings in the sycl lane, and core/src/compat/crt_portable.h's
C spelling vmaf_tmpfile_portable(void), which the sycl lane reports through
sycl/common.cpp, carries the ADR-1138 NOLINT that x86/avx512_warm_up.h has.

* fix(api): destroy a SYNC_FILE fence by closing its descriptor, as ADR-2091 item 6 decides

The WP3 lanes met with two rules for a SYNC_FILE fence: the SYCL lane
(ADR-2091 item 6) closes the descriptor in vmafx_fence_destroy(), the HIP
lane refused the call. rc4/integration kept the refusal when it merged
the lanes and restored item 6 only with the Vulkan import (#2375). This
PR carries item 6 itself, so no landed state contradicts its own ADR:
vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor in every
build (refused on Windows, which has no sync_files) for every lane, and
still refuses a GL_SYNC fence (the producer's, glDeleteSync()).

Test: test_vmafx_fence_kinds test_sync_file_states checks that destroy
closes the descriptor; it fails on the refusal. test_vmafx_import_sycl's
sync_file case destroys its acquire fence instead of closing it itself.

* fix(test): key the one-EGL-export contract by POSIX paths so it holds on Windows

test_vmafx_one_egl_export_contract.py keyed the library sources by
str(path.relative_to(SRC)), which is "vmafx\sync_object.c" on Windows, so
test_sync_objects_defined_once compared it with "vmafx/sync_object.c" and
failed there. The keys are now relative_to(SRC).as_posix(), as the other
contract tests do (#1850). scripts/dev/preflight.sh --stage msvcism, which
refuses that pattern, failed on the branch and passes now.

* ci(tidy): measure the SYCL device-frame import's units in the cpu, sycl, cuda and hip lanes

On master after #2367: every lane's baseline is master's, and the units
this pull request adds or changes were measured again in the dev
container (scripts/dev/tidy-lane.sh --write --only ..., clang-tidy
22.1.8), 0 findings and 0 uncited NOLINT in each lane:

- cpu: libvmaf.c, the shared core/src/vmafx/ units (device, device_context,
  egl_export, fence, frame_import, frame_import_admit, frame_pool, submit,
  sync_object) and test_vmafx_fence_kinds.c;
- sycl: the shared units, the SYCL import sources and runtime, the SYCL
  feature twins this pull request touches and the import tests;
- cuda: the shared units and core/src/cuda/import_gl.c;
- hip: the shared units, core/src/hip/import_frame.c and import_gl.c.

core/src/vmafx/gl_sync.c and sync_file.c, folded into sync_object.c, leave
the cuda and hip lanes: a scoped write drops the entries of deleted files
since #2636, so the whole-lane measurements the earlier record needed are
gone. This replaces the pull request's two earlier tidy commits.

* fix(build): compile the SYCL import tests' producer with the SYCL argument list

Master's test_sycl_math_constants_contract (#2654) requires every icpx
compile command to take sycl_common_args or sycl_feature_tail_args, which
carry the Windows math-constant define. The vmafx_sycl_producer.cpp
command of this pull request spelled out the parts of sycl_common_args
without that define and failed the contract on the rebased tree. It now
uses sycl_common_args itself; the contract test passes again.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris marked this pull request as ready for review October 10, 2026 01:18
@lusoris
lusoris force-pushed the rc4/api-wp3-vulkan-import branch from 8505d76 to 9c460c4 Compare October 10, 2026 01:25
lusoris added a commit that referenced this pull request Oct 10, 2026
…an tests build and run on CPU in every lane (#2714)

* fix(dev): put the Vulkan loader and lavapipe in the dev image so Vulkan tests build and run on CPU in every lane

The dev image had no Vulkan loader, so meson never found
dependency('vulkan') in it. The Vulkan frame-import tests of #2375
(ADR-2152) were not built there, and the hosted cuda, hip and sycl
clang-tidy lanes, which run in the image's pinned digest, failed closed
with "not measured" on the five test files #2375 records.

Per operator decision ci-config-19 (ADR-3137): the build-deps stage
installs libvulkan1, libvulkan-dev, mesa-vulkan-drivers (lavapipe, the CPU
Vulkan driver) and vulkan-tools from the archive of the digest-pinned
DEV_BASE, like the stage's other distribution packages, and the build
fails when vulkaninfo --summary lists no lavapipe. One image, no tidy
exception; real-GPU Vulkan runs stay a separate signal.

The hosted lanes pick it up when the tidy-lane action's pinned digest
moves to the published build of this change; #2375 moves it and
re-measures its baselines.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the rc4/api-wp3-vulkan-import branch from 9c460c4 to 4b8610d Compare October 10, 2026 06:30
…152) (#2375)

* feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152)

VMAFX_MEMORY_VULKAN and VMAFX_FENCE_VULKAN_SEMAPHORE with the handle
type, tiling, flags and PCI location of the producer;
VmafxFrameImport.acquire_more and VmafxDeviceInfo.pci;
vmafx_frame_signal_on_release(). CUDA imports opaque memory as arrays
or buffers and waits on and signals the producer's timeline on its
stream; SYCL and HIP read LINEAR and DRM-modifier frames through their
dma-buf paths. Memory of another GPU, multi-plane images, Windows
handles, OPTIMAL tiling off CUDA and Vulkan semaphores off CUDA are
refused by name. Import only (Q-011): the library links no Vulkan
loader and makes no Vulkan call.

A SYNC_FILE fence is destroyed by closing its descriptor (ADR-2091
item 6), which the SYCL lane (#2342) now lands; this lane adds the
VULKAN_SEMAPHORE refusals of vmafx_fence_wait() / vmafx_fence_destroy()
next to it.

Landing (Q-083): the branch merges the SYCL and HIP lanes, so its tip
is not squashed. The change is rc4/integration's squashed variant
(798c42c, 541a91a, bae72c1, a21d20c: the lane's cff8500,
3decc47, 43b2231, c101d50 on the merged lanes) plus the lane's
last commit 8505d76 (the release canary scores each frame in two
contexts), cherry-picked onto #2368's landing head. The twelve
definition additions are ABI 0.1.10 (0.1.5 on the lane, ADR-1897).
ADR-2152 is Accepted (maintainer, Q-089). The rebase note is a
fragment (ADR-2197); the Vulkan agent page is in the caveman register.
The touched units are measured in the cpu, cuda, hip and sycl lanes (dev
image with the Vulkan loader and headers): 0 findings.

cuda/import_vulkan.c included <unistd.h> and called dup() / close()
unconditionally, which the Windows MSVC+CUDA build cannot compile; it
now uses VMAF_DUP (new in compat/crt_portable.h, _dup() on Windows) and
VMAF_CLOSE, as the other C runtime calls of the tree do.

* test(vulkan): judge the fence test's skipped wait only where no import waited for its writer

The acquire arm of test_vmafx_import_vulkan_fence demanded bad frames from
the planted skipped wait whenever any import returned before its write had
finished. On the gfx1036 under ROCm 10.1 the import normally waits for the
writer; one landing run of the HIP lane had 2 of 16 imports ahead, both
frames still read finished (0 bad), and the arm failed although the library
behaved correctly: a copy queued behind an early import can run after the
write has landed.

The arm now demands bad frames on CUDA, whose import never waits, and
wherever every import ran ahead of its write (SYCL); it demands none where
no import ran ahead, and prints a partial count without judging it. The
real arm is unchanged.

Not reproduced on demand (one run in 79 had imports ahead). With the new
rule: gfx1036 40 of 40, RTX 4090 10 of 10 and Arc A380 10 of 10 (skipped
wait 8 bad, 16 imports ahead, every run).

* ci(tidy): measure the translation units of the Vulkan frame import in the cpu, cuda, hip and sycl lanes

The tidy coverage rule requires every tracked translation unit to be read
by a lane. The units this pull request adds or touches were measured on
the rebased tree in the dev image with the Vulkan loader and headers
(scripts/dev/tidy-lane.sh --image vmaf-dev-mcp:local-vulkan --write
--only ..., clang-tidy 22.1.8): cpu 8, cuda 16, hip 15 and sycl 15
translation units, 0 findings and 0 uncited NOLINT in each lane. The
earlier record of these units fell away when the branch was rebased onto
the restacked SYCL lanes.

* ci(impact): route the Vulkan import tests to the cuda, hip and sycl tidy lanes

Master's test_ci_impact (TidyLaneRouting, from #2704) requires
every file a GPU tidy lane measures and the cpu lane does not to match
that lane's tidy_<lane> patterns in .github/ci-impact.json. This pull
request's Vulkan import tests (core/test/test_vmafx_import_vulkan*.c and
vmafx_vulkan_producer.c) are measured by the cuda, hip and sycl lanes
only and matched none, so the test failed on the rebased tree.
core/test/*vulkan* joins the three lanes' patterns; the test passes.

* test(vulkan): skip with the reason and run no case when the lane has no GPU Vulkan device

On a CPU-only lane (a hosted runner whose only Vulkan device is lavapipe)
the Vulkan import tests exit 77, as maintainer decision Q-346 settles: the
import reads a GPU's Vulkan memory, which lavapipe does not have, and the
real-GPU runs are the device signal. The skip now says so, and it never
reads like a pass:

- vk_open() names the reason on stderr: "skipped: no CUDA device, so no GPU
  Vulkan memory to import (lavapipe is a CPU device; the producer needs the
  importing GPU's memory)", or that no GPU Vulkan device sits on the lane
  device; it names the lane (CUDA, SYCL, HIP).
- test_vmafx_import_vulkan, _bitexact and _fence run no case without the
  GPU, so no case prints "pass" for work it did not do; the harness reports
  "0 tests run, skipped" and exits 77. The FFmpeg test names its two other
  reasons (inputs not set, no Vulkan decode on the GPU) the same way.

With the GPU the tests run as before (CUDA lane: import 5/5, fence 2/2).

* ci(tidy): run the hosted tidy lanes in the dev image that carries the Vulkan loader

#2714 (ADR-3137, operator decision ci-config-19) put the Vulkan loader,
its headers and lavapipe into the dev image. dev-container-publish.yml
published that build of master 9d3734f as
ghcr.io/vmafx/vmafx-dev-mcp:sha-9d3734f0602a3114849d8edce4136d9c8e7c850a,
digest sha256:29b13030cacce60cec4209f645c6772feda5b2d936fa511fc807eddccf12722c
(read from the registry with docker buildx imagetools inspect).
.github/actions/tidy-lane pins it, as its description asks of the pull
request that re-measures the baselines after a dev/Containerfile change.

Measured in that image (scripts/dev/tidy-lane.sh --image <digest> --write
--only ...): the cuda (16), hip (15) and sycl (15) units of this pull
request, the five Vulkan test files among them, 0 findings, 0 uncited
NOLINT, no compile failure; the records equal the ones already committed.

* test(vulkan): say when the FFmpeg test's inputs are not readable

With VMAFX_TEST_VULKAN_REF / _DIST set to files that do not exist the
skip said "no Vulkan decode of the inputs on this GPU", which blamed the
device. It now says "skipped: the inputs are not readable (<ref>, <dist>)"
and keeps the decode reason for inputs that exist but do not decode. The
cuda, hip and sycl tidy lanes measured the file in the published dev image:
0 findings, records unchanged.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the rc4/api-wp3-vulkan-import branch from 4b8610d to db110bd Compare October 10, 2026 11:08
@lusoris
lusoris merged commit db110bd into master Oct 10, 2026
149 of 168 checks passed
@lusoris
lusoris deleted the rc4/api-wp3-vulkan-import branch October 10, 2026 11:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag type:feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant