Skip to content

feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023) - #2277

Merged
lusoris merged 2 commits into
masterfrom
rc4/api-wp3-cuda
Oct 8, 2026
Merged

lusoris merged 2 commits into
masterfrom
rc4/api-wp3-cuda

Conversation

@lusoris

@lusoris lusoris commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

RC4 work package 3, CUDA lane (label rc4): the VMAFx API scores frames that already live on a CUDA device, with no copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. It implements the CUDA lane of design section 2.7 of ADR-1852 / ADR-1929 and the CUDA row of the OpenGL interop item (#2238), and records the lane's choices in ADR-2023, Accepted by maintainer popup on 2026-10-06: "Accept, overlap as tuning later (Recommended)" (multi-stream overlap becomes an RC8 tuning row, measured, with the same fence tests).

What it adds:

  • CUDA devices. vmafx_device_create with VMAFX_BACKEND_CUDA by index (the device's retained primary context) or from the caller's context and stream (external[0], external[1]); vmafx_device_count / vmafx_device_info / vmafx_device_describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY, GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT, GL_SYNC. Devices below the ADR-1223 architecture floor are refused. vmafx_context_use_device imports the device into the context's engine (the successor of vmaf_cuda_import_state()), and a feature registered afterwards runs on its CUDA twin; a feature without one runs on the CPU with a warning, and admission then refuses device frames naming it.
  • One library stream per device (ADR-2023 item 1). Every imported or pooled frame is a CUDA picture on it, so every reader of a frame is ordered on one stream.
  • Import. A device-pointer plane is bound where it is when it starts 8-byte aligned with a pitch that is a multiple of 8 (the twins read rows with vector loads); any other layout is VMAFX_E_NOTSUP naming plane[i].offset or plane[i].pitch, or a copy on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (core/src/cuda/import_convert.cu: de-interleave and, for P010, a shift by 6, nothing else). CUDA arrays and GL textures are read out on the device. No path copies through the host; every host copy site calls vmafx_count_host_copy().
  • Fences. A CUDA_EVENT acquire fence is waited on by the library stream (cuStreamWaitEvent), never on the host. VMAFX_FENCE_GL_SYNC (new, kind 8) orders a GL producer: the sync is checked on the host (glClientWaitSync, resolved at run time), and an unsignalled one is VMAFX_E_BUSY, which the D8 import rule retries. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context: HOST by a host function enqueued on the library stream at the last unref, CUDA_EVENT by an event recorded there. The new VmafxFrameImport.release / user callback runs once the release event is recorded, so a producer makes its stream wait on it without a host stall (ADR-2023 item 4). The ADR-1199 barrier stays only for pictures that carry no fence ordering.
  • OpenGL interop (VMAFX_MEMORY_GL_TEXTURE, new, kind 8): GL 2D textures, one per plane (NV12 as an R8 luma and an RG8 chroma texture, the layout a screen-capture pipeline renders), are registered read-only, mapped on the library stream, read out, unmapped and unregistered when the frame is released.
  • CUDA frame pools. vmafx_frame_pool_create on a CUDA device hands out device frames; a returned frame is handed out again only after the device readers of its previous use ran (an idle event recorded at release, waited on at acquire).

ABI 0.1.3 -> 0.1.4 (additions only, node VMAFX_0.1, ADR-1897): VMAFX_MEMORY_GL_TEXTURE, VMAFX_FENCE_GL_SYNC, VmafxFrameImport.release (VmafxFrameReleaseCallback) and .user. A 0.1.2-sized VmafxFrameImport is still accepted. --abi-check --against-ref origin/rc4/api-generation-prototype: definition is an append-only successor of origin/rc4/api-generation-prototype (141 additions); against origin/rc4/api-wp3-common: append-only successor ... (4 additions).

Two CUDA twins fixed on the way (rows in docs/state.md):

  • integer_vif_cuda read both pictures with the pitch it computed at init (T-CUDA-VIF-PICTURE-PITCH-2026-10-06). An imported plane with another pitch (368 for 576x324) was read from the wrong rows: one import scored by two contexts gave 16 wrong values in the second. Scale 0 now reads each picture with its own pitch (VifBufferCuda.dis_stride, vif_vert_load_tiles). Not reachable on master, so it stays here. Every CUDA picture on master comes from vmaf_cuda_picture_alloc() (cuMemAllocPitch), and the FFmpeg libvmaf_cuda filter copies its frames into those pool pictures (copy_picture_data_cuda(), dstPitch = dst->stride[i]). A probe on the RTX 4090 found the cuMemAllocPitch pitch equal to the init formula for every width from 1 to 8192 at 1 and 2 bytes per sample: 16384 allocations, 0 differ.
  • float_ms_ssim_cuda copied every plane of every frame to the host (T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06, found by the vendor profiler trace): device to pinned host memory, a stream synchronisation per plane, picture_copy() on the host, upload. Level 0 is now picture_copy() on the device (ms_ssim_picture_to_float, the same samples, division by a power of two). A master defect, so it is master PR fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host #2282 (first in the train) with its own failing-first test test_cuda_float_ms_ssim_host_traffic: on master, 3538944 bytes uploaded and 18 plane copies with a host side while scoring; after the fix, 0 and 0. The restack onto master dropped it from this PR.
  • vmafx_context_use_feature registered the CPU extractor on a device context (T-VMAFX-DEVICE-CONTEXT-CPU-EXTRACTOR-2026-10-06, defect in the draft WP3-common code): it now picks the twin of the device's backend (vmaf_engine_feature_backend_twin).

Landing (Q-083)

Lands bottom-up per Q-083: v1.0.0-rc.3 is tagged, so RC4 lands through the merge train, one API PR at a time. Its base #2303 (WP6, the library split) is on master; this PR was squashed to its own change, rebased onto master 3d67718bb (the base's commits dropped), retargeted to master, and #2287 (WP4) follows once it lands. The float_ms_ssim_cuda commit this draft carried is on master as #2282 and was dropped. The API docs keep marking the VMAFx API as a preview (docs/api/vmafx/index.md, ABI 0.x) until the rc.4 cut. ADR-2023 is Accepted.

ABI check against master: python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/master -> definition is an append-only successor of origin/master (4 additions). Additions only, node VMAFX_0.1, abi_version 0.1.7 (patch bump over master's).

Rebase: the CUDA import sources join libvmafx_sources (the engine library since the WP6 split) and the lane's white-box tests link vmaf_test_link, as the merged RC4 integration branch did. Generated files take master's side and are regenerated once at the tip (vmafx-api.py --write, the AGENTS indexes); under render at landing (ADR-2197) this PR carries no rendered file (CHANGELOG.md, the ADR index, tag and title pages, docs/rebase-notes.md): its rebase note is the fragment docs/rebase-notes.d/vmafx-device-frames-cuda.md, and the citation map is derived from the tree (ADR-2200); docs/state.md by scripts/dev/resolve-state-md-conflict.py. The vif_cuda pitch note in core/src/feature/cuda/AGENTS.d/vif.md is written in the internal register for the praetor caveman lint.

Local gate on the rebased head (CPU, -Db_lto=false, -j4, warnings as errors): build 0 warnings; --suite=fast 411 OK, 0 fail (test_gpu_picture_pool_uaf on its own with MALLOC_PERTURB_=0: OK); codegen tests 132 passed; affected suites: tooling 2499 passed, 6 skipped; make test-netflix-golden GOLDEN_NINJA_JOBS=4 280 passed, 3 skipped; preflight.sh --stage msvcism pass; assertion density pass. CUDA build of the API stack through #2290 (CUDA 13.4, RTX 4090, -Denable_cuda=true, warnings as errors): 0 warnings, --suite=fast 489 OK, 0 fail. Platform check: the GL interop resolves glClientWaitSync through dlsym on POSIX and GetProcAddress under #ifdef _WIN32, and the device tests that use EGL are registered on Linux only. The merge train builds CPU and CUDA and runs its own gates (deliverables, state-md, silent revert, praetorctl audit).

Type

  • feat — new feature

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, FFmpeg patch stack, HISS audit, assertion density).
  • Unit tests pass: CPU build python3 scripts/ci/run_meson_test.py -- -C build-cpu --suite=fast --num-processes 4 → 364 OK, 0 fail, 1 skipped; CUDA build: the 63 fast+gpu CUDA tests, the 4 lane programs, the cells contract, test_cuda_float_ms_ssim_host_traffic and the extended ms_ssim contract → 70 OK, 0 fail; test_cuda_parity_gate_default_run OK.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. — scripts/ci/exact_twin_matrix.py --backends cuda --features vif float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chroma: 48 cells (8 / 10 / 12 / 16 bit x 4:2:0 / 4:2:2 / 4:4:4), 48 equal to the CPU (0 ULP); the parity gate's default run passes.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. — the two fixes are CUDA-only defects (the CPU, SIMD, SYCL and HIP paths do not stage through the host or assume a pitch the same way); the SYCL / HIP import lanes re-check their twins against foreign pitches.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (EUPL-1.2, fork-authored).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. — not a breaking change: additions only; VmafxFrameImport grew at its end and a 0.1.2-sized struct is accepted.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — ADR-2023 (number claimed with scripts/adr/next-free.sh --claim): docs/adr/_index_fragments/2023-vmafx-cuda-device-frames.md, slug appended, indexes regenerated.

tidy: cuda core/src/cuda/{import_device,import_fence,import_frame,import_gl,import_pool}.c, core/src/cuda/import_convert.cu, core/src/feature/cuda/{integer_vif_cuda,integer_ms_ssim_cuda}.c, core/src/libvmaf.c, core/src/vmafx/{context,device,device_context,fence,frame_host,frame_import,frame_import_admit,frame_pool,register,submit}.c, core/test/test_vmafx_import_cuda{,_bitexact,_fence,_gl}.c (and through them core/src/cuda/vmafx_cuda{,_internal}.h, core/test/vmafx_cuda_{test_util,cells}.h) 0 findings, 0 uncited NOLINT; cpu core/src/libvmaf.c, the touched core/src/vmafx/*.c, core/test/test_vmafx_{import_api,frame,abi_layout,import_bitexact,import_fence}.c 0 findings, 0 uncited NOLINT (dev container, clang-tidy 22.1.8, scripts/dev/tidy-lane.sh --only ... <lane>). ms_ssim_score.cu and filter1d.cu measure 39 and 30, their baselines on this stack's base: the kernel added to ms_ssim_score.cu adds none (it takes float *dst, so it holds no integer-to-pointer cast) and filter1d.cu's change adds none; master cleaned both files in #2109, which the restack of this stack resolves. The first runs found 54 (enum casts of CUresult values the driver headers lack, misplaced const on handle typedefs, integer-to-pointer casts of CUDA handles, analyzer array bounds, padding, widening, test leaks on failure paths) and one NOLINT without an ADR; all fixed, every remaining NOLINT cites its ADR.

scripts/dev/preflight.sh --stage msvcism: pass. core/test/test_win32_pthread_shim_contract.py: pass. scripts/ci/assertion-density.sh: pass. praetorctl audit: governance gates passed, no HISS finding in a touched file.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated — closed rows T-CUDA-VIF-PICTURE-PITCH-2026-10-06, T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06 and T-VMAFX-DEVICE-CONTEXT-CPU-EXTRACTOR-2026-10-06 (above).

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below — not applicable.

Golden gate (GOLDEN_NINJA_JOBS=4 make test-netflix-golden, core/build-golden built with gcc, pytest from the repository venv): 280 passed, 3 skipped.

Cross-backend numerical results

Measured on an RTX 4090 (sm_89), driver 615.71.09, CUDA 13.4, every device run under flock ~/.cache/vmafx-locks/cuda-4090.lock.

Exit evidence (ADR-1829) Test Result
Imported frames score bit for bit as host-uploaded frames for every exact CUDA twin test_vmafx_import_cuda_bitexact: every cell of scripts/ci/exact_twins.d/*.cuda (core/test/vmafx_cuda_cells.h, held to the declared list by test_vmafx_import_cuda_cells_contract.py) on the 576x324 golden pair, both checkerboards, Sparks 10-bit and 4K bbb (6 frames), as planar and as NV12 / P010, with row padding 236 cells, 10732 values, 0 differing; 6028 imports, 3014 conversions, 0 host copies
A skipped acquire wait gives wrong scores under load; the real wait 0 bad of N test_vmafx_import_cuda_fence test_acquire_order_under_load (modelled on the ADR-1199 harness: the producer writes each frame on its own stream behind a host function that holds the stream for a few milliseconds, a load thread keeps the device busy; psnr and adm against host runs) real wait: 0 bad of 16 (psnr, adm); planted VMAFX_TEST_SKIP_ACQUIRE_WAIT: 15 bad of 16
Release-fence canary test_release_canary (the library stream is held while the readers queue; the release callback makes the producer's stream wait on the CUDA_EVENT release fence and write a canary into the frame, no host wait) real release: 0 bad of 16, 0 early canaries; planted VMAFX_TEST_EARLY_RELEASE: 15 bad of 16
Host release after the device test_host_release_after_device a HOST release fence stays pending while the device still holds the frame
Host-copy counter 0 every lane program (vmafx_count_host_copy() at every host copy site, libvmaf.c's download included) 0; planted VMAFX_TEST_FORCE_HOST_COPY: counted
One import scored by two contexts test_vmafx_import_cuda test_one_import_two_contexts (vmaf_v0.6.1 and vmaf_v1.0.16_3d0h) each context equals its own run (56 values each, 0 differing); CUDA_EVENT and HOST release only after the second context finished
Pool frames without a fence test_pool_frames_ordered_by_barrier 0 bad of 16
GL textures with a GL sync test_vmafx_import_cuda_gl (headless EGL device, NV12 rendered into R8 + RG8 textures, GL sync acquire) 42 values, 0 differing, 0 host copies; an import without a GL context current on the thread is refused naming the fence
CUDA arrays, ALLOW_COPY, pools test_vmafx_import_cuda 56 values each, 0 differing

Vendor profiler trace (Nsight Systems 2026.3.2, --trace=cuda, the import sessions of test_vmafx_import_cuda_bitexact with VMAFX_TEST_IMPORT_ONLY=1): library stream 0 host-to-device, 0 device-to-host copies; 2340 device-to-device copies on it (2174 MB), all the twins' own packing of picture planes into their buffers. Producer stream: 650 host-to-device copies (435 MB), the test's uploads. Elsewhere: 60 host-to-device copies of at most 66 KB (the twins' constant tables) and the twins' result and term readbacks on their private streams. Before the float_ms_ssim_cuda fix the same trace showed 4227 MB host to device.

Gates shown failing on planted defects (measured on this branch)

Planted defect Test Result
acquire wait skipped (VMAFX_TEST_SKIP_ACQUIRE_WAIT) test_vmafx_import_cuda_fence 15 bad of 16 under load (asserted every run)
release recorded when the frame is submitted (VMAFX_TEST_EARLY_RELEASE) test_vmafx_import_cuda_fence 15 bad canaries of 16 (asserted every run)
planar import staged through the host (VMAFX_TEST_FORCE_HOST_COPY) test_vmafx_import_cuda host-copy counter 3 (asserted every run)
P010 shift 7 instead of 6 test_vmafx_import_cuda_bitexact psnr 480x270 10-bit P010: 5 values differ
Cb / Cr swapped in the de-interleave test_vmafx_import_cuda_bitexact psnr 576x324 NV12: 96 values differ
integer_vif_cuda reads with the init pitch (the defect fixed here) test_vmafx_import_cuda_bitexact, test_vmafx_import_cuda vif cell: 192 values differ; second context: 16 differ
release event recorded at request time test_vmafx_import_cuda, test_vmafx_import_cuda_fence two contexts: released while the second still holds the frame; canary 15 bad of 16
no twin picked on a device context (the defect fixed here) test_vmafx_import_cuda_bitexact adm session refused: CPU extractor would need a host copy
no alignment check on bound planes test_vmafx_import_cuda test_import_refusals: the unaligned start is not named (the twin then faults with CUDA_ERROR_MISALIGNED_ADDRESS)
GL sync ignored test_vmafx_import_cuda_gl test_gl_sync_needs_context fails
HOST release signalled at the last unref instead of after the device test_vmafx_import_cuda_fence test_host_release_after_device fails
pool frame handed out without the idle wait test_vmafx_import_cuda_fence 5 bad of 16
ADR-1199 barrier skipped for every frame test_vmafx_import_cuda_fence pool frames without a fence: 15 bad of 16
cells table missing / adding / mis-aliasing a declared exact twin test_vmafx_import_cuda_cells_contract.py each refused (planted cases run every time)

Performance (if perf or feat)

No timing claim (RC7). A planar import binds the producer's memory; a semi-planar import converts once on the device. float_ms_ssim_cuda no longer copies 4 planes per frame through the host and no longer synchronises the stream per plane. Frames of one device are read on one stream, so two frames' kernels do not overlap across streams: an RC8 tuning row (ADR-2023, accepted), measured and held to the same fence tests.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: implements Research-2158 section 2.7 (docs(adr): accept the VMAFx API redesign and VMAFx filter names for RC4 (ADR-1852) #2176) for CUDA; the measurements that decided the lane's open points (the gate-stream deadlock, the VIF pitch, the ms_ssim host staging, the alignment fault) are recorded in ADR-2023's Context and Alternatives and in the three bug-ledger rows above.
  • Decision matrix — docs/adr/2023-vmafx-cuda-device-frames.md ## Alternatives considered (stream per device or per frame, release event at release or behind a gate stream, barrier policy, unaligned planes, GL sync wait, array read-out).
  • AGENTS.md invariant note — core/src/cuda/AGENTS.md (new section: library stream, fences, release registry, alignment, pools, host-copy counter), core/src/AGENTS.d/vmafx-device-frames.md, core/src/feature/cuda/AGENTS.d/vif.md (per-picture pitch) and ms-ssim.md (level 0 on the device); docs/development/rebase-sensitive-invariants.md entry "VMAFx device frames on CUDA".
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/added/api-cuda-device-frames.md, changelog.d/fixed/cuda-vif-picture-pitch.md, changelog.d/changed/cuda-ms-ssim-device-level0.md.
  • Rebase note — docs/rebase-notes.md, "VMAFx device frames on CUDA (RC4 WP3 CUDA lane)".

User documentation: docs/api/vmafx/index.md gains "CUDA devices" (creating a device, memory kinds and their layout rules, fences, release callback with a decoder example, ordering, pools, profiling); docs/backends/cuda/overview.md gains "Importing frames through the VMAFx API"; the generated reference pages cover the new kinds and fields.

Reproducer

python3 scripts/codegen/vmafx-api.py --check
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/rc4/api-generation-prototype
meson setup build-cuda core -Denable_cuda=true -Denable_dnn=disabled -Db_lto=false
nice -n 10 ninja -C build-cuda -j6
# fixtures: python/test/resource/yuv (golden pair, checkerboards, Sparks) and testdata/bbb (4K)
flock ~/.cache/vmafx-locks/cuda-4090.lock timeout 300 \
  python3 scripts/ci/run_meson_test.py -- -C build-cuda --num-processes 1 \
  test_vmafx_import_cuda test_vmafx_import_cuda_bitexact test_vmafx_import_cuda_fence \
  test_vmafx_import_cuda_gl test_vmafx_import_cuda_cells_contract
python3 scripts/ci/exact_twin_matrix.py --vmaf-binary build-cuda/tools/vmaf --backends cuda \
  --features vif float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chroma   # takes the lock itself
# vendor trace of the import sessions
VMAFX_TEST_IMPORT_ONLY=1 nsys profile --trace=cuda --output import_only \
  build-cuda/test/test_vmafx_import_cuda_bitexact

The GL test needs an EGL device (headless works); without one it skips.

Known follow-ups

  • Restack onto master. This stack's base is behind origin/master and does not merge cleanly there (WP3-common already conflicts in 14 files). Once fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host #2282 has landed, the restack drops the carried first commit and takes master's side of ms_ssim_score.cu and integer_ms_ssim_cuda.c. filter1d.cu merges cleanly. Master's check-tidy-coverage hook will then require this PR's new translation units in scripts/ci/tidy-baseline-cuda.json measured_sources: tidy-lane.sh --write --only <unit> cuda at the restack. check-silent-revert.py against origin/master cannot run until then ("the merge does not resolve cleanly"); against origin/rc4/api-wp3-common it is clean.
  • GL: a texture is registered and unregistered per import (a registration cache is an RC7 tuning row); the GL sync is waited on the host, not on the device; the Windows GL path (wglGetProcAddress) is compiled but not run here.
  • One library stream per device: no cross-frame kernel overlap; multi-stream overlap is an RC8 tuning row, measured, held to test_vmafx_import_cuda_fence (acquire under load, release canary, pool frames without a fence).
  • Release events: at most 4096 frames with a pending CUDA_EVENT release fence at once; a device wait issued before the release callback ran is the caller's error (documented).
  • VMAFX_DEVICE_PROFILING stays VMAFX_E_NOTSUP on CUDA: the vendor profiler serves.
  • float_ms_ssim_cuda at 9, 11, 13–15 bits takes the one-byte path as picture_copy() does; the exact-twin matrix covers 8 / 10 / 12 / 16 only.
  • WP9 (filters): import AV_PIX_FMT_CUDA frames with a device made from the frames context's context and stream, and make the decoder's stream wait on the release event from the release callback.
  • SYCL / HIP lanes: reuse VMAFX_MEMORY_GL_TEXTURE, VMAFX_FENCE_GL_SYNC and the release callback; re-check their twins against foreign pitches.

@lusoris lusoris added type:feature New feature or request rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag labels Oct 6, 2026
@lusoris
lusoris force-pushed the rc4/api-wp3-common branch 4 times, most recently from 1b7f00d to 9476e9b Compare October 7, 2026 23:36
Base automatically changed from rc4/api-wp3-common to master October 7, 2026 23:40
@lusoris
lusoris marked this pull request as ready for review October 8, 2026 11:42
… role from the platform definition (ADR-2350) (#2605)

* feat(api): generate the custom resources, their CRDs and the operator role from the platform definition (ADR-2350)

api/vmafx-platform.toml now declares the vmafx.dev/v1 group and its four
resources: [[groups]], [[resources]], and [[messages]] / [[enums]] that
name a group instead of a protobuf file. Fields carry the JSON name, an
optional Go name and the OpenAPI validation (enum, items enum, minimum,
min/max length and items, pattern, format, default). The generator
validates the tables and writes api/vmafx/v1/groupversion_info.go and one
gofmt-clean <kind>_types.go per resource with the kubebuilder markers,
including the VmafxTenant type that had none. Each resource's
DeepCopyInto is one call of a generated deepCopyResource helper, so the
four copies are written once; controller-gen writes everything else.

scripts/codegen/crd_generate.py runs controller-gen, pinned as a go.mod
tool (v0.22.0), for the deepcopy code, the CRDs and the operator's RBAC
role, and writes or checks api/vmafx/v1/zz_generated.deepcopy.go,
deploy/helm/vmafx/crds/*.yaml and config/rbac/role.yaml. The chart's
crds/ directory is the only CRD tree: config/crd/bases/, the hand-written
deepcopy file and the per-kind roles are removed, and the operator's
envtest suite loads the chart's CRDs. The installed schemas are unchanged
apart from descriptions and the order of required lists.

Within v1 a resource only grows: --compat-against REF compares every
generated CRD with the one at REF and refuses a removed field, version or
short name, a new required field, a narrower enum, bound, type or format,
a changed pattern or default. Meson tests test_crd_generated_current and
test_crd_compat run both checks; the helm-chart workflow checks that the
chart grants the operator every rule of the generated role. The
operator's leader-election lease gets its RBAC marker, so the generated
role is complete. The output-tree sync is shared by the sqlc, buf and
controller-gen runners.

Signed-off-by: Lusoris <lusoris@proton.me>
…s (RC4 WP3, ADR-2023) (#2277)

* feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023)

The VMAFx API scores frames that already live on a CUDA device without a
copy through the host, and imported frames score bit for bit as the same
frames uploaded from the host for every CUDA twin declared exact.

- Devices: CUDA devices by index (retained primary context) or from the
  caller's context and stream (external[0], external[1]); count, info and
  describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and
  GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC.
  vmafx_context_use_device imports the device into the context's engine, and
  features registered afterwards run on their CUDA twins.
- Import: device pointers are bound where they are when each plane starts
  8-byte aligned with a pitch that is a multiple of 8 (the twins' vector
  loads); other layouts are refused naming the field, or copied on the
  device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on
  the device (import_convert.cu); CUDA arrays and GL textures are read out
  on the device. No import path copies through the host, and every host copy
  site calls vmafx_count_host_copy().
- Fences: a CUDA_EVENT acquire fence is waited on by the device's library
  stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL
  textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after
  the last reader in every context; the new VmafxFrameImport release callback
  (release, user; ABI 0.1.4) lets a producer make its stream wait on the
  release event without a host stall. The ADR-1199 barrier stays only for
  pictures that carry no fence ordering.
- Pools: CUDA frame pools hand out device frames; a returned frame is reused
  only after the device readers of its previous use finished.

integer_vif_cuda read both pictures with the pitch it computed at init, so
an imported plane with another pitch was read from the wrong rows; scale 0
now reads each picture with its own pitch. Master cannot hand the engine a
CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's
host round trip is fixed in the commit before this one (#2282 on master).

Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4):
test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both
checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0
differing, 6028 imports and 0 host copies. A planted skipped acquire wait
gives 15 bad frames of 16 under device load and the real wait 0; a planted
early release gives 15 bad canaries, the real release 0. One import scored
by two contexts equals each context's own run and is released only after
the second context finished. The exact-twin matrix keeps 48 of 48 vif and
float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3
skipped.

* docs(api): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197)

* docs(agents): write the vif_cuda pitch note in the internal register (praetor caveman lint)

Signed-off-by: Lusoris <lusoris@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag type:feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant