Repository navigation
feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023) - #2277
Merged
Merged
Conversation
13 of 18 tasks
lusoris
force-pushed
the
rc4/api-wp3-cuda
branch
from
October 6, 2026 12:43
ffaead6 to
15bd091
Compare
This was referenced Oct 6, 2026
Draft
lusoris
force-pushed
the
rc4/api-wp3-common
branch
4 times, most recently
from
October 7, 2026 23:36
1b7f00d to
9476e9b
Compare
lusoris
force-pushed
the
rc4/api-wp3-cuda
branch
from
October 8, 2026 11:42
15bd091 to
468f5cc
Compare
lusoris
marked this pull request as ready for review
October 8, 2026 11:42
… role from the platform definition (ADR-2350) (#2605) * feat(api): generate the custom resources, their CRDs and the operator role from the platform definition (ADR-2350) api/vmafx-platform.toml now declares the vmafx.dev/v1 group and its four resources: [[groups]], [[resources]], and [[messages]] / [[enums]] that name a group instead of a protobuf file. Fields carry the JSON name, an optional Go name and the OpenAPI validation (enum, items enum, minimum, min/max length and items, pattern, format, default). The generator validates the tables and writes api/vmafx/v1/groupversion_info.go and one gofmt-clean <kind>_types.go per resource with the kubebuilder markers, including the VmafxTenant type that had none. Each resource's DeepCopyInto is one call of a generated deepCopyResource helper, so the four copies are written once; controller-gen writes everything else. scripts/codegen/crd_generate.py runs controller-gen, pinned as a go.mod tool (v0.22.0), for the deepcopy code, the CRDs and the operator's RBAC role, and writes or checks api/vmafx/v1/zz_generated.deepcopy.go, deploy/helm/vmafx/crds/*.yaml and config/rbac/role.yaml. The chart's crds/ directory is the only CRD tree: config/crd/bases/, the hand-written deepcopy file and the per-kind roles are removed, and the operator's envtest suite loads the chart's CRDs. The installed schemas are unchanged apart from descriptions and the order of required lists. Within v1 a resource only grows: --compat-against REF compares every generated CRD with the one at REF and refuses a removed field, version or short name, a new required field, a narrower enum, bound, type or format, a changed pattern or default. Meson tests test_crd_generated_current and test_crd_compat run both checks; the helm-chart workflow checks that the chart grants the operator every rule of the generated role. The operator's leader-election lease gets its RBAC marker, so the generated role is complete. The output-tree sync is shared by the sqlc, buf and controller-gen runners. Signed-off-by: Lusoris <lusoris@proton.me>
…s (RC4 WP3, ADR-2023) (#2277) * feat(api): import CUDA device frames with event fences and GL textures (RC4 WP3, ADR-2023) The VMAFx API scores frames that already live on a CUDA device without a copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. - Devices: CUDA devices by index (retained primary context) or from the caller's context and stream (external[0], external[1]); count, info and describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC. vmafx_context_use_device imports the device into the context's engine, and features registered afterwards run on their CUDA twins. - Import: device pointers are bound where they are when each plane starts 8-byte aligned with a pitch that is a multiple of 8 (the twins' vector loads); other layouts are refused naming the field, or copied on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (import_convert.cu); CUDA arrays and GL textures are read out on the device. No import path copies through the host, and every host copy site calls vmafx_count_host_copy(). - Fences: a CUDA_EVENT acquire fence is waited on by the device's library stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context; the new VmafxFrameImport release callback (release, user; ABI 0.1.4) lets a producer make its stream wait on the release event without a host stall. The ADR-1199 barrier stays only for pictures that carry no fence ordering. - Pools: CUDA frame pools hand out device frames; a returned frame is reused only after the device readers of its previous use finished. integer_vif_cuda read both pictures with the pitch it computed at init, so an imported plane with another pitch was read from the wrong rows; scale 0 now reads each picture with its own pitch. Master cannot hand the engine a CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's host round trip is fixed in the commit before this one (#2282 on master). Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4): test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0 differing, 6028 imports and 0 host copies. A planted skipped acquire wait gives 15 bad frames of 16 under device load and the real wait 0; a planted early release gives 15 bad canaries, the real release 0. One import scored by two contexts equals each context's own run and is released only after the second context finished. The exact-twin matrix keeps 48 of 48 vif and float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3 skipped. * docs(api): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197) * docs(agents): write the vif_cuda pitch note in the internal register (praetor caveman lint) Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-wp3-cuda
branch
from
October 8, 2026 12:08
468f5cc to
c595358
Compare
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RC4 work package 3, CUDA lane (label
rc4): the VMAFx API scores frames that already live on a CUDA device, with no copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. It implements the CUDA lane of design section 2.7 of ADR-1852 / ADR-1929 and the CUDA row of the OpenGL interop item (#2238), and records the lane's choices in ADR-2023, Accepted by maintainer popup on 2026-10-06: "Accept, overlap as tuning later (Recommended)" (multi-stream overlap becomes an RC8 tuning row, measured, with the same fence tests).What it adds:
vmafx_device_createwithVMAFX_BACKEND_CUDAby index (the device's retained primary context) or from the caller's context and stream (external[0],external[1]);vmafx_device_count/vmafx_device_info/vmafx_device_describereport the memory kinds DEVICE_POINTER, DEVICE_ARRAY, GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT, GL_SYNC. Devices below the ADR-1223 architecture floor are refused.vmafx_context_use_deviceimports the device into the context's engine (the successor ofvmaf_cuda_import_state()), and a feature registered afterwards runs on its CUDA twin; a feature without one runs on the CPU with a warning, and admission then refuses device frames naming it.VMAFX_E_NOTSUPnamingplane[i].offsetorplane[i].pitch, or a copy on the device withVMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (core/src/cuda/import_convert.cu: de-interleave and, for P010, a shift by 6, nothing else). CUDA arrays and GL textures are read out on the device. No path copies through the host; every host copy site callsvmafx_count_host_copy().CUDA_EVENTacquire fence is waited on by the library stream (cuStreamWaitEvent), never on the host.VMAFX_FENCE_GL_SYNC(new, kind 8) orders a GL producer: the sync is checked on the host (glClientWaitSync, resolved at run time), and an unsignalled one isVMAFX_E_BUSY, which the D8 import rule retries. Release fences (HOST,CUDA_EVENT) are signalled after the last reader in every context:HOSTby a host function enqueued on the library stream at the last unref,CUDA_EVENTby an event recorded there. The newVmafxFrameImport.release/usercallback runs once the release event is recorded, so a producer makes its stream wait on it without a host stall (ADR-2023 item 4). The ADR-1199 barrier stays only for pictures that carry no fence ordering.VMAFX_MEMORY_GL_TEXTURE, new, kind 8): GL 2D textures, one per plane (NV12 as an R8 luma and an RG8 chroma texture, the layout a screen-capture pipeline renders), are registered read-only, mapped on the library stream, read out, unmapped and unregistered when the frame is released.vmafx_frame_pool_createon a CUDA device hands out device frames; a returned frame is handed out again only after the device readers of its previous use ran (an idle event recorded at release, waited on at acquire).ABI
0.1.3->0.1.4(additions only, nodeVMAFX_0.1, ADR-1897):VMAFX_MEMORY_GL_TEXTURE,VMAFX_FENCE_GL_SYNC,VmafxFrameImport.release(VmafxFrameReleaseCallback) and.user. A 0.1.2-sizedVmafxFrameImportis still accepted.--abi-check --against-ref origin/rc4/api-generation-prototype:definition is an append-only successor of origin/rc4/api-generation-prototype (141 additions); againstorigin/rc4/api-wp3-common:append-only successor ... (4 additions).Two CUDA twins fixed on the way (rows in
docs/state.md):integer_vif_cudaread both pictures with the pitch it computed at init (T-CUDA-VIF-PICTURE-PITCH-2026-10-06). An imported plane with another pitch (368 for 576x324) was read from the wrong rows: one import scored by two contexts gave 16 wrong values in the second. Scale 0 now reads each picture with its own pitch (VifBufferCuda.dis_stride,vif_vert_load_tiles). Not reachable on master, so it stays here. Every CUDA picture on master comes fromvmaf_cuda_picture_alloc()(cuMemAllocPitch), and the FFmpeglibvmaf_cudafilter copies its frames into those pool pictures (copy_picture_data_cuda(),dstPitch = dst->stride[i]). A probe on the RTX 4090 found thecuMemAllocPitchpitch equal to the init formula for every width from 1 to 8192 at 1 and 2 bytes per sample: 16384 allocations, 0 differ.float_ms_ssim_cudacopied every plane of every frame to the host (T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06, found by the vendor profiler trace): device to pinned host memory, a stream synchronisation per plane,picture_copy()on the host, upload. Level 0 is nowpicture_copy()on the device (ms_ssim_picture_to_float, the same samples, division by a power of two). A master defect, so it is master PR fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host #2282 (first in the train) with its own failing-first testtest_cuda_float_ms_ssim_host_traffic: on master, 3538944 bytes uploaded and 18 plane copies with a host side while scoring; after the fix, 0 and 0. The restack onto master dropped it from this PR.vmafx_context_use_featureregistered the CPU extractor on a device context (T-VMAFX-DEVICE-CONTEXT-CPU-EXTRACTOR-2026-10-06, defect in the draft WP3-common code): it now picks the twin of the device's backend (vmaf_engine_feature_backend_twin).Landing (Q-083)
Lands bottom-up per Q-083:
v1.0.0-rc.3is tagged, so RC4 lands through the merge train, one API PR at a time. Its base #2303 (WP6, the library split) is on master; this PR was squashed to its own change, rebased onto master3d67718bb(the base's commits dropped), retargeted tomaster, and #2287 (WP4) follows once it lands. Thefloat_ms_ssim_cudacommit this draft carried is on master as #2282 and was dropped. The API docs keep marking the VMAFx API as a preview (docs/api/vmafx/index.md, ABI 0.x) until the rc.4 cut. ADR-2023 is Accepted.ABI check against master:
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/master->definition is an append-only successor of origin/master (4 additions). Additions only, nodeVMAFX_0.1,abi_version0.1.7 (patch bump over master's).Rebase: the CUDA import sources join
libvmafx_sources(the engine library since the WP6 split) and the lane's white-box tests linkvmaf_test_link, as the merged RC4 integration branch did. Generated files take master's side and are regenerated once at the tip (vmafx-api.py --write, the AGENTS indexes); under render at landing (ADR-2197) this PR carries no rendered file (CHANGELOG.md, the ADR index, tag and title pages,docs/rebase-notes.md): its rebase note is the fragmentdocs/rebase-notes.d/vmafx-device-frames-cuda.md, and the citation map is derived from the tree (ADR-2200);docs/state.mdbyscripts/dev/resolve-state-md-conflict.py. The vif_cuda pitch note incore/src/feature/cuda/AGENTS.d/vif.mdis written in the internal register for the praetor caveman lint.Local gate on the rebased head (CPU,
-Db_lto=false,-j4, warnings as errors): build 0 warnings;--suite=fast411 OK, 0 fail (test_gpu_picture_pool_uafon its own withMALLOC_PERTURB_=0: OK); codegen tests 132 passed; affected suites: tooling 2499 passed, 6 skipped;make test-netflix-golden GOLDEN_NINJA_JOBS=4280 passed, 3 skipped;preflight.sh --stage msvcismpass; assertion density pass. CUDA build of the API stack through #2290 (CUDA 13.4, RTX 4090,-Denable_cuda=true, warnings as errors): 0 warnings,--suite=fast489 OK, 0 fail. Platform check: the GL interop resolvesglClientWaitSyncthroughdlsymon POSIX andGetProcAddressunder#ifdef _WIN32, and the device tests that use EGL are registered on Linux only. The merge train builds CPU and CUDA and runs its own gates (deliverables, state-md, silent revert,praetorctl audit).Type
feat— new featureChecklist
make format && make lintis green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, FFmpeg patch stack, HISS audit, assertion density).python3 scripts/ci/run_meson_test.py -- -C build-cpu --suite=fast --num-processes 4→ 364 OK, 0 fail, 1 skipped; CUDA build: the 63fast+gpuCUDA tests, the 4 lane programs, the cells contract,test_cuda_float_ms_ssim_host_trafficand the extended ms_ssim contract → 70 OK, 0 fail;test_cuda_parity_gate_default_runOK./cross-backend-diffand the worst ULP is ≤ 2. —scripts/ci/exact_twin_matrix.py --backends cuda --features vif float_ms_ssim float_ms_ssim_lcs float_ms_ssim_chroma: 48 cells (8 / 10 / 12 / 16 bit x 4:2:0 / 4:2:2 / 4:4:4), 48 equal to the CPU (0 ULP); the parity gate's default run passes..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (EUPL-1.2, fork-authored).!orBREAKING CHANGE:and the migration path is documented below. — not a breaking change: additions only;VmafxFrameImportgrew at its end and a 0.1.2-sized struct is accepted.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— ADR-2023 (number claimed withscripts/adr/next-free.sh --claim):docs/adr/_index_fragments/2023-vmafx-cuda-device-frames.md, slug appended, indexes regenerated.tidy: cuda
core/src/cuda/{import_device,import_fence,import_frame,import_gl,import_pool}.c,core/src/cuda/import_convert.cu,core/src/feature/cuda/{integer_vif_cuda,integer_ms_ssim_cuda}.c,core/src/libvmaf.c,core/src/vmafx/{context,device,device_context,fence,frame_host,frame_import,frame_import_admit,frame_pool,register,submit}.c,core/test/test_vmafx_import_cuda{,_bitexact,_fence,_gl}.c(and through themcore/src/cuda/vmafx_cuda{,_internal}.h,core/test/vmafx_cuda_{test_util,cells}.h) 0 findings, 0 uncited NOLINT; cpucore/src/libvmaf.c, the touchedcore/src/vmafx/*.c,core/test/test_vmafx_{import_api,frame,abi_layout,import_bitexact,import_fence}.c0 findings, 0 uncited NOLINT (dev container, clang-tidy 22.1.8,scripts/dev/tidy-lane.sh --only ... <lane>).ms_ssim_score.cuandfilter1d.cumeasure 39 and 30, their baselines on this stack's base: the kernel added toms_ssim_score.cuadds none (it takesfloat *dst, so it holds no integer-to-pointer cast) andfilter1d.cu's change adds none; master cleaned both files in #2109, which the restack of this stack resolves. The first runs found 54 (enum casts ofCUresultvalues the driver headers lack, misplaced const on handle typedefs, integer-to-pointer casts of CUDA handles, analyzer array bounds, padding, widening, test leaks on failure paths) and one NOLINT without an ADR; all fixed, every remaining NOLINT cites its ADR.scripts/dev/preflight.sh --stage msvcism: pass.core/test/test_win32_pthread_shim_contract.py: pass.scripts/ci/assertion-density.sh: pass.praetorctl audit: governance gates passed, no HISS finding in a touched file.Bug-status hygiene (ADR-0165)
docs/state.mdupdated — closed rows T-CUDA-VIF-PICTURE-PITCH-2026-10-06, T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06 and T-VMAFX-DEVICE-CONTEXT-CPU-EXTRACTOR-2026-10-06 (above).Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Golden gate (
GOLDEN_NINJA_JOBS=4 make test-netflix-golden,core/build-goldenbuilt with gcc, pytest from the repository venv): 280 passed, 3 skipped.Cross-backend numerical results
Measured on an RTX 4090 (sm_89), driver 615.71.09, CUDA 13.4, every device run under
flock ~/.cache/vmafx-locks/cuda-4090.lock.test_vmafx_import_cuda_bitexact: every cell ofscripts/ci/exact_twins.d/*.cuda(core/test/vmafx_cuda_cells.h, held to the declared list bytest_vmafx_import_cuda_cells_contract.py) on the 576x324 golden pair, both checkerboards, Sparks 10-bit and 4K bbb (6 frames), as planar and as NV12 / P010, with row paddingtest_vmafx_import_cuda_fencetest_acquire_order_under_load(modelled on the ADR-1199 harness: the producer writes each frame on its own stream behind a host function that holds the stream for a few milliseconds, a load thread keeps the device busy; psnr and adm against host runs)VMAFX_TEST_SKIP_ACQUIRE_WAIT: 15 bad of 16test_release_canary(the library stream is held while the readers queue; the release callback makes the producer's stream wait on theCUDA_EVENTrelease fence and write a canary into the frame, no host wait)VMAFX_TEST_EARLY_RELEASE: 15 bad of 16test_host_release_after_deviceHOSTrelease fence stays pending while the device still holds the framevmafx_count_host_copy()at every host copy site,libvmaf.c's download included)VMAFX_TEST_FORCE_HOST_COPY: countedtest_vmafx_import_cudatest_one_import_two_contexts(vmaf_v0.6.1andvmaf_v1.0.16_3d0h)CUDA_EVENTandHOSTrelease only after the second context finishedtest_pool_frames_ordered_by_barriertest_vmafx_import_cuda_gl(headless EGL device, NV12 rendered into R8 + RG8 textures, GL sync acquire)test_vmafx_import_cudaVendor profiler trace (Nsight Systems 2026.3.2,
--trace=cuda, the import sessions oftest_vmafx_import_cuda_bitexactwithVMAFX_TEST_IMPORT_ONLY=1): library stream 0 host-to-device, 0 device-to-host copies; 2340 device-to-device copies on it (2174 MB), all the twins' own packing of picture planes into their buffers. Producer stream: 650 host-to-device copies (435 MB), the test's uploads. Elsewhere: 60 host-to-device copies of at most 66 KB (the twins' constant tables) and the twins' result and term readbacks on their private streams. Before thefloat_ms_ssim_cudafix the same trace showed 4227 MB host to device.Gates shown failing on planted defects (measured on this branch)
VMAFX_TEST_SKIP_ACQUIRE_WAIT)test_vmafx_import_cuda_fenceVMAFX_TEST_EARLY_RELEASE)test_vmafx_import_cuda_fenceVMAFX_TEST_FORCE_HOST_COPY)test_vmafx_import_cudatest_vmafx_import_cuda_bitexacttest_vmafx_import_cuda_bitexactinteger_vif_cudareads with the init pitch (the defect fixed here)test_vmafx_import_cuda_bitexact,test_vmafx_import_cudatest_vmafx_import_cuda,test_vmafx_import_cuda_fencetest_vmafx_import_cuda_bitexacttest_vmafx_import_cudatest_import_refusals: the unaligned start is not named (the twin then faults withCUDA_ERROR_MISALIGNED_ADDRESS)test_vmafx_import_cuda_gltest_gl_sync_needs_contextfailsHOSTrelease signalled at the last unref instead of after the devicetest_vmafx_import_cuda_fencetest_host_release_after_devicefailstest_vmafx_import_cuda_fencetest_vmafx_import_cuda_fencetest_vmafx_import_cuda_cells_contract.pyPerformance (if
perforfeat)No timing claim (RC7). A planar import binds the producer's memory; a semi-planar import converts once on the device.
float_ms_ssim_cudano longer copies 4 planes per frame through the host and no longer synchronises the stream per plane. Frames of one device are read on one stream, so two frames' kernels do not overlap across streams: an RC8 tuning row (ADR-2023, accepted), measured and held to the same fence tests.Deep-dive deliverables (ADR-0108)
docs/adr/2023-vmafx-cuda-device-frames.md## Alternatives considered(stream per device or per frame, release event at release or behind a gate stream, barrier policy, unaligned planes, GL sync wait, array read-out).AGENTS.mdinvariant note —core/src/cuda/AGENTS.md(new section: library stream, fences, release registry, alignment, pools, host-copy counter),core/src/AGENTS.d/vmafx-device-frames.md,core/src/feature/cuda/AGENTS.d/vif.md(per-picture pitch) andms-ssim.md(level 0 on the device);docs/development/rebase-sensitive-invariants.mdentry "VMAFx device frames on CUDA".changelog.d/added/api-cuda-device-frames.md,changelog.d/fixed/cuda-vif-picture-pitch.md,changelog.d/changed/cuda-ms-ssim-device-level0.md.docs/rebase-notes.md, "VMAFx device frames on CUDA (RC4 WP3 CUDA lane)".User documentation:
docs/api/vmafx/index.mdgains "CUDA devices" (creating a device, memory kinds and their layout rules, fences, release callback with a decoder example, ordering, pools, profiling);docs/backends/cuda/overview.mdgains "Importing frames through the VMAFx API"; the generated reference pages cover the new kinds and fields.Reproducer
The GL test needs an EGL device (headless works); without one it skips.
Known follow-ups
origin/masterand does not merge cleanly there (WP3-common already conflicts in 14 files). Once fix(cuda): build float_ms_ssim's level 0 on the device instead of round-tripping every plane through the host #2282 has landed, the restack drops the carried first commit and takes master's side ofms_ssim_score.cuandinteger_ms_ssim_cuda.c.filter1d.cumerges cleanly. Master'scheck-tidy-coveragehook will then require this PR's new translation units inscripts/ci/tidy-baseline-cuda.jsonmeasured_sources:tidy-lane.sh --write --only <unit> cudaat the restack.check-silent-revert.pyagainstorigin/mastercannot run until then ("the merge does not resolve cleanly"); againstorigin/rc4/api-wp3-commonit is clean.wglGetProcAddress) is compiled but not run here.test_vmafx_import_cuda_fence(acquire under load, release canary, pool frames without a fence).CUDA_EVENTrelease fence at once; a device wait issued before the release callback ran is the caller's error (documented).VMAFX_DEVICE_PROFILINGstaysVMAFX_E_NOTSUPon CUDA: the vendor profiler serves.float_ms_ssim_cudaat 9, 11, 13–15 bits takes the one-byte path aspicture_copy()does; the exact-twin matrix covers 8 / 10 / 12 / 16 only.AV_PIX_FMT_CUDAframes with a device made from the frames context's context and stream, and make the decoder's stream wait on the release event from the release callback.VMAFX_MEMORY_GL_TEXTURE,VMAFX_FENCE_GL_SYNCand the release callback; re-check their twins against foreign pitches.