Skip to content

RC4 integration - #2343

Draft
lusoris wants to merge 77 commits into
masterfrom
rc4/integration
Draft

lusoris wants to merge 77 commits into
masterfrom
rc4/integration

Conversation

@lusoris

@lusoris lusoris commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

One branch, rc4/integration, carries every RC4 API draft so they can be built and tested together: WP6 #2303 (with WP5 #2288, WP8 #2258, WP3-common #2213, WP2 #2199, WP1 #2187 and the prototype #2173 below it), the WP3 CUDA lane #2277, WP4 window scores #2287, incremental motion #2290, and since 2026-10-07 the WP3 HIP lane #2341 with its GL follow-up #2360, the WP3 SYCL lane #2342, and 4:2:2 / 4:4:4 import #2367 / #2368, and the WP3 Vulkan import lane #2375 (picked without its own merge of the SYCL and HIP lanes). It starts from rc4/api-wp6-compat (24e4104de) and merges the three other stacks, resolving each conflict per hunk; generated files took one side and were regenerated once per merge. It is a test vehicle and a base for WP9: it never goes into train.order, and each lane still lands through its own PR. The merged branch needed seven fixes that no single lane shows; they are separate commits after the merges, each with a docs/state.md row, so the lanes can take them back.

Merged heads: rc4/api-wp3-cuda 15bd091e2, rc4/api-wp4-windows 4037b18b4, rc4/api-motion-incremental 4947774ba, rc4/api-wp3-hip 20a3f4d4d, rc4/api-wp3-hip-gl-dmabuf 657e74a53, rc4/api-wp3-formats-422-444 8998921df, rc4/api-wp3-sycl 5dafa1cbd, rc4/api-wp3-formats-422-444-sycl c50cbc8e0, rc4/api-wp3-vulkan-import c101d5030 (its four commits after 02aa3a7d2 picked: 798c42c5d, 541a91a01, bae72c148, a21d20ce1).

ABI numbering (ADR-1897)

Several lanes bumped abi_version from 0.1.3 to 0.1.4 independently. The integration gives each addition set the next free patch, in merge order:

Patch Addition set Lane
0.1.4 option groups WP8
0.1.5 provenance and reports WP5
0.1.6 libvmaf bridge, tiny-AI, MCP WP6
0.1.7 VMAFX_MEMORY_GL_TEXTURE, VMAFX_FENCE_GL_SYNC, VmafxFrameImport.release / user WP3 CUDA (0.1.4 on its branch)
0.1.8 window scores, window clock, vmafx_context_max_in_flight WP4 (0.1.4 on its branch); incremental motion adds no entry
0.1.9 thirteen 4:2:2 / 4:4:4 VmafxPixelFormat entries (NV16 to YUV444P_MSB) formats lanes #2367 / #2368 (0.1.5 on their branches)
0.1.10 VMAFX_MEMORY_VULKAN, VMAFX_FENCE_VULKAN_SEMAPHORE, VmafxVulkanHandleType, VmafxVulkanTiling, VmafxVulkanFlags, VmafxDeviceInfo.pci, VmafxFrameImport.acquire_more / vulkan_*, vmafx_frame_signal_on_release() Vulkan lane #2375 (0.1.5 on its branch)

The HIP and SYCL lanes add no entry of their own; their texts on the GL texture memory, the device backend and the release callback name all three device lanes, with integration's 0.1.7.

Every entry keeps since = "0.1" (node VMAFX_0.1); only the "Added in ABI 0.1.x" notes, the changelog fragments, the AGENTS page and the rebase note of the two lanes changed. ADR-2023's body still says 0.1.4: it is Accepted and frozen. --abi-check (scripts/codegen/vmafx-api.py): append-only successor of origin/rc4/api-generation-prototype (378 additions), of origin/rc4/api-wp6-compat (33), origin/rc4/api-wp3-cuda (237), origin/rc4/api-wp4-windows (212) and origin/rc4/api-motion-incremental (212). No break, so no new minor.

Conflicts resolved

Merge of the WP3 CUDA lane (ec4ae3023):

  • core/api/vmafx.toml: abi_version 0.1.7 and the lane's four "Added in ABI" notes, as above.
  • core/src/vmafx/fence.c: the lane's vmafx_fence_poll() refactor kept, with WP5's exported vmafx_monotonic_ns() in place of the static monotonic_ns().
  • core/src/vmafx/internal.h: VmafxContext keeps both lane_state (CUDA lane) and provenance (WP5).
  • docs/development/rebase-sensitive-invariants.md: both new entries kept.
  • docs/state.md: scripts/dev/resolve-state-md-conflict.py.
  • Generated (bindings/python/vmafx/_api.py, core/include/vmafx/version.h, core/src/vmafx_symbols.txt, core/test/test_vmafx_abi_layout.c, docs/api/vmafx/reference.md, ADR indexes, scripts/ci/source-adr-citations.json, CHANGELOG.md): one side, regenerated.

Merge of WP4 (76bb126ab):

  • core/api/vmafx.toml: both lanes appended a section before the libvmaf_bridge.h functions; the provenance section (WP5) is followed by the window section (WP4). abi_version 0.1.8, WP4's twenty notes 0.1.8, WP4's vmafx_context_frame_retention() doc kept.
  • core/src/vmafx/fence.c: one timed wait. The CUDA lane's vmafx_fence_poll() now carries WP4's rule for timeouts of 2^62 ns and more (no clock check, UINT64_MAX rounds; the icx zero-round fix of 3a042aac4), and vmafx_host_fence_wait(), exported for window.c, wraps it.
  • core/src/vmafx/internal.h: VmafxContext keeps lane_state, provenance and windows; the 0.1.6 and 0.1.8 minimum-size macros.
  • core/src/vmafx/context.c: a failed create closes the windows and releases the provenance state (a failed vmafx_windows_init() releases the provenance state too); destroy keeps WP5's assert(context->engine != NULL) and WP4's vmafx_windows_pause().
  • core/src/vmafx/register.c: WP5's registered_name() with WP4's vmafx_engine_leave(context, previous). The same new signature (WP4's engine lock) at the call sites the other lanes added: core/src/cuda/import_device.c, context_frames.c, report.c, tiny_model.c.
  • core/src/libvmaf.c: WP4's three engine functions (thread count, in-flight bound, subsample) kept; WP2's libvmaf forwarders stay removed, since WP6 moved the libvmaf functions into the compat library.
  • core/test/meson.build: WP4's window tests before WP5's public test list.
  • docs/development/rebase-sensitive-invariants.md: both entries kept. docs/state.md: the resolver. Generated files: one side, regenerated.

Merge of incremental motion (d1d5ae2b1):

  • core/api/vmafx.toml: the lane's vmafx_window_submit() doc (a window over a VMAF model completes one frame after its last frame, ADR-2090) with the integration's 0.1.8 note. changelog.d/added/api-window-scores.md: the lane's sentence, ABI 0.1.8.
  • docs/state.md: T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06 takes the lane's Recently closed row, as WP4's answer to request WP4-1 says; the rest by the resolver.
  • Generated files: one side, regenerated.

Second round (2026-10-07): HIP, SYCL and 4:2:2 / 4:4:4

Merges a9ff02915 (HIP), ff42580ee (HIP GL), d9513b463 (formats), a082db7a9 (SYCL), a03b8c1bf (SYCL formats). Per hunk:

  • Device dispatch (device.c, device_context.c, frame_import.c, frame_import_admit.c, frame_pool.c, submit.c): both backends, each in its own #ifdef block.
  • fence.c: the HIP lane's dispatch by device and kind with SYCL_EVENT added; SYNC_FILE / GL_SYNC polled through vmafx_fence_poll().
  • vmafx_device_cells.h and its contract test: the HIP lane's cells with the SYCL lane's shared helpers.
  • meson: both lanes' import sources in libvmafx_sources (the SYCL lane used libvmaf_sources); window and sync-object sources kept.
  • docs/api/vmafx/index.md is hand-written: three-way merged per lane (3b9a41457; the first pass had taken one side as if generated, which MkDocs strict caught on three anchors).
  • Research numbers: both lanes added Research-2160; the HIP GL digest is Research-2161 here.
  • docs/adr/_index_fragments/_order.txt: one ADR-2133 entry (both formats lanes appended it).

One implementation (HISS-19): the HIP lane's gl_sync.c / sync_file.c duplicated the SYCL lane's sync_object.c; sync_object.c stays. The SYCL lane's own EGL export is folded onto core/src/vmafx/egl_export.c (44cbb8cd6; vmafx_egl_export_planes() takes a REFUSE / COPY / KEEP mode for tiled exports). test_vmafx_one_egl_export_contract refuses a second copy of either, planted copies included.

Contract reconciled (third round): vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor (ADR-2091 item 6), as the Vulkan lane resolves the HIP / SYCL difference; a GL sync is still refused. The second round had taken the HIP lane's refusal; the API index, core/src/AGENTS.d/vmafx-device-frames.md, the rebase note and the SYCL sync_file test follow the new contract.

Third round (2026-10-07): Vulkan import

The lane's first commit 02aa3a7d2 is its own merge of the SYCL and HIP lanes; both are here already, so it is dropped and the four commits after it are cherry-picked (-x). Per hunk:

  • core/api/vmafx.toml: abi_version 0.1.10; the lane's twelve "Added in ABI" notes say 0.1.10. --abi-check --against-ref the previous tip: append-only successor (18 additions).
  • core/src/cuda/import_frame.c: the Vulkan check first; a packed or MSB layout (4:2:2 / 4:4:4 lane) is refused for an OPTIMAL Vulkan image as for CUDA arrays, since OPTIMAL images are read out of arrays.
  • core/src/hip/import_frame.c: the GL dma-buf export (ADR-2132) and the lane's Vulkan-as-dma-buf path both kept; the lane's VmafxHipGl registration struct is not taken (its GL interop is gone here). The shared vmafx_parse_pci_bus_id() replaces vmafx_hip_parse_pci().
  • core/src/vmafx/internal.h: WP4's pooling and window declarations and the Vulkan declarations both kept.
  • docs/api/vmafx/index.md: the lane's "Vulkan frames" section after "SYCL devices"; the lane's copy of the older HIP section is not taken.
  • docs/state.md: the lane's rows and bucket entries, except T-HIP-ROCM10-GL-TEXTURE-READ-2026-10-06, which stays closed here (feat(hip): import GL textures through EGL dma-buf export so the pinned ROCm 10.1 reads them (ADR-2132) #2360).
  • docs/rebase-notes.md: the merge had duplicated the SYCL lane's section; one copy kept. The changelog fragment and the rebase note give ABI 0.1.10 here.
  • core/test/test_vmafx_import_vulkan_api.c: designated initialisers for the layout, which gained the packed fields (-Wmissing-field-initializers).
  • Generated files: one side, regenerated.

Fixes the merged branch needed

Commit Defect Found by State row
2eef29685 Three white-box tests the lanes added (test_model_collection_score_repeat, test_cuda_float_ms_ssim_host_traffic, test_motion_window_incremental) linked libvmaf.get_static_lib(), which the WP6 split no longer offers: meson setup failed configure none (build wiring)
784f57801 advance_extractors() (incremental motion) called advance() without installing the extractor as feature producer (WP5): integer_motion2 had source unknown in the provenance record test_vmafx_provenance test_features_in_name_order T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06
3b957b8ee The CUDA lane added its import sources to libvmaf_sources, renamed libvmafx_sources by WP6: a CUDA build failed to configure meson setup -Denable_cuda=true T-RC4-CUDA-IMPORT-SOURCES-OLD-LIBRARY-VARIABLE-2026-10-06
885912a34 A CUDA build left the caller's VmafPicture structs set after vmaf_read_pictures() while the compat library clears them: test_compat_conformance failed in every CUDA build build-cuda fast suite T-RC4-COMPAT-CONFORMANCE-CUDA-CONSUMED-2026-10-06
df671fc3d test_device_target_header_dependencies expected 21 CUDA targets (the CUDA lane added a 22nd) and the Windows fallback check did not parse cuda_dir entries build-cuda fast suite T-RC4-CUDA-IMPORT-CONVERT-TARGET-COUNT-2026-10-06
50eab63cf The provenance report tests scored on the device in a GPU build (no --backend) and expected the CPU record build-cuda fast suite T-RC4-PROVENANCE-TESTS-SCORE-ON-GPU-2026-10-06
22db201a9 The tiny-AI conformance scenario passed a 16-byte buffer for a 7x7 plane: stack-buffer-overflow under AddressSanitizer build-asan T-RC4-COMPAT-TINY-7X7-STACK-OVERFLOW-2026-10-06
fba26df2b vmafx_context_use_model() accepted a second model of a taken name, whose scores were the first one's (vmaf_v0.6.1neg reported 76.667831, alone 75.073658) the WP9 FFmpeg filter T-RC4-MODEL-NAME-COLLISION-2026-10-06
07d9fceae The carried C-locale fix (#2351) lacked its feature-name part: ..._0,7 names under a decimal-comma locale the WP9 GStreamer element T-OPTION-NUMBERS-CALLER-LOCALE-2026-10-06
44cbb8cd6 sycl/import_device.c and hip/import_device.c used the pre-WP4 vmafx_engine_leave(previous); a Linux SYCL link failed (undefined version: VMAF_LEGACY_SYCL) because vmaf_sycl_import_d3d11_surface() was defined on Windows only, now a refusing stub elsewhere first SYCL and HIP builds none (build)
adf6121c9 The private duplicate's failure path returned -errno; a zero errno would read as success. Falls back to -EBADF, as the master fix #2377 does review of #2377 none (follow-up of 869ec15f6)
869ec15f6 The SYCL dma-buf import closed a descriptor Level Zero had already closed (compute runtime 26.35): the release closed whatever had taken the number, after ten frames the FFmpeg filter's reference mapping the WP9 VAAPI-to-SYCL run T-SYCL-DMABUF-IMPORT-DOUBLE-CLOSE-2026-10-07

Open: T-VMAFX-IMPORT-CUDA-FENCE-CANARY-TIMING-2026-10-06. test_vmafx_import_cuda_fence test_release_canary counts a correctly ordered canary as early when other device work outlasts its fixed HOLD_MS hold (1 of 16 with --num-processes 2; 0 bad scores). Serial runs pass. Owner: the CUDA lane.

Type

  • build / ci — tooling / infra (integration branch; the fixes are fix / test / build commits)

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this; every commit went through the hooks).
  • make format && make lint is green locally — the commit and pre-push hooks pass.
  • Unit tests pass (third round, tip adf6121c9): CPU fast suite 391 OK, 2 expected fail, 0 fail, 1 skipped; CUDA on the RTX 4090: test_vmafx_import_cuda, _bitexact, _formats, _fence, _gl, test_vmafx_import_vulkan_cuda, _bitexact (21 s), _fence, test_vmafx_import_vulkan_api, test_vmafx_fence_kinds 10 of 10 OK; SYCL on the A380: test_vmafx_import_sycl 17, _fence 6, _bitexact 3 (236 cells, 0 differing), _formats 2, _gl 2, test_vmafx_import_vulkan_sycl 5, _bitexact 3 (472 cells, 0 differing), _fence 2, test_vmafx_fence_kinds 5, all OK; HIP in the pinned ROCm 10.1.0 image on the gfx1036: test_vmafx_import_hip 16, _fence 8, _bitexact 3, _formats 2, _gl 5, test_vmafx_fence_kinds 5, all OK. test_vmafx_import_vulkan_<lane>_ffmpeg skips (77) on CUDA and SYCL: it needs decoded inputs the run did not give. The HIP Vulkan tests are not built in the pinned image, which has no Vulkan development files (dependency('vulkan') not found); the lane measured them in the same image with the Vulkan loader and headers added (feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) #2375). Second round: CPU fast suite 390 OK, 0 fail, 1 skipped; SYCL build (icpx 2026.0, dg2-g11) on the Arc A380: test_vmafx_import_sycl 17, _fence 6, _bitexact 3, _formats 2, _gl 2, test_vmafx_fence_kinds 5, all OK; HIP build in the pinned ROCm 10.1.0 image on the gfx1036: test_vmafx_import_hip 16, _fence 8, _formats 2, _bitexact 3, _gl 5, test_hip_shared_frame 9, all OK; CUDA: test_vmafx_import_cuda 12, _gl 2, _formats 2 OK. First round: CPU build (build-cpu, gcc 16.2.1, -Db_lto=false): python3 scripts/ci/run_meson_test.py -- -C build-cpu --suite=fast --num-processes 6 385 OK, 2 expected fail (the planted conformance defects), 0 fail, 1 skipped (test_vmafx_api_abi_append_only: the merge base with origin/master has no definition yet). CUDA build (build-cuda, release, -Denable_cuda=true): fast suite without the gpu suite 394 OK, 2 expected fail, 1 skipped; gpu suite under the CUDA lock with --num-processes 1 68 of 68 OK; test_vmaf_cuda_gpumask, test_vmaf_cuda_threads, test_vmaf_feature_backend_cuda, test_cuda_parity_gate_default_run OK. Not run: test_cuda_exact_twin_matrix and test_cuda_v1_models_no_fallback (slow suite) each take more than the 300 s a device lock may be held.
  • Sanitizers on the VMAFx tests (24 tests: every test_vmafx_*, test_compat_conformance, test_motion_window_incremental; build-asan with address,undefined, build-tsan with thread): ASan/UBSan 24 of 24 OK after 22db201a9, with LeakSanitizer reporting only two 256-byte blocks allocated inside libonnxruntime.so with no fork frame (suppressed for the run with leak:libonnxruntime.so); TSan 24 of 24 OK.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. — not applicable: no kernel changed; the CUDA exact-twin and parity tests above pass.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. — not applicable: no extractor changed.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header — no new native file.
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. — not a breaking change: the definition is an append-only successor of every merged lane.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md — no ADR added: the merges follow ADR-1897; the fixes are defect fixes.

tidy: not measured on this branch; no file is new, and the fixed native sources (core/src/libvmaf.c, core/src/vmafx/{fence,context,register,context_frames,report,tiny_model}.c, core/src/cuda/import_device.c, core/test/test_compat_conformance_api.c) change by merge resolutions and one-line edits. Each lane measures its own files before it lands.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated — the rows in the fixes table (closed) and one open row, listed above.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below — not applicable.

make test-netflix-golden (core/build-golden, gcc): 280 passed, 3 skipped after the merges (784f57801), at 22db201a9 (after the engine change of 885912a34) and again at the tip adf6121c9 (after the Vulkan round).

Cross-backend numerical results

Not applicable: no kernel or extractor changed. Imported device frames score as host frames on every exact cell (the bit-exactness counts above, Vulkan imports included). CUDA twins: test_cuda_exact_twins and the *_parity tests of the gpu suite pass on the RTX 4090.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: integration of drafts that carry their own; the fixes are defect fixes.
  • Decision matrix — no alternatives: the ABI numbering follows ADR-1897; the conflict resolutions keep both sides.
  • AGENTS.md invariant note — no rebase-sensitive invariants beyond the lanes' own: core/src/AGENTS.d/vmafx-provenance.md already requires the producer around a direct extractor callback, which 784f57801 follows.
  • Reproducer / smoke-test command — under "Reproducer".
  • CHANGELOG fragment — no fragment: an integration branch; each lane's PR carries its own, and two of them say ABI 0.1.7 / 0.1.8 here.
  • Rebase note — no rebase impact: merge resolutions only; the lanes' notes stand (the CUDA lane's note names both ABI numbers).

Reproducer

git switch rc4/integration
meson setup build-cpu core -Db_lto=false && nice -n 10 ninja -C build-cpu -j6
python3 scripts/ci/run_meson_test.py -- -C build-cpu --suite=fast --num-processes 6
meson setup build-cuda core -Dbuildtype=release -Db_lto=false -Denable_cuda=true -Denable_nvcc=true
nice -n 10 ninja -C build-cuda -j6
flock ~/.cache/vmafx-locks/cuda-4090.lock timeout 300 \
  meson test -C build-cuda --suite gpu --no-suite slow --no-rebuild --num-processes 1
GOLDEN_NINJA_JOBS=4 make test-netflix-golden
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/rc4/api-generation-prototype

Known follow-ups

  • Each lane takes its fix back (or rebases onto this branch): CUDA lane 3b957b8ee (after WP6), df671fc3d; WP5 50eab63cf; WP6 885912a34, 22db201a9; incremental motion 784f57801 (after WP5).
  • Request WP5-1 (provenance pointer on VmafxWindowResult) is still open: WP5 and WP4 meet only here.
  • WP13 (feat(api): host input of every layout, RGB with a stated matrix, raw CLI layouts (RC4 WP13) #2376) is not merged: its RGB colour enums (VmafxColorMatrix, VmafxColorRange, VmafxColorTransfer in vmafx/types.h) collide with WP6's of the same names here (duplicate type: VmafxColorRange), and its --rgb_* flags are hand-written where WP8 generates the command line. Request WP13-6 asks the lane to converge on WP6's enums as ADR-2146 already decides; integration then gives its additions ABI 0.1.11.
  • The Vulkan lane's PCI parser and SYNC_FILE contract are taken as feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) #2375 has them; the HIP Vulkan tests need a Vulkan loader in the pinned image to build here.
  • The branch sits 92 commits behind origin/master, like every RC4 lane; it moves when they rebase.

xxxxxxxxxxxxx and others added 30 commits October 5, 2026 21:46
…e, ADR-1852)

core/api/vmafx.toml describes the new VMAFx C API and
scripts/codegen/vmafx-api.py writes every surface from it: the public
headers vmafx/vmafx.h and vmafx/libvmaf_bridge.h, the status tables, the
libvmaf compat shims, a standard-library Python binding, the ABI layout
test and the reference page. Generated files are committed; the Meson test
test_vmafx_api_generated_current fails when one differs from the
definition, and --abi-check refuses a non-append-only change (HISS-14).

The slice proves the approach end to end: vmafx_context_create/destroy,
version and ABI queries, provenance, extractor info and vmafx_feature_score,
with errors that name what failed. libvmaf's vmaf_init, vmaf_close,
vmaf_version and vmaf_feature_score_at_index are now generated shims on it;
their bodies moved to internal vmaf_engine_* functions in libvmaf.c, and the
shims return the engine's own errno, so libvmaf behaviour is unchanged.

Verified on a CPU build: the fast suite passes (342 tests), the golden-data
gate passes through the shims (280 passed, 3 skipped), and each new gate is
shown refusing a planted defect.
…4 WP1, ADR-1852)

The VMAFx API definition and its generator now carry everything the later
RC4 work packages need, so they can add API entries without touching the
generator.

Headers follow the layout of the design: vmafx/vmafx.h includes version.h,
types.h, error.h, context.h, device.h, frame.h, model.h, score.h,
provenance.h, report.h, dnn.h and mcp.h. Each entry names its header, each
header includes what its declarations use (computed; an include cycle stops
generation), and optional headers such as libvmaf_bridge.h stay out of the
umbrella.

The definition format gains callbacks with a trailing void *user, u32/u64
flag sets, fixed arrays, nested sized structs whose struct_size the _INIT
macros set, handle, pointer and size fields, since on every entry,
deprecated = {since, replacement, removal} (VMAFX_DEPRECATED on functions)
and option groups with per-surface spellings for the later surface
emitters. Layouts are computed per data model (LP64 / LLP64) without
recursion.

New outputs: the linker version script core/src/vmafx.map (one node per
ABI minor), which libvmaf now links on ELF targets with
--no-undefined-version; the Windows export list core/src/vmafx.def; the
symbol list check_exported_symbols now compares the vmafx_ exports and
their version nodes against; the header install list; one reference page
per header; and a changelog draft (--changelog <ref>).

New gates: test_vmafx_api_abi_append_only checks the definition against
the merge base with origin/master and skips with the reason when it
cannot compare, and test_vmafx_api_generator runs the generator's own
tests. Every gate is shown refusing a planted defect. A 0.x break needs a
higher ABI minor; additions may join the current minor's node until 1.0
freezes shipped nodes. The prototype definition (schema 1) still reads
for comparisons.
Within the 0.x preview an addition to the VMAFx API definition raises the
abi_version patch and may join the current minor's linker version node,
never an older one; only a break raises the minor, and --abi-check lists
it. From 1.0.0 a shipped node is frozen and an addition needs a newer
minor. The maintainer chose this over a new minor and node per addition,
which would make every RC4 work package that adds a function claim a minor
and renumber on each rebase.

The API generation guide and the vmafx-api agent note cite the record.
…852)

A program can now score videos through vmafx/*.h alone. The definition
(core/api/vmafx.toml, ABI 0.1.1 per ADR-1897) gains contexts with a log
callback and options, option sets, extractor / model / model-set
registration, imported scores, feature resolution (ADR-1359), frame
retention, refcounted models and model sets with the SHA-256 of the bytes
as loaded, the CPU device, host frames allocated or borrowed without a
copy, submission and flush, and synchronous per-frame and pooled scores
for features, models and model sets. Errors also name the kind of subject
and the function; an input struct below its introduction size is the new
VMAFX_E_ABI.

The implementation in core/src/vmafx/ calls the engine through
vmaf_engine_* entry points: the bodies of 14 more libvmaf functions are
renamed in core/src/libvmaf.c and the libvmaf names forward to them until
the compat layer generates them as shims. A context with a log callback
gets the engine messages of its calls through a per-thread sink in
core/src/log.cpp and never changes the process log level. A frame
reference is one count of the engine picture's counter, so one frame is
scored by several contexts without a copy (ADR-1880). ADR-1906 records
these rules.

Scores equal the libvmaf calls bit for bit: test_vmafx_bitexact compares
every feature and model score, per frame and pooled with every method, of
the golden pair, both checkerboards and the 10-bit sparks pair for
vmaf_v1.0.16_3d0h and vmaf_v0.6.1. New tests cover every function's
success, named failures and NULL error path, struct size negotiation,
logging, lifetimes under ASan and UBSan, and SHA-256; each was shown
failing on a planted defect.

Two generator fixes for WP1 (requests/WP1-2): Python records keep to_c()
when they gain pointer fields, and the generator tests no longer assume
the definition's ABI version or function set.

Found on the way: a model set scored per frame and then pooled over that
frame fails in libvmaf too (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05).
…ng-tidy findings

The cpu lane measured 33 findings in the files of the core API commit
(dev container, clang-tidy 22.1.8). The VMAFx log level and pixel format
reach the engine through explicit mappings instead of enum casts; pointer
offsets in the SHA-256 code are computed in size_t; the model file reader
cannot reach malloc(0) on any analysed path; the held-reference array is
cast from realloc explicitly.

The tests compare doubles through their integer bits instead of memcmp,
take the model and fixture directories from Meson at build time instead
of getenv, release their clip buffers on every path, and test functions
over the branch threshold are split into smaller ones. A source contract
test (test_gpu_float_ssim_auto_scale_contract) now reads the
vmaf_engine_* bodies the libvmaf entry points forward to.

vmafx_device_create() and the two registration functions assert the state
their checks established, so every fork-added function of 20 lines or
more carries an assert (assertion-density gate).
… (ADR-1906)

The maintainer chose full routing for RC4 (ADR-1906, now Accepted): a
context with a log callback receives every message the library raises
for it, worker threads included, and nothing of it reaches the process
log.

Each job the engine runs on a worker thread now captures the log sink of
the call that submitted it and installs it while it runs
(ThreadDataBatch.log_sink in core/src/libvmaf.c, vmaf_log_thread_sink()
in core/src/log.cpp); core/src/thread_pool.c is unchanged. The error
prints of the float ADM, SSIM, MS-SSIM, motion and VIF code went to
stdout and now go through vmaf_log(). A model belongs to no context:
VmafxModelConfig gains log_level, log_callback and log_user, and a load
routes its messages there. VmafxLogCallback moves to vmafx/types.h and
documents that it runs on library threads, possibly several at once.

The engine's init no longer sets the process log level;
vmafx_context_create() sets it for a context without a callback only, so
a context with a callback never writes it. A ThreadSanitizer run of the
new test found that the level itself is a plain global every init writes
while vmaf_log() reads it on all threads; that is a master defect and its
fix is master PR #2207 (T-LOG-LEVEL-GLOBAL-DATA-RACE-2026-10-06), which
this branch gets on its rebase.

test_vmafx_log_routing covers a message raised on a worker thread with
n_threads = 4, two contexts on two threads in a deterministic
interleaving and concurrently with worker threads, the process log of a
context without a callback, and a model load. Planted defects (no job
sink, a process-wide sink, a model load without its sink, a plain
context that leaves the process level) each fail it; with #2207's atomic
level applied TSan reports nothing in 8 runs. A source contract
(test_engine_log_routing_contract.py) refuses a direct stdout / stderr
write in core/src outside a declared exception table.
…e on the CPU device (RC4 WP3, ADR-1929)

The VMAFx API gains the shared contract of zero-copy frame import that the
CUDA, SYCL, HIP and Metal lanes implement next, working end to end on the
CPU device. The definition (ABI 0.1.2, additions only, node VMAFX_0.1)
declares every memory kind and every fence kind now, so the backend lanes
add implementations without an ABI change.

New: device enumeration and information (vmafx_device_count,
vmafx_device_info, vmafx_device_describe, vmafx_device_profile; the
size-prefixed VmafxDeviceInfo grows by the RC6 / RC7 format envelope
later), external handles and flags in VmafxDeviceDesc,
vmafx_context_use_device, VmafxFence with vmafx_fence_create / signal /
wait / destroy and the status VMAFX_E_TIMEOUT, vmafx_frame_import with an
acquire fence (NV12, P010 and P016 de-interleaved and shifted on the
device, nothing else, never a host copy), vmafx_frame_release_fence,
frame pools, vmafx_context_admit (each refusing extractor named, ADR-1688
generalised) and vmafx_context_import_frame, the D8 rule: one retry after
a host wait on the acquire fence, then one failure naming the input,
backend, device, memory kind, format, modifiers and the cause.

An imported frame follows the frame-reference rule of ADR-1906: its
release fence is signalled where its last reference is dropped, so one
import scored by two contexts is released after the last reader of
either. Test-only switches and counters (core/src/vmafx/frame_import_hooks.h)
count host copies, which the backend lanes assert stay 0, and plant the
defects the fence tests must catch. ADR-1929 records the choices: a poll
answers VMAFX_PENDING, VMAFX_IMPORT_ALLOW_COPY so a zeroed descriptor means
zero copy, an unsignalled acquire fence is VMAFX_E_BUSY on the CPU device.

Tests: imported NV12 / P010 / P016 frames score bit for bit as host frames
on the golden pair, both checkerboards, the 10-bit pair and synthetic 4K;
a producer thread's acquire fence orders the read; release-fence canaries
land only in released memory, with worker threads too; D8 retries once and
names the import; admission names each refusing extractor. Each planted
defect makes its test fail.

Found on the way: VmafxError kept 95 bytes of its subject, so a long model
path was named cut off (T-VMAFX-ERROR-SUBJECT-TRUNCATED-2026-10-06);
subjects and messages now keep 1023 bytes. The generator spells an out
string parameter const char ** (requests/WP1-3).
…pt ADR-1929

The maintainer decided ADR-1929 by popup on 2026-10-06: the zero-copy flag
stays VMAFX_IMPORT_ALLOW_COPY (a zeroed descriptor requires zero copy), a
timeout-0 poll answers VMAFX_PENDING and an unsignalled acquire fence on the
CPU device answers VMAFX_E_BUSY, and the host wait before the import rule's
one retry becomes a per-context option with a 10 s default. The ADR is now
Accepted, with the three answers in its references and the design's
VMAFX_IMPORT_REQUIRE_ZERO_COPY spelling under its alternatives.

VmafxContextConfig gains import_retry_wait_ns (ABI 0.1.3, an addition per
ADR-1897; --abi-check is append-only against the WP2 base and the first
push). 0, the initialiser's value and what a 0.1.1-sized config gets, is the
10 s default; 1 ns to 10 minutes is used as given; a larger value is
refused with VMAFX_E_RANGE naming config.import_retry_wait_ns rather than
clamped. vmafx_context_import_frame() waits at most that long.

A virtual test clock in core/src/vmafx/frame_import_hooks.h lets the tests
measure the bound without sleeping: the default holds exactly 10 s, a 20 ms
option fails after 20 ms, 1 ns after one poll step. The range edges are
tested through the public API. Planted defects (option ignored, no upper
bound, 0 not mapped to the default) each fail a test.

The generated Python record of VmafxContextConfig has no field defaults, so
test_vmafx_python_binding passes the new field; request WP1-4 asks the
generator for defaults. The device-frame notes move to their own AGENTS.d
page (core/src/AGENTS.d/vmafx-device-frames.md), the old one was over its
size budget.
… and test it

Five of eight panels queried series no binary exposes (jobs_queued,
jobs_in_flight, nodes_active, frames_per_second, gpu_utilization).
Rename the three that have a registered equivalent, drop the two with no
producer, and fail the build when a panel names an unregistered series.

Refs #1251
…unusable

/readyz only tested that a scorer object existed, so a removed or
non-executable vmaf binary kept the pod ready while every Score failed.
Add a vmaf-binary readiness check on the status registry and serve the
legacy /readyz JSON from that registry.

Refs #1251
…d document config precedence

A 1080p ScoreStream frame pair is 6.2 MB against the framework's 4 MiB
gRPC cap, and a synchronous POST /v1/score outlives the 60 s write
deadline. Apply 64 MiB and 15 min unless the operator sets the key, pin
the effective limits of the production graph, and add
docs/server/configuration.md with the env, file, default precedence and
the underscore rule, each pinned by a test.

Refs #1251
…logs

Server log lines named the same things differently or not at all. Define
request_id, rpc, route, model, backend, duration_s and error in
pkg/observability, derive request_id from the trace id when a span
exists, and log the set on the legacy POST /v1/score path.

Refs #1251
TestRunHTTPGracefulShutdown from issue #1251 is gone, but nothing kept the
pattern from returning. Scan every Go test under cmd, pkg and internal for
literal-port binds, with a matcher self-test and a planted case.

Refs #1251
…ps (RC4 WP8)

The option groups of core/api/vmafx.toml now hold every scoring option and
the generator emits each surface from them: the vmaf command-line table and
usage text (core/tools/cli_options.gen.inc), the MCP input schemas and the
argument-vector spec both MCP servers read (options.gen.json), the proto
messages (proto/vmafx_api.proto), the OpenAPI components, the vmafx filter's
AVOption table (ffmpeg-patches/src/vf_vmafx_options.h) and the option tables
of the user documentation. Both MCP servers serve the generated schemas and
build their vmaf argument vectors from the spec; the CLI keeps every spelling
it accepted. The CLI JSON report carries the provenance record.
…ring response (#2155)

ScoreRequest gains the generated ScoreOptions and every scoring response
(gRPC Score, POST /v1/score, the REST adapter, the ScoreStream aggregate)
carries a ScoreProvenance: the library record, the model the server loaded
with its SHA-256, the backend receipt and the precision. The three HTTP and
gRPC entry points share one scoring path; options become vmaf flags through
pkg/scoreopts, unknown request fields and bad option values are refused, and
scores are lossless unless the request asks otherwise.

A contract test (meson test test_vmafx_score_contract) runs the same requests
through the vmaf CLI, the C API and the server and requires the same score
bit for bit and the same library build on every surface.
…DR-2044)

ADR-2044 records the option-group emitters and the scoring contract. New
page docs/server/api-contract.md (requests, provenance, compatibility policy,
generated option table); the CLI, MCP, FFmpeg, gRPC, REST and API-generation
pages describe the generated tables, the provenance record and the changed
MCP defaults. State rows record the defects found on the way, agent notes and
the rebase note record the invariants.

The server also gives running gRPC calls a 30-minute grace when the framework
rotates a connection after two minutes, so a long Score or ScoreStream is no
longer cut (#1251).
The provenance member of the vmaf JSON report moves from the C++ receipt
helpers into core/tools/cli_provenance.c, so no C++ CLI source includes the
generated C headers of the VMAFx API (they trip modernize-use-using and
performance-enum-size in a C++ translation unit). test_cli_provenance covers
the member of a record and of a real libvmaf context; the CLI option-table
test uses designated initialisers. clang-tidy cpu lane: 0 findings in the
touched translation units.
… location and the OpenAPI splice

The maintainer accepted both deviations from the work-package brief by popup
(2026-10-06): the generated proto stays next to proto/vmafx.proto, and the
OpenAPI components are spliced into the server spec as well as written to
components.gen.yaml. ADR-2044 is Accepted; both answers are cited in its
References; the index rows are regenerated.
…ing a frame twice (#2206)

* fix(model): read a model collection's stored score instead of predicting a frame twice

vmaf_score_at_index_model_collection() predicted every member model and
wrote the members' and the four named bootstrap scores of the frame into
the feature collector, which refuses a second write of a frame. A second
per-frame call of a frame, or vmaf_score_pooled_model_collection() over a
range holding a frame already scored per frame (its loop predicts every
frame of the range again), therefore failed with -EINVAL ("feature ...
cannot be overwritten"). No score was wrong; the call failed.

A frame whose four named bootstrap scores are already in the collector
now returns them, as vmaf_score_at_index() reads a single model's stored
score first. The values are the first prediction's, bit for bit, so a
pooled score equals a fresh session's. Upstream Netflix/vmaf has the same
code.

test_model_collection_score_repeat fails on master without the fix: the
second per-frame call and the pooled call after a per-frame call both
return -EINVAL. Golden gate 280 passed, 3 skipped.
…nd-tripping every plane through the host

float_ms_ssim_cuda copied every scored plane of both input pictures to
pinned host memory, waited for the copy, converted it with picture_copy()
on the host and uploaded the floats again: per plane and frame two
plane-sized copies to the host, two uploads and two host waits, for
pictures that were already on the device. Level 0 of the pyramids is now
picture_copy() on the device (ms_ssim_picture_to_float: the same samples,
division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference
picture's stream after the distorted picture's ready event, and the
private stream waits behind it. The pinned staging buffers are gone.

test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts
every copy with a host side while they are scored. On the old code, 8-bit
4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18
plane copies with a host side; now 0 and 0, and what goes to the host is
exactly the per-window term planes the host adds in raster order. The
device-free contract test refuses the host staging. Scores are unchanged:
every output equals the CPU's bits, and the exact-twin matrix keeps 36 of
36 float_ms_ssim cells equal.

This is #2282 on master, carried by the RC4 WP3 CUDA lane until the stack
is restacked onto a master that has it; the restack drops this commit and
takes master's side of both files.
…s (RC4 WP3, ADR-2023)

The VMAFx API scores frames that already live on a CUDA device without a
copy through the host, and imported frames score bit for bit as the same
frames uploaded from the host for every CUDA twin declared exact.

- Devices: CUDA devices by index (retained primary context) or from the
  caller's context and stream (external[0], external[1]); count, info and
  describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and
  GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC.
  vmafx_context_use_device imports the device into the context's engine, and
  features registered afterwards run on their CUDA twins.
- Import: device pointers are bound where they are when each plane starts
  8-byte aligned with a pitch that is a multiple of 8 (the twins' vector
  loads); other layouts are refused naming the field, or copied on the
  device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on
  the device (import_convert.cu); CUDA arrays and GL textures are read out
  on the device. No import path copies through the host, and every host copy
  site calls vmafx_count_host_copy().
- Fences: a CUDA_EVENT acquire fence is waited on by the device's library
  stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL
  textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after
  the last reader in every context; the new VmafxFrameImport release callback
  (release, user; ABI 0.1.4) lets a producer make its stream wait on the
  release event without a host stall. The ADR-1199 barrier stays only for
  pictures that carry no fence ordering.
- Pools: CUDA frame pools hand out device frames; a returned frame is reused
  only after the device readers of its previous use finished.

integer_vif_cuda read both pictures with the pitch it computed at init, so
an imported plane with another pitch was read from the wrong rows; scale 0
now reads each picture with its own pitch. Master cannot hand the engine a
CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's
host round trip is fixed in the commit before this one (#2282 on master).

Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4):
test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both
checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0
differing, 6028 imports and 0 host copies. A planted skipped acquire wait
gives 15 bad frames of 16 under device load and the real wait 0; a planted
early release gives 15 bad canaries, the real release 0. One import scored
by two contexts equals each context's own run and is released only after
the second context finished. The exact-twin matrix keeps 48 of 48 vif and
float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3
skipped.
… row

The maintainer accepted ADR-2023 by popup on 2026-10-06 ("Accept, overlap as
tuning later (Recommended)"). One library stream per CUDA device stays the
rule; overlapping two frames' kernels across streams becomes an RC8 tuning
row, measured and held to the same fence tests (acquire under load, release
canary, pool frames without a fence). The popup is cited in the ADR's
References, and the index row and generated pages now show it Accepted.
RC4 work package 5 (#2142) leaves three choices open in ADR-1852: how the
provenance record is canonicalised and digested, whether CSV and SUB reports
get a sidecar, and how a report is verified. ADR-2073 records them before the
implementation lands:

- canonical JSON in the RFC 8785 form of the proto JSON mapping, one digest
  over the configuration and a digest of every score's bits, timing excluded,
  so the 1.1 signed assertion (#2159) signs one value;
- JSON and XML reports embed the record, CSV and SUB keep their bytes and get
  an opt-in sidecar;
- `vmaf --verify-provenance` re-runs the recorded command line and names the
  first differing field; the encode-record digest of VMAFx/pelorus#81 has a
  field now.

Status: Proposed, pending the maintainer's decision.
…t bound (RC4 WP4, ADR-2074)

A window asks for a model, a model set or a feature pooled with a set of
methods over [first, last] and returns at once; vmafx_window_poll(),
vmafx_window_wait() and an optional callback deliver the result once
every frame of the range is final, with the synchronous pooled call's
values bit for bit: the sync calls and the windows share
vmafx_pool_engine(). vmafx_flush() completes open windows over the
frames the stream had (VMAFX_WINDOW_PARTIAL), vmafx_context_destroy()
completes the rest with VMAFX_E_INVALID, vmafx_window_release() cancels.

Completion is found on the thread that feeds the context, in submit,
flush, import and window submit, with a per-window cursor and fence-free
probes (vmaf_engine_try_score_at_index(), vmaf_predict_inputs_written()),
so no engine lock is added and the producer never waits for work in
flight. Callbacks run on one window thread per context.

VmafxWindowClock holds the n_stats / n_stats_frames rule of #2138 for
every consumer. vmafx_context_max_in_flight() reports the engine queue's
bound, R + 2 * T * (R + 1), measured by a live-plugin harness (#2238):
textures imported per frame at 60 fps, windows polled from another
thread, within two frame periods, equal to the offline CLI, no host copy.
The former statement of vmafx_context_frame_retention() (n_threads more
frames) was wrong: 6 frames were measured with T = 2.

Windows over motion2 / motion3, and so over VMAF models, complete at the
flush: the integer motion extractors derive them there (state row
T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). ABI 0.1.3 -> 0.1.4, node
VMAFX_0.1, additions only.
…ify it (RC4 WP5, ADR-2073)

Every score now says how it was made (#2142). The library builds the record
when it is asked for and writes it into every report, so the CLI, API users,
the scoring server and both MCP servers carry the same one:

- VmafxProvenance grows at the end (ABI 0.1.5): commit, build id, compilers,
  build flags, the strict FP policy, backends, Rust extractors, SIMD level,
  device, options, frames, counts, the encode-record digest of
  VMAFx/pelorus#81, a digest of every score's bits, timing and the record
  digest. Models, features and annotations are size-prefixed structs read by
  index; the proto Provenance message carries them as repeated members
  (generator key proto_repeated, numbers from 100).
- The feature collector records the producer of every feature vector (the
  extractor instance and its options, a model, an import), so option-decorated
  names such as mse_y resolve and imported scores are marked imported.
- Exactness classes come from a table generated from exact_twins.d and the
  parity gate's tables (scripts/codegen/vmafx_exactness.py).
- The record's JSON is the proto JSON mapping in RFC 8785 form; digest covers
  everything but itself and the timing, so the 1.1 signed assertion (#2159)
  signs one value.
- vmafx_report_write() embeds the record in JSON and XML, leaves CSV and SUB
  unchanged and writes an opt-in sidecar; the CLI's JSON splice and its
  receipt formatter are gone, the library writes backend_used and
  feature_backends.
- vmafx_report_open/field/verify and `vmaf --verify-provenance <report>`
  re-run the recorded command line and name the first field that differs.
…nance record

The maintainer accepted the four choices of ADR-2073 by popup on 2026-10-06,
each as recommended: the record's JSON is the proto JSON mapping in RFC 8785
form; the digest covers the record and the scores but not the timing; CSV and
SUB reports get the record only in an opt-in sidecar; verification re-runs
the recorded configuration and compares. The answers are cited in the ADR's
References; the status and the index row move to Accepted.
…query can run during a submit

Design section 2.5 lets vmafx_context_provenance() run while another thread
submits frames, but the record read the frame count, the first frame's
geometry and the first-frame / flush times from engine fields the submitting
thread writes without synchronisation. A ThreadSanitizer build of the new
test_vmafx_provenance_threads (one thread submits, the other queries)
reported ten data races on them.

The engine now publishes what the record reads in VmafContext.run: the
geometry and first-frame time are stored before the frame count leaves 0
(release), the count is read with acquire, and the flush time is a release
store. pic_cnt and pic_params stay the submitting thread's own; only the
lines that already counted a frame for the record change, so RC4 WP4's
submit work rebases onto one call, run_note_frame(), after each pic_cnt++.
Under TSan the test now reports no race (five runs, with and without worker
threads); every build checks each record it sees is consistent.
…and accept ADR-2074

The maintainer chose a dedicated completion thread: windows complete
independently of the thread that feeds the context. Each context starts
one with its first window. It sleeps on a condition variable; the
engine's worker jobs (a new frame listener at the end of
threaded_extract_batch_func()), vmafx_submit(), vmafx_flush(),
vmafx_context_import_score() and vmafx_window_submit() raise a
generation counter and wake it, and each pass runs the cursor check of
the open windows. Callbacks run on a separate callback thread in
completion order, so a slow callback never holds up the completion of
another window.

The completion thread calls the engine, so every engine entry of the API
now takes the context's engine lock (vmafx_engine_enter() /
vmafx_engine_leave(context, previous)). Without it the completion thread
and a caller predict the same model frame twice; the new test fails then
and TSan reports the races.

vmafx_context_destroy() pauses the thread after it has looked at every
change already signalled, so a window whose frames were final before the
destroy completes with its values; a failed close resumes it and keeps
every window open (ADR-1336).

New tests: a window completes while the feeder stalls (16 of 16 frames
still on a worker at submit return, about 2 ms later), destroy with a
pending wake (64 rounds), a failed destroy keeps windows working, the
engine shared with the completion thread. Live harness re-measured:
about 2.3 ms for windows submitted ahead, 19 ms for windows the clock
cuts. ADR-2074 is Accepted with the maintainer's answers; the
feeding-thread design moves to its alternatives.
…ws complete before the flush (ADR-2090)

The integer motion extractors and their GPU twins derived motion2 and
motion3 of every frame in flush(), so a VMAFx window over a VMAF model,
a per-frame model score and vmaf_score_pooled() over the frames read so
far waited for the end of the stream (state row
T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF
windows in the 1.0 scope (Q-038, #2138, #2238).

Frame i is now final once the SAD scores of frames 0 to max(i + 1,
min_idx) are in the collector. vmaf_motion_window_advance() derives every
complete frame with upstream's per-frame statements in index order and
carries the stamp and the moving-average value in a VmafMotionWindowState;
vmaf_motion_window_flush() derives the rest. The engine calls a new
optional VmafFeatureExtractor.advance() after every accepted frame and
read fence, and the VMAFx completion thread calls it through
vmaf_engine_advance() under the engine lock before each pass, so a window
completes as soon as a worker's SAD is in, also while the feeder stalls.
motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL,
HIP, Metal) register it. A synchronous read of a fed frame without a
collector slot now fences instead of answering -EINVAL
(T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06).

Every value is the flush-time value: 32 988 per-frame motion values
against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K;
default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU
backend against the CPU, 0 different; parity gate motion cells exact on
CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped.
… icx build does not return at once

In a build with icx 2026.0.0 at -O3, vmafx_window_wait(UINT64_MAX)
returned VMAFX_E_PENDING at once (test_callbacks_run_once on the
incremental-motion lane's SYCL build, requests/WP4-1.md), and
vmafx_fence_wait(UINT64_MAX) on an unsignalled HOST fence returned
VMAFX_E_TIMEOUT. Both wait in vmafx_host_fence_wait(): a round count of
timeout / poll interval + 1, then a loop that leaves when the clock
passes the timeout. icx unrolls that loop and ran no round of it for
UINT64_MAX and UINT64_MAX - 1; 2^62 and smaller waited, gcc and clang
wait for every value, and -fno-unroll-loops makes icx wait too. The
arithmetic has no overflow.

A timeout of 2^62 ns (146 years) or more now waits without a limit: the
round bound is UINT64_MAX and the loop reads no clock. Two live tests
wait with UINT64_MAX, UINT64_MAX - 1, 2^63, 2^62 + 1, 2^62, 2^62 - 1 and
600 s, through vmafx_window_wait() for a score imported 20 ms later and
through vmafx_fence_wait() for a fence signalled 20 ms later. Each fails
without the fix in the icx build and passes with it in the icx and gcc
builds; a planted wrap of the round count fails each under gcc.

State row T-VMAFX-WAIT-FOREVER-ICX-ZERO-ROUNDS-2026-10-06.
…ersioned

check_exported_symbols.py took version visibility from the undefined imports
only. The Ubuntu 26.04 toolchain of the pinned ROCm 10.1.0 image links
__cxa_finalize unversioned, so the checker skipped the version-node
comparison there although nm prints the exports' nodes, and a planted wrong
node went unseen. Any versioned symbol, import or export, now counts.
State row T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06 (found on the
HIP lane) closed.
…egration branch

Conflicts resolved per hunk:

- core/test/check_exported_symbols.py: WP6's docstring for the split
  libraries, with WP1's sentence that any versioned symbol shows nm prints
  versions; parse_nm() merged cleanly.
- docs/state.md: T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06 takes
  WP1's Recently closed row (scripts/dev/resolve-state-md-conflict.py).
The acquire test's producer zeroed each frame buffer, waited on its
queue, then held the queue with a host task before the copy the imports
depend on. In 8 of about 110 runs on an Arc A380 (DPC++ 2026.0 and
2026.1.1) the process died in the runtime's host-task cleanup
(Scheduler::NotifyHostTaskCompletion -> event_impl::cleanupDependencyEvents)
while the main thread waited on the same queue; no VMAFx frame was on
the stacks, and a barrier instead of the join made no difference.

Every buffer of an arm is now allocated, zeroed and finished before the
first host task, so the host never waits on the producer while one of
its host tasks completes. The test ran 40 of 40 times clean with each
toolchain and still sees a skipped acquire wait (16 bad of 16).
Research-2160 finding 8 and T-SYCL-RUNTIME-HOST-TASK-CLEANUP-RACE-2026-10-06
record the runtime defect.
… yet

With incremental motion (ADR-2090) motion3 of frame i is written when frame
i + 1 is scored. vmafx_score_frame() of the frame just submitted returned
VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the
end of the motion3 vector's storage and the collector answered -EINVAL.
engine_score_at_index() now checks the inputs of a fed frame before it
predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps
its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.
…he caller's

Carried from the master fix (fix/numeric-options-c-locale) until the RC4
branches restack: a decimal-comma locale refused fractional feature
options, so the GStreamer element (gst-launch sets the user's locale)
could not score the default model vmaf_v1.0.16_3d0h. State row
T-OPTION-NUMBERS-CALLER-LOCALE-2026-10-06.

(cherry picked from commit 2fca94e0c)
…d ROCm 10.1 reads them (ADR-2132)

The runtime's GL interop maps a texture but cannot read it on ROCm 10.1, so
every GL import on HIP was refused. A GL texture is now exported as a dma-buf
through EGL and imported as a DMABUF frame. radeonsi exports its own tiling
(measured: 97 percent of samples differ when read as linear rows), so a tiled
export is copied on the GPU into a linear GBM dma-buf with
VMAFX_IMPORT_ALLOW_COPY, the producer's GL state restored. The context's GPU
is checked against the device's. Closes T-HIP-ROCM10-GL-TEXTURE-READ.
…tegration branch

The master fix (#2351) gained a third part after the first carry: a feature
named after a fractional option (`..._0.7`) was formatted with `%g` in the
caller's locale and came out as `..._0,7`, so a model never found its scores.
feature_name.cpp now formats the number in the C locale on the calling thread.
This takes #2351's final dict.cpp, feature_name.cpp, test_locale_handling.c
(MuTest table, new feature-name and model-feature cases) and the matching
changelog, AGENTS.d page, rebase note and state row.

Test: test_locale_handling 10/10; fast suite 385 ok, 0 failed.
…ed, bit-exact with host upload (RC4 WP3, ADR-2133)

vmafx_frame_import() takes NV16, P210, P216, NV24, P410, P416, the packed
Y210, Y212, YUYV422, Y410 (XV30), XV36 and VUYX, and NVDEC's MSB-aligned
planar 4:4:4 words (YUV444P_MSB), the layouts the decoders of the development
host were measured to emit. One CPU reference (vmafx_import_read_plane() and
the per-layout plan) serves the host import and the Metal reader; CUDA and HIP
convert on the device with the existing de-interleave kernels plus one gather
kernel that performs the same plan. ABI 0.1.5, additive.
…ready scores

A context keeps a model's scores under the model's name, and models are named
`vmaf` unless the configuration names them. vmafx_context_use_model()
accepted a second model of a taken name, whose scores were then the first
one's: the FFmpeg vmafx filter with vmaf_v0.6.1 and vmaf_v0.6.1neg logged
76.667831 for both, where vmaf_v0.6.1neg alone scores 75.073658. The vmaf CLI
refuses the same pair.

vmafx_context_use_model() now refuses a model whose name another model of the
context has, and vmafx_context_use_model_set() a set whose name another set
has, with VMAFX_E_INVALID naming it. A model and a set may share a name: a
set writes its scores under <name>_bagging and the other bootstrap suffixes.

Test: test_vmafx_context test_use_model_name_taken failed at "model of the
same name" before the change and passes; fast suite 385 passed, 0 failed.
tidy: cpu register.c test_vmafx_context.c 0 findings.
…ith host upload (RC4 WP3, ADR-2133)

The CPU reference, the pixel formats and the tests of ADR-2133, on the SYCL
lane: NV16 .. P416 use the existing de-interleave and shift kernels; the packed
layouts (Y210, Y212, YUYV422, Y410, XV36, VUYX) and MSB planar words use one
scratch-free gather kernel that performs the plan of the CPU reference. A
packed or MSB layout in a GL texture or an Intel-tiled dma-buf is refused
by name.
lusoris added 13 commits October 6, 2026 22:22
Conflicts resolved per hunk:
- core/api/vmafx.toml, core/src/AGENTS.d/vmafx-device-frames.md: the HIP
  lane's text with integration's ABI number (the release callback and the GL
  texture memory are ABI 0.1.7 here, 0.1.4 on the lane).
- core/test/test_device_target_header_dependencies.py: the lane's counts
  (CUDA 22, HIP 21) and its kernel-source roots (feature_src_dir, cuda_dir,
  src_dir).
- docs/development/rebase-sensitive-invariants.md: both entries kept.
- docs/state.md: the resolver; T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU stays
  closed (fixed on the WP1 branch, merged here).
- Generated files took integration's side and were regenerated
  (vmafx-api.py --write, make docs-fragments-write, the ADR citation registry,
  where the lane's rename of core/test/vmafx_cuda_cells.h to
  vmafx_device_cells.h is carried).

Fixes for the integration base: the HIP import sources join libvmafx_sources
(the WP6 split, ADR-2094), and hip/import_device.c leaves the engine with
vmafx_engine_leave(context, previous).

CPU fast suite 387 passed, 0 failed.
…to the integration branch

Conflicts resolved per hunk: core/api/vmafx.toml takes the lane's GL texture
text with integration's ABI number (0.1.7); docs/state.md through the
resolver; generated files (frame.h, the ADR and research indexes, the API
pages, the citation registry) took integration's side and were regenerated.
The shared EGL export (core/src/vmafx/egl_export.c, ADR-2132) joins
libvmafx_sources. CPU build clean.
…the integration branch

ABI: the lane's thirteen VmafxPixelFormat additions (NV16 to YUV444P_MSB,
values 19 to 31) are ABI 0.1.9 here. 0.1.5 is the provenance record on this
branch, and ADR-1897 numbers an addition set with the next free patch where
the lanes meet. vmafx-api.py --abi-check against the previous head: an
append-only successor, 13 additions. ADR-2133 (Accepted) keeps its lane
number; the changelog fragment and the rebase note name both.

Conflicts resolved per hunk: abi_version 0.1.9; the device-frames invariant
page takes the lane's text with integration's ABI number for the release
callback (0.1.7); generated files (version.h, types.h, the symbol list, the
ABI layout test, the Python bindings, the API and ADR indexes, the citation
registry) took integration's side and were regenerated. CPU build clean.
The HIP and SYCL lanes both added a backend next to CUDA in the same places.
Conflicts resolved per hunk (a recorded rerere resolution was set aside,
each conflict recreated, and the recorded result taken only where it keeps
both sides):
- device.c, device_context.c, frame_import.c, frame_import_admit.c,
  frame_pool.c, submit.c: both backends, each in its own #ifdef block.
- fence.c: the HIP lane's dispatch by device and kind, with SYCL_EVENT added
  to create, wait and destroy; SYNC_FILE and GL_SYNC fences are polled
  through vmafx_fence_poll() (the virtual test clock applies).
- cuda/import_fence.c: the HIP lane's shared release-event table.
- vmafx_device_cells.h and its contract test: the HIP lane's device cells
  with the SYCL lane's shared helpers (vmafx_exact_cells.h) and table
  parameter.
- core/api/vmafx.toml: GL texture, device backend and release callback texts
  name all three lanes; ABI numbers stay integration's (0.1.7).
- meson: both lanes' sources in libvmafx_sources (the WP6 split; the SYCL
  lane used libvmaf_sources), window and sync-object sources both kept.
- invariant pages and docs/state.md: both sides; generated files
  regenerated.

One implementation (HISS-19): the HIP lane's gl_sync.c and sync_file.c and
the SYCL lane's sync_object.c defined the same four functions. sync_object.c
(the superset: dma-buf implicit fences too) stays; the HIP files, their
declarations in internal.h and their citation-registry entries go, and the
HIP callers include sync_object.h.

Contract reconciled: the lanes disagreed on vmafx_fence_destroy() of a
SYNC_FILE fence (the HIP lane documents the descriptor as borrowed and
refuses; the SYCL lane closed it). The API documents the descriptor as
borrowed and the destroy as for fences the library returned, so the destroy
refuses; test_vmafx_import_sycl's sync_file case now expects VMAFX_E_NOTSUP
and closes the descriptor itself.

Research numbers: both lanes added Research-2160. The SYCL digest keeps
2160 (cited by code and the accepted ADR-2091); the HIP GL digest becomes
Research-2161, with its link in ADR-2132's references and the state row.

CPU fast suite 389 passed, 0 failed; CUDA build clean. SYCL and HIP builds
follow after the SYCL formats lane.
…egration branch

The SYCL half of ADR-2133 repeats the format table its HIP-based sibling
brought: abi_version and the thirteen VmafxPixelFormat entries keep
integration's numbers (0.1.9); the changelog fragment names the SYCL device
and both ABI numbers; the import test helper cites ADR-2091 and ADR-2092;
ADR-2133's references keep the link to ADR-2092; the formats state row takes
the SYCL lane's text (only Metal still refuses the layouts); the ADR order
list keeps one entry for ADR-2133 (both lanes appended it). Generated files
took integration's side and were regenerated (vmafx-api.py --abi-check
against the previous head: 0 additions). CPU build clean.
…port.c

The HIP lane's GL follow-up (ADR-2132) added core/src/vmafx/egl_export.c and
asked for the SYCL lane's own EGL export (core/src/sycl/import_gl.c) to be
folded onto it when the lanes met (HISS-19). The two differ in one point: the
HIP device reads linear rows, so a tiled export is refused or copied on the
GPU, while the SYCL device de-tiles Intel tilings itself and wants the export
as the driver made it. vmafx_egl_export_planes() takes a VmafxEglTiled mode
(REFUSE, COPY, KEEP) instead of allow_copy; a target fourcc of 0 and a NULL
device_pci skip those checks. sycl/import_gl.c keeps only the translation of
the exported planes into a DMABUF descriptor. In KEEP mode no writers wait is
made: the SYCL dma-buf import honours the implicit fences (D8 retry).

test_vmafx_one_egl_export_contract (fast suite) refuses a lane that loads EGL
itself and a second definition of a sync-object function; it fails on the
SYCL lane's former import_gl.c and on a planted sync_file.c, and passes now.
test_vmafx_import_hip_contract's planted anchor follows the new open_display()
call.

Fixes the integration base found by the first SYCL and HIP builds:
- sycl/import_device.c leaves the engine with vmafx_engine_leave(context,
  previous) (the WP4 signature); the API invariant page names it.
- sycl/d3d11_import.cpp defines vmaf_sycl_import_d3d11_surface() off Windows
  as a refusal (-ENOSYS): libvmaf_sycl.h declares it everywhere and the WP6
  version script lists it, so a Linux SYCL link failed with "undefined
  version: VMAF_LEGACY_SYCL".

Evidence: CPU fast suite 390 passed, 0 failed. SYCL build (icpx 2026.0,
dg2-g11) on the Arc A380: test_vmafx_import_sycl 16, _fence 6, _bitexact 3,
_formats 2, _gl 2, test_vmafx_fence_kinds 5, all passed. HIP build in the
pinned ROCm 10.1.0 image on the gfx1036: test_vmafx_import_hip 16, _fence 8,
_formats 2, _bitexact 3, _gl 5 (through the shared export),
test_hip_shared_frame 9, test_vmafx_fence_kinds 5, all passed.
…AFx API guide

docs/api/vmafx/index.md is written by hand, not generated, but the lane
merges took integration's side of it as if it were (a regeneration does not
rebuild it). The HIP devices and SYCL devices sections, the formats rows and
the fence and import overviews of the five lanes were lost, and three links
from the backend pages pointed at missing anchors (MkDocs strict).

The file is now the three-way merge of each lane against its base, applied in
order (HIP with its GL follow-up, the formats lane, SYCL, the SYCL formats
half), with the overview, import and fence paragraphs written for all three
device backends and the HIP and SYCL sections kept side by side.
…escriptor

zeMemAllocDevice() with a dma-buf import descriptor closes the descriptor it
is given on compute runtime 26.35 (Arc A380, measured with an LD_PRELOAD
close() log) and the specification leaves ownership open. The SYCL import
duplicated the caller's descriptor, handed the duplicate to the driver and
closed it again at the frame's release: the second close hit whatever
descriptor had taken the number in between. Running the FFmpeg vmafx filter
on VAAPI frames mapped to DRM PRIME, that was the reference frame's mapped
descriptor, and the eleventh import failed.

vmaf_sycl_dmabuf_import_queue() now gives the driver a private duplicate taken
from a high floor and closes it afterwards unless the driver did; the floor
makes that check safe against a descriptor opened meanwhile. The caller's
descriptor stays the caller's on every driver. The legacy VA surface import
shares the function and so the fix.

Test: test_vmafx_import_sycl test_import_closes_only_its_own opens a pipe
after an import (it takes the number such a driver freed) and checks that it,
and the caller's dma-buf descriptor, survive the frame's release. It failed on
the A380 before the change and passes; the SYCL import tests pass.
…152)

Brings the Vulkan import lane (#2375) onto the integration branch. The
lane's first commit (02aa3a7, its own merge of the SYCL and HIP lanes)
is not taken: both lanes are here already.

Resolved per hunk:

- core/api/vmafx.toml: the lane's twelve additions are ABI 0.1.10 here
  (0.1.5 on the lane, ADR-1897); generated files regenerated.
- cuda/import_frame.c: the Vulkan check runs first; a packed or MSB layout
  is refused for an OPTIMAL Vulkan image as for CUDA arrays.
- hip/import_frame.c: the GL dma-buf export (ADR-2132) and the Vulkan
  dma-buf path are both kept; the lane's VmafxHipGl registration struct is
  dropped, since the GL interop it served is gone here.
- vmafx/internal.h: pooling and window declarations and the Vulkan
  declarations are both kept.
- SYNC_FILE fences: vmafx_fence_destroy() closes the descriptor (ADR-2091
  item 6), as the lane resolves it; the API index, the device-frames
  agents page, the rebase note and the SYCL sync_file test follow.
- test_vmafx_import_vulkan_api: designated initialisers for the layout,
  which gained the packed fields.

(cherry picked from commit cff8500)
…3, ADR-2152)

The lane's API section "Vulkan frames" follows the SYCL devices section;
the lane's copy of the older HIP devices section is not taken (the HIP GL
dma-buf export of ADR-2132 is described here already). docs/state.md
takes the Vulkan rows and T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER; the lane's
T-HIP-ROCM10-GL-TEXTURE-READ row stays closed here. The shared PCI parser
replaces vmafx_hip_parse_pci(). The changelog fragment and rebase note
give ABI 0.1.10 for the integration branch; the duplicate SYCL rebase
section the merge produced is dropped.

(cherry picked from commit 3decc47)
…rser (RC4 WP3, ADR-2152)

vmafx_sycl_rt.cpp hands the bus id string of the Intel device-info
extension to C, which parses it with vmafx_parse_pci_bus_id() as the
CUDA and HIP lanes do (no sscanf, cert-err34). The CUDA interop asserts
its plane and import indices, and the PCI parser's refused forms are one
table-driven test (readability-function-size).

(cherry picked from commit 43b2231)
…lease paths (RC4 WP3, ADR-2152)

core/test/meson.build lists each Vulkan test program's source by name,
so the stale-source-reference check (ADR-1135) resolves them; the CUDA
Vulkan release and the SYCL device query carry the asserts the
Power-of-10 assertion-density gate requires of a function of 20 lines
or more.

(cherry picked from commit c101d50)
vmaf_sycl_dmabuf_import_queue() returned -errno when driver_fd() failed. fcntl() sets errno on failure, but a zero errno would turn the failure into a return of 0 with no pointer; the import now falls back to -EBADF, as the master fix (#2377) does.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants