Repository navigation
RC4 integration - #2343
Draft
lusoris wants to merge 77 commits into
Draft
RC4 integration#2343lusoris wants to merge 77 commits into
lusoris wants to merge 77 commits into
Conversation
…e, ADR-1852) core/api/vmafx.toml describes the new VMAFx C API and scripts/codegen/vmafx-api.py writes every surface from it: the public headers vmafx/vmafx.h and vmafx/libvmaf_bridge.h, the status tables, the libvmaf compat shims, a standard-library Python binding, the ABI layout test and the reference page. Generated files are committed; the Meson test test_vmafx_api_generated_current fails when one differs from the definition, and --abi-check refuses a non-append-only change (HISS-14). The slice proves the approach end to end: vmafx_context_create/destroy, version and ABI queries, provenance, extractor info and vmafx_feature_score, with errors that name what failed. libvmaf's vmaf_init, vmaf_close, vmaf_version and vmaf_feature_score_at_index are now generated shims on it; their bodies moved to internal vmaf_engine_* functions in libvmaf.c, and the shims return the engine's own errno, so libvmaf behaviour is unchanged. Verified on a CPU build: the fast suite passes (342 tests), the golden-data gate passes through the shims (280 passed, 3 skipped), and each new gate is shown refusing a planted defect.
…4 WP1, ADR-1852)
The VMAFx API definition and its generator now carry everything the later
RC4 work packages need, so they can add API entries without touching the
generator.
Headers follow the layout of the design: vmafx/vmafx.h includes version.h,
types.h, error.h, context.h, device.h, frame.h, model.h, score.h,
provenance.h, report.h, dnn.h and mcp.h. Each entry names its header, each
header includes what its declarations use (computed; an include cycle stops
generation), and optional headers such as libvmaf_bridge.h stay out of the
umbrella.
The definition format gains callbacks with a trailing void *user, u32/u64
flag sets, fixed arrays, nested sized structs whose struct_size the _INIT
macros set, handle, pointer and size fields, since on every entry,
deprecated = {since, replacement, removal} (VMAFX_DEPRECATED on functions)
and option groups with per-surface spellings for the later surface
emitters. Layouts are computed per data model (LP64 / LLP64) without
recursion.
New outputs: the linker version script core/src/vmafx.map (one node per
ABI minor), which libvmaf now links on ELF targets with
--no-undefined-version; the Windows export list core/src/vmafx.def; the
symbol list check_exported_symbols now compares the vmafx_ exports and
their version nodes against; the header install list; one reference page
per header; and a changelog draft (--changelog <ref>).
New gates: test_vmafx_api_abi_append_only checks the definition against
the merge base with origin/master and skips with the reason when it
cannot compare, and test_vmafx_api_generator runs the generator's own
tests. Every gate is shown refusing a planted defect. A 0.x break needs a
higher ABI minor; additions may join the current minor's node until 1.0
freezes shipped nodes. The prototype definition (schema 1) still reads
for comparisons.
Within the 0.x preview an addition to the VMAFx API definition raises the abi_version patch and may join the current minor's linker version node, never an older one; only a break raises the minor, and --abi-check lists it. From 1.0.0 a shipped node is frozen and an addition needs a newer minor. The maintainer chose this over a new minor and node per addition, which would make every RC4 work package that adds a function claim a minor and renumber on each rebase. The API generation guide and the vmafx-api agent note cite the record.
…852) A program can now score videos through vmafx/*.h alone. The definition (core/api/vmafx.toml, ABI 0.1.1 per ADR-1897) gains contexts with a log callback and options, option sets, extractor / model / model-set registration, imported scores, feature resolution (ADR-1359), frame retention, refcounted models and model sets with the SHA-256 of the bytes as loaded, the CPU device, host frames allocated or borrowed without a copy, submission and flush, and synchronous per-frame and pooled scores for features, models and model sets. Errors also name the kind of subject and the function; an input struct below its introduction size is the new VMAFX_E_ABI. The implementation in core/src/vmafx/ calls the engine through vmaf_engine_* entry points: the bodies of 14 more libvmaf functions are renamed in core/src/libvmaf.c and the libvmaf names forward to them until the compat layer generates them as shims. A context with a log callback gets the engine messages of its calls through a per-thread sink in core/src/log.cpp and never changes the process log level. A frame reference is one count of the engine picture's counter, so one frame is scored by several contexts without a copy (ADR-1880). ADR-1906 records these rules. Scores equal the libvmaf calls bit for bit: test_vmafx_bitexact compares every feature and model score, per frame and pooled with every method, of the golden pair, both checkerboards and the 10-bit sparks pair for vmaf_v1.0.16_3d0h and vmaf_v0.6.1. New tests cover every function's success, named failures and NULL error path, struct size negotiation, logging, lifetimes under ASan and UBSan, and SHA-256; each was shown failing on a planted defect. Two generator fixes for WP1 (requests/WP1-2): Python records keep to_c() when they gain pointer fields, and the generator tests no longer assume the definition's ABI version or function set. Found on the way: a model set scored per frame and then pooled over that frame fails in libvmaf too (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05).
…ng-tidy findings The cpu lane measured 33 findings in the files of the core API commit (dev container, clang-tidy 22.1.8). The VMAFx log level and pixel format reach the engine through explicit mappings instead of enum casts; pointer offsets in the SHA-256 code are computed in size_t; the model file reader cannot reach malloc(0) on any analysed path; the held-reference array is cast from realloc explicitly. The tests compare doubles through their integer bits instead of memcmp, take the model and fixture directories from Meson at build time instead of getenv, release their clip buffers on every path, and test functions over the branch threshold are split into smaller ones. A source contract test (test_gpu_float_ssim_auto_scale_contract) now reads the vmaf_engine_* bodies the libvmaf entry points forward to. vmafx_device_create() and the two registration functions assert the state their checks established, so every fork-added function of 20 lines or more carries an assert (assertion-density gate).
… (ADR-1906) The maintainer chose full routing for RC4 (ADR-1906, now Accepted): a context with a log callback receives every message the library raises for it, worker threads included, and nothing of it reaches the process log. Each job the engine runs on a worker thread now captures the log sink of the call that submitted it and installs it while it runs (ThreadDataBatch.log_sink in core/src/libvmaf.c, vmaf_log_thread_sink() in core/src/log.cpp); core/src/thread_pool.c is unchanged. The error prints of the float ADM, SSIM, MS-SSIM, motion and VIF code went to stdout and now go through vmaf_log(). A model belongs to no context: VmafxModelConfig gains log_level, log_callback and log_user, and a load routes its messages there. VmafxLogCallback moves to vmafx/types.h and documents that it runs on library threads, possibly several at once. The engine's init no longer sets the process log level; vmafx_context_create() sets it for a context without a callback only, so a context with a callback never writes it. A ThreadSanitizer run of the new test found that the level itself is a plain global every init writes while vmaf_log() reads it on all threads; that is a master defect and its fix is master PR #2207 (T-LOG-LEVEL-GLOBAL-DATA-RACE-2026-10-06), which this branch gets on its rebase. test_vmafx_log_routing covers a message raised on a worker thread with n_threads = 4, two contexts on two threads in a deterministic interleaving and concurrently with worker threads, the process log of a context without a callback, and a model load. Planted defects (no job sink, a process-wide sink, a model load without its sink, a plain context that leaves the process level) each fail it; with #2207's atomic level applied TSan reports nothing in 8 runs. A source contract (test_engine_log_routing_contract.py) refuses a direct stdout / stderr write in core/src outside a declared exception table.
…e on the CPU device (RC4 WP3, ADR-1929) The VMAFx API gains the shared contract of zero-copy frame import that the CUDA, SYCL, HIP and Metal lanes implement next, working end to end on the CPU device. The definition (ABI 0.1.2, additions only, node VMAFX_0.1) declares every memory kind and every fence kind now, so the backend lanes add implementations without an ABI change. New: device enumeration and information (vmafx_device_count, vmafx_device_info, vmafx_device_describe, vmafx_device_profile; the size-prefixed VmafxDeviceInfo grows by the RC6 / RC7 format envelope later), external handles and flags in VmafxDeviceDesc, vmafx_context_use_device, VmafxFence with vmafx_fence_create / signal / wait / destroy and the status VMAFX_E_TIMEOUT, vmafx_frame_import with an acquire fence (NV12, P010 and P016 de-interleaved and shifted on the device, nothing else, never a host copy), vmafx_frame_release_fence, frame pools, vmafx_context_admit (each refusing extractor named, ADR-1688 generalised) and vmafx_context_import_frame, the D8 rule: one retry after a host wait on the acquire fence, then one failure naming the input, backend, device, memory kind, format, modifiers and the cause. An imported frame follows the frame-reference rule of ADR-1906: its release fence is signalled where its last reference is dropped, so one import scored by two contexts is released after the last reader of either. Test-only switches and counters (core/src/vmafx/frame_import_hooks.h) count host copies, which the backend lanes assert stay 0, and plant the defects the fence tests must catch. ADR-1929 records the choices: a poll answers VMAFX_PENDING, VMAFX_IMPORT_ALLOW_COPY so a zeroed descriptor means zero copy, an unsignalled acquire fence is VMAFX_E_BUSY on the CPU device. Tests: imported NV12 / P010 / P016 frames score bit for bit as host frames on the golden pair, both checkerboards, the 10-bit pair and synthetic 4K; a producer thread's acquire fence orders the read; release-fence canaries land only in released memory, with worker threads too; D8 retries once and names the import; admission names each refusing extractor. Each planted defect makes its test fail. Found on the way: VmafxError kept 95 bytes of its subject, so a long model path was named cut off (T-VMAFX-ERROR-SUBJECT-TRUNCATED-2026-10-06); subjects and messages now keep 1023 bytes. The generator spells an out string parameter const char ** (requests/WP1-3).
…pt ADR-1929 The maintainer decided ADR-1929 by popup on 2026-10-06: the zero-copy flag stays VMAFX_IMPORT_ALLOW_COPY (a zeroed descriptor requires zero copy), a timeout-0 poll answers VMAFX_PENDING and an unsignalled acquire fence on the CPU device answers VMAFX_E_BUSY, and the host wait before the import rule's one retry becomes a per-context option with a 10 s default. The ADR is now Accepted, with the three answers in its references and the design's VMAFX_IMPORT_REQUIRE_ZERO_COPY spelling under its alternatives. VmafxContextConfig gains import_retry_wait_ns (ABI 0.1.3, an addition per ADR-1897; --abi-check is append-only against the WP2 base and the first push). 0, the initialiser's value and what a 0.1.1-sized config gets, is the 10 s default; 1 ns to 10 minutes is used as given; a larger value is refused with VMAFX_E_RANGE naming config.import_retry_wait_ns rather than clamped. vmafx_context_import_frame() waits at most that long. A virtual test clock in core/src/vmafx/frame_import_hooks.h lets the tests measure the bound without sleeping: the default holds exactly 10 s, a 20 ms option fails after 20 ms, 1 ns after one poll step. The range edges are tested through the public API. Planted defects (option ignored, no upper bound, 0 not mapped to the default) each fail a test. The generated Python record of VmafxContextConfig has no field defaults, so test_vmafx_python_binding passes the new field; request WP1-4 asks the generator for defaults. The device-frame notes move to their own AGENTS.d page (core/src/AGENTS.d/vmafx-device-frames.md), the old one was over its size budget.
… and test it Five of eight panels queried series no binary exposes (jobs_queued, jobs_in_flight, nodes_active, frames_per_second, gpu_utilization). Rename the three that have a registered equivalent, drop the two with no producer, and fail the build when a panel names an unregistered series. Refs #1251
…unusable /readyz only tested that a scorer object existed, so a removed or non-executable vmaf binary kept the pod ready while every Score failed. Add a vmaf-binary readiness check on the status registry and serve the legacy /readyz JSON from that registry. Refs #1251
…d document config precedence A 1080p ScoreStream frame pair is 6.2 MB against the framework's 4 MiB gRPC cap, and a synchronous POST /v1/score outlives the 60 s write deadline. Apply 64 MiB and 15 min unless the operator sets the key, pin the effective limits of the production graph, and add docs/server/configuration.md with the env, file, default precedence and the underscore rule, each pinned by a test. Refs #1251
…logs Server log lines named the same things differently or not at all. Define request_id, rpc, route, model, backend, duration_s and error in pkg/observability, derive request_id from the trace id when a span exists, and log the set on the legacy POST /v1/score path. Refs #1251
…ps (RC4 WP8) The option groups of core/api/vmafx.toml now hold every scoring option and the generator emits each surface from them: the vmaf command-line table and usage text (core/tools/cli_options.gen.inc), the MCP input schemas and the argument-vector spec both MCP servers read (options.gen.json), the proto messages (proto/vmafx_api.proto), the OpenAPI components, the vmafx filter's AVOption table (ffmpeg-patches/src/vf_vmafx_options.h) and the option tables of the user documentation. Both MCP servers serve the generated schemas and build their vmaf argument vectors from the spec; the CLI keeps every spelling it accepted. The CLI JSON report carries the provenance record.
…ring response (#2155) ScoreRequest gains the generated ScoreOptions and every scoring response (gRPC Score, POST /v1/score, the REST adapter, the ScoreStream aggregate) carries a ScoreProvenance: the library record, the model the server loaded with its SHA-256, the backend receipt and the precision. The three HTTP and gRPC entry points share one scoring path; options become vmaf flags through pkg/scoreopts, unknown request fields and bad option values are refused, and scores are lossless unless the request asks otherwise. A contract test (meson test test_vmafx_score_contract) runs the same requests through the vmaf CLI, the C API and the server and requires the same score bit for bit and the same library build on every surface.
…DR-2044) ADR-2044 records the option-group emitters and the scoring contract. New page docs/server/api-contract.md (requests, provenance, compatibility policy, generated option table); the CLI, MCP, FFmpeg, gRPC, REST and API-generation pages describe the generated tables, the provenance record and the changed MCP defaults. State rows record the defects found on the way, agent notes and the rebase note record the invariants. The server also gives running gRPC calls a 30-minute grace when the framework rotates a connection after two minutes, so a long Score or ScoreStream is no longer cut (#1251).
The provenance member of the vmaf JSON report moves from the C++ receipt helpers into core/tools/cli_provenance.c, so no C++ CLI source includes the generated C headers of the VMAFx API (they trip modernize-use-using and performance-enum-size in a C++ translation unit). test_cli_provenance covers the member of a record and of a real libvmaf context; the CLI option-table test uses designated initialisers. clang-tidy cpu lane: 0 findings in the touched translation units.
… location and the OpenAPI splice The maintainer accepted both deviations from the work-package brief by popup (2026-10-06): the generated proto stays next to proto/vmafx.proto, and the OpenAPI components are spliced into the server spec as well as written to components.gen.yaml. ADR-2044 is Accepted; both answers are cited in its References; the index rows are regenerated.
…ing a frame twice (#2206) * fix(model): read a model collection's stored score instead of predicting a frame twice vmaf_score_at_index_model_collection() predicted every member model and wrote the members' and the four named bootstrap scores of the frame into the feature collector, which refuses a second write of a frame. A second per-frame call of a frame, or vmaf_score_pooled_model_collection() over a range holding a frame already scored per frame (its loop predicts every frame of the range again), therefore failed with -EINVAL ("feature ... cannot be overwritten"). No score was wrong; the call failed. A frame whose four named bootstrap scores are already in the collector now returns them, as vmaf_score_at_index() reads a single model's stored score first. The values are the first prediction's, bit for bit, so a pooled score equals a fresh session's. Upstream Netflix/vmaf has the same code. test_model_collection_score_repeat fails on master without the fix: the second per-frame call and the pooled call after a per-frame call both return -EINVAL. Golden gate 280 passed, 3 skipped.
…nd-tripping every plane through the host float_ms_ssim_cuda copied every scored plane of both input pictures to pinned host memory, waited for the copy, converted it with picture_copy() on the host and uploaded the floats again: per plane and frame two plane-sized copies to the host, two uploads and two host waits, for pictures that were already on the device. Level 0 of the pyramids is now picture_copy() on the device (ms_ssim_picture_to_float: the same samples, division by 4, 16 or 256 at 10, 12 or 16 bits), on the reference picture's stream after the distorted picture's ready event, and the private stream waits behind it. The pinned staging buffers are gone. test_cuda_float_ms_ssim_host_traffic feeds device pictures and counts every copy with a host side while they are scored. On the old code, 8-bit 4:4:4 with enable_chroma over 3 frames uploaded 3538944 bytes and made 18 plane copies with a host side; now 0 and 0, and what goes to the host is exactly the per-window term planes the host adds in raster order. The device-free contract test refuses the host staging. Scores are unchanged: every output equals the CPU's bits, and the exact-twin matrix keeps 36 of 36 float_ms_ssim cells equal. This is #2282 on master, carried by the RC4 WP3 CUDA lane until the stack is restacked onto a master that has it; the restack drops this commit and takes master's side of both files.
…s (RC4 WP3, ADR-2023) The VMAFx API scores frames that already live on a CUDA device without a copy through the host, and imported frames score bit for bit as the same frames uploaded from the host for every CUDA twin declared exact. - Devices: CUDA devices by index (retained primary context) or from the caller's context and stream (external[0], external[1]); count, info and describe report the memory kinds DEVICE_POINTER, DEVICE_ARRAY and GL_TEXTURE and the fence kinds NONE, HOST, CUDA_EVENT and GL_SYNC. vmafx_context_use_device imports the device into the context's engine, and features registered afterwards run on their CUDA twins. - Import: device pointers are bound where they are when each plane starts 8-byte aligned with a pitch that is a multiple of 8 (the twins' vector loads); other layouts are refused naming the field, or copied on the device with VMAFX_IMPORT_ALLOW_COPY. NV12, P010 and P016 are planarised on the device (import_convert.cu); CUDA arrays and GL textures are read out on the device. No import path copies through the host, and every host copy site calls vmafx_count_host_copy(). - Fences: a CUDA_EVENT acquire fence is waited on by the device's library stream; GL_SYNC (new, ABI 0.1.4) is waited on the host before the GL textures are mapped. Release fences (HOST, CUDA_EVENT) are signalled after the last reader in every context; the new VmafxFrameImport release callback (release, user; ABI 0.1.4) lets a producer make its stream wait on the release event without a host stall. The ADR-1199 barrier stays only for pictures that carry no fence ordering. - Pools: CUDA frame pools hand out device frames; a returned frame is reused only after the device readers of its previous use finished. integer_vif_cuda read both pictures with the pitch it computed at init, so an imported plane with another pitch was read from the wrong rows; scale 0 now reads each picture with its own pitch. Master cannot hand the engine a CUDA picture of another pitch, so this fix stays here. float_ms_ssim_cuda's host round trip is fixed in the commit before this one (#2282 on master). Evidence on an RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4): test_vmafx_import_cuda_bitexact compares 236 cells (576x324 pair, both checkerboards, 4K bbb; planar and NV12 / P010) with 10732 values, 0 differing, 6028 imports and 0 host copies. A planted skipped acquire wait gives 15 bad frames of 16 under device load and the real wait 0; a planted early release gives 15 bad canaries, the real release 0. One import scored by two contexts equals each context's own run and is released only after the second context finished. The exact-twin matrix keeps 48 of 48 vif and float_ms_ssim cells equal to the CPU. Netflix golden gate: 280 passed, 3 skipped.
… row
The maintainer accepted ADR-2023 by popup on 2026-10-06 ("Accept, overlap as
tuning later (Recommended)"). One library stream per CUDA device stays the
rule; overlapping two frames' kernels across streams becomes an RC8 tuning
row, measured and held to the same fence tests (acquire under load, release
canary, pool frames without a fence). The popup is cited in the ADR's
References, and the index row and generated pages now show it Accepted.
RC4 work package 5 (#2142) leaves three choices open in ADR-1852: how the provenance record is canonicalised and digested, whether CSV and SUB reports get a sidecar, and how a report is verified. ADR-2073 records them before the implementation lands: - canonical JSON in the RFC 8785 form of the proto JSON mapping, one digest over the configuration and a digest of every score's bits, timing excluded, so the 1.1 signed assertion (#2159) signs one value; - JSON and XML reports embed the record, CSV and SUB keep their bytes and get an opt-in sidecar; - `vmaf --verify-provenance` re-runs the recorded command line and names the first differing field; the encode-record digest of VMAFx/pelorus#81 has a field now. Status: Proposed, pending the maintainer's decision.
…t bound (RC4 WP4, ADR-2074) A window asks for a model, a model set or a feature pooled with a set of methods over [first, last] and returns at once; vmafx_window_poll(), vmafx_window_wait() and an optional callback deliver the result once every frame of the range is final, with the synchronous pooled call's values bit for bit: the sync calls and the windows share vmafx_pool_engine(). vmafx_flush() completes open windows over the frames the stream had (VMAFX_WINDOW_PARTIAL), vmafx_context_destroy() completes the rest with VMAFX_E_INVALID, vmafx_window_release() cancels. Completion is found on the thread that feeds the context, in submit, flush, import and window submit, with a per-window cursor and fence-free probes (vmaf_engine_try_score_at_index(), vmaf_predict_inputs_written()), so no engine lock is added and the producer never waits for work in flight. Callbacks run on one window thread per context. VmafxWindowClock holds the n_stats / n_stats_frames rule of #2138 for every consumer. vmafx_context_max_in_flight() reports the engine queue's bound, R + 2 * T * (R + 1), measured by a live-plugin harness (#2238): textures imported per frame at 60 fps, windows polled from another thread, within two frame periods, equal to the offline CLI, no host copy. The former statement of vmafx_context_frame_retention() (n_threads more frames) was wrong: 6 frames were measured with T = 2. Windows over motion2 / motion3, and so over VMAF models, complete at the flush: the integer motion extractors derive them there (state row T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). ABI 0.1.3 -> 0.1.4, node VMAFX_0.1, additions only.
…ify it (RC4 WP5, ADR-2073) Every score now says how it was made (#2142). The library builds the record when it is asked for and writes it into every report, so the CLI, API users, the scoring server and both MCP servers carry the same one: - VmafxProvenance grows at the end (ABI 0.1.5): commit, build id, compilers, build flags, the strict FP policy, backends, Rust extractors, SIMD level, device, options, frames, counts, the encode-record digest of VMAFx/pelorus#81, a digest of every score's bits, timing and the record digest. Models, features and annotations are size-prefixed structs read by index; the proto Provenance message carries them as repeated members (generator key proto_repeated, numbers from 100). - The feature collector records the producer of every feature vector (the extractor instance and its options, a model, an import), so option-decorated names such as mse_y resolve and imported scores are marked imported. - Exactness classes come from a table generated from exact_twins.d and the parity gate's tables (scripts/codegen/vmafx_exactness.py). - The record's JSON is the proto JSON mapping in RFC 8785 form; digest covers everything but itself and the timing, so the 1.1 signed assertion (#2159) signs one value. - vmafx_report_write() embeds the record in JSON and XML, leaves CSV and SUB unchanged and writes an opt-in sidecar; the CLI's JSON splice and its receipt formatter are gone, the library writes backend_used and feature_backends. - vmafx_report_open/field/verify and `vmaf --verify-provenance <report>` re-run the recorded command line and name the first field that differs.
…nance record The maintainer accepted the four choices of ADR-2073 by popup on 2026-10-06, each as recommended: the record's JSON is the proto JSON mapping in RFC 8785 form; the digest covers the record and the scores but not the timing; CSV and SUB reports get the record only in an opt-in sidecar; verification re-runs the recorded configuration and compares. The answers are cited in the ADR's References; the status and the index row move to Accepted.
…query can run during a submit Design section 2.5 lets vmafx_context_provenance() run while another thread submits frames, but the record read the frame count, the first frame's geometry and the first-frame / flush times from engine fields the submitting thread writes without synchronisation. A ThreadSanitizer build of the new test_vmafx_provenance_threads (one thread submits, the other queries) reported ten data races on them. The engine now publishes what the record reads in VmafContext.run: the geometry and first-frame time are stored before the frame count leaves 0 (release), the count is read with acquire, and the flush time is a release store. pic_cnt and pic_params stay the submitting thread's own; only the lines that already counted a frame for the record change, so RC4 WP4's submit work rebases onto one call, run_note_frame(), after each pic_cnt++. Under TSan the test now reports no race (five runs, with and without worker threads); every build checks each record it sees is consistent.
…and accept ADR-2074 The maintainer chose a dedicated completion thread: windows complete independently of the thread that feeds the context. Each context starts one with its first window. It sleeps on a condition variable; the engine's worker jobs (a new frame listener at the end of threaded_extract_batch_func()), vmafx_submit(), vmafx_flush(), vmafx_context_import_score() and vmafx_window_submit() raise a generation counter and wake it, and each pass runs the cursor check of the open windows. Callbacks run on a separate callback thread in completion order, so a slow callback never holds up the completion of another window. The completion thread calls the engine, so every engine entry of the API now takes the context's engine lock (vmafx_engine_enter() / vmafx_engine_leave(context, previous)). Without it the completion thread and a caller predict the same model frame twice; the new test fails then and TSan reports the races. vmafx_context_destroy() pauses the thread after it has looked at every change already signalled, so a window whose frames were final before the destroy completes with its values; a failed close resumes it and keeps every window open (ADR-1336). New tests: a window completes while the feeder stalls (16 of 16 frames still on a worker at submit return, about 2 ms later), destroy with a pending wake (64 rounds), a failed destroy keeps windows working, the engine shared with the completion thread. Live harness re-measured: about 2.3 ms for windows submitted ahead, 19 ms for windows the clock cuts. ADR-2074 is Accepted with the maintainer's answers; the feeding-thread design moves to its alternatives.
…ws complete before the flush (ADR-2090) The integer motion extractors and their GPU twins derived motion2 and motion3 of every frame in flush(), so a VMAFx window over a VMAF model, a per-frame model score and vmaf_score_pooled() over the frames read so far waited for the end of the stream (state row T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF windows in the 1.0 scope (Q-038, #2138, #2238). Frame i is now final once the SAD scores of frames 0 to max(i + 1, min_idx) are in the collector. vmaf_motion_window_advance() derives every complete frame with upstream's per-frame statements in index order and carries the stamp and the moving-average value in a VmafMotionWindowState; vmaf_motion_window_flush() derives the rest. The engine calls a new optional VmafFeatureExtractor.advance() after every accepted frame and read fence, and the VMAFx completion thread calls it through vmaf_engine_advance() under the engine lock before each pass, so a window completes as soon as a worker's SAD is in, also while the feeder stalls. motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL, HIP, Metal) register it. A synchronous read of a fed frame without a collector slot now fences instead of answering -EINVAL (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06). Every value is the flush-time value: 32 988 per-frame motion values against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K; default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU backend against the CPU, 0 different; parity gate motion cells exact on CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped.
… icx build does not return at once In a build with icx 2026.0.0 at -O3, vmafx_window_wait(UINT64_MAX) returned VMAFX_E_PENDING at once (test_callbacks_run_once on the incremental-motion lane's SYCL build, requests/WP4-1.md), and vmafx_fence_wait(UINT64_MAX) on an unsignalled HOST fence returned VMAFX_E_TIMEOUT. Both wait in vmafx_host_fence_wait(): a round count of timeout / poll interval + 1, then a loop that leaves when the clock passes the timeout. icx unrolls that loop and ran no round of it for UINT64_MAX and UINT64_MAX - 1; 2^62 and smaller waited, gcc and clang wait for every value, and -fno-unroll-loops makes icx wait too. The arithmetic has no overflow. A timeout of 2^62 ns (146 years) or more now waits without a limit: the round bound is UINT64_MAX and the loop reads no clock. Two live tests wait with UINT64_MAX, UINT64_MAX - 1, 2^63, 2^62 + 1, 2^62, 2^62 - 1 and 600 s, through vmafx_window_wait() for a score imported 20 ms later and through vmafx_fence_wait() for a fence signalled 20 ms later. Each fails without the fix in the icx build and passes with it in the icx and gcc builds; a planted wrap of the round count fails each under gcc. State row T-VMAFX-WAIT-FOREVER-ICX-ZERO-ROUNDS-2026-10-06.
…ersioned check_exported_symbols.py took version visibility from the undefined imports only. The Ubuntu 26.04 toolchain of the pinned ROCm 10.1.0 image links __cxa_finalize unversioned, so the checker skipped the version-node comparison there although nm prints the exports' nodes, and a planted wrong node went unseen. Any versioned symbol, import or export, now counts. State row T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06 (found on the HIP lane) closed.
…egration branch Conflicts resolved per hunk: - core/test/check_exported_symbols.py: WP6's docstring for the split libraries, with WP1's sentence that any versioned symbol shows nm prints versions; parse_nm() merged cleanly. - docs/state.md: T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU-2026-10-06 takes WP1's Recently closed row (scripts/dev/resolve-state-md-conflict.py).
The acquire test's producer zeroed each frame buffer, waited on its queue, then held the queue with a host task before the copy the imports depend on. In 8 of about 110 runs on an Arc A380 (DPC++ 2026.0 and 2026.1.1) the process died in the runtime's host-task cleanup (Scheduler::NotifyHostTaskCompletion -> event_impl::cleanupDependencyEvents) while the main thread waited on the same queue; no VMAFx frame was on the stacks, and a barrier instead of the join made no difference. Every buffer of an arm is now allocated, zeroed and finished before the first host task, so the host never waits on the producer while one of its host tasks completes. The test ran 40 of 40 times clean with each toolchain and still sees a skipped acquire wait (16 bad of 16). Research-2160 finding 8 and T-SYCL-RUNTIME-HOST-TASK-CLEANUP-RACE-2026-10-06 record the runtime defect.
… yet With incremental motion (ADR-2090) motion3 of frame i is written when frame i + 1 is scored. vmafx_score_frame() of the frame just submitted returned VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the end of the motion3 vector's storage and the collector answered -EINVAL. engine_score_at_index() now checks the inputs of a fed frame before it predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.
…he caller's Carried from the master fix (fix/numeric-options-c-locale) until the RC4 branches restack: a decimal-comma locale refused fractional feature options, so the GStreamer element (gst-launch sets the user's locale) could not score the default model vmaf_v1.0.16_3d0h. State row T-OPTION-NUMBERS-CALLER-LOCALE-2026-10-06. (cherry picked from commit 2fca94e0c)
11 of 18 tasks
…d ROCm 10.1 reads them (ADR-2132) The runtime's GL interop maps a texture but cannot read it on ROCm 10.1, so every GL import on HIP was refused. A GL texture is now exported as a dma-buf through EGL and imported as a DMABUF frame. radeonsi exports its own tiling (measured: 97 percent of samples differ when read as linear rows), so a tiled export is copied on the GPU into a linear GBM dma-buf with VMAFX_IMPORT_ALLOW_COPY, the producer's GL state restored. The context's GPU is checked against the device's. Closes T-HIP-ROCM10-GL-TEXTURE-READ.
…tegration branch The master fix (#2351) gained a third part after the first carry: a feature named after a fractional option (`..._0.7`) was formatted with `%g` in the caller's locale and came out as `..._0,7`, so a model never found its scores. feature_name.cpp now formats the number in the C locale on the calling thread. This takes #2351's final dict.cpp, feature_name.cpp, test_locale_handling.c (MuTest table, new feature-name and model-feature cases) and the matching changelog, AGENTS.d page, rebase note and state row. Test: test_locale_handling 10/10; fast suite 385 ok, 0 failed.
…ed, bit-exact with host upload (RC4 WP3, ADR-2133) vmafx_frame_import() takes NV16, P210, P216, NV24, P410, P416, the packed Y210, Y212, YUYV422, Y410 (XV30), XV36 and VUYX, and NVDEC's MSB-aligned planar 4:4:4 words (YUV444P_MSB), the layouts the decoders of the development host were measured to emit. One CPU reference (vmafx_import_read_plane() and the per-layout plan) serves the host import and the Metal reader; CUDA and HIP convert on the device with the existing de-interleave kernels plus one gather kernel that performs the same plan. ABI 0.1.5, additive.
…ready scores A context keeps a model's scores under the model's name, and models are named `vmaf` unless the configuration names them. vmafx_context_use_model() accepted a second model of a taken name, whose scores were then the first one's: the FFmpeg vmafx filter with vmaf_v0.6.1 and vmaf_v0.6.1neg logged 76.667831 for both, where vmaf_v0.6.1neg alone scores 75.073658. The vmaf CLI refuses the same pair. vmafx_context_use_model() now refuses a model whose name another model of the context has, and vmafx_context_use_model_set() a set whose name another set has, with VMAFX_E_INVALID naming it. A model and a set may share a name: a set writes its scores under <name>_bagging and the other bootstrap suffixes. Test: test_vmafx_context test_use_model_name_taken failed at "model of the same name" before the change and passes; fast suite 385 passed, 0 failed. tidy: cpu register.c test_vmafx_context.c 0 findings.
…ith host upload (RC4 WP3, ADR-2133) The CPU reference, the pixel formats and the tests of ADR-2133, on the SYCL lane: NV16 .. P416 use the existing de-interleave and shift kernels; the packed layouts (Y210, Y212, YUYV422, Y410, XV36, VUYX) and MSB planar words use one scratch-free gather kernel that performs the plan of the CPU reference. A packed or MSB layout in a GL texture or an Intel-tiled dma-buf is refused by name.
6 tasks done
Conflicts resolved per hunk: - core/api/vmafx.toml, core/src/AGENTS.d/vmafx-device-frames.md: the HIP lane's text with integration's ABI number (the release callback and the GL texture memory are ABI 0.1.7 here, 0.1.4 on the lane). - core/test/test_device_target_header_dependencies.py: the lane's counts (CUDA 22, HIP 21) and its kernel-source roots (feature_src_dir, cuda_dir, src_dir). - docs/development/rebase-sensitive-invariants.md: both entries kept. - docs/state.md: the resolver; T-VMAFX-SYMBOL-VERSIONS-UNSEEN-UBUNTU stays closed (fixed on the WP1 branch, merged here). - Generated files took integration's side and were regenerated (vmafx-api.py --write, make docs-fragments-write, the ADR citation registry, where the lane's rename of core/test/vmafx_cuda_cells.h to vmafx_device_cells.h is carried). Fixes for the integration base: the HIP import sources join libvmafx_sources (the WP6 split, ADR-2094), and hip/import_device.c leaves the engine with vmafx_engine_leave(context, previous). CPU fast suite 387 passed, 0 failed.
…to the integration branch Conflicts resolved per hunk: core/api/vmafx.toml takes the lane's GL texture text with integration's ABI number (0.1.7); docs/state.md through the resolver; generated files (frame.h, the ADR and research indexes, the API pages, the citation registry) took integration's side and were regenerated. The shared EGL export (core/src/vmafx/egl_export.c, ADR-2132) joins libvmafx_sources. CPU build clean.
…the integration branch ABI: the lane's thirteen VmafxPixelFormat additions (NV16 to YUV444P_MSB, values 19 to 31) are ABI 0.1.9 here. 0.1.5 is the provenance record on this branch, and ADR-1897 numbers an addition set with the next free patch where the lanes meet. vmafx-api.py --abi-check against the previous head: an append-only successor, 13 additions. ADR-2133 (Accepted) keeps its lane number; the changelog fragment and the rebase note name both. Conflicts resolved per hunk: abi_version 0.1.9; the device-frames invariant page takes the lane's text with integration's ABI number for the release callback (0.1.7); generated files (version.h, types.h, the symbol list, the ABI layout test, the Python bindings, the API and ADR indexes, the citation registry) took integration's side and were regenerated. CPU build clean.
The HIP and SYCL lanes both added a backend next to CUDA in the same places. Conflicts resolved per hunk (a recorded rerere resolution was set aside, each conflict recreated, and the recorded result taken only where it keeps both sides): - device.c, device_context.c, frame_import.c, frame_import_admit.c, frame_pool.c, submit.c: both backends, each in its own #ifdef block. - fence.c: the HIP lane's dispatch by device and kind, with SYCL_EVENT added to create, wait and destroy; SYNC_FILE and GL_SYNC fences are polled through vmafx_fence_poll() (the virtual test clock applies). - cuda/import_fence.c: the HIP lane's shared release-event table. - vmafx_device_cells.h and its contract test: the HIP lane's device cells with the SYCL lane's shared helpers (vmafx_exact_cells.h) and table parameter. - core/api/vmafx.toml: GL texture, device backend and release callback texts name all three lanes; ABI numbers stay integration's (0.1.7). - meson: both lanes' sources in libvmafx_sources (the WP6 split; the SYCL lane used libvmaf_sources), window and sync-object sources both kept. - invariant pages and docs/state.md: both sides; generated files regenerated. One implementation (HISS-19): the HIP lane's gl_sync.c and sync_file.c and the SYCL lane's sync_object.c defined the same four functions. sync_object.c (the superset: dma-buf implicit fences too) stays; the HIP files, their declarations in internal.h and their citation-registry entries go, and the HIP callers include sync_object.h. Contract reconciled: the lanes disagreed on vmafx_fence_destroy() of a SYNC_FILE fence (the HIP lane documents the descriptor as borrowed and refuses; the SYCL lane closed it). The API documents the descriptor as borrowed and the destroy as for fences the library returned, so the destroy refuses; test_vmafx_import_sycl's sync_file case now expects VMAFX_E_NOTSUP and closes the descriptor itself. Research numbers: both lanes added Research-2160. The SYCL digest keeps 2160 (cited by code and the accepted ADR-2091); the HIP GL digest becomes Research-2161, with its link in ADR-2132's references and the state row. CPU fast suite 389 passed, 0 failed; CUDA build clean. SYCL and HIP builds follow after the SYCL formats lane.
…egration branch The SYCL half of ADR-2133 repeats the format table its HIP-based sibling brought: abi_version and the thirteen VmafxPixelFormat entries keep integration's numbers (0.1.9); the changelog fragment names the SYCL device and both ABI numbers; the import test helper cites ADR-2091 and ADR-2092; ADR-2133's references keep the link to ADR-2092; the formats state row takes the SYCL lane's text (only Metal still refuses the layouts); the ADR order list keeps one entry for ADR-2133 (both lanes appended it). Generated files took integration's side and were regenerated (vmafx-api.py --abi-check against the previous head: 0 additions). CPU build clean.
…port.c The HIP lane's GL follow-up (ADR-2132) added core/src/vmafx/egl_export.c and asked for the SYCL lane's own EGL export (core/src/sycl/import_gl.c) to be folded onto it when the lanes met (HISS-19). The two differ in one point: the HIP device reads linear rows, so a tiled export is refused or copied on the GPU, while the SYCL device de-tiles Intel tilings itself and wants the export as the driver made it. vmafx_egl_export_planes() takes a VmafxEglTiled mode (REFUSE, COPY, KEEP) instead of allow_copy; a target fourcc of 0 and a NULL device_pci skip those checks. sycl/import_gl.c keeps only the translation of the exported planes into a DMABUF descriptor. In KEEP mode no writers wait is made: the SYCL dma-buf import honours the implicit fences (D8 retry). test_vmafx_one_egl_export_contract (fast suite) refuses a lane that loads EGL itself and a second definition of a sync-object function; it fails on the SYCL lane's former import_gl.c and on a planted sync_file.c, and passes now. test_vmafx_import_hip_contract's planted anchor follows the new open_display() call. Fixes the integration base found by the first SYCL and HIP builds: - sycl/import_device.c leaves the engine with vmafx_engine_leave(context, previous) (the WP4 signature); the API invariant page names it. - sycl/d3d11_import.cpp defines vmaf_sycl_import_d3d11_surface() off Windows as a refusal (-ENOSYS): libvmaf_sycl.h declares it everywhere and the WP6 version script lists it, so a Linux SYCL link failed with "undefined version: VMAF_LEGACY_SYCL". Evidence: CPU fast suite 390 passed, 0 failed. SYCL build (icpx 2026.0, dg2-g11) on the Arc A380: test_vmafx_import_sycl 16, _fence 6, _bitexact 3, _formats 2, _gl 2, test_vmafx_fence_kinds 5, all passed. HIP build in the pinned ROCm 10.1.0 image on the gfx1036: test_vmafx_import_hip 16, _fence 8, _formats 2, _bitexact 3, _gl 5 (through the shared export), test_hip_shared_frame 9, test_vmafx_fence_kinds 5, all passed.
…AFx API guide docs/api/vmafx/index.md is written by hand, not generated, but the lane merges took integration's side of it as if it were (a regeneration does not rebuild it). The HIP devices and SYCL devices sections, the formats rows and the fence and import overviews of the five lanes were lost, and three links from the backend pages pointed at missing anchors (MkDocs strict). The file is now the three-way merge of each lane against its base, applied in order (HIP with its GL follow-up, the formats lane, SYCL, the SYCL formats half), with the overview, import and fence paragraphs written for all three device backends and the HIP and SYCL sections kept side by side.
…escriptor zeMemAllocDevice() with a dma-buf import descriptor closes the descriptor it is given on compute runtime 26.35 (Arc A380, measured with an LD_PRELOAD close() log) and the specification leaves ownership open. The SYCL import duplicated the caller's descriptor, handed the duplicate to the driver and closed it again at the frame's release: the second close hit whatever descriptor had taken the number in between. Running the FFmpeg vmafx filter on VAAPI frames mapped to DRM PRIME, that was the reference frame's mapped descriptor, and the eleventh import failed. vmaf_sycl_dmabuf_import_queue() now gives the driver a private duplicate taken from a high floor and closes it afterwards unless the driver did; the floor makes that check safe against a descriptor opened meanwhile. The caller's descriptor stays the caller's on every driver. The legacy VA surface import shares the function and so the fix. Test: test_vmafx_import_sycl test_import_closes_only_its_own opens a pipe after an import (it takes the number such a driver freed) and checks that it, and the caller's dma-buf descriptor, survive the frame's release. It failed on the A380 before the change and passes; the SYCL import tests pass.
…152) Brings the Vulkan import lane (#2375) onto the integration branch. The lane's first commit (02aa3a7, its own merge of the SYCL and HIP lanes) is not taken: both lanes are here already. Resolved per hunk: - core/api/vmafx.toml: the lane's twelve additions are ABI 0.1.10 here (0.1.5 on the lane, ADR-1897); generated files regenerated. - cuda/import_frame.c: the Vulkan check runs first; a packed or MSB layout is refused for an OPTIMAL Vulkan image as for CUDA arrays. - hip/import_frame.c: the GL dma-buf export (ADR-2132) and the Vulkan dma-buf path are both kept; the lane's VmafxHipGl registration struct is dropped, since the GL interop it served is gone here. - vmafx/internal.h: pooling and window declarations and the Vulkan declarations are both kept. - SYNC_FILE fences: vmafx_fence_destroy() closes the descriptor (ADR-2091 item 6), as the lane resolves it; the API index, the device-frames agents page, the rebase note and the SYCL sync_file test follow. - test_vmafx_import_vulkan_api: designated initialisers for the layout, which gained the packed fields. (cherry picked from commit cff8500)
…3, ADR-2152) The lane's API section "Vulkan frames" follows the SYCL devices section; the lane's copy of the older HIP devices section is not taken (the HIP GL dma-buf export of ADR-2132 is described here already). docs/state.md takes the Vulkan rows and T-HIP-DMABUF-IMPORT-WAITS-FOR-WRITER; the lane's T-HIP-ROCM10-GL-TEXTURE-READ row stays closed here. The shared PCI parser replaces vmafx_hip_parse_pci(). The changelog fragment and rebase note give ABI 0.1.10 for the integration branch; the duplicate SYCL rebase section the merge produced is dropped. (cherry picked from commit 3decc47)
…rser (RC4 WP3, ADR-2152) vmafx_sycl_rt.cpp hands the bus id string of the Intel device-info extension to C, which parses it with vmafx_parse_pci_bus_id() as the CUDA and HIP lanes do (no sscanf, cert-err34). The CUDA interop asserts its plane and import indices, and the PCI parser's refused forms are one table-driven test (readability-function-size). (cherry picked from commit 43b2231)
…lease paths (RC4 WP3, ADR-2152) core/test/meson.build lists each Vulkan test program's source by name, so the stale-source-reference check (ADR-1135) resolves them; the CUDA Vulkan release and the SYCL device query carry the asserts the Power-of-10 assertion-density gate requires of a function of 20 lines or more. (cherry picked from commit c101d50)
vmaf_sycl_dmabuf_import_queue() returned -errno when driver_fd() failed. fcntl() sets errno on failure, but a zero errno would turn the failure into a return of 0 with no pointer; the import now falls back to -EBADF, as the master fix (#2377) does.
lusoris
force-pushed
the
rc4/api-wp6-compat
branch
2 times, most recently
from
October 8, 2026 09:54
0c6dcd8 to
14dc297
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
One branch,
rc4/integration, carries every RC4 API draft so they can be built and tested together: WP6 #2303 (with WP5 #2288, WP8 #2258, WP3-common #2213, WP2 #2199, WP1 #2187 and the prototype #2173 below it), the WP3 CUDA lane #2277, WP4 window scores #2287, incremental motion #2290, and since 2026-10-07 the WP3 HIP lane #2341 with its GL follow-up #2360, the WP3 SYCL lane #2342, and 4:2:2 / 4:4:4 import #2367 / #2368, and the WP3 Vulkan import lane #2375 (picked without its own merge of the SYCL and HIP lanes). It starts fromrc4/api-wp6-compat(24e4104de) and merges the three other stacks, resolving each conflict per hunk; generated files took one side and were regenerated once per merge. It is a test vehicle and a base for WP9: it never goes intotrain.order, and each lane still lands through its own PR. The merged branch needed seven fixes that no single lane shows; they are separate commits after the merges, each with adocs/state.mdrow, so the lanes can take them back.Merged heads:
rc4/api-wp3-cuda15bd091e2,rc4/api-wp4-windows4037b18b4,rc4/api-motion-incremental4947774ba,rc4/api-wp3-hip20a3f4d4d,rc4/api-wp3-hip-gl-dmabuf657e74a53,rc4/api-wp3-formats-422-4448998921df,rc4/api-wp3-sycl5dafa1cbd,rc4/api-wp3-formats-422-444-syclc50cbc8e0,rc4/api-wp3-vulkan-importc101d5030(its four commits after02aa3a7d2picked:798c42c5d,541a91a01,bae72c148,a21d20ce1).ABI numbering (ADR-1897)
Several lanes bumped
abi_versionfrom 0.1.3 to 0.1.4 independently. The integration gives each addition set the next free patch, in merge order:VMAFX_MEMORY_GL_TEXTURE,VMAFX_FENCE_GL_SYNC,VmafxFrameImport.release/uservmafx_context_max_in_flightVmafxPixelFormatentries (NV16toYUV444P_MSB)VMAFX_MEMORY_VULKAN,VMAFX_FENCE_VULKAN_SEMAPHORE,VmafxVulkanHandleType,VmafxVulkanTiling,VmafxVulkanFlags,VmafxDeviceInfo.pci,VmafxFrameImport.acquire_more/vulkan_*,vmafx_frame_signal_on_release()The HIP and SYCL lanes add no entry of their own; their texts on the GL texture memory, the device backend and the release callback name all three device lanes, with integration's 0.1.7.
Every entry keeps
since = "0.1"(nodeVMAFX_0.1); only the "Added in ABI 0.1.x" notes, the changelog fragments, the AGENTS page and the rebase note of the two lanes changed. ADR-2023's body still says 0.1.4: it is Accepted and frozen.--abi-check(scripts/codegen/vmafx-api.py): append-only successor oforigin/rc4/api-generation-prototype(378 additions), oforigin/rc4/api-wp6-compat(33),origin/rc4/api-wp3-cuda(237),origin/rc4/api-wp4-windows(212) andorigin/rc4/api-motion-incremental(212). No break, so no new minor.Conflicts resolved
Merge of the WP3 CUDA lane (
ec4ae3023):core/api/vmafx.toml:abi_version0.1.7 and the lane's four "Added in ABI" notes, as above.core/src/vmafx/fence.c: the lane'svmafx_fence_poll()refactor kept, with WP5's exportedvmafx_monotonic_ns()in place of the staticmonotonic_ns().core/src/vmafx/internal.h:VmafxContextkeeps bothlane_state(CUDA lane) andprovenance(WP5).docs/development/rebase-sensitive-invariants.md: both new entries kept.docs/state.md:scripts/dev/resolve-state-md-conflict.py.bindings/python/vmafx/_api.py,core/include/vmafx/version.h,core/src/vmafx_symbols.txt,core/test/test_vmafx_abi_layout.c,docs/api/vmafx/reference.md, ADR indexes,scripts/ci/source-adr-citations.json,CHANGELOG.md): one side, regenerated.Merge of WP4 (
76bb126ab):core/api/vmafx.toml: both lanes appended a section before thelibvmaf_bridge.hfunctions; the provenance section (WP5) is followed by the window section (WP4).abi_version0.1.8, WP4's twenty notes 0.1.8, WP4'svmafx_context_frame_retention()doc kept.core/src/vmafx/fence.c: one timed wait. The CUDA lane'svmafx_fence_poll()now carries WP4's rule for timeouts of 2^62 ns and more (no clock check,UINT64_MAXrounds; the icx zero-round fix of3a042aac4), andvmafx_host_fence_wait(), exported forwindow.c, wraps it.core/src/vmafx/internal.h:VmafxContextkeepslane_state,provenanceandwindows; the 0.1.6 and 0.1.8 minimum-size macros.core/src/vmafx/context.c: a failed create closes the windows and releases the provenance state (a failedvmafx_windows_init()releases the provenance state too); destroy keeps WP5'sassert(context->engine != NULL)and WP4'svmafx_windows_pause().core/src/vmafx/register.c: WP5'sregistered_name()with WP4'svmafx_engine_leave(context, previous). The same new signature (WP4's engine lock) at the call sites the other lanes added:core/src/cuda/import_device.c,context_frames.c,report.c,tiny_model.c.core/src/libvmaf.c: WP4's three engine functions (thread count, in-flight bound, subsample) kept; WP2's libvmaf forwarders stay removed, since WP6 moved the libvmaf functions into the compat library.core/test/meson.build: WP4's window tests before WP5's public test list.docs/development/rebase-sensitive-invariants.md: both entries kept.docs/state.md: the resolver. Generated files: one side, regenerated.Merge of incremental motion (
d1d5ae2b1):core/api/vmafx.toml: the lane'svmafx_window_submit()doc (a window over a VMAF model completes one frame after its last frame, ADR-2090) with the integration's 0.1.8 note.changelog.d/added/api-window-scores.md: the lane's sentence, ABI 0.1.8.docs/state.md:T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06takes the lane's Recently closed row, as WP4's answer to request WP4-1 says; the rest by the resolver.Second round (2026-10-07): HIP, SYCL and 4:2:2 / 4:4:4
Merges
a9ff02915(HIP),ff42580ee(HIP GL),d9513b463(formats),a082db7a9(SYCL),a03b8c1bf(SYCL formats). Per hunk:device.c,device_context.c,frame_import.c,frame_import_admit.c,frame_pool.c,submit.c): both backends, each in its own#ifdefblock.fence.c: the HIP lane's dispatch by device and kind withSYCL_EVENTadded; SYNC_FILE / GL_SYNC polled throughvmafx_fence_poll().vmafx_device_cells.hand its contract test: the HIP lane's cells with the SYCL lane's shared helpers.libvmafx_sources(the SYCL lane usedlibvmaf_sources); window and sync-object sources kept.docs/api/vmafx/index.mdis hand-written: three-way merged per lane (3b9a41457; the first pass had taken one side as if generated, which MkDocs strict caught on three anchors).docs/adr/_index_fragments/_order.txt: one ADR-2133 entry (both formats lanes appended it).One implementation (HISS-19): the HIP lane's
gl_sync.c/sync_file.cduplicated the SYCL lane'ssync_object.c;sync_object.cstays. The SYCL lane's own EGL export is folded ontocore/src/vmafx/egl_export.c(44cbb8cd6;vmafx_egl_export_planes()takes a REFUSE / COPY / KEEP mode for tiled exports).test_vmafx_one_egl_export_contractrefuses a second copy of either, planted copies included.Contract reconciled (third round):
vmafx_fence_destroy()closes a SYNC_FILE fence's descriptor (ADR-2091 item 6), as the Vulkan lane resolves the HIP / SYCL difference; a GL sync is still refused. The second round had taken the HIP lane's refusal; the API index,core/src/AGENTS.d/vmafx-device-frames.md, the rebase note and the SYCL sync_file test follow the new contract.Third round (2026-10-07): Vulkan import
The lane's first commit
02aa3a7d2is its own merge of the SYCL and HIP lanes; both are here already, so it is dropped and the four commits after it are cherry-picked (-x). Per hunk:core/api/vmafx.toml:abi_version0.1.10; the lane's twelve "Added in ABI" notes say 0.1.10.--abi-check --against-refthe previous tip: append-only successor (18 additions).core/src/cuda/import_frame.c: the Vulkan check first; a packed or MSB layout (4:2:2 / 4:4:4 lane) is refused for an OPTIMAL Vulkan image as for CUDA arrays, since OPTIMAL images are read out of arrays.core/src/hip/import_frame.c: the GL dma-buf export (ADR-2132) and the lane's Vulkan-as-dma-buf path both kept; the lane'sVmafxHipGlregistration struct is not taken (its GL interop is gone here). The sharedvmafx_parse_pci_bus_id()replacesvmafx_hip_parse_pci().core/src/vmafx/internal.h: WP4's pooling and window declarations and the Vulkan declarations both kept.docs/api/vmafx/index.md: the lane's "Vulkan frames" section after "SYCL devices"; the lane's copy of the older HIP section is not taken.docs/state.md: the lane's rows and bucket entries, exceptT-HIP-ROCM10-GL-TEXTURE-READ-2026-10-06, which stays closed here (feat(hip): import GL textures through EGL dma-buf export so the pinned ROCm 10.1 reads them (ADR-2132) #2360).docs/rebase-notes.md: the merge had duplicated the SYCL lane's section; one copy kept. The changelog fragment and the rebase note give ABI 0.1.10 here.core/test/test_vmafx_import_vulkan_api.c: designated initialisers for the layout, which gained the packed fields (-Wmissing-field-initializers).Fixes the merged branch needed
2eef29685test_model_collection_score_repeat,test_cuda_float_ms_ssim_host_traffic,test_motion_window_incremental) linkedlibvmaf.get_static_lib(), which the WP6 split no longer offers:meson setupfailed784f57801advance_extractors()(incremental motion) calledadvance()without installing the extractor as feature producer (WP5):integer_motion2had sourceunknownin the provenance recordtest_vmafx_provenancetest_features_in_name_orderT-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-063b957b8eelibvmaf_sources, renamedlibvmafx_sourcesby WP6: a CUDA build failed to configuremeson setup -Denable_cuda=trueT-RC4-CUDA-IMPORT-SOURCES-OLD-LIBRARY-VARIABLE-2026-10-06885912a34VmafPicturestructs set aftervmaf_read_pictures()while the compat library clears them:test_compat_conformancefailed in every CUDA buildbuild-cudafast suiteT-RC4-COMPAT-CONFORMANCE-CUDA-CONSUMED-2026-10-06df671fc3dtest_device_target_header_dependenciesexpected 21 CUDA targets (the CUDA lane added a 22nd) and the Windows fallback check did not parsecuda_direntriesbuild-cudafast suiteT-RC4-CUDA-IMPORT-CONVERT-TARGET-COUNT-2026-10-0650eab63cf--backend) and expected the CPU recordbuild-cudafast suiteT-RC4-PROVENANCE-TESTS-SCORE-ON-GPU-2026-10-0622db201a9build-asanT-RC4-COMPAT-TINY-7X7-STACK-OVERFLOW-2026-10-06fba26df2bvmafx_context_use_model()accepted a second model of a taken name, whose scores were the first one's (vmaf_v0.6.1negreported 76.667831, alone 75.073658)T-RC4-MODEL-NAME-COLLISION-2026-10-0607d9fceae..._0,7names under a decimal-comma localeT-OPTION-NUMBERS-CALLER-LOCALE-2026-10-0644cbb8cd6sycl/import_device.candhip/import_device.cused the pre-WP4vmafx_engine_leave(previous); a Linux SYCL link failed (undefined version: VMAF_LEGACY_SYCL) becausevmaf_sycl_import_d3d11_surface()was defined on Windows only, now a refusing stub elsewhereadf6121c9-errno; a zeroerrnowould read as success. Falls back to-EBADF, as the master fix #2377 does869ec15f6)869ec15f6T-SYCL-DMABUF-IMPORT-DOUBLE-CLOSE-2026-10-07Open:
T-VMAFX-IMPORT-CUDA-FENCE-CANARY-TIMING-2026-10-06.test_vmafx_import_cuda_fencetest_release_canarycounts a correctly ordered canary as early when other device work outlasts its fixedHOLD_MShold (1 of 16 with--num-processes 2; 0 bad scores). Serial runs pass. Owner: the CUDA lane.Type
build/ci— tooling / infra (integration branch; the fixes arefix/test/buildcommits)Checklist
make format && make lintis green locally — the commit and pre-push hooks pass.adf6121c9): CPU fast suite 391 OK, 2 expected fail, 0 fail, 1 skipped; CUDA on the RTX 4090:test_vmafx_import_cuda,_bitexact,_formats,_fence,_gl,test_vmafx_import_vulkan_cuda,_bitexact(21 s),_fence,test_vmafx_import_vulkan_api,test_vmafx_fence_kinds10 of 10 OK; SYCL on the A380:test_vmafx_import_sycl17,_fence6,_bitexact3 (236 cells, 0 differing),_formats2,_gl2,test_vmafx_import_vulkan_sycl5,_bitexact3 (472 cells, 0 differing),_fence2,test_vmafx_fence_kinds5, all OK; HIP in the pinned ROCm 10.1.0 image on the gfx1036:test_vmafx_import_hip16,_fence8,_bitexact3,_formats2,_gl5,test_vmafx_fence_kinds5, all OK.test_vmafx_import_vulkan_<lane>_ffmpegskips (77) on CUDA and SYCL: it needs decoded inputs the run did not give. The HIP Vulkan tests are not built in the pinned image, which has no Vulkan development files (dependency('vulkan')not found); the lane measured them in the same image with the Vulkan loader and headers added (feat(api): import Vulkan frames on CUDA, SYCL and HIP (RC4 WP3, ADR-2152) #2375). Second round: CPU fast suite 390 OK, 0 fail, 1 skipped; SYCL build (icpx 2026.0,dg2-g11) on the Arc A380:test_vmafx_import_sycl17,_fence6,_bitexact3,_formats2,_gl2,test_vmafx_fence_kinds5, all OK; HIP build in the pinned ROCm 10.1.0 image on the gfx1036:test_vmafx_import_hip16,_fence8,_formats2,_bitexact3,_gl5,test_hip_shared_frame9, all OK; CUDA:test_vmafx_import_cuda12,_gl2,_formats2 OK. First round: CPU build (build-cpu, gcc 16.2.1,-Db_lto=false):python3 scripts/ci/run_meson_test.py -- -C build-cpu --suite=fast --num-processes 6385 OK, 2 expected fail (the planted conformance defects), 0 fail, 1 skipped (test_vmafx_api_abi_append_only: the merge base withorigin/masterhas no definition yet). CUDA build (build-cuda, release,-Denable_cuda=true): fast suite without thegpusuite 394 OK, 2 expected fail, 1 skipped;gpusuite under the CUDA lock with--num-processes 168 of 68 OK;test_vmaf_cuda_gpumask,test_vmaf_cuda_threads,test_vmaf_feature_backend_cuda,test_cuda_parity_gate_default_runOK. Not run:test_cuda_exact_twin_matrixandtest_cuda_v1_models_no_fallback(slow suite) each take more than the 300 s a device lock may be held.test_vmafx_*,test_compat_conformance,test_motion_window_incremental;build-asanwithaddress,undefined,build-tsanwiththread): ASan/UBSan 24 of 24 OK after22db201a9, with LeakSanitizer reporting only two 256-byte blocks allocated insidelibonnxruntime.sowith no fork frame (suppressed for the run withleak:libonnxruntime.so); TSan 24 of 24 OK./cross-backend-diffand the worst ULP is ≤ 2. — not applicable: no kernel changed; the CUDA exact-twin and parity tests above pass..c/.cpp/.cu/.h/.hpp, it has the appropriate license header — no new native file.!orBREAKING CHANGE:and the migration path is documented below. — not a breaking change: the definition is an append-only successor of every merged lane.docs/adr/_index_fragments/<NNNN-slug>.md— no ADR added: the merges follow ADR-1897; the fixes are defect fixes.tidy: not measured on this branch; no file is new, and the fixed native sources (
core/src/libvmaf.c,core/src/vmafx/{fence,context,register,context_frames,report,tiny_model}.c,core/src/cuda/import_device.c,core/test/test_compat_conformance_api.c) change by merge resolutions and one-line edits. Each lane measures its own files before it lands.Bug-status hygiene (ADR-0165)
docs/state.mdupdated — the rows in the fixes table (closed) and one open row, listed above.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.make test-netflix-golden(core/build-golden, gcc): 280 passed, 3 skipped after the merges (784f57801), at22db201a9(after the engine change of885912a34) and again at the tipadf6121c9(after the Vulkan round).Cross-backend numerical results
Not applicable: no kernel or extractor changed. Imported device frames score as host frames on every exact cell (the bit-exactness counts above, Vulkan imports included). CUDA twins:
test_cuda_exact_twinsand the*_paritytests of thegpusuite pass on the RTX 4090.Deep-dive deliverables (ADR-0108)
AGENTS.mdinvariant note — no rebase-sensitive invariants beyond the lanes' own:core/src/AGENTS.d/vmafx-provenance.mdalready requires the producer around a direct extractor callback, which784f57801follows.Reproducer
Known follow-ups
3b957b8ee(after WP6),df671fc3d; WP550eab63cf; WP6885912a34,22db201a9; incremental motion784f57801(after WP5).VmafxWindowResult) is still open: WP5 and WP4 meet only here.VmafxColorMatrix,VmafxColorRange,VmafxColorTransferinvmafx/types.h) collide with WP6's of the same names here (duplicate type: VmafxColorRange), and its--rgb_*flags are hand-written where WP8 generates the command line. Request WP13-6 asks the lane to converge on WP6's enums as ADR-2146 already decides; integration then gives its additions ABI 0.1.11.origin/master, like every RC4 lane; it moves when they rebase.