Repository navigation
feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) - #2290
Merged
Merged
Conversation
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 6, 2026 14:48
939479a to
4947774
Compare
lusoris
added a commit
that referenced
this pull request
Oct 6, 2026
…tion lane's draft T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06 stays open on this branch, as the maintainer decided; the lane that derives motion2 / motion3 frame by frame now has a draft (#2290, rc4/api-motion-incremental, ADR-2090 there), which closes the row.
This was referenced Oct 6, 2026
Draft
lusoris
force-pushed
the
rc4/api-wp4-windows
branch
2 times, most recently
from
October 8, 2026 12:56
3eabc39 to
9926bfe
Compare
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 8, 2026 14:55
4947774 to
d5f18fd
Compare
lusoris
marked this pull request as ready for review
October 8, 2026 14:55
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 8, 2026 15:32
d5f18fd to
a7b218c
Compare
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
…re-commit hook (#2615) * fix(ci): list SPONSORS.md among the CI impact planner's known top-level files (#2619) * fix(ci): list SPONSORS.md among the CI impact planner's known top-level files top-level entry the planner may see, and test_ci_impact.py test_every_top_level_repo_entry_is_known fails on master since then (SPONSORS.md is neither a known file nor under a known prefix). The file joins known_files next to SECURITY.md and SUPPORT.md; it selects no suite of its own, as those do not. State row T-CI-IMPACT-SPONSORS-UNKNOWN-2026-10-08. * feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) (#2290) * feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) The integer motion extractors and their GPU twins derived motion2 and motion3 of every frame in flush(), so a VMAFx window over a VMAF model, a per-frame model score and vmaf_score_pooled() over the frames read so far waited for the end of the stream (state row T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF windows in the 1.0 scope (Q-038, #2138, #2238). Frame i is now final once the SAD scores of frames 0 to max(i + 1, min_idx) are in the collector. vmaf_motion_window_advance() derives every complete frame with upstream's per-frame statements in index order and carries the stamp and the moving-average value in a VmafMotionWindowState; vmaf_motion_window_flush() derives the rest. The engine calls a new optional VmafFeatureExtractor.advance() after every accepted frame and read fence, and the VMAFx completion thread calls it through vmaf_engine_advance() under the engine lock before each pass, so a window completes as soon as a worker's SAD is in, also while the feeder stalls. motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL, HIP, Metal) register it. A synchronous read of a fed frame without a collector slot now fences instead of answering -EINVAL (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06). Every value is the flush-time value: 32 988 per-frame motion values against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K; default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU backend against the CPU, 0 different; parity gate motion cells exact on CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped. * fix(api): name the extractor as producer of the scores advance() writes advance_extractors() called VmafFeatureExtractor.advance() (ADR-2090) without installing the extractor as the thread's feature producer (ADR-2073), so the integer_motion2 vector advance() creates first carried source unknown and the provenance record named no extractor for it. The threaded flush installs the producer around flush(); advance() now gets the same. test_vmafx_provenance test_features_in_name_order failed on the merged tip and passes. State row T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06. * fix(api): answer pending for a fed frame whose motion3 is not derived yet With incremental motion (ADR-2090) motion3 of frame i is written when frame i + 1 is scored. vmafx_score_frame() of the frame just submitted returned VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the end of the motion3 vector's storage and the collector answered -EINVAL. engine_score_at_index() now checks the inputs of a fed frame before it predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06. * docs(motion): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197) * docs(agents): write the incremental motion page in the internal register (praetor caveman lint) * ci(tidy): measure the translation units of the incremental motion window in the cpu lane The clang-tidy coverage rule on master requires every tracked translation unit to be read by a lane. The new and touched units of this pull request were measured in the dev container (scripts/dev/tidy-lane.sh --write --only ... cpu, clang-tidy 22.1.8): 0 findings, 0 uncited NOLINT; they join the cpu lane's measured sources. * feat(rust): give Rust twins their own advance callback (ADR-2090, MI-1) The shim copies the C extractor's descriptor whole (add_twin()), so after incremental motion (ADR-2090) a Rust twin inherited the C advance(): for motion_rust it derived motion2 / motion3 from the Rust SAD scores into the C-layout priv while frames came in, and the Rust flush then appended every frame again, -EINVAL at the flush; with worker threads the inherited advance also marked the registered context initialised, so the Rust flush found no instance. Maintainer answer Q-093: Rust implements advance. VmafxRsTwin gains advance as its last field. VMAFX_RS_ABI_VERSION stays 1: no release has shipped the Rust ABI, so the layout changes in place. The header is regenerated verbatim with cbindgen 0.29.4 (scripts/dev/rust-abi-header.sh, --check up to date). vmafx_fex::Extractor gains advance(&mut self, host), default nothing, with its trampoline. The shim sets advance to twin_advance() when the C extractor has one and refuses a twin without the entry point. advance_one_extractor() in libvmaf.c runs init_shared_rust_twin() before a pooled Rust twin's first advance, as the threaded flush does before its flush. test_rust_abi_layout checks the new offset and the entry point. * feat(rust): derive motion_rust's motion2 and motion3 frame by frame (ADR-2090, MI-1) motion_rust now implements Extractor::advance with the statements of integer_motion.c's window: frame i is derived once the SAD scores of frames 0 to max(i + 1, min_idx) are in, the stamp and the moving-average value are carried in a State as in VmafMotionWindowState, and the flush continues where the last advance stopped (motion_window_stamp(), motion_window_count_sads(), motion_window_derive(), vmaf_motion_window_advance(), vmaf_motion_window_flush(), ported in window.rs). A VMAF window over motion_rust completes before the flush, as over the C motion. test_rust_motion_window_incremental (suite rust): motion_rust on a context, both windows, 0, 2 and 3 worker threads; frame i final after frame i + 1 and before the flush, values unchanged after they became final, every SAD, motion2 and motion3 equal to the C motion bit for bit. It failed on the merged tree (flush -22: the inherited C advance and the Rust flush appended frame 0 twice), failed with scoring -22 at 2 threads before the engine initialised the pooled twin, and fails on a planted off-by-one in the Rust advance ("final too early"). Cargo tests cover advance + flush against the C derivation with the SAD scores in order, swapped in pairs and reversed, and short streams of 0 to 4 frames. rust_twin_diff.py --feature motion on netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45 EQUAL. State row T-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07. Carried onto the 2290 landing: the rebase note joins the motion-window-incremental fragment (ADR-2197), which also names read_pictures_owned(), the read helper's name since the WP6 split, as does the picture-ownership page. The Rust framework page is rewritten for the caveman lint and names the engine library (libvmafx) as the archive's only target. * docs(agents): restore the Metal motion twins' advance note on the AGENTS.d pages The incremental motion window registers .advance on integer_motion_metal and motion_v2_metal (ADR-2090). The rebase across the conversion of core/src/feature/metal/AGENTS.md into AGENTS.d pages took master's generated index and lost this pull request's two hunks of the old single file. The motion-fps-weight and motion3-v2 pages now name vmaf_motion_window_advance() next to the flush, the advance contract test, and integer_motion_v2_metal.mm as a path the motion3-v2 page covers; the index is regenerated. * docs(community): link the live Patreon page and its euro tiers (#2620) * docs(community): link the live Patreon page and its euro tiers * fix(ci): skip docs/rebase-notes.d in the Markdown Lint job like the pre-commit hook (#2615) * fix(ci): skip docs/rebase-notes.d in the Markdown Lint job like the pre-commit hook Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 8, 2026 17:06
a7b218c to
d5f18fd
Compare
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
7a3a7d0 so #2290 lands as its own commit The merge train's batch push of 17:41 failed after it had pushed the stacked squashes of the batch to their pull request branches. The fixer of skip docs/rebase-notes.d in the Markdown Lint job", #2615) also carries the whole change of #2290 (incremental motion2 / motion3, ADR-2090) and the CI-impact entry of #2619. #2290 never landed as its own commit, and its code, tests, docs, changelog fragments and state rows sit under #2615's title. This reverts #2290's own diff (its gated head d5f18fd against its base cfdcfb5), reverse-applied on master 0bf2996: 66 of its 68 files return to their content before #2290; docs/state.md keeps the rows of the commits that landed since. ADR-2090 and its index fragment stay: the ADR indexes are written only by the landing render and link to it, so deleting it fails the ADR link gate; #2290 re-lands without them. #2615's own change (the Markdown Lint fragment scope) and the SPONSORS.md entry of .github/ci-impact.json (#2626, also #2619's) stay. #2290 re-lands next as its own squash. docs/state.md also folds the duplicate row of the SPONSORS.md defect: T-CI-IMPACT-SPONSORS-MD-UNKNOWN-2026-10-08 (#2626) stays and names T-CI-IMPACT-SPONSORS-UNKNOWN-2026-10-08 (#2619, closed as a duplicate) as its alias. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
…so #2620 lands as its own commit 7a3a7d0 (#2615's squash, built on the stacked heads of the failed batch push of 17:41) also carries #2620 (the live Patreon page and its euro tiers). This reverts #2620's own diff (its head 66be228 against its base cfdcfb5): .github/FUNDING.yml, GOVERNANCE.md, README.md, SPONSORS.md, docs/support-vmafx.md and changelog.d/changed/patreon-live.md return to their content before #2620. #2620 re-lands after #2290, as its own squash. Signed-off-by: Lusoris <lusoris@proton.me>
This was referenced Oct 8, 2026
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 8, 2026 22:22
d5f18fd to
4613fc6
Compare
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 9, 2026 08:13
4613fc6 to
f6db3be
Compare
…ce (#2647) * feat(cuda): evaluate two ADM viewing distances in one adm_cuda instance adm_cuda takes adm_norm_view_dist_extra and the merge callback: the scale-0 and scale-1..3 DWT run once per frame, and the denominator, CSF, contrast-masking and AIM kernels run once per distance into that distance's result block (tmp_res and results_host hold two). The host concludes each distance with its own CPU contexts and files the second under the <base>:nvde keys. Both distances return the CPU's bits (ADR-2795; test_adm_two_views_exact, test_adm_merged_registrations_exact). The merge callback and the second distance's names move into adm_view_dist.c, which reads the options by name at each descriptor's own offsets, so the CPU extractor, the Rust twin and adm_cuda share one implementation; integer_adm.c drops its copies. Signed-off-by: Lusoris <lusoris@proton.me>
…ws complete before the flush (ADR-2090) (#2290) * feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) The integer motion extractors and their GPU twins derived motion2 and motion3 of every frame in flush(), so a VMAFx window over a VMAF model, a per-frame model score and vmaf_score_pooled() over the frames read so far waited for the end of the stream (state row T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF windows in the 1.0 scope (Q-038, #2138, #2238). Frame i is now final once the SAD scores of frames 0 to max(i + 1, min_idx) are in the collector. vmaf_motion_window_advance() derives every complete frame with upstream's per-frame statements in index order and carries the stamp and the moving-average value in a VmafMotionWindowState; vmaf_motion_window_flush() derives the rest. The engine calls a new optional VmafFeatureExtractor.advance() after every accepted frame and read fence, and the VMAFx completion thread calls it through vmaf_engine_advance() under the engine lock before each pass, so a window completes as soon as a worker's SAD is in, also while the feeder stalls. motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL, HIP, Metal) register it. A synchronous read of a fed frame without a collector slot now fences instead of answering -EINVAL (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06). Every value is the flush-time value: 32 988 per-frame motion values against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K; default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU backend against the CPU, 0 different; parity gate motion cells exact on CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped. * fix(api): name the extractor as producer of the scores advance() writes advance_extractors() called VmafFeatureExtractor.advance() (ADR-2090) without installing the extractor as the thread's feature producer (ADR-2073), so the integer_motion2 vector advance() creates first carried source unknown and the provenance record named no extractor for it. The threaded flush installs the producer around flush(); advance() now gets the same. test_vmafx_provenance test_features_in_name_order failed on the merged tip and passes. State row T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06. * fix(api): answer pending for a fed frame whose motion3 is not derived yet With incremental motion (ADR-2090) motion3 of frame i is written when frame i + 1 is scored. vmafx_score_frame() of the frame just submitted returned VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the end of the motion3 vector's storage and the collector answered -EINVAL. engine_score_at_index() now checks the inputs of a fed frame before it predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06. * docs(motion): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197) * docs(agents): write the incremental motion page in the internal register (praetor caveman lint) * feat(rust): give Rust twins their own advance callback (ADR-2090, MI-1) The shim copies the C extractor's descriptor whole (add_twin()), so after incremental motion (ADR-2090) a Rust twin inherited the C advance(): for motion_rust it derived motion2 / motion3 from the Rust SAD scores into the C-layout priv while frames came in, and the Rust flush then appended every frame again, -EINVAL at the flush; with worker threads the inherited advance also marked the registered context initialised, so the Rust flush found no instance. Maintainer answer Q-093: Rust implements advance. VmafxRsTwin gains advance as its last field. VMAFX_RS_ABI_VERSION stays 1: no release has shipped the Rust ABI, so the layout changes in place. The header is regenerated verbatim with cbindgen 0.29.4 (scripts/dev/rust-abi-header.sh, --check up to date). vmafx_fex::Extractor gains advance(&mut self, host), default nothing, with its trampoline. The shim sets advance to twin_advance() when the C extractor has one and refuses a twin without the entry point. advance_one_extractor() in libvmaf.c runs init_shared_rust_twin() before a pooled Rust twin's first advance, as the threaded flush does before its flush. test_rust_abi_layout checks the new offset and the entry point. * feat(rust): derive motion_rust's motion2 and motion3 frame by frame (ADR-2090, MI-1) motion_rust now implements Extractor::advance with the statements of integer_motion.c's window: frame i is derived once the SAD scores of frames 0 to max(i + 1, min_idx) are in, the stamp and the moving-average value are carried in a State as in VmafMotionWindowState, and the flush continues where the last advance stopped (motion_window_stamp(), motion_window_count_sads(), motion_window_derive(), vmaf_motion_window_advance(), vmaf_motion_window_flush(), ported in window.rs). A VMAF window over motion_rust completes before the flush, as over the C motion. test_rust_motion_window_incremental (suite rust): motion_rust on a context, both windows, 0, 2 and 3 worker threads; frame i final after frame i + 1 and before the flush, values unchanged after they became final, every SAD, motion2 and motion3 equal to the C motion bit for bit. It failed on the merged tree (flush -22: the inherited C advance and the Rust flush appended frame 0 twice), failed with scoring -22 at 2 threads before the engine initialised the pooled twin, and fails on a planted off-by-one in the Rust advance ("final too early"). Cargo tests cover advance + flush against the C derivation with the SAD scores in order, swapped in pairs and reversed, and short streams of 0 to 4 frames. rust_twin_diff.py --feature motion on netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45 EQUAL. State row T-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07. Carried onto the 2290 landing: the rebase note joins the motion-window-incremental fragment (ADR-2197), which also names read_pictures_owned(), the read helper's name since the WP6 split, as does the picture-ownership page. The Rust framework page is rewritten for the caveman lint and names the engine library (libvmafx) as the archive's only target. * docs(agents): restore the Metal motion twins' advance note on the AGENTS.d pages The incremental motion window registers .advance on integer_motion_metal and motion_v2_metal (ADR-2090). The rebase across the conversion of core/src/feature/metal/AGENTS.md into AGENTS.d pages took master's generated index and lost this pull request's two hunks of the old single file. The motion-fps-weight and motion3-v2 pages now name vmaf_motion_window_advance() next to the flush, the advance contract test, and integer_motion_v2_metal.mm as a path the motion3-v2 page covers; the index is regenerated. * docs(adr): leave the rendered ADR by-tag pages as master renders them Replayed onto the revert, this branch's earlier commits first add and then drop the ADR-2090 lines of docs/adr/by-tag/, which master's landing render already wrote (ADR-2197). The pages are master's again, so the pull request edits no rendered file. * ci(tidy): measure the translation units of the incremental motion window in the cpu lane The clang-tidy coverage rule on master requires every tracked translation unit to be read by a lane. The new and touched units of this pull request were measured in the dev container (scripts/dev/tidy-lane.sh --write --only ... cpu, clang-tidy 22.1.8): 0 findings, 0 uncited NOLINT; they join the cpu lane's measured sources. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
rc4/api-motion-incremental
branch
from
October 9, 2026 08:47
f6db3be to
e584c79
Compare
lusoris
added a commit
that referenced
this pull request
Oct 9, 2026
… tests The Cppcheck job fails on master since #2290: cppcheck 2.19.0 reports uninitvar on sad in test_motion_window_incremental.c at lines 256 and 289. The loops write every entry the check reads (the arrival order is a permutation of 0..n-1 for every stream length used) and stop early only on an error, after which nothing is read: a false positive that cppcheck cannot see through the permuted index. Both tables are zero-initialised, so they are whole on every path. No suppression; both findings are gone with the CI flags. The state row is T-CPPCHECK-MOTION-WINDOW-SAD-UNINITVAR-2026-10-09. Signed-off-by: Lusoris <lusoris@proton.me>
This was referenced Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RC4 (label
rc4):motion2/motion3are now final one frame after their frame instead of at the flush, with every value bit-identical to the flush-time derivation. A VMAFx window over a VMAF model completes once the frame after its last is scored, also while the producer stalls (#2138 liven_stats, #2238 OBS-ready API). Maintainer decision Q-038. Written on WP4 #2287 (the completion-thread rework, now on master). ADR-2090 records the contract and amends ADR-2074 decision 9.The rule. Frame
i'smotion2/motion3are final once the SAD scores of frames0tomax(i + 1, min_idx)are in the collector (min_idx1, or 2 withmotion_five_frame_window). The last frame, and every frame of a stream withmin_idxframes or fewer, are derived by the flush.Same statements, same order.
vmaf_motion_window_advance()(core/src/feature/integer_motion.c) derives every complete frame andvmaf_motion_window_flush()the rest. Both run upstream's per-frame statements (motion_flush_one(), unchanged) in index order, and carry upstream's stamp and moving-average value in aVmafMotionWindowState. A retried flush and an advance after the flush append nothing.Who calls it.
VmafFeatureExtractorgains an optionaladvance()callback.advance_extractors()incore/src/libvmaf.ccalls it:vmaf_engine_read_pictures(), now a wrapper aroundread_pictures_frame(), andvmaf_read_pictures_sycl());fence_for_read());vmaf_engine_advance()under the engine lock before each pass (advance_engine()incore/src/vmafx/window.c).Every call is serialised with the context's other engine calls and never runs on a worker. With worker threads the registered context of a pooled extractor is advanced; it never extracts, so it builds its dictionary on first use and is marked initialised, as the threaded flush does.
Frame-final signal (for WP4). The signal is the collector write itself. Without workers it happens before
vmaf_engine_read_pictures()returns, so beforevmafx_submit()wakes the completion thread. With workers, the worker's frame listener wakes the thread, which advances the engine before it probes. RequestWP4-1lists the WP4 files this commit touches.Twins. Every twin that derived at the flush now registers
.advanceon its own state:motion_v2_cuda,motion_v2_sycl,motion_v2_hipandmotion_v2_metal(both windows);integer_motion_metal(both windows);motion_cuda,motion_syclandmotion_hip(the five-frame window).The three-frame paths of
motion_{cuda,sycl,hip}andfloat_motionon every backend already wrote framei - 1while collecting framei, so they are unchanged.motion_cudareads its SADs back in batches of eight (ADR-0845), so its frames complete in batches (open rowT-CUDA-MOTION-BATCH-LIVE-LATENCY-2026-10-06, RC7).Found and fixed on the way.
vmaf_feature_score_at_index()answered-EINVALat once for a fed frame that had no collector slot yet (no score of that feature written, or an index past the vector capacity of 8) while a worker still held it. It now fences first (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06, ADR-2090 rule 5).Can master take it on its own? Yes, except for the window parts. The extractor and engine change does not depend on the VMAFx API: everything except
vmaf_engine_advance(),window.c, the window tests and the window docs could be split onto master.Landing (Q-083)
Re-lands as its own commit after #2640 (maintainer decision Q-307). This change first reached master inside the squash of #2615 (
7a3a7d0e9): the merge train's batch push of 17:41 failed after it had pushed the batch's stacked squashes to the pull request branches, and #2615 was recovered onto that stacked head. #2640 takes this change back out (all of it except ADR-2090 and its index fragment, which the rendered indexes link to); this pull request puts it back as its own squash, so its code, tests, docs, changelog fragments and state rows land under this title.reverts: #2640 (this pull request restores what #2640 took out; the files return to their content at
7a3a7d0e9, which the silent-revert guard reads as a rewind).Its base #2287 (WP4, window scores) is on master (
9926bfecc). The branch is its gated headd5f18fd30(own change plus the fixes below, first rebased onto mastercfdcfb51f) rebased onto master4007c9801, which has #2640 (c25412a38) and #2651 (thegolang.org/x/netfix the pre-push vulnerability check asked for): every file this PR changes equals its content atd5f18fd30, exceptscripts/ci/tidy-baseline-cpu.json, whose record of this PR's 9 translation units was measured again on this tree (master's baseline had moved). ADR-2090 and its index fragment stay as master has them (#2640 kept them). One commit restores master's rendered ADR by-tag pages, which the replay of the branch's commits had touched.core/src/AGENTS.d/rust-extractor-framework.mdconflicted with #2650's ADR-2795 note (twins copymergeandextend_name_dict): both bullets are kept, master's rewritten in the internal register the caveman lint asks for. ADR-2090 is Accepted (maintainer answer Q-088; status line andQ-088in its## References).ABI check against master:
python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/master->definition is an append-only successor of origin/master (0 additions);abi_version0.1.8 (no addition of its own).Carried fixes (each failed on the stack first)
fix(api): name the extractor as producer of the scores advance() writes: without ittest_vmafx_provenancetest_features_in_name_orderfails ("every score has a source":integer_motion2had source unknown). State rowT-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06.fix(api): answer pending for a fed frame whose motion3 is not derived yet:test_vmaf_frame_score_pending_until_finalfails without it ("frame 8 not pending",VMAFX_E_INVALID). State rowT-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.MI-1: Rust twins and
advance(maintainer answer Q-093)#2096 (
motion_rust) is on master, so this PR carries MI-1 (two commits):feat(rust): give Rust twins their own advance callback:VmafxRsTwin.advance(last field;VMAFX_RS_ABI_VERSIONstays 1, no release has shipped the Rust ABI;core/src/rust/include/vmafx_rs.hregenerated verbatim with cbindgen 0.29.4,scripts/dev/rust-abi-header.sh --checkup to date),Extractor::advance(default nothing) with its trampoline, the shim mapsadvancetotwin_advance()only when the C extractor has one, andadvance_one_extractor()initialises a pooled Rust twin before its first advance.feat(rust): derive motion_rust's motion2 and motion3 frame by frame:window.rsportsmotion_window_stamp(),motion_window_count_sads(),motion_window_derive(),vmaf_motion_window_advance()andvmaf_motion_window_flush(); new testtest_rust_motion_window_incremental(suiterust); state rowT-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07.Failing first on this tree: with the MI-1 sources reverted,
test_rust_motion_window_incrementalfails withmotion_rust (five 0, 0 threads): flush returned -22. With MI-1, on a Rust build (-Denable_rust_features=true, warnings as errors): build 0 warnings; fast + rust suites 427 OK, 0 fail;rust_twin_diff.py --feature motionon netflix, checker1, checker10, sparks10 and bbb4k at--threads 0,1,4: 45 EQUAL;cargo fmt,cargo clippy -D warningsclean;cargo test -p vmafx-fex -p vmafx-fex-motion17 passed.Rebase
Conflicts resolved per hunk:
core/src/libvmaf.ckeeps master'sread_pictures_owned()and caller-struct clearing (#2303), thenadvance_extractors(); the Metal motion twins keep master's anonymous-namespace style (#2222);docs/state.mdbyscripts/dev/resolve-state-md-conflict.py;core/test/meson.buildlinks the new tests throughvmaf_test_link. The rebase note is the fragmentdocs/rebase-notes.d/motion-window-incremental.md(ADR-2197), which also namesread_pictures_owned();core/src/AGENTS.d/rust-extractor-framework.mdis rewritten for the caveman lint and names the engine library (libvmafx) as the Rust archive's only target. Master splitcore/src/feature/metal/AGENTS.mdintoAGENTS.dpages while this PR was open; the rebase took master's generated index and lost this PR's two hunks of the old file, sodocs(agents): restore the Metal motion twins' advance noteputs them on themotion-fps-weightandmotion3-v2pages.Gates (rebased head)
4007c9801;-Db_lto=false,-j4, warnings as errors): build 0 warnings;--suite=fast426 OK, 1 failed:test_vmafx_window_cli, which passed 3 of 3 direct runs right after (the WP4 load flake under Known follow-ups; on the previous base the whole suite passed, 423 OK);test_gpu_picture_pool_uafon its own withMALLOC_PERTURB_=0: OK; codegen tests 207 passed;make test-netflix-golden GOLDEN_NINJA_JOBS=4280 passed, 3 skipped;preflight.sh --stage msvcismpass; affected suites: tooling 2648 passed, 6 skipped.-Denable_rust_features=true): build 0 warnings; fast + rust suites 434 OK, 0 fail (with master's Rust ADM view-distance change);rust_twin_diff.py --feature motionon netflix, checker1, checker10, sparks10 and bbb4k at--threads 0,1,4: 45 EQUAL;cargo fmt,cargo clippy -D warningsclean;cargo test -p vmafx-fex -p vmafx-fex-motion17 passed. The MI-1 failing-first result above was measured on the first rebase.--suite=fast489 OK, 0 fail, includingtest_cuda_motion_five_frame_windowand the motion, motion3 and motion_v2 parity tests. HIP (pinned ROCm 10.1.0 image, gfx1036) and SYCL (pinned oneAPI 2026.1.1, A380), on the lane before the restacks: builds 0 warnings, the five-frame window tests and every motion parity test passed; the motion twins are unchanged since (same blobs).praetorctl audit).Type
feat— new featuresycl/cuda/simd— backend-specific (every motion twin)Checklist
make format && make lintis green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, HISS audit).praetorctl audit: pass, HISS 6 within baseline 6, touched files clean.python3 scripts/ci/run_meson_test.py -- -C build-cpu --num-processes 6→ 387 OK, 0 fail, 2 skipped (test_cuda_parity_gate_default_runwithout CUDA,test_vmafx_api_abi_append_onlyon a base without a definition)./cross-backend-diffand the worst ULP is ≤ 2 — worst ULP 0: every cell is==(see Cross-backend numerical results)..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (core/test/test_motion_window_incremental.c, EUPL-1.2).!orBREAKING CHANGE:and the migration path is documented below. — not a breaking change: no public type or function changed; scores arrive earlier with the same values.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt—2090-motion-window-incremental, indexes regenerated.tidy (dev container, clang-tidy 22.1.8,
scripts/dev/tidy-lane.sh --only ...):integer_motion.c,integer_motion_v2.c,libvmaf.c,vmafx/window.c,test_motion_window_incremental.c,test_score_pooled_eagain.c,test_vmafx_window.c,test_metal_integer_motion_parity.candtest_metal_motion_v2_parity.c, with the headers they include (motion_window.h,feature_extractor.h): 0 findings.integer_motion_cuda.c,integer_motion_v2_cuda.candtest_cuda_motion_five_frame_window.c(+motion_five_frame_twin_parity.h): 0 findings.integer_motion_hip.c,integer_motion_v2_hip.candtest_hip_motion_five_frame_window.c: 0 findings.integer_motion_sycl.cpp,integer_motion_v2_sycl.cppandtest_sycl_motion_five_frame_window.c: 0 findings in these files. The lane reports 7 in the untouchedcore/include/libvmaf/picture.h, which are inherited: an untouched SYCL TU (integer_psnr_sycl.cpp) reproduces the same 7.scripts/dev/preflight.sh --stage msvcism: pass.core/test/test_win32_pthread_shim_contract.py: pass. ThreadSanitizer (-Db_sanitize=thread,halt_on_error=1):test_vmafx_window,test_motion_window_incremental,test_vmafx_window_liveandtest_score_pooled_eagaineach clean in 3 runs.Bug-status hygiene (ADR-0165)
docs/state.mdupdated:Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Golden gate (
make test-netflix-golden GOLDEN_NINJA_JOBS=6,core/build-goldenwith gcc, on the rebased head): 280 passed, 3 skipped, the same as the base. No golden assertion touched.Cross-backend numerical results
Per-frame values compared as IEEE bits (
--precision max). Keys: everymotionkey under four option sets (default;motion_five_frame_window;motion_moving_average+motion_blend_factor=0.5+motion_blend_offset=2; five-frame + moving average +motion_fps_weight=0.7+motion_max_val=5),motion_v2under two (default; five-frame + moving average),float_motion, plus thevmaf_v0.6.1score. Fixtures: Netflix golden pair (48 frames), both 1080p checkerboards,sparks10-bit, BBB 4K (200 frames).5c32bde1f(flush-time derivation)*_cudatwins (RTX 4090) vs CPU, named explicitly*_sycltwins (Arc A380, xe) vs CPU*_hiptwins (gfx1036) vs CPUscripts/ci/cross_backend_parity_gate.py --hold-exact: the cellsmotion,motion_debug,motion_mffw,motion_v2,motion_v2_mffwandfloat_motionhold==(max abs diff 0) on CUDA, SYCL and HIP, on four fixtures each (golden pair, both checkerboards, sparks 10-bit). The exact-twin matrix is unchanged.SYCL:
test_sycl_kernel_scratchpasses (no kernel changed, host code only).VMAF_SYCL_AOT_JOBS=6 meson test --suite sycl-aotpasses in 217.6 s, compiling every SYCL TU for the 19 default targets.Performance (if
perforfeat)No hot-path change.
advance()runs once per frame on the feeding thread and per completion pass. It reads the collector from the last derived frame on (amortised O(1) per frame) and appends two scores per frame. The CUDA readback batching is untouched.Deep-dive deliverables (ADR-0108)
Research digest — no digest needed: the contract is upstream's flush statements run frame by frame; the measurements are in ADR-2090 and this body.
Decision matrix —
docs/adr/2090-motion-window-incremental.md## Alternatives considered: where the advance runs (engine hook vs extract / collect vs worker listener), stateless recompute, keeping the flush, a lookahead in window code, CUDA per-frame readback, deduplicating the three-frame twin paths.AGENTS.mdinvariant note — new pagecore/src/feature/AGENTS.d/motion-window-incremental.md, plus updates to:core/src/AGENTS.d/picture-ownership-and-dispatch.md(engine advance points, fence on-EINVAL);core/src/AGENTS.d/vmafx-windows.md(completion thread advances first);motion-five-frame-windowpages;core/src/AGENTS.d/rust-extractor-framework.md(Rust twins andadvance, MI-1);core/src/feature/metal/AGENTS.md.Indexes regenerated.
Reproducer / smoke-test command — below.
CHANGELOG fragment —
changelog.d/changed/motion-window-incremental.md,changelog.d/changed/rust-motion-twin-advance.md,changelog.d/fixed/engine-read-fed-frame-fence.md, and the last sentence ofchangelog.d/added/api-window-scores.md(WP4's).Rebase note — fragment
docs/rebase-notes.d/motion-window-incremental.md(ADR-2197), "motion2 / motion3 derived frame by frame".User documentation:
docs/metrics/motion.md, new section "Whenmotion2andmotion3are final" and themotion_v2paragraph;docs/api/vmafx/windows.md, bullet "Features that read the next frame";docs/api/vmafx/index.md, flush sentence;vmafx_window_submitdoc incore/api/vmafx.toml, with the reference pages regenerated.Reproducer
Tests
test_motion_window_incremental(new, 5 cases)motionandmotion_v2(default and five-frame, serial and with 2 / 3 workers) andfloat_motionhave frameifinal after framei + 1is read and not before, the early values unchanged after the flush and equal to upstream's flush over their SADstest_vmafx_window(WP4's, 2 new cases)test_vmaf_window_completes_after_the_frame_after_last: serial, window [2, 5] ofvmaf_v0.6.1open after each submit up to frame 5, complete after the submit of frame 6, equal to a flushed session for every method.test_vmaf_window_completes_while_the_feeder_stalls: 2 workers, frames 0 to 6 submitted, then no call and no flush: completes, equal to a flushed session. The motion window test now asserts completion in the submit of frame 4, its frame afterlasttest_{cuda,sycl,hip}_motion_five_frame_window,test_metal_selftest_{integer_motion,motion_v2}(sharedmotion_five_frame_twin_parity.h)ifinal before the flush by readmax(i + 1, 2) + lag - 1(CPU 1; SYCL / HIP / Metal /motion_v2_cuda2;motion_cuda9, its readback batch), early values unchanged after the flush, every output==the CPUtest_motion_window_advance_contract.py(new, device-free, 8 cases)vmaf_motion_window_flush()(CPU, CUDA, SYCL, HIP, Metal) registers.advancecallingvmaf_motion_window_advance()on aVmafMotionWindowState; the engine's four advance points; the completion thread'sadvance_engine()test_score_pooled_eagainread_pictures(i), framei - 1pools, frameiis-EAGAIN; after the flush every frame pools (the old test pinned "nothing before the flush")Gates shown failing on planted defects (measured)
ifinal once the SAD of frameiis in (state->n_sadinstead ofn_sad - 1)test_motion_window_incremental(advance/flush vs upstream),test_vmafx_window(motion window),test_score_pooled_eagain,test_motion_five_frame_windowtest_vmafx_window(motion window step),test_motion_window_advance_contract.pytest_vmaf_window_completes_while_the_feeder_stalls(3 of 3),test_motion_window_advance_contract.pytest_motion_window_incrementalthreaded runs, 20 of 20.advanceremoved from the CUDA twinstest_cuda_motion_five_frame_window: "integer_motion2_mffw of frame 0 not final after frame 10" (motion), "... after frame 3" (motion_v2).advanceremoved from the SYCL twinstest_sycl_motion_five_frame_window(A380): frame 0 not final after frame 3.advanceremoved from the HIP twinstest_hip_motion_five_frame_window(gfx1036): frame 0 not final after frame 3.advance, state member or.stateremoved; the engine call removedtest_motion_window_advance_contract.pyplanted casesKnown follow-ups
test_vmafx_window_live'stest_unpaced_producer_is_held_backasserts a latency of at most two frame periods (33 ms) on the wall clock. Eight parallel copies at host load 35 fail it in 6 of 24 runs on master without this change and in 3 of 24 with it, so it is a load flake of the test, not of this PR;test_vmafx_window_cli, which drives the same binary, fails with it. Reported for a fix.test_metal_integer_motion_exact_contract.py,test_metal_motion_v2_exact_contract.py,test_motion_window_advance_contract.py, and the Metal parity tests as self-tests. Metal compile on macOS dispatched on this branch, not waited for (queue saturated): Tidy Metal https://github.com/VMAFx/vmafx/actions/runs/37482028525 (head4947774ba; compiles every Metal TU with-Denable_metal=enabledand measures the metal lane), libvmaf build matrix incl. themacOS Metalleg https://github.com/VMAFx/vmafx/actions/runs/37481464110 (939479a20, the same sources; the head differs only indocs/state.md).motion_rust(RC4 lane M): the shim copies the C descriptor (slot.fex = *c_fex) and would inherit.advance. RequestMI-1(rc4-requests) asks F for a trampoline or a cleared callback, and asks M to implement the frame-by-frame derivation.WP4-1):advance_engine()inwindow.cand the doc / test edits;test_callbacks_run_oncefails, becausevmafx_window_wait(UINT64_MAX)returnsPENDINGat once. It reproduces with this lane's window changes removed.WP9-2): drop the motion caveat of WP9-1;n_statsofvmafis live.motion_cudalive latency: frames complete per readback batch of 8 (T-CUDA-MOTION-BATCH-LIVE-LATENCY-2026-10-06, RC7).make test-netflix-golden-arm64) not run: the NEON code is SAD only and unchanged, and the derivation is shared scalar C.