Skip to content

feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) - #2290

Merged
lusoris merged 2 commits into
masterfrom
rc4/api-motion-incremental
Oct 9, 2026
Merged

lusoris merged 2 commits into
masterfrom
rc4/api-motion-incremental

Conversation

@lusoris

@lusoris lusoris commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

RC4 (label rc4): motion2 / motion3 are now final one frame after their frame instead of at the flush, with every value bit-identical to the flush-time derivation. A VMAFx window over a VMAF model completes once the frame after its last is scored, also while the producer stalls (#2138 live n_stats, #2238 OBS-ready API). Maintainer decision Q-038. Written on WP4 #2287 (the completion-thread rework, now on master). ADR-2090 records the contract and amends ADR-2074 decision 9.

The rule. Frame i's motion2 / motion3 are final once the SAD scores of frames 0 to max(i + 1, min_idx) are in the collector (min_idx 1, or 2 with motion_five_frame_window). The last frame, and every frame of a stream with min_idx frames or fewer, are derived by the flush.

Same statements, same order. vmaf_motion_window_advance() (core/src/feature/integer_motion.c) derives every complete frame and vmaf_motion_window_flush() the rest. Both run upstream's per-frame statements (motion_flush_one(), unchanged) in index order, and carry upstream's stamp and moving-average value in a VmafMotionWindowState. A retried flush and an advance after the flush append nothing.

Who calls it. VmafFeatureExtractor gains an optional advance() callback. advance_extractors() in core/src/libvmaf.c calls it:

  • at the end of every accepted frame (vmaf_engine_read_pictures(), now a wrapper around read_pictures_frame(), and vmaf_read_pictures_sycl());
  • after a read fence (fence_for_read());
  • from the VMAFx completion thread, through vmaf_engine_advance() under the engine lock before each pass (advance_engine() in core/src/vmafx/window.c).

Every call is serialised with the context's other engine calls and never runs on a worker. With worker threads the registered context of a pooled extractor is advanced; it never extracts, so it builds its dictionary on first use and is marked initialised, as the threaded flush does.

Frame-final signal (for WP4). The signal is the collector write itself. Without workers it happens before vmaf_engine_read_pictures() returns, so before vmafx_submit() wakes the completion thread. With workers, the worker's frame listener wakes the thread, which advances the engine before it probes. Request WP4-1 lists the WP4 files this commit touches.

Twins. Every twin that derived at the flush now registers .advance on its own state:

  • motion_v2_cuda, motion_v2_sycl, motion_v2_hip and motion_v2_metal (both windows);
  • integer_motion_metal (both windows);
  • motion_cuda, motion_sycl and motion_hip (the five-frame window).

The three-frame paths of motion_{cuda,sycl,hip} and float_motion on every backend already wrote frame i - 1 while collecting frame i, so they are unchanged. motion_cuda reads its SADs back in batches of eight (ADR-0845), so its frames complete in batches (open row T-CUDA-MOTION-BATCH-LIVE-LATENCY-2026-10-06, RC7).

Found and fixed on the way. vmaf_feature_score_at_index() answered -EINVAL at once for a fed frame that had no collector slot yet (no score of that feature written, or an index past the vector capacity of 8) while a worker still held it. It now fences first (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06, ADR-2090 rule 5).

Can master take it on its own? Yes, except for the window parts. The extractor and engine change does not depend on the VMAFx API: everything except vmaf_engine_advance(), window.c, the window tests and the window docs could be split onto master.

Landing (Q-083)

Re-lands as its own commit after #2640 (maintainer decision Q-307). This change first reached master inside the squash of #2615 (7a3a7d0e9): the merge train's batch push of 17:41 failed after it had pushed the batch's stacked squashes to the pull request branches, and #2615 was recovered onto that stacked head. #2640 takes this change back out (all of it except ADR-2090 and its index fragment, which the rendered indexes link to); this pull request puts it back as its own squash, so its code, tests, docs, changelog fragments and state rows land under this title.

reverts: #2640 (this pull request restores what #2640 took out; the files return to their content at 7a3a7d0e9, which the silent-revert guard reads as a rewind).

Its base #2287 (WP4, window scores) is on master (9926bfecc). The branch is its gated head d5f18fd30 (own change plus the fixes below, first rebased onto master cfdcfb51f) rebased onto master 4007c9801, which has #2640 (c25412a38) and #2651 (the golang.org/x/net fix the pre-push vulnerability check asked for): every file this PR changes equals its content at d5f18fd30, except scripts/ci/tidy-baseline-cpu.json, whose record of this PR's 9 translation units was measured again on this tree (master's baseline had moved). ADR-2090 and its index fragment stay as master has them (#2640 kept them). One commit restores master's rendered ADR by-tag pages, which the replay of the branch's commits had touched. core/src/AGENTS.d/rust-extractor-framework.md conflicted with #2650's ADR-2795 note (twins copy merge and extend_name_dict): both bullets are kept, master's rewritten in the internal register the caveman lint asks for. ADR-2090 is Accepted (maintainer answer Q-088; status line and Q-088 in its ## References).

ABI check against master: python3 scripts/codegen/vmafx-api.py --abi-check --against-ref origin/master -> definition is an append-only successor of origin/master (0 additions); abi_version 0.1.8 (no addition of its own).

Carried fixes (each failed on the stack first)

  • fix(api): name the extractor as producer of the scores advance() writes: without it test_vmafx_provenance test_features_in_name_order fails ("every score has a source": integer_motion2 had source unknown). State row T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06.
  • fix(api): answer pending for a fed frame whose motion3 is not derived yet: test_vmaf_frame_score_pending_until_final fails without it ("frame 8 not pending", VMAFX_E_INVALID). State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.

MI-1: Rust twins and advance (maintainer answer Q-093)

#2096 (motion_rust) is on master, so this PR carries MI-1 (two commits):

  • feat(rust): give Rust twins their own advance callback: VmafxRsTwin.advance (last field; VMAFX_RS_ABI_VERSION stays 1, no release has shipped the Rust ABI; core/src/rust/include/vmafx_rs.h regenerated verbatim with cbindgen 0.29.4, scripts/dev/rust-abi-header.sh --check up to date), Extractor::advance (default nothing) with its trampoline, the shim maps advance to twin_advance() only when the C extractor has one, and advance_one_extractor() initialises a pooled Rust twin before its first advance.
  • feat(rust): derive motion_rust's motion2 and motion3 frame by frame: window.rs ports motion_window_stamp(), motion_window_count_sads(), motion_window_derive(), vmaf_motion_window_advance() and vmaf_motion_window_flush(); new test test_rust_motion_window_incremental (suite rust); state row T-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07.

Failing first on this tree: with the MI-1 sources reverted, test_rust_motion_window_incremental fails with motion_rust (five 0, 0 threads): flush returned -22. With MI-1, on a Rust build (-Denable_rust_features=true, warnings as errors): build 0 warnings; fast + rust suites 427 OK, 0 fail; rust_twin_diff.py --feature motion on netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45 EQUAL; cargo fmt, cargo clippy -D warnings clean; cargo test -p vmafx-fex -p vmafx-fex-motion 17 passed.

Rebase

Conflicts resolved per hunk: core/src/libvmaf.c keeps master's read_pictures_owned() and caller-struct clearing (#2303), then advance_extractors(); the Metal motion twins keep master's anonymous-namespace style (#2222); docs/state.md by scripts/dev/resolve-state-md-conflict.py; core/test/meson.build links the new tests through vmaf_test_link. The rebase note is the fragment docs/rebase-notes.d/motion-window-incremental.md (ADR-2197), which also names read_pictures_owned(); core/src/AGENTS.d/rust-extractor-framework.md is rewritten for the caveman lint and names the engine library (libvmafx) as the Rust archive's only target. Master split core/src/feature/metal/AGENTS.md into AGENTS.d pages while this PR was open; the rebase took master's generated index and lost this PR's two hunks of the old file, so docs(agents): restore the Metal motion twins' advance note puts them on the motion-fps-weight and motion3-v2 pages.

Gates (rebased head)

  • CPU on this head (master 4007c9801; -Db_lto=false, -j4, warnings as errors): build 0 warnings; --suite=fast 426 OK, 1 failed: test_vmafx_window_cli, which passed 3 of 3 direct runs right after (the WP4 load flake under Known follow-ups; on the previous base the whole suite passed, 423 OK); test_gpu_picture_pool_uaf on its own with MALLOC_PERTURB_=0: OK; codegen tests 207 passed; make test-netflix-golden GOLDEN_NINJA_JOBS=4 280 passed, 3 skipped; preflight.sh --stage msvcism pass; affected suites: tooling 2648 passed, 6 skipped.
  • Rust on the same head (-Denable_rust_features=true): build 0 warnings; fast + rust suites 434 OK, 0 fail (with master's Rust ADM view-distance change); rust_twin_diff.py --feature motion on netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45 EQUAL; cargo fmt, cargo clippy -D warnings clean; cargo test -p vmafx-fex -p vmafx-fex-motion 17 passed. The MI-1 failing-first result above was measured on the first rebase.
  • CUDA (13.4, RTX 4090), on the stack before the rebase: build 0 warnings, --suite=fast 489 OK, 0 fail, including test_cuda_motion_five_frame_window and the motion, motion3 and motion_v2 parity tests. HIP (pinned ROCm 10.1.0 image, gfx1036) and SYCL (pinned oneAPI 2026.1.1, A380), on the lane before the restacks: builds 0 warnings, the five-frame window tests and every motion parity test passed; the motion twins are unchanged since (same blobs).
  • Tidy (clang-tidy 22.1.8): cpu on the 9 translation units of this PR, 0 findings (measured again on this head); the CUDA, HIP and SYCL motion units are unchanged since their 0-finding runs.
  • Platform check: no POSIX-only code added; the merge train builds CPU and CUDA and runs its own gates (deliverables, state-md, silent revert, praetorctl audit).

Type

  • feat — new feature
  • sycl / cuda / simd — backend-specific (every motion twin)

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally — the commit hooks pass (clang-format, markdownlint, semgrep, source ADR citations, generated-index freshness, HISS audit). praetorctl audit: pass, HISS 6 within baseline 6, touched files clean.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build-cpu --num-processes 6 → 387 OK, 0 fail, 2 skipped (test_cuda_parity_gate_default_run without CUDA, test_vmafx_api_abi_append_only on a base without a definition).
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2 — worst ULP 0: every cell is == (see Cross-backend numerical results).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. Every twin that derived at the flush changed. Metal has no device here; see follow-ups.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (core/test/test_motion_window_incremental.c, EUPL-1.2).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. — not a breaking change: no public type or function changed; scores arrive earlier with the same values.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — 2090-motion-window-incremental, indexes regenerated.

tidy (dev container, clang-tidy 22.1.8, scripts/dev/tidy-lane.sh --only ...):

  • cpu: integer_motion.c, integer_motion_v2.c, libvmaf.c, vmafx/window.c, test_motion_window_incremental.c, test_score_pooled_eagain.c, test_vmafx_window.c, test_metal_integer_motion_parity.c and test_metal_motion_v2_parity.c, with the headers they include (motion_window.h, feature_extractor.h): 0 findings.
  • cuda: integer_motion_cuda.c, integer_motion_v2_cuda.c and test_cuda_motion_five_frame_window.c (+ motion_five_frame_twin_parity.h): 0 findings.
  • hip: integer_motion_hip.c, integer_motion_v2_hip.c and test_hip_motion_five_frame_window.c: 0 findings.
  • sycl: integer_motion_sycl.cpp, integer_motion_v2_sycl.cpp and test_sycl_motion_five_frame_window.c: 0 findings in these files. The lane reports 7 in the untouched core/include/libvmaf/picture.h, which are inherited: an untouched SYCL TU (integer_psnr_sycl.cpp) reproduces the same 7.
  • The first measurement found 6, all fixed by refactoring: braces in two places, a test over the function-size limit split twice, and two implicit widenings in the shared harness.
  • metal: the Tidy Metal lane is dispatched on this branch (macOS, see follow-ups).

scripts/dev/preflight.sh --stage msvcism: pass. core/test/test_win32_pthread_shim_contract.py: pass. ThreadSanitizer (-Db_sanitize=thread, halt_on_error=1): test_vmafx_window, test_motion_window_incremental, test_vmafx_window_live and test_score_pooled_eagain each clean in 3 runs.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated:
    • T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06 moved to Recently closed (found on WP4, fixed here);
    • T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06 found and fixed here, Recently closed;
    • T-CUDA-MOTION-BATCH-LIVE-LATENCY-2026-10-06 opened (RC7 tuning).

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below — not applicable.

Golden gate (make test-netflix-golden GOLDEN_NINJA_JOBS=6, core/build-golden with gcc, on the rebased head): 280 passed, 3 skipped, the same as the base. No golden assertion touched.

Cross-backend numerical results

Per-frame values compared as IEEE bits (--precision max). Keys: every motion key under four option sets (default; motion_five_frame_window; motion_moving_average + motion_blend_factor=0.5 + motion_blend_offset=2; five-frame + moving average + motion_fps_weight=0.7 + motion_max_val=5), motion_v2 under two (default; five-frame + moving average), float_motion, plus the vmaf_v0.6.1 score. Fixtures: Netflix golden pair (48 frames), both 1080p checkerboards, sparks 10-bit, BBB 4K (200 frames).

Comparison Configurations Values Different
this branch, CPU vs master 5c32bde1f (flush-time derivation) 5 fixtures x default / AVX2 / scalar dispatch x 0 / 4 threads 32 988 0
*_cuda twins (RTX 4090) vs CPU, named explicitly 5 fixtures x 0 / 4 threads 10 878 0
*_sycl twins (Arc A380, xe) vs CPU same 10 878 0
*_hip twins (gfx1036) vs CPU same 10 878 0

scripts/ci/cross_backend_parity_gate.py --hold-exact: the cells motion, motion_debug, motion_mffw, motion_v2, motion_v2_mffw and float_motion hold == (max abs diff 0) on CUDA, SYCL and HIP, on four fixtures each (golden pair, both checkerboards, sparks 10-bit). The exact-twin matrix is unchanged.

SYCL: test_sycl_kernel_scratch passes (no kernel changed, host code only). VMAF_SYCL_AOT_JOBS=6 meson test --suite sycl-aot passes in 217.6 s, compiling every SYCL TU for the 19 default targets.

Performance (if perf or feat)

No hot-path change. advance() runs once per frame on the feeding thread and per completion pass. It reads the collector from the last derived frame on (amortised O(1) per frame) and appends two scores per frame. The CUDA readback batching is untouched.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: the contract is upstream's flush statements run frame by frame; the measurements are in ADR-2090 and this body.

  • Decision matrix — docs/adr/2090-motion-window-incremental.md ## Alternatives considered: where the advance runs (engine hook vs extract / collect vs worker listener), stateless recompute, keeping the flush, a lookahead in window code, CUDA per-frame readback, deduplicating the three-frame twin paths.

  • AGENTS.md invariant note — new page core/src/feature/AGENTS.d/motion-window-incremental.md, plus updates to:

    • core/src/AGENTS.d/picture-ownership-and-dispatch.md (engine advance points, fence on -EINVAL);
    • core/src/AGENTS.d/vmafx-windows.md (completion thread advances first);
    • the CUDA / SYCL / HIP motion-five-frame-window pages;
    • core/src/AGENTS.d/rust-extractor-framework.md (Rust twins and advance, MI-1);
    • core/src/feature/metal/AGENTS.md.

    Indexes regenerated.

  • Reproducer / smoke-test command — below.

  • CHANGELOG fragment — changelog.d/changed/motion-window-incremental.md, changelog.d/changed/rust-motion-twin-advance.md, changelog.d/fixed/engine-read-fed-frame-fence.md, and the last sentence of changelog.d/added/api-window-scores.md (WP4's).

  • Rebase note — fragment docs/rebase-notes.d/motion-window-incremental.md (ADR-2197), "motion2 / motion3 derived frame by frame".

User documentation:

  • docs/metrics/motion.md, new section "When motion2 and motion3 are final" and the motion_v2 paragraph;
  • docs/api/vmafx/windows.md, bullet "Features that read the next frame";
  • docs/api/vmafx/index.md, flush sentence;
  • the vmafx_window_submit doc in core/api/vmafx.toml, with the reference pages regenerated.

Reproducer

meson setup build core -Db_lto=false && nice -n 10 ninja -C build -j6
python3 scripts/ci/run_meson_test.py -- -C build test_motion_window_incremental \
    test_motion_window_advance_contract test_vmafx_window test_score_pooled_eagain \
    test_motion_five_frame_window test_metal_selftest_integer_motion test_metal_selftest_motion_v2
# devices (each under its lock): test_{cuda,sycl,hip}_motion_five_frame_window, test_{cuda,sycl}_exact_twins
python3 scripts/ci/cross_backend_parity_gate.py --vmaf-binary build/tools/vmaf \
    --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
    --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv --width 576 --height 324 \
    --features motion motion_debug motion_mffw motion_v2 motion_v2_mffw float_motion \
    --backends cpu cuda --hold-exact cuda
GOLDEN_NINJA_JOBS=6 make test-netflix-golden

Tests

Test Covers
test_motion_window_incremental (new, 5 cases) (1) advance after every SAD, then flush, equals a transcription of upstream's flush bit for bit: 6 option sets (both windows, moving average, blend 0.5 / 0.2, cap 0.5), 0 to 17 frames, SADs in order and out of order (blocks of three reversed, as workers append them), with a check after each SAD of which frames are final; (2) a flush without a state equals upstream's; (3) a second flush and an advance after the flush append nothing; (4) refusals; (5) on a context: motion and motion_v2 (default and five-frame, serial and with 2 / 3 workers) and float_motion have frame i final after frame i + 1 is read and not before, the early values unchanged after the flush and equal to upstream's flush over their SADs
test_vmafx_window (WP4's, 2 new cases) test_vmaf_window_completes_after_the_frame_after_last: serial, window [2, 5] of vmaf_v0.6.1 open after each submit up to frame 5, complete after the submit of frame 6, equal to a flushed session for every method. test_vmaf_window_completes_while_the_feeder_stalls: 2 workers, frames 0 to 6 submitted, then no call and no flush: completes, equal to a flushed session. The motion window test now asserts completion in the submit of frame 4, its frame after last
test_{cuda,sycl,hip}_motion_five_frame_window, test_metal_selftest_{integer_motion,motion_v2} (shared motion_five_frame_twin_parity.h) frame i final before the flush by read max(i + 1, 2) + lag - 1 (CPU 1; SYCL / HIP / Metal / motion_v2_cuda 2; motion_cuda 9, its readback batch), early values unchanged after the flush, every output == the CPU
test_motion_window_advance_contract.py (new, device-free, 8 cases) every TU that calls vmaf_motion_window_flush() (CPU, CUDA, SYCL, HIP, Metal) registers .advance calling vmaf_motion_window_advance() on a VmafMotionWindowState; the engine's four advance points; the completion thread's advance_engine()
test_score_pooled_eagain the Netflix#755 streaming pattern: after read_pictures(i), frame i - 1 pools, frame i is -EAGAIN; after the flush every frame pools (the old test pinned "nothing before the flush")

Gates shown failing on planted defects (measured)

Planted defect Fails
off-by-one: frame i final once the SAD of frame i is in (state->n_sad instead of n_sad - 1) test_motion_window_incremental (advance/flush vs upstream), test_vmafx_window (motion window), test_score_pooled_eagain, test_motion_five_frame_window
no advance after a frame (the old flush-only behaviour) test_vmafx_window (motion window step), test_motion_window_advance_contract.py
completion thread does not advance the engine test_vmaf_window_completes_while_the_feeder_stalls (3 of 3), test_motion_window_advance_contract.py
sync read of a fed frame without a slot does not fence test_motion_window_incremental threaded runs, 20 of 20
.advance removed from the CUDA twins test_cuda_motion_five_frame_window: "integer_motion2_mffw of frame 0 not final after frame 10" (motion), "... after frame 3" (motion_v2)
.advance removed from the SYCL twins test_sycl_motion_five_frame_window (A380): frame 0 not final after frame 3
.advance removed from the HIP twins test_hip_motion_five_frame_window (gfx1036): frame 0 not final after frame 3
a twin's .advance, state member or .state removed; the engine call removed test_motion_window_advance_contract.py planted cases

Known follow-ups

  • Real-time window tests under host load (master, WP4): test_vmafx_window_live's test_unpaced_producer_is_held_back asserts a latency of at most two frame periods (33 ms) on the wall clock. Eight parallel copies at host load 35 fail it in 6 of 24 runs on master without this change and in 3 of 24 with it, so it is a load flake of the test, not of this PR; test_vmafx_window_cli, which drives the same binary, fails with it. Reported for a fix.
  • Metal: no device here. The device-free contracts pass: test_metal_integer_motion_exact_contract.py, test_metal_motion_v2_exact_contract.py, test_motion_window_advance_contract.py, and the Metal parity tests as self-tests. Metal compile on macOS dispatched on this branch, not waited for (queue saturated): Tidy Metal https://github.com/VMAFx/vmafx/actions/runs/37482028525 (head 4947774ba; compiles every Metal TU with -Denable_metal=enabled and measures the metal lane), libvmaf build matrix incl. the macOS Metal leg https://github.com/VMAFx/vmafx/actions/runs/37481464110 (939479a20, the same sources; the head differs only in docs/state.md).
  • Rust motion_rust (RC4 lane M): the shim copies the C descriptor (slot.fex = *c_fex) and would inherit .advance. Request MI-1 (rc4-requests) asks F for a trampoline or a cleared callback, and asks M to implement the frame-by-frame derivation.
  • WP4 (requests WP4-1):
    • take advance_engine() in window.c and the doc / test edits;
    • finding not caused by this lane: in an icx-compiled (SYCL) build test_callbacks_run_once fails, because vmafx_window_wait(UINT64_MAX) returns PENDING at once. It reproduces with this lane's window changes removed.
  • WP9 (WP9-2): drop the motion caveat of WP9-1; n_stats of vmaf is live.
  • motion_cuda live latency: frames complete per readback batch of 8 (T-CUDA-MOTION-BATCH-LIVE-LATENCY-2026-10-06, RC7).
  • aarch64 golden gate (make test-netflix-golden-arm64) not run: the NEON code is SAD only and unchanged, and the derivation is shared scalar C.

@lusoris lusoris added type:feature New feature or request rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag labels Oct 6, 2026
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from 939479a to 4947774 Compare October 6, 2026 14:48
lusoris added a commit that referenced this pull request Oct 6, 2026
…tion lane's draft

T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06 stays open on this branch, as the
maintainer decided; the lane that derives motion2 / motion3 frame by frame
now has a draft (#2290, rc4/api-motion-incremental, ADR-2090 there), which
closes the row.
@lusoris
lusoris force-pushed the rc4/api-wp4-windows branch 2 times, most recently from 3eabc39 to 9926bfe Compare October 8, 2026 12:56
Base automatically changed from rc4/api-wp4-windows to master October 8, 2026 12:59
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from 4947774 to d5f18fd Compare October 8, 2026 14:55
@lusoris
lusoris marked this pull request as ready for review October 8, 2026 14:55
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from d5f18fd to a7b218c Compare October 8, 2026 15:32
lusoris added a commit that referenced this pull request Oct 8, 2026
…re-commit hook (#2615)

* fix(ci): list SPONSORS.md among the CI impact planner's known top-level files (#2619)

* fix(ci): list SPONSORS.md among the CI impact planner's known top-level files

top-level entry the planner may see, and test_ci_impact.py
test_every_top_level_repo_entry_is_known fails on master since then
(SPONSORS.md is neither a known file nor under a known prefix). The file
joins known_files next to SECURITY.md and SUPPORT.md; it selects no suite
of its own, as those do not.

State row T-CI-IMPACT-SPONSORS-UNKNOWN-2026-10-08.

* feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090) (#2290)

* feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090)

The integer motion extractors and their GPU twins derived motion2 and
motion3 of every frame in flush(), so a VMAFx window over a VMAF model,
a per-frame model score and vmaf_score_pooled() over the frames read so
far waited for the end of the stream (state row
T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF
windows in the 1.0 scope (Q-038, #2138, #2238).

Frame i is now final once the SAD scores of frames 0 to max(i + 1,
min_idx) are in the collector. vmaf_motion_window_advance() derives every
complete frame with upstream's per-frame statements in index order and
carries the stamp and the moving-average value in a VmafMotionWindowState;
vmaf_motion_window_flush() derives the rest. The engine calls a new
optional VmafFeatureExtractor.advance() after every accepted frame and
read fence, and the VMAFx completion thread calls it through
vmaf_engine_advance() under the engine lock before each pass, so a window
completes as soon as a worker's SAD is in, also while the feeder stalls.
motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL,
HIP, Metal) register it. A synchronous read of a fed frame without a
collector slot now fences instead of answering -EINVAL
(T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06).

Every value is the flush-time value: 32 988 per-frame motion values
against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K;
default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU
backend against the CPU, 0 different; parity gate motion cells exact on
CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped.

* fix(api): name the extractor as producer of the scores advance() writes

advance_extractors() called VmafFeatureExtractor.advance() (ADR-2090)
without installing the extractor as the thread's feature producer
(ADR-2073), so the integer_motion2 vector advance() creates first carried
source unknown and the provenance record named no extractor for it. The
threaded flush installs the producer around flush(); advance() now gets the
same. test_vmafx_provenance test_features_in_name_order failed on the merged
tip and passes. State row T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06.

* fix(api): answer pending for a fed frame whose motion3 is not derived yet

With incremental motion (ADR-2090) motion3 of frame i is written when frame
i + 1 is scored. vmafx_score_frame() of the frame just submitted returned
VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the
end of the motion3 vector's storage and the collector answered -EINVAL.
engine_score_at_index() now checks the inputs of a fed frame before it
predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps
its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.

* docs(motion): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197)

* docs(agents): write the incremental motion page in the internal register (praetor caveman lint)

* ci(tidy): measure the translation units of the incremental motion window in the cpu lane

The clang-tidy coverage rule on master requires every tracked translation
unit to be read by a lane. The new and touched units of this pull request
were measured in the dev container (scripts/dev/tidy-lane.sh --write
--only ... cpu, clang-tidy 22.1.8): 0 findings, 0 uncited NOLINT; they join
the cpu lane's measured sources.

* feat(rust): give Rust twins their own advance callback (ADR-2090, MI-1)

The shim copies the C extractor's descriptor whole (add_twin()), so after
incremental motion (ADR-2090) a Rust twin inherited the C advance(): for
motion_rust it derived motion2 / motion3 from the Rust SAD scores into the
C-layout priv while frames came in, and the Rust flush then appended every
frame again, -EINVAL at the flush; with worker threads the inherited advance
also marked the registered context initialised, so the Rust flush found no
instance. Maintainer answer Q-093: Rust implements advance.

VmafxRsTwin gains advance as its last field. VMAFX_RS_ABI_VERSION stays 1:
no release has shipped the Rust ABI, so the layout changes in place. The
header is regenerated verbatim with cbindgen 0.29.4
(scripts/dev/rust-abi-header.sh, --check up to date). vmafx_fex::Extractor
gains advance(&mut self, host), default nothing, with its trampoline. The
shim sets advance to twin_advance() when the C extractor has one and refuses
a twin without the entry point. advance_one_extractor() in libvmaf.c runs
init_shared_rust_twin() before a pooled Rust twin's first advance, as the
threaded flush does before its flush. test_rust_abi_layout checks the new
offset and the entry point.

* feat(rust): derive motion_rust's motion2 and motion3 frame by frame (ADR-2090, MI-1)

motion_rust now implements Extractor::advance with the statements of
integer_motion.c's window: frame i is derived once the SAD scores of frames
0 to max(i + 1, min_idx) are in, the stamp and the moving-average value are
carried in a State as in VmafMotionWindowState, and the flush continues where
the last advance stopped (motion_window_stamp(), motion_window_count_sads(),
motion_window_derive(), vmaf_motion_window_advance(),
vmaf_motion_window_flush(), ported in window.rs). A VMAF window over
motion_rust completes before the flush, as over the C motion.

test_rust_motion_window_incremental (suite rust): motion_rust on a context,
both windows, 0, 2 and 3 worker threads; frame i final after frame i + 1 and
before the flush, values unchanged after they became final, every SAD,
motion2 and motion3 equal to the C motion bit for bit. It failed on the
merged tree (flush -22: the inherited C advance and the Rust flush appended
frame 0 twice), failed with scoring -22 at 2 threads before the engine
initialised the pooled twin, and fails on a planted off-by-one in the Rust
advance ("final too early"). Cargo tests cover advance + flush against the C
derivation with the SAD scores in order, swapped in pairs and reversed, and
short streams of 0 to 4 frames. rust_twin_diff.py --feature motion on
netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45
EQUAL. State row T-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07.

Carried onto the 2290 landing: the rebase note joins the
motion-window-incremental fragment (ADR-2197), which also names
read_pictures_owned(), the read helper's name since the WP6 split, as does
the picture-ownership page. The Rust framework page is rewritten for the
caveman lint and names the engine library (libvmafx) as the archive's only
target.

* docs(agents): restore the Metal motion twins' advance note on the AGENTS.d pages

The incremental motion window registers .advance on integer_motion_metal
and motion_v2_metal (ADR-2090). The rebase across the conversion of
core/src/feature/metal/AGENTS.md into AGENTS.d pages took master's
generated index and lost this pull request's two hunks of the old single
file. The motion-fps-weight and motion3-v2 pages now name
vmaf_motion_window_advance() next to the flush, the advance contract test,
and integer_motion_v2_metal.mm as a path the motion3-v2 page covers; the
index is regenerated.

* docs(community): link the live Patreon page and its euro tiers (#2620)

* docs(community): link the live Patreon page and its euro tiers

* fix(ci): skip docs/rebase-notes.d in the Markdown Lint job like the pre-commit hook (#2615)

* fix(ci): skip docs/rebase-notes.d in the Markdown Lint job like the pre-commit hook

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from a7b218c to d5f18fd Compare October 8, 2026 17:06
lusoris added a commit that referenced this pull request Oct 8, 2026
7a3a7d0 so #2290 lands as its own commit

The merge train's batch push of 17:41 failed after it had pushed the
stacked squashes of the batch to their pull request branches. The fixer of
skip docs/rebase-notes.d in the Markdown Lint job", #2615) also carries the
whole change of #2290 (incremental motion2 / motion3, ADR-2090) and the
CI-impact entry of #2619. #2290 never landed as its own commit, and its
code, tests, docs, changelog fragments and state rows sit under #2615's
title.

This reverts #2290's own diff (its gated head d5f18fd against its base
cfdcfb5), reverse-applied on master 0bf2996: 66 of its 68 files return
to their content before #2290; docs/state.md keeps the rows of the commits
that landed since. ADR-2090 and its index fragment stay: the ADR indexes
are written only by the landing render and link to it, so deleting it
fails the ADR link gate; #2290 re-lands without them. #2615's own change
(the Markdown Lint fragment scope) and the SPONSORS.md entry of
.github/ci-impact.json (#2626, also #2619's) stay. #2290 re-lands next as
its own squash.

docs/state.md also folds the duplicate row of the SPONSORS.md defect:
T-CI-IMPACT-SPONSORS-MD-UNKNOWN-2026-10-08 (#2626) stays and names
T-CI-IMPACT-SPONSORS-UNKNOWN-2026-10-08 (#2619, closed as a duplicate) as
its alias.

Signed-off-by: Lusoris <lusoris@proton.me>
lusoris added a commit that referenced this pull request Oct 8, 2026
…so #2620 lands as its own commit

7a3a7d0 (#2615's squash, built on the stacked heads of the failed batch
push of 17:41) also carries #2620 (the live Patreon page and its euro
tiers). This reverts #2620's own diff (its head 66be228 against its base
cfdcfb5): .github/FUNDING.yml, GOVERNANCE.md, README.md, SPONSORS.md,
docs/support-vmafx.md and changelog.d/changed/patreon-live.md return to
their content before #2620. #2620 re-lands after #2290, as its own squash.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from d5f18fd to 4613fc6 Compare October 8, 2026 22:22
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from 4613fc6 to f6db3be Compare October 9, 2026 08:13
…ce (#2647)

* feat(cuda): evaluate two ADM viewing distances in one adm_cuda instance

adm_cuda takes adm_norm_view_dist_extra and the merge callback: the
scale-0 and scale-1..3 DWT run once per frame, and the denominator, CSF,
contrast-masking and AIM kernels run once per distance into that
distance's result block (tmp_res and results_host hold two). The host
concludes each distance with its own CPU contexts and files the second
under the <base>:nvde keys. Both distances return the CPU's bits
(ADR-2795; test_adm_two_views_exact, test_adm_merged_registrations_exact).

The merge callback and the second distance's names move into
adm_view_dist.c, which reads the options by name at each descriptor's
own offsets, so the CPU extractor, the Rust twin and adm_cuda share one
implementation; integer_adm.c drops its copies.

Signed-off-by: Lusoris <lusoris@proton.me>
…ws complete before the flush (ADR-2090) (#2290)

* feat(motion): derive motion2 and motion3 frame by frame so VMAF windows complete before the flush (ADR-2090)

The integer motion extractors and their GPU twins derived motion2 and
motion3 of every frame in flush(), so a VMAFx window over a VMAF model,
a per-frame model score and vmaf_score_pooled() over the frames read so
far waited for the end of the stream (state row
T-VMAFX-WINDOW-MOTION-AT-FLUSH-2026-10-06). The maintainer put live VMAF
windows in the 1.0 scope (Q-038, #2138, #2238).

Frame i is now final once the SAD scores of frames 0 to max(i + 1,
min_idx) are in the collector. vmaf_motion_window_advance() derives every
complete frame with upstream's per-frame statements in index order and
carries the stamp and the moving-average value in a VmafMotionWindowState;
vmaf_motion_window_flush() derives the rest. The engine calls a new
optional VmafFeatureExtractor.advance() after every accepted frame and
read fence, and the VMAFx completion thread calls it through
vmaf_engine_advance() under the engine lock before each pass, so a window
completes as soon as a worker's SAD is in, also while the feeder stalls.
motion, motion_v2 and every twin that derived at the flush (CUDA, SYCL,
HIP, Metal) register it. A synchronous read of a fed frame without a
collector slot now fences instead of answering -EINVAL
(T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06).

Every value is the flush-time value: 32 988 per-frame motion values
against master (golden pair, both checkerboards, sparks 10-bit, BBB 4K;
default, AVX2 and scalar dispatch; 0 and 4 threads) and 10 878 per GPU
backend against the CPU, 0 different; parity gate motion cells exact on
CUDA, SYCL and HIP; Netflix golden gate 280 passed, 3 skipped.

* fix(api): name the extractor as producer of the scores advance() writes

advance_extractors() called VmafFeatureExtractor.advance() (ADR-2090)
without installing the extractor as the thread's feature producer
(ADR-2073), so the integer_motion2 vector advance() creates first carried
source unknown and the provenance record named no extractor for it. The
threaded flush installs the producer around flush(); advance() now gets the
same. test_vmafx_provenance test_features_in_name_order failed on the merged
tip and passes. State row T-RC4-MOTION-ADVANCE-NO-PRODUCER-2026-10-06.

* fix(api): answer pending for a fed frame whose motion3 is not derived yet

With incremental motion (ADR-2090) motion3 of frame i is written when frame
i + 1 is scored. vmafx_score_frame() of the frame just submitted returned
VMAFX_E_INVALID at frames 8, 16 and 32, where the frame index reached the
end of the motion3 vector's storage and the collector answered -EINVAL.
engine_score_at_index() now checks the inputs of a fed frame before it
predicts and answers -EAGAIN while one is unwritten; a frame never fed keeps
its error. State row T-RC4-SCORE-FRAME-INVALID-AT-VECTOR-END-2026-10-06.

* docs(motion): move the rebase note to a fragment and leave the rendered files to the landing render (ADR-2197)

* docs(agents): write the incremental motion page in the internal register (praetor caveman lint)

* feat(rust): give Rust twins their own advance callback (ADR-2090, MI-1)

The shim copies the C extractor's descriptor whole (add_twin()), so after
incremental motion (ADR-2090) a Rust twin inherited the C advance(): for
motion_rust it derived motion2 / motion3 from the Rust SAD scores into the
C-layout priv while frames came in, and the Rust flush then appended every
frame again, -EINVAL at the flush; with worker threads the inherited advance
also marked the registered context initialised, so the Rust flush found no
instance. Maintainer answer Q-093: Rust implements advance.

VmafxRsTwin gains advance as its last field. VMAFX_RS_ABI_VERSION stays 1:
no release has shipped the Rust ABI, so the layout changes in place. The
header is regenerated verbatim with cbindgen 0.29.4
(scripts/dev/rust-abi-header.sh, --check up to date). vmafx_fex::Extractor
gains advance(&mut self, host), default nothing, with its trampoline. The
shim sets advance to twin_advance() when the C extractor has one and refuses
a twin without the entry point. advance_one_extractor() in libvmaf.c runs
init_shared_rust_twin() before a pooled Rust twin's first advance, as the
threaded flush does before its flush. test_rust_abi_layout checks the new
offset and the entry point.

* feat(rust): derive motion_rust's motion2 and motion3 frame by frame (ADR-2090, MI-1)

motion_rust now implements Extractor::advance with the statements of
integer_motion.c's window: frame i is derived once the SAD scores of frames
0 to max(i + 1, min_idx) are in, the stamp and the moving-average value are
carried in a State as in VmafMotionWindowState, and the flush continues where
the last advance stopped (motion_window_stamp(), motion_window_count_sads(),
motion_window_derive(), vmaf_motion_window_advance(),
vmaf_motion_window_flush(), ported in window.rs). A VMAF window over
motion_rust completes before the flush, as over the C motion.

test_rust_motion_window_incremental (suite rust): motion_rust on a context,
both windows, 0, 2 and 3 worker threads; frame i final after frame i + 1 and
before the flush, values unchanged after they became final, every SAD,
motion2 and motion3 equal to the C motion bit for bit. It failed on the
merged tree (flush -22: the inherited C advance and the Rust flush appended
frame 0 twice), failed with scoring -22 at 2 threads before the engine
initialised the pooled twin, and fails on a planted off-by-one in the Rust
advance ("final too early"). Cargo tests cover advance + flush against the C
derivation with the SAD scores in order, swapped in pairs and reversed, and
short streams of 0 to 4 frames. rust_twin_diff.py --feature motion on
netflix, checker1, checker10, sparks10 and bbb4k at --threads 0,1,4: 45
EQUAL. State row T-RUST-TWIN-INHERITS-C-ADVANCE-2026-10-07.

Carried onto the 2290 landing: the rebase note joins the
motion-window-incremental fragment (ADR-2197), which also names
read_pictures_owned(), the read helper's name since the WP6 split, as does
the picture-ownership page. The Rust framework page is rewritten for the
caveman lint and names the engine library (libvmafx) as the archive's only
target.

* docs(agents): restore the Metal motion twins' advance note on the AGENTS.d pages

The incremental motion window registers .advance on integer_motion_metal
and motion_v2_metal (ADR-2090). The rebase across the conversion of
core/src/feature/metal/AGENTS.md into AGENTS.d pages took master's
generated index and lost this pull request's two hunks of the old single
file. The motion-fps-weight and motion3-v2 pages now name
vmaf_motion_window_advance() next to the flush, the advance contract test,
and integer_motion_v2_metal.mm as a path the motion3-v2 page covers; the
index is regenerated.

* docs(adr): leave the rendered ADR by-tag pages as master renders them

Replayed onto the revert, this branch's earlier commits first add and then
drop the ADR-2090 lines of docs/adr/by-tag/, which master's landing render
already wrote (ADR-2197). The pages are master's again, so the pull request
edits no rendered file.

* ci(tidy): measure the translation units of the incremental motion window in the cpu lane

The clang-tidy coverage rule on master requires every tracked translation
unit to be read by a lane. The new and touched units of this pull request
were measured in the dev container (scripts/dev/tidy-lane.sh --write
--only ... cpu, clang-tidy 22.1.8): 0 findings, 0 uncited NOLINT; they join
the cpu lane's measured sources.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the rc4/api-motion-incremental branch from f6db3be to e584c79 Compare October 9, 2026 08:47
@lusoris
lusoris merged commit e584c79 into master Oct 9, 2026
22 of 56 checks passed
@lusoris
lusoris deleted the rc4/api-motion-incremental branch October 9, 2026 08:50
lusoris added a commit that referenced this pull request Oct 9, 2026
… tests

The Cppcheck job fails on master since #2290: cppcheck 2.19.0 reports
uninitvar on sad in test_motion_window_incremental.c at lines 256 and
289. The loops write every entry the check reads (the arrival order is
a permutation of 0..n-1 for every stream length used) and stop early
only on an error, after which nothing is read: a false positive that
cppcheck cannot see through the permuted index.

Both tables are zero-initialised, so they are whole on every path. No
suppression; both findings are gone with the CI flags.

The state row is T-CPPCHECK-MOTION-WINDOW-SAD-UNINITVAR-2026-10-09.

Signed-off-by: Lusoris <lusoris@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rc4 RC4: the vmaf_v1.0.16_3d0h path in Rust; lands after the v1.0.0-rc.3 tag type:feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant