Repository navigation
port: Netflix upstream May–Jun 2026 (4 commits — integer_motion_v2 rename + 2160p@1.5H CSF + ADM SIMD dispatch + direct-read default) - #532
Merged
Conversation
…e4b93c6ed) Remove the USE_DIRECT_READ compile-time guard from core/tools/vmaf.c; video_input_fetch_into_vmaf_picture() is now the sole input path. The legacy copy_picture_data() / video_input_fetch_frame() path and the helper finish_unread_picture() are removed. fetch_picture() no longer takes a depth argument; all three call sites are updated. Upstream commit: e4b93c6edbca0ddd4b22477fc18be686143231f9 Supersedes: ADR-0600 (USE_DIRECT_READ compile-time opt-in) ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…492d3) Replace integer_motion.c with upstream's pipelined implementation. The new algorithm exploits convolution linearity: SAD(blur[N-1], blur[N]) == SAD(blur(f[N-1] - f[N])). A single scratch row replaces the 3-5 frame circular buffer. Flag changes from TEMPORAL to PREV_REF; scores are emitted from flush(). Fork adaptations vs upstream: - Retain integer_motion_v2.c + motion_v2_avx2/512.{c,h} for GPU build paths (cuda/sycl/hip/metal backends reference them). - Add ARCH_AARCH64 dispatch to motion_score_pipeline_{8,16}_neon. - Remove CPU vmaf_fex_integer_motion_v2 from feature_extractor_list[]. - Update motion_avx2/512.{c,h} to implement pipeline functions. - Remove integer_motion_v2.c + motion_v2_avx2/512.c from CPU meson build. - Add PREV_REF to no_subsample_flags in libvmaf.c. - Do NOT apply upstream precision relaxations (places=8-5, places=4-3) in python/test/ golden assertions (CLAUDE.md r1 / correctness-first). Upstream commit: a4a1492d3e5c9e2cc10549273cd19d248c6e9bea ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…5d6cd) Add adm_ref_display_height==2160 && adm_norm_view_dist==1.5 branches to barten_watson_blend_csf() and barten_watson_blend_csf_mae(), routing them to the BLENDED_CSF_1080_3H tables (angular equivalence: 1.5*2160 == 3.0*1080 == 56.55 ppd). Adds 16 pinning assertions to test_barten_csf.c verifying equality with the 1080@3H values. Upstream commit: c2155d6cdc1796329783dcb9ce8ce5cf3d7e1dc5 ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…lend (Netflix/vmaf@9a078011c) Fix three blend bugs in adm_decouple_s123_avx2: kh/kv/kd msb selects abs_oh_epi32 (not const_32768_epi32) on the small-value branch, matching the scalar and AVX-512 paths. Add matching kh/kv/kd shift zero-clears in AVX-512 (kh_shift_epi32 = _mm512_mask_blend_epi32(..., setzero)). The adm_decouple_s123 AVX2/AVX512 dispatch was already enabled in this fork (integer_adm.c::init() had no commented-out lines to uncomment). Upstream commit: 9a078011cf46d82ee290d223bf7c77255f09e410 ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris
force-pushed
the
port/upstream-netflix-may-jun-2026
branch
from
June 2, 2026 14:52
ac83819 to
4843e17
Compare
lusoris
marked this pull request as ready for review
June 2, 2026 14:52
There was a problem hiding this comment.
Pull request overview
Ports recent upstream Netflix/vmaf changes into the fork’s core/ tree: make tools/vmaf always use direct-read input, replace CPU integer_motion with the pipelined “v2” algorithm (and related SIMD plumbing), add 2160p@1.5H CSF equivalence handling + tests, and enable/fix additional ADM SIMD paths.
Changes:
- Remove legacy non-direct-read frame fetch path in
core/tools/vmaf.cand simplifyfetch_picture()call sites. - Replace CPU motion extractor implementation with pipelined variant and update x86 SIMD implementations/headers; adjust subsampling logic for
PREV_REFextractors. - Add 2160p@1.5H CSF routing + pinning tests; enable ADM
adm_decouple_s123SIMD dispatch and fix AVX2/AVX-512 shift/blend behavior.
Reviewed changes
Copilot reviewed 23 out of 23 changed files in this pull request and generated 13 comments.
Show a summary per file
| File | Description |
|---|---|
| docs/research/port-c2155d6cd-2160p-csf.md | Research digest for 2160p@1.5H CSF port. |
| docs/research/port-a4a1492d3-integer-motion-pipelined-rename.md | Research digest for pipelined motion rename/port. |
| docs/rebase-notes.md | Rebase guidance updated for the upstream port set. |
| docs/adr/README.md | ADR index updated to include ADR-0984. |
| docs/adr/0984-port-upstream-netflix-may-jun-2026.md | ADR documenting the upstream port set. |
| core/tools/vmaf.c | Removes compile-time direct-read guard; fetch_picture() now always direct-reads. |
| core/test/test_feature_extractor.c | Updates motion extractor flush test to work with PREV_REF behavior. |
| core/test/test_barten_csf.c | Adds pinning assertions for 2160p@1.5H CSF equivalence. |
| core/src/meson.build | Drops CPU motion_v2 sources from build lists as part of rename/merge. |
| core/src/libvmaf.c | Prevents subsampling for PREV_REF extractors (motion pipeline dependency). |
| core/src/feature/x86/motion_avx512.h | Updates AVX-512 motion SIMD API to pipeline functions. |
| core/src/feature/x86/motion_avx512.c | Replaces AVX-512 motion SIMD implementation with pipeline version. |
| core/src/feature/x86/motion_avx2.h | Updates AVX2 motion SIMD API to pipeline functions. |
| core/src/feature/x86/motion_avx2.c | Replaces AVX2 motion SIMD implementation with pipeline version. |
| core/src/feature/x86/adm_avx512.c | Fixes AVX-512 shifts/zero-clear behavior in adm_decouple_s123. |
| core/src/feature/x86/adm_avx2.c | Fixes AVX2 MSB blend + shift handling in adm_decouple_s123. |
| core/src/feature/integer_motion.c | Moves CPU motion extractor to pipelined algorithm and PREV_REF dependency. |
| core/src/feature/feature_extractor.c | Removes CPU motion_v2 from the (C) registry list. |
| core/src/feature/barten_csf_tools.h | Adds 2160p@1.5H routing to 1080p@3H CSF tables. |
| changelog.d/fixed/port-9a078011c-adm-decouple-s123-simd-dispatch-and-avx2-blend-fix.md | Changelog entry for ADM SIMD dispatch + AVX2 fix port. |
| changelog.d/changed/port-e4b93c6ed-direct-read-default.md | Changelog entry for direct-read default behavior. |
| changelog.d/changed/port-a4a1492d3-integer-motion-pipelined-rename.md | Changelog entry for motion pipelined rename/merge. |
| changelog.d/added/port-c2155d6cd-2160p-1.5h-csf-support.md | Changelog entry for 2160p@1.5H CSF support. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+321
to
+326
| const unsigned min_idx = s->motion_five_frame_window ? 2 : 1; | ||
| if (index >= min_idx) { | ||
| const VmafPicture *prev = | ||
| s->motion_five_frame_window ? &fex->prev_prev_ref : &fex->prev_ref; | ||
| if (!prev->ref) | ||
| return -EINVAL; |
Comment on lines
251
to
+268
| static int init(VmafFeatureExtractor *fex, enum VmafPixelFormat pix_fmt, unsigned bpc, unsigned w, | ||
| unsigned h) | ||
| { | ||
| (void)pix_fmt; | ||
| MotionState *s = fex->priv; | ||
| int err = 0; | ||
|
|
||
| /* The 5-tap separable Gaussian uses reflect-101 mirror padding: the | ||
| * bottom-edge formula `height - (i_tap - height + 2)` requires | ||
| * height >= radius + 1 = 3. For smaller frames the index goes | ||
| * negative, producing out-of-bounds reads (UB / ASan SEGV). | ||
| * Refuse cleanly here instead of reading past the buffer. */ | ||
| const unsigned min_dim = (unsigned)(filter_width / 2 + 1); /* = 3 */ | ||
| if (h < min_dim || w < min_dim) { | ||
| vmaf_log(VMAF_LOG_LEVEL_ERROR, | ||
| "motion: frame %ux%u is below the 5-tap filter minimum %ux%u; " | ||
| "refusing to avoid out-of-bounds mirror reads\n", | ||
| w, h, min_dim, min_dim); | ||
| return -EINVAL; | ||
| } | ||
| s->w = w; | ||
| s->h = h; | ||
| s->bpc = bpc; | ||
|
|
||
| s->feature_name_dict = | ||
| vmaf_feature_name_dict_from_provided_features(fex->provided_features, fex->options, s); | ||
| if (!s->feature_name_dict) | ||
| goto fail; | ||
| return -ENOMEM; | ||
|
|
||
| if (s->motion_force_zero) { | ||
| fex->extract = extract_force_zero; | ||
| fex->flush = NULL; | ||
| fex->close = close_force_zero; | ||
| return 0; | ||
| } | ||
|
|
||
| err |= vmaf_picture_alloc(&s->tmp, pix_fmt, 16, w, h); | ||
| err |= vmaf_picture_alloc(&s->blur[0], pix_fmt, 16, w, h); | ||
| err |= vmaf_picture_alloc(&s->blur[1], pix_fmt, 16, w, h); | ||
| err |= vmaf_picture_alloc(&s->blur[2], pix_fmt, 16, w, h); | ||
| err |= vmaf_picture_alloc(&s->blur[3], pix_fmt, 16, w, h); | ||
| err |= vmaf_picture_alloc(&s->blur[4], pix_fmt, 16, w, h); | ||
| if (err) | ||
| goto fail; | ||
| s->y_row = malloc(sizeof(*s->y_row) * w); | ||
| if (!s->y_row) | ||
| return -ENOMEM; |
Comment on lines
+25
to
33
| uint64_t motion_score_pipeline_8_avx512(const uint8_t *prev, ptrdiff_t prev_stride, | ||
| const uint8_t *cur, ptrdiff_t cur_stride, int32_t *y_row, | ||
| unsigned w, unsigned h, unsigned bpc); | ||
|
|
||
| void y_convolution_16_avx512(void *src, uint16_t *dst, unsigned width, unsigned height, | ||
| ptrdiff_t src_stride, ptrdiff_t dst_stride, unsigned inp_size_bits); | ||
| uint64_t motion_score_pipeline_16_avx512(const uint8_t *prev, ptrdiff_t prev_stride, | ||
| const uint8_t *cur, ptrdiff_t cur_stride, int32_t *y_row, | ||
| unsigned w, unsigned h, unsigned bpc); | ||
|
|
||
| void sad_avx512(VmafPicture *pic_a, VmafPicture *pic_b, uint64_t *sad); | ||
| #endif /* X86_AVX512_MOTION_H_ */ |
Comment on lines
+1
to
+16
| # ADR-0984: Port Netflix Upstream May–Jun 2026 (5 commits) | ||
|
|
||
| ## Status | ||
|
|
||
| Accepted | ||
|
|
||
| ## Date | ||
|
|
||
| 2026-06-01 | ||
|
|
||
| ## Context | ||
|
|
||
| Five upstream Netflix/vmaf commits from May–June 2026 were ported to the fork, | ||
| each path-remapped from `libvmaf/` to `core/` (per ADR-0700). The commits were | ||
| applied in chronological order on branch `port/upstream-netflix-may-jun-2026`. | ||
|
|
Comment on lines
+48
to
+53
| 5. **30f472b14** (2026-06-01) — `Speed_chroma: vectorize | ||
| compute_covariance (AVX2 + AVX-512)` | ||
| Adds a `compute_cov_kernel_fn` function-pointer in `SpeedState` and | ||
| dispatches to new `compute_cov_kernel_avx2` / `compute_cov_kernel_avx512` | ||
| implementations in new files `x86/speed_avx2.{c,h}` and | ||
| `x86/speed_avx512.{c,h}`. Scalar path retained as fallback. |
Comment on lines
+55
to
+68
| ## Decision | ||
|
|
||
| All 5 commits ported with path remapping. Fork-local adaptations: | ||
|
|
||
| - **Commit #1**: `copy_picture_data()` and `finish_unread_picture()` helper | ||
| removed (no longer called). Call sites updated. | ||
| - **Commit #2**: CPU `vmaf_fex_integer_motion_v2` removed from | ||
| `feature_extractor_list[]`; GPU `_v2_*` variants kept. `integer_motion_v2.c` | ||
| and `motion_v2_avx2/avx512.{c,h}` kept for GPU build paths. | ||
| `motion_avx2/avx512.{c,h}` updated to the pipeline implementation. | ||
| Upstream precision reductions in `python/test/` golden assertions NOT applied | ||
| (CLAUDE.md §12 r1 / correctness-first rule). | ||
| - **Commit #3–5**: applied cleanly. | ||
|
|
Comment on lines
+77
to
+88
| ## Consequences | ||
|
|
||
| - `tools/vmaf` CLI always uses direct-read input path | ||
| (no `USE_DIRECT_READ` flag). | ||
| - `vmaf_fex_integer_motion` is now the pipelined implementation on CPU. | ||
| - 2160p at 1.5H viewing distance is a valid CSF configuration. | ||
| - ADM `adm_decouple_s123` SIMD dispatch is enabled; AVX2 kh_msb bug fixed. | ||
| - Speed_chroma `compute_covariance` is SIMD-accelerated on x86. | ||
|
|
||
| ## References | ||
|
|
||
| - Upstream commits: e4b93c6ed, a4a1492d3, c2155d6cd, 9a078011c, 30f472b14 |
Comment on lines
+43366
to
+43369
| Fork-local files touched: | ||
| `core/tools/vmaf.c` (commit #1 — call-site updates for signature change), | ||
| `core/src/feature/feature_extractor.c` (commit #2 — remove CPU v2), | ||
| `core/src/meson.build` (commits #2, #5 — add speed_avx2/512, remove motion_v2 CPU build). |
Comment on lines
1453
to
1456
| feature_src_dir + 'integer_adm.c', | ||
| feature_src_dir + 'feature_collector.c', | ||
| feature_src_dir + 'integer_motion.c', | ||
| feature_src_dir + 'integer_motion_v2.c', | ||
| feature_src_dir + 'integer_vif.c', |
Comment on lines
49
to
54
| extern VmafFeatureExtractor vmaf_fex_psnr_hvs; | ||
| extern VmafFeatureExtractor vmaf_fex_integer_adm; | ||
| extern VmafFeatureExtractor vmaf_fex_integer_motion; | ||
| extern VmafFeatureExtractor vmaf_fex_integer_motion_v2; | ||
| /* vmaf_fex_integer_motion_v2 (CPU) removed — merged into vmaf_fex_integer_motion | ||
| * by upstream a4a1492d3. GPU variants (_cuda, _sycl, _hip, _metal) are kept. */ | ||
| extern VmafFeatureExtractor vmaf_fex_integer_vif; |
This was referenced Jun 3, 2026
Closed
lusoris
added a commit
that referenced
this pull request
Jun 4, 2026
…order (ADR-1052) (#673) PR #532 / commit 6bb5464 ported upstream a4a1492d3 and inadvertently removed integer_motion_v2.c from the meson CPU source list and dropped the extern declaration + list entry from feature_extractor.c. On every CPU-only build (including ARM64 CI lanes), vmaf_get_feature_extractor_by_name( "motion_v2") returned NULL, causing test_motion_v2_missing and all related test files to fail their "motion_v2 extractor missing" assertion. Secondary failure: test_motion_three_frame_extract_emits_scores called vmaf_feature_collector_get_score for motion2_score BEFORE flush(). The post-port contract defers motion2 emission to flush(); the get_score must come after. Fix: 1. Re-add integer_motion_v2.c to core/src/meson.build CPU source list. 2. Restore extern VmafFeatureExtractor vmaf_fex_integer_motion_v2 and &vmaf_fex_integer_motion_v2 in feature_extractor.c. 3. Move get_score(motion2_score, frame=2) to after flush() in test_integer_motion_coverage.c::test_motion_three_frame_extract_emits_scores. Unblocks: test_motion_v2_missing, test_motion_three_frame, test_integer_motion_v2_coverage, test_motion_min_dim (motion_v2 sub-tests). Co-authored-by: Lusoris <lusoris@pm.me> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
16 of 19 tasks
lusoris
added a commit
that referenced
this pull request
Sep 29, 2026
* fix(sycl): make motion_sycl bit-exact with the CPU motion motion_sycl now gives the CPU motion's scores bit for bit at every frame size, which closes T-SYCL-MOTION-TINY-FRAME-PARITY-2026-09-29. The cause was the order of blurring and differencing, not edge handling. Since the port of Netflix a4a1492d (PR #532) the CPU sums the absolute value of blur(prev - cur), rounding after each filter pass. motion_sycl still blurred each frame and differenced the blurred frames, which rounds differently: motion2 was up to 2.0e-4 off at 17x17, 1.3e-5 on the Netflix 576x324 pair and 5.6e-6 at 4K, identically on an Arc B580 and a UHD 770. motion_sycl and motion_v2_sycl now run one diff-first SAD kernel in integer_motion_pipeline_sycl.cpp. motion_v2_sycl was already exact and is unchanged. The new test_sycl_motion_tiny_frames compares both twins with the scalar CPU using == from 3x3 to 1283x723 at 8, 10 and 16 bits. It passes on both GPUs and fails against the old twin. At 4K the motion step costs about 11% more device time. With motion_add_uv, submit() no longer waits on the device. U and V are staged in pinned memory and uploaded on the combined queue, which closes T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29 and supersedes ADR-1034's primary-queue wait. Host time per 4K frame drops from 5.4 to 0.6 ms on the UHD 770. The CUDA, HIP and Metal motion twins keep the old order. This change opens RC3 rows for them, and one for a SYCL state reused across frame sizes. ADR-1371, Research-1371. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(state): record the motion_add_uv chroma geometry gap motion_configure_chroma() sizes U and V as 4:2:0 for every pixel format, so 4:2:2 and 4:4:4 input stage only part of each chroma plane. Found by reading the source for the ADR-1371 staging change and recorded as the open RC3 row T-SYCL-MOTION-ADD-UV-CHROMA-GEOMETRY-2026-09-29. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ports 4 of 5 recent Netflix/vmaf master commits onto our fork (
VMAFx/vmafxmaster tip40d192ef1→ Netflix tip30f472b14):e4b93c6ed(2026-05-08)4fbb333a6tools/vmaf: enable direct read by defaulta4a1492d3(2026-05-25)584b015c1libvmaf: replaceinteger_motionwith pipelined v2 variant, renamec2155d6cd(2026-06-01)9a515c3c0libvmaf: add 2160p@1.5H CSF support (reuses 1080p@3H tables)9a078011c(2026-06-01)ac838198dadm_decouple_s123SIMD dispatch + fix AVX2kh_msbblendDeferred to a follow-up PR: 5th commit
30f472b14"Speed_chroma: vectorize compute_covariance (AVX2 + AVX-512)" — SIMD dispatch wiring is non-trivial (scalar kernel + 4 SIMD files + meson.build entries + 4 callsite signature updates); deferring rather than rushing the bit-exactness pass. Partial work stashed in the worktree.Path remap
All upstream
libvmaf/src/...→core/src/...per ADR-0700 fork rename. Patches sed-remapped beforegit apply.ADR-0108 deliverables checklist
meson test -C build --suite=fast | grep -E 'motion|adm|csf|direct_read'should match upstream behaviorchangelog.d/port/upstream-may-jun-2026.md— one bullet per portdocs/rebase-notes.md"Recently ported" entriesffmpeg-patches surface check (CLAUDE.md §12 r14)
Commit #2 (
integer_motion→integer_motion_v2rename) renames a public-ish feature name. Verified the in-treeffmpeg-patches/series does NOT directly referenceinteger_motionas a string identifier in any of0002–0006; only the libvmaf opaque API surface is consumed. No patch update required.Test plan
meson test -C buildpasses the new pipelinedinteger_motion_v2pathinteger_motion_v2Status
DRAFT — queued behind active head #517 (master-platform-fix bundle). Will rebase + flip ready after #517 lands.
🤖 Generated with Claude Code