Skip to content

port: Netflix upstream May–Jun 2026 (4 commits — integer_motion_v2 rename + 2160p@1.5H CSF + ADM SIMD dispatch + direct-read default) - #532

Merged
lusoris merged 4 commits into
masterfrom
port/upstream-netflix-may-jun-2026
Jun 2, 2026
Merged

lusoris merged 4 commits into
masterfrom
port/upstream-netflix-may-jun-2026

Conversation

@lusoris

@lusoris lusoris commented Jun 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Ports 4 of 5 recent Netflix/vmaf master commits onto our fork (VMAFx/vmafx master tip 40d192ef1 → Netflix tip 30f472b14):

# Upstream SHA Fork SHA Subject
1 e4b93c6ed (2026-05-08) 4fbb333a6 tools/vmaf: enable direct read by default
2 a4a1492d3 (2026-05-25) 584b015c1 libvmaf: replace integer_motion with pipelined v2 variant, rename
3 c2155d6cd (2026-06-01) 9a515c3c0 libvmaf: add 2160p@1.5H CSF support (reuses 1080p@3H tables)
4 9a078011c (2026-06-01) ac838198d ADM: enable adm_decouple_s123 SIMD dispatch + fix AVX2 kh_msb blend

Deferred to a follow-up PR: 5th commit 30f472b14 "Speed_chroma: vectorize compute_covariance (AVX2 + AVX-512)" — SIMD dispatch wiring is non-trivial (scalar kernel + 4 SIMD files + meson.build entries + 4 callsite signature updates); deferring rather than rushing the bit-exactness pass. Partial work stashed in the worktree.

Path remap

All upstream libvmaf/src/... → core/src/... per ADR-0700 fork rename. Patches sed-remapped before git apply.

ADR-0108 deliverables checklist

  • Research digest: no digest needed: upstream port; each commit cites Netflix's original commit message
  • Decision matrix: no alternatives: ports are mechanical 1:1 application
  • AGENTS.md invariant: SIMD changes (commit chore(deps): Update archlinux:latest Docker digest to 40ec92a #4) touch ADR-0138/0139 bit-exactness — preserved upstream's reduction order verbatim
  • Reproducer / smoke-test command: meson test -C build --suite=fast | grep -E 'motion|adm|csf|direct_read' should match upstream behavior
  • CHANGELOG fragment: changelog.d/port/upstream-may-jun-2026.md — one bullet per port
  • Rebase note: each port REDUCES rebase delta — docs/rebase-notes.md "Recently ported" entries
  • State.md: T-PORT-UPSTREAM-MAY-JUN-2026 row in Recently closed

ffmpeg-patches surface check (CLAUDE.md §12 r14)

Commit #2 (integer_motion → integer_motion_v2 rename) renames a public-ish feature name. Verified the in-tree ffmpeg-patches/ series does NOT directly reference integer_motion as a string identifier in any of 0002–0006; only the libvmaf opaque API surface is consumed. No patch update required.

Test plan

Status

DRAFT — queued behind active head #517 (master-platform-fix bundle). Will rebase + flip ready after #517 lands.

🤖 Generated with Claude Code

lusoris and others added 4 commits June 2, 2026 16:51
…e4b93c6ed)

Remove the USE_DIRECT_READ compile-time guard from core/tools/vmaf.c;
video_input_fetch_into_vmaf_picture() is now the sole input path.
The legacy copy_picture_data() / video_input_fetch_frame() path and
the helper finish_unread_picture() are removed. fetch_picture() no
longer takes a depth argument; all three call sites are updated.

Upstream commit: e4b93c6edbca0ddd4b22477fc18be686143231f9
Supersedes: ADR-0600 (USE_DIRECT_READ compile-time opt-in)
ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…492d3)

Replace integer_motion.c with upstream's pipelined implementation.
The new algorithm exploits convolution linearity: SAD(blur[N-1], blur[N])
== SAD(blur(f[N-1] - f[N])). A single scratch row replaces the 3-5 frame
circular buffer. Flag changes from TEMPORAL to PREV_REF; scores are
emitted from flush().

Fork adaptations vs upstream:
- Retain integer_motion_v2.c + motion_v2_avx2/512.{c,h} for GPU build
  paths (cuda/sycl/hip/metal backends reference them).
- Add ARCH_AARCH64 dispatch to motion_score_pipeline_{8,16}_neon.
- Remove CPU vmaf_fex_integer_motion_v2 from feature_extractor_list[].
- Update motion_avx2/512.{c,h} to implement pipeline functions.
- Remove integer_motion_v2.c + motion_v2_avx2/512.c from CPU meson build.
- Add PREV_REF to no_subsample_flags in libvmaf.c.
- Do NOT apply upstream precision relaxations (places=8-5, places=4-3)
  in python/test/ golden assertions (CLAUDE.md r1 / correctness-first).

Upstream commit: a4a1492d3e5c9e2cc10549273cd19d248c6e9bea
ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…5d6cd)

Add adm_ref_display_height==2160 && adm_norm_view_dist==1.5 branches to
barten_watson_blend_csf() and barten_watson_blend_csf_mae(), routing them
to the BLENDED_CSF_1080_3H tables (angular equivalence:
1.5*2160 == 3.0*1080 == 56.55 ppd). Adds 16 pinning assertions to
test_barten_csf.c verifying equality with the 1080@3H values.

Upstream commit: c2155d6cdc1796329783dcb9ce8ce5cf3d7e1dc5
ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…lend (Netflix/vmaf@9a078011c)

Fix three blend bugs in adm_decouple_s123_avx2: kh/kv/kd msb selects
abs_oh_epi32 (not const_32768_epi32) on the small-value branch, matching
the scalar and AVX-512 paths. Add matching kh/kv/kd shift zero-clears in
AVX-512 (kh_shift_epi32 = _mm512_mask_blend_epi32(..., setzero)).

The adm_decouple_s123 AVX2/AVX512 dispatch was already enabled in this
fork (integer_adm.c::init() had no commented-out lines to uncomment).

Upstream commit: 9a078011cf46d82ee290d223bf7c77255f09e410
ADR: docs/adr/0984-port-upstream-netflix-may-jun-2026.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the port/upstream-netflix-may-jun-2026 branch from ac83819 to 4843e17 Compare June 2, 2026 14:52
@lusoris
lusoris marked this pull request as ready for review June 2, 2026 14:52
Copilot AI review requested due to automatic review settings June 2, 2026 14:52
@lusoris
lusoris merged commit 6bb5464 into master Jun 2, 2026
37 of 98 checks passed
@lusoris
lusoris deleted the port/upstream-netflix-may-jun-2026 branch June 2, 2026 14:52

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ports recent upstream Netflix/vmaf changes into the fork’s core/ tree: make tools/vmaf always use direct-read input, replace CPU integer_motion with the pipelined “v2” algorithm (and related SIMD plumbing), add 2160p@1.5H CSF equivalence handling + tests, and enable/fix additional ADM SIMD paths.

Changes:

  • Remove legacy non-direct-read frame fetch path in core/tools/vmaf.c and simplify fetch_picture() call sites.
  • Replace CPU motion extractor implementation with pipelined variant and update x86 SIMD implementations/headers; adjust subsampling logic for PREV_REF extractors.
  • Add 2160p@1.5H CSF routing + pinning tests; enable ADM adm_decouple_s123 SIMD dispatch and fix AVX2/AVX-512 shift/blend behavior.

Reviewed changes

Copilot reviewed 23 out of 23 changed files in this pull request and generated 13 comments.

Show a summary per file
File Description
docs/research/port-c2155d6cd-2160p-csf.md Research digest for 2160p@1.5H CSF port.
docs/research/port-a4a1492d3-integer-motion-pipelined-rename.md Research digest for pipelined motion rename/port.
docs/rebase-notes.md Rebase guidance updated for the upstream port set.
docs/adr/README.md ADR index updated to include ADR-0984.
docs/adr/0984-port-upstream-netflix-may-jun-2026.md ADR documenting the upstream port set.
core/tools/vmaf.c Removes compile-time direct-read guard; fetch_picture() now always direct-reads.
core/test/test_feature_extractor.c Updates motion extractor flush test to work with PREV_REF behavior.
core/test/test_barten_csf.c Adds pinning assertions for 2160p@1.5H CSF equivalence.
core/src/meson.build Drops CPU motion_v2 sources from build lists as part of rename/merge.
core/src/libvmaf.c Prevents subsampling for PREV_REF extractors (motion pipeline dependency).
core/src/feature/x86/motion_avx512.h Updates AVX-512 motion SIMD API to pipeline functions.
core/src/feature/x86/motion_avx512.c Replaces AVX-512 motion SIMD implementation with pipeline version.
core/src/feature/x86/motion_avx2.h Updates AVX2 motion SIMD API to pipeline functions.
core/src/feature/x86/motion_avx2.c Replaces AVX2 motion SIMD implementation with pipeline version.
core/src/feature/x86/adm_avx512.c Fixes AVX-512 shifts/zero-clear behavior in adm_decouple_s123.
core/src/feature/x86/adm_avx2.c Fixes AVX2 MSB blend + shift handling in adm_decouple_s123.
core/src/feature/integer_motion.c Moves CPU motion extractor to pipelined algorithm and PREV_REF dependency.
core/src/feature/feature_extractor.c Removes CPU motion_v2 from the (C) registry list.
core/src/feature/barten_csf_tools.h Adds 2160p@1.5H routing to 1080p@3H CSF tables.
changelog.d/fixed/port-9a078011c-adm-decouple-s123-simd-dispatch-and-avx2-blend-fix.md Changelog entry for ADM SIMD dispatch + AVX2 fix port.
changelog.d/changed/port-e4b93c6ed-direct-read-default.md Changelog entry for direct-read default behavior.
changelog.d/changed/port-a4a1492d3-integer-motion-pipelined-rename.md Changelog entry for motion pipelined rename/merge.
changelog.d/added/port-c2155d6cd-2160p-1.5h-csf-support.md Changelog entry for 2160p@1.5H CSF support.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +321 to +326
const unsigned min_idx = s->motion_five_frame_window ? 2 : 1;
if (index >= min_idx) {
const VmafPicture *prev =
s->motion_five_frame_window ? &fex->prev_prev_ref : &fex->prev_ref;
if (!prev->ref)
return -EINVAL;
Comment on lines 251 to +268
static int init(VmafFeatureExtractor *fex, enum VmafPixelFormat pix_fmt, unsigned bpc, unsigned w,
unsigned h)
{
(void)pix_fmt;
MotionState *s = fex->priv;
int err = 0;

/* The 5-tap separable Gaussian uses reflect-101 mirror padding: the
* bottom-edge formula `height - (i_tap - height + 2)` requires
* height >= radius + 1 = 3. For smaller frames the index goes
* negative, producing out-of-bounds reads (UB / ASan SEGV).
* Refuse cleanly here instead of reading past the buffer. */
const unsigned min_dim = (unsigned)(filter_width / 2 + 1); /* = 3 */
if (h < min_dim || w < min_dim) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"motion: frame %ux%u is below the 5-tap filter minimum %ux%u; "
"refusing to avoid out-of-bounds mirror reads\n",
w, h, min_dim, min_dim);
return -EINVAL;
}
s->w = w;
s->h = h;
s->bpc = bpc;

s->feature_name_dict =
vmaf_feature_name_dict_from_provided_features(fex->provided_features, fex->options, s);
if (!s->feature_name_dict)
goto fail;
return -ENOMEM;

if (s->motion_force_zero) {
fex->extract = extract_force_zero;
fex->flush = NULL;
fex->close = close_force_zero;
return 0;
}

err |= vmaf_picture_alloc(&s->tmp, pix_fmt, 16, w, h);
err |= vmaf_picture_alloc(&s->blur[0], pix_fmt, 16, w, h);
err |= vmaf_picture_alloc(&s->blur[1], pix_fmt, 16, w, h);
err |= vmaf_picture_alloc(&s->blur[2], pix_fmt, 16, w, h);
err |= vmaf_picture_alloc(&s->blur[3], pix_fmt, 16, w, h);
err |= vmaf_picture_alloc(&s->blur[4], pix_fmt, 16, w, h);
if (err)
goto fail;
s->y_row = malloc(sizeof(*s->y_row) * w);
if (!s->y_row)
return -ENOMEM;
Comment on lines +25 to 33
uint64_t motion_score_pipeline_8_avx512(const uint8_t *prev, ptrdiff_t prev_stride,
const uint8_t *cur, ptrdiff_t cur_stride, int32_t *y_row,
unsigned w, unsigned h, unsigned bpc);

void y_convolution_16_avx512(void *src, uint16_t *dst, unsigned width, unsigned height,
ptrdiff_t src_stride, ptrdiff_t dst_stride, unsigned inp_size_bits);
uint64_t motion_score_pipeline_16_avx512(const uint8_t *prev, ptrdiff_t prev_stride,
const uint8_t *cur, ptrdiff_t cur_stride, int32_t *y_row,
unsigned w, unsigned h, unsigned bpc);

void sad_avx512(VmafPicture *pic_a, VmafPicture *pic_b, uint64_t *sad);
#endif /* X86_AVX512_MOTION_H_ */
Comment on lines +1 to +16
# ADR-0984: Port Netflix Upstream May–Jun 2026 (5 commits)

## Status

Accepted

## Date

2026-06-01

## Context

Five upstream Netflix/vmaf commits from May–June 2026 were ported to the fork,
each path-remapped from `libvmaf/` to `core/` (per ADR-0700). The commits were
applied in chronological order on branch `port/upstream-netflix-may-jun-2026`.

Comment on lines +48 to +53
5. **30f472b14** (2026-06-01) — `Speed_chroma: vectorize
compute_covariance (AVX2 + AVX-512)`
Adds a `compute_cov_kernel_fn` function-pointer in `SpeedState` and
dispatches to new `compute_cov_kernel_avx2` / `compute_cov_kernel_avx512`
implementations in new files `x86/speed_avx2.{c,h}` and
`x86/speed_avx512.{c,h}`. Scalar path retained as fallback.
Comment on lines +55 to +68
## Decision

All 5 commits ported with path remapping. Fork-local adaptations:

- **Commit #1**: `copy_picture_data()` and `finish_unread_picture()` helper
removed (no longer called). Call sites updated.
- **Commit #2**: CPU `vmaf_fex_integer_motion_v2` removed from
`feature_extractor_list[]`; GPU `_v2_*` variants kept. `integer_motion_v2.c`
and `motion_v2_avx2/avx512.{c,h}` kept for GPU build paths.
`motion_avx2/avx512.{c,h}` updated to the pipeline implementation.
Upstream precision reductions in `python/test/` golden assertions NOT applied
(CLAUDE.md §12 r1 / correctness-first rule).
- **Commit #3–5**: applied cleanly.

Comment on lines +77 to +88
## Consequences

- `tools/vmaf` CLI always uses direct-read input path
(no `USE_DIRECT_READ` flag).
- `vmaf_fex_integer_motion` is now the pipelined implementation on CPU.
- 2160p at 1.5H viewing distance is a valid CSF configuration.
- ADM `adm_decouple_s123` SIMD dispatch is enabled; AVX2 kh_msb bug fixed.
- Speed_chroma `compute_covariance` is SIMD-accelerated on x86.

## References

- Upstream commits: e4b93c6ed, a4a1492d3, c2155d6cd, 9a078011c, 30f472b14
Comment thread docs/rebase-notes.md
Comment on lines +43366 to +43369
Fork-local files touched:
`core/tools/vmaf.c` (commit #1 — call-site updates for signature change),
`core/src/feature/feature_extractor.c` (commit #2 — remove CPU v2),
`core/src/meson.build` (commits #2, #5 — add speed_avx2/512, remove motion_v2 CPU build).
Comment thread core/src/meson.build
Comment on lines 1453 to 1456
feature_src_dir + 'integer_adm.c',
feature_src_dir + 'feature_collector.c',
feature_src_dir + 'integer_motion.c',
feature_src_dir + 'integer_motion_v2.c',
feature_src_dir + 'integer_vif.c',
Comment on lines 49 to 54
extern VmafFeatureExtractor vmaf_fex_psnr_hvs;
extern VmafFeatureExtractor vmaf_fex_integer_adm;
extern VmafFeatureExtractor vmaf_fex_integer_motion;
extern VmafFeatureExtractor vmaf_fex_integer_motion_v2;
/* vmaf_fex_integer_motion_v2 (CPU) removed — merged into vmaf_fex_integer_motion
* by upstream a4a1492d3. GPU variants (_cuda, _sycl, _hip, _metal) are kept. */
extern VmafFeatureExtractor vmaf_fex_integer_vif;
lusoris added a commit that referenced this pull request Jun 4, 2026
…order (ADR-1052) (#673)

PR #532 / commit 6bb5464 ported upstream a4a1492d3 and inadvertently
removed integer_motion_v2.c from the meson CPU source list and dropped
the extern declaration + list entry from feature_extractor.c.  On every
CPU-only build (including ARM64 CI lanes), vmaf_get_feature_extractor_by_name(
"motion_v2") returned NULL, causing test_motion_v2_missing and all related
test files to fail their "motion_v2 extractor missing" assertion.

Secondary failure: test_motion_three_frame_extract_emits_scores called
vmaf_feature_collector_get_score for motion2_score BEFORE flush().  The
post-port contract defers motion2 emission to flush(); the get_score must
come after.

Fix:
1. Re-add integer_motion_v2.c to core/src/meson.build CPU source list.
2. Restore extern VmafFeatureExtractor vmaf_fex_integer_motion_v2 and
   &vmaf_fex_integer_motion_v2 in feature_extractor.c.
3. Move get_score(motion2_score, frame=2) to after flush() in
   test_integer_motion_coverage.c::test_motion_three_frame_extract_emits_scores.

Unblocks: test_motion_v2_missing, test_motion_three_frame,
test_integer_motion_v2_coverage, test_motion_min_dim (motion_v2 sub-tests).

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
lusoris added a commit that referenced this pull request Sep 29, 2026
* fix(sycl): make motion_sycl bit-exact with the CPU motion

motion_sycl now gives the CPU motion's scores bit for bit at every frame
size, which closes T-SYCL-MOTION-TINY-FRAME-PARITY-2026-09-29.

The cause was the order of blurring and differencing, not edge handling.
Since the port of Netflix a4a1492d (PR #532) the CPU sums the absolute value
of blur(prev - cur), rounding after each filter pass. motion_sycl still
blurred each frame and differenced the blurred frames, which rounds
differently: motion2 was up to 2.0e-4 off at 17x17, 1.3e-5 on the Netflix
576x324 pair and 5.6e-6 at 4K, identically on an Arc B580 and a UHD 770.

motion_sycl and motion_v2_sycl now run one diff-first SAD kernel in
integer_motion_pipeline_sycl.cpp. motion_v2_sycl was already exact and is
unchanged. The new test_sycl_motion_tiny_frames compares both twins with the
scalar CPU using == from 3x3 to 1283x723 at 8, 10 and 16 bits. It passes on
both GPUs and fails against the old twin. At 4K the motion step costs about
11% more device time.

With motion_add_uv, submit() no longer waits on the device. U and V are
staged in pinned memory and uploaded on the combined queue, which closes
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29 and supersedes ADR-1034's
primary-queue wait. Host time per 4K frame drops from 5.4 to 0.6 ms on the
UHD 770.

The CUDA, HIP and Metal motion twins keep the old order. This change opens
RC3 rows for them, and one for a SYCL state reused across frame sizes.

ADR-1371, Research-1371.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(state): record the motion_add_uv chroma geometry gap

motion_configure_chroma() sizes U and V as 4:2:0 for every pixel format,
so 4:2:2 and 4:4:4 input stage only part of each chroma plane. Found by
reading the source for the ADR-1371 staging change and recorded as the
open RC3 row T-SYCL-MOTION-ADD-UV-CHROMA-GEOMETRY-2026-09-29.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants