Skip to content

port(upstream): libvmaf/speed: use fused antialias filter on all platforms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f) - #2632

Merged
lusoris merged 1 commit into
masterfrom
port/upstream-9cb9479f
Oct 8, 2026
Merged

lusoris merged 1 commit into
masterfrom
port/upstream-9cb9479f

Conversation

@lusoris

@lusoris lusoris commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Ports the last two upstream SpEED commits (upstream branch speed-fused-avx2, landed 2026-10-07). After this PR, speed_chroma and speed_temporal evaluate the anti-alias filter only at the samples the 16x decimation keeps on x86 too, and x86 computes that filter's vertical pass with AVX2. No score changes: 141 of 141 x86 reports are byte-identical before and after, at the scalar, AVX2 and AVX-512 dispatch levels.

Upstream commit Result
ad42c532 libvmaf/speed: use fused antialias filter on all platforms Ported. filter_and_downscale() (speed.c) and its mirror speed_internal_filter_and_downscale() (speed_internal.c) lose the #if ARCH_X86 two-call branch.
9cb9479f libvmaf/speed: add AVX2 fused antialias filter and decimation Ported in the fork's form. Only the vertical row pass is vectorised (convolution_f32_avx_rows_s()); the decimated horizontal pass stays the one scalar implementation in vif_tools.c. The checkasm case's sizes are rows of test_speed_filter.c.

Maintainer decision Q-297 (2026-10-08): "PR A (SpEED fused + AVX2) and PR B (ADM NEON 8bc5a5c6a + b41d2340a): start now."

What changed

  • core/src/feature/vif_tools.c: vif_filter1d_dec16_s() takes its vertical pass from the new vif_filter1d_vertical_dispatch_s(). On x86 with AVX2 and at most 17 taps (the same gate as vif_filter1d_s()'s AVX2 convolution), that function resolves the mirrored row pointers with vif_mirror_index() and calls convolution_f32_avx_rows_s(). Otherwise it calls the existing scalar vif_filter1d_vertical_s().

  • core/src/feature/common/convolution_avx.c / convolution.h: convolution_f32_avx_rows_s() computes one output row of the vertical pass. Each lane holds one column, and the products are added in tap order to a sum that starts at zero, which is the scalar order. It takes no stride and needs no alignment: loads are unaligned and the tail is masked, so nothing is read or written past the width.

  • Not taken from upstream, on purpose (see the rebase note):

    • upstream's convolution_f32_avx_dec16_s() includes a second scalar copy of the decimated horizontal pass (HISS-19);
    • the vif_filter1d_dec16_scalar_s() split is not needed, because the CPU mask already selects the scalar pass;
    • the VMAF_NO_FUSE asm barrier is not needed, because every translation unit already builds with contraction off (ADR-1461).
  • core/src/feature/speed.c, core/src/feature/speed_internal.c: the fused call runs on every target.

  • CUDA, HIP and SYCL SpEED twins: comments only. They already evaluate the filter at the decimated samples in the scalar arithmetic. Each file keeps its line count, and the comment-stripped sources (gcc -x c++ -fpreprocessed -dD -E -P) are byte-identical to master's. The Rust twin already called filter_dec16.

  • core/test/test_speed_filter.c: each case now runs at three dispatch settings, and every result is compared with memcmp against the scalar two-call reference:

    • the fused call at scalar dispatch;
    • the fused call with the host's flags;
    • the x86 path before this PR (AVX2 convolution, then vif_dec16_s()).

    A third, aligned layout covers that last path, because its convolution loads aligned. Also added: upstream's checkasm size 79x48, and test_avx_rows, which compares the AVX2 row pass with the scalar sum on every column of widths 1 to 41. The decimated horizontal pass never reads a row's last columns, so the masked tail is only visible in test_avx_rows.

  • Docs: docs/metrics/speed.md (CPU SIMD dispatch table) and docs/backends/arm/overview.md. docs/development/known-upstream-bugs.md records the port. Its parity heading stays at 9e48141b until every upstream commit between it and 9cb9479f is in.

Evidence

Measured on ryzen-4090-arc (Ryzen 9 9950X3D, AVX-512), GCC 16.2.1 release builds of master a7a9120c1 and of this branch, -Db_lto=false.

Check Result
x86 before/after, --precision max, --cpumask 63 / 48 / 0 (scalar, AVX2, AVX-512) 141 of 141 reports identical (frames, pooled and aggregate metrics, scores_digest): 2655 frames, 14 283 per-frame scores
Across dispatch levels on this branch 47 of 47 fixture/option groups identical at 63 / 48 / 0
Path actually taken (gdb breakpoint count on convolution_f32_avx_rows_s, Netflix pair, 2 frames) --cpumask 63: 0 calls; 48: 120; 0: 120
test_speed_filter pass (fused scalar, fused with host flags and the old x86 path, on 3 layouts × 22 sizes × 5 patterns × 8 kernel scales, plus whole-frame and test_avx_rows)
Proof the test catches defects reversed tap order in the vector body: fails (fused, host flags); +1e-3 on the last tap of the masked tail: test_avx_rows fails 216 cases; full-width store in the tail: aborts (guard)
test_speed, test_speed_upstream_form, test_speed_chroma, test_speed_simd, test_speed_frame_buffers pass
Netflix golden gate (core/build-golden, gcc, GOLDEN_NINJA_JOBS=4) 280 passed, 3 skipped
scripts/dev/preflight.sh --stage msvcism pass

Fixtures: the Netflix 576x324 pair (8-bit 4:2:0, 48 frames; 10-bit 4:2:0; 10-bit 4:2:2), both 1080p checkerboard pairs and 6 frames of BBB 3840x2160. Option sets: the defaults, speed_kernelscale 0.5 / 2.0 / 3.0 / 360/97, speed_prescale 1.5 / 0.5, and the default model (vmaf_v1.0.16_3d0h, which reads speed_chroma). The 1 px checkerboard at speed_prescale=1.5 fails in the same way before and after (non-finite score, already known), so its 3 runs are not in the count.

The golden gate's pytest step ran with the main checkout's .venv interpreter (pytest 9.1.1). The worktree's bootstrap venv carries only the build lock and has no pytest. The build was the Makefile's own build-golden.

Device parity was not re-run. It cannot move, because:

  • the CPU reference is byte-identical before and after on every report above;
  • the CUDA, HIP and SYCL sources are byte-identical to master once comments are stripped.

The exact-twin declarations speed_{chroma,temporal}.{cuda,hip,sycl} stand unchanged.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • Every commit is signed off (git commit -s; fix a branch with git rebase --signoff origin/master). See DCO sign-off.
  • make format && make lint is green locally. Not run as one target. These are green: clang-format on the touched files, the commit hooks, the msvcism preflight stage, and the clang-tidy cpu lane on the touched .c files (tidy: cpu convolution_avx.c, vif_tools.c, speed.c, speed_internal.c, test_speed_filter.c: 0 findings).
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. The SpEED tests listed above pass. The merge train runs the fast suite on the stack.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. Not applicable: the GPU sources change in comments only, and the CPU reference does not move.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md). No new file.
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not a breaking change.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and nothing else is touched for the index. No ADR: the port follows upstream. The fork-local shape (one horizontal pass, no asm barrier) applies HISS-19 and ADR-1461.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred). Not needed — no state delta: upstream port with no score change; it opens and closes no bug.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests, except by porting Netflix's own updated assertion verbatim from upstream (value and places as upstream has them, measured against the fork's CPU build first; ADR-1828).
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. No golden value changes.

Cross-backend numerical results

Not applicable: the twins' code is unchanged and the CPU reference is byte-identical before and after (see Evidence).

Performance (if perf or feat)

Informational only: no performance claim in RC4. speed_chroma + speed_temporal, one thread, median of 3. Load average was 80 to 85 on 32 threads while this ran, so the absolute times are noisy; the ratios are large enough to show.

Fixture --cpumask Before After
Netflix 576x324, 48 frames 63 (scalar) 13.3 ms/frame 2.5 ms/frame
Netflix 576x324, 48 frames 0 (default) 2.1 ms/frame 1.35 ms/frame
BBB 3840x2160, 12 frames 63 (scalar) 491 ms/frame 55 ms/frame
BBB 3840x2160, 12 frames 0 (default) 52.6 ms/frame 27.8 ms/frame

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: upstream port with no score change; the measurements are in this body and in the rebase note.
  • Decision matrix — no alternatives: the port follows upstream. The fork-local shape is fixed by HISS-19 and ADR-1461 (rebase note).
  • AGENTS.md invariant note — core/src/feature/AGENTS.d/speed.md (fused on every target, the vertical dispatch, what the tests hold) and core/src/feature/AGENTS.d/convolution-dispatch.md (convolution_f32_avx_rows_s()).
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/changed/speed-fused-antialias-x86-avx2.md.
  • Rebase note — docs/rebase-notes.d/port-upstream-speed-fused-avx2.md.

Reproducer

meson setup build core -Dbuildtype=release -Db_lto=false
ninja -C build tools/vmaf test/test_speed_filter
build/test/test_speed_filter
for m in 63 48 0; do
  build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
      -d python/test/resource/yuv/src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 \
      --no_prediction --feature speed_chroma --feature speed_temporal \
      --cpumask $m --precision max --json -o after-$m.json   # compare with a master build's reports
done

Known follow-ups

  • aarch64 is untouched by this PR: everything new sits under #if ARCH_X86, and the scalar passes are unchanged. test_speed_filter from an aarch64 cross build of this branch (GCC 16.1, build-aux/aarch64-linux-gnu.ini) passes under qemu-aarch64-static (2 of 2 tests; test_avx_rows is x86-only).
  • The upstream parity pin moves after the remaining commits between 9e48141b and 9cb9479f land: ADM NEON (PR B), the MSVC naming and FFmpeg leg, the M_PI cleanup, and the ADM shared viewing distances.

…forms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f) (#2632)

* port(upstream): libvmaf/speed: use fused antialias filter on all platforms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f)

SpEED evaluates its anti-alias filter only at the samples the 16x
decimation keeps on x86 too (Netflix/vmaf ad42c532), and the vertical
pass of that filter runs on AVX2 where the host has it (9cb9479f).

The fork vectorises only the vertical row pass
(convolution_f32_avx_rows_s(), called from vif_tools.c with mirrored row
pointers). The decimated horizontal pass stays the one scalar
implementation in vif_tools.c, and upstream's VMAF_NO_FUSE barrier is not
needed because every translation unit builds with contraction off
(ADR-1461).

Scores are byte-identical before and after: 141 of 141 x86 reports at
--cpumask 63, 48 and 0 on the Netflix pair, the 1080p checkerboards and
BBB 4K. test_speed_filter compares the fused call, its AVX2 form and the
old x86 two-call path with the scalar two calls, and test_avx_rows holds
the AVX2 row pass to the scalar sum on every column, its masked tail
included. The GPU twins change in comments only.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the port/upstream-9cb9479f branch from a8b1264 to 87248a1 Compare October 8, 2026 20:25
@lusoris
lusoris merged commit 87248a1 into master Oct 8, 2026
11 of 37 checks passed
@lusoris
lusoris deleted the port/upstream-9cb9479f branch October 8, 2026 20:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:port Port from upstream Netflix/vmaf

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant