Repository navigation
port(upstream): libvmaf/speed: use fused antialias filter on all platforms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f) - #2632
Merged
Conversation
…forms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f) (#2632) * port(upstream): libvmaf/speed: use fused antialias filter on all platforms, add AVX2 fused antialias filter and decimation (ad42c532, 9cb9479f) SpEED evaluates its anti-alias filter only at the samples the 16x decimation keeps on x86 too (Netflix/vmaf ad42c532), and the vertical pass of that filter runs on AVX2 where the host has it (9cb9479f). The fork vectorises only the vertical row pass (convolution_f32_avx_rows_s(), called from vif_tools.c with mirrored row pointers). The decimated horizontal pass stays the one scalar implementation in vif_tools.c, and upstream's VMAF_NO_FUSE barrier is not needed because every translation unit builds with contraction off (ADR-1461). Scores are byte-identical before and after: 141 of 141 x86 reports at --cpumask 63, 48 and 0 on the Netflix pair, the 1080p checkerboards and BBB 4K. test_speed_filter compares the fused call, its AVX2 form and the old x86 two-call path with the scalar two calls, and test_avx_rows holds the AVX2 row pass to the scalar sum on every column, its masked tail included. The GPU twins change in comments only. Signed-off-by: Lusoris <lusoris@proton.me>
lusoris
force-pushed
the
port/upstream-9cb9479f
branch
from
October 8, 2026 20:25
a8b1264 to
87248a1
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ports the last two upstream SpEED commits (upstream branch
speed-fused-avx2, landed 2026-10-07). After this PR,speed_chromaandspeed_temporalevaluate the anti-alias filter only at the samples the 16x decimation keeps on x86 too, and x86 computes that filter's vertical pass with AVX2. No score changes: 141 of 141 x86 reports are byte-identical before and after, at the scalar, AVX2 and AVX-512 dispatch levels.ad42c532libvmaf/speed: use fused antialias filter on all platformsfilter_and_downscale()(speed.c) and its mirrorspeed_internal_filter_and_downscale()(speed_internal.c) lose the#if ARCH_X86two-call branch.9cb9479flibvmaf/speed: add AVX2 fused antialias filter and decimationconvolution_f32_avx_rows_s()); the decimated horizontal pass stays the one scalar implementation invif_tools.c. The checkasm case's sizes are rows oftest_speed_filter.c.Maintainer decision Q-297 (2026-10-08): "PR A (SpEED fused + AVX2) and PR B (ADM NEON 8bc5a5c6a + b41d2340a): start now."
What changed
core/src/feature/vif_tools.c:vif_filter1d_dec16_s()takes its vertical pass from the newvif_filter1d_vertical_dispatch_s(). On x86 with AVX2 and at most 17 taps (the same gate asvif_filter1d_s()'s AVX2 convolution), that function resolves the mirrored row pointers withvif_mirror_index()and callsconvolution_f32_avx_rows_s(). Otherwise it calls the existing scalarvif_filter1d_vertical_s().core/src/feature/common/convolution_avx.c/convolution.h:convolution_f32_avx_rows_s()computes one output row of the vertical pass. Each lane holds one column, and the products are added in tap order to a sum that starts at zero, which is the scalar order. It takes no stride and needs no alignment: loads are unaligned and the tail is masked, so nothing is read or written past the width.Not taken from upstream, on purpose (see the rebase note):
convolution_f32_avx_dec16_s()includes a second scalar copy of the decimated horizontal pass (HISS-19);vif_filter1d_dec16_scalar_s()split is not needed, because the CPU mask already selects the scalar pass;VMAF_NO_FUSEasm barrier is not needed, because every translation unit already builds with contraction off (ADR-1461).core/src/feature/speed.c,core/src/feature/speed_internal.c: the fused call runs on every target.CUDA, HIP and SYCL SpEED twins: comments only. They already evaluate the filter at the decimated samples in the scalar arithmetic. Each file keeps its line count, and the comment-stripped sources (
gcc -x c++ -fpreprocessed -dD -E -P) are byte-identical to master's. The Rust twin already calledfilter_dec16.core/test/test_speed_filter.c: each case now runs at three dispatch settings, and every result is compared withmemcmpagainst the scalar two-call reference:vif_dec16_s()).A third, aligned layout covers that last path, because its convolution loads aligned. Also added: upstream's checkasm size 79x48, and
test_avx_rows, which compares the AVX2 row pass with the scalar sum on every column of widths 1 to 41. The decimated horizontal pass never reads a row's last columns, so the masked tail is only visible intest_avx_rows.Docs:
docs/metrics/speed.md(CPU SIMD dispatch table) anddocs/backends/arm/overview.md.docs/development/known-upstream-bugs.mdrecords the port. Its parity heading stays at9e48141buntil every upstream commit between it and9cb9479fis in.Evidence
Measured on
ryzen-4090-arc(Ryzen 9 9950X3D, AVX-512), GCC 16.2.1 release builds of mastera7a9120c1and of this branch,-Db_lto=false.--precision max,--cpumask 63/48/0(scalar, AVX2, AVX-512)scores_digest): 2655 frames, 14 283 per-frame scoresconvolution_f32_avx_rows_s, Netflix pair, 2 frames)--cpumask 63: 0 calls;48: 120;0: 120test_speed_filtertest_avx_rows)fused, host flags); +1e-3 on the last tap of the masked tail:test_avx_rowsfails 216 cases; full-width store in the tail: aborts (guard)test_speed,test_speed_upstream_form,test_speed_chroma,test_speed_simd,test_speed_frame_bufferscore/build-golden, gcc,GOLDEN_NINJA_JOBS=4)scripts/dev/preflight.sh --stage msvcismFixtures: the Netflix 576x324 pair (8-bit 4:2:0, 48 frames; 10-bit 4:2:0; 10-bit 4:2:2), both 1080p checkerboard pairs and 6 frames of BBB 3840x2160. Option sets: the defaults,
speed_kernelscale0.5 / 2.0 / 3.0 / 360/97,speed_prescale1.5 / 0.5, and the default model (vmaf_v1.0.16_3d0h, which readsspeed_chroma). The 1 px checkerboard atspeed_prescale=1.5fails in the same way before and after (non-finite score, already known), so its 3 runs are not in the count.The golden gate's pytest step ran with the main checkout's
.venvinterpreter (pytest 9.1.1). The worktree's bootstrap venv carries only the build lock and has no pytest. The build was the Makefile's ownbuild-golden.Device parity was not re-run. It cannot move, because:
The exact-twin declarations
speed_{chroma,temporal}.{cuda,hip,sycl}stand unchanged.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
git commit -s; fix a branch withgit rebase --signoff origin/master). See DCO sign-off.make format && make lintis green locally. Not run as one target. These are green:clang-formaton the touched files, the commit hooks, themsvcismpreflight stage, and the clang-tidycpulane on the touched.cfiles (tidy: cpuconvolution_avx.c,vif_tools.c,speed.c,speed_internal.c,test_speed_filter.c: 0 findings).python3 scripts/ci/run_meson_test.py -- -C build. The SpEED tests listed above pass. The merge train runs the fast suite on the stack./cross-backend-diffand the worst ULP is ≤ 2. Not applicable: the GPU sources change in comments only, and the CPU reference does not move..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md). No new file.!orBREAKING CHANGE:and the migration path is documented below. Not a breaking change.docs/adr/_index_fragments/<NNNN-slug>.mdand nothing else is touched for the index. No ADR: the port follows upstream. The fork-local shape (one horizontal pass, no asm barrier) applies HISS-19 and ADR-1461.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred). Not needed — no state delta: upstream port with no score change; it opens and closes no bug.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests, except by porting Netflix's own updated assertion verbatim from upstream (value and places as upstream has them, measured against the fork's CPU build first; ADR-1828).Cross-backend numerical results
Not applicable: the twins' code is unchanged and the CPU reference is byte-identical before and after (see Evidence).
Performance (if
perforfeat)Informational only: no performance claim in RC4.
speed_chroma+speed_temporal, one thread, median of 3. Load average was 80 to 85 on 32 threads while this ran, so the absolute times are noisy; the ratios are large enough to show.--cpumaskDeep-dive deliverables (ADR-0108)
AGENTS.mdinvariant note —core/src/feature/AGENTS.d/speed.md(fused on every target, the vertical dispatch, what the tests hold) andcore/src/feature/AGENTS.d/convolution-dispatch.md(convolution_f32_avx_rows_s()).changelog.d/changed/speed-fused-antialias-x86-avx2.md.docs/rebase-notes.d/port-upstream-speed-fused-avx2.md.Reproducer
Known follow-ups
#if ARCH_X86, and the scalar passes are unchanged.test_speed_filterfrom an aarch64 cross build of this branch (GCC 16.1,build-aux/aarch64-linux-gnu.ini) passes underqemu-aarch64-static(2 of 2 tests;test_avx_rowsis x86-only).9e48141band9cb9479fland: ADM NEON (PR B), the MSVC naming and FFmpeg leg, the M_PI cleanup, and the ADM shared viewing distances.