Skip to content

adm: add NEON scale-zero decoupling - #1656

Merged
kylophone merged 1 commit into
Netflix:masterfrom
dantrapp:perf/adm-neon
Oct 2, 2026
Merged

kylophone merged 1 commit into
Netflix:masterfrom
dantrapp:perf/adm-neon

Conversation

@dantrapp

@dantrapp dantrapp commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

ADM scale-zero decoupling runs scalar on AArch64. This adds a four-sample NEON implementation for integer enhancement limits, retaining the scalar path for fractional limits and regions narrower than four samples.

At an enhancement limit of 1, the Q15 reconstruction already lies between zero and the distorted coefficient, so the angle-dependent clamp cannot change it. The NEON path skips that angle calculation. Other integer limits retain the scalar angle test's intermediate float rounding and double-precision comparisons.

On an Apple M4, ADM extraction with the VMAF v1 model options improved from 0.5315 s to 0.4327 s (1.228x). Full vmaf_v1.0.16_3d0h timings:

Threads Before After Speedup
1 1.6405 s 1.5464 s 1.061x
2 0.8674 s 0.8068 s 1.075x
4 0.5451 s 0.5148 s 1.059x
8 0.4187 s 0.3971 s 1.054x

These are medians of seven alternating runs after warmup on macOS 15.6.1 with Apple Clang 17. The input is the upstream five-frame 1920x1080 8-bit YUV420 fixture repeated ten times; file reads are included, decoding is not. No other ARM hardware was benchmarked. With #1653 also applied, this change adds 1.077–1.130x full-model speedup across the same thread counts; it does not depend on that PR.

Validation:

  • Checkasm coverage for small regions, odd widths, vector tails, padding, full-range coefficients, near-aligned and enhanced inputs, the one-degree threshold, and integer/fractional enhancement limits.
  • Release and ASan/UBSan suites: 23/23 passing each. Scalar-only suite: 21/21 passing.
  • 8,280 full-precision feature/prediction comparisons matched across 48 cases, including 8/10/12/16-bit inputs and 1/2/4/8 threads. Higher-bit-depth inputs were synthetic; VMAF v1 gain variants were used for numerical validation alongside the unmodified v1 and v0.6.1 models.
  • A separate exhaustive check confirmed that the gain-1 clamp leaves all 4,294,967,296 int16 reference/distorted coefficient pairs unchanged.

@dantrapp
dantrapp marked this pull request as ready for review October 1, 2026 23:04
@kylophone
kylophone merged commit 9e48141 into Netflix:master Oct 2, 2026
13 checks passed
@lusoris

lusoris commented Oct 2, 2026

Copy link
Copy Markdown

Checked this after the merge (9e48141 against cea2b4d). The NEON scale-zero decouple matches the scalar adm_decouple bit for bit in everything I ran, and I found no memory-safety problem. Details, since the thread has none:

Builds and tests. x86-64 GCC 16.2 and Clang 22.1, release and ASan/UBSan: 25/25 pass in release on both revisions. The ASan/UBSan runs fail the same three tests on both revisions: test_predict and test_pic_preallocation (LeakSanitizer reports), and checkasm, which aborts in the existing adm_dwt2_16 test (heap overflow at integer_adm.c:2603 with GCC, a UBSan report in check_adm.c with Clang), not in the new code. aarch64 cross build (GCC 16.1, release) under qemu-user: 26/26 on both revisions; checkasm --test=adm goes from 3 to 363 checks, all pass.

Kernel against scalar. A standalone harness calls adm_decouple_neon() and adm_decouple() on identical band buffers and compares all six output bands over the whole buffer, padding included. Every dimension 1..40 plus 48, 63, 64, 65, 97, 129 (so tails of 1 to 3 samples and the right - left < 4 fallback), strides equal to the width, rounded up to 8, and width plus a random pad. Gains 1, 2, 3, 4, 5, 7, 10, 50, 99, 100 (vector path) and 1.2, 1.5 (scalar fallback). Nine input patterns: random full-range int16; reference scaled copies (angle passes, so the gain limit binds); the 22916 coefficient bound; the special values 0, ±1, ±2, 32767, -32768, ±16384, ±4095; near-aligned small magnitudes; (h, v) rotated by 1 degree ± 0.5 % at magnitudes 1 to 30000; zeros on either side; and -32768/32767 in every band. About 1.6 million cases, 0 mismatches; qemu-user, GCC 16.1 -O2 and Clang 22 -O3 -Wall -Wextra builds of adm_neon.c (no warnings). In the aggregate the gain limit changes the scalar result for 2216 to 220021 samples per pattern at gain 100 against gain 1, so the branch is exercised. Four deliberate defects in a copy of the kernel (truncating instead of rounding rst, the angle threshold scaled by 1.0000001, > for >= 0, +1 on the limited sample) are each caught by the harness.

Whole extractor. C API at %.17g, integer_adm2 and integer_adm_scale0..3, aarch64: the three Netflix pairs at limits 1, 1.2, 1.5, 2, 3, 5, 10, 50 and 100; synthetic frames from 17x17 to 641x359; 10/12/16-bit; 2, 4, 8 threads. 132 cases, 831 frames: old against new, new NEON against scalar (--cpumask 1), and old NEON against scalar are bit-identical in every frame.

Edges. Same harness with every band in its own mapping, ending at (or starting after) an inaccessible page, so any access outside [0, h * stride) faults: no fault in any case.

One coverage gap in the checkasm test: the div_lookup it passes is the static copy in integer_adm.h that check_adm.c never fills (only adm_buffer_alloc() fills its own copy), so the reconstruction ratio is zero for every sample with a non-zero reference coefficient and the gain-limited branch of the new kernel never runs there. Replacing the limited sample with min(scaled + 1, dis) in the kernel still passes checkasm --test=adm. #1633, rebased on this commit, keeps the cases added here and fills the lookup (div_lookup_generator()), adds the limits 1.2 and 1.5 and one pattern whose distorted bands are derived from the reference; with that test checkasm --test=adm fails the +1 mutation (55 of 507 tests on aarch64 under qemu) and passes with the kernel as merged (507 tests). Boundary mutations of the angle test are only caught by the dedicated harness.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants