Repository navigation
adm: add NEON scale-zero decoupling - #1656
Conversation
|
Checked this after the merge (9e48141 against cea2b4d). The NEON scale-zero decouple matches the scalar Builds and tests. x86-64 GCC 16.2 and Clang 22.1, release and ASan/UBSan: 25/25 pass in release on both revisions. The ASan/UBSan runs fail the same three tests on both revisions: Kernel against scalar. A standalone harness calls Whole extractor. C API at Edges. Same harness with every band in its own mapping, ending at (or starting after) an inaccessible page, so any access outside One coverage gap in the checkasm test: the |
ADM scale-zero decoupling runs scalar on AArch64. This adds a four-sample NEON implementation for integer enhancement limits, retaining the scalar path for fractional limits and regions narrower than four samples.
At an enhancement limit of 1, the Q15 reconstruction already lies between zero and the distorted coefficient, so the angle-dependent clamp cannot change it. The NEON path skips that angle calculation. Other integer limits retain the scalar angle test's intermediate float rounding and double-precision comparisons.
On an Apple M4, ADM extraction with the VMAF v1 model options improved from 0.5315 s to 0.4327 s (1.228x). Full
vmaf_v1.0.16_3d0htimings:These are medians of seven alternating runs after warmup on macOS 15.6.1 with Apple Clang 17. The input is the upstream five-frame 1920x1080 8-bit YUV420 fixture repeated ten times; file reads are included, decoding is not. No other ARM hardware was benchmarked. With #1653 also applied, this change adds 1.077–1.130x full-model speedup across the same thread counts; it does not depend on that PR.
Validation: