Skip to content
18 changes: 9 additions & 9 deletions .standards-baseline.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"version": 1,
"generated_at": "2026-09-19T09:52:52Z",
"generated_at": "2026-09-19T11:30:29Z",
"repository": "",
"commit_sha": "",
"total_infractions": 1622,
Expand Down Expand Up @@ -2735,30 +2735,30 @@
{
"rule_id": "HISS-01",
"file_path": "core/src/feature/cambi.c",
"line_number": 857,
"line_number": 868,
"message": "Legacy non-DAG control flow jump (goto)",
"fingerprint": "core/src/feature/cambi.c:857:HISS-01"
"fingerprint": "core/src/feature/cambi.c:868:HISS-01"
},
{
"rule_id": "HISS-01",
"file_path": "core/src/feature/cambi.c",
"line_number": 863,
"line_number": 874,
"message": "Legacy non-DAG control flow jump (goto)",
"fingerprint": "core/src/feature/cambi.c:863:HISS-01"
"fingerprint": "core/src/feature/cambi.c:874:HISS-01"
},
{
"rule_id": "HISS-01",
"file_path": "core/src/feature/cambi.c",
"line_number": 870,
"line_number": 881,
"message": "Legacy non-DAG control flow jump (goto)",
"fingerprint": "core/src/feature/cambi.c:870:HISS-01"
"fingerprint": "core/src/feature/cambi.c:881:HISS-01"
},
{
"rule_id": "HISS-01",
"file_path": "core/src/feature/cambi.c",
"line_number": 874,
"line_number": 885,
"message": "Legacy non-DAG control flow jump (goto)",
"fingerprint": "core/src/feature/cambi.c:874:HISS-01"
"fingerprint": "core/src/feature/cambi.c:885:HISS-01"
},
{
"rule_id": "HISS-04",
Expand Down
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -14425,6 +14425,25 @@ for HD/UHD frames. Four AVX-512 parity cases added to
and ADR-0754 (SSIM vert_combine). ADR-0763.


- **CAMBI's SIMD paths cover every stage and skip the work that changes
nothing**: the anti-dithering filter, derivative row, decimate, mode filter
and c-values stage now have AVX-512 kernels, and all but the mode filter have
NEON kernels in use (for the mode filter and the mask row the compilers
already vectorise the scalar code on aarch64). Several of these kernels had
been built but not called since the upstream CAMBI optimisation batch was
ported. The c-values stage on AVX2, AVX-512 and NEON now skips the pixels
that leave its sliding histogram unchanged, which is most of them on the flat
content CAMBI looks at; upstream's AVX2 c-values driver, which was slower
than the scalar code in icx builds (icx builds the published container), is
no longer used. On a Zen 5 core a whole CAMBI frame runs 1.23–1.32x faster
than before with GCC at the default AVX-512 dispatch, and a CPU without
AVX-512 gets 1.21–1.78x depending on the compiler; under `qemu-aarch64` the
NEON stages execute 9–44 % of the scalar instructions. Scores are
byte-identical on every path. See
[CAMBI CPU SIMD paths](docs/metrics/cambi.md#cpu-simd-paths) and
[Research-2065](docs/research/2065-cambi-simd-gaps.md).


- **CAMBI spatial-mask rows use SIMD on x86 and aarch64**: the per-frame
summed-area (`compute_dp_row`) and box-sum threshold (`compute_mask_row`)
rows of CAMBI's spatial mask now dispatch to AVX2, AVX-512 and — for the dp
Expand Down
17 changes: 17 additions & 0 deletions changelog.d/changed/perf-cambi-simd-gaps.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
- **CAMBI's SIMD paths cover every stage and skip the work that changes
nothing**: the anti-dithering filter, derivative row, decimate, mode filter
and c-values stage now have AVX-512 kernels, and all but the mode filter have
NEON kernels in use (for the mode filter and the mask row the compilers
already vectorise the scalar code on aarch64). Several of these kernels had
been built but not called since the upstream CAMBI optimisation batch was
ported. The c-values stage on AVX2, AVX-512 and NEON now skips the pixels
that leave its sliding histogram unchanged, which is most of them on the flat
content CAMBI looks at; upstream's AVX2 c-values driver, which was slower
than the scalar code in icx builds (icx builds the published container), is
no longer used. On a Zen 5 core a whole CAMBI frame runs 1.23–1.32x faster
than before with GCC at the default AVX-512 dispatch, and a CPU without
AVX-512 gets 1.21–1.78x depending on the compiler; under `qemu-aarch64` the
NEON stages execute 9–44 % of the scalar instructions. Scores are
byte-identical on every path. See
[CAMBI CPU SIMD paths](docs/metrics/cambi.md#cpu-simd-paths) and
[Research-2065](docs/research/2065-cambi-simd-gaps.md).
18 changes: 17 additions & 1 deletion core/src/feature/arm64/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,13 +82,29 @@ to other halves in **same PR**:
| **MS-SSIM decimate LPF** (ADR-0125) | `ms_ssim_decimate_neon.c` + `../x86/ms_ssim_decimate_avx2.c` + `../x86/ms_ssim_decimate_avx512.c` + scalar `../ms_ssim_decimate.c`. The 9-tap filter table appears verbatim in all four. |
| **PSNR-HVS DCT** (ADR-0160) | `psnr_hvs_neon.c` + `../x86/psnr_hvs_avx2.c` + scalar `../third_party/xiph/psnr_hvs.c`. Butterfly block byte-identical across the three; threading `ret` by pointer is load-bearing. |
| **SSIMULACRA 2 SIMD** (ADR-0161 / 0162 / 0163 / 0213 / 0252) | `ssimulacra2_neon.c` + `ssimulacra2_sve2.c` + `../x86/ssimulacra2_avx2.c` + `../x86/ssimulacra2_avx512.c` + `ssimulacra2_host_neon.c` + `../x86/ssimulacra2_host_avx2.c` + scalar `../ssimulacra2.c` + Vulkan host-path `../vulkan/ssimulacra2_vulkan.c` |
| **CAMBI calculate_c_values_row** (ADR-0452) | `cambi_neon.c` (`calculate_c_values_row_neon`) + `../x86/cambi_avx2.c` (`calculate_c_values_row_avx2`) + `../x86/cambi_avx512.c` (`calculate_c_values_row_avx512`) + scalar in `../cambi.c` (`calculate_c_values_row`). Every cambi inner-loop function ported to AVX2 **must** have AVX-512 + NEON siblings in the **same PR**. NEON uses vectorised mask-zero detection (vmaxvq_u16) + scalar per-pixel inner loop for bit-exact output (no gather instruction on NEON). Tested in `../../test/test_cambi_simd.c`. |
| **CAMBI stage kernels** (ADR-1256, Research-2065) | `cambi_neon.c` + `../x86/cambi_avx2.c` (upstream mirror) + `../x86/cambi_avx512.c` + scalar `../cambi.c`; NEON / AVX-512 c-values drivers share walk `../cambi_c_values_frame.h`. Dispatched on NEON: anti-dither, derivative, decimate, dp row, c-values. Kept scalar: mask row, mode filter. Tests: `test_cambi_stage_simd.c`, `test_cambi_dispatch_invariance.c`, `test_cambi_simd.c` (run under `qemu-aarch64` without aarch64 host). |
| **CAMBI spatial-mask rows** (ADR-1256) | `cambi_neon.c` (`compute_dp_row_neon`, `compute_mask_row_neon`) + `../x86/cambi_avx2.c` + `../x86/cambi_avx512.c` twins + scalar reference in `../cambi.c`. Only the dp row is dispatched on aarch64: GCC and Clang auto-vectorize the scalar mask row into the same `cmhi` / `uzp1` sequence, so `compute_mask_row_neon` stays built, parity-tested and undispatched — re-check the compiled scalar before wiring it. The dp row keeps a single add on the loop-carried chain; keep that shape. Tested in `../../test/test_cambi_spatial_mask_simd.c` (run under `qemu-aarch64` when no aarch64 host is available). |
| **Motion v2 NEON** (ADR-0145) | `motion_v2_neon.c` uses **arithmetic** right-shift (`vshrq_n_s64(v, 16)` / `vshlq_s64(v, -(int64_t)bpc)`); matches scalar. Sister `../x86/motion_v2_avx2.c` uses `_mm256_srlv_epi64` (logical) — knowingly out-of-spec until the AVX2 audit. **Do NOT port the AVX2 logical pattern here.** 4-lane stride + scalar tails on both sides of the row are load-bearing for the x_conv edge-mirror contract. |

Complete invariants in [../AGENTS.md
§"Rebase-sensitive invariants"](../AGENTS.md).

## CAMBI NEON invariants (Research-2065)

- `filter_mode_neon`, `compute_mask_row_neon`: built, parity-tested, not
dispatched. GCC + Clang vectorise scalar loop same way (insn count 0.93x /
1.02x). Re-count with qemu insn plugin before wiring.
- NEON c-values driver uses plain C range updaters on purpose: compilers emit
same 8-lane adds; intrinsic versions cost +0.3–0.9 % insns, retired. No
`cambi_*_range_neon`.
- Scans: no masked load → scalar tail < 8 cols via shared
`cambi_column_*` predicates (`../cambi_c_values_frame.h`, also AVX2). Never
vector-load past last column (last row may end at buffer end).
- Scan may over-flag, never under-flag; mirrors `uh_slide` skip + band test.
- `cambi_neon.c` lives in integer lib `arm64_v8` (no `-ffp-contract=off`):
c-value is one mul, no add, so nothing to fuse. Adding `a * b + c` float math
here → move TU to `arm64_v8_fp`.

## SVE2 invariants (ADR-0213, ADR-0584)

`ssimulacra2_sve2.c` and `moment_sve2.c` = SVE2 consumers in directory.
Expand Down
Loading
Loading