Skip to content

fix(simd): take sad_avx512's differences in unsigned lanes so 16-bit samples do not wrap - #2200

Merged
lusoris merged 4 commits into
masterfrom
fix/sad-avx512-16bit-difference
Oct 5, 2026
Merged

lusoris merged 4 commits into
masterfrom
fix/sad-avx512-16bit-difference

Conversation

@lusoris

@lusoris lusoris commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Summary

sad_avx512() (core/src/feature/x86/motion_avx512.c) sums the absolute differences of two 16-bit luma planes. It formed each difference with _mm512_sub_epi16 and _mm512_abs_epi16, in signed 16-bit lanes, so 16-bit samples that differ by more than 32767 wrapped: 65535 against 0 gave 1, and 40000 against 0 gave 25536. Its scalar tail was right, so the vector columns and the tail disagreed.

Only test_motion_avx512_parity calls the function, and only with 10-bit samples. No extractor uses it, so no score changes. Found by the RC3 integer-overflow audit of every accumulator (CPU and SIMD rows; the wrap does not depend on the picture size).

The difference is now _mm512_max_epu16(a, b) - _mm512_min_epu16(a, b), which is exact for every 16-bit pair and leaves 10-bit results unchanged. The per-row lane sums stay in range: at most 32768 x 65535 = 2,147,450,880 per row, below INT32_MAX.

Evidence

The host is ryzen-4090-arc (AVX-512).

Check Result
test_motion_avx512_parity, new test_sad_avx512_16bit_random and test_sad_avx512_16bit_extremes 12 of 12 pass
The same test against master's motion_avx512.c fails: bpc16 random scalar 21056890, AVX-512 17263056

clang-tidy ran in the container (scripts/dev/tidy-lane.sh --only, cpu lane). motion_avx512.c and test_motion_avx512_parity.c have 0 findings. scripts/dev/preflight.sh --stage msvcism passes.

Checklist

  • Conventional Commits.
  • docs/state.md: T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05 opened and closed.
  • No Netflix golden assertion modified; no score changes.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: trivial bug fix.
  • Decision matrix — no alternatives: only-one-way fix (the unsigned difference the scalar tail computes).
  • AGENTS.md invariant note — core/src/feature/x86/AGENTS.d/motion.md.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — no changelog needed: no extractor calls the function and no output changes.
  • Rebase note — no rebase impact: sad_avx512() is fork-only (upstream's motion_avx512.c has no such function).

Reproducer

meson test -C build test_motion_avx512_parity

…ing a frame twice (#2206)

* fix(model): read a model collection's stored score instead of predicting a frame twice

vmaf_score_at_index_model_collection() predicted every member model and
wrote the members' and the four named bootstrap scores of the frame into
the feature collector, which refuses a second write of a frame. A second
per-frame call of a frame, or vmaf_score_pooled_model_collection() over a
range holding a frame already scored per frame (its loop predicts every
frame of the range again), therefore failed with -EINVAL ("feature ...
cannot be overwritten"). No score was wrong; the call failed.

A frame whose four named bootstrap scores are already in the collector
now returns them, as vmaf_score_at_index() reads a single model's stored
score first. The values are the first prediction's, bit for bit, so a
pooled score equals a fresh session's. Upstream Netflix/vmaf has the same
code.

test_model_collection_score_repeat fails on master without the fix: the
second per-frame call and the pooled call after a per-frame call both
return -EINVAL. Golden gate 280 passed, 3 skipped.
…t index of vif_tools.c (#2204)

* fix(feature): refuse a float_vif or SpEED prescaled plane past the int index of vif_tools.c

vif_tools.c indexes a plane as y * stride + x in int. With vif_prescale or
speed_prescale above about 1.414 at the 32768x32768 picture cap the
prescaled plane passes INT_MAX samples, the index overflows and the
resamplers write outside the plane (16K stays inside at every accepted
prescale, up to 4.0). vif_plane_fits_int_index() in vif_tools.h is now
checked by float_vif's init (init_scaled_plane()), speed.c's
speed_init_dimensions() and speed_internal_init_dimensions(), the geometry
of every SpEED device twin; such a plane fails init with -EINVAL before any
allocation. No accepted input and no score changes.

T-PRESCALED-PLANE-INT-INDEX-2026-10-05. Found by the RC3 accumulator audit.
test_prescaled_plane_int_index fails on master's SpEED geometry.
* fix(dnn): preserve explicit int8 paths in session open

When vmaf_dnn_session_open() was called with an explicit .int8.onnx path,
resolve_load_path() lacked the kInt8Suffix early return present in
dnn_attach_api.c, causing it to append a redundant .int8 suffix and
derive <name>.int8.int8.onnx before falling back to the fp32 path.

Preserve explicit int8 paths directly by returning 0 when the onnx_path
already ends in kInt8Suffix, matching the behavior in dnn_attach_api.c.

Closes T-DNN-SESSION-INT8-EXPLICIT-PATH-2026-10-05.
…samples do not wrap (#2200)

* fix(simd): take sad_avx512's differences in unsigned lanes so 16-bit samples do not wrap

sad_avx512() formed each sample difference with _mm512_sub_epi16 and
_mm512_abs_epi16, in signed 16-bit lanes: 16-bit samples that differ by more
than 32767 wrapped (65535 against 0 gave 1) and the vector columns disagreed
with the scalar tail. The difference is now max_epu16 minus min_epu16, exact
for every 16-bit pair; 10-bit results are unchanged. Only
test_motion_avx512_parity calls the function, so no score changes.

T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05. Found by the RC3 accumulator
audit. The two new 16-bit cases fail on master's kernel.
@lusoris
lusoris force-pushed the fix/sad-avx512-16bit-difference branch from 3f02eb2 to 16eafa6 Compare October 5, 2026 23:12
@lusoris
lusoris merged commit 16eafa6 into master Oct 5, 2026
5 of 79 checks passed
@lusoris
lusoris deleted the fix/sad-avx512-16bit-difference branch October 5, 2026 23:12
@github-actions github-actions Bot added the type:bug Something isn't working label Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants