Repository navigation
fix(simd): take sad_avx512's differences in unsigned lanes so 16-bit samples do not wrap - #2200
Merged
Merged
Conversation
8 of 10 tasks
…ing a frame twice (#2206) * fix(model): read a model collection's stored score instead of predicting a frame twice vmaf_score_at_index_model_collection() predicted every member model and wrote the members' and the four named bootstrap scores of the frame into the feature collector, which refuses a second write of a frame. A second per-frame call of a frame, or vmaf_score_pooled_model_collection() over a range holding a frame already scored per frame (its loop predicts every frame of the range again), therefore failed with -EINVAL ("feature ... cannot be overwritten"). No score was wrong; the call failed. A frame whose four named bootstrap scores are already in the collector now returns them, as vmaf_score_at_index() reads a single model's stored score first. The values are the first prediction's, bit for bit, so a pooled score equals a fresh session's. Upstream Netflix/vmaf has the same code. test_model_collection_score_repeat fails on master without the fix: the second per-frame call and the pooled call after a per-frame call both return -EINVAL. Golden gate 280 passed, 3 skipped.
…t index of vif_tools.c (#2204) * fix(feature): refuse a float_vif or SpEED prescaled plane past the int index of vif_tools.c vif_tools.c indexes a plane as y * stride + x in int. With vif_prescale or speed_prescale above about 1.414 at the 32768x32768 picture cap the prescaled plane passes INT_MAX samples, the index overflows and the resamplers write outside the plane (16K stays inside at every accepted prescale, up to 4.0). vif_plane_fits_int_index() in vif_tools.h is now checked by float_vif's init (init_scaled_plane()), speed.c's speed_init_dimensions() and speed_internal_init_dimensions(), the geometry of every SpEED device twin; such a plane fails init with -EINVAL before any allocation. No accepted input and no score changes. T-PRESCALED-PLANE-INT-INDEX-2026-10-05. Found by the RC3 accumulator audit. test_prescaled_plane_int_index fails on master's SpEED geometry.
* fix(dnn): preserve explicit int8 paths in session open When vmaf_dnn_session_open() was called with an explicit .int8.onnx path, resolve_load_path() lacked the kInt8Suffix early return present in dnn_attach_api.c, causing it to append a redundant .int8 suffix and derive <name>.int8.int8.onnx before falling back to the fp32 path. Preserve explicit int8 paths directly by returning 0 when the onnx_path already ends in kInt8Suffix, matching the behavior in dnn_attach_api.c. Closes T-DNN-SESSION-INT8-EXPLICIT-PATH-2026-10-05.
…samples do not wrap (#2200) * fix(simd): take sad_avx512's differences in unsigned lanes so 16-bit samples do not wrap sad_avx512() formed each sample difference with _mm512_sub_epi16 and _mm512_abs_epi16, in signed 16-bit lanes: 16-bit samples that differ by more than 32767 wrapped (65535 against 0 gave 1) and the vector columns disagreed with the scalar tail. The difference is now max_epu16 minus min_epu16, exact for every 16-bit pair; 10-bit results are unchanged. Only test_motion_avx512_parity calls the function, so no score changes. T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05. Found by the RC3 accumulator audit. The two new 16-bit cases fail on master's kernel.
lusoris
force-pushed
the
fix/sad-avx512-16bit-difference
branch
from
October 5, 2026 23:12
3f02eb2 to
16eafa6
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
sad_avx512()(core/src/feature/x86/motion_avx512.c) sums the absolute differences of two 16-bit luma planes. It formed each difference with_mm512_sub_epi16and_mm512_abs_epi16, in signed 16-bit lanes, so 16-bit samples that differ by more than 32767 wrapped: 65535 against 0 gave 1, and 40000 against 0 gave 25536. Its scalar tail was right, so the vector columns and the tail disagreed.Only
test_motion_avx512_paritycalls the function, and only with 10-bit samples. No extractor uses it, so no score changes. Found by the RC3 integer-overflow audit of every accumulator (CPU and SIMD rows; the wrap does not depend on the picture size).The difference is now
_mm512_max_epu16(a, b) - _mm512_min_epu16(a, b), which is exact for every 16-bit pair and leaves 10-bit results unchanged. The per-row lane sums stay in range: at most 32768 x 65535 = 2,147,450,880 per row, below INT32_MAX.Evidence
The host is
ryzen-4090-arc(AVX-512).test_motion_avx512_parity, newtest_sad_avx512_16bit_randomandtest_sad_avx512_16bit_extremesmotion_avx512.cbpc16 randomscalar 21056890, AVX-512 17263056clang-tidy ran in the container (
scripts/dev/tidy-lane.sh --only, cpu lane).motion_avx512.candtest_motion_avx512_parity.chave 0 findings.scripts/dev/preflight.sh --stage msvcismpasses.Checklist
docs/state.md:T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05opened and closed.Deep-dive deliverables (ADR-0108)
AGENTS.mdinvariant note —core/src/feature/x86/AGENTS.d/motion.md.sad_avx512()is fork-only (upstream'smotion_avx512.chas no such function).Reproducer
meson test -C build test_motion_avx512_parity