Skip to content

fix: 6 real-code-fix bundle (matrix-v2 follow-up + NEON FMA re-dispatch) - #853

Merged
lusoris merged 6 commits into
masterfrom
chore/bundle-b-6-matrix-v2-real-fixes
Jun 8, 2026
Merged

lusoris merged 6 commits into
masterfrom
chore/bundle-b-6-matrix-v2-real-fixes

Conversation

@lusoris

@lusoris lusoris commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Summary

Bundles 6 real production-code fixes:

  1. perf(neon): re-dispatch float-ADM via FMA-safe DWT2 kernel (closes ADR-1057)
  2. fix(cuda): adm_cm aim signal shift mismatch — fixes 7.66e-03 / 3.83e-03 CUDA parity failures vs CPU
  3. fix(cuda): synchronize async D-to-H download in translate_picture_device (fixes libvmaf_cuda identical-input score 69→100 — host-device race)
  4. fix(vmaftune): corpus parse_feature_aggregates — silent NaN for all 12 feature columns (canonical-name → integer_* JSON path mapping; +test)
  5. fix(vmaftune): recommend --with-uncertainty honored on --from-corpus path (was silently ignored; +test)
  6. chore(test): fetch-test-yuvs.sh — add 1080p checkerboard fixtures for golden gate (CLAUDE.md §8)

Test plan

  • meson fast suite passes (84/84 — test_pic_preallocation + full suite green)
  • vmaf-tune unit tests pass (1938 passed, 17 skipped, 2 xfailed; 20 new test cases from commits 4+5 pass)
  • git grep for conflict markers empty
  • No Netflix golden assertions touched
  • Note: test_format_both_json.py::test_compare_format_both_writes_json is a pre-existing failure on master (not introduced by this PR)

Deep-dive deliverables (ADR-0108)

  • Research digest: no digest needed: matrix-v2 surfaced bugs each with concrete root-cause analysis
  • Decision matrix: no alternatives: only-one-way fix for each surfaced bug
  • AGENTS.md invariant note: no rebase-sensitive invariants
  • Reproducer / smoke-test command: per-commit verify commands in body
  • changelog.d fragment: no changelog fragment needed: bug-fix bundle (PR description carries the per-fix summary)
  • docs/rebase-notes.md: no rebase impact: bug fixes only

state.md touch

  • state.md: opens + closes T-CUDA-ADM-CM-AIM-SHIFT-MISMATCH, T-FFMPEG-CUDA-HOST-DEVICE-RACE, T-VMAFTUNE-CORPUS-NAN-COLUMNS, T-VMAFTUNE-RECOMMEND-UNCERTAINTY-FROM-CORPUS, T-NEON-FMA-FLOAT-ADM-DWT2; closes T-CONTAINER-MISSING-CHECKERBOARD-FIXTURES

lusoris and others added 6 commits June 8, 2026 19:17
…-1057)

PR #695 / ADR-1057 reverted the float-ADM NEON dispatch because
vmlaq_laneq_f32 is a hardware FMA intrinsic that produces a 1-ULP
divergence from the scalar reference; #pragma clang fp contract(off)
alone cannot suppress intrinsic-encoded FMA.

Fix: extract the DWT2 kernel into float_adm_dwt2_neon.c compiled in its
own static library (arm64_adm_dwt2_neon_lib) with -ffp-contract=off.
Replace every vmlaq_laneq_f32(acc, v, coeff, lane) with the explicit
two-step vaddq_f32(acc, vmulq_laneq_f32(v, coeff, lane)), emitting
separate fmul + fadd instructions that match the scalar reference
rounding.  Belt-and-suspenders: add #pragma clang fp contract(off) at TU
level (Clang) and __attribute__((optimize("-ffp-contract=off"))) on the
function definition (GCC) to protect the scalar tail.

Wire the dispatch in adm.c::adm_dwt2_dispatch(): selects the NEON kernel
when VMAF_ARM_CPU_FLAG_NEON is set, falls back to adm_dwt2_s() otherwise.
The #define adm_dwt2 adm_dwt2_dispatch replaces the direct scalar macro.

Verified: aarch64-linux-gnu-gcc -c compiles all three TUs cleanly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…tching conclude_adm_cm offsets

The adm_cm_aim_line_kernel<rows_per_thread> for scale-0 int16 AIM path
incorrectly used inline_s0_csf_a() to compute the signal, which applies
a >>15 pre-shift before returning. The threshold path (conclude_adm_cm)
and the shift_sub[] constant array {10, 10, 12} are calibrated for the
unshifted signal magnitude (abs(rfactor * a_val)), producing near-zero
AIM scores because the signal was ~32768x too small.

Fix mirrors the DLM kernel pattern (adm_cm_line_kernel lines 349-355):
use inline_s0_decouple_r() to obtain r_val, load the distorted band
value t_val via __ldg(), compute a_val = t_val - r_val, then form the
signal as abs(int32_t(i_rfactor[blockIdx.z] * a_val)) without the >>15
shift. No API or header changes; CUDA-only internal kernel fix.

Build verified: meson setup core/build-aim-fix -Denable_cuda=true &&
ninja -C core/build-aim-fix (exit 0, adm_cm PTX compiled without errors).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ice before return (host-device race)

vmaf_cuda_picture_download_async issues a cuMemcpy2DAsync on the
picture's per-picture CUstream (CU_STREAM_NON_BLOCKING, allocated in
vmaf_cuda_picture_alloc per ADR-0378).  The function returned
immediately without waiting for the DMA to complete, so CPU-side
feature extractors (motion, adm, vif, etc.) racing in on the same
thread pool read partially-written host buffers — producing VMAF
scores of ~69-71 instead of ~100 on identical-input pairs.

Fix: after the async download succeeds, call
cu_f->cuStreamSynchronize(vmaf_cuda_picture_get_stream(pic)) to drain
the per-picture stream before returning.  The synchronize is O(1)
overhead on the critical path (the DMA is already in-flight; this
just adds a CPU-side wait) and is consistent with the
cuStreamSynchronize call in vmaf_cuda_picture_free and the
cuStreamSynchronize in flush_context_cuda.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…to integer_* JSON paths

Modern libvmaf emits pooled_metrics keys with an ``integer_`` prefix for
the integer-pipeline feature extractors (e.g. ``integer_adm2`` instead of
``adm2``, ``integer_vif_scale0`` instead of ``vif_scale0``).
``parse_feature_aggregates`` was looking up the bare canonical names
directly, so every corpus row's 12 per-feature columns (adm2_mean,
vif_scale[0..3]_mean, motion2_mean and their _std counterparts) were
silently NaN against any real libvmaf output.

Fix: add ``_CANONICAL_TO_POOLED_KEY`` mapping in score.py and update
``parse_feature_aggregates`` to resolve each canonical name through the
map first, falling back to the bare name for non-integer features (cambi,
future additions) and for synthetic test fixtures that use bare keys.

Regression guard: add test_parse_feature_aggregates_integer_keys.py with
8 cases covering real integer_* payloads, stddev presence/absence, bare-key
fallback, partial presence, empty pooled_metrics, and backward compat with
the existing synthetic-fixture shape used by test_corpus_schema_v3.py.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…path

_run_recommend_from_corpus unconditionally called the point-estimate
recommend() path, silently ignoring args.with_uncertainty. When
--from-corpus and --with-uncertainty are combined, the function now
routes through pick_target_vmaf_with_uncertainty (the interval-aware
predictor from ADR-0279), emitting decision= and visited= fields in
human-readable output.

--with-uncertainty + --target-bitrate is unsupported (no interval-aware
bitrate predicate exists); the function now warns on stderr and falls
back to the point-estimate path rather than silently ignoring the flag.

Tests added to test_recommend_uncertainty.py covering:
- plain vs uncertainty-aware output differ when tight interval fires
- --json output emits the winning row correctly on the uncertainty path
- --with-uncertainty + --target-bitrate warns and falls back gracefully

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
… golden gate

The script only fetched the src01 576x324 normal pair. The Containerfile
(dev/Containerfile, added in PR #850) calls this script at build time, but the
two 1920x1080 checkerboard pairs required by CLAUDE.md §8 golden gate
(python/test/quality_runner_test.py::test_run_vmaf_runner_checkerboard) were
never downloaded, so the checkerboard golden tests always failed inside the
container.

Add the three missing YUV files to the FIXTURES array with md5sums verified
2026-06-08 against Netflix/vmaf_resource HEAD:
  - checkerboard_1920_1080_10_3_0_0.yuv  ad14c75d1897e7f2bc72a882c32e49a6
  - checkerboard_1920_1080_10_3_1_0.yuv  0e290566458d800c534ab18103619d43
  - checkerboard_1920_1080_10_3_10_0.yuv 289119e1168ab656ac79487df8d307b9

Local verify: script run against a fresh tmp dir — all 5 fixtures fetched and
md5-verified; re-run with all files present — all 5 report "ok" (no network I/O).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris marked this pull request as ready for review June 8, 2026 17:56
@lusoris
lusoris merged commit a6c4dff into master Jun 8, 2026
103 of 112 checks passed
@lusoris
lusoris deleted the chore/bundle-b-6-matrix-v2-real-fixes branch June 8, 2026 17:56
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
lusoris added a commit that referenced this pull request Sep 6, 2026
The ledger row for T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06 was
self-contradictory: a 2026-06-27 note inside the
T-MASTER-CI-TSAN-ARM-GOLDEN-2026-06-27 row said the bug "stays Open"
for a correct re-attempt, while its Recently-closed row already
recorded a 2026-08-30 closure — and that closure narrative described
only the integer-path dropped-tap defect and cited a branch name
(`fix/adr-1057-arm-fma-drift`) instead of a commit.

Re-verified against origin/master, with no defective code left to point
at:

- The float-ADM follow-up the row asked for landed in `a6c4dfffb`
  (PR #853): the kernel lives in a dedicated non-contracting TU
  `core/src/feature/arm64/float_adm_dwt2_neon.c`, built with
  `-ffp-contract=off` (`core/src/meson.build`), and
  `adm_dwt2_dispatch()` in `core/src/feature/adm.c` calls
  `float_adm_dwt2_neon()` under `VMAF_ARM_CPU_FLAG_NEON`. The scalar
  `adm_dwt2_s` carries the matching function-scoped guard, so both
  sides of the comparison are non-contracting and the 1-ULP FMA gap
  that motivated the row cannot arise.
- The integer `idx < 3` dropped filter tap in `adm_dwt2_8_neon` was
  fixed by `a013c1410` (PR #1134) and hardened to bit-exactness by
  `89a8e3258` (PR #1154); `6d61106ed` (PR #1156) added NEON parity
  tests for the remaining uncovered kernels.
- `core/test/test_float_adm_dwt2_neon.c` gates NEON-vs-scalar
  bit-exactness by bit pattern and is registered in the default `fast`
  suite.

The row stays in "Recently closed" (exactly one row for the id),
rewritten in past tense with real commit shas, and the stale "stays
Open" note is marked superseded so the next session does not
re-investigate. Documentation only — no code or behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
The ledger row for T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06 was
self-contradictory: a 2026-06-27 note inside the
T-MASTER-CI-TSAN-ARM-GOLDEN-2026-06-27 row said the bug "stays Open"
for a correct re-attempt, while its Recently-closed row already
recorded a 2026-08-30 closure — and that closure narrative described
only the integer-path dropped-tap defect and cited a branch name
(`fix/adr-1057-arm-fma-drift`) instead of a commit.

Re-verified against origin/master, with no defective code left to point
at:

- The float-ADM follow-up the row asked for landed in `a6c4dfffb`
  (PR #853): the kernel lives in a dedicated non-contracting TU
  `core/src/feature/arm64/float_adm_dwt2_neon.c`, built with
  `-ffp-contract=off` (`core/src/meson.build`), and
  `adm_dwt2_dispatch()` in `core/src/feature/adm.c` calls
  `float_adm_dwt2_neon()` under `VMAF_ARM_CPU_FLAG_NEON`. The scalar
  `adm_dwt2_s` carries the matching function-scoped guard, so both
  sides of the comparison are non-contracting and the 1-ULP FMA gap
  that motivated the row cannot arise.
- The integer `idx < 3` dropped filter tap in `adm_dwt2_8_neon` was
  fixed by `a013c1410` (PR #1134) and hardened to bit-exactness by
  `89a8e3258` (PR #1154); `6d61106ed` (PR #1156) added NEON parity
  tests for the remaining uncovered kernels.
- `core/test/test_float_adm_dwt2_neon.c` gates NEON-vs-scalar
  bit-exactness by bit pattern and is registered in the default `fast`
  suite.

The row stays in "Recently closed" (exactly one row for the id),
rewritten in past tense with real commit shas, and the stale "stays
Open" note is marked superseded so the next session does not
re-investigate. Documentation only — no code or behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 6, 2026
The ledger row for T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06 was
self-contradictory: a 2026-06-27 note inside the
T-MASTER-CI-TSAN-ARM-GOLDEN-2026-06-27 row said the bug "stays Open"
for a correct re-attempt, while its Recently-closed row already
recorded a 2026-08-30 closure — and that closure narrative described
only the integer-path dropped-tap defect and cited a branch name
(`fix/adr-1057-arm-fma-drift`) instead of a commit.

Re-verified against origin/master, with no defective code left to point
at:

- The float-ADM follow-up the row asked for landed in `a6c4dfffb`
  (PR #853): the kernel lives in a dedicated non-contracting TU
  `core/src/feature/arm64/float_adm_dwt2_neon.c`, built with
  `-ffp-contract=off` (`core/src/meson.build`), and
  `adm_dwt2_dispatch()` in `core/src/feature/adm.c` calls
  `float_adm_dwt2_neon()` under `VMAF_ARM_CPU_FLAG_NEON`. The scalar
  `adm_dwt2_s` carries the matching function-scoped guard, so both
  sides of the comparison are non-contracting and the 1-ULP FMA gap
  that motivated the row cannot arise.
- The integer `idx < 3` dropped filter tap in `adm_dwt2_8_neon` was
  fixed by `a013c1410` (PR #1134) and hardened to bit-exactness by
  `89a8e3258` (PR #1154); `6d61106ed` (PR #1156) added NEON parity
  tests for the remaining uncovered kernels.
- `core/test/test_float_adm_dwt2_neon.c` gates NEON-vs-scalar
  bit-exactness by bit pattern and is registered in the default `fast`
  suite.

The row stays in "Recently closed" (exactly one row for the id),
rewritten in past tense with real commit shas, and the stale "stays
Open" note is marked superseded so the next session does not
re-investigate. Documentation only — no code or behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 7, 2026
The ledger row for T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06 was
self-contradictory: a 2026-06-27 note inside the
T-MASTER-CI-TSAN-ARM-GOLDEN-2026-06-27 row said the bug "stays Open"
for a correct re-attempt, while its Recently-closed row already
recorded a 2026-08-30 closure — and that closure narrative described
only the integer-path dropped-tap defect and cited a branch name
(`fix/adr-1057-arm-fma-drift`) instead of a commit.

Re-verified against origin/master, with no defective code left to point
at:

- The float-ADM follow-up the row asked for landed in `a6c4dfffb`
  (PR #853): the kernel lives in a dedicated non-contracting TU
  `core/src/feature/arm64/float_adm_dwt2_neon.c`, built with
  `-ffp-contract=off` (`core/src/meson.build`), and
  `adm_dwt2_dispatch()` in `core/src/feature/adm.c` calls
  `float_adm_dwt2_neon()` under `VMAF_ARM_CPU_FLAG_NEON`. The scalar
  `adm_dwt2_s` carries the matching function-scoped guard, so both
  sides of the comparison are non-contracting and the 1-ULP FMA gap
  that motivated the row cannot arise.
- The integer `idx < 3` dropped filter tap in `adm_dwt2_8_neon` was
  fixed by `a013c1410` (PR #1134) and hardened to bit-exactness by
  `89a8e3258` (PR #1154); `6d61106ed` (PR #1156) added NEON parity
  tests for the remaining uncovered kernels.
- `core/test/test_float_adm_dwt2_neon.c` gates NEON-vs-scalar
  bit-exactness by bit pattern and is registered in the default `fast`
  suite.

The row stays in "Recently closed" (exactly one row for the id),
rewritten in past tense with real commit shas, and the stale "stays
Open" note is marked superseded so the next session does not
re-investigate. Documentation only — no code or behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 7, 2026
…#1331)

The ledger row for T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06 was
self-contradictory: a 2026-06-27 note inside the
T-MASTER-CI-TSAN-ARM-GOLDEN-2026-06-27 row said the bug "stays Open"
for a correct re-attempt, while its Recently-closed row already
recorded a 2026-08-30 closure — and that closure narrative described
only the integer-path dropped-tap defect and cited a branch name
(`fix/adr-1057-arm-fma-drift`) instead of a commit.

Re-verified against origin/master, with no defective code left to point
at:

- The float-ADM follow-up the row asked for landed in `a6c4dfffb`
  (PR #853): the kernel lives in a dedicated non-contracting TU
  `core/src/feature/arm64/float_adm_dwt2_neon.c`, built with
  `-ffp-contract=off` (`core/src/meson.build`), and
  `adm_dwt2_dispatch()` in `core/src/feature/adm.c` calls
  `float_adm_dwt2_neon()` under `VMAF_ARM_CPU_FLAG_NEON`. The scalar
  `adm_dwt2_s` carries the matching function-scoped guard, so both
  sides of the comparison are non-contracting and the 1-ULP FMA gap
  that motivated the row cannot arise.
- The integer `idx < 3` dropped filter tap in `adm_dwt2_8_neon` was
  fixed by `a013c1410` (PR #1134) and hardened to bit-exactness by
  `89a8e3258` (PR #1154); `6d61106ed` (PR #1156) added NEON parity
  tests for the remaining uncovered kernels.
- `core/test/test_float_adm_dwt2_neon.c` gates NEON-vs-scalar
  bit-exactness by bit pattern and is registered in the default `fast`
  suite.

The row stays in "Recently closed" (exactly one row for the id),
rewritten in past tense with real commit shas, and the stale "stays
Open" note is marked superseded so the next session does not
re-investigate. Documentation only — no code or behaviour change.

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant