Skip to content

feat(cuda): evaluate two ADM viewing distances in one adm_cuda instance - #2647

Merged
lusoris merged 1 commit into
masterfrom
port/upstream-cffd5b77-cuda
Oct 9, 2026
Merged

lusoris merged 1 commit into
masterfrom
port/upstream-cffd5b77-cuda

Conversation

@lusoris

@lusoris lusoris commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Summary

The CUDA twin of integer ADM evaluates two viewing distances in one instance, the second step of the Q-298 stack (ADR-2795). This covers adm_norm_view_dist_extra and two models such as vmaf_v1.0.16_3d0h and _5d0h on --backend cuda. The wavelet transform runs once per scale. The denominator, CSF, contrast-masking and AIM kernels run once per distance. Both distances return the CPU's scores bit for bit.

The merge callback and the name dictionary move into a shared helper. It reads the options by name, so the CPU extractor, adm_rust and adm_cuda use one implementation (HISS-19). The SYCL, HIP and Metal twins will plug into it.

Stacked on #2646 (C + Rust). This branch contains its commit; it rebases to one commit once #2646 lands.

Maintainer decision Q-298 (2026-10-08): "D4 nvde, cffd5b77d + 33e5f0aca: full scope — C + Rust mirror (ADR-1713) AND bit-identical on CUDA, SYCL, HIP and Metal now (no DEFAULT_ONLY). Stack it as needed (C+Rust first, then one PR per backend is fine), each with == parity tests and exact_twins entries; device runs under the locks; Metal evidence is outside hardware — mark pending with the tester path if no Mac is available."

What changed

  • core/src/feature/adm_view_dist.{c,h} (new): vmaf_adm_merge_view_dist(), vmaf_adm_extend_name_dict(), vmaf_adm_view_count(), vmaf_adm_view_dist() and the key tables.
  • core/src/feature/integer_adm.c: drops its static copies of these and points .merge / .extend_name_dict at the helpers. The scores are unchanged.
  • core/src/feature/cuda/integer_adm_cuda.c:
    • Option. adm_norm_view_dist_extra (nvde), the same table entry as the CPU's.
    • Per-frame driver. It is split into adm_scale0_transform() / adm_scale123_transform() (DWT) and adm_scale0_weigh() / adm_scale123_weigh() (per distance, through adm_view_buffer(), which swaps in the second block's result slots).
    • Results. tmp_res and results_host hold two RES_BUFFER_SIZE blocks.
    • Per-distance code. Fixed parameters are built per distance. The denominator contexts and the host conclusion (adm_cm_scale_result(), adm_csf_den_scale_result(), adm_dlm_terms(), adm_aim_num()) take the distance as an argument. The CSF configuration is checked for every distance before device resources are claimed.
    • Filing. write_scores() files each distance, the second under <base>:nvde, with no debug scores. Init extends the name dictionary through the helper.
  • Tests:
    • test_cuda_adm_parity.c: the recorded option gap is gone, and MAX_KEYS goes to 32. test_adm_two_views_exact covers 8-bit and 10-bit with debug, and the model options at 3H/5H. test_adm_merged_registrations_exact registers adm_cuda twice, as two models do, against one CPU context. Both compare with ==.
    • test_adm_view_merge.c: test_every_adm_descriptor_merges. A registered ADM descriptor must merge exactly when its row says so (adm, adm_cuda yes; SYCL/HIP/Metal flip in their pull requests), and each merging descriptor folds nvd 5 into nvd 3 on its own option layout.
    • test_adm_view_dist_contract.py: keys read from the helper. The non-feature-parameter option set is checked for every merging table (adm, adm_cuda).
    • test_cuda_adm_exact_contract.py: the pinned denominator-context calls take the distance argument nvd instead of s->adm_norm_view_dist. The launch and the result must still both use them.
  • Docs and notes:
    • docs/metrics/adm.md: option row and backends.
    • core/src/feature/AGENTS.d/adm-view-distance-merge.md and core/src/feature/cuda/AGENTS.d/adm.md.
    • The changelog fragment and the rebase note.

Evidence

Check Result
test_cuda_adm_parity on the RTX 4090 (CUDA 13.4.92, the pinned release) 14 of 14. The existing 12 still hold the CPU's bits. Both new tests compare 25 or 14 values with ==
Planted defects (5): adm_cuda without merge; every view's kernels at the first distance; second view writing the first view's slots; host DLM conclusion at the first distance; no second-distance names each fails test_cuda_adm_parity or test_adm_view_merge
CLI on CUDA, --precision max, vmaf_v1.0.16_3d0h + _5d0h together vs each alone, 576x324 pair, 48 frames 1560 of 1560 identical
Same CUDA two-model run vs the CPU two-model run 1352 of 1352 identical
Run receipt (feature_backends) of the two-model run one adm_cuda (master registers two adm)
test_adm_view_merge, test_integer_adm_view_dist, test_adm_view_dist_contract.py 11 of 11, 3 of 3, 4 of 4
rust_twin_diff.py --feature adm, Netflix + checkerboard, threads 0 and 4 28 cells EQUAL (the Rust shim now calls the shared helper)
Fast suite (CPU + CUDA + Rust build) 504 OK, 2 expected fail, 0 fail
Netflix golden gate (GOLDEN_NINJA_JOBS=4) 280 passed, 3 skipped
tidy: cpu adm_view_dist.c integer_adm.c test_adm_view_merge.c 0 findings
tidy: cuda integer_adm_cuda.c test_cuda_adm_parity.c 0 findings (two bugprone-multi-level-implicit-pointer-conversion and one braces finding fixed)
scripts/dev/preflight.sh --stage msvcism pass

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • Every commit is signed off (git commit -s; fix a branch with git rebase --signoff origin/master). See DCO sign-off.
  • make format && make lint is green locally. Not run as one target. These are green: the commit hooks and tidy above.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. Fast suite, see Evidence.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. 0 ULP: the twin's outputs equal the CPU's (== tests and the CLI comparison).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not a breaking change.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and nothing else is touched for the index. No new ADR: ADR-2795 (port(upstream): libvmaf/adm: share computation across viewing distances (cffd5b77d + 33e5f0aca) #2646) decides the GPU twins' step.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred). Not needed — no state delta: a feature step of the upstream port; it opens and closes no bug.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests, except by porting Netflix's own updated assertion verbatim from upstream (value and places as upstream has them, measured against the fork's CPU build first; ADR-1828).
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. No golden value changes.

Cross-backend numerical results

adm_cuda against the CPU, both distances: 0 differences. That holds in test_cuda_adm_parity (25 values with debug at 8 and 10 bits, 14 values with the model options) and on the 48-frame CLI run (1352 values). adm.cuda stays declared exact (scripts/ci/exact_twins.d/adm.cuda).

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: ADR-2795 (port(upstream): libvmaf/adm: share computation across viewing distances (cffd5b77d + 33e5f0aca) #2646) covers the design; this pull request applies it to the CUDA twin.
  • Decision matrix — no alternatives: ADR-2795's matrix covers the choice; the per-distance driver mirrors the CPU's split.
  • AGENTS.md invariant note — core/src/feature/cuda/AGENTS.d/adm.md and core/src/feature/AGENTS.d/adm-view-distance-merge.md.
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/added/adm-cuda-shared-viewing-distances.md.
  • Rebase note — docs/rebase-notes.d/adm-cuda-shared-viewing-distances.md.

Reproducer

meson setup build core -Db_lto=false -Denable_cuda=true && ninja -C build
flock ~/.cache/vmafx-locks/cuda-4090.lock timeout 300 build/test/test_cuda_adm_parity
build/test/test_adm_view_merge
build/tools/vmaf --precision max -r testdata/ref_576x324_48f.yuv -d testdata/dis_576x324_48f.yuv \
  -w 576 -h 324 -p 420 -b 8 --backend cuda \
  --model version=vmaf_v1.0.16_3d0h --model version=vmaf_v1.0.16_5d0h --json -o both-cuda.json

Known follow-ups

  • SYCL, HIP and Metal adm twins: one pull request each, stacked on this one.

…ce (#2647)

* feat(cuda): evaluate two ADM viewing distances in one adm_cuda instance

adm_cuda takes adm_norm_view_dist_extra and the merge callback: the
scale-0 and scale-1..3 DWT run once per frame, and the denominator, CSF,
contrast-masking and AIM kernels run once per distance into that
distance's result block (tmp_res and results_host hold two). The host
concludes each distance with its own CPU contexts and files the second
under the <base>:nvde keys. Both distances return the CPU's bits
(ADR-2795; test_adm_two_views_exact, test_adm_merged_registrations_exact).

The merge callback and the second distance's names move into
adm_view_dist.c, which reads the options by name at each descriptor's
own offsets, so the CPU extractor, the Rust twin and adm_cuda share one
implementation; integer_adm.c drops its copies.

Signed-off-by: Lusoris <lusoris@proton.me>
@lusoris
lusoris force-pushed the port/upstream-cffd5b77-cuda branch from 10d0013 to 8cc34b5 Compare October 9, 2026 08:47
@lusoris
lusoris merged commit 8cc34b5 into master Oct 9, 2026
17 of 40 checks passed
@lusoris
lusoris deleted the port/upstream-cffd5b77-cuda branch October 9, 2026 08:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant