Skip to content

perf(cuda): compact nonzero terms before psnr_hvs readback (ADR-1397) - #1730

Merged
lusoris merged 1 commit into
masterfrom
perf/cuda-psnr-hvs-tune
Oct 1, 2026
Merged

lusoris merged 1 commit into
masterfrom
perf/cuda-psnr-hvs-tune

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Recovers psnr_hvs_cuda throughput on RTX 4090 (3.39 ms at 3840x2160, beating sixteen CPU threads at 6.70 ms and recovering from 12.16 ms) without altering a single bit of any score. Compacting nonzero terms on device via intra-block mask and prefix scans shrinks readback from 64.8 MB to 11.01 MB per 4K frame, eliminating host dependency overhead while strictly preserving the CPU's exact running float sum.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred), OR no state delta: REASON.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception.

Cross-backend numerical results

Bit-identical across all test pairs at --precision max:

  • Netflix 576x324 8-bit: diff = 0.0 (0 ULP)
  • Netflix 576x324 10-bit: diff = 0.0 (0 ULP)
  • 1080p checkerboard 1-px: diff = 0.0 (0 ULP)
  • 1080p checkerboard 10-px: diff = 0.0 (0 ULP)
  • BBB 3840x2160 (200 frames): max_abs_diff=0.000e+00 OK (0 ULP)

Performance (if perf or feat)

Measured on RTX 4090 (zeus, uptime 23:15, load average ~14.5), median of three runs of (t(N) - t(2)) / (N - 2):

  • 576x324 (N = 960): 0.29 ms -> 0.11 ms/frame (CPU 16 threads: 0.15 ms), readback: 1.45 MB -> 0.22 MB (84.8% reduction)
  • 1920x1080 (N = 102): 3.06 ms -> 0.84 ms/frame (CPU 16 threads: 1.42 ms), readback: 16.2 MB -> 1.39 MB (91.4% reduction)
  • 3840x2160 (N = 102): 12.16 ms -> 3.39 ms/frame (CPU 16 threads: 6.70 ms), readback: 64.8 MB -> 11.01 MB (83.0% reduction)

Target met: 4K frame time is 3.39 ms $\le 6.5$ ms (CPU reference).

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: research captured in ADR-1397 and already landed on master in fix(cuda): reproduce the CPU's running float sum in psnr_hvs_cuda #1666.
  • Decision matrix — captured in docs/state.md row T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01 and ADR-1397.
  • AGENTS.md invariant note — no rebase-sensitive invariants: internal CUDA kernel tuning and device compaction preserving bit-exact CPU sum contract.
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/changed/perf-cuda-psnr-hvs-device-compact.md.
  • Rebase note — no rebase impact: internal CUDA kernel optimization preserving exact twin ABI and CPU float sum contract.

Reproducer

./build/test/test_psnr_hvs_score && ./build/test/test_cuda_psnr_hvs_parity && python3 core/test/test_psnr_hvs_twin_exact_sum_contract.py

Known follow-ups

SYCL and HIP twins (T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01) remain tracked for similar compaction tuning once PR #1692 lands.

@lusoris
lusoris force-pushed the perf/cuda-psnr-hvs-tune branch from ea7b383 to 418c0e9 Compare October 1, 2026 16:39
Compacts nonzero terms on device before readback to recover throughput while
preserving bit-for-bit identity with CPU running float sum:
- psnr_hvs_score.cu: computes a 64-bit mask and popcount per block,
  hvs_scan_reduce / hvs_scan_prefix / hvs_compact scatters nonzero terms
  into a contiguous device buffer and prepends a 16-byte header with count.
- integer_psnr_hvs_cuda.c: transfers 16-byte header asynchronously, then
  reads back only total_terms * sizeof(float).
- psnr_hvs_score.c: added vmaf_psnr_hvs_plane_score_compacted to add
  packed terms in CPU order.
- test_psnr_hvs_score.c: added unit tests for all-zero block, all-zero plane,
  plane starting with zeros, and sign-of-zero behavior.
- docs/state.md: closed T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01.
@lusoris
lusoris force-pushed the perf/cuda-psnr-hvs-tune branch from 418c0e9 to d52d35d Compare October 1, 2026 17:00
@lusoris
lusoris merged commit 71e8b7f into master Oct 1, 2026
66 of 76 checks passed
@lusoris
lusoris deleted the perf/cuda-psnr-hvs-tune branch October 1, 2026 17:00
@github-actions github-actions Bot added the type:perf Performance improvement label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant