Skip to content

perf(cuda): psnr_hvs — __ldg() + __restrict__ + __launch_bounds__(64) (ADR-0764) - #563

Merged
lusoris merged 1 commit into
masterfrom
perf/cuda-psnr-hvs-ldg-launch-bounds-rb107
Jun 3, 2026
Merged

lusoris merged 1 commit into
masterfrom
perf/cuda-psnr-hvs-ldg-launch-bounds-rb107

Conversation

@lusoris

@lusoris lusoris commented Jun 3, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Apply the F3 __ldg() + __restrict__ pointer extraction + __launch_bounds__(64) pattern to the psnr_hvs CUDA kernel (core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu), PR docs(cuda): ADR-0756 F3 struct-by-value audit — 20 kernels, top-5 dispatch #96 candidate chore(deps): Update fedora:45 Docker digest to 0c1f63e #5.
  • Extract const float *__restrict__ ref_buf and dist_buf from VmafCudaBuffer struct args before the 64-thread cooperative tile load, routing all 128 reads through the L1 read-only texture cache via __ldg().
  • Add __launch_bounds__(64) matching the actual 8×8 block dispatch.
  • No arithmetic change; bit-identical scores (ADR-0214 places=4). Predicted -3 to -5% kernel duration at >=1080p, mirroring ADR-0754 1080p -4.2% result.

Rebased from original #107 branch onto current master.

Six deep-dive deliverables

  • Research digest: docs/research/0764-cuda-psnr-hvs-ldg-launch-bounds-2026-05-29.md
  • Decision matrix: ADR-0764 ## Alternatives considered
  • AGENTS.md invariant: core/src/feature/cuda/AGENTS.md — psnr_hvs __ldg() tile load invariant
  • Smoke-test: meson test -C build --suite=fast (CPU) — no CUDA gate required; arithmetic is unchanged
  • Changelog: changelog.d/perf/cuda-psnr-hvs-ldg-launch-bounds.md
  • Rebase notes: docs/rebase-notes.md — ADR-0764 section added

Test plan

  • CPU fast suite passes (no arithmetic change)
  • CUDA psnr_hvs bit-exact check: scores identical at places=4 vs pre-patch (ADR-0214 gate)
  • CI: Linux GCC + Coverage + Perf + CodeQL + Aggregator

no rebase impact: psnr_hvs_score.cu is fork-local; no upstream C/Python changes

Closes #107

🤖 Generated with Claude Code

… (ADR-0764)

Apply the F3 struct-by-value fix (PR #96 candidate #5) to the psnr_hvs
CUDA kernel in core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu:

- Extract `const float *__restrict__ ref_buf` and `dist_buf` from the
  VmafCudaBuffer struct args before the cooperative 64-thread tile load.
- Apply `__ldg()` to both per-thread element reads, routing them through
  the L1 read-only texture cache (LDG.E.CONSTANT in SASS).
- Add `__launch_bounds__(64)` matching the actual 8x8 block dispatch.

No arithmetic change; bit-identical scores (ADR-0214 places=4). Predicted
-3 to -5% kernel duration at >=1080p (mirrors ADR-0754 1080p -4.2% result).

Six deep-dive deliverables: ADR-0764, Research-0764, changelog.d/perf,
AGENTS.md invariant extended, rebase-notes.md entry, state.md row.

Rebased from #107 onto current master.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings June 3, 2026 15:36
@lusoris
lusoris merged commit 4c2c3f0 into master Jun 3, 2026
19 of 38 checks passed
@lusoris
lusoris deleted the perf/cuda-psnr-hvs-ldg-launch-bounds-rb107 branch June 3, 2026 15:37
@lusoris
lusoris removed the request for review from Copilot June 3, 2026 15:59
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant