Repository navigation
docs(research): hardware measurement of CUDA ms_ssim_decimate + adm_cm (Research-0749 / ADR-0750) - #89
Merged
Conversation
9 tasks done
lusoris
enabled auto-merge (squash)
May 28, 2026 22:14
lusoris
added a commit
that referenced
this pull request
May 28, 2026
…op ms_ssim smem tiling Per PR #89 hardware measurement on RTX 4090 (vmaf-dev-mcp:cuda13.3): - REVERT: ms_ssim_decimate smem tiling. The baseline kernel was already L1-resident (95% hit rate). Cooperative load + __syncthreads() broke the hardware prefetcher and raised kernel duration +8–24% at all tested resolutions (576p and 1080p). The non-tiled kernel is restored. - KEEP: adm_cm_line_kernel_8 __launch_bounds__(128, 8). Confirmed −9.3% kernel duration at 1080p; registers 114→64 per thread. End-to-end net: +3.9–4.8% fps driven solely by this occupancy gain. ADR-0744 updated from stub to Accepted with full revert rationale and alternatives table. docs/state.md row added. Correctness: ADR-0214 places=4 parity gate confirmed bit-exact before and after (max_diff=0.0, per ADR-0750 Research-0749). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
May 28, 2026
…op ms_ssim smem tiling Per PR #89 hardware measurement on RTX 4090 (vmaf-dev-mcp:cuda13.3): - REVERT: ms_ssim_decimate smem tiling. The baseline kernel was already L1-resident (95% hit rate). Cooperative load + __syncthreads() broke the hardware prefetcher and raised kernel duration +8–24% at all tested resolutions (576p and 1080p). The non-tiled kernel is restored. - KEEP: adm_cm_line_kernel_8 __launch_bounds__(128, 8). Confirmed −9.3% kernel duration at 1080p; registers 114→64 per thread. End-to-end net: +3.9–4.8% fps driven solely by this occupancy gain. ADR-0744 updated from stub to Accepted with full revert rationale and alternatives table. docs/state.md row added. Correctness: ADR-0214 places=4 parity gate confirmed bit-exact before and after (max_diff=0.0, per ADR-0750 Research-0749). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
May 28, 2026
…(ADR-0744) (#79) * perf(cuda): ms_ssim_decimate smem tiling + adm_cm register reduction (ADR-0744) Opt A: ms_ssim_decimate shared-memory tile Converts 81 global/L2 reads per output pixel to L1 hits via a (2*BLOCK_X+2*LPF_HALF+1) x (2*BLOCK_Y+2*LPF_HALF) = 41x24 float smem tile per CTA (3936 B). Cooperative load applies mirror_idx once in the load phase only; the hot 9x9 LPF convolution loop reads smem at tile[2*ty+kv][2*tx+ku] unconditionally. TILE_W_PAD=41 (+1 pad) prevents bank aliasing per the ADR-0454 / filter1d.cu convention. Estimated: -30 to -40% DRAM throughput; +68-93% local speedup at 1080p+. Opt B: adm_cm_line_kernel_8 __launch_bounds__(128, 8) Hints ptxas to target <=64 regs/thread (65536/(8x128)) matching the fused scale 1-3 kernel, raising theoretical occupancy from 33% to ~67% on Ampere. Block size 128 (BLOCKX=32 x BLOCKY=4) is fixed by the host launcher in integer_adm_cuda.c. Estimated: +66.7% local kernel throughput at 1080p+. Both changes are correctness-neutral; verified by ADR-0214 parity gate at places=4. ncu measurement commands documented in Research-0744. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * perf(cuda): partial revert PR #79 — keep adm_cm __launch_bounds__, drop ms_ssim smem tiling Per PR #89 hardware measurement on RTX 4090 (vmaf-dev-mcp:cuda13.3): - REVERT: ms_ssim_decimate smem tiling. The baseline kernel was already L1-resident (95% hit rate). Cooperative load + __syncthreads() broke the hardware prefetcher and raised kernel duration +8–24% at all tested resolutions (576p and 1080p). The non-tiled kernel is restored. - KEEP: adm_cm_line_kernel_8 __launch_bounds__(128, 8). Confirmed −9.3% kernel duration at 1080p; registers 114→64 per thread. End-to-end net: +3.9–4.8% fps driven solely by this occupancy gain. ADR-0744 updated from stub to Accepted with full revert rationale and alternatives table. docs/state.md row added. Correctness: ADR-0214 places=4 parity gate confirmed bit-exact before and after (max_diff=0.0, per ADR-0750 Research-0749). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
…m (Research-0749 / ADR-0750) Add hardware benchmark results for PR #79 (ms_ssim_decimate smem tiling + adm_cm register reduction) on 1080p content, research digest 0749, and ADR-0750 documenting the measurement methodology. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris
force-pushed
the
docs/cuda-ms-ssim-decimate-adm-cm-measure-0749
branch
from
May 28, 2026 22:51
6801d45 to
02ef22f
Compare
8 of 11 tasks
lusoris
added a commit
that referenced
this pull request
Oct 7, 2026
…report CSV with _wfsopen (ADR-1113) (#2424) * refactor(interop): re-vendor Pelorus at the commit that opens the qp-report CSV with _wfsopen (ADR-1113) VMAFx/pelorus #89 (fixing #88) moves open_utf8() from the deprecated _wfopen() to _wfsopen(..., _SH_DENYNO), the local edit the MSVC zero-warnings series carried in core/src/interop/pelorus_qp_report_csv.c. PELORUS_VENDOR_SHA moves to 4aae30711c65 and --update re-renders the ten vendored files; the mirror is byte-identical to pelorus again apart from the banner and the include rewrite, and the drift check passes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
perf/cuda-ms-ssim-decimate-adm-cm-ncu-driven-20260528adm_cm_line_kernel_8 __launch_bounds__: confirmed -9.3% kernel duration at 1080p, registers 114→64 (-44%). Keep.ms_ssim_decimatesmem tiling: measured +8-24% kernel regression (baseline L1 hit rate already 95%; cooperative-load breaks hardware prefetcher). Recommend revert.Test plan
--privilegedcontainer (RmProfilingAdminOnly=1 on host)float_ms_ssim_cuda+ vmaf model, max_diff=0.0 for all 48f/3fReproducer
Deep-dive checklist (ADR-0108)
docs/research/0749-cuda-ms-ssim-decimate-adm-cm-1080p-measure.md## Alternatives consideredchangelog.d/perf/cuda-ms-ssim-adm-cm-measure-0749.mddocs/rebase-notes.mdupdated with adm_cm register budget invariantno rebase impact: doc-only, no source changes
🤖 Generated with Claude Code