Repository navigation
docs(research): cross-backend 4K baseline + PR #79 adm_cm A/B at 4K (Research-0751) - #90
Merged
Merged
Conversation
…Research-0751) Establishes the first measured 4K (3840x2160) CUDA throughput baseline on RTX 4090 and A/B tests the PR #79 adm_cm __launch_bounds__(128,8) change at 4K resolution. Key findings: - vif CUDA: 147 fps, adm CUDA: 161 fps, motion CUDA: 176 fps (24-frame medians) - filter1d_8_horizontal: fully saturated at 4K (253 waves, 69.7% active warps) versus 0.84 waves at 576p. PR #76 optimization is fully expressed at 4K. - adm_cm __launch_bounds__: zero gain at 4K (-0.3%, noise) vs -9.3% at 1080p. The register-bound regime ends at ~32 waves (1080p boundary). At 4K (32.2 waves) the scheduler is wave-saturated regardless of register count. - ms_ssim_decimate scale 0: 88.1% active warps at 4K (126 waves) -- smem-tiling revert from Research-0749 confirmed correct at 4K. Deliverables: Research-0751, changelog fragment, rebase-notes sentinel, state.md update, research README entry (also resolves pre-existing conflict marker). no rebase impact: research/docs-only change, no source code modified. no per-surface docs needed: no user-discoverable surface changed. no ADR needed: measurement digest, no architectural decision. no AGENTS.md update needed: no rebase-sensitive invariants. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
vmaf_bench)adm_cm_line_kernel_8 __launch_bounds__(128,8)at 4K using the same container-patch methodology as Research-0749Key findings
filter1d_8_horizontal (PR #76 target): 32,400 CTAs / 128 SMs = 253 waves at 4K, 69.7% active warps. The 0.84-wave launch-width limit from 576p is completely gone. PR #76 is fully expressed at production 4K.
adm_cm_line_kernel_8 (PR #79
__launch_bounds__): −0.3% kernel duration at 4K (noise) vs −9.3% at 1080p. The register-bound regime that the optimization targets is 8–32 waves; at 4K (32.2 waves) the scheduler is wave-saturated. The optimization remains beneficial and zero-cost at 4K — recommendation unchanged: ship the__launch_bounds__change.ms_ssim_decimate (scale 0 at 4K): 88.1% active warps, 126 waves. Smem-tiling revert from Research-0749 is confirmed correct — the kernel is already L1-resident at all scales (>99.5% hit rate).
Deliverables
docs/research/0751-cross-backend-4k-baseline-and-pr79-adm-cm-4k-measure.mdchangelog.d/changed/cross-backend-4k-baseline.mdReproducer
Full ncu reproducer in Research-0751 §7.
🤖 Generated with Claude Code