Skip to content

docs(research): cross-backend 4K baseline + PR #79 adm_cm A/B at 4K (Research-0751) - #90

Merged
lusoris merged 1 commit into
masterfrom
research/cross-backend-4k-baseline-20260529
May 29, 2026
Merged

lusoris merged 1 commit into
masterfrom
research/cross-backend-4k-baseline-20260529

Conversation

@lusoris

@lusoris lusoris commented May 28, 2026

Copy link
Copy Markdown
Contributor

Summary

Key findings

Feature CUDA fps (4K, 24f median) vs CPU
vif 147 fps 7.0×
adm 161 fps 2.3×
motion 176 fps 0.6× (launch overhead)

filter1d_8_horizontal (PR #76 target): 32,400 CTAs / 128 SMs = 253 waves at 4K, 69.7% active warps. The 0.84-wave launch-width limit from 576p is completely gone. PR #76 is fully expressed at production 4K.

adm_cm_line_kernel_8 (PR #79 __launch_bounds__): −0.3% kernel duration at 4K (noise) vs −9.3% at 1080p. The register-bound regime that the optimization targets is 8–32 waves; at 4K (32.2 waves) the scheduler is wave-saturated. The optimization remains beneficial and zero-cost at 4K — recommendation unchanged: ship the __launch_bounds__ change.

ms_ssim_decimate (scale 0 at 4K): 88.1% active warps, 126 waves. Smem-tiling revert from Research-0749 is confirmed correct — the kernel is already L1-resident at all scales (>99.5% hit rate).

Deliverables

  • Research digest: docs/research/0751-cross-backend-4k-baseline-and-pr79-adm-cm-4k-measure.md
  • Changelog fragment: changelog.d/changed/cross-backend-4k-baseline.md
  • Rebase notes: sentinel added (no rebase impact)
  • state.md updated
  • research/README.md: 0751 entry added + pre-existing merge-conflict marker resolved
  • no ADR needed: measurement digest, no architectural decision
  • no AGENTS.md update needed: no rebase-sensitive invariants
  • no per-surface docs needed: no user-discoverable surface changed

Reproducer

docker run --rm --gpus all --entrypoint bash \
  -v /path/to/.corpus:/corpus:ro \
  vmaf-dev-mcp:cuda13.3 -c '
    mkdir -p /tmp/vmaf_test && \
    ln -sf /corpus/netflix/ref/BigBuckBunny_25fps.yuv /tmp/vmaf_test/ref_3840x2160.yuv && \
    ln -sf /corpus/netflix/dis/BigBuckBunny_85_1080_3800.yuv /tmp/vmaf_test/dis_3840x2160.yuv && \
    /build/vmaf/core/build/tools/vmaf_bench --resolution 3840x2160 --frames 24
  '

Full ncu reproducer in Research-0751 §7.

🤖 Generated with Claude Code

…Research-0751)

Establishes the first measured 4K (3840x2160) CUDA throughput baseline on RTX 4090
and A/B tests the PR #79 adm_cm __launch_bounds__(128,8) change at 4K resolution.

Key findings:
- vif CUDA: 147 fps, adm CUDA: 161 fps, motion CUDA: 176 fps (24-frame medians)
- filter1d_8_horizontal: fully saturated at 4K (253 waves, 69.7% active warps)
  versus 0.84 waves at 576p. PR #76 optimization is fully expressed at 4K.
- adm_cm __launch_bounds__: zero gain at 4K (-0.3%, noise) vs -9.3% at 1080p.
  The register-bound regime ends at ~32 waves (1080p boundary). At 4K (32.2 waves)
  the scheduler is wave-saturated regardless of register count.
- ms_ssim_decimate scale 0: 88.1% active warps at 4K (126 waves) -- smem-tiling
  revert from Research-0749 confirmed correct at 4K.

Deliverables: Research-0751, changelog fragment, rebase-notes sentinel,
state.md update, research README entry (also resolves pre-existing conflict marker).

no rebase impact: research/docs-only change, no source code modified.
no per-surface docs needed: no user-discoverable surface changed.
no ADR needed: measurement digest, no architectural decision.
no AGENTS.md update needed: no rebase-sensitive invariants.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris enabled auto-merge (squash) May 28, 2026 22:57
@lusoris
lusoris merged commit 8930853 into master May 29, 2026
10 of 15 checks passed
@lusoris
lusoris deleted the research/cross-backend-4k-baseline-20260529 branch May 29, 2026 06:22
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant