Repository navigation
perf(cuda): compact nonzero terms before psnr_hvs readback (ADR-1397) - #1730
Merged
Merged
Conversation
lusoris
force-pushed
the
perf/cuda-psnr-hvs-tune
branch
from
October 1, 2026 16:39
ea7b383 to
418c0e9
Compare
Compacts nonzero terms on device before readback to recover throughput while preserving bit-for-bit identity with CPU running float sum: - psnr_hvs_score.cu: computes a 64-bit mask and popcount per block, hvs_scan_reduce / hvs_scan_prefix / hvs_compact scatters nonzero terms into a contiguous device buffer and prepends a 16-byte header with count. - integer_psnr_hvs_cuda.c: transfers 16-byte header asynchronously, then reads back only total_terms * sizeof(float). - psnr_hvs_score.c: added vmaf_psnr_hvs_plane_score_compacted to add packed terms in CPU order. - test_psnr_hvs_score.c: added unit tests for all-zero block, all-zero plane, plane starting with zeros, and sign-of-zero behavior. - docs/state.md: closed T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01.
lusoris
force-pushed
the
perf/cuda-psnr-hvs-tune
branch
from
October 1, 2026 17:00
418c0e9 to
d52d35d
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Recovers
psnr_hvs_cudathroughput on RTX 4090 (3.39 ms at 3840x2160, beating sixteen CPU threads at 6.70 ms and recovering from 12.16 ms) without altering a single bit of any score. Compacting nonzero terms on device via intra-block mask and prefix scans shrinks readback from 64.8 MB to 11.01 MB per 4K frame, eliminating host dependency overhead while strictly preserving the CPU's exact running float sum.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally.python3 scripts/ci/run_meson_test.py -- -C build./cross-backend-diffand the worst ULP is ≤ 2..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred), ORno state delta: REASON.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
Bit-identical across all test pairs at
--precision max:max_abs_diff=0.000e+00 OK(0 ULP)Performance (if
perforfeat)Measured on RTX 4090 (
zeus, uptime 23:15, load average ~14.5), median of three runs of(t(N) - t(2)) / (N - 2):Target met: 4K frame time is 3.39 ms$\le 6.5$ ms (CPU reference).
Deep-dive deliverables (ADR-0108)
docs/state.mdrowT-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01and ADR-1397.AGENTS.mdinvariant note — no rebase-sensitive invariants: internal CUDA kernel tuning and device compaction preserving bit-exact CPU sum contract.changelog.d/changed/perf-cuda-psnr-hvs-device-compact.md.Reproducer
Known follow-ups
SYCL and HIP twins (
T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01) remain tracked for similar compaction tuning once PR #1692 lands.