Repository navigation
perf(hip): upload native samples and convert on device in psnr_hvs - #1658
Merged
Merged
Conversation
17 of 26 tasks
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before, it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the ADR-1361 parity tolerance (3.34e-3 dB). calc_psnrhvs() adds every masked coefficient error of a plane into one running float, so its result depends on the order of the additions. The twin summed the 64 terms of a block on the device and the blocks on the host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation; a double accumulator on the CPU was ruled out because it moves a Netflix golden value past places=4. - The kernel stores the 64 terms of every block, computed in the CPU's arithmetic: masking table and threshold in double, integer coefficient difference, fatbin built with --fmad=false. - psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms the combined score and the dB value with the CPU's expressions. It is built with the strict floating-point arguments of the scalar reference. - The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0 at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361. - test_cuda_psnr_hvs_parity asserts equality on all four outputs, including two 3840x2160 cases that fail on master by 1.6e-2 dB. test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are device-free. Measured on an RTX 4090 at --precision max: every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores are unchanged. The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of 2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09), and 256 bytes of readback per block. Tuning is tracked as T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins are unchanged until #1658 and #1657 land (T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01). Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris
force-pushed
the
perf/hip-psnr-hvs-device-convert
branch
from
October 1, 2026 08:18
a731cee to
f2881b1
Compare
Port ADR-1369 design to the HIP psnr_hvs feature extractor, closing T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29. - core/src/feature/hip/integer_psnr_hvs_hip.c: replace host float conversion loops with native sample upload via vmaf_hip_picture_upload(). Discard unused pinned staging allocations (h_uint_ref/dist). - core/src/feature/hip/integer_psnr_hvs/psnr_hvs_score.hip: take raw sample planes and convert to integers directly on device. Resolves a latent scaling bug where 9-bit and 11-bit depths were multiplied by 16. - core/test/test_hip_psnr_hvs_parity.c: add test_psnr_hvs_deep_parity covering 9, 10, 11, and 12-bit input parity against CPU reference. - docs/state.md: close T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29 with measured parity and timing evidence on ryzen-4090-arc.
Fuse all plane dispatches into a single kernel launch in psnr_hvs_score.hip (n_dispatches_per_frame = 1), matching the SYCL twin's ADR-1369 architecture: - Pass plane descriptors and buffer pointers via PsnrHvsHipKernelArgs - Launch one grid covering all blocks across all 3 planes - Reduce 4K frame time from 221.79 ms to 18.90 ms/frame on AMD gfx1036 - Keep 576x324 delta within 8.37e-05 dB and 4K delta within area-scaled tolerance - Update docs, rebase notes, CHANGELOG and docs/state.md
lusoris
force-pushed
the
perf/hip-psnr-hvs-device-convert
branch
from
October 1, 2026 08:29
f2881b1 to
640e5c0
Compare
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before, it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the ADR-1361 parity tolerance (3.34e-3 dB). calc_psnrhvs() adds every masked coefficient error of a plane into one running float, so its result depends on the order of the additions. The twin summed the 64 terms of a block on the device and the blocks on the host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation; a double accumulator on the CPU was ruled out because it moves a Netflix golden value past places=4. - The kernel stores the 64 terms of every block, computed in the CPU's arithmetic: masking table and threshold in double, integer coefficient difference, fatbin built with --fmad=false. - psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms the combined score and the dB value with the CPU's expressions. It is built with the strict floating-point arguments of the scalar reference. - The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0 at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361. - test_cuda_psnr_hvs_parity asserts equality on all four outputs, including two 3840x2160 cases that fail on master by 1.6e-2 dB. test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are device-free. Measured on an RTX 4090 at --precision max: every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores are unchanged. The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of 2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09), and 256 bytes of readback per block. Tuning is tracked as T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins are unchanged until #1658 and #1657 land (T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01). Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before, it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the ADR-1361 parity tolerance (3.34e-3 dB). calc_psnrhvs() adds every masked coefficient error of a plane into one running float, so its result depends on the order of the additions. The twin summed the 64 terms of a block on the device and the blocks on the host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation; a double accumulator on the CPU was ruled out because it moves a Netflix golden value past places=4. - The kernel stores the 64 terms of every block, computed in the CPU's arithmetic: masking table and threshold in double, integer coefficient difference, fatbin built with --fmad=false. - psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms the combined score and the dB value with the CPU's expressions. It is built with the strict floating-point arguments of the scalar reference. - The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0 at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361. - test_cuda_psnr_hvs_parity asserts equality on all four outputs, including two 3840x2160 cases that fail on master by 1.6e-2 dB. test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are device-free. Measured on an RTX 4090 at --precision max: every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores are unchanged. The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of 2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09), and 256 bytes of readback per block. Tuning is tracked as T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins are unchanged until #1658 and #1657 land (T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01). Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris
added a commit
that referenced
this pull request
Oct 1, 2026
) * fix(cuda): reproduce the CPU's running float sum in psnr_hvs_cuda psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before, it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the ADR-1361 parity tolerance (3.34e-3 dB). calc_psnrhvs() adds every masked coefficient error of a plane into one running float, so its result depends on the order of the additions. The twin summed the 64 terms of a block on the device and the blocks on the host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy the CPU's accumulation; a double accumulator on the CPU was ruled out because it moves a Netflix golden value past places=4. - The kernel stores the 64 terms of every block, computed in the CPU's arithmetic: masking table and threshold in double, integer coefficient difference, fatbin built with --fmad=false. - psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms the combined score and the dB value with the CPU's expressions. It is built with the strict floating-point arguments of the scalar reference. - The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0 at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361. - test_cuda_psnr_hvs_parity asserts equality on all four outputs, including two 3840x2160 cases that fail on master by 1.6e-2 dB. test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are device-free. Measured on an RTX 4090 at --precision max: every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores are unchanged. The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of 2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09), and 256 bytes of readback per block. Tuning is tracked as T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins are unchanged until #1658 and #1657 land (T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01). Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30. * docs: regenerate the indexes and the citation map after rebasing
17 of 26 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ports the ADR-1369 psnr_hvs design to the HIP backend (closing
T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29):h_uint_refandh_uint_dist).vmaf_hip_picture_upload()/ stage plane).uint8_t) and wide 9–12 bpc (uint16_t) inpsnr_hvs_score.hip.n_dispatches_per_frame = 1), reducing 4K frame time from 221.79 ms to 18.90 ms/frame on AMD gfx1036.test_psnr_hvs_deep_parity.Type
perf— performance improvementhip/cuda/simd— backend-specificChecklist
make format && make lintis green locally.python3 scripts/ci/run_meson_test.py -- -C build./cross-backend-diffand the worst ULP is ≤ 2..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred), ORno state delta: REASON.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
Performance (if
perforfeat)Measured on AMD gfx1036 (
(t(22) - t(2)) / 20ms/frame, median of 3):Deep-dive deliverables (ADR-0108)
## Alternatives considered(or in the digest), OR "no alternatives: only-one-way fix".AGENTS.mdinvariant note — added to the relevant package'sAGENTS.md, OR "no rebase-sensitive invariants".changelog.d/<section>/<topic>.md(added/changed/deprecated/removed/fixed/security). Do not editCHANGELOG.mddirectly —scripts/release/concat-changelog-fragments.shrenders the Unreleased block from the fragment tree (see ADR-0221).docs/rebase-notes.mdunder a new ID, ORno rebase impact: REASON.Reproducer
python3 scripts/dev/speed_gpu_parity.py --backend hip --feature psnr_hvs --max-abs-diff 0.02 --vmaf $(pwd)/build/tools/vmaf --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbbKnown follow-ups
None. Closes
T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29.Breaking changes / migration
None.