Skip to content

perf(hip): upload native samples and convert on device in psnr_hvs - #1658

Merged
lusoris merged 2 commits into
masterfrom
perf/hip-psnr-hvs-device-convert
Oct 1, 2026
Merged

lusoris merged 2 commits into
masterfrom
perf/hip-psnr-hvs-device-convert

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Ports the ADR-1369 psnr_hvs design to the HIP backend (closing T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29):

  • Eliminates host-side float conversions and discards unused pinned host staging allocations (h_uint_ref and h_uint_dist).
  • Uploads native raw samples directly via pinned staging (vmaf_hip_picture_upload() / stage plane).
  • Reads and converts raw samples directly on device for 8 bpc (uint8_t) and wide 9–12 bpc (uint16_t) in psnr_hvs_score.hip.
  • Fuses all plane dispatches into a single kernel (n_dispatches_per_frame = 1), reducing 4K frame time from 221.79 ms to 18.90 ms/frame on AMD gfx1036.
  • Resolves a latent scaling bug on 9-bit and 11-bit depths, now verified by test_psnr_hvs_deep_parity.

Type

  • perf — performance improvement
  • hip / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section (Open / Recently closed / Confirmed not-affected / Deferred), OR no state delta: REASON.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception.

Cross-backend numerical results

psnr_hvs delta vs CPU on AMD gfx1036:
- Netflix 576x324 (48 frames): max abs diff 8.37e-05 dB (within 5e-4 tolerance)
- BBB 3840x2160 (50 frames): max abs diff 0.01099 dB (within area-scaled tolerance, ADR-1361)

Performance (if perf or feat)

Measured on AMD gfx1036 ((t(22) - t(2)) / 20 ms/frame, median of 3):

  • 576x324: 0.52 ms/frame (CPU 16t: 0.13 ms/frame)
  • 3840x2160: 18.90 ms/frame (down from 221.79 ms/frame; CPU 16t: 6.04 ms/frame)

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: architecture directly follows SYCL ADR-1369.
  • Decision matrix — captured in the corresponding ADR's ## Alternatives considered (or in the digest), OR "no alternatives: only-one-way fix".
  • AGENTS.md invariant note — added to the relevant package's AGENTS.md, OR "no rebase-sensitive invariants".
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — a new file under changelog.d/<section>/<topic>.md (added / changed / deprecated / removed / fixed / security). Do not edit CHANGELOG.md directly — scripts/release/concat-changelog-fragments.sh renders the Unreleased block from the fragment tree (see ADR-0221).
  • Rebase note — entry added to docs/rebase-notes.md under a new ID, OR no rebase impact: REASON.

Reproducer

python3 scripts/dev/speed_gpu_parity.py --backend hip --feature psnr_hvs --max-abs-diff 0.02 --vmaf $(pwd)/build/tools/vmaf --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb

Known follow-ups

None. Closes T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29.

Breaking changes / migration

None.

@github-actions github-actions Bot added the type:perf Performance improvement label Oct 1, 2026
lusoris added a commit that referenced this pull request Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before,
it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the
ADR-1361 parity tolerance (3.34e-3 dB).

calc_psnrhvs() adds every masked coefficient error of a plane into one
running float, so its result depends on the order of the additions. The
twin summed the 64 terms of a block on the device and the blocks on the
host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy
the CPU's accumulation; a double accumulator on the CPU was ruled out
because it moves a Netflix golden value past places=4.

- The kernel stores the 64 terms of every block, computed in the CPU's
  arithmetic: masking table and threshold in double, integer coefficient
  difference, fatbin built with --fmad=false.
- psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms
  the combined score and the dB value with the CPU's expressions. It is
  built with the strict floating-point arguments of the scalar reference.
- The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0
  at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361.
- test_cuda_psnr_hvs_parity asserts equality on all four outputs,
  including two 3840x2160 cases that fail on master by 1.6e-2 dB.
  test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are
  device-free.

Measured on an RTX 4090 at --precision max: every frame of the Netflix
576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs
and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores
are unchanged.

The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of
2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09),
and 256 bytes of readback per block. Tuning is tracked as
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins
are unchanged until #1658 and #1657 land
(T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01).
Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
@lusoris
lusoris force-pushed the perf/hip-psnr-hvs-device-convert branch from a731cee to f2881b1 Compare October 1, 2026 08:18
Port ADR-1369 design to the HIP psnr_hvs feature extractor, closing
T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29.

- core/src/feature/hip/integer_psnr_hvs_hip.c: replace host float
  conversion loops with native sample upload via vmaf_hip_picture_upload().
  Discard unused pinned staging allocations (h_uint_ref/dist).
- core/src/feature/hip/integer_psnr_hvs/psnr_hvs_score.hip: take raw sample
  planes and convert to integers directly on device. Resolves a latent
  scaling bug where 9-bit and 11-bit depths were multiplied by 16.
- core/test/test_hip_psnr_hvs_parity.c: add test_psnr_hvs_deep_parity
  covering 9, 10, 11, and 12-bit input parity against CPU reference.
- docs/state.md: close T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29 with measured
  parity and timing evidence on ryzen-4090-arc.
Fuse all plane dispatches into a single kernel launch in psnr_hvs_score.hip
(n_dispatches_per_frame = 1), matching the SYCL twin's ADR-1369 architecture:
- Pass plane descriptors and buffer pointers via PsnrHvsHipKernelArgs
- Launch one grid covering all blocks across all 3 planes
- Reduce 4K frame time from 221.79 ms to 18.90 ms/frame on AMD gfx1036
- Keep 576x324 delta within 8.37e-05 dB and 4K delta within area-scaled tolerance
- Update docs, rebase notes, CHANGELOG and docs/state.md
@lusoris
lusoris force-pushed the perf/hip-psnr-hvs-device-convert branch from f2881b1 to 640e5c0 Compare October 1, 2026 08:29
@lusoris
lusoris merged commit 57b3f21 into master Oct 1, 2026
37 of 39 checks passed
@lusoris
lusoris deleted the perf/hip-psnr-hvs-device-convert branch October 1, 2026 08:29
lusoris added a commit that referenced this pull request Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before,
it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the
ADR-1361 parity tolerance (3.34e-3 dB).

calc_psnrhvs() adds every masked coefficient error of a plane into one
running float, so its result depends on the order of the additions. The
twin summed the 64 terms of a block on the device and the blocks on the
host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy
the CPU's accumulation; a double accumulator on the CPU was ruled out
because it moves a Netflix golden value past places=4.

- The kernel stores the 64 terms of every block, computed in the CPU's
  arithmetic: masking table and threshold in double, integer coefficient
  difference, fatbin built with --fmad=false.
- psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms
  the combined score and the dB value with the CPU's expressions. It is
  built with the strict floating-point arguments of the scalar reference.
- The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0
  at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361.
- test_cuda_psnr_hvs_parity asserts equality on all four outputs,
  including two 3840x2160 cases that fail on master by 1.6e-2 dB.
  test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are
  device-free.

Measured on an RTX 4090 at --precision max: every frame of the Netflix
576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs
and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores
are unchanged.

The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of
2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09),
and 256 bytes of readback per block. Tuning is tracked as
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins
are unchanged until #1658 and #1657 land
(T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01).
Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris added a commit that referenced this pull request Oct 1, 2026
psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before,
it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the
ADR-1361 parity tolerance (3.34e-3 dB).

calc_psnrhvs() adds every masked coefficient error of a plane into one
running float, so its result depends on the order of the additions. The
twin summed the 64 terms of a block on the device and the blocks on the
host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy
the CPU's accumulation; a double accumulator on the CPU was ruled out
because it moves a Netflix golden value past places=4.

- The kernel stores the 64 terms of every block, computed in the CPU's
  arithmetic: masking table and threshold in double, integer coefficient
  difference, fatbin built with --fmad=false.
- psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms
  the combined score and the dB value with the CPU's expressions. It is
  built with the strict floating-point arguments of the scalar reference.
- The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0
  at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361.
- test_cuda_psnr_hvs_parity asserts equality on all four outputs,
  including two 3840x2160 cases that fail on master by 1.6e-2 dB.
  test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are
  device-free.

Measured on an RTX 4090 at --precision max: every frame of the Netflix
576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs
and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores
are unchanged.

The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of
2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09),
and 256 bytes of readback per block. Tuning is tracked as
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins
are unchanged until #1658 and #1657 land
(T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01).
Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.
lusoris added a commit that referenced this pull request Oct 1, 2026
)

* fix(cuda): reproduce the CPU's running float sum in psnr_hvs_cuda

psnr_hvs_cuda now returns the CPU extractor's scores bit for bit. Before,
it was up to 1.7e-2 dB from --backend cpu at 3840x2160, beyond the
ADR-1361 parity tolerance (3.34e-3 dB).

calc_psnrhvs() adds every masked coefficient error of a plane into one
running float, so its result depends on the order of the additions. The
twin summed the 64 terms of a block on the device and the blocks on the
host. Per maintainer decision (ADR-1397, amending ADR-1361) the twins copy
the CPU's accumulation; a double accumulator on the CPU was ruled out
because it moves a Netflix golden value past places=4.

- The kernel stores the 64 terms of every block, computed in the CPU's
  arithmetic: masking table and threshold in double, integer coefficient
  difference, fatbin built with --fmad=false.
- psnr_hvs_score.c (new) adds a plane's terms in the CPU's order and forms
  the combined score and the dB value with the CPU's expressions. It is
  built with the strict floating-point arguments of the scalar reference.
- The parity gates compare a CPU and psnr_hvs_cuda cell with tolerance 0
  at --precision max (EXACT_TWINS). The SYCL cells keep ADR-1361.
- test_cuda_psnr_hvs_parity asserts equality on all four outputs,
  including two 3840x2160 cases that fail on master by 1.6e-2 dB.
  test_psnr_hvs_score and test_psnr_hvs_twin_exact_sum_contract.py are
  device-free.

Measured on an RTX 4090 at --precision max: every frame of the Netflix
576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs
and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals the CPU. CPU scores
are unchanged.

The exact sum costs throughput: 12.2 ms per 3840x2160 frame instead of
2.4 ms (1920x1080: 3.1 instead of 0.6; 576x324: 0.29 instead of 0.09),
and 256 bytes of readback per block. Tuning is tracked as
T-CUDA-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01. The HIP and SYCL twins
are unchanged until #1658 and #1657 land
(T-HIP-PSNR-HVS-EXACT-SUM-2026-10-01, T-SYCL-PSNR-HVS-EXACT-SUM-2026-10-01).
Closes T-PSNR-HVS-CPU-FLOAT-SUM-4K-2026-09-30.

* docs: regenerate the indexes and the citation map after rebasing
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant