Repository navigation
perf(sycl): read the shared frame in psnr, psnr_hvs and motion_v2 - #1634
Merged
Merged
Conversation
lusoris
force-pushed
the
perf/sycl-psnr-hvs-light-twins-4k
branch
from
September 29, 2026 16:20
f5ef1b1 to
07d617f
Compare
The SYCL psnr_hvs, psnr and motion_v2 twins now read the planes the SYCL state uploads once per frame instead of converting and uploading their own copies. Scores are bit-identical to the previous twins on an Arc B580 and a UHD 770. - Opt-in shared Cb/Cr planes in common.cpp (vmaf_sycl_shared_chroma_init / _upload, vmaf_sycl_get_shared_plane): the first chroma-reading twin of a frame packs the chroma into pinned staging and uploads it with one DMA per plane; later twins reuse it. Luma-only runs never allocate chroma. - vmaf_sycl_queue_after_upload() gives twins on their own queue the input barriers the combined graph gets; a device-side slot fence orders each upload after the last readers of the slot it overwrites. - psnr_hvs: no host float conversion or private upload; two work-items per 8x8 block, one dispatch for all planes, per-block float expressions unchanged. 9- and 11-bit input now scores the raw sample like the CPU. - motion_v2: runs the ADR-1371 pipeline on the shared luma and keeps the frame through its cur_copy; no host copy or private upload. - psnr: chroma from the shared planes; one atomic per work-group. At 3840x2160, psnr_hvs drops from 17.1 to 7.6 ms per frame on the B580 and from 124 to 60 on the UHD 770, where psnr drops from 25.3 to 12.3. ADR-1369, Research-1369. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
force-pushed
the
perf/sycl-psnr-hvs-light-twins-4k
branch
from
September 30, 2026 09:45
07d617f to
ed5b7b3
Compare
This was referenced Sep 30, 2026
lusoris
added a commit
that referenced
this pull request
Oct 7, 2026
… to #1668 Add a "Confirmed not-affected" row for each of the fork's open upstream pull requests from #1631 to #1668 (25 rows), naming the fork file, test or ADR that shows the fork already carries the fix, covers it another way or is not affected. #1643 is the one open item (a test-only x87 comparison with the same line in the fork). #1634 is recorded as closed: the fork keeps integer AIM unclipped, as upstream defines it. Also records that the ten pull requests that conflicted with upstream master acdd9376e were rebased on 2026-10-07. Documentation only.
lusoris
added a commit
that referenced
this pull request
Oct 7, 2026
… to #1668 (#2404) * docs(state): reconcile the fork's open Netflix/vmaf pull requests #1631 to #1668 Add a "Confirmed not-affected" row for each of the fork's open upstream pull requests from #1631 to #1668 (25 rows), naming the fork file, test or ADR that shows the fork already carries the fix, covers it another way or is not affected. #1643 is the one open item (a test-only x87 comparison with the same line in the fork). #1634 is recorded as closed: the fork keeps integer AIM unclipped, as upstream defines it. Also records that the ten pull requests that conflicted with upstream master acdd9376e were rebased on 2026-10-07. Documentation only. * docs: regenerate the indexes and the citation map after rebasing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The SYCL
psnr_hvs,psnrandmotion_v2twins now read the planes the SYCL state uploads once per frame instead of converting and uploading their own copies, and their kernels are reworked. At 3840x2160,psnr_hvsgoes from 17.1 to 7.6 ms per frame on an Arc B580 and from 124 to 60 on a UHD 770; on the UHD 770psnrgoes from 25.3 to 12.3. Every per-frame score is bit-identical to master on both devices. Rebased onto #1624 and #1628; #1628 (ADR-1371) movedmotion_v2_syclonto the shared motion pipeline, so this PR keeps only its host-side change there. Design: ADR-1369; profile and evidence: Research-1369.What the profile found (master, 4K, B580, event timing per phase):
psnr_hvs_syclspent 14.5 ms converting all six planes to float on the host and 7.2 ms uploading 99.5 MB for a 6.3 ms kernel;motion_v2_syclre-uploaded the luma the shared frame already held;psnr_syclandpsnr_hvs_sycleach uploaded their own chroma. What changed:core/src/sycl/common.cpp): opt-in Cb/Cr planes (vmaf_sycl_shared_chroma_init/_upload,vmaf_sycl_get_shared_plane), uploaded once per frame by the first twin that reads chroma, packed into pinned staging with one DMA per plane;vmaf_sycl_queue_after_uploadfor twins on their own queue; a device-side slot fence so an upload never overwrites a slot whose readers are still running (it matters forn_subsample).cur_copy; no host copy, no private upload. (A kernel variant with the vertical taps once per tile column, 13.2 → 6.3 ms on the UHD 770, bit-exact, was dropped on rebase and is recorded as a pipeline follow-up.)Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. — Ran the parts this change touches, not the full targets: clang-format 23.1.1 on every changed C/C++ file; clang-tidy 22.1.8 throughscripts/ci/clang-tidy-sycl.shon the four SYCL TUs (common.cpp 60 findings on master and on this branch; the three feature TUs 0) and on the three changed tests (0); markdownlint on the changed docs;make docs-fragments-check;check-source-adr-citations.py;check-state-md-rows.sh;assertion-density.sh. No cppcheck (the CI job builds with-Denable_sycl=false, so none of these files are in its database).python3 scripts/ci/run_meson_test.py -- -C build --suite faston the AOT build before the fix(sycl): make motion_sycl bit-exact with the CPU motion #1628 rebase (235 OK, 1 skipped;test_gpu_picture_pool_uafwas SIGKILLed under memory pressure in the parallel run and passes alone). After the rebase, on both GPUs:test_sycl_shared_planes,test_sycl_psnr_hvs_parity{,_large},test_sycl_psnr_parity,test_sycl_motion_v2_parity{,_large},test_sycl_motion_tiny_frames,test_sycl_motion3_parity,test_sycl_twin_option_parity,test_sycl_init_unwind, andtest_sycl_kernel_source_contract.py(12 passed)./cross-backend-diffand the worst ULP is ≤ 2. — Bit-identical to master (0 ULP) on the B580 and the UHD 770; psnr and motion_v2 equal the CPU; psnr_hvs keeps its ADR-1361 distance (table below)..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. — Not breaking: internalcommon.hAPI only; public headers unchanged.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR. Closed:T-SYCL-LIGHT-TWIN-HOST-ROUNDTRIPS-2026-09-29,T-SYCL-PSNR-HVS-ODD-BPC-SCALE-2026-09-29,T-SYCL-SHARED-SLOT-SUBSAMPLE-WAR-2026-09-29. Opened RC3:T-CUDA-PSNR-HVS-HOST-ROUNDTRIP-2026-09-29,T-HIP-PSNR-HVS-HOST-CONVERT-2026-09-29,T-GPU-MOTION-V2-INT64-VERTICAL-2026-09-29,T-HIP-TWIN-PRIVATE-PLANE-UPLOADS-2026-09-29,T-SYCL-PAGEABLE-UPLOAD-HOST-STAGING-2026-09-29,T-SYCL-PSNR-HVS-XE-LP-THROUGHPUT-2026-09-29; RC2:T-SYCL-SHARED-FRAME-STICKY-GEOMETRY-2026-09-29. The RC2 dispositions cell that master had emptied lists its open rows again.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
--precision max, every metric of every frame, this branch against master's SYCL twin and against--backend cpu:Performance (if
perforfeat)3840x2160 Big Buck Bunny, ms per frame, (t(22) − t(2)) / 20, median of 5 interleaved reps (base and new back to back),
--backend sycl -n --feature <cpu-name>, WSL2, CPU shared with other builds:b20472dcb, ADR-1371 kernel)The 22-frame differential swings by about ±3 ms on the B580 at 4K, where these twins sit at the CLI's frame-read rate. Over 100 frames (t(102) − t(2)) / 100, same method: motion_v2 against
b20472dcb6.11 → 5.39 on the B580 and 14.73 → 15.11 on the UHD 770 (median of 3; its kernel is ADR-1371's and dominates there), motion 6.57 → 6.82, float_psnr 7.52 → 7.35, default model 45.41 → 45.46 — no regression. The other rows compare against2d9d5b069; #1624 and #1628 do not touch psnr_hvs, psnr, float_psnr or the default model's twins except motion. At 576x324 (48 frames) every wall-time difference is inside the timer noise; the per-phase profile shows each twin's per-frame work shrinking there too (psnr_hvs on the UHD 770: kernel 2.42 → 1.14 ms).Per-phase profile, 4K, ms per frame (event timing; full tables in Research-1369):
A first version copied chroma straight from the pageable picture with
ext_oneapi_memcpy2d; the timing pass caught it costing 5.1 ms per 576x324 frame on the B580 and 168 ms on the UHD 770, and the upload now packs into pinned staging.Crash check for T-SYCL-PSNR-HVS-B580-SIGSEGV: the AOT build for
bmg-g21,adl-scompiles the new kernel (also a scratch variant withreqd_sub_group_size(32)); a SPIR-V JIT-only build runstest_sycl_psnr_hvs_parityon both devices by default, withIGC_ForceOCLSIMDWidth=32and with=16(NEO_CACHE_PERSISTENT=0), and the JIT SIMD32 4K output equals the AOT output.Deep-dive deliverables (ADR-0108)
docs/research/1369-sycl-shared-planes-light-twins.md.## Alternatives considered.AGENTS.mdinvariant note —core/src/sycl/AGENTS.md(shared planes, slot fence, no direct pageable chroma copy) andcore/src/feature/sycl/AGENTS.md(psnr_hvs kernel shape, motion_v2 32-bit bounds, psnr reduction).changelog.d/changed/perf-sycl-shared-planes-light-twins.md,changelog.d/fixed/sycl-psnr-hvs-odd-bit-depth.md.docs/rebase-notes.mdentryperf/sycl-psnr-hvs-light-twins-4k.Reproducer
Known follow-ups
docs/state.mdwith file:function, the SYCL design to port and a verify-and-time command forryzen-4090-arc):psnr_hvs_cudaround-trips every plane device → host → float → device (integer_psnr_hvs_cuda.cupload_frame);psnr_hvs_hipconverts on the host; both keep the thread-0-serial kernel; CUDA/HIPmotion_v2kernels recompute the vertical taps per pixel in 64 bits; every HIP twin uploads its own copy of the planes.psnr_cudaalready reads the device picture and reduces per warp.T-GPU-MOTION-V2-INT64-VERTICAL-2026-09-29for SYCL, CUDA and HIP.T-SYCL-PAGEABLE-UPLOAD-HOST-STAGING-2026-09-29, touches the pool the CLI read-ahead work also changes).T-SYCL-PSNR-HVS-XE-LP-THROUGHPUT-2026-09-29).T-SYCL-SHARED-FRAME-STICKY-GEOMETRY-2026-09-29(pre-existing, API only): a SYCL state reused for a second frame size keeps the first size's shared frame.🤖 Generated with Claude Code