Repository navigation
perf(sycl): run the SpEED twins on the device, bit-exact with the CPU - #1620
Merged
Merged
Conversation
lusoris
force-pushed
the
perf/sycl-speed-device-resident
branch
from
September 29, 2026 06:58
205930c to
ebbec87
Compare
speed_chroma_sycl and speed_temporal_sycl no longer round-trip through the host inside a frame. Before, they filtered and decimated each plane on the CPU and ran the 25x25 eigenvalue problem, the QR factorisation and the Q^T B multiply there too, waiting on the queue 11 and 9 times per frame. Now each frame is one upload of the raw planes, one replayed SYCL graph and one read of the result. The whole chain lives in the new speed_sycl_pipeline.cpp, shared by both extractors, and mirrors speed.c operation for operation in fp32. To make that exact on the device the SpEED TUs build with -ffp-contract=off, and division, square root and log2 are rounded correctly in the source: -fp-model=precise alone still contracts FMAs and leaves fp32 division non-correctly-rounded (measured with icpx 2026.1). The fp64 comparisons of the reference have exact fp32 forms. Every per-frame speed_chroma_u/v/uv and speed_temporal score now equals --backend cpu on the Netflix pair and BBB 4K, on an Arc B580 and a UHD 770; before, 1 to 9 frames per run matched. On the B580, speed_temporal at 4K drops from 60.4 to 7.6 ms per frame and speed_chroma from 23.3 to 7.5; at 576x324 both run under a millisecond. scripts/dev/speed_gpu_parity.py checks and times any GPU twin against the CPU. The CUDA and HIP twins keep the host split; porting them is tracked in docs/state.md with a verify command for the maintainer's machine. ADR-1358. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Register the ADR-1358 citations the SpEED device pipeline adds, merge the RC3 disposition row with the cambi IDs that landed on master, and regenerate the ADR nav, by-tag pages and CHANGELOG after the rebase. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s exit status The assertion-density gate flagged five host functions in speed_sycl_pipeline.cpp with no assert. They now state the invariants the kernels rely on: the mean and covariance windows end exactly at the truncated plane, blocks tile it, channels come in reference/distorted pairs, the graph cache stays within its slots, and the sample width is one of the two picture_copy() paths. The parity-script test returned the importlib-loaded module's main() result, which mypy types as Any; it now returns int(status). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
force-pushed
the
perf/sycl-speed-device-resident
branch
from
September 29, 2026 09:01
587a81b to
3a18a23
Compare
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
speed_chroma_syclandspeed_temporal_syclnow run the whole per-frame SpEED chain on the device, with no host compute and no queue wait inside a frame, and their scores are bit-identical to--backend cpu. Each frame is one upload of the raw planes, one replayed SYCL graph and one read of the result (ADR-1358). This answers the maintainer's RC3 requirement: "there shouldnt be any gpu cpu rountrips" / "remove the host roundtrips and then test again".Before, both twins filtered and decimated every plane on the CPU and ran the 25x25 eigenvalue problem, the QR factorisation and the Q^T B multiply there, waiting on the queue 11 and 9 times per frame. They matched the CPU on 1 to 9 of 48/50 frames per run, by up to 4.2e-5.
What made exactness possible, measured with icpx 2026.1 on an Arc B580 and a UHD 770 (details in Research-1358):
-fp-model=precisedoes not stop FMA contraction inside kernel lambdas (28% of randoma * b + cdiffer from the host). The SpEED TUs now build with-ffp-contract=offafter it./andsqrtare not correctly rounded, and-foffload-fp32-prec-div/-sqrtonly work on the final image link, which every SYCL extractor shares. The pipeline rounds both itself: the hardware approximation, one refinement, then an exact-residual choice between the two adjacent floats, with the oneAPI math extension'sfdiv_rn/fsqrt_rnas the fallback (0 mismatches on 16.7M operands). Usingfdiv_rneverywhere was exact but 15x slower.log2fis evaluated in fp32 pairs and rounded once. The fp64 constantEIGENVALUE_EPSequals0x1.0c6f7ap-20f + 0x1.6bdb1ap-49fexactly, so every fp64 comparison inspeed.chas an exact fp32 form. The covariance sum is carried in exact fp32 pairs.This also explains the one-ulp mean that
T-SYCL-SPEED-CHROMA-V-PARITY-2026-09-16could not: the device divided with a non-correctly-rounded division.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. Run per file instead, all clean: clang-format 23.1.1, clang-tidy 22.1.8 throughscripts/ci/clang-tidy-sycl.sh(0 findings in the four SpEED TUs andspeed_internal.c), icpx warnings (none), lizard CCN <= 10 / 60 LOC, ruff + black, markdownlint. cppcheck was not available in the container.python3 scripts/ci/run_meson_test.py -- -C build. The SpEED tests were run (below); the full suite was not, because this workstation's LTO link of the static test executables fails on master as well (test_context,test_pool_percentile)./cross-backend-diffand the worst ULP is ≤ 2. Ran the per-frame equivalent,scripts/dev/speed_gpu_parity.py --backend sycl: 0 ULP on every output of every frame, both GPUs..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. Not a breaking change.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-SYCL-SPEED-HOST-ROUNDTRIPS-2026-09-29closed; RC3 rows opened forT-CUDA-SPEED-HOST-RESIDUAL-2026-09-29andT-HIP-SPEED-HOST-RESIDUAL-2026-09-29(exact call sites, port reference, contract, verify-and-time command forryzen-4090-arc) andT-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29; the Metal gap row points at the new port reference.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.speed.cis untouched.Cross-backend numerical results
Per-frame comparison against
--backend cpu --threads 16at--precision max, Netflixsrc01_hrc00/01_576x324(48 frames) and BBB 3840x2160 (50 frames). Max abs diff, bit-identical frames in brackets:Options on the B580 at 576x324 (12 frames): bit-identical for
speed_prescale=2.0withnearest,bilinearandbicubic,speed_prescale=0.7:bilinear(temporal),speed_kernelscale=2.0,speed_nn_floor=0.3:speed_weight_var_mode=5,speed_weight_var_mode=3,speed_use_ref_diff=true, and 10-bit input.speed_prescale_method=lanczos4differs by at most 1.8e-5 (the CPU weights use fp64sin), inside the ADR-0214 tolerance; documented.Performance (if
perforfeat)Milliseconds per frame,
(t(22) - t(2)) / 20, median of 3; at 576x324(t(48) - t(2)) / 46, median of 5, because the 20-frame difference is inside run-to-run noise there. i9-12900K, Arc B580 and UHD 770 through WSL2 Level Zero,vmaf-dev-mcpimage (icx/icpx 2026.1), GPU runs under the shared/f/gpu.lock. CPU rows:--backend cpu --threads 16 --feature speed_*; GPU rows:--backend sycl --feature speed_*_sycl.speed_chromaspeed_chromaspeed_temporalspeed_temporalAt 4K both twins are now at roughly the rate the CLI reads 4K frames.
speed_temporal, which the CPU cannot run frames in parallel for, is about five times faster than the CPU. At 576x324 the eigenvalue sweep, which has to run on one work item to keep the reference's operation order, holds the B580 at about 1 ms per frame.Measurement note:
--feature speed_chroma --backend syclruns the CPU extractor (single-threaded): 18.35 ms per frame at 4K, which is the figure previously quoted as the B580 baseline. The twins must be named,speed_chroma_sycl/speed_temporal_sycl.Deep-dive deliverables (ADR-0108)
docs/research/1358-sycl-speed-device-resident.md.## Alternatives considered.AGENTS.mdinvariant note —core/src/feature/sycl/AGENTS.md(SpEED pipeline arithmetic contract,-fp-model=precisecorrection, kernel-identity note),core/src/sycl/AGENTS.md, index entry indocs/development/rebase-sensitive-invariants.md.changelog.d/changed/perf-sycl-speed-device-resident.md.docs/rebase-notes.md, "perf/sycl-speed-device-resident — SYCL SpEED twins are device-resident (ADR-1358)".Reproducer
Local gates:
standardsctl compile-context --verifyin sync;standardsctl hiss coverage --verify18/18 claims hold;make dedupe-check100%;standardsctl auditreports 0 findings in touched files and fails on 147 pre-existing unbaselined findings in untouched files, identical on a pristineorigin/mastertree.Known follow-ups
T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29,T-HIP-SPEED-HOST-RESIDUAL-2026-09-29(RC3; each row lists the call sites to replace, the port reference inspeed_sycl_pipeline.cpp, the contract, andpython3 scripts/dev/speed_gpu_parity.py --backend cuda|hipforryzen-4090-arc). Not changed here: they cannot be tested on this machine.-fp-model=preciseonly, so it inherits the contraction and division rounding differences:T-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29.ocloc, so both builds JIT-compile SPIR-V), cppcheck, and the full meson suite.🤖 Generated with Claude Code