Skip to content

perf(sycl): run the SpEED twins on the device, bit-exact with the CPU - #1620

Merged
lusoris merged 3 commits into
masterfrom
perf/sycl-speed-device-resident
Sep 29, 2026
Merged

lusoris merged 3 commits into
masterfrom
perf/sycl-speed-device-resident

Conversation

@lusoris

@lusoris lusoris commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

speed_chroma_sycl and speed_temporal_sycl now run the whole per-frame SpEED chain on the device, with no host compute and no queue wait inside a frame, and their scores are bit-identical to --backend cpu. Each frame is one upload of the raw planes, one replayed SYCL graph and one read of the result (ADR-1358). This answers the maintainer's RC3 requirement: "there shouldnt be any gpu cpu rountrips" / "remove the host roundtrips and then test again".

Before, both twins filtered and decimated every plane on the CPU and ran the 25x25 eigenvalue problem, the QR factorisation and the Q^T B multiply there, waiting on the queue 11 and 9 times per frame. They matched the CPU on 1 to 9 of 48/50 frames per run, by up to 4.2e-5.

What made exactness possible, measured with icpx 2026.1 on an Arc B580 and a UHD 770 (details in Research-1358):

  • -fp-model=precise does not stop FMA contraction inside kernel lambdas (28% of random a * b + c differ from the host). The SpEED TUs now build with -ffp-contract=off after it.
  • Device fp32 / and sqrt are not correctly rounded, and -foffload-fp32-prec-div/-sqrt only work on the final image link, which every SYCL extractor shares. The pipeline rounds both itself: the hardware approximation, one refinement, then an exact-residual choice between the two adjacent floats, with the oneAPI math extension's fdiv_rn / fsqrt_rn as the fallback (0 mismatches on 16.7M operands). Using fdiv_rn everywhere was exact but 15x slower.
  • log2f is evaluated in fp32 pairs and rounded once. The fp64 constant EIGENVALUE_EPS equals 0x1.0c6f7ap-20f + 0x1.6bdb1ap-49f exactly, so every fp64 comparison in speed.c has an exact fp32 form. The covariance sum is carried in exact fp32 pairs.

This also explains the one-ulp mean that T-SYCL-SPEED-CHROMA-V-PARITY-2026-09-16 could not: the device divided with a non-correctly-rounded division.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. Run per file instead, all clean: clang-format 23.1.1, clang-tidy 22.1.8 through scripts/ci/clang-tidy-sycl.sh (0 findings in the four SpEED TUs and speed_internal.c), icpx warnings (none), lizard CCN <= 10 / 60 LOC, ruff + black, markdownlint. cppcheck was not available in the container.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. The SpEED tests were run (below); the full suite was not, because this workstation's LTO link of the static test executables fails on master as well (test_context, test_pool_percentile).
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. Ran the per-frame equivalent, scripts/dev/speed_gpu_parity.py --backend sycl: 0 ULP on every output of every frame, both GPUs.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not a breaking change.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-SYCL-SPEED-HOST-ROUNDTRIPS-2026-09-29 closed; RC3 rows opened for T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29 and T-HIP-SPEED-HOST-RESIDUAL-2026-09-29 (exact call sites, port reference, contract, verify-and-time command for ryzen-4090-arc) and T-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29; the Metal gap row points at the new port reference.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. No golden value changes; speed.c is untouched.

Cross-backend numerical results

Per-frame comparison against --backend cpu --threads 16 at --precision max, Netflix src01_hrc00/01_576x324 (48 frames) and BBB 3840x2160 (50 frames). Max abs diff, bit-identical frames in brackets:

feature            size       B580 before        B580 after      UHD 770 before     UHD 770 after
speed_chroma_u     576x324    4.0e-05 (1/48)     0 (48/48)       4.0e-05 (1/48)     0 (48/48)
speed_chroma_v     576x324    3.1e-05 (5/48)     0 (48/48)       3.1e-05 (5/48)     0 (48/48)
speed_chroma_uv    576x324    3.0e-05 (2/48)     0 (48/48)       3.0e-05 (2/48)     0 (48/48)
speed_temporal     576x324    4.2e-05 (5/48)     0 (48/48)       4.2e-05 (5/48)     0 (48/48)
speed_chroma_u     3840x2160  6.2e-06 (6/50)     0 (50/50)       6.2e-06 (6/50)     0 (50/50)
speed_chroma_v     3840x2160  4.8e-06 (4/50)     0 (50/50)       4.8e-06 (4/50)     0 (50/50)
speed_chroma_uv    3840x2160  4.3e-06 (4/50)     0 (50/50)       4.3e-06 (4/50)     0 (50/50)
speed_temporal     3840x2160  2.9e-06 (9/50)     0 (50/50)       2.9e-06 (9/50)     0 (50/50)

Options on the B580 at 576x324 (12 frames): bit-identical for speed_prescale=2.0 with nearest, bilinear and bicubic, speed_prescale=0.7:bilinear (temporal), speed_kernelscale=2.0, speed_nn_floor=0.3:speed_weight_var_mode=5, speed_weight_var_mode=3, speed_use_ref_diff=true, and 10-bit input. speed_prescale_method=lanczos4 differs by at most 1.8e-5 (the CPU weights use fp64 sin), inside the ADR-0214 tolerance; documented.

Performance (if perf or feat)

Milliseconds per frame, (t(22) - t(2)) / 20, median of 3; at 576x324 (t(48) - t(2)) / 46, median of 5, because the 20-frame difference is inside run-to-run noise there. i9-12900K, Arc B580 and UHD 770 through WSL2 Level Zero, vmaf-dev-mcp image (icx/icpx 2026.1), GPU runs under the shared /f/gpu.lock. CPU rows: --backend cpu --threads 16 --feature speed_*; GPU rows: --backend sycl --feature speed_*_sycl.

Feature Size CPU16 B580 before B580 after UHD 770 before UHD 770 after
speed_chroma 576x324 0.16 3.60 0.89 17.31 3.44
speed_chroma 3840x2160 7.24 23.31 7.51 37.98 14.48
speed_temporal 576x324 0.85 2.19 0.83 8.32 3.47
speed_temporal 3840x2160 37.27 60.37 7.58 72.15 18.27

At 4K both twins are now at roughly the rate the CLI reads 4K frames. speed_temporal, which the CPU cannot run frames in parallel for, is about five times faster than the CPU. At 576x324 the eigenvalue sweep, which has to run on one work item to keep the reference's operation order, holds the B580 at about 1 ms per frame.

Measurement note: --feature speed_chroma --backend sycl runs the CPU extractor (single-threaded): 18.35 ms per frame at 4K, which is the figure previously quoted as the B580 baseline. The twins must be named, speed_chroma_sycl / speed_temporal_sycl.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1358-sycl-speed-device-resident.md.
  • Decision matrix — ADR-1358 ## Alternatives considered.
  • AGENTS.md invariant note — core/src/feature/sycl/AGENTS.md (SpEED pipeline arithmetic contract, -fp-model=precise correction, kernel-identity note), core/src/sycl/AGENTS.md, index entry in docs/development/rebase-sensitive-invariants.md.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/changed/perf-sycl-speed-device-resident.md.
  • Rebase note — docs/rebase-notes.md, "perf/sycl-speed-device-resident — SYCL SpEED twins are device-resident (ADR-1358)".

Reproducer

# Per-frame parity (exit 0 = every output of every frame identical) and timing, CPU vs SYCL twin:
ONEAPI_DEVICE_SELECTOR=level_zero:0 python3 scripts/dev/speed_gpu_parity.py --backend sycl \
  --vmaf build/tools/vmaf --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb
# SYCL parity gates (ran on both devices, 3/3 each) and the source contract:
python3 scripts/ci/run_meson_test.py -- -C build --print-errorlogs \
  test_sycl_speed_chroma_parity test_sycl_speed_temporal_parity test_sycl_speed_singular_parity \
  test_speed test_speed_simd test_speed_clamp_score test_speed_singular_tally test_sycl_kernel_source_contract
python3 -B scripts/dev/tests/test_speed_gpu_parity.py

Local gates: standardsctl compile-context --verify in sync; standardsctl hiss coverage --verify 18/18 claims hold; make dedupe-check 100%; standardsctl audit reports 0 findings in touched files and fails on 147 pre-existing unbaselined findings in untouched files, identical on a pristine origin/master tree.

Known follow-ups

  • CUDA and HIP twins still use the ADR-0567 host split and are not bit-exact: T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29, T-HIP-SPEED-HOST-RESIDUAL-2026-09-29 (RC3; each row lists the call sites to replace, the port reference in speed_sycl_pipeline.cpp, the contract, and python3 scripts/dev/speed_gpu_parity.py --backend cuda|hip for ryzen-4090-arc). Not changed here: they cannot be tested on this machine.
  • Every other SYCL twin is still built with -fp-model=precise only, so it inherits the contraction and division rounding differences: T-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29.
  • Not verified here: AdaptiveCpp builds (no math-extension fallback; they degrade to the ADR-0214 tolerance by design), the AOT device images (this container has no ocloc, so both builds JIT-compile SPIR-V), cppcheck, and the full meson suite.

🤖 Generated with Claude Code

@github-actions github-actions Bot added the type:perf Performance improvement label Sep 29, 2026
@lusoris
lusoris force-pushed the perf/sycl-speed-device-resident branch from 205930c to ebbec87 Compare September 29, 2026 06:58
lusoris and others added 3 commits September 29, 2026 10:59
speed_chroma_sycl and speed_temporal_sycl no longer round-trip through the
host inside a frame. Before, they filtered and decimated each plane on the
CPU and ran the 25x25 eigenvalue problem, the QR factorisation and the Q^T B
multiply there too, waiting on the queue 11 and 9 times per frame. Now each
frame is one upload of the raw planes, one replayed SYCL graph and one read
of the result.

The whole chain lives in the new speed_sycl_pipeline.cpp, shared by both
extractors, and mirrors speed.c operation for operation in fp32. To make that
exact on the device the SpEED TUs build with -ffp-contract=off, and division,
square root and log2 are rounded correctly in the source: -fp-model=precise
alone still contracts FMAs and leaves fp32 division non-correctly-rounded
(measured with icpx 2026.1). The fp64 comparisons of the reference have exact
fp32 forms. Every per-frame speed_chroma_u/v/uv and speed_temporal score now
equals --backend cpu on the Netflix pair and BBB 4K, on an Arc B580 and a
UHD 770; before, 1 to 9 frames per run matched.

On the B580, speed_temporal at 4K drops from 60.4 to 7.6 ms per frame and
speed_chroma from 23.3 to 7.5; at 576x324 both run under a millisecond.

scripts/dev/speed_gpu_parity.py checks and times any GPU twin against the
CPU. The CUDA and HIP twins keep the host split; porting them is tracked in
docs/state.md with a verify command for the maintainer's machine.

ADR-1358.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Register the ADR-1358 citations the SpEED device pipeline adds, merge the
RC3 disposition row with the cambi IDs that landed on master, and
regenerate the ADR nav, by-tag pages and CHANGELOG after the rebase.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s exit status

The assertion-density gate flagged five host functions in
speed_sycl_pipeline.cpp with no assert. They now state the invariants the
kernels rely on: the mean and covariance windows end exactly at the
truncated plane, blocks tile it, channels come in reference/distorted
pairs, the graph cache stays within its slots, and the sample width is one
of the two picture_copy() paths.

The parity-script test returned the importlib-loaded module's main() result,
which mypy types as Any; it now returns int(status).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant