Skip to content

perf(cuda): run cambi and SpEED entirely on the device - #1639

Merged
lusoris merged 9 commits into
masterfrom
perf/cuda-rc3-device-resident
Oct 1, 2026
Merged

lusoris merged 9 commits into
masterfrom
perf/cuda-rc3-device-resident

Conversation

@lusoris

@lusoris lusoris commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

cambi_cuda, speed_chroma_cuda and speed_temporal_cuda now run every per-frame stage on the GPU: each frame reads back one small result block and waits once, in collect(), and nothing is uploaded beyond the planes the CUDA engine already has on the device. On the RTX 4090 of ryzen-4090-arc, every frame of the Netflix 576x324 pair and of BBB 3840x2160 equals --backend cpu (icx build) to the last bit, and a 4K frame of any of the three now costs about what the trivial psnr_cuda costs. The rewrite also fixes two bugs on master: speed_temporal_cuda failed at 1920x1080 and above, and cambi_cuda crashed on wide, short frames.

  • cambi_cuda is device-resident (ADR-1379). The old twin downloaded the distorted picture, preprocessed it on the host, uploaded it again, and at each of five scales waited and read the image and mask back for host c-values and top-K pooling. Now twelve kernels (the ADR-1357 design) do all of it, and 88 bytes come back per frame. The twin also gains cambi.c's guard against windows above 65 x 65.
  • The CUDA SpEED twins are device-resident (ADR-1380). The host filtering, the 25x25 eigenvalue problem, the QR solve and the host score combine are gone: nine kernels in speed/speed_score.cu, shared by both extractors through speed_cuda_pipeline.c, and 40 bytes back per frame. Every rounding the CPU performs is spelled with __fadd_rn / __fmul_rn / __fdiv_rn / __fsqrt_rn, and the fatbin builds with --fmad=false.
  • Exact log2 on the device. The fp32-pair speed_log2() misrounds exactly 48 positive floats. Both device twins (CUDA and SYCL) now look those up in speed_log2_hard_cases.h. An exhaustive run over all 2 139 095 039 positive finite floats finds 0 misrounds on the RTX 4090 and 0 on the Arc A380.
  • One host setup (HISS-19). cambi.c exports the helpers both device twins need, and speed_internal_gpu_configure() sets up the SYCL and CUDA SpEED pipelines alike; the SYCL twins drop their private copies. CPU output does not change.
  • Bugs fixed on master's CUDA twins. speed_temporal_cuda launched its solve kernel with ((blocks + 7) / 8) * 32 threads per block, over the 1024 limit above 256 SpEED blocks, so at 1920x1080 and 3840x2160 every frame failed with CUDA_ERROR_INVALID_VALUE (T-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30). The new test test_cuda_speed_temporal_parity_1080p fails on 10f27efe2 with that error and passes here. cambi_cuda ran cambi.c's host c-values walk, so on the wide, short frames of cambi: out-of-bounds read and write on wide, short frames (e.g. 1920x128), C and AVX2 paths Netflix/vmaf#1628 it scored frame 0 wrongly and then segfaulted in close_fex_cuda(). The device twin runs those sizes clean under compute-sanitizer and equals the fixed CPU of fix(cambi): keep the c-values walks inside short frames #1642.
  • Tidy Ratchet. speed_gpu_common.h reaches core/tools/vmaf.cpp through feature_dimensions.h, so clang-tidy parsed its typedef structs as C++ (modernize-use-using, +6 on the CPU lane). The header now carries the cited, file-scoped NOLINTBEGIN(modernize-use-using) bracket that ADR-1138 prescribes for C headers parsed as C++.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally, run piecewise: clang-format on every touched C, C++, CUDA and header file; the full pre-commit set over origin/master..HEAD and the pre-push stage; the CPU tidy lane scoped to core/tools/vmaf.cpp (the TU that parses speed_gpu_common.h as C++) at 0 for speed_gpu_common.h; the CUDA lane baseline scripts/ci/tidy-baseline-cuda.json tightened for the rewritten TUs by the ratchet's scoped write.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build-cuda-icx --no-rebuild test_cuda_cambi_parity test_cuda_cambi_parity_large test_cuda_device_resident_contract test_cuda_speed_chroma_parity test_cuda_speed_temporal_parity test_cuda_speed_temporal_parity_1080p test_cuda_speed_singular_parity test_cuda_speed_chroma_smoke test_cuda_speed_temporal_smoke, all OK on the RTX 4090, none skipped.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. Per frame at --precision max against an icx build: 0 ULP on every cambi, speed_chroma_u/v/uv and speed_temporal output on both fixtures (tables below).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not breaking: no public API, CLI or option change.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Hooks: earlier pushes of this branch skipped hooks. This push ran the full lefthook and pre-commit set, pre-push stage included, with the repo-pinned praetorctl (f41e74d8, the engine CI's standards gate installs). The newer praetorctl build installed on the host during this run (ff7ea2c4) failed praetorctl audit on origin/master itself (428 infractions against the 185 baseline), so it could not gate a branch.

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR. Closed: T-CUDA-CAMBI-HOST-RESIDUAL-2026-09-29, T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29 and T-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30, with the measurements below. Opened: T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30 (RC3), T-SPEED-TEMPORAL-PRESCALE-UP-OVERFLOW-2026-09-30 (RC2), T-DEV-IMAGE-ICX-NATIVE-FMA-DRIFT-2026-09-30 (RC3) and T-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30 (RC3, verification blocked, see below), each in its disposition row.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. Not applicable: the cambi.c and speed_internal.c changes only move code into shared functions; the CPU CAMBI and SpEED scores are unchanged.

Cross-backend numerical results

Measured on ryzen-4090-arc: RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4), icx 2026.0 release build with -Db_lto=false and without -march=native (the state rows' reference build). "Before" is origin/master 10f27efe2 built the same way. python3 scripts/dev/speed_gpu_parity.py --backend cuda --vmaf $PWD/build-cuda-icx/tools/vmaf --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb (and --feature cambi), identical frames per output at --precision max:

Output 576x324 before 576x324 after 3840x2160 before 3840x2160 after
cambi 48/48 48/48 50/50 50/50
speed_chroma_u 5/48 48/48 9/50 50/50
speed_chroma_v 3/48 48/48 8/50 50/50
speed_chroma_uv 1/48 48/48 10/50 50/50
speed_temporal 8/48 48/48 run fails 50/50

Before, the largest difference was 4.0e-5. A gcc 16.2.1 build of this branch against its own CPU extractor (glibc 2.44 log2f) differs on 0 to 2 speed_chroma frames per output, by at most 1.4e-6, and matches on speed_temporal; the device rounds log2 correctly and glibc does not on those inputs.

  • Prescale (speed_prescale=0.5, BBB 3840x2160, 6 frames): nearest, bilinear and bicubic identical on every frame. lanczos4 is not: at most 1.9e-5 apart on BBB, but 2.1e-2 (5.7e-4 relative, speed_chroma_v) on an ffmpeg gradients clip with noise at 1920x1080, beyond the ADR-0214 tolerance (T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30, being fixed in this PR, see Known follow-ups).
  • Wide, short frames (banded 8-bit ramps at 1920x64, 1920x128, 1920x160, 3840x128): cambi_cuda equals the CPU build of fix(cambi): keep the c-values walks inside short frames #1642 on every frame (6.0056, 7.7220, 8.0805, 6.0946), compute-sanitizer --tool memcheck 0 errors. master's twin scored frame 0 differently, the later frames 0, and then segfaulted in close_fex_cuda() → vmaf_picture_unref() (gdb backtrace).
  • Sanitizers: compute-sanitizer memcheck, racecheck and synccheck on test_cuda_cambi_parity, test_cuda_cambi_parity_large, test_cuda_speed_singular_parity, test_cuda_speed_temporal_parity and test_cuda_speed_chroma_parity: 0 errors and 0 hazards on all 15 runs.
  • SYCL (Arc A380, shared host code changed): cambi_sycl equals --backend cpu on 48/48 (576x324) and 50/50 (BBB 4K) frames. The SYCL SpEED re-check is blocked: since its 18:04 boot this host drives the A380 with the xe kernel driver, under which, as measured by the fix(sycl): give every SYCL feature kernel the CPU's fp32 arithmetic #1630 track, SYCL kernels that use scratch memory return wrong values (standalone probes without vmafx code too, and origin/master fails the same 16 of 55 SYCL tests). The SpEED twins emit 0 there on master and on this branch alike (T-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30).

Performance (if perf or feat)

Per frame, counted with a CUPTI driver-API callback over frames 13 to 22 of the Netflix 576x324 pair. The engine's own calls (six picture uploads, one context and two event synchronisations) are the same in every row and left out:

Twin Launches Device-to-host Host-to-device Stream syncs
cambi_cuda before 19 11 copies, 1.1 MiB 1 7
cambi_cuda after 65 1 copy, 88 B 0 1
speed_chroma_cuda before 18 20 copies, 336 KiB 16 6
speed_chroma_cuda after 7 (8 with prescale) 1 copy, 40 B 0 1
speed_temporal_cuda before 9 10 copies, 659 KiB 8 6
speed_temporal_cuda after 7 (8 with prescale) 1 copy, 40 B 0 1

Milliseconds per frame, (t(N) - t(2)) / (N - 2), median of 3, before and after runs alternating so both see the same host load (other sessions kept the load average between 13 and 41). The rows' N = 22 was below the run-to-run noise at 576x324 on this host, so the table uses N = 102 at 3840x2160 and N = 402 at 576x324 (the Netflix pair looped ten times):

Twin Size Before After CPU, 16 threads
cambi_cuda 3840x2160 64.71 6.01 19.65
speed_chroma_cuda 3840x2160 24.90 6.89 9.11
speed_temporal_cuda 3840x2160 fails 5.88 26.74
cambi_cuda 576x324 1.59 0.38 0.09
speed_chroma_cuda 576x324 1.51 0.44 0.12
speed_temporal_cuda 576x324 2.81 0.42 0.92

With the rows' own N = 22 at 3840x2160: cambi_cuda 68.39 → 3.67, speed_chroma_cuda 14.29 → 5.81, speed_temporal_cuda fails → 5.06. The absolute numbers move with the host load; the relation to a trivial twin does not. In a run at load average 13 with N = 102, psnr_cuda took 2.56 and 2.41 ms per 4K frame, cambi_cuda 2.55, speed_chroma_cuda 2.75 and speed_temporal_cuda 2.73: at 4K the three twins now run at the cost of reading and uploading the pictures.

Deep-dive deliverables (ADR-0108)

  • Research digest — Research-1379: the CUDA rounding contract, the CPU reference's dependence on log2f and FMA contraction, the exact top-K sum, the lanczos4 and speed_temporal prescale findings, the exhaustive log2 replay (finding 7) and the RTX 4090 measurements with their commands (finding 8).
  • Decision matrix — in the ## Alternatives considered tables of ADR-1379 and ADR-1380.
  • AGENTS.md invariant note — core/src/feature/cuda/AGENTS.md "Device-resident CAMBI and SpEED (ADR-1379, ADR-1380)", including the speed_log2() hard-case table; index entry in docs/development/rebase-sensitive-invariants.md.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/changed/perf-cuda-cambi-device-resident.md and changelog.d/changed/perf-cuda-speed-device-resident.md.
  • Rebase note — docs/rebase-notes.md "ADR-1379 / ADR-1380 — CUDA CAMBI and SpEED run entirely on the device".

Reproducer

# Device-free, on any host:
python3 core/test/test_cuda_device_resident_contract.py

# On an NVIDIA host, from the repository root, with icx and without -march=native
# (a gcc build differs from its own CPU extractor in the last bits on a few SpEED frames):
. /opt/intel/oneapi/setvars.sh
CC=icx CXX=icpx meson setup build-cuda-icx core -Denable_cuda=true -Denable_nvcc=true \
  -Denable_sycl=false -Denable_hip=false --buildtype=release -Db_lto=false
ninja -C build-cuda-icx
python3 scripts/ci/run_meson_test.py -- -C build-cuda-icx test_cuda_cambi_parity \
  test_cuda_cambi_parity_large test_cuda_speed_chroma_parity test_cuda_speed_temporal_parity \
  test_cuda_speed_temporal_parity_1080p test_cuda_speed_singular_parity test_cuda_device_resident_contract
python3 scripts/dev/speed_gpu_parity.py --backend cuda --vmaf $PWD/build-cuda-icx/tools/vmaf \
  --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb
python3 scripts/dev/speed_gpu_parity.py --backend cuda --feature cambi --vmaf $PWD/build-cuda-icx/tools/vmaf \
  --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb

Known follow-ups

  • lanczos4 prescale (T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30): the device twins compute the kernel weights in fp32 with sinpif(), the CPU in fp64 with sin(). The weights depend only on the output column and row, so computing them once on the host with vif_tools.c's own function makes the CUDA twin exact; that change is being added to this PR. The SYCL twin keeps its fp32 weights until it can be verified on an Intel GPU with a working driver.
  • SYCL SpEED re-check of this PR's shared host setup (T-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30): run speed_gpu_parity.py --backend sycl on an Arc with a working driver (the A380 back on i915, or the office B580 / UHD 770).
  • HIP and Metal CAMBI and HIP SpEED keep their host residuals (T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29, T-METAL-CAMBI-HOST-RESIDUAL-2026-09-29, T-HIP-SPEED-HOST-RESIDUAL-2026-09-29).
  • The CPU speed_temporal prescale overflow (T-SPEED-TEMPORAL-PRESCALE-UP-OVERFLOW-2026-09-30, fix in fix(speed): size speed_temporal and speed_chroma frame buffers correctly #1643) and the dev image's -march=native FMA drift (T-DEV-IMAGE-ICX-NATIVE-FMA-DRIFT-2026-09-30).
  • fix(cuda): match the CPU motion order, option tables and tiny-frame guards #1637 also adds CUDA source-contract tests; this PR's test lives in its own file, test_cuda_device_resident_contract.py, so the two do not collide.

@github-actions github-actions Bot added the type:perf Performance improvement label Sep 30, 2026
lusoris added a commit that referenced this pull request Sep 30, 2026
…t code

The Release Script Contract failed on #1639: check-issue-reference-provenance
pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them
in integer_cambi_cuda.c (the dispatch helpers and the host download step the
device-resident rewrite removed) and one in docs/metrics/cambi.md. The four
code contracts go with the code they described. The CAMBI page keeps its
history as an "Implementation note (before ADR-1379)" block that still names
lusoris/vmaf#870, so that contract stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…t code

The Release Script Contract failed on #1639: check-issue-reference-provenance
pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them
in integer_cambi_cuda.c (the dispatch helpers and the host download step the
device-resident rewrite removed) and one in docs/metrics/cambi.md. The four
code contracts go with the code they described. The CAMBI page keeps its
history as an "Implementation note (before ADR-1379)" block that still names
lusoris/vmaf#870, so that contract stays.
@lusoris
lusoris force-pushed the perf/cuda-rc3-device-resident branch from 8616195 to d756336 Compare September 30, 2026 17:54
lusoris and others added 9 commits October 1, 2026 02:12
cambi_cuda, speed_chroma_cuda and speed_temporal_cuda no longer hand work
back to the host inside a frame. Each frame reads back one small result
block (88 bytes for cambi, 40 for SpEED) and waits once, in collect(); the
twins read the planes the CUDA engine already uploaded and upload nothing
of their own.

cambi_cuda used to download the distorted picture, preprocess it on the
host and read the image and mask back at every scale for host c-values and
top-K pooling. It now runs the ADR-1357 design on CUDA: twelve kernels for
preprocessing, the spatial mask, decimation, the mode filter, the
column-histogram c-values and an exact 128-bit top-K sum. Scores equal the
CPU's to the last bit whenever cambi.c's own double sum is exact. The twin
also gains cambi.c's guard against windows above 65 x 65.

The SpEED twins run the ADR-1358 chain, 25x25 eigenvalues and QR
included, in nine kernels shared through speed_cuda_pipeline.c. Every
rounding the CPU performs is spelled with a round-to-nearest intrinsic and
the fatbin builds with --fmad=false, so the scores equal a CPU build that
rounds log2f correctly and does not fuse multiply-adds.

cambi.c exports the host helpers both device twins need, and
speed_internal_gpu_configure() sets up the SYCL and CUDA SpEED pipelines;
the SYCL twins drop their private copies. CPU output is unchanged.

There is no NVIDIA device on this host. The kernels were built for sm_80
to sm_120 and checked frame by frame through a host emulation of the CUDA
driver; the two state rows stay open with the commands to verify and time
the port on ryzen-4090-arc. Three rows are opened for what the checks
found: the lanczos4 prescale drift of both device twins, the CPU
speed_temporal buffer overflow with speed_prescale above 1, and the dev
image's -march=native FMA drift of the CPU SpEED scores.

ADR-1379, ADR-1380, Research-1379.
…t code

The Release Script Contract failed on #1639: check-issue-reference-provenance
pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them
in integer_cambi_cuda.c (the dispatch helpers and the host download step the
device-resident rewrite removed) and one in docs/metrics/cambi.md. The four
code contracts go with the code they described. The CAMBI page keeps its
history as an "Implementation note (before ADR-1379)" block that still names
lusoris/vmaf#870, so that contract stays.
The fp32-pair log2 evaluation misrounds 48 positive finite floats whose
exact log2 falls closer than 2^-45 to a rounding boundary. Both device
twins now look the input up in a shared table (speed_log2_hard_cases.h)
and return the correctly rounded result. The common path costs one
fraction-field compare.

The CUDA cambi twin also drops a score < 0.0 clamp that the CPU does
not perform.

Contract tests verify the table entries against quad-precision log2 and
detect removal of the correction call in both twins.
…on.h NOLINT bracket

The bracket suppresses modernize-use-using for the header's typedef
structs, which clang-tidy parses as C++ under core/tools/vmaf.cpp
(through feature_dimensions.h and speed_internal.h) and under the SYCL
SpEED translation units; C cannot spell the `using` alias it proposes.
The CPU lane of the Tidy Ratchet had counted six findings there (0 -> 6).
The comment now names the C and C++ includers and cites ADR-1138, which
prescribes this file-scoped shape for C code parsed as C++. A CPU-lane
ratchet run scoped to vmaf.cpp measures 0 findings in the header. The
source ADR citation registry follows the new citation.
test_cuda_speed_temporal_parity_1080p builds the temporal parity test at
1920x1080, where the luma plane has 312 SpEED blocks. The host-split
twin on master launched its solve kernel with ((blocks + 7) / 8) * 32
threads per block, 1248 > 1024, so every frame at 1080p and above failed
with CUDA_ERROR_INVALID_VALUE (speed_temporal_cuda.c:425 at 10f27ef).
The 768x432 and 960x540 fixtures stay under 256 blocks and never reached
it. On the RTX 4090 the test fails against 10f27ef with that error and
passes with the ADR-1380 twin.
Measured on ryzen-4090-arc (RTX 4090, icx release build without
-march=native; before = origin/master 10f27ef built the same way):
- every per-frame cambi, speed_chroma_u/v/uv and speed_temporal equals
  --backend cpu at --precision max on the Netflix 576x324 pair (48/48)
  and on BBB 3840x2160 (50/50);
- compute-sanitizer memcheck, racecheck and synccheck report no error or
  hazard on the five CAMBI and SpEED parity tests;
- one readback and one stream synchronisation per frame (CUPTI count);
- 4K ms/frame before -> after: cambi_cuda 64.71 -> 6.01,
  speed_chroma_cuda 24.90 -> 6.89, speed_temporal_cuda fails -> 5.88;
- speed_log2() correctly rounded on every positive finite float, on the
  RTX 4090 and on the Arc A380.

T-CUDA-CAMBI-HOST-RESIDUAL-2026-09-29, T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29
and T-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30 are closed. The
lanczos4 row now carries device numbers (up to 2.1e-2 on a smooth 1080p
gradient), and the Arc A380 SpEED row records that its check is blocked
by the xe kernel driver, which returns wrong values from SYCL scratch
memory, rather than a vmafx bug. ADR-1379, ADR-1380, Research-1379, the
CAMBI and SpEED pages, the CUDA backend page, the changelog fragments,
the CUDA AGENTS.md and the rebase notes replace the emulation-only
statements with these measurements.
…es()

#1643 replaced speed_internal.c's SI_ALMOST_EQUAL with the shared speed_prescale_resamples() helper. The device geometry spelled the same rule out by hand with the removed macro; it now calls the helper, so the CPU and the device twins decide the resample from one function.
@lusoris
lusoris force-pushed the perf/cuda-rc3-device-resident branch from 005d417 to d4f37c0 Compare October 1, 2026 00:22
@lusoris
lusoris merged commit d056ef1 into master Oct 1, 2026
71 of 76 checks passed
@lusoris
lusoris deleted the perf/cuda-rc3-device-resident branch October 1, 2026 00:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant