Repository navigation
perf(cuda): run cambi and SpEED entirely on the device - #1639
Merged
Merged
Conversation
lusoris
added a commit
that referenced
this pull request
Sep 30, 2026
…t code The Release Script Contract failed on #1639: check-issue-reference-provenance pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them in integer_cambi_cuda.c (the dispatch helpers and the host download step the device-resident rewrite removed) and one in docs/metrics/cambi.md. The four code contracts go with the code they described. The CAMBI page keeps its history as an "Implementation note (before ADR-1379)" block that still names lusoris/vmaf#870, so that contract stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 30, 2026
lusoris
added a commit
that referenced
this pull request
Sep 30, 2026
…t code The Release Script Contract failed on #1639: check-issue-reference-provenance pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them in integer_cambi_cuda.c (the dispatch helpers and the host download step the device-resident rewrite removed) and one in docs/metrics/cambi.md. The four code contracts go with the code they described. The CAMBI page keeps its history as an "Implementation note (before ADR-1379)" block that still names lusoris/vmaf#870, so that contract stays.
lusoris
force-pushed
the
perf/cuda-rc3-device-resident
branch
from
September 30, 2026 17:54
8616195 to
d756336
Compare
17 of 26 tasks
cambi_cuda, speed_chroma_cuda and speed_temporal_cuda no longer hand work back to the host inside a frame. Each frame reads back one small result block (88 bytes for cambi, 40 for SpEED) and waits once, in collect(); the twins read the planes the CUDA engine already uploaded and upload nothing of their own. cambi_cuda used to download the distorted picture, preprocess it on the host and read the image and mask back at every scale for host c-values and top-K pooling. It now runs the ADR-1357 design on CUDA: twelve kernels for preprocessing, the spatial mask, decimation, the mode filter, the column-histogram c-values and an exact 128-bit top-K sum. Scores equal the CPU's to the last bit whenever cambi.c's own double sum is exact. The twin also gains cambi.c's guard against windows above 65 x 65. The SpEED twins run the ADR-1358 chain, 25x25 eigenvalues and QR included, in nine kernels shared through speed_cuda_pipeline.c. Every rounding the CPU performs is spelled with a round-to-nearest intrinsic and the fatbin builds with --fmad=false, so the scores equal a CPU build that rounds log2f correctly and does not fuse multiply-adds. cambi.c exports the host helpers both device twins need, and speed_internal_gpu_configure() sets up the SYCL and CUDA SpEED pipelines; the SYCL twins drop their private copies. CPU output is unchanged. There is no NVIDIA device on this host. The kernels were built for sm_80 to sm_120 and checked frame by frame through a host emulation of the CUDA driver; the two state rows stay open with the commands to verify and time the port on ryzen-4090-arc. Three rows are opened for what the checks found: the lanczos4 prescale drift of both device twins, the CPU speed_temporal buffer overflow with speed_prescale above 1, and the dev image's -march=native FMA drift of the CPU SpEED scores. ADR-1379, ADR-1380, Research-1379.
…t code The Release Script Contract failed on #1639: check-issue-reference-provenance pins five historical blocks that cite lusoris/vmaf#857 and #870, four of them in integer_cambi_cuda.c (the dispatch helpers and the host download step the device-resident rewrite removed) and one in docs/metrics/cambi.md. The four code contracts go with the code they described. The CAMBI page keeps its history as an "Implementation note (before ADR-1379)" block that still names lusoris/vmaf#870, so that contract stays.
The fp32-pair log2 evaluation misrounds 48 positive finite floats whose exact log2 falls closer than 2^-45 to a rounding boundary. Both device twins now look the input up in a shared table (speed_log2_hard_cases.h) and return the correctly rounded result. The common path costs one fraction-field compare. The CUDA cambi twin also drops a score < 0.0 clamp that the CPU does not perform. Contract tests verify the table entries against quad-precision log2 and detect removal of the correction call in both twins.
…e hardware evidence
…on.h NOLINT bracket The bracket suppresses modernize-use-using for the header's typedef structs, which clang-tidy parses as C++ under core/tools/vmaf.cpp (through feature_dimensions.h and speed_internal.h) and under the SYCL SpEED translation units; C cannot spell the `using` alias it proposes. The CPU lane of the Tidy Ratchet had counted six findings there (0 -> 6). The comment now names the C and C++ includers and cites ADR-1138, which prescribes this file-scoped shape for C code parsed as C++. A CPU-lane ratchet run scoped to vmaf.cpp measures 0 findings in the header. The source ADR citation registry follows the new citation.
test_cuda_speed_temporal_parity_1080p builds the temporal parity test at 1920x1080, where the luma plane has 312 SpEED blocks. The host-split twin on master launched its solve kernel with ((blocks + 7) / 8) * 32 threads per block, 1248 > 1024, so every frame at 1080p and above failed with CUDA_ERROR_INVALID_VALUE (speed_temporal_cuda.c:425 at 10f27ef). The 768x432 and 960x540 fixtures stay under 256 blocks and never reached it. On the RTX 4090 the test fails against 10f27ef with that error and passes with the ADR-1380 twin.
Measured on ryzen-4090-arc (RTX 4090, icx release build without -march=native; before = origin/master 10f27ef built the same way): - every per-frame cambi, speed_chroma_u/v/uv and speed_temporal equals --backend cpu at --precision max on the Netflix 576x324 pair (48/48) and on BBB 3840x2160 (50/50); - compute-sanitizer memcheck, racecheck and synccheck report no error or hazard on the five CAMBI and SpEED parity tests; - one readback and one stream synchronisation per frame (CUPTI count); - 4K ms/frame before -> after: cambi_cuda 64.71 -> 6.01, speed_chroma_cuda 24.90 -> 6.89, speed_temporal_cuda fails -> 5.88; - speed_log2() correctly rounded on every positive finite float, on the RTX 4090 and on the Arc A380. T-CUDA-CAMBI-HOST-RESIDUAL-2026-09-29, T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29 and T-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30 are closed. The lanczos4 row now carries device numbers (up to 2.1e-2 on a smooth 1080p gradient), and the Arc A380 SpEED row records that its check is blocked by the xe kernel driver, which returns wrong values from SYCL scratch memory, rather than a vmafx bug. ADR-1379, ADR-1380, Research-1379, the CAMBI and SpEED pages, the CUDA backend page, the changelog fragments, the CUDA AGENTS.md and the rebase notes replace the emulation-only statements with these measurements.
…es() #1643 replaced speed_internal.c's SI_ALMOST_EQUAL with the shared speed_prescale_resamples() helper. The device geometry spelled the same rule out by hand with the removed macro; it now calls the helper, so the CPU and the device twins decide the resample from one function.
lusoris
force-pushed
the
perf/cuda-rc3-device-resident
branch
from
October 1, 2026 00:22
005d417 to
d4f37c0
Compare
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
cambi_cuda,speed_chroma_cudaandspeed_temporal_cudanow run every per-frame stage on the GPU: each frame reads back one small result block and waits once, incollect(), and nothing is uploaded beyond the planes the CUDA engine already has on the device. On the RTX 4090 ofryzen-4090-arc, every frame of the Netflix 576x324 pair and of BBB 3840x2160 equals--backend cpu(icx build) to the last bit, and a 4K frame of any of the three now costs about what the trivialpsnr_cudacosts. The rewrite also fixes two bugs onmaster:speed_temporal_cudafailed at 1920x1080 and above, andcambi_cudacrashed on wide, short frames.cambi_cudais device-resident (ADR-1379). The old twin downloaded the distorted picture, preprocessed it on the host, uploaded it again, and at each of five scales waited and read the image and mask back for host c-values and top-K pooling. Now twelve kernels (the ADR-1357 design) do all of it, and 88 bytes come back per frame. The twin also gainscambi.c's guard against windows above 65 x 65.speed/speed_score.cu, shared by both extractors throughspeed_cuda_pipeline.c, and 40 bytes back per frame. Every rounding the CPU performs is spelled with__fadd_rn/__fmul_rn/__fdiv_rn/__fsqrt_rn, and the fatbin builds with--fmad=false.log2on the device. The fp32-pairspeed_log2()misrounds exactly 48 positive floats. Both device twins (CUDA and SYCL) now look those up inspeed_log2_hard_cases.h. An exhaustive run over all 2 139 095 039 positive finite floats finds 0 misrounds on the RTX 4090 and 0 on the Arc A380.cambi.cexports the helpers both device twins need, andspeed_internal_gpu_configure()sets up the SYCL and CUDA SpEED pipelines alike; the SYCL twins drop their private copies. CPU output does not change.master's CUDA twins.speed_temporal_cudalaunched its solve kernel with((blocks + 7) / 8) * 32threads per block, over the 1024 limit above 256 SpEED blocks, so at 1920x1080 and 3840x2160 every frame failed withCUDA_ERROR_INVALID_VALUE(T-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30). The new testtest_cuda_speed_temporal_parity_1080pfails on10f27efe2with that error and passes here.cambi_cudarancambi.c's host c-values walk, so on the wide, short frames of cambi: out-of-bounds read and write on wide, short frames (e.g. 1920x128), C and AVX2 paths Netflix/vmaf#1628 it scored frame 0 wrongly and then segfaulted inclose_fex_cuda(). The device twin runs those sizes clean undercompute-sanitizerand equals the fixed CPU of fix(cambi): keep the c-values walks inside short frames #1642.speed_gpu_common.hreachescore/tools/vmaf.cppthroughfeature_dimensions.h, so clang-tidy parsed itstypedef structs as C++ (modernize-use-using, +6 on the CPU lane). The header now carries the cited, file-scopedNOLINTBEGIN(modernize-use-using)bracket that ADR-1138 prescribes for C headers parsed as C++.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally, run piecewise: clang-format on every touched C, C++, CUDA and header file; the full pre-commit set overorigin/master..HEADand the pre-push stage; the CPU tidy lane scoped tocore/tools/vmaf.cpp(the TU that parsesspeed_gpu_common.has C++) at 0 forspeed_gpu_common.h; the CUDA lane baselinescripts/ci/tidy-baseline-cuda.jsontightened for the rewritten TUs by the ratchet's scoped write.python3 scripts/ci/run_meson_test.py -- -C build-cuda-icx --no-rebuild test_cuda_cambi_parity test_cuda_cambi_parity_large test_cuda_device_resident_contract test_cuda_speed_chroma_parity test_cuda_speed_temporal_parity test_cuda_speed_temporal_parity_1080p test_cuda_speed_singular_parity test_cuda_speed_chroma_smoke test_cuda_speed_temporal_smoke, all OK on the RTX 4090, none skipped./cross-backend-diffand the worst ULP is ≤ 2. Per frame at--precision maxagainst an icx build: 0 ULP on everycambi,speed_chroma_u/v/uvandspeed_temporaloutput on both fixtures (tables below)..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. Not breaking: no public API, CLI or option change.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Hooks: earlier pushes of this branch skipped hooks. This push ran the full lefthook and pre-commit set, pre-push stage included, with the repo-pinned
praetorctl(f41e74d8, the engine CI's standards gate installs). The newerpraetorctlbuild installed on the host during this run (ff7ea2c4) failedpraetorctl auditonorigin/masteritself (428 infractions against the 185 baseline), so it could not gate a branch.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR. Closed:T-CUDA-CAMBI-HOST-RESIDUAL-2026-09-29,T-CUDA-SPEED-HOST-RESIDUAL-2026-09-29andT-CUDA-SPEED-TEMPORAL-SOLVE-LAUNCH-1080P-2026-09-30, with the measurements below. Opened:T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30(RC3),T-SPEED-TEMPORAL-PRESCALE-UP-OVERFLOW-2026-09-30(RC2),T-DEV-IMAGE-ICX-NATIVE-FMA-DRIFT-2026-09-30(RC3) andT-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30(RC3, verification blocked, see below), each in its disposition row.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.cambi.candspeed_internal.cchanges only move code into shared functions; the CPU CAMBI and SpEED scores are unchanged.Cross-backend numerical results
Measured on
ryzen-4090-arc: RTX 4090 (sm_89, driver 615.71.09, CUDA 13.4), icx 2026.0 release build with-Db_lto=falseand without-march=native(the state rows' reference build). "Before" isorigin/master10f27efe2built the same way.python3 scripts/dev/speed_gpu_parity.py --backend cuda --vmaf $PWD/build-cuda-icx/tools/vmaf --netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb(and--feature cambi), identical frames per output at--precision max:cambispeed_chroma_uspeed_chroma_vspeed_chroma_uvspeed_temporalBefore, the largest difference was 4.0e-5. A gcc 16.2.1 build of this branch against its own CPU extractor (glibc 2.44
log2f) differs on 0 to 2speed_chromaframes per output, by at most 1.4e-6, and matches onspeed_temporal; the device roundslog2correctly and glibc does not on those inputs.speed_prescale=0.5, BBB 3840x2160, 6 frames): nearest, bilinear and bicubic identical on every frame.lanczos4is not: at most 1.9e-5 apart on BBB, but 2.1e-2 (5.7e-4 relative,speed_chroma_v) on an ffmpeggradientsclip with noise at 1920x1080, beyond the ADR-0214 tolerance (T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30, being fixed in this PR, see Known follow-ups).cambi_cudaequals the CPU build of fix(cambi): keep the c-values walks inside short frames #1642 on every frame (6.0056, 7.7220, 8.0805, 6.0946),compute-sanitizer --tool memcheck0 errors.master's twin scored frame 0 differently, the later frames 0, and then segfaulted inclose_fex_cuda()→vmaf_picture_unref()(gdb backtrace).compute-sanitizermemcheck, racecheck and synccheck ontest_cuda_cambi_parity,test_cuda_cambi_parity_large,test_cuda_speed_singular_parity,test_cuda_speed_temporal_parityandtest_cuda_speed_chroma_parity: 0 errors and 0 hazards on all 15 runs.cambi_syclequals--backend cpuon 48/48 (576x324) and 50/50 (BBB 4K) frames. The SYCL SpEED re-check is blocked: since its 18:04 boot this host drives the A380 with thexekernel driver, under which, as measured by the fix(sycl): give every SYCL feature kernel the CPU's fp32 arithmetic #1630 track, SYCL kernels that use scratch memory return wrong values (standalone probes without vmafx code too, andorigin/masterfails the same 16 of 55 SYCL tests). The SpEED twins emit 0 there onmasterand on this branch alike (T-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30).Performance (if
perforfeat)Per frame, counted with a CUPTI driver-API callback over frames 13 to 22 of the Netflix 576x324 pair. The engine's own calls (six picture uploads, one context and two event synchronisations) are the same in every row and left out:
cambi_cudabeforecambi_cudaafterspeed_chroma_cudabeforespeed_chroma_cudaafterspeed_temporal_cudabeforespeed_temporal_cudaafterMilliseconds per frame,
(t(N) - t(2)) / (N - 2), median of 3, before and after runs alternating so both see the same host load (other sessions kept the load average between 13 and 41). The rows' N = 22 was below the run-to-run noise at 576x324 on this host, so the table uses N = 102 at 3840x2160 and N = 402 at 576x324 (the Netflix pair looped ten times):cambi_cudaspeed_chroma_cudaspeed_temporal_cudacambi_cudaspeed_chroma_cudaspeed_temporal_cudaWith the rows' own N = 22 at 3840x2160:
cambi_cuda68.39 → 3.67,speed_chroma_cuda14.29 → 5.81,speed_temporal_cudafails → 5.06. The absolute numbers move with the host load; the relation to a trivial twin does not. In a run at load average 13 with N = 102,psnr_cudatook 2.56 and 2.41 ms per 4K frame,cambi_cuda2.55,speed_chroma_cuda2.75 andspeed_temporal_cuda2.73: at 4K the three twins now run at the cost of reading and uploading the pictures.Deep-dive deliverables (ADR-0108)
log2fand FMA contraction, the exact top-K sum, thelanczos4andspeed_temporalprescale findings, the exhaustivelog2replay (finding 7) and the RTX 4090 measurements with their commands (finding 8).## Alternatives consideredtables of ADR-1379 and ADR-1380.AGENTS.mdinvariant note —core/src/feature/cuda/AGENTS.md"Device-resident CAMBI and SpEED (ADR-1379, ADR-1380)", including thespeed_log2()hard-case table; index entry indocs/development/rebase-sensitive-invariants.md.changelog.d/changed/perf-cuda-cambi-device-resident.mdandchangelog.d/changed/perf-cuda-speed-device-resident.md.docs/rebase-notes.md"ADR-1379 / ADR-1380 — CUDA CAMBI and SpEED run entirely on the device".Reproducer
Known follow-ups
lanczos4prescale (T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30): the device twins compute the kernel weights in fp32 withsinpif(), the CPU in fp64 withsin(). The weights depend only on the output column and row, so computing them once on the host withvif_tools.c's own function makes the CUDA twin exact; that change is being added to this PR. The SYCL twin keeps its fp32 weights until it can be verified on an Intel GPU with a working driver.T-SYCL-SPEED-A380-SINGULAR-COVARIANCE-2026-09-30): runspeed_gpu_parity.py --backend syclon an Arc with a working driver (the A380 back oni915, or the office B580 / UHD 770).T-HIP-CAMBI-HOST-RESIDUAL-2026-09-29,T-METAL-CAMBI-HOST-RESIDUAL-2026-09-29,T-HIP-SPEED-HOST-RESIDUAL-2026-09-29).speed_temporalprescale overflow (T-SPEED-TEMPORAL-PRESCALE-UP-OVERFLOW-2026-09-30, fix in fix(speed): size speed_temporal and speed_chroma frame buffers correctly #1643) and the dev image's-march=nativeFMA drift (T-DEV-IMAGE-ICX-NATIVE-FMA-DRIFT-2026-09-30).test_cuda_device_resident_contract.py, so the two do not collide.