Skip to content

fix(ci): skip the CUDA parity-gate default run before the device lock on a build without CUDA - #1746

Merged
lusoris merged 4 commits into
masterfrom
fix/cuda-parity-gate-skip-before-lock
Oct 1, 2026
Merged

lusoris merged 4 commits into
masterfrom
fix/cuda-parity-gate-skip-before-lock

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01 (found running --suite=fast --suite=gpu in a HIP-only build dir: 293 OK, 1 timeout).

test_cuda_parity_gate_default_run runs the gate under the CUDA device lock and decided that a binary has no CUDA from the gate's output (#1698), which is after it has the lock. On a build configured without CUDA it waited for a device it cannot use, and timed out after 120 s whenever another job held the lock. Run alone in that state it used 0.17 s of CPU time in 120 s.

The test now reads enable_cuda from the build directory's meson-info/intro-buildoptions.json and skips (exit 77) before the command that takes the lock. A build directory without that record goes on as before, and the CLI's refusal still decides there. On the same HIP-only build the test exits 77 in 0.02 s.

core/test/test_cuda_parity_gate_skip.py gets four cases: a build without CUDA, one with, a record that does not say (missing, malformed, no such option, a non-boolean value), and the skip ahead of the lock. They fail on the parent commit (4 errors).

Not changed: the lock path and the flock call are specific to this host and stay as #1676 wrote them.

Type

  • fix — bug fix

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: trivial (a skip condition in one test).
  • Decision matrix — no alternatives: only-one-way fix.
  • AGENTS.md invariant note — no rebase-sensitive invariants.
  • Reproducer / smoke-test command — under "Reproducer" below.
  • CHANGELOG fragment — changelog.d/fixed/ci-parity-gate-default-run-skip-before-lock.md.
  • Rebase note — no rebase impact: fork-only test files.

Reproducer

python3 core/test/test_cuda_parity_gate_skip.py
# with a build dir configured without CUDA, whatever holds the CUDA lock:
VMAF_BUILD_DIR=build-hip python3 core/test/test_cuda_parity_gate_default_run.py; echo $?   # 77

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01 added to Recently closed.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.

…1142) (#1742)

* refactor(hip): bring hip tests part 2 to lint and HISS standard (ADR-1142)

Refactor the second batch of HIP test sources to eliminate clang-tidy debt
and satisfy HISS complexity limits:
- core/test/test_hip_wavefront_reduce.c
- core/test/test_hip_psnr_hvs_parity.c
- core/test/test_hip_speed_chroma_parity.c
- core/test/test_hip_speed_temporal_parity.c
- core/test/test_hip_speed_singular_parity.c
- core/test/test_hip_ciede_parity.c
- core/test/test_hip_motion_parity.c
- core/test/test_hip_vif_parity.c
- core/test/test_hip_psnr_parity.c
- core/test/test_hip_adm_parity.c

Functions exceeding statement and branch thresholds are decomposed into
focused helpers to conform with HISS-04 / NASA JPL Rule 4. Single-statement
loops and branches receive explicit braces. Portable C spelling of NULL
is preserved via ADR-1138 NOLINT brackets for MSVC compatibility.

Ratchet baselines in scripts/ci/tidy-baseline-*.json are tightened
(-156 warnings in the HIP lane down to 0 for all touched files, and -4
warnings across arm64, cpu, cuda, and sycl for wavefront reduce).
…gate (#1743)

* ci(lint): wire actionlint pre-commit hook and document workflow lint gate
…take its debug default (#1744)

* fix(sycl): round integer vif's per-scale sums where the CPU does and take its debug default

integer_vif.c stores each scale's numerator and denominator sum in a
float (vif_store_residuals()), adds those rounded values for the debug
outputs and divides in single precision. vif_sycl kept the sums in
double and divided in double, so no score of any frame was the CPU's:
up to 3.5e-7 on the Netflix pair. vif_cuda and vif_hip already round as
the CPU does.

- integer_vif_sycl.cpp: vif_scale_sums() rounds each scale's sums to
  float; vif_score_set() mirrors integer_vif.c::write_scores() and asks
  the emitter for the single-precision ratio.
- The twin's debug option defaults to false, as on the CPU. It defaulted
  to true and added eleven debug outputs to every run.

Measured on an Arc A380 at --precision max against --backend cpu,
frames with identical scores on scales 0 / 1 / 2 / 3, before and after:
Netflix 576x324 0 of 48, then 41 / 31 / 16 / 12; checkerboard 1 px 0 of
3, then 3 / 3 / 3 / 2; checkerboard 10 px 1 / 3 / 3 / 3, then 3 of 3;
BBB 3840x2160 0 of 20, then 196 / 174 / 184 / 140 of 200. The
denominator sums are identical on every frame.

The largest difference is unchanged (3.6e-7): on the remaining frames a
numerator sum is one or a few fp32 steps off, because the kernel forms
the per-pixel gain in fp32 where the CPU uses fp64
(T-SYCL-VIF-FP32-GAIN-2026-10-01). The kernels are not touched.

Tests: test_sycl_vif_parity asserts equal denominators, single-precision
outputs and the CPU's default output set at 8 and 10 bits, and fails on
the old twin; test_sycl_vif_float_sums_contract.py plants four
regressions.
… on a build without CUDA (#1746)

* fix(ci): skip the CUDA parity-gate default run before the device lock on a build without CUDA

test_cuda_parity_gate_default_run runs the gate under the CUDA device
lock and learned that a binary has no CUDA only from the gate's output,
that is after it had the lock. On a build configured without CUDA it
therefore waited for a device it cannot use: with another job holding
the lock, --suite=fast --suite=gpu in a HIP-only build dir reported one
TIMEOUT (120 s, 0.17 s of CPU time) next to 293 passing tests.

build_has_cuda() reads enable_cuda from the build directory's
meson-info/intro-buildoptions.json and skip_reason() returns the skip
ahead of the command that takes the lock. A build directory without
that record goes on as before and the CLI's refusal decides. The same
build now exits 77 in 0.02 s.

test_cuda_parity_gate_skip.py gets four cases: a build without CUDA, one
with, a record that does not say, and the skip ahead of the lock. They
fail on the parent commit (4 errors).

docs/state.md: T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01
found and closed.
@lusoris
lusoris force-pushed the fix/cuda-parity-gate-skip-before-lock branch from 03a9f85 to f70591c Compare October 1, 2026 18:52
@lusoris
lusoris merged commit f70591c into master Oct 1, 2026
60 of 69 checks passed
@lusoris
lusoris deleted the fix/cuda-parity-gate-skip-before-lock branch October 1, 2026 18:52
@github-actions github-actions Bot added the type:bug Something isn't working label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant