Repository navigation
fix(ci): skip the CUDA parity-gate default run before the device lock on a build without CUDA - #1746
Merged
Merged
Conversation
…1142) (#1742) * refactor(hip): bring hip tests part 2 to lint and HISS standard (ADR-1142) Refactor the second batch of HIP test sources to eliminate clang-tidy debt and satisfy HISS complexity limits: - core/test/test_hip_wavefront_reduce.c - core/test/test_hip_psnr_hvs_parity.c - core/test/test_hip_speed_chroma_parity.c - core/test/test_hip_speed_temporal_parity.c - core/test/test_hip_speed_singular_parity.c - core/test/test_hip_ciede_parity.c - core/test/test_hip_motion_parity.c - core/test/test_hip_vif_parity.c - core/test/test_hip_psnr_parity.c - core/test/test_hip_adm_parity.c Functions exceeding statement and branch thresholds are decomposed into focused helpers to conform with HISS-04 / NASA JPL Rule 4. Single-statement loops and branches receive explicit braces. Portable C spelling of NULL is preserved via ADR-1138 NOLINT brackets for MSVC compatibility. Ratchet baselines in scripts/ci/tidy-baseline-*.json are tightened (-156 warnings in the HIP lane down to 0 for all touched files, and -4 warnings across arm64, cpu, cuda, and sycl for wavefront reduce).
…gate (#1743) * ci(lint): wire actionlint pre-commit hook and document workflow lint gate
…take its debug default (#1744) * fix(sycl): round integer vif's per-scale sums where the CPU does and take its debug default integer_vif.c stores each scale's numerator and denominator sum in a float (vif_store_residuals()), adds those rounded values for the debug outputs and divides in single precision. vif_sycl kept the sums in double and divided in double, so no score of any frame was the CPU's: up to 3.5e-7 on the Netflix pair. vif_cuda and vif_hip already round as the CPU does. - integer_vif_sycl.cpp: vif_scale_sums() rounds each scale's sums to float; vif_score_set() mirrors integer_vif.c::write_scores() and asks the emitter for the single-precision ratio. - The twin's debug option defaults to false, as on the CPU. It defaulted to true and added eleven debug outputs to every run. Measured on an Arc A380 at --precision max against --backend cpu, frames with identical scores on scales 0 / 1 / 2 / 3, before and after: Netflix 576x324 0 of 48, then 41 / 31 / 16 / 12; checkerboard 1 px 0 of 3, then 3 / 3 / 3 / 2; checkerboard 10 px 1 / 3 / 3 / 3, then 3 of 3; BBB 3840x2160 0 of 20, then 196 / 174 / 184 / 140 of 200. The denominator sums are identical on every frame. The largest difference is unchanged (3.6e-7): on the remaining frames a numerator sum is one or a few fp32 steps off, because the kernel forms the per-pixel gain in fp32 where the CPU uses fp64 (T-SYCL-VIF-FP32-GAIN-2026-10-01). The kernels are not touched. Tests: test_sycl_vif_parity asserts equal denominators, single-precision outputs and the CPU's default output set at 8 and 10 bits, and fails on the old twin; test_sycl_vif_float_sums_contract.py plants four regressions.
… on a build without CUDA (#1746) * fix(ci): skip the CUDA parity-gate default run before the device lock on a build without CUDA test_cuda_parity_gate_default_run runs the gate under the CUDA device lock and learned that a binary has no CUDA only from the gate's output, that is after it had the lock. On a build configured without CUDA it therefore waited for a device it cannot use: with another job holding the lock, --suite=fast --suite=gpu in a HIP-only build dir reported one TIMEOUT (120 s, 0.17 s of CPU time) next to 293 passing tests. build_has_cuda() reads enable_cuda from the build directory's meson-info/intro-buildoptions.json and skip_reason() returns the skip ahead of the command that takes the lock. A build directory without that record goes on as before and the CLI's refusal decides. The same build now exits 77 in 0.02 s. test_cuda_parity_gate_skip.py gets four cases: a build without CUDA, one with, a record that does not say, and the skip ahead of the lock. They fail on the parent commit (4 errors). docs/state.md: T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01 found and closed.
lusoris
force-pushed
the
fix/cuda-parity-gate-skip-before-lock
branch
from
October 1, 2026 18:52
03a9f85 to
f70591c
Compare
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes
T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01(found running--suite=fast --suite=gpuin a HIP-only build dir: 293 OK, 1 timeout).test_cuda_parity_gate_default_runruns the gate under the CUDA device lock and decided that a binary has no CUDA from the gate's output (#1698), which is after it has the lock. On a build configured without CUDA it waited for a device it cannot use, and timed out after 120 s whenever another job held the lock. Run alone in that state it used 0.17 s of CPU time in 120 s.The test now reads
enable_cudafrom the build directory'smeson-info/intro-buildoptions.jsonand skips (exit 77) before the command that takes the lock. A build directory without that record goes on as before, and the CLI's refusal still decides there. On the same HIP-only build the test exits 77 in 0.02 s.core/test/test_cuda_parity_gate_skip.pygets four cases: a build without CUDA, one with, a record that does not say (missing, malformed, no such option, a non-boolean value), and the skip ahead of the lock. They fail on the parent commit (4 errors).Not changed: the lock path and the
flockcall are specific to this host and stay as #1676 wrote them.Type
fix— bug fixDeep-dive deliverables (ADR-0108)
AGENTS.mdinvariant note — no rebase-sensitive invariants.changelog.d/fixed/ci-parity-gate-default-run-skip-before-lock.md.Reproducer
Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR:T-CI-PARITY-GATE-DEFAULT-RUN-WAITS-FOR-CUDA-LOCK-2026-10-01added to Recently closed.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.