Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 6 additions & 40 deletions .github/workflows/libvmaf-build-matrix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -107,46 +107,12 @@ jobs:
name: Build — macOS clang (CPU) + DNN
experimental: true

# --- Vulkan T5-1b runtime (ADR-0175) ---
# Real volk + VMA + glslc + compute dispatch. Needs the
# Vulkan loader + headers + glslc shader compiler before
# meson setup; lavapipe is the software ICD that exposes
# a compute-capable VkPhysicalDevice on headless runners
# so the smoke test can run end-to-end. Wraps for volk +
# VMA are sha256-pinned in core/subprojects/.
- os: ubuntu-latest
CC: ccache gcc
CXX: ccache g++
vulkan: true
meson_extra: -Denable_vulkan=enabled
name: Build — Ubuntu Vulkan (T5-1b runtime)

# --- macOS Vulkan via MoltenVK (ADR-0338, advisory) ---
# Validates the MoltenVK passthrough route on Apple
# Silicon: brew installs `molten-vk` (the Vulkan-on-Metal
# translation layer ICD) plus the Khronos `vulkan-loader`
# and `vulkan-headers`. `VK_ICD_FILENAMES` points the
# loader at MoltenVK's `MoltenVK_icd.json` so the loader
# enumerates exactly one ICD (Apple GPU via Metal). The
# smoke test pins the runtime contract; the cross-backend
# gate runs at `places=4` to confirm shader correctness.
#
# `experimental: true` + `continue-on-error` mark the lane
# advisory until one green run on master, per ADR-0338.
# Primary risk: `moment.comp` uses
# `GL_EXT_shader_atomic_int64` (atomicAdd on int64), which
# MoltenVK maps to Metal Tier-2 argument buffers — well-
# supported on M1+ but the most fragile dependency. Other
# shaders only use non-atomic int64 arithmetic
# (`GL_EXT_shader_explicit_arithmetic_types_int64`) which
# lowers to native Metal `long` and should pass.
- os: macos-latest
CC: ccache clang
CXX: ccache clang++
moltenvk: true
meson_extra: -Denable_vulkan=enabled
experimental: true
name: Build — macOS Vulkan via MoltenVK (advisory)
# --- Vulkan backend removed (ADR-0726) ---
# The Vulkan backend was removed in ADR-0726; the Ubuntu Vulkan
# (T5-1b runtime) and macOS MoltenVK (advisory) matrix rows that
# passed -Denable_vulkan=enabled have been dropped. meson.build
# no longer exports that option and configure would fail with
# "Unknown option: enable_vulkan".

# --- HIP T7-10b runtime (ADR-0212) ---
# Real HIP runtime build — compiles + links against
Expand Down
16 changes: 16 additions & 0 deletions changelog.d/fixed/build-matrix-macos-windows-fixes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
- Fix NEON `neon_any_nonzero_s32` uint64-truncation bug that incorrectly
skipped rows with alternating zero/non-zero y_row values on arm64,
causing the checkerboard motion score to return 0.0 instead of 12.55
on macOS arm64 CI runners.
- Fix Python harness `disable_avx` emitting `--cpumask -1` which the
CLI (ADR-1088) now rejects; updated to `4294967295` (0xFFFFFFFF).
- Fix Windows MSVC/CUDA/SYCL builds: add `pthread_dependency` to
`picture_pool_cpp23_lib` and `gpu_picture_pool_cpp23_lib` so the
win32 pthreads shim include path is resolved.
- Remove stale Vulkan matrix rows (`Build — Ubuntu Vulkan` and
`Build — macOS Vulkan via MoltenVK`) from the CI workflow; the Vulkan
backend was removed in ADR-0726 and `enable_vulkan` is no longer a
valid meson option.
- Update `test_run_preserves_user_env` to expect `LC_ALL=C` and
`LANG=C` that `ProcessRunner` unconditionally stamps for deterministic
subprocess error messages.
9 changes: 7 additions & 2 deletions compat/python-vmaf/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -230,7 +230,10 @@ def call_vmafexec_multi_features(
if options is not None and "disable_avx" in options:
assert isinstance(options["disable_avx"], bool)
if options["disable_avx"] is True:
cmd += ["--cpumask", "-1"]
# 0xFFFFFFFF disables all CPU ISA extensions (all mask bits set).
# parse_unsigned() now rejects negative strings such as "-1"
# (ADR-1088); pass the unsigned equivalent instead.
cmd += ["--cpumask", "4294967295"]

if options is not None and "n_threads" in options:
assert isinstance(options["n_threads"], int) and options["n_threads"] >= 1
Expand Down Expand Up @@ -373,7 +376,9 @@ def call_vmafexec(
vmafexec_cmd += " --threads {}".format(n_threads)

if disable_avx:
vmafexec_cmd += " --cpumask -1"
# 0xFFFFFFFF disables all CPU ISA extensions (all mask bits set).
# parse_unsigned() rejects negative strings (ADR-1088).
vmafexec_cmd += " --cpumask 4294967295"

if logger:
logger.info(vmafexec_cmd)
Expand Down
23 changes: 17 additions & 6 deletions core/src/feature/arm64/motion_v2_neon.c
Original file line number Diff line number Diff line change
Expand Up @@ -58,15 +58,26 @@ static inline uint32_t neon_hadd_u32(uint32x4_t v)
return vaddvq_u32(v);
}

/* Non-zero test across 4×int32 lanes: returns non-zero iff any bit is set
* in any lane. Uses bitwise OR-fold rather than arithmetic sum so that
* positive and negative signed values cannot cancel each other to give a
* false zero (e.g. checkerboard frames trigger this on the signed-sum path).
/* Non-zero test across 4×int32 lanes: returns non-zero iff any lane is
* non-zero. Folds all four uint32 values with bitwise OR so that positive
* and negative signed values cannot cancel each other to give a false zero.
*
* The previous implementation reinterpreted int32x4 as uint64x2 and then
* cast the OR of the two uint64 lanes to uint32. That loses the upper 32
* bits of each uint64 lane: if lane-0 = (int32[0] | int32[1]<<32) and
* int32[0]==0 while int32[1]!=0, the uint64 is non-zero but the uint32
* cast truncates it to 0. On checkerboard frames — where every other
* y_row entry is zero — this false-zero caused the x_conv phase to be
* skipped for every row, yielding a motion SAD of 0.0. Fix: OR all four
* uint32 lanes directly at uint32 width; no truncation possible.
*/
static inline uint32_t neon_any_nonzero_s32(int32x4_t v)
{
const uint64x2_t u64 = vreinterpretq_u64_s32(v);
return (uint32_t)(vgetq_lane_u64(u64, 0) | vgetq_lane_u64(u64, 1));
const uint32x4_t u32 = vreinterpretq_u32_s32(v);
const uint32x2_t lo = vget_low_u32(u32);
const uint32x2_t hi = vget_high_u32(u32);
const uint32x2_t folded = vorr_u32(lo, hi);
return vget_lane_u32(folded, 0) | vget_lane_u32(folded, 1);
}

/* Compute abs(x_conv) for 4 int32 lanes starting at y_row[j] using
Expand Down
6 changes: 6 additions & 0 deletions core/src/meson.build
Original file line number Diff line number Diff line change
Expand Up @@ -1765,23 +1765,29 @@ opt_cpp23_lib = static_library(

# ADR-0768: picture_pool.cpp compiled as an isolated C++23 static lib
# (same isolation pattern as metadata_handler_cpp20_lib / ADR-0708).
# pthread_dependency: on Windows MSVC (no pthread.h in system headers) this
# adds src/compat/win32/ to the include path so the win32 pthreads shim is
# found. On POSIX it is an empty list and has no effect.
picture_pool_cpp23_lib = static_library(
'picture_pool_cpp23',
src_dir + 'picture_pool.cpp',
include_directories : [vmaf_base_include, libvmaf_include],
# libvmaf_cpu_cpp_std: compiler-aware token (ADR-0860 follow-up, PR #692 fix).
override_options : ['cpp_std=' + libvmaf_cpu_cpp_std],
dependencies : [pthread_dependency],
pic : true,
install : false,
)

# ADR-0768: gpu_picture_pool.cpp compiled as an isolated C++23 static lib.
# pthread_dependency: same win32 shim rationale as picture_pool_cpp23_lib above.
gpu_picture_pool_cpp23_lib = static_library(
'gpu_picture_pool_cpp23',
src_dir + 'gpu_picture_pool.cpp',
include_directories : [vmaf_base_include, libvmaf_include],
# libvmaf_cpu_cpp_std: compiler-aware token (ADR-0860 follow-up, PR #692 fix).
override_options : ['cpp_std=' + libvmaf_cpu_cpp_std],
dependencies : [pthread_dependency],
pic : true,
install : false,
)
Expand Down
21 changes: 21 additions & 0 deletions docs/rebase-notes.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,27 @@
<!-- markdownlint-disable MD001 MD003 MD004 MD007 MD013 MD018 MD022 MD024 MD025 MD026 MD028 MD029 MD031 MD032 MD033 MD036 MD037 MD038 MD040 MD041 MD046 MD049 MD050 MD051 MD052 MD053 MD055 MD056 MD058 MD059 -->
# Rebase notes

## fix/build-matrix-macos-windows-fixes (2026-06-07)

**Rebase-sensitive (meson.build):** `core/src/meson.build` gains
`dependencies : [pthread_dependency]` on both `picture_pool_cpp23_lib` and
`gpu_picture_pool_cpp23_lib` static library targets (~lines 1768–1788). If a
concurrent branch adds other fields to those `static_library()` calls, merge
both sets of fields.

Other changes are not rebase-sensitive:
- `core/src/feature/arm64/motion_v2_neon.c`: rewrite of `neon_any_nonzero_s32`
(isolated function, no surrounding context).
- `compat/python-vmaf/__init__.py`: two call-sites of `--cpumask` changed from
`"-1"` to `"4294967295"`.
- `.github/workflows/libvmaf-build-matrix.yml`: two Vulkan matrix rows removed;
if a concurrent branch also removes the Vulkan step bodies (`Install Vulkan SDK`,
`Cache meson subprojects (Vulkan wraps)`, `Run Vulkan smoke tests (macOS
MoltenVK)`, etc.), take both removals.
- `python/test/python_harness_coverage_test.py`: test expectation update
(`--cpumask -1` → `--cpumask 4294967295`; `test_run_preserves_user_env`
expected dict gains `LC_ALL`/`LANG`).

## fix/nightly-bisect-tracker-issue (2026-06-07)
no rebase impact: changes confined to `.github/workflows/nightly-bisect.yml`,
`scripts/ci/post-bisect-comment.py`, `docs/state.md`, and
Expand Down
Loading
Loading