Skip to content

cuda: enable CUDA feature extraction on Windows (MSYS2/MinGW) - #1472

Closed
birkdev wants to merge 4 commits into
Netflix:masterfrom
birkdev:windows-cuda-support
Closed

birkdev wants to merge 4 commits into
Netflix:masterfrom
birkdev:windows-cuda-support

Conversation

@birkdev

@birkdev birkdev commented Mar 16, 2026 •

Copy link
Copy Markdown

Summary

Enables building libvmaf with -Denable_cuda=true on Windows using MSYS2/MinGW, closing the gap described in #1154.

On Windows, nvcc requires MSVC's cl.exe as its host compiler for preprocessing .cu files, even when the rest of the project is built with MinGW GCC. This creates two categories of issues that this PR addresses:

Source portability (commit 1): Headers included by .cu files contain constructs that cl.exe cannot process — <pthread.h> (POSIX-only), <ffnvcodec/dynlink_*.h> (installed in MinGW paths), C99 designated initializers (not supported by nvcc in C++ mode with MSVC). These are resolved with #ifdef DEVICE_CODE / #ifndef __CUDACC__ guards that are no-ops on Linux.

Build system (commit 2): The meson build auto-detects cl.exe via vswhere (without polluting PATH, which would cause meson to pick MSVC as the default compiler), discovers MSVC and Windows SDK include directories, and passes them to nvcc as -I flags since cl.exe runs outside a vcvars environment.

Prerequisites on Windows

  • MSYS2 with MinGW GCC toolchain
  • NVIDIA CUDA Toolkit (provides nvcc, bin2c)
  • nv-codec-headers built from git master (commit 876af32 or later). The latest release (n13.0.19.0) is missing several CudaFunctions members that vmaf uses (cuMemFreeHost, cuStreamCreateWithPriority, cuLaunchHostFunc, etc.). This is a pre-existing issue, not specific to this PR.
  • MSVC Build Tools (provides cl.exe, needed by nvcc for preprocessing)
  • Windows SDK (provides UCRT headers)

Test results

Tested on Windows 11 with:

  • CUDA Toolkit 13.2
  • MSVC Build Tools (VS 18), Windows 11 SDK
  • MinGW GCC 15.2.0 (MSYS2)
  • RTX 5090

The full pipeline works end-to-end:

  • meson setup -Denable_cuda=true configures successfully
  • ninja compiles all 7 .cu files to fatbin, library links
  • FFmpeg libvmaf_cuda filter scores at 1,135 fps on 1080p60 content (37-minute video scored in 2 minutes)

Transparency note

This patch was developed with assistance from Claude (Anthropic's AI), as reflected in the Co-Authored-By lines. All changes were manually tested on real hardware and real video content.

Closes #1154

🤖 Generated with Claude Code

birkdev and others added 2 commits March 16, 2026 03:09
On Windows, nvcc uses MSVC's cl.exe as its host compiler for
preprocessing. Several headers included by .cu files contain
MinGW/GCC-specific constructs that cl.exe cannot process:

- Remove unnecessary #include <pthread.h> from cuda/common.h.
  No declarations in this header use pthread types; consumers
  that need pthread (ring_buffer.c, libvmaf.c) include it directly.

- In cuda_helper.cuh, use <cuda.h> (from CUDA toolkit) instead of
  <ffnvcodec/dynlink_cuda.h> (from MinGW) for the DEVICE_CODE path.
  Device code only needs CUDA driver API types, not the dynamic
  loader machinery.

- In picture.h, use <cuda.h> and forward-declare VmafCudaState for
  the DEVICE_CODE path, avoiding ffnvcodec and libvmaf_cuda.h which
  pull in host-only dependencies.

- Guard C99 designated initializers in integer_adm.h with
  #ifndef __CUDACC__, as nvcc compiles .cu files as C++ where
  this syntax is not portable across host compilers.

- Guard #include "feature_collector.h" in ADM .cu files with
  #ifndef DEVICE_CODE. This header is host-only (contains pthread
  usage) and is never referenced by device kernel code.

These changes are no-ops on Linux where nvcc uses GCC (which handles
all of the above natively). They enable CUDA compilation on Windows
where nvcc must use MSVC's cl.exe.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Enable CUDA feature extraction on Windows (MSYS2/MinGW) by handling
the nvcc + MSVC toolchain requirements in the meson build:

- Auto-detect MSVC cl.exe via vswhere without adding it to PATH
  (which would cause meson to pick MSVC as the default C compiler
  instead of GCC). Pass it to nvcc via -ccbin.

- Discover MSVC and Windows SDK include directories and pass them
  as -I flags to nvcc, since cl.exe runs outside a vcvars
  environment and cannot find system headers otherwise.

- Add -D_USE_MATH_DEFINES so MSVC's math.h exposes M_PI.

On Linux, all new code paths are skipped (guarded by
host_machine.system() == 'windows').

Tested with CUDA 13.2, MSVC Build Tools (VS 18), Windows 11 SDK,
and MinGW GCC 15.2.0 on MSYS2. The full pipeline (nvcc -> fatbin ->
bin2c -> static lib -> shared lib) completes successfully, and the
resulting libvmaf integrates with FFmpeg's libvmaf_cuda filter.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@kylophone

Copy link
Copy Markdown
Collaborator

I believe nv-codec-headers should work on Windows.

@birkdev

birkdev commented Mar 16, 2026

Copy link
Copy Markdown
Author

You're right, nv-codec-headers works fine on Windows for the host code (compiled by GCC/MinGW). The <cuda.h> substitution only applies to the DEVICE_CODE path, where nvcc uses MSVC's cl.exe for preprocessing. cl.exe can't find headers installed in MinGW's include paths, so we use <cuda.h> from the CUDA toolkit for the device code path instead.

Separately, the CUDA code in vmaf requires nv-codec-headers built from git master (specifically commit 876af32 "add a few missing functions"), not the latest release n13.0.19.0. The release is missing several CudaFunctions members that vmaf uses (cuMemFreeHost, cuStreamCreateWithPriority, cuLaunchHostFunc, etc.). This is a pre-existing dependency issue, not specific to this PR.

@birkdev

birkdev commented Apr 13, 2026

Copy link
Copy Markdown
Author

Any thoughts on whether this approach is worth pursuing?

@1480c1

1480c1 commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

I fail to see how the contents of the PR is linked to the title. I do not see how this would allow someone to use any mingw-w64 based toolchain to compile libvmaf with cuda components without requiring msvc, which would imply you are already going to use msvc with with meson, so this whole setup looks a bit convoluted.

@birkdev

birkdev commented Apr 19, 2026

Copy link
Copy Markdown
Author

nvcc on Windows can't use GCC as its host compiler. You hit this yourself in #1154 (size_t redeclaration, "cuda sdk's nvcc doesn't understand gcc/g++ on Windows at all"). No PR removes that.

What this enables is libvmaf building in an MSYS2/MinGW environment with CUDA on. .cu files still route through cl.exe via nvcc (unavoidable), host code stays MinGW (matching how FFmpeg is typically cross-built with CUDA on Windows).

"Just use MSVC end to end" isn't a smaller alternative (see #1477, which does exactly that: ~700 net lines, pthread-win32 submodule, four dependency PRs). This is the minimum bridge, #1477 is the full rewrite.

Happy to rename the title.

@lusoris

lusoris commented Sep 8, 2026

Copy link
Copy Markdown

We have a concrete native Windows/MSVC + CUDA build sample from VMAFx that may help with the native-versus-MinGW discussion here. We carried the CUDA header separation and compiler-discovery work from this PR into our fork, with additional Windows compatibility code.

The successful Windows MSVC+CUDA job from today built checkout 491d8d2c34e6a3c96c74db408be484a54a0ce26f (the PR test merge of 2422c608962a7026fdec404def9f77ba827df2e1 into 7bafbb8cc9df61e372b075a934433db20515c23e). Its transcript contains NVCC fatbin generation, the native link /MACHINE:x64 /OUT:tools/vmaf.exe command, and installation of libvmaf.a, vmaf.exe and libvmaf_cuda.h.

The recorded configuration is windows-2025, Meson 1.12.0, host cl 19.51.36256, Windows SDK 10.0.26100.0, and the CUDA 13.3.1 installer reporting NVCC 13.3.73. One detail matters when reproducing this: the discovery code selected MSVC 14.29.30133 for NVCC's -ccbin, while Meson used the newer host compiler. The NVCC command also includes --allow-unsupported-compiler; this is the actual build transcript, not a claim that every VS/CUDA combination is vendor-supported.

The complete CI recipe at that checkout includes the SDK installation and environment setup. After those prerequisites, its cmd.exe build is:

set INCLUDE=%CD%\nvcodec-prefix\include;%INCLUDE%
set CFLAGS=/experimental:c11atomics
set CXXFLAGS=/experimental:c11atomics
meson setup core core\build --buildtype release ^
  --prefix %CD%\install -Denable_float=true ^
  --default-library=static -Denable_cuda=true ^
  -Denable_nvcc=true -Denable_sycl=false
ninja -v -C core\build install

The potentially reusable source pieces are our native Win32 pthread subset and CUDA custom-target include/dependency wiring. Our fork now has substantial C/C++ and dependency changes, so this recipe applies to that source checkout; it is not a drop-in build command for unmodified upstream v3.2.0 or independent validation of this PR's MinGW path. Also, that CI recipe clones nv-codec-headers HEAD without recording its SHA, so the transcript does not establish a fully pinned header-version matrix.

The scope of the evidence is native compilation, linking and installation. That Windows job has no GPU test step and does not upload its CUDA binaries; the run's MINGW64-vmaf artifact comes from a separate job. I cannot use it to confirm native Windows GPU scoring, numerical parity or the RTX 5090 throughput reported above.

One small issue in this PR's current b7b65e64 build logic: when the vswhere search returns no path but find_program('cl') succeeds, the fallback fills nvcc_ccbin_flags without assigning cl_path. The following MSVC include-directory command still concatenates cl_path. Assigning cl_path = cl_exe.full_path() in that fallback before constructing the flags would make that branch consistent. Our imported copy has the same issue.

When vswhere finds nothing but cl.exe is on PATH, the fallback filled
nvcc_ccbin_flags without assigning cl_path, so the MSVC include-directory
lookup below referenced an undefined variable and configure failed.
The kernel custom_target appends nvcc_ccbin_flags and nvcc_host_includes
unconditionally, but only the enable_nvcc branch defines them, so
configuring with -Denable_nvcc=false fails with an unknown variable.
Initialise both to empty lists before the branch.
@birkdev

birkdev commented Sep 17, 2026

Copy link
Copy Markdown
Author

Superseded by #1596, which builds the kernels with clang and needs neither nvcc nor MSVC on Windows. Closing this in favor of that approach.

@birkdev birkdev closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Request: CUDA feature extraction for Windows

4 participants