Skip to content

fix(sycl): register device images in Windows MSVC builds - #1626

Merged
lusoris merged 4 commits into
masterfrom
fix/windows-sycl-native-run
Sep 30, 2026
Merged

lusoris merged 4 commits into
masterfrom
fix/windows-sycl-native-run

Conversation

@lusoris

@lusoris lusoris commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

The native Windows MSVC+SYCL build now runs its SYCL kernels. It had never run on a GPU (the CI leg only builds), and on its first run on an Arc B580 and a UHD 770 every kernel submit failed with No kernel named ... was found, because Meson links MSVC builds with link.exe, which never registered the device images. MSVC builds now run one icpx -fsycl -fsycl-link step whose object vmaf.lib pulls into every SYCL program (ADR-1364). Running the tests natively also showed that scripts/ci/run_meson_test.py returned 0 on Windows before Meson finished, so the Windows MinGW64 and ARM64 MSVC lanes stopped gating after 17 to 19 tests; the runner now waits, and the four Windows test-harness failures it hid are fixed.

What changed:

  • core/src/meson.build: on MSVC-syntax toolchains the SYCL TUs compile as relocatable device code, one device link (AOT device list, --offload-compress, -fp-model=precise, 8 parallel link jobs) wraps every image, and core/src/sycl/coff_add_anchor.py gives that object the external symbol vmaf_sycl_device_images, which core/src/sycl/common.cpp /includes. -fsycl leaves the MSVC link line (/IGNORE:4078 instead, as on the driver's own link). Linux keeps the ADR-1360 per-TU code generation.
  • test_sycl_kernel_registration asks the SYCL runtime for its kernel IDs (0 before, 96 now) without a GPU; the Windows MSVC+SYCL leg runs it.
  • scripts/ci/run_meson_test.py: Windows runs Meson as a child and returns its status; POSIX keeps the ADR-1333 exec.
  • Windows harness fixes: test_vmaf_per_shot_input (CRT invalid-parameter fast-fail), test_device_target_header_dependencies (shell syntax in Ninja rules), test_vmaf_per_shot (/dev/zero and FIFO cases are POSIX-only), dnn tests (WSL bash.exe, backslash path spliced into Python).
  • Parity gate: the cambi cell read Cambi_feature_cambi_score, which vmaf --json writes as cambi.
  • Docs: new SYCL on Windows how-to; overview, backends index, oneAPI page and nav link to it.

Results on the i9-12900K host (details in Research-2125):

Check Arc B580 UHD 770
sycl suite before / after 3/50 / 51/51 not run / 51/51
whole suite (runner fixed) 240 OK, 1 skipped (fixture) 240 OK, 1 skipped (fixture)
default model, pooled VMAF vs Windows CPU (576x324 / 4K) 7.95e-7 / 2.85e-7 same, bit-identical to B580
vmaf_v0.6.1, pooled VMAF vs Windows CPU (576x324 / 4K) 2.48e-5 / 2.05e-5 same
19 SYCL extractors vs CPU (parity gate, 576x324 and 4K) all within tolerance all within tolerance
Windows SYCL vs Linux container SYCL (same revision) features bit-identical, VMAF ≤ 6.1e-12 same

Type

  • fix — bug fix
  • build / ci — tooling / infra
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. Run per tool on the changed files (clang-format 23.1.1, black 26.5.1, ruff 0.16.9, shellcheck 0.11.0.1, shfmt 3.13.1, markdownlint-cli2 0.23.3), plus make docs-fragments-check, the state-row, ADR-link, research-ID and citation gates; the full make lint is POSIX-only on this Windows host.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. No kernel changed; SYCL output is bit-identical to the Linux build of the same revision, and the parity gate results are in the table above.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. No extractor changed.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not breaking.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR: closed T-SYCL-WINDOWS-MSVC-KERNELS-UNREGISTERED-2026-09-29, T-CI-WINDOWS-MESON-TEST-RUNNER-EXIT-0-2026-09-29, T-TEST-WINDOWS-HARNESS-MASKED-FAILURES-2026-09-29, T-CI-PARITY-GATE-CAMBI-KEY-2026-09-29; opened (RC3) T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29, T-SYCL-FLOAT-SSIM-XE2-XELP-CALIBRATION-2026-09-29.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. No golden value changes.

Cross-backend numerical results

parity gate, --backends cpu sycl, Netflix pair (576x324) / BBB 4K, Arc B580 = UHD 770
adm 1e-6/2e-6  cambi 0/0  ciede 1.2e-5/9.7e-5  float_adm 2.5e-5/1.6e-5  float_moment 0/0
float_motion 3e-6/2.7e-5  float_ms_ssim(+lcs) 1e-6/1e-6  float_psnr 0/0  float_ssim 1e-6/-(scale 8)
float_vif 2.8e-5/3e-6  motion_v2 0/0  psnr 0/0  psnr_hvs 8.3e-5/8.43e-4 (tol 3.34e-3)  ssimulacra2 0/0  vif 1e-6/1e-6
by name at --precision max: motion_sycl 1.26e-5/5.6e-6  integer_ssim_sycl 1.5e-8/7.8e-8  speed_chroma/temporal 0/0

Performance (if perf or feat)

Diagnostic only (native builds are not canonical bench rows, ADR-1102). Default model, (t(N) - t(2)) / (N - 2), minimum of five interleaved runs: 576x324 B580 2.56 ms, UHD 770 5.30 ms, CPU 2.48 ms (1 thread) / 0.51 ms (16); 4K B580 44.8 ms, UHD 770 66.9 ms, CPU 77.5 / 20.4 ms. The Linux container build of the same revision measures 2.38 / 10.56 ms and 52.3 / 76.1 ms. At 4K the B580 waits on the CPU ADM branch (no AIM device pass, T-GPU-ADM-AIM-DEVICE-PASS-MISSING-SYCL-HIP-2026-09-05); vmaf_v0.6.1, all on the device, takes 9.3 ms per 4K frame on the B580.

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/2125-windows-native-sycl-run.md.
  • Decision matrix — in ADR-1364 ## Alternatives considered.
  • AGENTS.md invariant note — core/src/sycl/AGENTS.md "MSVC builds register images through one explicit device link (ADR-1364)".
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/fixed/windows-sycl-native-run.md.
  • Rebase note — docs/rebase-notes.md "ADR-1364 — Windows MSVC SYCL device link, Windows test runner exit status".

Reproducer

# Windows, cmd, after the environment in docs/backends/sycl/windows.md
meson setup build core --buildtype release --default-library=static -Denable_float=true -Dcpp_std=c++latest -Denable_cuda=false -Denable_sycl=true
ninja -C build
set ONEAPI_DEVICE_SELECTOR=level_zero:0
python scripts\ci\run_meson_test.py -- -C build --suite sycl --print-errorlogs
build\tools\vmaf.exe -r src01_hrc00_576x324.yuv -d src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 --backend sycl --json -o out.json

Known follow-ups

  • The two Windows test lanes now gate their whole suites for the first time. Failures specific to MinGW-w64 gcc or the ARM64 MSVC toolset could not be reproduced on this x64 host and will show up in this PR's checks.
  • The device link was verified with oneAPI 2025.1.1; the CI leg's 2025.3.0 builds it and checks the kernel registry, but no GPU has run a 2025.3.0 Windows build.
  • T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29 and T-SYCL-FLOAT-SSIM-XE2-XELP-CALIBRATION-2026-09-29 (RC3).

🤖 Generated with Claude Code

@github-actions github-actions Bot added the type:bug Something isn't working label Sep 29, 2026
@lusoris
lusoris force-pushed the fix/windows-sycl-native-run branch 2 times, most recently from 988ab17 to cbff095 Compare September 29, 2026 16:48
lusoris and others added 4 commits September 30, 2026 09:40
The Windows MSVC+SYCL build never ran its kernels. Meson links MSVC
builds with link.exe, which ignored -fsycl and never wrapped or
registered the SYCL device images, so every kernel submit failed with
"No kernel named ... was found" and 47 of 50 SYCL tests failed on an
Arc B580. The CI leg only builds, so nothing had noticed.

MSVC builds now compile the SYCL translation units as relocatable
device code and run one `icpx -fsycl -fsycl-link` step over all SYCL
objects. A small COFF patcher gives the resulting object an external
anchor symbol that sycl/common.cpp asks for with /include, so link.exe
pulls it out of vmaf.lib into every program that uses SYCL, including
static consumers such as FFmpeg. Linux keeps the ADR-1360 per-TU code
generation. The CI leg now checks the kernel registry without a GPU.

On an Arc B580 and a UHD 770 all 51 SYCL tests and the whole suite
pass, the default model and vmaf_v0.6.1 agree with the CPU within the
gate on the Netflix pair and 50 frames of 4K, all 19 SYCL extractors
pass the parity gate, and the SYCL output is bit-identical to the
Linux container build of the same revision.

Running the tests natively exposed that scripts/ci/run_meson_test.py
returned 0 on Windows as soon as it started Meson (os.execvp has no
exec semantics there), so the MinGW64 and ARM64 MSVC lanes passed
after 17 to 19 tests. The runner now waits on Windows. Four Windows
test-harness failures it hid are fixed, and the parity gate reads the
cambi score by the name vmaf writes.

ADR-1364, Research-2125.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the fix/windows-sycl-native-run branch from cbff095 to c26a68f Compare September 30, 2026 07:42
@lusoris
lusoris merged commit 8228503 into master Sep 30, 2026
102 checks passed
@lusoris
lusoris deleted the fix/windows-sycl-native-run branch September 30, 2026 08:20
lusoris added a commit that referenced this pull request Sep 30, 2026
The rebase onto #1626 kept master's open copy of
T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29 next to this branch's
closing row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…d Intel GPU runtime (#1629)

* fix(docker): run the oneAPI release image on Debian 13 with the pinned Intel GPU runtime

The published oneAPI image crashed every `vmaf --backend sycl` run on an
Arc B580 (exit 139 right after device selection). The cause is the Intel
GPU compute runtime (NEO 25.18) that Intel's oneapi-runtime:2025.3.1 image
carries: swapping only that runtime for NEO 26.35 in the unchanged image
fixes the crash, and swapping only the Level Zero loader does not.

The image now follows the design build-config.env already recorded
(ADR-1368). Builder and final stage start from the release track's
debian:13-slim. scripts/ci/install-intel-oneapi.sh installs Intel's oneAPI
2026.1 compiler or SYCL runtime plus UMF at apt build 2026.1.1-325, with
the repository key pinned by fingerprint. install-intel-ocloc.sh gains
`--components build|runtime`, which adds the NEO GPU runtime at
INTEL_NEO_VERSION (the package set the dev container uses) and the Level
Zero loader at LEVEL_ZERO_VERSION, checked against GitHub's asset digests.
The base-image gate now requires ONEAPI_BUILDER and ONEAPI_RUNTIME to equal
RELEASE_BUILDER_BASE, and its distro exemption list is empty.

The image is published as `-oneapi2026`. The `-oneapi2025` tag and the
`final-oneapi2025` stage stay as aliases of the same image.

In the rebuilt image the B580 and a UHD 770 match `--backend cpu` within
the parity gate for the default model, psnr_hvs and ssimulacra2 on the
Netflix pair and BBB 4K. The image shrinks from 5.86 GB to 2.39 GB.

Closes T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(state): drop the open oneAPI B580 row this PR closes

The rebase onto #1626 kept master's open copy of
T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29 next to this branch's
closing row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
Merge the RC2/RC3 disposition rows three-way by bug id: master (#1626)
had two RC3 rows from an earlier keep-both resolution, and this branch
adds three RC2 and six RC3 ids. Drop the open oneAPI B580 row #1629
closed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
)

* perf(sycl): read the shared frame in psnr, psnr_hvs and motion_v2

The SYCL psnr_hvs, psnr and motion_v2 twins now read the planes the SYCL
state uploads once per frame instead of converting and uploading their own
copies. Scores are bit-identical to the previous twins on an Arc B580 and a
UHD 770.

- Opt-in shared Cb/Cr planes in common.cpp (vmaf_sycl_shared_chroma_init /
  _upload, vmaf_sycl_get_shared_plane): the first chroma-reading twin of a
  frame packs the chroma into pinned staging and uploads it with one DMA per
  plane; later twins reuse it. Luma-only runs never allocate chroma.
- vmaf_sycl_queue_after_upload() gives twins on their own queue the input
  barriers the combined graph gets; a device-side slot fence orders each
  upload after the last readers of the slot it overwrites.
- psnr_hvs: no host float conversion or private upload; two work-items per
  8x8 block, one dispatch for all planes, per-block float expressions
  unchanged. 9- and 11-bit input now scores the raw sample like the CPU.
- motion_v2: runs the ADR-1371 pipeline on the shared luma and keeps the
  frame through its cur_copy; no host copy or private upload.
- psnr: chroma from the shared planes; one atomic per work-group.

At 3840x2160, psnr_hvs drops from 17.1 to 7.6 ms per frame on the B580 and
from 124 to 60 on the UHD 770, where psnr drops from 25.3 to 12.3.
ADR-1369, Research-1369.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Merge the RC2/RC3 disposition rows three-way by bug id: master (#1626)
had two RC3 rows from an earlier keep-both resolution, and this branch
adds three RC2 and six RC3 ids. Drop the open oneAPI B580 row #1629
closed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant