Repository navigation
fix(sycl): register device images in Windows MSVC builds - #1626
Merged
Merged
Conversation
lusoris
force-pushed
the
fix/windows-sycl-native-run
branch
2 times, most recently
from
September 29, 2026 16:48
988ab17 to
cbff095
Compare
The Windows MSVC+SYCL build never ran its kernels. Meson links MSVC builds with link.exe, which ignored -fsycl and never wrapped or registered the SYCL device images, so every kernel submit failed with "No kernel named ... was found" and 47 of 50 SYCL tests failed on an Arc B580. The CI leg only builds, so nothing had noticed. MSVC builds now compile the SYCL translation units as relocatable device code and run one `icpx -fsycl -fsycl-link` step over all SYCL objects. A small COFF patcher gives the resulting object an external anchor symbol that sycl/common.cpp asks for with /include, so link.exe pulls it out of vmaf.lib into every program that uses SYCL, including static consumers such as FFmpeg. Linux keeps the ADR-1360 per-TU code generation. The CI leg now checks the kernel registry without a GPU. On an Arc B580 and a UHD 770 all 51 SYCL tests and the whole suite pass, the default model and vmaf_v0.6.1 agree with the CPU within the gate on the Netflix pair and 50 frames of 4K, all 19 SYCL extractors pass the parity gate, and the SYCL output is bit-identical to the Linux container build of the same revision. Running the tests natively exposed that scripts/ci/run_meson_test.py returned 0 on Windows as soon as it started Meson (os.execvp has no exec semantics there), so the MinGW64 and ARM64 MSVC lanes passed after 17 to 19 tests. The runner now waits on Windows. Four Windows test-harness failures it hid are fixed, and the parity gate reads the cambi score by the name vmaf writes. ADR-1364, Research-2125. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
force-pushed
the
fix/windows-sycl-native-run
branch
from
September 30, 2026 07:42
cbff095 to
c26a68f
Compare
lusoris
added a commit
that referenced
this pull request
Sep 30, 2026
The rebase onto #1626 kept master's open copy of T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29 next to this branch's closing row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 30, 2026
…d Intel GPU runtime (#1629) * fix(docker): run the oneAPI release image on Debian 13 with the pinned Intel GPU runtime The published oneAPI image crashed every `vmaf --backend sycl` run on an Arc B580 (exit 139 right after device selection). The cause is the Intel GPU compute runtime (NEO 25.18) that Intel's oneapi-runtime:2025.3.1 image carries: swapping only that runtime for NEO 26.35 in the unchanged image fixes the crash, and swapping only the Level Zero loader does not. The image now follows the design build-config.env already recorded (ADR-1368). Builder and final stage start from the release track's debian:13-slim. scripts/ci/install-intel-oneapi.sh installs Intel's oneAPI 2026.1 compiler or SYCL runtime plus UMF at apt build 2026.1.1-325, with the repository key pinned by fingerprint. install-intel-ocloc.sh gains `--components build|runtime`, which adds the NEO GPU runtime at INTEL_NEO_VERSION (the package set the dev container uses) and the Level Zero loader at LEVEL_ZERO_VERSION, checked against GitHub's asset digests. The base-image gate now requires ONEAPI_BUILDER and ONEAPI_RUNTIME to equal RELEASE_BUILDER_BASE, and its distro exemption list is empty. The image is published as `-oneapi2026`. The `-oneapi2025` tag and the `final-oneapi2025` stage stay as aliases of the same image. In the rebuilt image the B580 and a UHD 770 match `--backend cpu` within the parity gate for the default model, psnr_hvs and ssimulacra2 on the Netflix pair and BBB 4K. The image shrinks from 5.86 GB to 2.39 GB. Closes T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: regenerate generated docs after rebasing onto master Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: regenerate generated docs after rebasing onto master Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(state): drop the open oneAPI B580 row this PR closes The rebase onto #1626 kept master's open copy of T-RELEASE-ONEAPI-IMAGE-B580-SIGSEGV-2026-09-29 next to this branch's closing row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
added a commit
that referenced
this pull request
Sep 30, 2026
) * perf(sycl): read the shared frame in psnr, psnr_hvs and motion_v2 The SYCL psnr_hvs, psnr and motion_v2 twins now read the planes the SYCL state uploads once per frame instead of converting and uploading their own copies. Scores are bit-identical to the previous twins on an Arc B580 and a UHD 770. - Opt-in shared Cb/Cr planes in common.cpp (vmaf_sycl_shared_chroma_init / _upload, vmaf_sycl_get_shared_plane): the first chroma-reading twin of a frame packs the chroma into pinned staging and uploads it with one DMA per plane; later twins reuse it. Luma-only runs never allocate chroma. - vmaf_sycl_queue_after_upload() gives twins on their own queue the input barriers the combined graph gets; a device-side slot fence orders each upload after the last readers of the slot it overwrites. - psnr_hvs: no host float conversion or private upload; two work-items per 8x8 block, one dispatch for all planes, per-block float expressions unchanged. 9- and 11-bit input now scores the raw sample like the CPU. - motion_v2: runs the ADR-1371 pipeline on the shared luma and keeps the frame through its cur_copy; no host copy or private upload. - psnr: chroma from the shared planes; one atomic per work-group. At 3840x2160, psnr_hvs drops from 17.1 to 7.6 ms per frame on the B580 and from 124 to 60 on the UHD 770, where psnr drops from 25.3 to 12.3. ADR-1369, Research-1369. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: regenerate generated docs after rebasing onto master Merge the RC2/RC3 disposition rows three-way by bug id: master (#1626) had two RC3 rows from an earlier keep-both resolution, and this branch adds three RC2 and six RC3 ids. Drop the open oneAPI B580 row #1629 closed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The native Windows MSVC+SYCL build now runs its SYCL kernels. It had never run on a GPU (the CI leg only builds), and on its first run on an Arc B580 and a UHD 770 every kernel submit failed with
No kernel named ... was found, because Meson links MSVC builds withlink.exe, which never registered the device images. MSVC builds now run oneicpx -fsycl -fsycl-linkstep whose objectvmaf.libpulls into every SYCL program (ADR-1364). Running the tests natively also showed thatscripts/ci/run_meson_test.pyreturned 0 on Windows before Meson finished, so the Windows MinGW64 and ARM64 MSVC lanes stopped gating after 17 to 19 tests; the runner now waits, and the four Windows test-harness failures it hid are fixed.What changed:
core/src/meson.build: on MSVC-syntax toolchains the SYCL TUs compile as relocatable device code, one device link (AOT device list,--offload-compress,-fp-model=precise, 8 parallel link jobs) wraps every image, andcore/src/sycl/coff_add_anchor.pygives that object the external symbolvmaf_sycl_device_images, whichcore/src/sycl/common.cpp/includes.-fsyclleaves the MSVC link line (/IGNORE:4078instead, as on the driver's own link). Linux keeps the ADR-1360 per-TU code generation.test_sycl_kernel_registrationasks the SYCL runtime for its kernel IDs (0 before, 96 now) without a GPU; theWindows MSVC+SYCLleg runs it.scripts/ci/run_meson_test.py: Windows runs Meson as a child and returns its status; POSIX keeps the ADR-1333 exec.test_vmaf_per_shot_input(CRT invalid-parameter fast-fail),test_device_target_header_dependencies(shell syntax in Ninja rules),test_vmaf_per_shot(/dev/zeroand FIFO cases are POSIX-only),dnntests (WSLbash.exe, backslash path spliced into Python).cambicell readCambi_feature_cambi_score, whichvmaf --jsonwrites ascambi.Results on the i9-12900K host (details in Research-2125):
syclsuite before / aftervmaf_v0.6.1, pooled VMAF vs Windows CPU (576x324 / 4K)Type
fix— bug fixbuild/ci— tooling / infrasycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. Run per tool on the changed files (clang-format 23.1.1, black 26.5.1, ruff 0.16.9, shellcheck 0.11.0.1, shfmt 3.13.1, markdownlint-cli2 0.23.3), plusmake docs-fragments-check, the state-row, ADR-link, research-ID and citation gates; the fullmake lintis POSIX-only on this Windows host.python3 scripts/ci/run_meson_test.py -- -C build./cross-backend-diffand the worst ULP is ≤ 2. No kernel changed; SYCL output is bit-identical to the Linux build of the same revision, and the parity gate results are in the table above..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md).!orBREAKING CHANGE:and the migration path is documented below. Not breaking.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR: closedT-SYCL-WINDOWS-MSVC-KERNELS-UNREGISTERED-2026-09-29,T-CI-WINDOWS-MESON-TEST-RUNNER-EXIT-0-2026-09-29,T-TEST-WINDOWS-HARNESS-MASKED-FAILURES-2026-09-29,T-CI-PARITY-GATE-CAMBI-KEY-2026-09-29; opened (RC3)T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29,T-SYCL-FLOAT-SSIM-XE2-XELP-CALIBRATION-2026-09-29.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
Performance (if
perforfeat)Diagnostic only (native builds are not canonical bench rows, ADR-1102). Default model,
(t(N) - t(2)) / (N - 2), minimum of five interleaved runs: 576x324 B580 2.56 ms, UHD 770 5.30 ms, CPU 2.48 ms (1 thread) / 0.51 ms (16); 4K B580 44.8 ms, UHD 770 66.9 ms, CPU 77.5 / 20.4 ms. The Linux container build of the same revision measures 2.38 / 10.56 ms and 52.3 / 76.1 ms. At 4K the B580 waits on the CPU ADM branch (no AIM device pass,T-GPU-ADM-AIM-DEVICE-PASS-MISSING-SYCL-HIP-2026-09-05);vmaf_v0.6.1, all on the device, takes 9.3 ms per 4K frame on the B580.Deep-dive deliverables (ADR-0108)
docs/research/2125-windows-native-sycl-run.md.## Alternatives considered.AGENTS.mdinvariant note —core/src/sycl/AGENTS.md"MSVC builds register images through one explicit device link (ADR-1364)".changelog.d/fixed/windows-sycl-native-run.md.docs/rebase-notes.md"ADR-1364 — Windows MSVC SYCL device link, Windows test runner exit status".Reproducer
Known follow-ups
T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29andT-SYCL-FLOAT-SSIM-XE2-XELP-CALIBRATION-2026-09-29(RC3).🤖 Generated with Claude Code