Repository navigation
Conversation
Record the GCC C + icpx C++ / SYCL toolchain for dev/Containerfile and Containerfile.vmafx. Master's ADR-1461 (strict FP on every translation unit) and ADR-1495 (-no-intel-lib=libimf on every Intel LLVM link) already make an icx build return a GCC build's CPU scores, so this ADR keeps only the delta: the faster GCC CPU path (8-12 % single-thread on the old base), icpx's strict spelling for the C++ units of a mixed toolchain, and C++ linking of the SYCL-on test executables. Adds the SYCL overview section on choosing the C compiler, the digest, perf changelog fragment and AGENTS.d topic pages. Drafted as ADR-1440; renumbered because master took 1440. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
- dev/Containerfile: CC=gcc CXX=icpx, -Db_lto=false (GCC LTO objects
cannot go through the icpx link). glibc libm comes from ADR-1495's
per-language link policy, so no -no-intel-lib option is passed.
- core/src/meson.build: with an Intel LLVM C++ compiler and another C
compiler, every C++ translation unit takes icpx's strict spelling
(-fp-model=precise -ffp-contract=off) as a project argument, next to
ADR-1461's policy. Replaces the libsvm-only -fp-model=precise, which
after ADR-1461's project argument re-enabled contraction.
- core/test/test_strict_fp_compiler_args.py: executes the new block per
compiler pair, pins it as a project argument, and checks each compile
command against its own language's compiler (icx/icpx must carry
-fp-model=precise).
- core/test/meson.build: test_link_kwargs ({'link_language': 'cpp'} with
SYCL on) on every test executable that does not set link_language;
sycl_dependency's -fsycl / AOT link args are icpx-only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Three-stage image (runtime / build / prod) with the Intel GPU stack from build-config.env, libvmaf (SYCL, ADR-1593 toolchain), FFmpeg at the series.txt tag with every fork patch, and a build-time Netflix CPU golden gate. Not a release image; complements docker/Dockerfile.* which ship no FFmpeg and run no goldens. Adds the docker-production.md section and the changelog fragment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…594) Port of the multistage image from chore/containerfile-vmafx-multistage, adapted to master: GPU stack via scripts/ci/install-intel-ocloc.sh (runtime/build sets, build-config.env pins) so ocloc matches the driver for AOT (ADR-1360/1368); FFmpeg via checkout-annotated-tag.sh at FFMPEG_TAG=n9.0.2 (the series.txt tag); GCC + icpx toolchain; pytest cache_dir instead of -p no:cacheprovider; EUPL-1.2 SPDX headers. run-all-tests.sh runs Meson through scripts/ci/run_meson_test.py and is added to the runner inventory of test_meson_secret_env_sanitization.py; reference_report.py pins version=vmaf_v0.6.1 (the CLI default model is not v0.6.1 on master). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
`$SUDO DEBIAN_FRONTEND=noninteractive apt-get ...` fails with an empty $SUDO: bash identifies assignment words before expanding $SUDO, so the assignment becomes the command name. Pass it through env. Also fix the 26.04 codename (Resolute). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Regenerated by make docs-fragments-write for ADR-1439, ADR-1593 and ADR-1594 and the three new changelog fragments; rebase notes for the toolchain, test link kwargs, the mixed-toolchain C++ strict FP block, Containerfile.vmafx and the golden sync. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
… ADR refs) - Copy scripts/ci/check-msvc-clz-shim.sh before meson setup; core/test/meson.build find_program()s it at configure time. - AOT targets: the option's default list, Xe2 included. Master fixed the Xe2 AOT build (VMAFx#1842, ADR-1468: no kernel requires sub-group size 8), so no exclusion is needed. - Drop -Dcpp_link_args=-no-intel-lib=libimf: ADR-1495's link policy passes it on every icpx link. - Install the golden-gate Python stack from the hash lock python/requirements-test-lock.txt with --require-hashes, as the check-python-dependency-locks hook requires. - Replace citations of the old branch's ADR numbers, which name other decisions on master (1134/1135/1140/1141/1142/1145), with ADR-1439, ADR-1593, ADR-1594 or plain descriptions; refresh scripts/ci/source-adr-citations.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
vmaf.config downloads missing test videos from github.com/Netflix/vmaf_resource with a single urlretrieve; on the NAS build host the image golden gate lost 3-4 tests per run to timeouts and dropped connections. Retry URLError / HTTPException / OSError up to 4 attempts with 2/4/8 s back-off; HTTP errors (404) still fail at once. Unit-tested (config_download_retry_test.py, 3 cases). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
ADR-1594 consequences and research digest with the NAS results: build golden gate 271 passed, prod image 980 MB, Netflix pairs PASS on CPU and SYCL, SYCL == CPU bit for bit on src01, FFmpeg filters present, math on glibc libm. That run excluded the Xe2 AOT targets, which master could not compile then; state.md records T-SYCL-XE2-SUBGROUP8-AOT-2026-10-02 as closed, fixed on master by VMAFx#1842 (ADR-1468), and the image now uses the default target list. Rebase notes and rendered CHANGELOG updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…nerfile.vmafx Only the oneVPL dispatcher (libvpl2) was installed, so every *_qsv decoder/encoder failed with "Error creating a MFX session: -9" and the libvmaf_sycl zero-copy path could not be exercised. Add libmfx-gen1.2. Verified on an Arc A380: hevc_qsv works, and libvmaf_sycl zero-copy on QSV-decoded 8-bit and 10-bit (P010) HEVC scores bit for bit like the CPU path with no NaN frames (ADR-1594 digest). Found and recorded as open: the zero-copy path silently drops requested non-SYCL features (T-SYCL-ZEROCOPY-DROPS-NON-SYCL-FEATURES-2026-10-02). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Both reproduce on master b01ffe4 built with icx in the vmafx:build-ocloc image (oneAPI 2026.1.1, Arc A380), so neither is caused by this branch: - T-SYCL-LD-BIND-NOW-LIBIMF-IFUNC-2026-10-03: under LD_BIND_NOW=1 a SYCL build's vmaf segfaults ("Relink libimf.so with libm.so.6 for IFUNC symbol cosf"); libimf comes in through the oneAPI runtime's libur_loader.so.0, not through libvmaf. test_icx_system_libm fails. - T-SYCL-ORDERED-SUM-MALLOC-PERTURB-2026-10-03: test_sycl_ordered_sum's device walk fails whenever MALLOC_PERTURB_ is non-zero (meson test sets it) and passes without it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…olchain The short post-rebase run on the 576x324 pair suggested the GCC speed gain was gone. A 1080p run (60 frames, 9 interleaved rounds, 1-3 % spread, bit-identical outputs across icx, hybrid and GCC) shows the hybrid 4.9 %, 10.9 % and 6.5 % faster than icx single-threaded for vmaf_v0.6.1, vmaf_float_v0.6.1 and cambi, equal to pure GCC, and within noise for SpEED, 16 threads and SYCL. ADR, digest and changelog fragment now say so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…ADR-1593) The rebase onto 2889f96 brought two test executables that do not pass kwargs test_link_kwargs: test_float_bits and test_hip_device_selection. ADR-1593's invariant (core/test/AGENTS.d/hybrid-toolchain-test-link.md) is that every test executable without its own link_language passes them, so a hybrid CC=gcc CXX=icpx SYCL build links every test with icpx. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
The fork no longer patches Intel's libimf.so (the ADR-1563 commits are dropped from this branch): the Intel EULA for Developer Tools forbids modifying the Materials (3.1(iv)) and allows modifying Redistributables only in Source Code form (2.1.C/2.1.D). Operator decision: "Да, убрать (Recommended)". - docs/state.md: T-SYCL-LD-BIND-NOW-LIBIMF-IFUNC-2026-10-03 stays open as the one row for the runtime defect (cause, affected programs, no patching per the EULA, no LD_PRELOAD workaround, Intel report drafted); test_icx_system_libm itself is fixed on master by the list-mode trace (T-ICX-LIBM-TEST-BIND-NOW-CRASH-2026-10-04). New open row T-ADR-ALLOCATOR-SHALLOW-GRAFTS-LOCAL-TIPS-2026-10-04: in an already shallow clone, next-free.sh --claim grafted local branch tips into .git/shallow. - SYCL overview: a known-issue section (do not run SYCL builds with LD_BIND_NOW=1; LD_PRELOAD=libimf.so changes scores, ADR-1495). - Research-1593 notes how the two failures it measured ended. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Master f224b1b took ADR-1562 and other branches claimed 1561 to 1568, so the branch's ADR-1561 and ADR-1562 moved to 1593 and 1594. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…rt main Standards (standardsctl audit) flagged two new findings in files from the Containerfile.vmafx harness: run-all-tests.sh ran without set -e (HISS-07) and reference_report.py's main() was 66 lines (HISS-04, limit 60). run-all-tests.sh now uses set -euo pipefail. Commands whose non-zero status is an expected outcome (a failing section, grep counting zero matches, the REPORT section's exit code) are guarded with || true / || rc=$?, and the image build failure is checked with an if around the pipeline, so a failing section is still reported in the summary instead of stopping the run. reference_report.py's main() is split into golden_table() and parity_table() with unchanged output and exit status. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
standardsctl audit reported HISS-07 findings in run-all-tests.sh: '|| true' discards any failure. - a section's exit status is kept in c_rc / golden_rc / material_rc and shown in the summary; - grep_opt treats grep's 'no match' (status 1) as an answer and still stops on a real error (status 2). run-all-tests.sh exercised with stub engines that fail every section and that produce empty logs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
The merge-base mypy gate flagged calls to the untyped golden_table() / parity_table() split out of main(); every function in the script is now annotated, and mypy reports no issues for the file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…ging The first full build of Containerfile.vmafx after the rebase failed in the build stage: apt's meson pulls python3-packaging, which ships no pip RECORD file, so 'pip install --require-hashes -r python/requirements-test-lock.txt' could not uninstall it to put the locked version in place (uninstall-no-record-file). --ignore-installed installs the locked versions over it. Verified with a full build: the image builds, the in-build Netflix golden gate passes (280 passed, 3 skipped), vmaf-selftest passes on the Arc A380 (CPU and SYCL equal), and a QSV zero-copy libvmaf_sycl run matches the CPU libvmaf filter on all 20 pooled metrics. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…L on (ADR-1593) Nineteen test executables added on master since the branch was last rebased (the portable Metal arithmetic tests, test_integer_vif_sv_sq and the two SYCL zero-copy admission tests) take kwargs : test_link_kwargs like every other test executable; without it a CC=gcc CXX=icpx build with SYCL on fails to link them (gcc: error: unrecognized command-line option '-fsycl'). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
… it reads The three NETFLIX_VMAF rows pointed at quality_runner_test.py lines 151, 380 and 378; on the current master the assertions are on 152, 384 and 382. The src01 provenance note no longer names the golden-sync ADR, which is a change of its own: the value is Netflix upstream's current one, and the fork's code already produces it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…master VMAFx/vmafx master has its own ADR-1593 (Helm node FUSE and eBPF) and ADR-1594 (zstd images, zopfli zips), and its open branches claim numbers up to 1713. The hybrid-toolchain record becomes ADR-1714 and the Containerfile.vmafx record ADR-1715, with their research digests; every reference this branch added follows. The golden-sync bullet leaves this branch's rebase note (that change is a pull request of its own), and the zero-copy finding in Research-1715 now points at ADR-1688, which refuses those features on master. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
… map Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…lists Ported from the Tualua fork's ADR-1596. With batched Level Zero command lists an Arc A380 silently dropped the zero-copy VA import from a random frame on once each frame's DMA-BUF import reused the GPU address the previous import had just freed; only the primary queue's mode matters. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Zero-copy cambi differed from host upload in 3 to 7 of 10 FFmpeg runs on an Arc A380 under UR_L0_USE_IMMEDIATE_COMMANDLISTS=0. Root cause in the driver stack: every frame the VA import maps a surface with zeMemAllocDevice (DMA-BUF) at the GPU address the previous frame's import just freed, and from a random frame on the batched primary queue's submissions (the import copy / de-tile and anything batched with it) are silently dropped. The shared slots keep two old frames, every extractor scores them, and wait_and_throw() reports nothing. The primary queue, which runs the import, now carries the DPC++ immediate_command_list property (under SYCL_EXT_INTEL_QUEUE_IMMEDIATE_COMMAND_LIST; ignored off Level Zero). An import-only immediate queue beside a batched primary queue does not help (8 of 12 runs failed); other queues keep the process's mode. Ported from the Tualua fork (ADR-1596 there). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
test_sycl_ordered_sum failed on an Arc A380 whenever UR_L0_USE_IMMEDIATE_COMMANDLISTS=0 was set (the dev launcher sets it): each case allocates and frees its device blocks, and under batched command lists the driver silently drops work against memory mapped at a just-freed address (ADR-1596), so the walk read the previous case's data. The probe's queue now carries immediate_command_list like the library's primary queue. 3/3 failures under =0 before, 5/5 passes after; the ordered-sum algorithm (ADR-1446) was never at fault. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…wait Corrects 73a720240, whose root cause was wrong. The probe built two device blocks from std::vector temporaries: DeviceBlock only enqueues the copy, and the temporaries were destroyed before q.wait_and_throw(). With deferred (batched) command lists the copy runs at the wait and reads freed heap memory, so the walk summed garbage (undefined behaviour under the SYCL spec; with immediate command lists the copy ran at enqueue and hid it). It was never a driver drop and device address reuse plays no part: a standalone repro keeps failing with fresh device addresses and passes with reused ones once the source outlives the copy. The sources are now named and live until the wait; the immediate command list property added by 73a720240 is removed. 5/5 under UR_L0_USE_IMMEDIATE_COMMANDLISTS=0 via meson test, 3/3 with MALLOC_PERTURB_=165. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
The function only enqueued the copy-queue memcpy from the caller's host pointer. vmaf_sycl_import_d3d11_surface() unmapped and released its staging texture straight after the call, so the copy could read unmapped memory, and nothing ordered the copy before the extractors: vmaf_read_pictures_sycl() waits on the primary queue only and vmaf_sycl_wait_copy_queue() had no caller. Same freed-source pattern as the ordered-sum probe; a race even with immediate command lists. The function now waits on the copy queue before it returns. The public header and docs/api/gpu.md state the contract. The D3D11 path was read from source and not run on Windows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…og level vmaf_sycl_print_timing() wrote "[vmaf-sycl] timing: ..." to stderr with fprintf, and every SYCL flush calls it, so FFmpeg's libvmaf_sycl filter printed it even at -loglevel error (the filter maps -loglevel to VmafConfiguration.log_level). It now goes through vmaf_log at VMAF_LOG_LEVEL_INFO. The logging invariant no longer lists it among the print-on-request exceptions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…contracts Ported from the Tualua fork and fitted to the luma-only zero-copy path of ADR-1688: - scripts/test/sycl-dev-container.sh builds libvmaf and FFmpeg (with the patch series) from the worktree in localhost/vmafx:build-ocloc and runs meson tests or commands on the Intel GPU; - scripts/test/zerocopy-e2e.sh QSV-encodes the Netflix src01 pair and the 1080p checkerboard at 8 and 10 bit and runs every case through a CPU, a host-upload and a zero-copy leg at score_fmt=%.17g, with --repeat N; - scripts/test/zerocopy_e2e_compare.py turns the legs into verdicts. Stage 1 is what ADR-1688 admits and must equal host upload exactly; the stage-2/3 cases need chroma (the post-1.0 import of ADR-1685) and must be refused loudly, naming the extractor; - ffmpeg-patches/test/check-sycl-feature-routing.sh checks the twin routing and the NV12 / P010 guard in patch 0005 (the import-failure behaviour is VMAFx#2110's and is not checked here); - make sycl-zerocopy-contract runs both device-free contracts in the FFmpeg Patch Stack job and as a pre-commit hook. The harness never cuts a leg with -frames:v or -t (FFmpeg applies them after the filtergraph, so a libvmaf* filter can score an extra pair). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…(ADR-1764) Patch 0005 resolves each feature= name through vmaf_feature_backend_twin() before registering it, as the vmaf CLI does (ADR-1359). On QSV zero-copy input a feature without a usable twin fails at configuration, naming the feature and the reason; on software input the filter warns and computes it on the CPU. A twin that needs chroma is still refused at the first frame by ADR-1688's admission check, now under the twin's name. QSV zero-copy accepts NV12 and P010 surfaces only. Ported from the Tualua fork (ADR-1595 there). Its CPU-extractor pre-pass, per-extractor -ENOTSUP returns and fatal VA import are left out: ADR-1688 and VMAFx#2110 cover them. The series replays on n9.0.2 (20 of 20). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
…or the zero-copy hardening - docs/usage/ffmpeg.md: how libvmaf_sycl resolves feature= names and its messages; "Score exactly N frames": -frames:v / -t are applied after the filtergraph, so any libvmaf* filter can score one more pair; trim both inputs instead. The timing summary is an info-level message. - docs/backends/sycl/zero-copy.md: routing and NV12 / P010 section; the QSV example opens VA-API on the render node and trims instead of -frames:v. - docs/development/sycl-zerocopy-testing.md: fitted to the luma-only path (stage 1 = ADR-1688's admitted set) and the real refusal messages. - docs/state.md: closes T-SYCL-UPLOAD-PLANE-NO-COMPUTE-FENCE-2026-10-05 and T-SYCL-ORDERED-SUM-MALLOC-PERTURB-2026-10-03, records the fork's T-SYCL-ZEROCOPY-IMPORT-DROPPED, T-SYCL-TIMING-LINE-IGNORES-LOGLEVEL and T-SYCL-FILTER-FEATURE-NAMES-NOT-ROUTED as fixed, and the -frames:v finding as not a fork bug (T-FFMPEG-OUTPUT-LIMIT-EXTRA-SCORED-FRAME). - Research-1763, the ADR-1764 index row, changelog fragments, rebase note, core/src/sycl/AGENTS.md invariants (synchronous upload_plane, timing log). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
… map Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RrrFjFBShTNNqDhdnZRtGt
|
Thanks. See my comment on #2216: both get a full review right after the rc.3 tag, as RC4 scope. The immediate-command-list finding (ADR-1763 in this branch) is the part we most want to take. |
|
Thanks for this hardening. Here is how it relates to the RC4 import work, so it doesn't get lost. #2342 (RC4 WP3) replaces the SYCL import layer around this code, but it leaves the luma-only VA path of ADR-1688 unchanged. That is the path this PR hardens, so the PR is still relevant. It overlaps #2342 in 17 files ( |
…GL textures (RC4 WP3, ADR-2091) The VMAFx API now scores frames that already live on a SYCL device without a copy through the host, and they score bit for bit as the same frames uploaded from the host for every SYCL twin declared exact. - SYCL devices are the Level Zero GPUs, opened by index or in the context of the caller's queue; each has one in-order library queue with immediate command lists, the precaution of draft PR #2217 (evaluated: its i915 defect did not reproduce on xe). - USM planes of any address and pitch and linear dma-bufs are bound; Intel Y-tiled and Tile4 dma-bufs are de-tiled on the device with the address math the VA-surface import now shares (core/src/sycl/detile.h); NV12 / P010 / P016 are planarised on the device. GL textures are imported through an EGL dma-buf export. - Every SYCL twin that stages its own planes, the chroma twins included, copies a device picture on the device through vmaf_sycl_picture_read_plane(), which records the read on the frame. - Acquire and release are joins (an empty kernel with depends_on()), not barriers: a barrier on the event of a command behind a host task waits on the host (Research-2159). SYCL_EVENT acquire fences are waited on by the device; HOST, SYNC_FILE, GL_SYNC and a dma-buf's implicit write fences are checked on the host for the import rule's retry. HOST and SYCL_EVENT release fences and the release callback follow the last reader in every context. - The sync_file, dma-buf fence and GL sync helpers move to core/src/vmafx/sync_object.c, shared with the CUDA lane. - Windows shared textures and SYNC_FILE release fences are refused, tracked as T-SYCL-VMAFX-IMPORT-REMAINDER-2026-10-06. Landing (Q-083): the branch's diff (15bd091..5dafa1c) squashed onto (a082db7): the HIP lane's vmafx/gl_sync.c and sync_file.c folded into sync_object.c; SYNC_FILE fences borrowed (vmafx_fence_destroy() refuses them); dmabuf_import.cpp keeps master's driver_fd() (#2377) inside vmaf_sycl_dmabuf_import_queue(). Carried from rc4/integration: 44cbb8c (the SYCL EGL export folded onto vmafx/egl_export.c, VmafxEglTiled; the engine_leave signature) without its d3d11 stub, which the WP6 generator fix replaces, and the test of 869ec15 (an import closes only its own descriptors), whose library half master's #2377 already holds, and the SYCL parts of docs/api/vmafx/index.md that the integration merge dropped (restored there by 3b9a414), written for CUDA, SYCL and HIP. The rebase note is a fragment (ADR-2197). vmafx_sycl_producer.cpp is brought to 0 clang-tidy findings in the sycl lane, and core/src/compat/crt_portable.h's C spelling vmaf_tmpfile_portable(void), which the sycl lane reports through sycl/common.cpp, carries the ADR-1138 NOLINT that x86/avx512_warm_up.h has. Signed-off-by: Lusoris <lusoris@proton.me>
…GL textures (RC4 WP3, ADR-2091) (#2342) * feat(api): import SYCL device frames with event fences, dma-bufs and GL textures (RC4 WP3, ADR-2091) The VMAFx API now scores frames that already live on a SYCL device without a copy through the host, and they score bit for bit as the same frames uploaded from the host for every SYCL twin declared exact. - SYCL devices are the Level Zero GPUs, opened by index or in the context of the caller's queue; each has one in-order library queue with immediate command lists, the precaution of draft PR #2217 (evaluated: its i915 defect did not reproduce on xe). - USM planes of any address and pitch and linear dma-bufs are bound; Intel Y-tiled and Tile4 dma-bufs are de-tiled on the device with the address math the VA-surface import now shares (core/src/sycl/detile.h); NV12 / P010 / P016 are planarised on the device. GL textures are imported through an EGL dma-buf export. - Every SYCL twin that stages its own planes, the chroma twins included, copies a device picture on the device through vmaf_sycl_picture_read_plane(), which records the read on the frame. - Acquire and release are joins (an empty kernel with depends_on()), not barriers: a barrier on the event of a command behind a host task waits on the host (Research-2159). SYCL_EVENT acquire fences are waited on by the device; HOST, SYNC_FILE, GL_SYNC and a dma-buf's implicit write fences are checked on the host for the import rule's retry. HOST and SYCL_EVENT release fences and the release callback follow the last reader in every context. - The sync_file, dma-buf fence and GL sync helpers move to core/src/vmafx/sync_object.c, shared with the CUDA lane. - Windows shared textures and SYNC_FILE release fences are refused, tracked as T-SYCL-VMAFX-IMPORT-REMAINDER-2026-10-06. Landing (Q-083): the branch's diff (15bd091..5dafa1c) squashed onto (a082db7): the HIP lane's vmafx/gl_sync.c and sync_file.c folded into sync_object.c; SYNC_FILE fences borrowed (vmafx_fence_destroy() refuses them); dmabuf_import.cpp keeps master's driver_fd() (#2377) inside vmaf_sycl_dmabuf_import_queue(). Carried from rc4/integration: 44cbb8c (the SYCL EGL export folded onto vmafx/egl_export.c, VmafxEglTiled; the engine_leave signature) without its d3d11 stub, which the WP6 generator fix replaces, and the test of 869ec15 (an import closes only its own descriptors), whose library half master's #2377 already holds, and the SYCL parts of docs/api/vmafx/index.md that the integration merge dropped (restored there by 3b9a414), written for CUDA, SYCL and HIP. The rebase note is a fragment (ADR-2197). vmafx_sycl_producer.cpp is brought to 0 clang-tidy findings in the sycl lane, and core/src/compat/crt_portable.h's C spelling vmaf_tmpfile_portable(void), which the sycl lane reports through sycl/common.cpp, carries the ADR-1138 NOLINT that x86/avx512_warm_up.h has. * fix(api): destroy a SYNC_FILE fence by closing its descriptor, as ADR-2091 item 6 decides The WP3 lanes met with two rules for a SYNC_FILE fence: the SYCL lane (ADR-2091 item 6) closes the descriptor in vmafx_fence_destroy(), the HIP lane refused the call. rc4/integration kept the refusal when it merged the lanes and restored item 6 only with the Vulkan import (#2375). This PR carries item 6 itself, so no landed state contradicts its own ADR: vmafx_fence_destroy() closes a SYNC_FILE fence's descriptor in every build (refused on Windows, which has no sync_files) for every lane, and still refuses a GL_SYNC fence (the producer's, glDeleteSync()). Test: test_vmafx_fence_kinds test_sync_file_states checks that destroy closes the descriptor; it fails on the refusal. test_vmafx_import_sycl's sync_file case destroys its acquire fence instead of closing it itself. * fix(test): key the one-EGL-export contract by POSIX paths so it holds on Windows test_vmafx_one_egl_export_contract.py keyed the library sources by str(path.relative_to(SRC)), which is "vmafx\sync_object.c" on Windows, so test_sync_objects_defined_once compared it with "vmafx/sync_object.c" and failed there. The keys are now relative_to(SRC).as_posix(), as the other contract tests do (#1850). scripts/dev/preflight.sh --stage msvcism, which refuses that pattern, failed on the branch and passes now. * ci(tidy): measure the SYCL device-frame import's units in the cpu, sycl, cuda and hip lanes On master after #2367: every lane's baseline is master's, and the units this pull request adds or changes were measured again in the dev container (scripts/dev/tidy-lane.sh --write --only ..., clang-tidy 22.1.8), 0 findings and 0 uncited NOLINT in each lane: - cpu: libvmaf.c, the shared core/src/vmafx/ units (device, device_context, egl_export, fence, frame_import, frame_import_admit, frame_pool, submit, sync_object) and test_vmafx_fence_kinds.c; - sycl: the shared units, the SYCL import sources and runtime, the SYCL feature twins this pull request touches and the import tests; - cuda: the shared units and core/src/cuda/import_gl.c; - hip: the shared units, core/src/hip/import_frame.c and import_gl.c. core/src/vmafx/gl_sync.c and sync_file.c, folded into sync_object.c, leave the cuda and hip lanes: a scoped write drops the entries of deleted files since #2636, so the whole-lane measurements the earlier record needed are gone. This replaces the pull request's two earlier tidy commits. * fix(build): compile the SYCL import tests' producer with the SYCL argument list Master's test_sycl_math_constants_contract (#2654) requires every icpx compile command to take sycl_common_args or sycl_feature_tail_args, which carry the Windows math-constant define. The vmafx_sycl_producer.cpp command of this pull request spelled out the parts of sycl_common_args without that define and failed the contract on the rebased tree. It now uses sycl_common_args itself; the contract test passes again. Signed-off-by: Lusoris <lusoris@proton.me>
|
Thank you for this hardening work. #2342 has landed, so here is the review on its own merits, run on 2026-10-10 at your head What works (measured)
What we could not confirm ADR-1763's dropped import did not show on our host. I removed the immediate property from the primary queue locally and re-ran the nine stage-1 cases on both clips with What the project's rules need before it can land
Our offer You do not have to do this rework. We can do it on top of your 11 commits, with your authorship kept, and show you the result before it lands:
Is that fine with you? Say if you would rather do some or all of it yourself. What only you can answer or do
|
Summary
Depends on the hybrid-toolchain PR (
build(container): GCC for CPU C, icpx for SYCL, and a Containerfile.vmafx ..., headTualua:pr/hybrid-toolchain-containerfile). Review and merge it first; this branch is stacked on it. Until it merges, the diff here also shows its 24 commits; this PR's own commits are the last 11. It is stacked rather than based onmasterbecause it closesT-SYCL-ORDERED-SUM-MALLOC-PERTURB-2026-10-03, thedocs/state.mdrow that PR opens, and its new tests follow that PR'stest_link_kwargsrule.SYCL zero-copy hardening for the luma-only path of ADR-1688, ported from the Tualua fork without its chroma import (that is post-1.0, ADR-1685, and comes as a separate draft):
UR_L0_USE_IMMEDIATE_COMMANDLISTS=0) an Arc A380 silently dropped the per-frame VA import from a random frame on; zero-copycambidiffered from host upload in 3 to 7 of 10 runs. Only the primary queue's mode matters (an import-only immediate queue still fails 8 of 12).vmaf_sycl_upload_plane()is synchronous: it waits on the copy queue before it returns, so the D3D11 import may unmap its staging texture at once and the plane is in place before the next read. ClosesT-SYCL-UPLOAD-PLANE-NO-COMPUTE-FENCE-2026-10-05.test_sycl_ordered_sumprobe keeps its host copy sources alive until the wait (it read freed memory under deferred command lists; fails onmasterundermeson test'sMALLOC_PERTURB_).feature=names to SYCL twins throughvmaf_feature_backend_twin(), as the CLI does: on QSV zero-copy input a feature without a usable twin fails at configuration naming the reason; on software input the filter warns and uses the CPU. QSV zero-copy accepts NV12 / P010 only.[vmaf-sycl] timing:follows the log level (vmaf_logat INFO instead offprintf(stderr));libvmaf_syclno longer prints it at-loglevel error.scripts/test/zerocopy-e2e.sh, comparator,sycl-dev-container.sh) andmake sycl-zerocopy-contractin the FFmpeg Patch Stack job and pre-commit. Stage 1 is ADR-1688's admitted set and must equal host upload exactly; stage-2/3 cases need chroma and must be refused by name.-frames:v/-t: not a fork bug. FFmpeg n9.0.2 applies output limits after the filtergraph, so anylibvmaf*filter (CPU included) can score one more pair (-frames:v 20: 7 of 12 runs withlibvmaf, 4 of 12 withlibvmaf_sycl;-t 1: 12 of 12 both). Documented withtrim=end_frame=Non both inputs (Score exactly N frames);docs/state.mdConfirmed not-affected row; the zero-copy docs example now trims.What the fork had that this PR leaves to master (Research-1763 has the table): the CPU-extractor pre-pass and per-extractor
-ENOTSUPguards (ADR-1688's admission check covers both), and the fatal VA import (#2110 decides it: retry, then fail naming the frame). Patch 0005's import-failure branch is unchanged here.ADR numbers: the fork's ADR-1596 becomes ADR-1763; ADR-1764 holds what is left of the fork's ADR-1595. Both are above the numbers
VMAFx/vmafxbranches claim (up to 1762).Type
fix— bug fixtest— test-onlysycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. (pre-commit over the branch diff; clang-tidy oncore/src/sycl/common.cppandtest_sycl_ordered_sum_probe.cppreports no finding thatmasterdoes not have; the 17 old ones incommon.cppare the subject of refactor(sycl): bring the SYCL host file, psnr_hvs host and SYCL tests to the lint and HISS standard (ADR-1142) #2092.)--suite fastand--suite syclon the Arc A380 (below)./cross-backend-diffand the worst ULP is ≤ 2. (No kernel arithmetic changed; the e2e harness compares host upload with the CPU and zero-copy with host upload at%.17g: 0 ULP.).c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md). (None added; new scripts carry SPDX headers.)!orBREAKING CHANGE:and the migration path is documented below. (Not breaking:vmaf_sycl_upload_plane()'s contract only gets stricter; the timing line keeps its text with thelibvmaf INFOprefix.)docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt.Bug-status hygiene (ADR-0165)
docs/state.mdupdated: Recently closedT-SYCL-UPLOAD-PLANE-NO-COMPUTE-FENCE-2026-10-05(moved from Open, removed from the RC3 list),T-SYCL-ORDERED-SUM-MALLOC-PERTURB-2026-10-03(moved from Open),T-SYCL-ZEROCOPY-IMPORT-DROPPED-2026-10-02,T-SYCL-TIMING-LINE-IGNORES-LOGLEVEL-2026-10-05,T-SYCL-FILTER-FEATURE-NAMES-NOT-ROUTED-2026-10-02; Confirmed not-affectedT-FFMPEG-OUTPUT-LIMIT-EXTRA-SCORED-FRAME-2026-10-05(with the evidence).Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
Deep-dive deliverables (ADR-0108)
docs/research/1763-sycl-zerocopy-hardening.md(reconciliation with ADR-1688 / fix(ffmpeg): retry a failed libvmaf_sycl VA import, then stop naming the frame, and import each input with its own VA display #2110, command-list measurements, the frame-count evidence table, validation).## Alternatives considered; the frame-count options in Research-1763.AGENTS.mdinvariant note —core/src/sycl/AGENTS.md(immediate primary queue, synchronousupload_plane, timing throughvmaf_log),core/src/AGENTS.d/logging-and-diagnostics.md.changelog.d/fixed/sycl-zerocopy-import-immediate-cmdlist.md,sycl-upload-plane-synchronous.md,sycl-timing-line-follows-log-level.md,sycl-filter-feature-twin-routing.md,changelog.d/changed/ffmpeg-score-exactly-n-frames.md.docs/rebase-notes.md: "ADR-1763 — SYCL primary queue on immediate command lists" and "SYCL zero-copy hardening: twin routing, synchronous upload, timing log, e2e harness".User docs:
docs/usage/ffmpeg.md(twin routing and its messages; Score exactly N frames),docs/backends/sycl/zero-copy.md,docs/backends/sycl/bundling.md,docs/api/gpu.md,docs/development/sycl-zerocopy-testing.md. FFmpeg patch 0005 changes in this PR (rule 14); the series replays onn9.0.2, 20 of 20.Reproducer
Results (Arc A380, i915/iHD,
localhost/vmafx:build-ocloc, fresh cache, measured on base52e265fc0with PR 1 underneath; since rebased onto2f95aa87f, whose 8 newer commits change CUDA/HIP ADM, Python, arm64 SIMD, the vendored interop and CPU SpEED registration, not the SYCL host code or patch 0005):--suite sycl69 OK / 0 fail;--suite fast397 OK / 0 fail / 1 skipped (test_sycl_ordered_sumpasses; it fails onmaster).--repeat 3: 8 bit and 10 bit eachpass=50 fail=0 nonexact=0.make sycl-zerocopy-contract: routing check passes, comparator 26 of 26.n9.0.2: 20 of 20 withgit am --3way; FFmpeg builds with--enable-libvmaf-sycl.feature=name=psnron QSV →feature 'psnr' -> psnr_sycl, refused at frame 0 namingpsnr_sycl;feature=name=niqe→ configuration error on QSV, CPU fallback with a warning on software input;[vmaf-sycl] timingabsent at-loglevel error, present atinfo;trim=end_frame=20on both QSV inputs → 20 frames.make docs-fragments-check, ADR link / numbering / source-citation checks,check-state-md-rows, assertion density, the mypy gate, andstandardsctl audit --base VMAFx/master(43 within 43 baselined on9aee0bc9e, touched files clean, no baseline growth) pass.Known follow-ups
float_motionmotion_add_uvon zero-copy (post-1.0, ADR-1685): prepared as a separate draft on top of this branch.core/src/sycl/common.cpp.🤖 Generated with Claude Code