Skip to content

fix(sycl): make motion_sycl bit-exact with the CPU motion - #1628

Merged
lusoris merged 2 commits into
masterfrom
fix/sycl-motion-tiny-frame-parity
Sep 29, 2026
Merged

lusoris merged 2 commits into
masterfrom
fix/sycl-motion-tiny-frame-parity

Conversation

@lusoris

@lusoris lusoris commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

motion_sycl now matches the CPU motion bit for bit at every frame size. This closes T-SYCL-MOTION-TINY-FRAME-PARITY-2026-09-29.

The cause was the order of blurring and differencing, not edge handling. Since the port of Netflix a4a1492d (PR #532), the CPU integer_motion.c sums the absolute value of blur(prev - cur), rounding after each filter pass. motion_sycl still blurred each frame and differenced the blurred frames. The two orders round differently, and the error averages out over large frames but not small ones. The reflection, the taps and the normalisation were already correct.

motion_sycl and motion_v2_sycl now share one diff-first SAD kernel in integer_motion_pipeline_sycl.cpp. motion_v2_sycl was already exact, so its output does not change.

The same PR removes the host wait inside motion_sycl's submit() when motion_add_uv=true (T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29). U and V are now staged in pinned memory and uploaded on the combined queue, so each frame keeps a single wait, in collect(). PR #1627 opens this row as open. Whichever of the two PRs merges second should keep only the closed row added here.

The decision is recorded in ADR-1371, which supersedes ADR-1034's primary-queue wait, and the measurements are in Research-1371.

Type

  • fix — bug fix
  • perf — performance improvement
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. I did not run the full targets because this Windows host has no make. I ran the individual checks instead: clang-format 23.1.1, black and ruff on the touched files; clang-tidy-22 (through scripts/ci/clang-tidy-sycl.sh for the SYCL TUs), which is file-clean on all touched TUs and tests; markdownlint-cli2 0.23.3 on the touched docs; make docs-fragments-check; check-state-md-rows.sh; check-source-adr-citations.py; assertion-density.sh.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. I did not run the whole suite. I ran the affected SYCL tests on both GPUs (list below).
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. Motion and motion_v2 are 0 against the scalar CPU (table below).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR. Two rows move to Recently closed: T-SYCL-MOTION-TINY-FRAME-PARITY-2026-09-29 and T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29. Five RC3 rows are opened: T-CUDA-MOTION-BLUR-THEN-DIFF-2026-09-29, T-HIP-MOTION-BLUR-THEN-DIFF-2026-09-29, T-METAL-MOTION-BLUR-THEN-DIFF-2026-09-29, T-SYCL-SHARED-FRAME-GEOMETRY-REUSE-2026-09-29 and T-SYCL-MOTION-ADD-UV-CHROMA-GEOMETRY-2026-09-29.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. (Not applicable: no golden value changes, and CPU scores are untouched.)

Cross-backend numerical results

Worst per-frame difference in integer_motion2 between SYCL and the scalar CPU (--cpumask 4294967295). integer_motion3 is the same or smaller. The Arc B580 (level_zero:0) and the UHD 770 (level_zero:1) gave identical numbers. The synthetic inputs are 6 frames of ramp plus noise in 4:4:4.

Fixture 8-bit before 10-bit before After (both depths, both GPUs)
17x17 2.03e-4 1.76e-4 0
19x23 1.70e-4 1.97e-4 0
33x33 1.33e-4 8.61e-5 0
64x64 4.86e-5 3.15e-5 0
257x145 1.52e-5 3.75e-5 0
576x324 synthetic 6.89e-6 3.83e-6 0
Netflix 576x324, 48 frames 1.26e-5 — 0
BBB 3840x2160, 24 frames 5.56e-6 — 0
  • The "after" column also holds with the combined graph forced on (VMAF_SYCL_USE_GRAPH=1) and forced off (VMAF_SYCL_NO_GRAPH=1).
  • motion_v2_sycl is 0 at every fixture, both before and after.
  • float_motion_sycl is unchanged. Its differences are float rounding that grows with the frame size (5.6e-7 at 17x17, 2.7e-5 at 4K), not the edge pattern.
  • The default-model VMAF score at 4K moves by at most 4e-6.
  • The CPU SIMD and scalar paths agree exactly on every fixture.

Performance

Motion kernel at 4K. The kernel now reads two raw planes instead of one blurred plane. Figures are from a scratch micro-benchmark with profiling events, over 50 runs.

GPU Before After (kernel + copy)
Arc B580 0.484 ms 0.525 + 0.013 ms
UHD 770 6.15 ms 6.65 + 0.17 ms

Both GPUs pay about 11% more for the motion step. I also tried storing the frame from inside the kernel instead of copying it: 0.517 ms and 6.88 ms, no better.

Default model at 4K. The ms per frame of three 60-frame runs is unchanged within run-to-run noise:

  • B580: 52.8 / 53.4 / 52.8 before, 53.2 / 52.5 / 57.3 after.
  • UHD 770: 90.5 / 84.6 / 78.6 before, 92.2 / 81.8 / 81.9 after.

motion_sycl=motion_add_uv=true, BBB 4K. 39 frames fed from memory through the C API, median of 4 runs, from the SYCL timing summary. The middle column is the new kernel with the old wait still in place, which isolates the effect of removing the wait.

GPU Host time per frame Frame time: before / new kernel + wait / after
Arc B580 0.89 → 0.55 ms 8.58 / 8.71 / 7.95 ms
UHD 770 5.4 → 0.6 ms 23.0 / 25.0 / 21.8 ms

Deep-dive deliverables (ADR-0108)

  • Research digest — docs/research/1371-sycl-motion-diff-first-pipeline.md.
  • Decision matrix — in ADR-1371 ## Alternatives considered.
  • AGENTS.md invariant note — core/src/feature/sycl/AGENTS.md (the motion SAD pipeline invariant, the rewritten motion_add_uv queue-sync invariant and the tile-loader list); core/src/feature/metal/AGENTS.md (stale SYCL symbol reference).
  • Reproducer / smoke-test command — see "Reproducer" below.
  • CHANGELOG fragment — changelog.d/fixed/sycl-motion-tiny-frame-parity.md.
  • Rebase note — docs/rebase-notes.md, "ADR-1371 — SYCL motion differences the frames before the blur, in one shared kernel".

Reproducer

# In a SYCL build (-Denable_sycl=true), on each device:
ONEAPI_DEVICE_SELECTOR=level_zero:0 build/test/test_sycl_motion_tiny_frames
ONEAPI_DEVICE_SELECTOR=level_zero:1 build/test/test_sycl_motion_tiny_frames
python3 core/test/test_sycl_kernel_source_contract.py

# CLI check on the Netflix pair (prints 0.0; before this PR 1.26e-05):
Y=python/test/resource/yuv
for b in cpu sycl; do
  vmaf -r $Y/src01_hrc00_576x324.yuv -d $Y/src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 \
    --no_prediction --feature motion --backend $b --precision=max --json -q -o /tmp/motion_$b.json
done
python3 -c "import json; a, b = (json.load(open(f'/tmp/motion_{x}.json'))['frames'] for x in ('cpu', 'sycl')); print(max(abs(p['metrics']['integer_motion2'] - q['metrics']['integer_motion2']) for p, q in zip(a, b)))"

I ran these tests on the Arc B580 and the UHD 770, and all pass: test_sycl_motion_tiny_frames, test_sycl_motion_add_uv_parity (and _large), test_sycl_motion3_parity (and _large), test_sycl_motion_v2_parity (and _large), test_sycl_twin_option_parity, test_sycl_adm_tiny_frames, and test_sycl_kernel_source_contract.py. Built against the old integer_motion_sycl.cpp, test_sycl_motion_tiny_frames and the updated test_sycl_motion_add_uv_parity both fail.

Known follow-ups

  • The CUDA, HIP and Metal motion twins still blur each frame. Their rows in docs/state.md carry a copy-paste verify command for ryzen-4090-arc (CUDA and HIP) and for Apple Silicon (Metal). Their motion_v2 twins already difference first.
  • T-SYCL-SHARED-FRAME-GEOMETRY-REUSE-2026-09-29: a VmafSyclState imported into a second context with a different frame size keeps the first context's shared frame buffers. I found this while writing the test, which now opens one state per case.
  • T-SYCL-MOTION-ADD-UV-CHROMA-GEOMETRY-2026-09-29: motion_sycl's motion_add_uv sizes its chroma planes as 4:2:0 ((w + 1) / 2) whatever the input format. I found this by reading the source and did not run it; it is outside the scope of this PR.
  • make verify-all did not run: standardsctl / praetorctl and make are not installed on this Windows host.

🤖 Generated with Claude Code

lusoris and others added 2 commits September 29, 2026 16:26
motion_sycl now gives the CPU motion's scores bit for bit at every frame
size, which closes T-SYCL-MOTION-TINY-FRAME-PARITY-2026-09-29.

The cause was the order of blurring and differencing, not edge handling.
Since the port of Netflix a4a1492d (PR #532) the CPU sums the absolute value
of blur(prev - cur), rounding after each filter pass. motion_sycl still
blurred each frame and differenced the blurred frames, which rounds
differently: motion2 was up to 2.0e-4 off at 17x17, 1.3e-5 on the Netflix
576x324 pair and 5.6e-6 at 4K, identically on an Arc B580 and a UHD 770.

motion_sycl and motion_v2_sycl now run one diff-first SAD kernel in
integer_motion_pipeline_sycl.cpp. motion_v2_sycl was already exact and is
unchanged. The new test_sycl_motion_tiny_frames compares both twins with the
scalar CPU using == from 3x3 to 1283x723 at 8, 10 and 16 bits. It passes on
both GPUs and fails against the old twin. At 4K the motion step costs about
11% more device time.

With motion_add_uv, submit() no longer waits on the device. U and V are
staged in pinned memory and uploaded on the combined queue, which closes
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29 and supersedes ADR-1034's
primary-queue wait. Host time per 4K frame drops from 5.4 to 0.6 ms on the
UHD 770.

The CUDA, HIP and Metal motion twins keep the old order. This change opens
RC3 rows for them, and one for a SYCL state reused across frame sizes.

ADR-1371, Research-1371.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
motion_configure_chroma() sizes U and V as 4:2:0 for every pixel format,
so 4:2:2 and 4:4:4 input stage only part of each chroma plane. Found by
reading the source for the ADR-1371 staging change and recorded as the
open RC3 row T-SYCL-MOTION-ADD-UV-CHROMA-GEOMETRY-2026-09-29.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the type:bug Something isn't working label Sep 29, 2026
@lusoris
lusoris merged commit b20472d into master Sep 29, 2026
115 of 116 checks passed
@lusoris
lusoris deleted the fix/sycl-motion-tiny-frame-parity branch September 29, 2026 16:07
lusoris added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
gitleaks ran git log --all over a fetch-depth: 0 checkout, so a finding on
any branch, even a force-pushed-away commit still in the fetch, failed every
other pull request (#1628 on 2026-09-29). Pass --log-opts so each run scans
HEAD's history: master plus the PR's commits on a pull request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
The rebase onto #1628 left two _function_body definitions in the kernel
source contract; the second, static-only one shadowed the first and could
not find ssimulacra2_sycl's non-static collect_fex_sycl. Keep the
brace-matched helper for both callers (HISS-19). #1628 also closed
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29, so drop this branch's open copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…CPU (#1625)

* perf(sycl): compute ADM's AIM pass on the device, bit-exact with the CPU

adm_sycl now emits VMAF_integer_feature_aim_score and
VMAF_integer_feature_adm3_score, so under --backend sycl the default model
vmaf_v1.0.16_3d0h no longer scores ADM with the CPU extractor through the
feature-name fallback. A frame is still one upload, one graph replay and one
small copy back, with no host wait in between.

The decouple/CSF kernel stores the AIM neighbourhood band |csf(r)| / 30 next
to the DLM one, and the reduction becomes one work-group per row that computes
all three bands of the CSF denominator, the DLM contrast measure and the AIM
contrast measure, folding each row total once through
adm_cm_round_row_total(). aim and adm3 are finalised in the CPU's own float
arithmetic and equal --backend cpu on every frame of the Netflix pair, 50
frames of BBB 4K, 853x480 and 17x17 crops and 10-bit input, with the default
and the default model's options, on an Arc B580 and a UHD 770. adm2 and the
scale outputs keep their double finalisation.

The decouple quotient is now clamped in int64, as the CPU's tmp_k is. The old
kernel narrowed it to int32 first and wrapped at scales 1-3 once |t / o|
exceeded 2^16, which left integer_adm_scale2 up to 1.40e-6 from the CPU on 4K
content; every DLM output is now within 2.9e-7.

With the default model on the B580, a 4K frame drops from 48.4 to 9.2 ms at
--threads 0 and from 24.3 to 9.7 ms at --threads 16. On the UHD 770 it gets
slower (75.7 to 87.5 ms at --threads 0, 40.5 to 79.4 ms at --threads 16)
because the iGPU now does the ADM work the CPU used to do beside it.

ADR-1362. The HIP half of T-GPU-ADM-AIM-DEVICE-PASS-MISSING-SYCL-HIP-2026-09-05
stays open, with a verify command for the HIP port.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR tag pages after the rebase onto #1619

The rebase merged the by-tag pages of ADR-1359 and ADR-1362 line by line,
which left the fork-local, gpu and index counts stale; they are regenerated
from the fragments. The HIP verify note in docs/state.md no longer claims that
`--feature adm --backend hip` runs the CPU extractor, which ADR-1359 changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): make adm2 and the ADM scale outputs bit-exact with the CPU

The SYCL integer ADM twin finalised adm2, integer_adm_scale0..3 and the debug
num / den values in double, which left them up to 2.9e-7 from the CPU even
though the device accumulators match it exactly. They now go through the same
float finalisation as aim and adm3 (adm_scale_cpu / adm_terms / adm_finalise,
mirroring integer_compute_adm and adm_result_finalise), and the double
finaliser is removed.

Every ADM output now equals --backend cpu at --precision max on every frame of
the Netflix 576x324 pair and 50 frames of BBB 3840x2160, with the default and
the default model's options, on an Arc B580 and a UHD 770.
test_sycl_adm_parity and test_sycl_adm_tiny_frames compare every key bit for
bit; ADR-1362, the changelog fragment, docs/state.md and the backend guide
follow the maintainer's decision that bit-exact with the CPU is the contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes and citations after the rebase onto #1624

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes after the rebase onto #1628

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
* fix(sycl): register device images in Windows MSVC builds

The Windows MSVC+SYCL build never ran its kernels. Meson links MSVC
builds with link.exe, which ignored -fsycl and never wrapped or
registered the SYCL device images, so every kernel submit failed with
"No kernel named ... was found" and 47 of 50 SYCL tests failed on an
Arc B580. The CI leg only builds, so nothing had noticed.

MSVC builds now compile the SYCL translation units as relocatable
device code and run one `icpx -fsycl -fsycl-link` step over all SYCL
objects. A small COFF patcher gives the resulting object an external
anchor symbol that sycl/common.cpp asks for with /include, so link.exe
pulls it out of vmaf.lib into every program that uses SYCL, including
static consumers such as FFmpeg. Linux keeps the ADR-1360 per-TU code
generation. The CI leg now checks the kernel registry without a GPU.

On an Arc B580 and a UHD 770 all 51 SYCL tests and the whole suite
pass, the default model and vmaf_v0.6.1 agree with the CPU within the
gate on the Netflix pair and 50 frames of 4K, all 19 SYCL extractors
pass the parity gate, and the SYCL output is bit-identical to the
Linux container build of the same revision.

Running the tests natively exposed that scripts/ci/run_meson_test.py
returned 0 on Windows as soon as it started Meson (os.execvp has no
exec semantics there), so the MinGW64 and ARM64 MSVC lanes passed
after 17 to 19 tests. The runner now waits on Windows. Four Windows
test-harness failures it hid are fixed, and the parity gate reads the
cambi score by the name vmaf writes.

ADR-1364, Research-2125.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes and citations after the rebase onto #1624

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after the rebase onto #1628

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
gitleaks ran git log --all over a fetch-depth: 0 checkout, so a finding on
any branch, even a force-pushed-away commit still in the fetch, failed every
other pull request (#1628 on 2026-09-29). Pass --log-opts so each run scans
HEAD's history: master plus the PR's commits on a pull request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…1633)

* ci(security): scan only the checked-out history in the Gitleaks job

gitleaks ran git log --all over a fetch-depth: 0 checkout, so a finding on
any branch, even a force-pushed-away commit still in the fetch, failed every
other pull request (#1628 on 2026-09-29). Pass --log-opts so each run scans
HEAD's history: master plus the PR's commits on a pull request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: render the Gitleaks changelog fragment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
The rebase onto #1628 left two _function_body definitions in the kernel
source contract; the second, static-only one shadowed the first and could
not find ssimulacra2_sycl's non-static collect_fex_sycl. Keep the
brace-matched helper for both callers (HISS-19). #1628 also closed
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29, so drop this branch's open copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…ms_ssim (#1627)

* perf(sycl): run ssimulacra2 on the device and wait once per frame in ms_ssim

ssimulacra2_sycl no longer round-trips through the host inside a frame.
Before, it blurred on the device and did everything else on the CPU: at
each of six scales it computed XYB on the host, uploaded it, copied five
full-size buffers back, waited, and combined the SSIM and edge maps and
downsampled on the host, about 4 GB of copies per 4K frame. Now submit()
uploads the raw Y/U/V planes once and the device runs colour conversion,
XYB, the products and blurs, the per-channel sums and the downsample;
collect() waits once for an 864-byte block of sums.

Everything up to the sums is bit-identical to the CPU: the TU builds
with contraction off and the shared cube root divides correctly rounded.
The CPU adds the per-pixel fp64 terms one after another, which no
parallel reduction can replay and the device has no fp64 for, so the
terms are exact fp32 pairs summed in a fixed tree. The score moves from
bit-identical to within 6.7e-12 of the CPU, the same on every device.
On an Arc B580 a 4K frame takes 33 ms instead of 963.

float_ms_ssim_sycl gives every scale its own partials span, enqueues
the whole frame in submit() and waits once in collect(); its output is
unchanged bit for bit.

The correctly rounded division and fp32-pair helpers move from the
SpEED pipeline to sycl_exact_fp.h, shared by both TUs; their slow path
uses the oneAPI math extension in device code only, so the static test
links do not pull its host fallback. The SYCL tidy baseline for
ssimulacra2_sycl.cpp drops from 19 to 0. speed_gpu_parity.py takes
--feature and --max-abs-diff to check any twin. The CUDA, HIP and Metal
twins keep the host combine; each is an RC3 row in docs/state.md.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): send inputs ssimulacra2_sycl rejects to the CPU extractor

Since ADR-1359 `--backend sycl --feature ssimulacra2` maps to
ssimulacra2_sycl when the twin can run the input. The twin now declares
that through the ADR-1324 context check: 4:0:0 input (no chroma to
convert) and frames below 8x8, which its init rejects, fall back to the
CPU extractor instead of failing. The parity test covers the fallback
name and the 8x8 / 7x8 / 8x7 / 4:0:0 boundaries.

Docs no longer claim that a CPU feature name always runs on the CPU, the
state rows name the mapped path, and the ADR tag pages are regenerated
after the rebase onto #1619.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(adr): note the ssimulacra2_sycl context check in ADR-1363

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): build sycl_exact_fp.h and ssimulacra2_sycl under MSVC+SYCL

<windows.h>, which compat/win32/pthread.h pulls in first, defines min, max,
near and far as macros. They broke Intel's math.hpp declarations, the
round_quotient locals and two sycl::min calls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes and citations after the rebase onto #1624

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(sycl): keep one source-body helper and drop the row #1628 closed

The rebase onto #1628 left two _function_body definitions in the kernel
source contract; the second, static-only one shadowed the first and could
not find ssimulacra2_sycl's non-static collect_fex_sycl. Keep the
brace-matched helper for both callers (HISS-19). #1628 also closed
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29, so drop this branch's open copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(state): merge the RC2/RC3 disposition rows the rebase duplicated

Three-way by bug-id set against master b61b4d1: master's ids, plus the
ids this branch added since its base, minus the ones it removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant