Skip to content

fix(sycl): emit motion3 from float_motion_sycl and pin the SYCL twin-parity cases - #1914

Merged
lusoris merged 1 commit into
masterfrom
fix/sycl-float-motion3-and-option-tests
Oct 3, 2026
Merged

lusoris merged 1 commit into
masterfrom
fix/sycl-float-motion3-and-option-tests

Conversation

@lusoris

@lusoris lusoris commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

float_motion_sycl now emits the CPU's motion3 and takes motion_blend_factor / motion_blend_offset, which closes the SYCL part of T-GPU-FLOAT-MOTION3-MISSING-2026-09-30. Before this, a --backend sycl --feature float_motion run wrote no motion3 and gave no warning, and a request with a blend option ran on the CPU. The parity gate's float_motion cell now compares motion3 too, so a CUDA, SYCL or HIP twin that loses it fails the cell. test_sycl_twin_option_parity gains the motion3 cases and the #1645 regression cases it never had, which is the SYCL test-only part of T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30.

Base: master 9aa990455, one commit.

What changed

  • core/src/feature/sycl/float_motion_sycl.cpp changes host code only. No kernel is touched, so no kernel changed. It provides VMAF_feature_motion3_score and declares motion_blend_factor (mbf) and motion_blend_offset (mbo) exactly as the CPU table does. motion3 is motion_blend_clip() of motion2, as in float_motion.c, float_motion_cuda.c (fix(cuda): match the CPU motion order, option tables and tiny-frame guards #1637) and float_motion_hip.c (ADR-1404):

    • frame 0 takes the first SAD;
    • the tail comes from flush();
    • a one-frame run gets 0;
    • motion_force_zero publishes 0.

    collect() is split into motion_emit_first() and motion_emit() so it stays inside HISS-04. motion_five_frame_window is an option of the integer motion / motion_v2 extractors only; the CPU float_motion has none, and neither does the twin.

  • FEATURE_METRICS["float_motion"] lists motion3 in scripts/ci/cross_backend_parity_gate.py and scripts/ci/cross_backend_vif_diff.py.

  • core/test/test_sycl_twin_option_parity.c gets these cases, mirroring test_cuda_twin_option_parity.c and the motion_v2 cases of test_hip_twin_option_parity.c. The float-motion and integer-motion cases compare with ==.

    • motion3: defaults on 8-bit 4:2:0 and 10-bit 4:2:2; a blend at the midpoint of the fixture's motion, with a check that the blend moves the CPU's motion3; all four score options together; motion_fps_weight=2:motion_max_val=4; one frame; motion3_force_0; option-table rows for motion_blend_factor / mbf / motion_blend_offset / mbo.
    • fix(sycl): RC3 parity follow-ups #1645 regressions: 64x64 flat identical frames for float_ssim (8 and 10 bit, enable_db with and without clip_db) and ssim (+inf); a 1x1 10-bit ssim frame; apsnr_* with --subsample 2; motion_v2 with motion_fps_weight, with a clipping motion_max_val and with both; a one-frame motion_v2 run, with and without the options.
    • Option tables: test_twin_options_are_cpu_options fails when a SYCL twin declares an option that its CPU extractor lacks or declares differently.
  • Docs: docs/metrics/motion.md, docs/backends/sycl/overview.md (new section), docs/development/cross-backend-gate.md, docs/development/rebase-sensitive-invariants.md (the float_motion_sycl entry; it also fixes the stale sub-group size, now 16 per ADR-1468), and core/src/feature/sycl/AGENTS.d/float-motion.md with its index row.

Verification

All on ryzen-4090-arc, with each device run under its lock.

Check Result
Arc A380, --precision max, --backend cpu against --backend sycl --feature float_motion=... (feature_backends names float_motion_sycl): Netflix 576x324 and both 1080p checkerboards. Each with the defaults, motion_blend_factor=0.5:motion_blend_offset=2, motion_fps_weight=2:motion_max_val=4, motion_fps_weight=2:motion_blend_factor=0.25:motion_blend_offset=3:motion_max_val=5 and motion_force_zero every motion / motion2 / motion3 value and every pooled value identical: 48 of 48 and 3 of 3
The same five option sets, one frame (--frame_cnt 1) 1 of 1, motion3 = 0
BBB 3840x2160, defaults and blend 200 of 200, all three scores
Gate float_motion cell, now with motion3, exact, Netflix pair and both checkerboards CPU↔SYCL 0, CPU↔CUDA (RTX 4090) 0, CPU↔HIP (gfx1036) 0
Same cell against the master SYCL twin ERROR (missing metrics: ... sycl lacks ['motion3']): the gate now sees the gap
test_sycl_twin_option_parity 22 of 22 pass on the A380
Each new case with its fix undone in a scratch copy: the old float_motion_sycl.cpp; #1645's changes undone (identical windows forced to 1 in both SSIM kernels, no TEMPORAL flag on psnr_sycl, raw SAD stored, no flush() output below two frames); a planted declaration mismatch (mbo default 41) all fail. Values: float_ssim inf against 72.247198959355487, ssim inf against 156.53559774527022, apsnr_y 40.764497719405341 against 40.770127717829908, SAD mfw_2 22.34 against 44.68; motion3 and the one-frame scores missing; test_twin_options_are_cpu_options names float_motion_sycl.motion_blend_offset. The flat-frame ssim +inf sub-case passes either way, because the CPU's value is +inf too
SYCL build (icx, JIT): device-free fast suite 261 pass
SYCL build: every gpu-suite test alone under the A380 lock 67 pass. test_cuda_parity_gate_default_run is not run here: this build has no CUDA
VMAF_SYCL_AOT_JOBS=8 meson test --suite sycl-aot OK: 35 SYCL translation units for the 19 default targets, 104 s
CUDA + HIP build (GCC): test_cuda_parity_gate_default_run; full default gate CPU↔CUDA pass; 24 of 24 cells OK
scripts/ci/test_cross_backend_parity_gate.py, test_parity_gate_metric_names.py, test_parity_gate_covers_registered_twins.py, test_cross_backend_feature_names.py 101 passed
test_sycl_kernel_source_contract.py (planted regressions), test_sycl_kernel_scratch, test_sycl_sub_group_size_contract, test_feature_extractor pass
praetorctl audit PASS, 14 touched files clean
scripts/dev/preflight.sh --stage msvcism, test_win32_pthread_shim_contract.py pass
praetorctl caveman check on AGENTS.d/float-motion.md PASS at 0.7 articles per 100 words (master's page failed at 2.2)
clang-tidy sycl lane in the dev container (scripts/dev/tidy-lane.sh sycl) neither touched file has a finding; the lane exits 2 on a master finding outside this PR, core/src/feature/motion_window.h:36 modernize-use-using (155 against baseline 154). It comes from the C++ includers integer_motion{,_v2}_sycl.cpp that #1893 added and is fixed in #1915 (with it applied the lane matches its baseline, 154, exit 0)

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. (clang-format, black, ruff and markdownlint ran through the commit hooks; praetorctl audit PASS; the clang-tidy sycl lane is in the table above.)
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. (SYCL build: 261 device-free tests and 67 device tests.)
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. (Gate cells exact, 0 on CUDA, SYCL and HIP.)
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. (CUDA and HIP already emit motion3; Metal is listed.)
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md). (No new source file.)
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. (Not breaking: the twin gains an output and two options the CPU has.)
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt. (No ADR: this applies ADR-0219 / ADR-1404's existing decision to the SYCL twin.)

Bug-status hygiene (ADR-0165)

  • docs/state.md is updated. T-GPU-FLOAT-MOTION3-MISSING-2026-09-30: SYCL part fixed with the A380 evidence; it stays open for Metal. T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30: the SYCL regression cases have landed, with the fail-without-fix evidence; it stays open for Metal.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests. (No CPU extractor and no Python test changed.)
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. (None changes.)

Cross-backend numerical results

float_motion (motion, motion2, motion3)  cpu-vs-cuda 0  cpu-vs-sycl 0  cpu-vs-hip 0   (Netflix 48 frames, checkerboards 3 + 3)
float_motion motion3, option sets        cpu-vs-sycl 0 on Netflix, checkerboards, one frame; BBB 4K 200 frames (defaults, blend)

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: trivial port of the CUDA / HIP host derivation; the measurements are in the state rows and the SYCL overview.
  • Decision matrix — no alternatives: only-one-way fix (the CPU extractor defines motion3; the twin reproduces it).
  • AGENTS.md invariant note — core/src/feature/sycl/AGENTS.d/float-motion.md (index row regenerated), plus the float_motion_sycl entry in docs/development/rebase-sensitive-invariants.md.
  • Reproducer / smoke-test command — pasted below under "Reproducer".
  • CHANGELOG fragment — changelog.d/fixed/sycl-float-motion-motion3.md.
  • Rebase note — docs/rebase-notes.md, "float_motion_sycl emits motion3 (2026-10-03)".

Reproducer

# On master the receipt names float_motion_sycl but the frames carry no motion3,
# and with a blend option the receipt names the CPU extractor.
build/tools/vmaf --backend sycl \
    --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
    --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
    --width 576 --height 324 --pixel_format 420 --bitdepth 8 --no_prediction \
    --feature float_motion=motion_blend_factor=0.5:motion_blend_offset=2 \
    --precision max --json --output /dev/stdout

python3 scripts/ci/cross_backend_parity_gate.py --vmaf-binary build/tools/vmaf \
    --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
    --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
    --width 576 --height 324 --features float_motion --backends cpu sycl

python3 scripts/ci/run_meson_test.py -- -C build test_sycl_twin_option_parity

Known follow-ups

  • refactor(motion): keep motion_window.h's typedef inside the SYCL clang-tidy baseline #1915 (separate, independent): the motion_window.h modernize-use-using finding that keeps the master sycl tidy lane one above its baseline.
  • Metal: float_motion_metal still emits no motion3, and the Metal twins still have the four items of T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30. Not run: they need an Apple device.
  • float_motion_sycl still does not declare motion_add_scale1, motion_add_uv or motion_filter_size. A request that sets one of them runs on the CPU (ADR-1183), as before; float_motion_hip has them (ADR-1404).

@github-actions github-actions Bot added the type:bug Something isn't working label Oct 3, 2026
@lusoris
lusoris force-pushed the fix/sycl-float-motion3-and-option-tests branch from 68e1c4d to e428c64 Compare October 3, 2026 11:37
…parity cases (#1914)

* fix(sycl): emit motion3 from float_motion_sycl and pin the SYCL twin-parity cases

float_motion_sycl provided motion and motion2 only. A `--backend sycl
--feature float_motion` run wrote no motion3, with no warning, and a request
with motion_blend_factor or motion_blend_offset fell back to the CPU. The
twin now provides VMAF_feature_motion3_score and declares both blend
options, as CUDA (#1637) and HIP (ADR-1404) do. motion3 is the CPU's
motion_blend_clip() of motion2: frame 0 from the first SAD, the tail from
flush(), 0 for a one-frame run, 0 under motion_force_zero. Only host code
changed; no kernel was touched.

FEATURE_METRICS["float_motion"] in both cross-backend gates now lists
motion3, so a twin that loses it fails the cell.

test_sycl_twin_option_parity gains the motion3 cases, plus the #1645
regression cases it never had: flat identical frames, a single-pixel frame,
apsnr with --subsample 2, and motion_v2 weight, cap and one frame. It also
gains test_twin_options_are_cpu_options.

Arc A380, --precision max, CPU against SYCL: every output is identical on
the Netflix pair, both 1080p checkerboards and BBB 4K, with and without the
options, and for one frame. The gate's float_motion cell is 0 for CUDA, SYCL
and HIP.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant