Skip to content

fix(cli): run --feature on the explicit --backend's twin and report what ran - #1619

Merged
lusoris merged 3 commits into
masterfrom
fix/cli-feature-backend-twin
Sep 29, 2026
Merged

lusoris merged 3 commits into
masterfrom
fix/cli-feature-backend-twin

Conversation

@lusoris

@lusoris lusoris commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

With an explicit --backend cuda|sycl|hip|metal, vmaf --feature <cpu-name> now runs that backend's twin: vmaf --backend sycl --feature ciede runs ciede_sycl instead of the CPU extractor. When there is no twin, or the twin cannot honour an option or the input size, the CPU extractor runs and the CLI prints one warning line that names the feature and the reason. The JSON backend_used now reports the backend the extractors actually ran on, and a new feature_backends array lists each extractor's backend, so a mixed run says so.

Stacked on #1616 (perf/sycl-ciede-throughput), rebased onto its current head 39c7880e5. It must merge after #1616: it moves #1616's T-CLI-FEATURE-NAME-BYPASSES-GPU-BACKEND-2026-09-29 row to Recently closed and rewrites the #1616 CLI-guide paragraph and changelog bullet that describe the old behaviour. The base is master, so until #1616 merges the diff also shows its commits.

This implements the maintainer's popup decision of 2026-09-29, "CLI picks the twin (Recommended)" (ADR-1359).

How the twin is chosen

The CLI links the shared libvmaf, which hides its registry (-fvisibility=hidden, ADR-0379), so this PR adds two additive public functions to libvmaf.h:

  • vmaf_feature_backend_twin() pairs a CPU extractor with the imported backend's twin through the lookup model dispatch already uses (its provided features through vmaf_get_feature_extractor_by_feature_name(), restricted to the backend's flag, so no name table). It accepts the twin only if it honours every option (ADR-1183 / ADR-1316) and passes the ADR-1324 geometry check for the input. It registers nothing.
  • vmaf_registered_feature_extractor() lists the registered extractors and the backend each runs on. The CLI reads it after the final flush to build the receipt.

Unchanged: twin names such as --feature ciede_sycl (and their ADR-0543 exit 100), --backend cpu, --backend auto, no --backend, vmaf_use_feature(), and the FFmpeg filters. ffmpeg-patches/ needs no change: no patch references the new symbols or backend_used, and the filters keep registering by exact name.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. Run piecewise, see "Verification" below: clang-format 23.1.1, the scoped clang-tidy 22.1.8 CPU-lane ratchet, shellcheck, shfmt, markdownlint and make docs-fragments-check are clean on the touched files.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. The fast suite passes except four tests that fail identically on the base without this change (see "Verification").
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. No kernel changed; the mapped run is bit-identical to the explicit twin on all 48 Netflix frames.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. No extractor changed.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md).
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not breaking: the C API and the JSON schema only grow.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section. T-CLI-FEATURE-NAME-BYPASSES-GPU-BACKEND-2026-09-29 moves to Recently closed. T-SYCL-PSNR-HVS-B580-SIGSEGV-2026-09-29 opens under RC2 stabilisation: psnr_hvs_sycl crashes on the Arc B580 with or without this change, and --backend sycl --feature psnr_hvs now reaches it.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. Not applicable.

Cross-backend numerical results

No kernel changed. --feature ciede --backend sycl against --feature ciede_sycl --backend sycl on the B580:

ciede2000  Netflix 576x324, 48 frames, --precision max  mapped == explicit on every frame; max |sycl - cpu| = 1.18e-5
ciede2000  BBB 3840x2160, 22 frames                      mapped == explicit on every frame

Performance (if perf or feat)

Not a perf change, but the reported symptom was throughput. BBB 3840x2160 on the B580, wall time including start-up: 22 frames in 395 ms with --feature ciede --backend sycl (504 ms with the explicit twin name, same scores), against 787 ms per frame for serial CPU ciede on the same host while other builds ran. Research-2120 measured 1714 ms per frame for the old --feature ciede --backend sycl.

Deep-dive deliverables (ADR-0108)

  • Research digest — Research-2121: why the CLI needs a library query, why the per-feature lookup must be restricted to the backend flag and scan every provided feature, the measured SYCL pairing table for 25 CPU extractors, and the receipt design. Research-2120's open question now points to it.
  • Decision matrix — in ADR-1359 ## Alternatives considered: resolving in vmaf_use_feature(), fixing only the receipt, a CLI name table, exporting the internal registry, failing without a twin, a "mixed" value, and requiring full feature coverage.
  • AGENTS.md invariant note — core/tools/AGENTS.md "--feature twin routing and backend receipt (ADR-1359)": the pairing lives in libvmaf only, the ADR-0543 gate runs first, the warning prints before the options are consumed, and backend_used gains fields, never values.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/fixed/cli-feature-backend-twin.md and changelog.d/added/libvmaf-feature-backend-twin-api.md; CHANGELOG.md re-rendered.
  • Rebase note — docs/rebase-notes.md "fix/cli-feature-backend-twin — --feature runs the explicit --backend's twin (ADR-1359)".

Reproducer

# Needs a SYCL build and an Intel GPU; swap sycl for cuda / hip / metal.
vmaf --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
     --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
     --width 576 --height 324 --pixel_format 420 --bitdepth 8 \
     --no_prediction --backend sycl --feature ciede --feature brisque \
     --json --output /tmp/out.json
# stderr: vmaf: warning: --feature brisque: the sycl backend has no twin of this extractor; computing it on the CPU
# /tmp/out.json ends with:
#   "backend_used": "sycl", "feature_backends": [{"extractor": "ciede_sycl", "backend": "sycl"}, {"extractor": "brisque", "backend": "cpu"}]

# Device-free tests, and the device run (skips with 77 without a GPU):
python3 scripts/ci/run_meson_test.py -- -C build test_cli_feature_backend test_feature_backend_twin \
    test_vmaf_feature_backend_cpu test_vmaf_feature_backend_sycl

Verification

Run in vmaf-dev-mcp:local on the WSL2 workstation, B580 as level_zero:0:

  • SYCL build (icx, AOT bmg-g21,adl-s, -Db_lto=false; with LTO the tests that link only libvmaf.a fail to link on the base too, because GNU ar does not index LLVM bitcode): test_cli_feature_backend (14 cases, libvmaf replaced by fakes), test_feature_backend_twin (9 cases, white-box against the real SYCL registry), test_vmaf_feature_backend_cpu, test_vmaf_feature_backend_sycl, test_feature_collector, test_feature_extractor, test_context, test_cli_parse, test_sycl_ciede_parity and test_vmaf_sycl_threads pass.
  • Fast suite: 227 of 232 pass. test_sycl_psnr_hvs_parity and _large (SIGSEGV), test_sycl_adm_tiny_frames and test_gpu_picture_pool_uaf (timeout) fail the same way in a build with this PR's library diff reverted; test_meson_secret_env_sanitization failed once on container clock skew.
  • Netflix golden subset: vmafexec_test.py and vmafexec_feature_extractor_test.py, 141 passed against a gcc CPU build.
  • clang-tidy 22.1.8, CPU lane, scoped to libvmaf.c, feature_extractor.cpp, vmaf.cpp and the new files: 0 findings in each; libvmaf.h stays at its baseline of 7.
  • make docs-fragments-check in python:3.14-slim, check-source-adr-citations.py, check-state-md-rows.sh, standardsctl compile-context --verify, standardsctl hiss coverage --verify and make dedupe-check pass. standardsctl audit fails on this Windows host with 148 unbaselined findings, reported as "0 in touched files"; a Linux build of praetor's current source reports none in touched files.

Known follow-ups

  • T-SYCL-PSNR-HVS-B580-SIGSEGV-2026-09-29: psnr_hvs_sycl crashes on the B580 (the UHD 770 runs it). Until it is fixed, --backend sycl --feature psnr_hvs crashes on that GPU where it used to run on the CPU; --backend cpu avoids it.
  • CUDA, HIP and Metal use the same lookup and are covered by the device-free tests; test_vmaf_feature_backend_{cuda,hip,metal} have not been run on those GPUs.
  • The SYCL clang-tidy lane was not re-measured; the CPU lane covers every touched host TU.

🤖 Generated with Claude Code

no ffmpeg-patches update needed: the patches register extractors by exact name through vmaf_use_feature(), which is unchanged; no patch references vmaf_feature_backend_twin(), vmaf_registered_feature_extractor() or backend_used.

lusoris added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the type:bug Something isn't working label Sep 29, 2026
lusoris added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the fix/cli-feature-backend-twin branch from 9ea0339 to 7355e89 Compare September 29, 2026 10:11
Comment thread core/tools/test/test_cli_feature_backend.c Fixed

#include "test.h"
#include "mu_table.h"
#include "libvmaf.c" // NOLINT(bugprone-suspicious-include): white-box seam (ADR-0141).
lusoris added a commit that referenced this pull request Sep 29, 2026
…st fake

cppcheck 2.19 flagged two things on #1619:
- sink_puts() looped `i < max_text && text[i]`, so for a 14-byte literal
  cppcheck assumed an index of 4095. Bound the loop by strnlen() instead,
  which makes the NUL stop explicit.
- The vmaf_feature_backend_twin() fake in test_cli_feature_backend.c wrote
  *twin_name unconditionally; cross-TU analysis paired it with the real
  function's NULL-argument test. The fake now returns -EINVAL for a NULL
  result pointer, the real function's contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris and others added 3 commits September 29, 2026 13:22
…hat ran

`vmaf --backend sycl --feature ciede` now runs `ciede_sycl`. It used to
initialise the SYCL device, compute ciede2000 with the CPU extractor one
frame at a time, and report `"backend_used": "sycl"` (1714 ms per 4K
frame on an Arc B580, against 8.4 ms for the twin).

With an explicit `--backend cuda|sycl|hip|metal`, each `--feature`
that names a CPU extractor is resolved through a new libvmaf query,
`vmaf_feature_backend_twin()`. It pairs the extractor with a twin the
way model dispatch does (its provided features through
`vmaf_get_feature_extractor_by_feature_name()`, restricted to the
backend's flag) and accepts the twin only if it honours every option
(ADR-1183 / ADR-1316) and can run the input geometry (ADR-1324).
Otherwise the CPU extractor runs and the CLI prints one warning line
that names the feature and the reason. Twin names such as
`--feature ciede_sycl`, `--backend cpu`, `--backend auto`, no
`--backend`, `vmaf_use_feature()` and the FFmpeg filters are unchanged.

`backend_used` now reports the backend the registered extractors ran
on, read after the final flush through a second new query,
`vmaf_registered_feature_extractor()`, and a new `feature_backends`
array lists each extractor's backend, so a mixed run says so. The value
set of `backend_used` is unchanged.

On the B580 the mapped run matches `--feature ciede_sycl` on all 48
Netflix frames at --precision max. `--backend sycl --feature psnr_hvs`
now reaches `psnr_hvs_sycl`, which crashes on the B580 with or without
this change; that is filed as T-SYCL-PSNR-HVS-B580-SIGSEGV-2026-09-29.

Closes T-CLI-FEATURE-NAME-BYPASSES-GPU-BACKEND-2026-09-29. ADR-1359,
Research-2121.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…st fake

cppcheck 2.19 flagged two things on #1619:
- sink_puts() looped `i < max_text && text[i]`, so for a 14-byte literal
  cppcheck assumed an index of 4095. Bound the loop by strnlen() instead,
  which makes the NUL stop explicit.
- The vmaf_feature_backend_twin() fake in test_cli_feature_backend.c wrote
  *twin_name unconditionally; cross-TU analysis paired it with the real
  function's NULL-argument test. The fake now returns -EINVAL for a NULL
  result pointer, the real function's contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
twin_fake.calls++;
twin_fake.seen_name = feature_name;
twin_fake.seen_opts = opts_dict;
twin_fake.seen_pic_cfg = pic_cfg;
@lusoris
lusoris force-pushed the fix/cli-feature-backend-twin branch from 86fee9a to 81a3453 Compare September 29, 2026 12:15
@lusoris
lusoris merged commit 2d9d5b0 into master Sep 29, 2026
110 checks passed
@lusoris
lusoris deleted the fix/cli-feature-backend-twin branch September 29, 2026 12:56
lusoris added a commit that referenced this pull request Sep 29, 2026
…SYCL twins

The SYCL twins psnr_sycl, integer_ssim_sycl, float_ssim_sycl and
float_motion_sycl now take their CPU extractor's option tables
(T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26, SYCL part; ADR-1365).
Before, a model that set one of these options computed the feature on the
CPU, and naming the twin with the option failed with "unknown option".

- psnr_sycl: enable_mse, enable_apsnr, reduced_hbd_peak and min_sse,
  bit-exact with the CPU (apsnr_* aggregates via a new flush). The option
  math moves verbatim from integer_psnr.c into psnr_score.h, which the CPU
  extractor and the twin both call; CPU output is bit-identical and the
  non-slow Netflix golden gate passes.
- integer_ssim_sycl / float_ssim_sycl: enable_db and clip_db on the host
  (vmaf_ssim_max_db), enable_lcs as a second pass-2 kernel that reduces
  L, C and S per work-group. Both twins now score identical windows
  exactly 1, so identical frames report the CPU's +inf / clip_db ceiling;
  default linear scores move by at most 1.1e-8.
- float_motion_sycl: motion_max_val (alias mmxv) through the CPU's
  motion_clip() on every emitted score; motion_force_zero, declared but
  ignored before, now emits zeros; the debug motion score carries
  motion_fps_weight as on the CPU.

With #1619 (ADR-1359), `--backend sycl --feature psnr=enable_mse=true` and
the like now run on the twins; test_vmaf_feature_backend.sh's fallback case
moves from float_ssim enable_lcs (now honoured) to float_motion
motion_filter_size, which no twin implements, and the CLI guide's fallback
example follows.

test_sycl_twin_option_parity checks every option against the CPU
(positive, negative, boundary) and passes on an Arc B580 and a UHD 770;
test_sycl_init_unwind gains the enable_lcs allocation case. Docs, state
rows, research digest 2127 and rebase notes updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
…SYCL twins (#1624)

The SYCL twins psnr_sycl, integer_ssim_sycl, float_ssim_sycl and
float_motion_sycl now take their CPU extractor's option tables
(T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26, SYCL part; ADR-1365).
Before, a model that set one of these options computed the feature on the
CPU, and naming the twin with the option failed with "unknown option".

- psnr_sycl: enable_mse, enable_apsnr, reduced_hbd_peak and min_sse,
  bit-exact with the CPU (apsnr_* aggregates via a new flush). The option
  math moves verbatim from integer_psnr.c into psnr_score.h, which the CPU
  extractor and the twin both call; CPU output is bit-identical and the
  non-slow Netflix golden gate passes.
- integer_ssim_sycl / float_ssim_sycl: enable_db and clip_db on the host
  (vmaf_ssim_max_db), enable_lcs as a second pass-2 kernel that reduces
  L, C and S per work-group. Both twins now score identical windows
  exactly 1, so identical frames report the CPU's +inf / clip_db ceiling;
  default linear scores move by at most 1.1e-8.
- float_motion_sycl: motion_max_val (alias mmxv) through the CPU's
  motion_clip() on every emitted score; motion_force_zero, declared but
  ignored before, now emits zeros; the debug motion score carries
  motion_fps_weight as on the CPU.

With #1619 (ADR-1359), `--backend sycl --feature psnr=enable_mse=true` and
the like now run on the twins; test_vmaf_feature_backend.sh's fallback case
moves from float_ssim enable_lcs (now honoured) to float_motion
motion_filter_size, which no twin implements, and the CLI guide's fallback
example follows.

test_sycl_twin_option_parity checks every option against the CPU
(positive, negative, boundary) and passes on an Arc B580 and a UHD 770;
test_sycl_init_unwind gains the enable_lcs allocation case. Docs, state
rows, research digest 2127 and rebase notes updated.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
The rebase merged the by-tag pages of ADR-1359 and ADR-1362 line by line,
which left the fork-local, gpu and index counts stale; they are regenerated
from the fragments. The HIP verify note in docs/state.md no longer claims that
`--feature adm --backend hip` runs the CPU extractor, which ADR-1359 changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
Since ADR-1359 `--backend sycl --feature ssimulacra2` maps to
ssimulacra2_sycl when the twin can run the input. The twin now declares
that through the ADR-1324 context check: 4:0:0 input (no chroma to
convert) and frames below 8x8, which its init rejects, fall back to the
CPU extractor instead of failing. The parity test covers the fallback
name and the 8x8 / 7x8 / 8x7 / 4:0:0 boundaries.

Docs no longer claim that a CPU feature name always runs on the CPU, the
state rows name the mapped path, and the ADR tag pages are regenerated
after the rebase onto #1619.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
The rebase merged the by-tag pages of ADR-1359 and ADR-1362 line by line,
which left the fork-local, gpu and index counts stale; they are regenerated
from the fragments. The HIP verify note in docs/state.md no longer claims that
`--feature adm --backend hip` runs the CPU extractor, which ADR-1359 changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 29, 2026
Since ADR-1359 `--backend sycl --feature ssimulacra2` maps to
ssimulacra2_sycl when the twin can run the input. The twin now declares
that through the ADR-1324 context check: 4:0:0 input (no chroma to
convert) and frames below 8x8, which its init rejects, fall back to the
CPU extractor instead of failing. The parity test covers the fallback
name and the 8x8 / 7x8 / 8x7 / 4:0:0 boundaries.

Docs no longer claim that a CPU feature name always runs on the CPU, the
state rows name the mapped path, and the ADR tag pages are regenerated
after the rebase onto #1619.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…CPU (#1625)

* perf(sycl): compute ADM's AIM pass on the device, bit-exact with the CPU

adm_sycl now emits VMAF_integer_feature_aim_score and
VMAF_integer_feature_adm3_score, so under --backend sycl the default model
vmaf_v1.0.16_3d0h no longer scores ADM with the CPU extractor through the
feature-name fallback. A frame is still one upload, one graph replay and one
small copy back, with no host wait in between.

The decouple/CSF kernel stores the AIM neighbourhood band |csf(r)| / 30 next
to the DLM one, and the reduction becomes one work-group per row that computes
all three bands of the CSF denominator, the DLM contrast measure and the AIM
contrast measure, folding each row total once through
adm_cm_round_row_total(). aim and adm3 are finalised in the CPU's own float
arithmetic and equal --backend cpu on every frame of the Netflix pair, 50
frames of BBB 4K, 853x480 and 17x17 crops and 10-bit input, with the default
and the default model's options, on an Arc B580 and a UHD 770. adm2 and the
scale outputs keep their double finalisation.

The decouple quotient is now clamped in int64, as the CPU's tmp_k is. The old
kernel narrowed it to int32 first and wrapped at scales 1-3 once |t / o|
exceeded 2^16, which left integer_adm_scale2 up to 1.40e-6 from the CPU on 4K
content; every DLM output is now within 2.9e-7.

With the default model on the B580, a 4K frame drops from 48.4 to 9.2 ms at
--threads 0 and from 24.3 to 9.7 ms at --threads 16. On the UHD 770 it gets
slower (75.7 to 87.5 ms at --threads 0, 40.5 to 79.4 ms at --threads 16)
because the iGPU now does the ADM work the CPU used to do beside it.

ADR-1362. The HIP half of T-GPU-ADM-AIM-DEVICE-PASS-MISSING-SYCL-HIP-2026-09-05
stays open, with a verify command for the HIP port.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR tag pages after the rebase onto #1619

The rebase merged the by-tag pages of ADR-1359 and ADR-1362 line by line,
which left the fork-local, gpu and index counts stale; they are regenerated
from the fragments. The HIP verify note in docs/state.md no longer claims that
`--feature adm --backend hip` runs the CPU extractor, which ADR-1359 changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): make adm2 and the ADM scale outputs bit-exact with the CPU

The SYCL integer ADM twin finalised adm2, integer_adm_scale0..3 and the debug
num / den values in double, which left them up to 2.9e-7 from the CPU even
though the device accumulators match it exactly. They now go through the same
float finalisation as aim and adm3 (adm_scale_cpu / adm_terms / adm_finalise,
mirroring integer_compute_adm and adm_result_finalise), and the double
finaliser is removed.

Every ADM output now equals --backend cpu at --precision max on every frame of
the Netflix 576x324 pair and 50 frames of BBB 3840x2160, with the default and
the default model's options, on an Arc B580 and a UHD 770.
test_sycl_adm_parity and test_sycl_adm_tiny_frames compare every key bit for
bit; ADR-1362, the changelog fragment, docs/state.md and the backend guide
follow the maintainer's decision that bit-exact with the CPU is the contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes and citations after the rebase onto #1624

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes after the rebase onto #1628

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
Since ADR-1359 `--backend sycl --feature ssimulacra2` maps to
ssimulacra2_sycl when the twin can run the input. The twin now declares
that through the ADR-1324 context check: 4:0:0 input (no chroma to
convert) and frames below 8x8, which its init rejects, fall back to the
CPU extractor instead of failing. The parity test covers the fallback
name and the 8x8 / 7x8 / 8x7 / 4:0:0 boundaries.

Docs no longer claim that a CPU feature name always runs on the CPU, the
state rows name the mapped path, and the ADR tag pages are regenerated
after the rebase onto #1619.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Sep 30, 2026
…ms_ssim (#1627)

* perf(sycl): run ssimulacra2 on the device and wait once per frame in ms_ssim

ssimulacra2_sycl no longer round-trips through the host inside a frame.
Before, it blurred on the device and did everything else on the CPU: at
each of six scales it computed XYB on the host, uploaded it, copied five
full-size buffers back, waited, and combined the SSIM and edge maps and
downsampled on the host, about 4 GB of copies per 4K frame. Now submit()
uploads the raw Y/U/V planes once and the device runs colour conversion,
XYB, the products and blurs, the per-channel sums and the downsample;
collect() waits once for an 864-byte block of sums.

Everything up to the sums is bit-identical to the CPU: the TU builds
with contraction off and the shared cube root divides correctly rounded.
The CPU adds the per-pixel fp64 terms one after another, which no
parallel reduction can replay and the device has no fp64 for, so the
terms are exact fp32 pairs summed in a fixed tree. The score moves from
bit-identical to within 6.7e-12 of the CPU, the same on every device.
On an Arc B580 a 4K frame takes 33 ms instead of 963.

float_ms_ssim_sycl gives every scale its own partials span, enqueues
the whole frame in submit() and waits once in collect(); its output is
unchanged bit for bit.

The correctly rounded division and fp32-pair helpers move from the
SpEED pipeline to sycl_exact_fp.h, shared by both TUs; their slow path
uses the oneAPI math extension in device code only, so the static test
links do not pull its host fallback. The SYCL tidy baseline for
ssimulacra2_sycl.cpp drops from 19 to 0. speed_gpu_parity.py takes
--feature and --max-abs-diff to check any twin. The CUDA, HIP and Metal
twins keep the host combine; each is an RC3 row in docs/state.md.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): send inputs ssimulacra2_sycl rejects to the CPU extractor

Since ADR-1359 `--backend sycl --feature ssimulacra2` maps to
ssimulacra2_sycl when the twin can run the input. The twin now declares
that through the ADR-1324 context check: 4:0:0 input (no chroma to
convert) and frames below 8x8, which its init rejects, fall back to the
CPU extractor instead of failing. The parity test covers the fallback
name and the 8x8 / 7x8 / 8x7 / 4:0:0 boundaries.

Docs no longer claim that a CPU feature name always runs on the CPU, the
state rows name the mapped path, and the ADR tag pages are regenerated
after the rebase onto #1619.

ADR-1363.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(adr): note the ssimulacra2_sycl context check in ADR-1363

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sycl): build sycl_exact_fp.h and ssimulacra2_sycl under MSVC+SYCL

<windows.h>, which compat/win32/pthread.h pulls in first, defines min, max,
near and far as macros. They broke Intel's math.hpp declarations, the
round_quotient locals and two sycl::min calls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate ADR indexes and citations after the rebase onto #1624

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(sycl): keep one source-body helper and drop the row #1628 closed

The rebase onto #1628 left two _function_body definitions in the kernel
source contract; the second, static-only one shadowed the first and could
not find ssimulacra2_sycl's non-static collect_fex_sycl. Keep the
brace-matched helper for both callers (HISS-19). #1628 also closed
T-SYCL-MOTION-ADD-UV-SUBMIT-WAIT-2026-09-29, so drop this branch's open copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: regenerate generated docs after rebasing onto master

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(state): merge the RC2/RC3 disposition rows the rebase duplicated

Three-way by bug-id set against master b61b4d1: master's ids, plus the
ids this branch added since its base, minus the ones it removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants