Skip to content

fix(gpu-dispatch): enforce token boundary on strategy-name match - #316

Closed
lusoris wants to merge 1 commit into
masterfrom
fix/gpu-dispatch-parse-strict-strategy-match
Closed

lusoris wants to merge 1 commit into
masterfrom
fix/gpu-dispatch-parse-strict-strategy-match

Conversation

@lusoris

@lusoris lusoris commented May 30, 2026

Copy link
Copy Markdown
Contributor

Summary

VMAF_<BACKEND>_DISPATCH strategy-name matching used
strncmp(v, strategy_names[idx], slen) without a token-boundary
check, so VMAF_CUDA_DISPATCH=feature:directx silently routed to
the valid direct strategy instead of being treated as unknown.
A typo in any backend's dispatch env-var was indistinguishable
from a valid override. This PR adds a single boundary check and
pins down the strict-match contract with 9 dedicated test cases.

Type

  • fix — bug fix

Checklist

  • Commits follow Conventional Commits.
  • pre-commit run --files <touched> is green locally.
  • Unit tests pass: meson test -C build-cpu --no-rebuild test_gpu_dispatch_parse.
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. — N/A: this is pure CPU parsing of an env variable; numeric outputs are unaffected.
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. — N/A: not a feature extractor change.
  • New .c file has the appropriate license header.
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE:. — N/A: fix is a strict tightening, no breaking semantics for previously-valid input.
  • If this PR adds an ADR. — N/A per CLAUDE.md §12 r8 ("Bug fixes and implementation details do NOT need an ADR").

Bug-status hygiene (ADR-0165)

  • docs/state.md updated, OR no state delta: REASON. — no state delta: bug was not tracked in state.md.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. — N/A: no golden value changed.

Cross-backend numerical results

N/A — this PR changes only how the parser interprets one class of
malformed env values. Strategy-name routing for valid inputs is
byte-identical to before. No backend kernel was touched.

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: trivial bug fix with a small, targeted reproducer.
  • Decision matrix — no alternatives: only-one-way fix. The terminator set is mechanically derived from the doc-commented grammar plus \n for line-buffered tests.
  • AGENTS.md invariant note — no rebase-sensitive invariants: file is fork-added (ADR-0483) and absent in upstream Netflix/vmaf.
  • Reproducer / smoke-test command — see below.
  • CHANGELOG fragment — changelog.d/fixed/gpu-dispatch-parse-strict-match.md.
  • Rebase note — docs/rebase-notes.md entry gpu-dispatch-parse-strict-strategy-match (2026-05-30).

Reproducer

meson setup build-cpu core \
  -Denable_cuda=false -Denable_sycl=false -Denable_hip=false \
  -Denable_metal=disabled -Denable_dnn=disabled
ninja -C build-cpu test/test_gpu_dispatch_parse -j4
meson test -C build-cpu --no-rebuild test_gpu_dispatch_parse
# Expected: 1/1 OK

Negative-control verification was performed locally: temporarily
reverting the boundary check causes
test_directx_does_not_match_direct to fail (the parser silently
routes directx to direct); restoring the check makes all 9
cases pass.

Known follow-ups

  • Audit any per-backend dispatch_strategy.* TUs that did NOT
    migrate to vmaf_gpu_dispatch_parse_env for the same class of
    bug — out of scope for this PR; a separate sweep ADR if any are
    found.

🤖 Generated with Claude Code

The shared VMAF_<BACKEND>_DISPATCH env-variable parser in
core/src/gpu_dispatch_parse.h matched strategy names with
strncmp(v, strategy_names[idx], slen) and accepted any string with
the strategy name as a prefix. So
VMAF_CUDA_DISPATCH=feature:directx silently routed to the valid
"direct" strategy instead of being treated as unknown. A typo in
any backend's dispatch env-var was indistinguishable from a valid
override at the parse layer.

Add a token-boundary check after the strncmp: the byte at
v[slen] must be one of '\0', ',', '\n', ' ', or '\t' for a match
to succeed. The terminator set mirrors the grammar documented in
the header's leading doc comment plus newline (for env values
read line-by-line in tests).

Adds core/test/test_gpu_dispatch_parse.c (9 cases) wired into
core/test/meson.build under the `fast` suite. The negative-control
(prefix-without-boundary) cases fail against the previous code and
pass against the fix, so any regression is caught locally and in
CI.

**Research digest** — no digest needed: trivial bug fix with a
small, targeted reproducer.
**Decision matrix** — no alternatives: only-one-way fix. The
terminator set is mechanically derived from the doc-commented
grammar plus '\n' for line-buffered tests.
**AGENTS.md invariant note** — no rebase-sensitive invariants:
file is fork-added (ADR-0483) and absent in upstream.
**Reproducer / smoke-test command** —
  meson setup build-cpu core -Denable_cuda=false -Denable_sycl=false \
    -Denable_hip=false -Denable_metal=disabled -Denable_dnn=disabled
  ninja -C build-cpu test/test_gpu_dispatch_parse -j4
  meson test -C build-cpu --no-rebuild test_gpu_dispatch_parse
**CHANGELOG fragment** — changelog.d/fixed/gpu-dispatch-parse-strict-match.md
**Rebase note** — docs/rebase-notes.md entry
gpu-dispatch-parse-strict-strategy-match (2026-05-30).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@lusoris
lusoris marked this pull request as ready for review May 31, 2026 13:22
@lusoris
lusoris marked this pull request as draft May 31, 2026 13:54
@lusoris
lusoris marked this pull request as ready for review May 31, 2026 14:00
@lusoris

lusoris commented May 31, 2026

Copy link
Copy Markdown
Contributor Author

Closing as part of marathon cleanup 2026-05-31 (150 PRs merged today). Content likely superseded by sibling merges. Reopen if specific finding still needs work; bigger PRs preferred going forward per session feedback.

@lusoris lusoris closed this May 31, 2026
@lusoris
lusoris deleted the fix/gpu-dispatch-parse-strict-strategy-match branch May 31, 2026 14:08
@lusoris
lusoris restored the fix/gpu-dispatch-parse-strict-strategy-match branch May 31, 2026 18:41
@lusoris lusoris reopened this May 31, 2026
@lusoris
lusoris marked this pull request as draft May 31, 2026 18:49
@lusoris

lusoris commented Jun 1, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #525 — bundled per 2026-06-01 triage.

@lusoris lusoris closed this Jun 1, 2026
lusoris added a commit that referenced this pull request Jun 2, 2026


Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 2, 2026
#308 ENOSYS stubs + #316 dispatch token-boundary) (#525)

* fix(core): propagate correct errno codes — PR #125 defects

Five actionable errno defects from the PR #125 code review:

1. vmaf_write_output_with_format: capture errno immediately after
   open(2) / fdopen(3) failure; return -errno instead of hardcoded
   -EINVAL so callers receive the OS-precise error code.

2. vmaf_cuda_state_init: map driver-library-missing → -ENOSYS and
   cuInit(0) failure → -ENODEV, mirroring the SYCL/HIP pattern
   already established in sycl/common.cpp and cuda/cuda_helper.cuh.

3. vmaf_close: propagate return values from vmaf_thread_pool_wait()
   and vmaf_framesync_destroy() (CERT ERR33-C / Power-of-10 #7).
   NULL-pool case (n_threads == 0) correctly returns 0.

4. vmaf_cuda_preallocate_pictures: return -EBUSY on double-call to
   prevent silent ring-buffer leak and in-flight picture corruption.

5. vmaf_init: return the actual sub-init error code rather than a
   hardcoded -ENOMEM for every error path.

Fast test suite: 49/49. Pre-commit: all checks pass.
ADR-0214 parity gate: unaffected (errno-only changes, no algorithm delta).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools): check fseek / vmaf_picture_alloc / vmaf_read_pictures returns in vmaf_bench

Discharge three S9 (JPL Power-of-10 r7) unchecked-return findings in
`core/tools/vmaf_bench.c` surfaced by the 2026-05-30 audit. Independent
of PR #304 (the vmaf.c/vmaf_bench.c CUDA/SYCL state-leak fix); no
overlapping hunks.

* `yuv_pair_read_frame` — the two `fseek` calls were `(void)`-cast,
  silencing the lint but hiding the semantic bug: a failed `fseek`
  leaves the FILE position undefined, after which `fread` silently
  feeds the wrong bytes into the benchmark. Now checks each `fseek`
  and returns `-EIO` with a `perror` diagnostic on failure.
* `run_sycl_gpu_profile` per-frame loop — `vmaf_picture_alloc`
  returns were discarded; on allocation failure the subsequent
  `yuv_pair_read_frame -> ref->data[0]` dereference would crash on
  a sentinel-zero `VmafPicture`. Now captures the return code,
  logs, unrefs any already-allocated sibling, and breaks the loop.
* `run_sycl_gpu_profile` end-of-stream block — the final
  `vmaf_read_pictures(vmaf, NULL, NULL, 0)` flush surfaces pooling /
  aggregation errors via its int return that the previous code
  discarded; now captured and propagated. `printf` / `vmaf_close`
  returns explicitly `(void)`-cast to match the surrounding file
  convention.

Adds `#include <errno.h>` for `EIO`. No behavior change on success
paths. Builds clean (`meson setup build-cpu core -Denable_cuda=false
-Denable_sycl=false && ninja -C build-cpu`); clang-tidy on the
touched file produces no new warnings vs master (line numbers shift
only).

* fix(core): emit -ENOSYS stubs for libvmaf_hip.h / libvmaf_metal.h when backend OFF

Both `core/include/libvmaf/libvmaf_hip.h` and
`core/include/libvmaf/libvmaf_metal.h` document the contract that every
public entry point returns `-ENOSYS` when libvmaf is built without the
relevant backend. The real bodies in `core/src/hip/common.c`,
`core/src/metal/common.mm`, `core/src/metal/picture_import.mm`, and the
`vmaf_hip_import_state` / `vmaf_metal_import_state` /
`vmaf_metal_read_imported_pictures` definitions inside `core/src/libvmaf.c`
all sit behind `#ifdef HAVE_HIP` / `#ifdef HAVE_METAL`, so a default
`-Denable_hip=false -Denable_metal=disabled` build emitted none of those
symbols into `libvmaf.so` and any downstream link that referenced them
failed.

Fix: add `core/src/hip/stubs.c` + `core/src/metal/stubs.c` that mirror the
canonical `core/src/dnn/dnn_api.c` `VMAF_HAVE_DNN` stub pattern and wire
each TU into `libvmaf_feature_static_lib` via `hip_sources` /
`metal_sources` only when the backend is disabled. The stubs return
`-ENOSYS`, set out-params to NULL on the pointer-returning entry points,
and `vmaf_hip_available()` / `vmaf_metal_available()` correctly return 0.

Verified locally with:
  meson setup build-cpu core -Denable_hip=false -Denable_metal=disabled \\
                              -Denable_cuda=false -Denable_sycl=false   \\
                              -Denable_dnn=disabled
  ninja -C build-cpu
  nm -D build-cpu/src/libvmaf.so | grep -E "vmaf_(hip|metal)_"
all 14 documented public symbols are present (T-type) and a smoke main
linking against libvmaf.so observes -ENOSYS / NULL out-params on every
entry point and 0 from the availability probes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(gpu-dispatch): enforce token boundary on strategy-name match

The shared VMAF_<BACKEND>_DISPATCH env-variable parser in
core/src/gpu_dispatch_parse.h matched strategy names with
strncmp(v, strategy_names[idx], slen) and accepted any string with
the strategy name as a prefix. So
VMAF_CUDA_DISPATCH=feature:directx silently routed to the valid
"direct" strategy instead of being treated as unknown. A typo in
any backend's dispatch env-var was indistinguishable from a valid
override at the parse layer.

Add a token-boundary check after the strncmp: the byte at
v[slen] must be one of '\0', ',', '\n', ' ', or '\t' for a match
to succeed. The terminator set mirrors the grammar documented in
the header's leading doc comment plus newline (for env values
read line-by-line in tests).

Adds core/test/test_gpu_dispatch_parse.c (9 cases) wired into
core/test/meson.build under the `fast` suite. The negative-control
(prefix-without-boundary) cases fail against the previous code and
pass against the fix, so any regression is caught locally and in
CI.

**Research digest** — no digest needed: trivial bug fix with a
small, targeted reproducer.
**Decision matrix** — no alternatives: only-one-way fix. The
terminator set is mechanically derived from the doc-commented
grammar plus '\n' for line-buffered tests.
**AGENTS.md invariant note** — no rebase-sensitive invariants:
file is fork-added (ADR-0483) and absent in upstream.
**Reproducer / smoke-test command** —
  meson setup build-cpu core -Denable_cuda=false -Denable_sycl=false \
    -Denable_hip=false -Denable_metal=disabled -Denable_dnn=disabled
  ninja -C build-cpu test/test_gpu_dispatch_parse -j4
  meson test -C build-cpu --no-rebuild test_gpu_dispatch_parse
**CHANGELOG fragment** — changelog.d/fixed/gpu-dispatch-parse-strict-match.md
**Rebase note** — docs/rebase-notes.md entry
gpu-dispatch-parse-strict-strategy-match (2026-05-30).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(changelog): add bundle-core-c fragment for PRs #148 #307 #308 #316

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Lusoris <lusoris@pm.me>
@lusoris
lusoris deleted the fix/gpu-dispatch-parse-strict-strategy-match branch June 4, 2026 08:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant