Skip to content

chore(meta): post-cutover URL sweep — lusoris/vmaf → VMAFx/vmafx - #1

Merged
lusoris merged 1 commit into
masterfrom
chore/post-cutover-url-sweep
May 28, 2026
Merged

lusoris merged 1 commit into
masterfrom
chore/post-cutover-url-sweep

Conversation

@lusoris

@lusoris lusoris commented May 28, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Replace all 279 occurrences of the lusoris/vmaf repository slug across 113 files with VMAFx/vmafx following the GitHub org cutover (master at 384d97d037).
  • GHCR image registry paths updated to vmafx/vmafx (lowercase per OCI convention, e.g. ghcr.io/vmafx/vmafx).
  • Version tag scheme (v3.x.y-lusoris.N), lusoris as a personal identifier in commit authorship, ADR ## References verbatim citations, and lusoris@pm.me email are intentionally left unchanged.

Research reference: docs/research/0731-vmafx-org-migration-plan.md

Test plan

  • git grep lusoris/vmaf returns only the two intentional "what was replaced" prose references in docs/state.md and docs/rebase-notes.md.
  • git grep lusoris/vmafx returns zero results.
  • All pre-commit hooks pass (clang-format, black, ruff, semgrep, copyright, ADR collision).
  • mkdocs build --strict — run in CI (no local mkdocs install in worktree).
  • Sample URL spot-check: curl -sI https://github.com/VMAFx/vmafx should return 200.

Deliverables checklist (ADR-0108)

  • Research digest: no digest needed: trivial — pure mechanical URL replacement post-org-cutover; plan already documented in docs/research/0731-vmafx-org-migration-plan.md.
  • Decision matrix: no alternatives: only-one-way fix — the org has moved; all slug references must track it.
  • AGENTS.md invariant note: no rebase-sensitive invariants — only fork-local URL strings changed; upstream Netflix/vmaf does not contain any of these references.
  • Reproducer / smoke-test: git grep lusoris/vmaf | grep -v 'lusoris@\|lusoris\.N\|Author: lusoris' — should return zero lines after merge.
  • Changelog fragment: changelog.d/changed/post-cutover-url-sweep.md
  • Rebase-notes entry: docs/rebase-notes.md §chore/post-cutover-url-sweep — no Netflix conflict.

docs/state.md

Row added to Recently closed: T-POST-CUTOVER-URL-SWEEP-2026-05-28.

Scope of changes

Category Files Replacements
docs/ (ADRs, guides, research) 71 ~170
.github/workflows + templates 3 8
changelog.d fragments 7 7
CHANGELOG.md 1 4
README.md / SECURITY.md / SUPPORT.md 3 9
LICENSES/ 1 3
deploy/helm + docker/ 5 12
dev/Containerfile 1 1
core/src/dnn/model_loader.c 1 2
scripts/ 4 6
model/ 2 4
mcp-server/ 1 1
.claude/agents/ 1 2
mkdocs.yml 1 1
changelog.d/changed/post-cutover-url-sweep.md 1 (new) —
docs/state.md + docs/rebase-notes.md 2 —

Sample replacements

  1. README.md badge URLs: https://github.com/lusoris/vmaf/actions/workflows/tests-and-quality-gates.yml → https://github.com/VMAFx/vmafx/actions/workflows/tests-and-quality-gates.yml
  2. deploy/helm/vmafx/values.yaml: ghcr.io/lusoris/vmafx → ghcr.io/vmafx/vmafx
  3. dev/Containerfile OCI label: https://github.com/lusoris/vmaf → https://github.com/VMAFx/vmafx
  4. core/src/dnn/model_loader.c OIDC Sigstore identity pattern: https://github.com/lusoris/vmaf/.github/workflows/.+ → https://github.com/VMAFx/vmafx/.github/workflows/.+
  5. .github/workflows/docker-publish-production.yml: IMAGE_NAME: lusoris/vmafx → IMAGE_NAME: vmafx/vmafx

🤖 Generated with Claude Code

Replace all in-tree references to the `lusoris/vmaf` repository slug with
`VMAFx/vmafx` following the GitHub org cutover (master at 384d97d). GHCR
image paths updated to `vmafx/vmafx` (lowercase per OCI registry convention).
113 files updated, 279 occurrences replaced. Version tag scheme
(v3.x.y-lusoris.N) and the `lusoris` user identity in commit authorship and
ADR References sections are left unchanged.

Covers: README.md, mkdocs.yml, docs/**/*.md, CHANGELOG.md, .github/workflows,
changelog.d fragments, core/src/dnn/model_loader.c OIDC URL, deploy/helm,
docker/Dockerfile.production*, dev/Containerfile, scripts/*, model/tiny/,
mcp-server/vmaf-mcp/, SECURITY.md, SUPPORT.md, AGENTS.md (no changes needed),
CLAUDE.md (no repo slug present), LICENSES/Apache-2.0-u2netp.txt.

Research ref: docs/research/0731-vmafx-org-migration-plan.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris marked this pull request as ready for review May 28, 2026 10:36
@lusoris
lusoris merged commit 3ec4af7 into master May 28, 2026
34 of 35 checks passed
@lusoris
lusoris deleted the chore/post-cutover-url-sweep branch May 28, 2026 10:36
lusoris added a commit that referenced this pull request May 28, 2026
…0735)

Fixes 8 pre-existing required-aggregator failures that blocked every PR
post-#60. All classified as workflow config bugs, post-rename path drift,
or code bugs from the float_ansnr drop in PR #38.

1. Windows MSVC CUDA + SYCL build: ninja -C libvmaf\build -> core\build
   (stale ADR-0700 rename leftover in libvmaf-build-matrix.yml)
2. Ubuntu HIP smoke: remove test_float_ansnr_hip_extractor_registered
   (float_ansnr_hip dropped by PR #38 / ADR-0720; test not updated)
3. Netflix CPU Golden (D24): remove float_ansnr from
   VmafIntegerFeatureExtractor._generate_result() and its
   ATOM_FEATURES_TO_VMAFEXEC_KEY_DICT — the CLI exited 255 on the
   removed extractor before any score was computed; the CI gate tests
   already expect ansnr to be absent (assertRaises KeyError)
4. CodeQL Python: update codeql-config.yml paths libvmaf/ -> core/;
   add explicit no-op build step to suppress C++ autobuild
5. Gitleaks: add go.sum, Cargo.lock, gen/go/*.pb.go, h1: regex, and
   stopwords to .gitleaks.toml allowlist (package-manager hash FPs)
6. Semgrep: add compat/python-vmaf/matlab/ and
   compat/python-vmaf/resource/ to .semgrepignore (post-ADR-0700
   rename; python/vmaf/matlab/ no longer matched the file);
   verified locally: 0 findings
7. Tiny AI: implement missing dumps_jsonl_row (aiutils.jsonl_utils)
   and dumps_registry_json + write_registry_json (vmaf_train.registry)
   which tests imported but were never implemented
8. core/AGENTS.md + docs/state.md: document the ansnr-removal invariant
   and track T-LEGACY-RUNNER-ANSNR-BROKEN as an open bug

Residual: VmafFeatureExtractor (legacy float path) still requests
float_ansnr; legacy runner tests fail locally. The CI golden gate only
runs test_run_vmaf_runner + checkerboard (integer path), now fixed.
Netflix golden assertion values are unchanged (GLOBAL PROJECT RULES #1).

Research-0735: docs/research/0735-ci-required-failures-round-3-2026-05-28.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request May 28, 2026
…0735) (#86)

Fixes 8 pre-existing required-aggregator failures that blocked every PR
post-#60. All classified as workflow config bugs, post-rename path drift,
or code bugs from the float_ansnr drop in PR #38.

1. Windows MSVC CUDA + SYCL build: ninja -C libvmaf\build -> core\build
   (stale ADR-0700 rename leftover in libvmaf-build-matrix.yml)
2. Ubuntu HIP smoke: remove test_float_ansnr_hip_extractor_registered
   (float_ansnr_hip dropped by PR #38 / ADR-0720; test not updated)
3. Netflix CPU Golden (D24): remove float_ansnr from
   VmafIntegerFeatureExtractor._generate_result() and its
   ATOM_FEATURES_TO_VMAFEXEC_KEY_DICT — the CLI exited 255 on the
   removed extractor before any score was computed; the CI gate tests
   already expect ansnr to be absent (assertRaises KeyError)
4. CodeQL Python: update codeql-config.yml paths libvmaf/ -> core/;
   add explicit no-op build step to suppress C++ autobuild
5. Gitleaks: add go.sum, Cargo.lock, gen/go/*.pb.go, h1: regex, and
   stopwords to .gitleaks.toml allowlist (package-manager hash FPs)
6. Semgrep: add compat/python-vmaf/matlab/ and
   compat/python-vmaf/resource/ to .semgrepignore (post-ADR-0700
   rename; python/vmaf/matlab/ no longer matched the file);
   verified locally: 0 findings
7. Tiny AI: implement missing dumps_jsonl_row (aiutils.jsonl_utils)
   and dumps_registry_json + write_registry_json (vmaf_train.registry)
   which tests imported but were never implemented
8. core/AGENTS.md + docs/state.md: document the ansnr-removal invariant
   and track T-LEGACY-RUNNER-ANSNR-BROKEN as an open bug

Residual: VmafFeatureExtractor (legacy float path) still requests
float_ansnr; legacy runner tests fail locally. The CI golden gate only
runs test_run_vmaf_runner + checkerboard (integer path), now fixed.
Netflix golden assertion values are unchanged (GLOBAL PROJECT RULES #1).

Research-0735: docs/research/0735-ci-required-failures-round-3-2026-05-28.md

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request May 30, 2026
…B-MISSING

Both rows referenced classes / imports that no longer exist on master:

1. T-LEGACY-RUNNER-ANSNR-BROKEN — AnsnrFeatureExtractor was deleted
   by PR #283 (merged at faede7a). The class is absent from
   compat/python-vmaf/core/feature_extractor.py on master tip
   d45d503. The Netflix golden assertions still cover the
   integer-path VMAF score (Rule #1 preserved).

2. T-LEGACY-RUNNER-STUB-MISSING-2026-05-29 — VmafLegacyQualityRunner
   import-time failure: the class was removed in ADR-0749 / PR #87
   and python/test/quality_runner_test.py was updated by the
   ADR-0749 sunset PR to drop the import (only a removed-comment
   placeholder at line 49 remains).

Moved both rows from Open to Recently closed with verified-on-master
reproducer commands.

No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- [x] **Research digest**: no digest needed: state.md hygiene
- [x] **Decision matrix**: no alternatives: only-one-way fix
- [x] **AGENTS.md invariant**: no rebase-sensitive invariants
- [x] **Reproducer**: `git show origin/master:compat/python-vmaf/core/feature_extractor.py | grep -c 'class AnsnrFeatureExtractor'` returns 0; `git show origin/master:python/test/quality_runner_test.py | grep -c '^from.*VmafLegacy'` returns 0.
- [x] **Changelog**: changelog.d/changed/state-md-drift-sweep-20260530.md
- [x] **Rebase notes**: no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request May 30, 2026
…B-MISSING

Both rows referenced classes / imports that no longer exist on master:

1. T-LEGACY-RUNNER-ANSNR-BROKEN — AnsnrFeatureExtractor was deleted
   by PR #283 (merged at faede7a). The class is absent from
   compat/python-vmaf/core/feature_extractor.py on master tip
   d45d503. The Netflix golden assertions still cover the
   integer-path VMAF score (Rule #1 preserved).

2. T-LEGACY-RUNNER-STUB-MISSING-2026-05-29 — VmafLegacyQualityRunner
   import-time failure: the class was removed in ADR-0749 / PR #87
   and python/test/quality_runner_test.py was updated by the
   ADR-0749 sunset PR to drop the import (only a removed-comment
   placeholder at line 49 remains).

Moved both rows from Open to Recently closed with verified-on-master
reproducer commands.

No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- [x] **Research digest**: no digest needed: state.md hygiene
- [x] **Decision matrix**: no alternatives: only-one-way fix
- [x] **AGENTS.md invariant**: no rebase-sensitive invariants
- [x] **Reproducer**: `git show origin/master:compat/python-vmaf/core/feature_extractor.py | grep -c 'class AnsnrFeatureExtractor'` returns 0; `git show origin/master:python/test/quality_runner_test.py | grep -c '^from.*VmafLegacy'` returns 0.
- [x] **Changelog**: changelog.d/changed/state-md-drift-sweep-20260530.md
- [x] **Rebase notes**: no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 1, 2026
PR #38 (ADR-0709) dropped float_ansnr from the C library on 2026-05-28
because Netflix never adopted it as a VMAF component (Research-0733:
"pre-VMAF 2001 era; legacy only"). The sunset was correct — ANSNR has
no production use today, isn't part of the VMAF v0.6.1 model, and was
only kept by Netflix upstream for libvmaf C-surface completeness.

But the corresponding Netflix golden Python test assertions in
python/test/feature_extractor_test.py (20 assertAlmostEqual calls
expecting VMAF_feature_anpsnr_score / VMAF_integer_feature_anpsnr_score
values) were left in place. macOS CI now KeyErrors on every
test_run_vmaf_(integer_)fextractor* variant because the C library no
longer produces the key the test reads.

Per ADR-0709 / PR #38 sunset (user-authorized 2026-06-01: "we dont use
it for anything and decided to"), complete the Python-side sunset:

- Strip 20 anpsnr assertions from python/test/feature_extractor_test.py
  (11 single-line + 9 multi-line VMAF_feature_anpsnr_score /
  VMAF_integer_feature_anpsnr_score variants)
- Drop 2 anpsnr fixture rows from compat/python-vmaf/core/result.py
- Drop 1 anpsnr legacy comment in compat/python-vmaf/core/feature_extractor.py
- Drop the "float_anpsnr → float_ansnr" mapping in ai/data/feature_extractor.py

This is the canonical case where CLAUDE.md Global Rule #1 ("never modify
Netflix golden-data assertions") yields to ADR-0709's explicit sunset:
the assertions test a feature the fork no longer ships. They are not
score-correctness goldens — they are completeness checks against a
removed code path.

Net behavioural change: macOS feature_extractor_test.py::test_run_vmaf_*
variants stop KeyError-ing; all other golden assertions (VMAF_score,
vif_score, adm_score, motion_score, integer_*, ssim, etc.) untouched.

Closes T-LEGACY-RUNNER-ANSNR-BROKEN.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 2, 2026
…ed_score_ptx, Metal scaffold (#517)

* fix(test): portable setenv/unsetenv wrappers for Windows MinGW64

MinGW64's default headers do not expose setenv/unsetenv (POSIX-only).
Add #ifdef _WIN32 wrappers in test_gpu_dispatch_runtime.c that delegate
to _putenv_s / _putenv("") so the test builds and runs on both POSIX
and Windows targets without _GNU_SOURCE or compatibility magic.

Fixes: Build — Windows MinGW64 (CPU) CI failure on master tip
40d192e (run 26726137428).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(compat): add pthread_once_t + pthread_once to Win32 pthread shim

MSVC has no <pthread.h> so ssim_simd.h:84, float_ssim.c, float_ms_ssim.c,
and cuda/dispatch_strategy.c all failed to build (error C2143: syntax
error: missing ')' before '*') when pthread_once_t appeared in the
function signature.

Extend core/src/compat/win32/pthread.h to define pthread_once_t as
INIT_ONCE, PTHREAD_ONCE_INIT as INIT_ONCE_STATIC_INIT, and a static
inline pthread_once() that delegates to InitOnceExecuteOnce. Mirrors
the pattern already used in cuda/dispatch_strategy.c (ADR-0181).

Fixes: Build — Windows MSVC + CUDA CI failure on master tip
40d192e (run 26726137428 / 26726137416).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(cuda): register speed/speed_score.cu in cuda_cu_sources

speed_chroma_cuda.c and speed_temporal_cuda.c both declare
  extern const char speed_score_ptx[];
and pass it to cuModuleLoadData. The .cu file exists at
core/src/feature/cuda/speed/speed_score.cu but was never added to
cuda_cu_sources in core/src/meson.build, so the bin2c pipeline never
generated the PTX blob array, causing a linker error:

  undefined reference to 'speed_score_ptx'

Add the entry so the standard nvcc + bin2c pipeline generates
speed_score.fatbin and the speed_score_ptx C array symbol.

Fixes: Build — Ubuntu CUDA + Build — Ubuntu CUDA Static CI failures
on master tip 40d192e (run 26726137428 / 26726137416).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(metal): rename integer_motion_v2_metal extractor from legacy name

The extractor struct in integer_motion_v2_metal.mm had
  .name = "motion_v2_metal"
which is the legacy T8-1 scaffold name. The coverage-audit test
test_metal_kernel_coverage_audit looks up "<basename>_metal" for
each basename in the canonical list — "integer_motion_v2_metal" —
and got NULL back because the stored name did not match.

Rename the .name field to "integer_motion_v2_metal" and update:
  - core/src/metal/dispatch_strategy.c (g_metal_features[] entry)
  - core/test/test_metal_smoke.c (extractor lookup + dispatch check)
  - core/test/test_metal_kernel_registration.c (kRegisteredMetalExtractors + kTemporal)
  - core/test/test_metal_motion_v2_parity.c (vmaf_use_feature call)

No behaviour change: the TEMPORAL flag, provided_features array, and
all kernel dispatch paths are unchanged; only the public name string
is corrected to match the canonical integer_motion_v2_metal convention
used by every other integer_* Metal extractor in the tree.

Fixes: Build — macOS Metal (T8-1 scaffold) CI failure on master tip
40d192e (run 26726137428 / 26726137416).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(sycl): add close_fex_sycl forward decls in integer_adm + integer_vif

The close_fex_sycl calls in init_fex_sycl error paths (introduced by the
SY-2a leak fix) appear before the function definition at the bottom of
the file. Builds fail under strict C++ modes with "use of undeclared
identifier 'close_fex_sycl'". Add a forward decl right before
init_fex_sycl.

Master Ubuntu SYCL + CUDA build failed with this error; bundled here
into the master-unblock PR so a single PR restores green CI on all 5
affected platforms.

This is the build-fix portion of PR #516. The full leak-fix (calling
close_chroma_sycl / close_temporal_sycl from speed extractor error paths)
remains in #516 to land separately.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(iqa): replace ATOMIC_VAR_INIT(NULL) with NULL — MSVC C2099

ATOMIC_VAR_INIT is absent from MSVC's <stdatomic.h> (undefined identifier,
treated as extern returning int), making the file-scope static initialiser
of g_ssim_dispatch_installer non-constant and causing C2099 on the Windows
MSVC + CUDA build matrix.

C11 §7.17.2.1 p3 explicitly allows plain NULL as an initial value for any
atomic type; the macro was deprecated in C17 for exactly this reason.
Replace ATOMIC_VAR_INIT(NULL) with NULL — semantically identical on all
three compilers we target (GCC, Clang, MSVC).

Fixes CI job "Build — Windows MSVC + CUDA (build only)".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): portable temp-file handling in test_svm_api — MSVC + MinGW64

Two separate Windows failures in test_svm_api.c:

1. MSVC + SYCL build (C compilation): `fatal error: 'unistd.h' file not found`
   at line 47.  Guard the include with `#ifndef _WIN32 / #endif`.

2. MinGW64 runtime: `test_save_load_roundtrip` failed at assertion
   "mkstemp ok" — `mkstemp("/tmp/vmaf-test-svm-XXXXXX", ...)` returned -1
   because MSYS2/MinGW64 in the GitHub Actions runner does not expose a
   usable /tmp from the MINGW64 shell.

   Fix: add `make_svm_temp_path()` portable helper (mirrors the
   `make_temp_output_path()` pattern in `test_public_api_score.c` per
   `core/test/AGENTS.md §6`).  On `_WIN32`: query `GetTempPathA()` +
   embed PID for uniqueness + pre-create the file.  On POSIX: mkstemp on
   the existing /tmp template.  Replace `unlink()` with `remove()` so no
   unistd.h dependency is needed for the cleanup path.

Fixes CI jobs "Build — Windows MSVC + oneAPI SYCL (build only)" and
"Build — Windows MinGW64 (CPU)".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): portable mkdtemp in test_mkdirp — MinGW64 /tmp missing

test_mkdirp_single_level failed on Windows MinGW64 at assertion
"mkdtemp must succeed" because MSYS2/MinGW64 in the GitHub Actions runner
does not expose a usable /tmp from the MINGW64 shell — hardcoded
/tmp/vmaf_mkdirp_*_XXXXXX templates return NULL from mkdtemp.

Fix: add `portable_mkdtemp()` + MKDTEMP/RMDIR macros (mirrors the
approach documented in `core/test/AGENTS.md §6`).  On `_WIN32`:
query `GetTempPathA()`, build a PID-unique path, and
`CreateDirectoryA()`.  On POSIX: delegate to `mkdtemp(3)` unchanged.
Replace `(void)rmdir(p)` with the `RMDIR()` macro so Windows builds
use `_rmdir()` from `<direct.h>`.  Guard `<unistd.h>` with
`#ifndef _WIN32`.

All three test functions (single_level, idempotent_eexist,
normalize_double_slash) updated consistently.

Fixes CI job "Build — Windows MinGW64 (CPU)".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changelog): add Layer-2 platform-breakage fix entries

Document three new fixes (MSVC C2099, MSVC+SYCL unistd.h, MinGW64
mkstemp/mkdtemp runtime) in the changelog fragment for PR #517.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): test_mkdirp.c — POSIX path needs /tmp/...XXXXXX template

The previous portable_mkdtemp() helper only existed inside the _WIN32
branch and the POSIX MKDTEMP macro expanded to bare mkdtemp(tmpl) —
which was called with the uninitialised char tmpl[260] buffer.
mkdtemp(3) requires a NUL-terminated template ending in 6 'X' chars;
without it, it returns NULL with EINVAL and every single test_mkdirp_*
case failed on Linux (test_mkdirp_single_level: "mkdtemp must succeed").

Unify portable_mkdtemp across platforms: on POSIX it now initialises
the buffer with "/tmp/vmaf_mkdirp_XXXXXX" via snprintf before
delegating to mkdtemp(3). The Win32 branch is unchanged.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(test): test_mkdirp.c — define S_ISDIR shim on Windows

Windows MSVC + Clang/SYCL ship <sys/stat.h> with _S_IFMT / _S_IFDIR
but no POSIX S_ISDIR(mode) macro, causing:
  error: call to undeclared function 'S_ISDIR'
on lines 105 and 153 of test_mkdirp.c (Layer-3 follow-up to the
portable_mkdtemp helper from Layer-2).

Define the standard expansion under #ifdef _WIN32 / #ifndef S_ISDIR.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(mkdirp): normalize POSIX '/' to '\\' on Windows in path_normalize

mkdirp.c uses PATH_SEPARATOR == '\\' on Windows when walking the path
backwards to strip leaves. But path_normalize() only collapses '/'
sequences and never converts them, so a mixed-separator path like
  C:\Users\runneradmin\AppData\Local\Temp\vmaf_mkdirp_1234/a/b/c
walked backwards stops at the LAST '\\' (i.e. \\vmaf_mkdirp_1234),
leaving '/a/b/c' as a single un-split leaf. _mkdir() then tries to
create 'a/b/c' inside the parent in one shot and fails with ENOENT.

Concrete symptom on master tip + Layer-3 commits of #517: MinGW64
runner reports `test_mkdirp_normalize_double_slash` FAIL with
"mkdirp with redundant '/' must succeed".

Under _WIN32, convert each '/' to '\\' as we copy + collapse runs of
either separator. POSIX path remains unchanged. The Windows kernel
treats both separators interchangeably at the API surface, so this
only affects mkdirp's internal walker.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(backends/sycl): document close_fex_sycl forward-decl pattern (SY-2a)

The Layer-1 commit added forward declarations of close_fex_sycl in
integer_adm_sycl.cpp + integer_vif_sycl.cpp so init_fex_sycl's
USM-failure cleanup compiles under strict C++ modes. ADR-0167
doc-substance gate requires a corresponding edit under
docs/backends/sycl/ when SYCL feature kernels are touched.

Add a "## Design notes" bullet explaining the pattern + why each TU
needs the local forward declaration.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(test): correct metal coverage audit basename + fix go-ci meson setup path

Two pre-existing master CI failures preventing merges via Required Checks
Aggregator:

1. test_metal_kernel_coverage_audit (all macOS jobs): g_metal_kernel_basenames[]
   listed "integer_motion_v2" expecting extractor "integer_motion_v2_metal", but
   the registered name is "motion_v2_metal" (ADR-0421 T8-1c short alias).
   Fix: change entry to "motion_v2"; add clarifying comment that entries are
   registered extractor name prefixes, not .mm file stems.

2. Go CI: meson setup core/build-cpu passed only one positional arg, so meson
   interpreted it as the source dir (no meson.build there) and exited 1.
   Fix: meson setup core core/build-cpu matching the build.yml pattern.

Research: docs/research/preexisting-macos-python-tinyai-failures-2026-06-01.md
Changelog: changelog.d/fixed/preexisting-macos-python-tinyai.md
state.md: T-PREEXISTING-MACOS-TINYAI-CI-FAILURES-2026-06-01 added to Recently closed
No ADR: bug fixes (CLAUDE.md §12 r8)
no rebase impact: test-only and CI workflow changes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ai/test): repair 9 Tiny AI pytest failures on master tip

Pre-existing failures in the Tiny AI (DNN Suite + ai/ Pytests) required
check (run 26726137418):

1. test_data_datasets_branches.py (3 failures): ManifestEntry._sha256_shape
   validator (added PR #506) rejects sha256 values shorter than 64 hex chars.
   Test fixtures used abbreviated stubs "deadbeef", "cafebabe", "s". Fix:
   replace with valid 64-char hex constants _SHA256_A / _SHA256_B / _SHA256_K.

2. test_frame_loader.py (4 failures): iter_frames now passes stderr=subprocess.PIPE
   to Popen (for ffmpeg diagnostic capture on non-zero exit). The fake_popen stub
   only accepted stdout, causing TypeError on every call. _FakeProcess also lacked
   a stderr attribute accessed by iter_frames cleanup path. Fix: add stderr param
   to fake_popen signature and stderr=None to _FakeProcess.

3. test_parquet_utils.py (1 failure): write_parquet_atomic was refactored to use
   pyarrow directly via _write_v2(); it no longer calls df.to_parquet(). The test
   override of that method never fired, so no RuntimeError was raised. Fix: inject
   via monkeypatch on aiutils.parquet_utils._write_v2 instead.

No ADR: test-only fixes (CLAUDE.md §12 r8)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* revert: undo Layer-1 Metal extractor rename — kept basenames fix instead

Layer-1 commit 2b61f9b renamed the integer_motion_v2_metal extractor's
.name field from "motion_v2_metal" to "integer_motion_v2_metal" to satisfy
the kernel coverage audit. The later bundled commit 902d8c2 (from #518)
applied the opposite resolution — change the audit's basename list from
"integer_motion_v2" to "motion_v2" — because "motion_v2_metal" is the
canonical short alias chosen in ADR-0421 / T8-1c.

Both fixes individually closed the audit; together they re-broke it.
Revert the .name rename so the registered name matches the short-alias
convention and the bundled audit fix.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(ci/sanitizers): exclude test_y4m_alloc_failure from ASan/TSan/MSan runs

test_y4m_alloc_failure uses setrlimit(RLIMIT_AS) to cap the process
address space at 256 MiB, forcing malloc failure in y4m_input_open_impl.
ASan, TSan, and MSan runtimes each require hundreds of MiB of shadow-memory
mmap during startup; the sanitizer's own mmap fails before any test logic
runs, producing "Failed to mmap" / "internal allocator is out of memory"
SIGABRT on every sanitizer build.

The bug the test guards (dst_buf-NULL regression) is exercised by every
unsanitized fast-suite run, so no correctness gap results from the exclusion.

Added to EXCLUDE regex in:
- sanitizers.yml (ASan+UBSan PR gate + TSan master gate)
- tests-and-quality-gates.yml (address/undefined/thread matrix)

Tracked as T-Y4M-ALLOC-RLIMIT-AS-SANITIZER-INCOMPATIBLE in docs/state.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci/lint): exclude core/src/compat/win32/ from clang-tidy changed-files scan

The Win32 pthread shim (core/src/compat/win32/pthread.h) contains an
intentional #error "win32/pthread.h shim included on a non-Windows target"
guard. When the file appears in the PR diff, clang-tidy on the Linux CI
runner hits this guard as a clang-diagnostic-error and exits with code 1,
blocking the required Clang-Tidy check.

Add core/src/compat/win32/ to all four grep -v exclusion blocks in the
"Run clang-tidy on changed files" step (PR/push/zero-sha push/dispatch
variants), matching the existing pattern for arm64, cuda, sycl, hip, mcp,
and fuzz paths that are likewise excluded because they require non-Linux
toolchains or non-CPU build configurations.

Tracked as T-CLANG-TIDY-WIN32-PTHREAD-SHIM-2026-06-01 in docs/state.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ai/test): repair 5 remaining Tiny AI pytest failures on master tip

Two root causes:

1. test_manifest_entry_is_frozen — ManifestEntry was migrated from
   @DataClass(frozen=True) to pydantic.BaseModel(ConfigDict(frozen=True))
   in PR #506. Pydantic raises ValidationError (ValueError subclass) on
   frozen-field assignment, not dataclasses.FrozenInstanceError
   (AttributeError subclass). Updated assertion to pydantic.ValidationError.

2. test_iter_frames_gray_keeps_2d_shape + 4 × test_iter_frames_packed_color_*
   — iter_frames calls proc.wait(timeout=_FFMPEG_WAIT_TIMEOUT_S) but
   _FakeProcess.wait() only accepted a positional argument. Added
   timeout: float | None = None keyword parameter to _FakeProcess.wait().

Tracked as T-AI-TEST-PYDANTIC-FROZEN-INSTANCE-2026-06-01 and
T-AI-TEST-FRAME-LOADER-FAKE-WAIT-TIMEOUT-2026-06-01 in docs/state.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(mcp): remove dead _run_benchmark() + fix test_bug3 for new raise semantics

server.py contained two definitions of async def _run_benchmark:
- Line 824: no parameters, returned dict on failure (pre-ADR-0608)
- Line 1409: takes progress_token, raises RuntimeError on failure

Python uses the last definition; the first was dead code. Its presence
was the source of confusion that led to test_bug3_run_benchmark_surfaces_
silent_pipefail being written for the old return-dict contract.

Remove the dead first definition (80 lines). Update the test to:
- Use monkeypatch.setattr(asyncio, "create_subprocess_exec", ...) like
  the other tests in test_path_and_bench_env.py, to avoid patching via
  patch.object(srv.asyncio, ...) which resolves to the wrong binding
- Assert pytest.raises(RuntimeError, match="benchmark failed.*no output")
  matching the ADR-0608 E-1 raise semantics of the current implementation

Tracked as T-MCP-RUN-BENCHMARK-DEAD-CODE-2026-06-01 in docs/state.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(state/changelog): document Layer-5 residual CI failure fixes

Add 5 new entries to docs/state.md under Recently Closed:
- T-Y4M-ALLOC-RLIMIT-AS-SANITIZER-INCOMPATIBLE-2026-06-01
- T-CLANG-TIDY-WIN32-PTHREAD-SHIM-2026-06-01
- T-AI-TEST-PYDANTIC-FROZEN-INSTANCE-2026-06-01
- T-AI-TEST-FRAME-LOADER-FAKE-WAIT-TIMEOUT-2026-06-01
- T-MCP-RUN-BENCHMARK-DEAD-CODE-2026-06-01

Append 5 fix entries to changelog.d/fixed/preexisting-macos-python-tinyai.md.
Update docs/rebase-notes.md Layer-5 entry with newly touched files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(mcp/tools): note Layer-5 dead-code removal of legacy _run_benchmark

ADR-0167 Doc-Substance Gate flags the Layer-5 MCP server.py edit as a
surface change without matching docs/mcp/ entry. Add an "Error contract"
callout under run_benchmark explaining that the legacy partial-dict
fallback was removed in PR #517 Layer-5 and that MCP clients should
branch on isError=True (ADR-0608).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(state): rephrase "this PR" → "PR #517" in clang-tidy row

ADR-0165 state.md Touch Gate rejects placeholder strings like "this PR"
in newly-inserted rows because they never get rewritten to numeric refs
after merge. Replace the only such phrase in #517's Layer-5 row
T-CLANG-TIDY-WIN32-PTHREAD-SHIM-2026-06-01 with the explicit PR number.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat!: complete ANSNR sunset — drop 20 stale Netflix golden assertions

PR #38 (ADR-0709) dropped float_ansnr from the C library on 2026-05-28
because Netflix never adopted it as a VMAF component (Research-0733:
"pre-VMAF 2001 era; legacy only"). The sunset was correct — ANSNR has
no production use today, isn't part of the VMAF v0.6.1 model, and was
only kept by Netflix upstream for libvmaf C-surface completeness.

But the corresponding Netflix golden Python test assertions in
python/test/feature_extractor_test.py (20 assertAlmostEqual calls
expecting VMAF_feature_anpsnr_score / VMAF_integer_feature_anpsnr_score
values) were left in place. macOS CI now KeyErrors on every
test_run_vmaf_(integer_)fextractor* variant because the C library no
longer produces the key the test reads.

Per ADR-0709 / PR #38 sunset (user-authorized 2026-06-01: "we dont use
it for anything and decided to"), complete the Python-side sunset:

- Strip 20 anpsnr assertions from python/test/feature_extractor_test.py
  (11 single-line + 9 multi-line VMAF_feature_anpsnr_score /
  VMAF_integer_feature_anpsnr_score variants)
- Drop 2 anpsnr fixture rows from compat/python-vmaf/core/result.py
- Drop 1 anpsnr legacy comment in compat/python-vmaf/core/feature_extractor.py
- Drop the "float_anpsnr → float_ansnr" mapping in ai/data/feature_extractor.py

This is the canonical case where CLAUDE.md Global Rule #1 ("never modify
Netflix golden-data assertions") yields to ADR-0709's explicit sunset:
the assertions test a feature the fork no longer ships. They are not
score-correctness goldens — they are completeness checks against a
removed code path.

Net behavioural change: macOS feature_extractor_test.py::test_run_vmaf_*
variants stop KeyError-ing; all other golden assertions (VMAF_score,
vif_score, adm_score, motion_score, integer_*, ssim, etc.) untouched.

Closes T-LEGACY-RUNNER-ANSNR-BROKEN.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat!: extend ANSNR sunset round 2 - strip ansnr from VMAF_feature lists

* test: ANSNR sunset round 3

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
…B-MISSING

Both rows referenced classes / imports that no longer exist on master:

1. T-LEGACY-RUNNER-ANSNR-BROKEN — AnsnrFeatureExtractor was deleted
   by PR #283 (merged at faede7a). The class is absent from
   compat/python-vmaf/core/feature_extractor.py on master tip
   d45d503. The Netflix golden assertions still cover the
   integer-path VMAF score (Rule #1 preserved).

2. T-LEGACY-RUNNER-STUB-MISSING-2026-05-29 — VmafLegacyQualityRunner
   import-time failure: the class was removed in ADR-0749 / PR #87
   and python/test/quality_runner_test.py was updated by the
   ADR-0749 sunset PR to drop the import (only a removed-comment
   placeholder at line 49 remains).

Moved both rows from Open to Recently closed with verified-on-master
reproducer commands.

No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- [x] **Research digest**: no digest needed: state.md hygiene
- [x] **Decision matrix**: no alternatives: only-one-way fix
- [x] **AGENTS.md invariant**: no rebase-sensitive invariants
- [x] **Reproducer**: `git show origin/master:compat/python-vmaf/core/feature_extractor.py | grep -c 'class AnsnrFeatureExtractor'` returns 0; `git show origin/master:python/test/quality_runner_test.py | grep -c '^from.*VmafLegacy'` returns 0.
- [x] **Changelog**: changelog.d/changed/state-md-drift-sweep-20260530.md
- [x] **Rebase notes**: no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
…291 state.md drift sweep + #233 changelog concat fix) (#527)

* docs(lint): NOLINT cluster audit and refactor plan (ADR-0780)

Swept all 218 NOLINT annotations in core/src/ for clusters of five or more
identical suppressions per file. Identified five clusters (71 annotations):

- 21 bare performance-no-int-to-ptr in GPU slab allocators — ADR-0278
  non-compliant (no citations); scheduled for SLAB_FIELD macro (PR B).
- 12 bugprone-implicit-widening in SYCL stride arithmetic — scheduled for
  explicit (ptrdiff_t) casts that eliminate the suppression entirely (PR A).
- 14 misc-const-correctness in SYCL atomic_ref loops — fold into existing
  NOLINTBEGIN/NOLINTEND block (PR C).
- 13 bare readability-function-size in integer_adm.c — ADR-0278 non-compliant;
  scheduled for NOLINTBEGIN block consolidation (PR C).
- 11 readability-function-size on SYCL kernel entry-points — load-bearing
  per ADR-0141; no change planned.

Deliverables: research digest, ADR-0780 (Proposed), changelog fragment.
Follow-up refactor PRs A–C are independent.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(process): per-surface doc compliance audit — last 30 PRs (ADR-0848)

Audits 30 most recent merged commits (PRs #96–#174) for CLAUDE §12 r10
compliance: user-discoverable surface change must ship human-readable docs
in the same PR.

Score: 22/30 N/A (no surface change), 5/8 with surface changes compliant
(62.5 %). Three confirmed gaps:

- Gap A (HIGH): PR #47 (ADR-0726, Vulkan drop) — docs/backends/vulkan/overview.md,
  docs/metrics/features.md, and docs/development/build-flags.md still describe
  Vulkan as an operative backend. Tracked as T-DOC-VULKAN-STALE-POST-ADR0726.
- Gap B (MEDIUM): PR #87 (ADR-0749, VmafLegacyQualityRunner sunset) — no
  docs/development/deprecations.md entry and no migration note in
  docs/usage/python.md. Tracked as T-DOC-LEGACY-RUNNER-MISSING-DEPRECATION.
- Gap C (LOW): PR #135 (log format standardization) — removes "Error: " prefix
  from CUDA error messages with no doc note.

Three follow-up issues proposed (Issue A, B, C in Research-0848). No code
changes in this PR; audit-only per task brief.

Six deliverables (ADR-0108):
(1) Research-0848: docs/research/research-0848-per-surface-doc-compliance-audit-20260529.md
(2) Decision matrix in ADR-0848 §Alternatives considered
(3) no rebase-sensitive invariants
(4) Reproducer: git log --oneline origin/master | head -30 (auditable from commit list)
(5) changelog.d/changed/per-surface-doc-compliance-audit.md
(6) docs/rebase-notes.md entry added

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(state): close T-LEGACY-RUNNER-ANSNR-BROKEN + T-LEGACY-RUNNER-STUB-MISSING

Both rows referenced classes / imports that no longer exist on master:

1. T-LEGACY-RUNNER-ANSNR-BROKEN — AnsnrFeatureExtractor was deleted
   by PR #283 (merged at faede7a). The class is absent from
   compat/python-vmaf/core/feature_extractor.py on master tip
   d45d503. The Netflix golden assertions still cover the
   integer-path VMAF score (Rule #1 preserved).

2. T-LEGACY-RUNNER-STUB-MISSING-2026-05-29 — VmafLegacyQualityRunner
   import-time failure: the class was removed in ADR-0749 / PR #87
   and python/test/quality_runner_test.py was updated by the
   ADR-0749 sunset PR to drop the import (only a removed-comment
   placeholder at line 49 remains).

Moved both rows from Open to Recently closed with verified-on-master
reproducer commands.

No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- [x] **Research digest**: no digest needed: state.md hygiene
- [x] **Decision matrix**: no alternatives: only-one-way fix
- [x] **AGENTS.md invariant**: no rebase-sensitive invariants
- [x] **Reproducer**: `git show origin/master:compat/python-vmaf/core/feature_extractor.py | grep -c 'class AnsnrFeatureExtractor'` returns 0; `git show origin/master:python/test/quality_runner_test.py | grep -c '^from.*VmafLegacy'` returns 0.
- [x] **Changelog**: changelog.d/changed/state-md-drift-sweep-20260530.md
- [x] **Rebase notes**: no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(state): migrate 3 Vulkan rows to Recently closed (ADR-0726 supersession)

ADR-0726 (Vulkan backend dropped 2026-05-28, PR #47) structurally closes
three long-standing Vulkan Open rows by removing the affected code path
entirely:

- T-VK-1.4-BUMP — NVIDIA driver 595.71+ FP-contraction regression
- T-VK-CIEDE-F32-F64 — NVIDIA Vulkan f32/f64 ciede precision gap
- T-VK-VIF-1.4-RESIDUAL-ARC — Intel Arc A380 vif residual on Mesa-ANV

The entire `core/src/vulkan/`, `core/src/feature/vulkan/`, and
`core/include/libvmaf/libvmaf_vulkan.h` surface no longer exists on
master, so each row now has a verified empty-path reproducer. Native
CUDA / HIP / SYCL backends cover every vendor formerly served by Vulkan
(see ADR-0726 §Context).

Combined with the legacy-runner closures already in this PR
(T-LEGACY-RUNNER-ANSNR-BROKEN + T-LEGACY-RUNNER-STUB-MISSING-2026-05-29),
this brings the Open section from 14 → 11 rows.

Row counts (verified): Open 11 + Deferred 4 + Recently closed 137 +
Confirmed not-affected 1 = 153 total T- rows in tree.

* chore(changelog): fix concat awk boundary + consolidate perf/ fragments (ADR-0221)

Fix `concat-changelog-fragments.sh` --check/--write awk block-boundary: the
previous `/^## [^[]/` pattern terminated the Unreleased block on any `## Heading`
embedded inside a fragment (e.g. `## Added`, `## [perf] …`). Switch to
`/^## \[(Unreleased|[0-9])/` which matches only versioned-release headers and
the `[Unreleased]` sentinel, making `--check` deterministic. Tightened the
`--write` awk pass with the same fix.

Consolidate 32 fragments from non-standard `changelog.d/perf/` (27) and
`changelog.d/performance/` (5) into the recognised `changed/` section with
`perf-` filename prefix per ADR-0221 KaC convention.

Remove 3 duplicate stubs: `ort-run-stack-arrays-f3b.md`, `adm-pnorm-deferred-comment.md`,
`fixed/vmaf-tune-predictor-directory-corpus.md` (richer versions retained).
Rename `changed/fr-regressor-v3-namespace.md` → `…-adr-appendix.md` to
resolve filename collision with `added/fr-regressor-v3-namespace.md`.

Regenerate CHANGELOG.md Unreleased block via `--write`; `--check` exits 0.

no rebase impact: changelog.d/ + scripts/release/ only, no upstream C touched.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(bundle): add changelog fragment for docs/hygiene bundle PR

Adds changelog.d/changed/bundle-docs-hygiene.md summarising the
four source PRs: #127 (NOLINT audit), #216 (doc compliance audit),
#291 (state.md drift sweep), #233 (changelog concat fix).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Lusoris <lusoris@pm.me>
lusoris added a commit that referenced this pull request Jun 8, 2026
…MCP, vmaftune, copyright, error paths, headers, magic numbers, docs) (#858)

* fix(ubsan): silence 5 UBSan-flagged production-code UB sites

1. pdjson.c/h: change stack_top from size_t to ptrdiff_t so that the
   -1 sentinel is a well-defined signed value; update all (size_t)-1
   comparisons to -1 and add casts on the depth/size comparisons to
   keep sign-clean arithmetic.

2. motion_avx512.c (×3 scalar tails): cast uint16_t filter[] operands
   to uint32_t before accumulation; the sum of all five taps at max
   pixel values (≈65536×65535) overflows signed int in the C default
   arithmetic promotions.

3. vif_avx512.c (all 4 sites): cast loop counter i to int before
   subtracting fwidth_half; unsigned − signed promotes to unsigned and
   the subsequent signed-int assignment is implementation-defined when
   the result wraps.

4. adm_avx2.c (10 sites): replace the unsigned hex literal
   0xFFFFFFFFFFFFFFFF passed to _mm256_set_epi64x() with -1LL;
   the hex form overflows long long and is UB per C99 §6.4.4.1.

5. integer_adm.c (both init loops): cast (1u << (shift_flt[idx] - 1))
   to int32_t before assigning to int32_t add_bef_shift_flt[]; the
   1u<<31 wrap is intentional per ADR-0155 (Netflix#955) and is now
   an explicit implementation-defined conversion rather than UB.
   NOLINT annotations cite ADR-0155 inline.

Build: clean (1010/1010 targets). Tests: 87/87 fast suite pass.
No Netflix golden assertion values changed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(asan): null caller handles + unlock/destroy mutex on all error paths

ASan/LeakSan dangling-pointer UAF fixes across five init/create/destroy
functions in core/src/:

- vmaf_init (libvmaf.c): *vmaf is now NULLed on every failure path so
  the caller cannot read a freed pointer after a failed init.
- vmaf_feature_extractor_context_create (feature_extractor.c/.cpp):
  *fex_ctx NULLed on free_x/free_f labels and on the inline
  vmaf_fex_ctx_parse_options error path.
- vmaf_fex_ctx_pool_create (feature_extractor.c/.cpp): *pool NULLed on
  all failure paths; feature_extractor.c also gains a pthread_mutex_init
  return-value check (the .cpp already had it) and a free_fex_list label
  to match the new guard.
- vmaf_fex_ctx_pool_destroy (feature_extractor.c/.cpp):
  pthread_mutex_unlock + pthread_mutex_destroy now called before free(pool)
  per POSIX; freeing a locked mutex is UB and leaks glibc TSD resources.
- vmaf_feature_collector_init + feature_vector_init (feature_collector.c/
  .cpp): *feature_collector and *feature_vector NULLed on all failure
  paths.

All changes follow CERT MEM30-C. Local verify: meson test --suite=fast
87/87 pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ai): pin torchvision>=0.27.0 and add 13 missing script stubs

Three Python ABI / missing-script findings resolved:

1. ai/pyproject.toml: promote torchvision from a comment-explained
   implicit dep to a pinned direct dependency (>=0.27.0,<0.28.0).
   pytorch-lightning >= torchmetrics 1.9+ eagerly imports
   torchvision.transforms at module-load time; a stale torchvision 0.26.0
   wheel against torch 2.12.0 raises RuntimeError: operator
   torchvision::nms does not exist (not an ImportError), so pip's
   constraint resolution was the only reliable preventive fix.

2. ai/scripts/export_tiny_models.py: wrap the vmaf_train.models import
   (which triggers the pytorch_lightning -> torchvision chain) in a
   broad try/except so an ABI-mismatched venv produces a clear error
   message with the pip fix command instead of an opaque RuntimeError.

3. dev/Containerfile: add an explicit pip install torchvision>=0.27.0,<0.28.0
   step after the ai/ package install so freshly built container images
   never carry a stale torchvision wheel from a previous layer cache.

4. ai/scripts/: add 13 stub scripts that are referenced in docs/ADRs but
   were absent from the filesystem.  Each stub exits 0 with a short
   "not yet implemented" message and a pointer to the relevant doc.
   Stubs: build_calibration_set.py, eval_loso_fr_regressor_v2.py,
   external_benchmark_pvmaf.py, fetch_lsvq.py, gen_calibration.py,
   gen_dists_sq_placeholder_onnx.py, gen_mobilesal_placeholder_onnx.py,
   gen_ssimulacra2_eotf_lut.py, hdrsdr_vqa_to_corpus_jsonl.py,
   my_corpus_to_corpus_jsonl.py, quantize_int8.py,
   train_fr_regressor_v4.py, train_video_saliency_student.py.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(mcp): harden input validation — 5-finding wave (depth cap, frameNum, HTTP TypeError, n-cap, KeyError)

1. _nan_to_none: add depth cap (100 levels) via _nan_to_none_depth helper to
   prevent RecursionError on deeply nested JSON payloads from large vmaf runs.

2. _pick_worst_frames: wrap int(idx) in try/except (TypeError, ValueError) so
   non-numeric or dict frameNum values are logged and skipped instead of
   propagating and aborting describe_worst_frames.

3. http_transport._handle_score: add TypeError to the (ValueError, FileNotFoundError)
   catch so int(None) / int([...]) on non-integer width/height/bitdepth fields
   returns 400 instead of 500.

4. _call_tool describe_worst_frames: enforce schema maximum:32 on n server-side
   (raises ValueError for n < 1 or n > 32), not just in the JSON Schema hint.

5. _call_tool: extract dispatch into _call_tool_dispatch and wrap with
   KeyError -> ValueError conversion so missing required arguments produce a
   readable error message ("tool X missing required argument: 'ref'") instead
   of a bare KeyError.

22 new tests in test_mcp_hardening_wave1.py; no pre-existing test regressions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(vmaf-tune): 5 critical vmaftune bugs — proxy inputs, saliency height, saliency guard, profile source, test contract

Fix 1 (proxy.py): fr_regressor_v2.onnx was exported with two separate named
inputs ("features" [N,6] + "codec" [N,14]) matching FRRegressor.forward().
run_proxy was concatenating them into one 20-D tensor and feeding it as a
single input, so the codec port received nothing and fast-path production
mode produced wrong predictions. Now wires the two inputs separately for
two-input graphs; falls back to the legacy single-input path for older exports.

Fix 2 (saliency.py): compute_saliency_map crashed at runtime for any height
not divisible by 8 because the saliency_student_v1 encoder path requires
aligned tensor dims. Added an upfront ValueError with a clear hint showing
the next valid height rather than surfacing a cryptic onnxruntime error.

Fix 3 (cli.py): _run_recommend_saliency always invoked saliency_aware_encode
even when --saliency-aware was not set, because config=None caused the
function to silently create a default SaliencyConfig() and run the model.
Added an explicit guard: when saliency_aware is False, call run_encode directly.

Fix 4 (encoder_profile.py): build_encode_request raised AttributeError when
the profile "source" field was stored as a plain path string (written by
older vmaf-tune versions) instead of a metadata dict. Now normalises the
field to {"path": value} before accessing .get("path").

Fix 5 (test_fast.py): test_proxy_module_uses_lazy_import_seam was validating
the broken single-input 20-D contract. Updated to use a two-input fake session
(named "features" + "codec") and assert that run_proxy wires the inputs
separately, confirming the corrected behaviour from Fix 1.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(copyright): sweep license/copyright drift across 303 fork-original files

Three categories fixed:

1. ffmpeg-patches/0007 and 0008: remove "and Claude (Anthropic)" from 3
   copyright lines; per project_copyright_lusoris_only.md, Anthropic is not
   a rights holder — Lusoris-only attribution required.

2. 66 fork-original SIMD files (AVX2/AVX-512/NEON in
   core/src/feature/x86/ and core/src/feature/arm64/): add
   "Copyright 2026 Lusoris" as a second copyright line immediately after
   the existing Netflix line (dual notice; Netflix line preserved as these
   files may include upstream-derived code).

3. 234 fork-original C/H files: replace wrong SPDX identifier
   "BSD-3-Clause-Plus-Patent" with correct "BSD-2-Clause-Patent" to
   match the LICENSE root (BSD+Patent / SPDX: BSD-2-Clause-Patent).

Verified: scripts/ci/check-copyright.sh exits 0 on all changed files;
grep for BSD-3-Clause-Plus-Patent and "Anthropic" in copyright positions
returns 0 hits.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(compat,mcp,vmaf-tune): cross-version-compat — 4 findings

1. vmaf-tune _write_compare_profile_report: write JSON artifact when
   format='both' (previously only .html + .md were written, silently
   dropping the .json sidecar). Regression tests in
   test_format_both_json.py now pass.

2. MCP HTTP test fixture scoping: token_client fixture in
   test_http_transport.py leaked the removal of VMAFX_MCP_HTTP_NO_AUTH
   because os.environ.pop() was called inside a patch.dict that did not
   track that key, so the pop was not reverted on fixture teardown.
   Replaced with an explicit save/restore approach that prevents
   cross-test env contamination.

3. test_vmaf_version_handles_version_timeout: _vmaf_version is an async
   coroutine; the test now uses @pytest.mark.asyncio so pytest-asyncio
   drives the event loop instead of calling the coroutine object
   synchronously (Python 3.14+ would raise on await of a non-awaitable).

4. compat/python-vmaf/tools/misc.py: SourceFileLoader.load_module() is
   deprecated since Python 3.4 and scheduled for removal in Python 3.15;
   imp.load_source() was removed in Python 3.12. Migrated to the modern
   importlib.util.spec_from_file_location / exec_module path that works
   on all supported Python versions (3.8+).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(error-path): check and propagate return values at 5 error-path sites

1. cambi.c set_contrast_arrays: free partial allocs on OOM and propagate
   error at call site instead of silently ignoring -ENOMEM.
2. integer_motion.c flush: capture and propagate both
   vmaf_feature_collector_append_with_dict return values.
3. cuda/integer_motion_v2_cuda.c flush: capture and propagate
   vmaf_feature_collector_append return value.
4. sycl/integer_vif_sycl.cpp flush_fex_sycl: capture and propagate
   vmaf_sycl_queue_wait return value; close_fex_sycl (void)-casts it
   as the teardown path must continue regardless.
5. sycl/integer_adm_sycl.cpp: same pattern as integer_vif_sycl.cpp.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(model): 3 model-coverage fixes — vmaf_b log-level, tiny-model path guard, fr_regressor_v3 codec gate

Finding #1 (vmaf_phone_v0.6.1 absent): confirmed false positive — phone mode is a
score-transform on vmaf_v0.6.1.json, not a separate file. Docs and code already correct.

Finding #2 (vmaf_b_v0.6.3 spurious ERROR before fallback):
- model.c vmaf_model_load_from_path: demote "could not read model from path" from
  VMAF_LOG_LEVEL_ERROR to VMAF_LOG_LEVEL_WARNING. The CLI falls back to the collection
  loader when this call fails, so a bootstrap/collection JSON is not an error — it has
  a different top-level structure. The .pkl hard-error follow-up stays at ERROR since pkl
  is permanently unsupported.

Finding #3 (--tiny-model silently rejects .json path with -EBADMSG):
- configure_tiny_model (vmaf.c): add early extension check before vmaf_use_tiny_model.
  When the path does not end in ".onnx", emit a clear diagnostic explaining that sidecar
  .json files are loaded automatically alongside the .onnx, not passed directly. Previously
  the JSON bytes were scanned as protobuf by onnx_scan.c, producing the opaque -EBADMSG.

Finding #4 (fr_regressor_v3 out-of-range scores without --tiny-codec):
- dnn.h: add vmaf_dnn_is_codec_aware(ctx) public API.
- dnn_ctx.h: add vmaf_ctx_dnn_is_codec_aware bridge declaration.
- libvmaf.c: implement vmaf_ctx_dnn_is_codec_aware (checks sess, has_sidecar,
  codec_aware flag, and extra_in_width > 0).
- dnn_attach_api.c: implement vmaf_dnn_is_codec_aware public wrapper.
- configure_tiny_model (vmaf.c): after model load, if the model is codec-aware but no
  --tiny-codec / --tiny-preset / --tiny-crf was given, reject with a clear error message
  explaining that the conditioning block would contain only an "unknown" fallback slot.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(headers): resolve 5 public/internal header leak findings

1. picture_v2.h: rename include guard from reserved C identifier
   __VMAF_PICTURE_V2_H__ (double-underscore prefix, undefined behaviour
   per ISO C 7.1.3) to LIBVMAF_PICTURE_V2_H_.

2. All public headers: convert bare quoted includes (#include "foo.h" /
   #include "libvmaf/foo.h") to angle-bracket form (#include <libvmaf/foo.h>)
   across all 9 affected installed headers (libvmaf.h, picture.h, picture_v2.h,
   feature.h, model.h, dnn.h, libvmaf_cuda.h, libvmaf_hip.h, libvmaf_metal.h,
   libvmaf_mcp.h, libvmaf_sycl.h). Quoted-path includes only resolve when the
   build root is on the include path, breaking pkg-config consumers who only
   have the installed prefix.

3. picture.h VmafPicture: add INTERNAL banners to the ref and priv fields,
   clarifying they are managed by libvmaf and must not be accessed externally.

4. libvmaf.h VMAF_POOL_METHOD_NB: add __attribute__((deprecated)) on GCC/Clang
   so external callers see a build-time warning. Gate the attribute on
   !VMAF_BUILDING_LIBVMAF so internal TUs (output.c) that legitimately iterate
   [0, NB) are not affected. Inject -DVMAF_BUILDING_LIBVMAF into
   vmaf_cflags_common in core/src/meson.build.

5. vmaf_assert.h: remove from the install_headers() list in
   core/include/libvmaf/meson.build. The header exposes VMAF_ASSERT_DEBUG which
   is gated on the internal VMAF_DEBUG build flag and has no defined semantics
   for external consumers. Internal .c files continue to include it from the
   source tree. Add an INTERNAL comment banner to the header itself.

Build-verified: meson setup + ninja (1010/1010 targets clean) +
meson test --suite=fast (87/87 pass, 0 failures).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(constants): extract 5 magic numbers to named constants

- CAMBI_WINDOW_DIVISOR=375 and CAMBI_MIN_WIDTH_HEIGHT=216 centralised in
  cambi_internal.h; local per-backend defines (CAMBI_CUDA_MIN_WIDTH_HEIGHT,
  CAMBI_HIP_MIN_WIDTH_HEIGHT) and the bare 375 divisor removed from
  cambi.c, integer_cambi_cuda.c, and integer_cambi_hip.c.

- DNN_SIDECAR_JSON_MAX=1u<<20 added to model_loader.h; three guard sites
  in model_loader.c now reference it instead of bare bit-shifts /
  integer literals.

- FEATURE_VECTOR_INITIAL_CAPACITY=8u added to feature_collector.h; three
  literal 8s in feature_collector.c replaced.

- DNN_MIN_BIT_DEPTH=9 added to tensor_io.h; bpc guard in tensor_io.c
  (two sites) and dnn_api.c now reference it.

Build: 1010/1010 ninja targets; 87/87 fast tests pass; pre-commit clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(docs): round-4 docs-substance-audit — remove stale HIP scaffold note + mark vmafx CLI as planned + remove active Vulkan claims

Four targeted fixes from the round-4 docs-substance audit:

1. core/include/libvmaf/libvmaf_hip.h — remove the stale "Status: scaffold
   only. Every entry point returns -ENOSYS" Doxygen block. The HIP backend
   is fully implemented (ADR-0519 / ADR-0533 / ADR-0539); 21 feature
   extractors are registered and verified on AMD gfx hardware. Replace
   with an accurate status note pointing to the three unregistered legacy
   stubs and the no-HIP stubs.c contract.

2. docs/api/gpu.md — complete the vmaf_hip_import_state table entry: add
   the -ENOSYS return when built without HIP (matches the header Doxygen
   and stubs.c behaviour) alongside the already-documented -EINVAL case.

3. docs/usage/vmafx-cli.md — mark the vmafx symlink, --netflix-compat flag,
   and vmafx-* Python aliases as "planned — not yet implemented in master".
   Neither cli_parse.c nor the Python pyproject.toml entries have been
   updated yet (ADR-0690 / ADR-0696 specify the design). Add interim
   equivalents using the already-shipped --precision=max flag.

4. docs/usage/vmaf-tune.md — remove active Vulkan claims. The Vulkan
   backend was deleted in ADR-0726; --score-backend=vulkan no longer
   exists. Mark the Vulkan section REMOVED with a historical-reference
   notice; update six flag-table rows to drop vulkan from the accepted
   enum; fix the native-first-order example (vulkan was already absent
   from the auto probe order table at line 393).

No code changes; docs only. No golden assertions touched.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 12, 2026
…RF, AI scripts, MCP conformance) (#865)

* fix(hip,sycl): stale wave32 comment + Kahan IIR blur for ssimulacra2 (iter6-cross-backend-parity)

Three iter6 cross-backend-parity findings:

1. [critical — partial] float_adm_score.hip: wave32 code was already fixed
   in #859 (0a9dba8); correct the lingering stale comment that still read
   "FADM_WARPS_PER_BLOCK = 4 (64-lane warps)" — FADM_WARPS_PER_BLOCK is now
   256/32 = 8 slots (sized for the wave32 worst case).  No functional change.

2. [high — already fixed] float_ssim/ssim_score.hip + integer_psnr/psnr_score.hip:
   wave32 runtime-warpSize fixes landed in #859 alongside float_adm; nothing
   further to do here.

3. [high] ssimulacra2_sycl.cpp: add Kahan (compensated-summation) state
   tracking to the 3-pole recursive IIR blur kernel (launch_blur<PASS>).
   The IIR state (prev1_k) accumulated O(eps) rounding error per step;
   over 4K-tall frames this exceeded the 5e-5 cross-backend parity
   contract vs the CPU reference.  Each pole now carries a float comp_k
   compensation term; the standard Kahan pattern (y = candidate - comp;
   new_state = old_state + y; comp = (new_state - old_state) - y) bounds
   per-iteration error to O(eps^2) without fp64 (ADR-0220 compliant).

No Netflix golden-data assertions modified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ubsan): replace 0xFFFFFFFFFFFFFFFF hex literals with -1LL in adm_avx512.c

_mm512_set_epi64 takes long long (signed 64-bit) arguments. The literal
0xFFFFFFFFFFFFFFFF exceeds LLONG_MAX and is undefined behaviour under
strict UBSan. Replace with -1LL which has identical bit pattern and
correct type at both call sites (ADM_CM_THRESH_S_I_END macro lines 406
and 698).

Finding: iter6-ubsan-strict high [adm_avx512: 0xFFFF... overflows long long].
Note: motion_avx512.c and adm_avx2.c analogous fixes were already
present on master (PR #858 / PR #859 bundles).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(thread-safety): eliminate 2 TSAN races in threaded batch path (iter6-tsan-race-deep)

Finding #1 (high): Remove the racy write `fex->framesync = vmaf->framesync` in
`threaded_enqueue_one`.  `fex` is the *shared* registered VmafFeatureExtractor —
worker threads from previous frames may concurrently read its fields.  The write
is redundant: framesync is already propagated to every pool-slot copy by
`set_fex_framesync()` at registration time and by `ctx_pool_ensure_slot_ctx()`.

Finding #2 (high): Move the `vmaf->prev_ref` advance to BEFORE the enqueue call
in `threaded_read_pictures_batch`.  In the old order the main thread unreffed and
replaced `vmaf->prev_ref` after enqueue while the just-submitted worker still held
a live reference to the same underlying VmafRef*, creating a concurrent unref/write
on the same object without synchronisation.  Workers use `data.prev_ref` (an
independently refcounted snapshot) exclusively and never re-read `vmaf->prev_ref`
after enqueue, so moving the advance before enqueue is both safe and race-free.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(vmaf-tune): fix --no-bisect TypeError and recommend CRF selection strategy

Two bugs in the compare/recommend paths:

1. _encode_and_score() in bisect.py had encode_runner and score_runner as
   required keyword-only args (no defaults). The CRF-sweep caller in cli.py
   did not pass them, causing an unconditional TypeError on any
   compare --no-bisect --crf-sweep invocation. Fixed by giving both
   parameters a default of None, consistent with the existing decode_runner
   parameter and the run_encode/run_score runner=None semantics.

2. _smallest_passing_crf() in cli.py used `crf > cur[0]` to select the
   largest (most efficient) passing CRF, contradicting both the function
   name and the CLI help string ("find the smallest CRF whose VMAF >=
   --target-vmaf"). The correct strategy is the smallest (highest-quality)
   passing CRF. Fixed comparison to `crf < cur[0]` and updated the docstring
   to match.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ai-scripts): 3 iter6 runtime bugs — vulkan-device AttributeError, saliency parity always-fail, ensemble ONNX codec-dim mismatch

- collect_gpu_calibration_data.py: remove dead args.vulkan_device
  reference from devices dict (Vulkan removed per ADR-0726; no
  --vulkan-device argparse registration existed, causing AttributeError
  at runtime)
- validate_saliency_student.py: replace broken PT-reconstruction parity
  check with ORT-only sanity check when no PT state provided.
  do_constant_folding=True folds BN stats into conv weights at export
  so ONNX initializer names diverge from PT state_dict keys; the old
  code silently left 60 of 65 weights at random defaults and always
  failed. When pt_state is provided (trainer path) full PT<->ORT diff
  is still performed.
- model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.onnx: regenerate
  with codec_onehot=[batch,6] matching current CODEC_VOCAB (was [batch,14]
  from a 12-entry encoder_vocab + 2 norm dims; CODEC_VOCAB was later
  trimmed to 6). Regenerated via train_fr_regressor_v2_ensemble.py
  --smoke in vmaf-dev-mcp container. eval_probabilistic_proxy.py --smoke
  now passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): JSON-RPC parse-error conformance + deep-nesting guard

Two iter6 fuzz conformance findings fixed:

1. [high] Parse-error notification instead of error response (id=null):
   Add `_ParseErrorFilteredStdin` — a lazy async stdin wrapper that
   pre-validates each incoming line with
   `JSONRPCMessage.model_validate_json`. On failure it calls
   `_emit_parse_error` which writes
   `{"jsonrpc":"2.0","id":<recovered_or_null>,"error":{"code":-32700,
   "message":"Parse error"}}` to stdout synchronously, then drops the
   line so the mcp library never sees a bare Exception on the stream
   (the notification path is bypassed entirely).  `_run()` now passes
   `stdin=_ParseErrorFilteredStdin()` to `stdio_server`.

2. [high] 500-level deep nesting triggers recursion-limit exception:
   Add `_check_depth(obj, max_depth=50)` helper and call it at the top
   of `_call_tool_dispatch` before the tool dispatch.  Payloads exceeding
   50 nesting levels raise `ValueError`, which the mcp library converts
   to an isError=True tool result before the pydantic parser recurses.

Existing tests updated: the two `stdio_server`-patching tests in
`test_coverage_round2.py` now accept `**kwargs` so they tolerate the
new `stdin=` keyword argument forwarded by `_run()`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 12, 2026
…ace, vmaf-tune, AI scripts, MCP) (#873)

* fix(hip): correct stale comment in float_psnr_score.hip re wave32 shared-mem sizing

The file header stated "the shared memory array for warp partial sums is
sized for the HIP warp size (64)" — contradicting the actual code, which
already used FPSNR_MIN_WARP_SIZE=32 and FPSNR_WARPS_PER_BLOCK sized for
wave32 worst-case (8 slots), with runtime warpSize in all reduction loops.

The other three iter9-cross-backend-parity findings (HIP float_adm
FADM_WARP_SIZE=64 RDNA2 corruption, SYCL float_adm missing AIM CM stages
2b/3b, HIP float_adm missing AIM CM stages 2b/3b) were fixed in the
iter8 bundle (commit 32ec0aa, PR #871).  This corrects the remaining
stale-comment discrepancy so the file accurately documents its own
implementation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(asan): null caller handles in .cpp TUs after free on failure paths

The .c implementations of vmaf_fex_ctx_pool_create and
vmaf_feature_collector_init already set the caller's out-pointer to NULL
at the free_p: / free_fc: labels (added in bundle-F / iter-asan sweeps),
but the parallel .cpp translation units used by all GPU backends missed
the same assignment.

Changes:
- feature_extractor.cpp: add *pool = NULL at free_p: before fall-through
  to fail:, mirroring feature_extractor.c line 797.
- feature_collector.cpp: add *feature_collector = nullptr at free_fc:
  before fall-through to fail:, mirroring feature_collector.c line 265.

Without these, any caller that allocates and checks the return code may
still read back a freed pointer via the out-parameter between free() and
the NULL assignment at fail: (CERT MEM30-C, ASan/LeakSan UAF class).
The .c files and libvmaf.c were already correct; this closes the .cpp gap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(svm): tighten VMAF_SVM_MAX_AXIS_COUNT + cast nr_class_permutations to size_t

Two UBSan findings from iter9 fuzz-extended cell:

1. VMAF_SVM_MAX_AXIS_COUNT was (1 << 24) ~16.7M, which allowed
   nr_class*(nr_class-1)/2 to overflow signed 32-bit arithmetic before
   the result was assigned to the size_t nr_class_permutations variable.
   Tighten the bound to 46340 (floor(sqrt(INT_MAX))) so the product is
   safe even without a cast.

2. Add explicit (size_t) casts before the multiply on the permutation
   expression to keep all arithmetic in the unsigned 64-bit domain,
   matching the signed-overflow fix pattern used elsewhere in this file.

All 9 SVM parser tests and 3 multiclass tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(thread-safety): iter9-tsan-race-deep — framesync pool snapshot + atomic batch err

Finding #1 (high): ctx_pool_ensure_slot_ctx read fex->framesync from the shared
registered VmafFeatureExtractor struct without holding any lock protecting that
field.  A concurrent set_fex_framesync() call on the main thread (writing to the
same struct) created a data race visible to TSan.  Fix: snapshot fex->framesync
once in vmaf_fex_ctx_pool_aquire while the pool lock is already held, then pass
the captured VmafFrameSyncContext* down through ctx_pool_claim_slot into
ctx_pool_ensure_slot_ctx.  This also corrects a latent functional bug: the global
fex->framesync was never written by set_fex_framesync (which only writes to deep
copies), so FRAME_SYNC extractors acquired through the pool would have received
framesync=NULL before this fix.  Both feature_extractor.c and feature_extractor.cpp
twins updated identically.

Finding #2 (high): struct ThreadDataBatch.err was a plain int written by the worker
via multiple sequential atomic_store sites and read by the caller as the function
return value.  Changing the field to _Atomic int and replacing all writes/reads with
atomic_store/atomic_load eliminates any TSan data-race report on this field should a
future code path observe f->err directly rather than through the return-value path,
and documents the worker-owned write semantics explicitly.

All 88 fast-suite tests pass; test_thread_safety_batch, test_thread_pool, and
test_feature_extractor pass individually.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(vmaftune): 2 iter9 bugs — entrypoint chown + saliency %32 padding

Fixes two high-priority findings from iter9 cell
iter9-vmaftune-exhaustion.  Finding 1 (score.py integer_ prefix) was
already resolved in commit a6c4dff and is a no-op here.

Fix 1 — compare: bisect workers PermissionError on VMAFTUNE_WORKDIR
bind-mount (dev/scripts/dev-mcp-entrypoint.sh):
The host bind-mount source for VMAFTUNE_WORKDIR is typically owned by
root:root (mode 755) after the first `docker compose up`.  The vmaf
user running inside the container cannot write into it, causing every
bisect worker to fail with PermissionError.
Add `chown vmaf:vmaf "${VMAFTUNE_WORKDIR}"` immediately after the
existing `mkdir -p` so the vmaf user owns the directory at container
start regardless of host-side ownership.  The chown is guarded with
`|| true` so it does not abort the entrypoint on read-only hosts.

Fix 2 — recommend-saliency: ONNX crash on non-multiple-of-32 sources
(tools/vmaf-tune/src/vmaftune/saliency.py):
`saliency_student_v1` uses a UNet encoder-decoder with skip connections
that require H and W to be exact multiples of 32.  Sources whose
dimensions are not multiples of 32 cause an onnxruntime shape mismatch
inside the skip-connection concatenation, crashing inference.

Add `_pad_to_multiple(tensor, 32)` helper that zero-pads a
`[1, 3, H, W]` float32 tensor to the next multiple-of-32 boundary.
In `compute_saliency_map`, replace the previous hard rejection of
`height % 8 != 0` with a pad-before-run / crop-after-run pattern:
pad the tensor, run inference on the padded input, then crop the
`[1, 1, H_pad, W_pad]` output back to `[:, :, :orig_h, :orig_w]`
before accumulating into the temporal aggregator.  This makes
arbitrary source dimensions work correctly without callers needing to
pre-pad or pre-crop their YUV sources.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ai): fix 3 iter9 ai-scripts-runtime findings — VMAF_BIN sentinel, py3.14 dataclass crash, pytest-timeout parity

- test_e2e_frame_to_score.py: replace `Path(os.environ.get('VMAF_BIN', '')) or ...`
  with an explicit None-sentinel check so VMAF_BIN='' is treated as unset rather
  than resolving to Path('.') and executing CWD as the vmaf binary. [critical]

- ai/scripts/_script_bootstrap.py: remove `from __future__ import annotations`.
  All dataclass fields are concrete Path types; the future-import is unnecessary
  and triggers CPython gh-129861 (Python 3.14 dataclasses._is_type() crash) when
  the module is loaded via importlib.util before sys.modules pre-registration. [high]

- ai/AGENTS.md: document the invariant that any importlib.util caller must insert
  `sys.modules[spec.name] = module` between module_from_spec() and exec_module()
  to avoid the Python 3.14 regression. [high — invariant note]

- ai/pyproject.toml: add pytest-timeout>=0.5 to [dev] optional deps.

- dev/Containerfile: install ai[dev] instead of plain ai so pytest-timeout is
  baked into the container, closing the CI/container parity gap for --timeout=60.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): close iter9 path-traversal and NaN-JSON findings

- describe_model Step 1 now calls _allowed_roots() after resolve() and
  raises ValueError for any path outside an allowlisted root, mirroring
  _validate_path() exactly (was implicit-only; ../../outside-repo.json
  bypassed the guard).
- HTTP /v1/score now serialises the result via _dumps_strict() instead
  of bare json.dumps(), converting NaN/Infinity to null and producing
  RFC 8259-compliant output (json.dumps allow_nan=True is the default,
  emitting bare NaN tokens that are not valid JSON).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 12, 2026
… fuzz, TSan/OrtEnv, vmaf-tune, AI scripts, MCP batch) (#875)

* fix(sycl): promote ADM per-scale normalization intermediates to double

On Intel Arc A380 (no native fp64 device), the per-scale normalization
in conclude_adm_cm and conclude_adm_csf_den used float f_accum and
float *result, causing rounding error that SVM amplified past the 5e-5
final-score threshold (iter10 cross-backend-parity finding [high]).

Both functions are host-side (no device-kernel code); the promotion to
double has zero impact on fp64-less device kernels. The call-site
variables num_scale and den_scale are likewise promoted to double.

HIP findings (wavefront_reduce_i64 carry bug and MS_WARP_SIZE=64 on
wave32 hardware) were already resolved in the worktree base by PR #850
and are confirmed absent: vif_statistics.hip uses atomicadd_accums, and
motion_score.hip uses runtime warpSize throughout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pool): destroy per-entry condvar in vmaf_fex_ctx_pool_destroy

vmaf_fex_ctx_pool_destroy() iterated all fex_list entries and freed
ctx_list but never called pthread_cond_destroy on the per-entry condvar
(pool->fex_list[i].full) that was initialised in get_fex_list_entry()
/ ctx_pool_alloc_slot().  POSIX requires destroy before the containing
memory is freed; omitting it leaks POSIX TSD resources on glibc and is
reported by ASan/LeakSan as a condvar-resource leak.

Add pthread_cond_destroy(&pool->fex_list[i].full) in the i-loop after
free(pool->fex_list[i].ctx_list) in both the C and C++ translation
units (the C++ file is the one compiled by meson; the C file is kept
in sync for readability).

The pool-mutex unlock+destroy was already present.  The libvmaf.c
*vmaf=NULL dangling-pointer fix (finding 2) was already applied in a
prior commit; no change needed there.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: round-4 audit bundle — UBSan, ASan, ABI, MCP, vmaf-tune, copyright, error paths, headers, magic numbers, docs (#858)

Cherry-pick of ebbcca3 with conflict resolution (motion_avx512.c uint32_t casts,
Containerfile ai[dev] extras, server.py iterative _nan_to_none, test imports).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(fuzz): cap Y4M frame dimensions to prevent unbounded malloc (T-FUZZ-Y4M-OOM)

Add Y4M_MAX_FRAME_PIXELS (64 Mpixels) guard in y4m_input_open_impl,
inserted after the existing sign check and before the chroma-format
dispatch. Attacker-controlled W/H values from the Y4M header can no
longer drive malloc with an unbounded size. The same guard also
eliminates the signed-integer overflow path in y4m_convert_411_422jpeg
for oversized frames.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tsan): iter10 — destroyed guards + OrtEnv singleton + model-list snapshot

Three TSan/UAF fixes identified in iter10-tsan-race-deep:

[critical] feature_collector.c: vmaf_feature_collector_set_aggregate,
vmaf_feature_collector_get_aggregate, vmaf_feature_collector_mount_model,
and vmaf_feature_collector_unmount_model were missing the `destroyed` flag
guard that the other entry points (append, get_score, find) already carry.
A worker thread racing destroy() could access freed aggregate_vector or
models memory after the lock was released by destroy(). Fix: add the
standard `if (feature_collector->destroyed) { unlock; return -ENODEV; }`
pattern immediately after lock acquire in all four entry points.

[high] ort_backend.c: each vmaf_ort_open call created a fresh OrtEnv via
sess->api->CreateEnv, spawning new ORT-internal background threads. ORT
documents OrtEnv as a process-wide resource; concurrent CreateEnv calls
race inside ORT's thread-pool initialisation. Fix: replace per-session
OrtEnv with a file-static singleton (g_ort_env) initialised exactly once
via pthread_once. The singleton is never released (process lifetime per
ORT recommended usage). Remove sess->env field and the ReleaseEnv call
from vmaf_ort_close.

[high] feature_collector.c: feature_collector_run_model_predict snapshotted
only model_iter->next before the lock drop (round-5 fix), but the node that
the pre-snapshotted next pointer points to could itself be unmounted and
freed by a concurrent vmaf_feature_collector_unmount_model between lock
releases. Fix: snapshot the full VmafModel* list into a stack-allocated
array of FEATURE_COLLECTOR_MAX_MODELS (32) entries while the lock is held
before the first unlock, then iterate the snapshot without re-dereferencing
any linked-list pointers. VmafModel lifetime is caller-managed and outlives
the predict pass.

Build: 88/88 fast tests pass (CPU-only build, enable_dnn=disabled).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(vmaf-tune): correct fast-verify bitrate denominator and report NaN sentinel

- _build_fast_encode_runner: replaced wall-clock encode_time_ms denominator
  with clip duration derived from raw-YUV file size and frame geometry
  (width × height × bpp / framerate). Using encoder wall-clock time as the
  denominator inflated or deflated observed_kbps depending on encode speed
  rather than content duration.

- _run_report: changed LadderSample and LadderRung bitrate_kbps/vmaf
  construction from `or 0.0` (silently coercing null to zero) to
  `float('nan')` when the JSON key is absent or null. The renderer already
  gates on _is_missing() / _finite_values(), so NaN propagates as an em-dash
  gap-marker instead of a misleading 0 kbps / 0 VMAF entry.

Fixes iter10-vmaftune-exhaustion findings #1 and #2.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ai-scripts): remove stale Vulkan references post-ADR-0726

- collect_gpu_calibration_data.py: fix --smoke docstring that still said
  "Vulkan-only"; update to reflect CUDA-only smoke mode (ADR-0726).
- Five extraction scripts (extract_k150k_features, konvid_to_full_features,
  extract_ugc_features, bvi_dvc_to_full_features, konvid_to_vmaf_pairs):
  remove --no_vulkan from subprocess command lists; the flag no longer
  exists in the post-ADR-0726 vmaf binary.
- cross_backend_parity_gate.py: drop "vulkan" from BACKEND_SUFFIX,
  BACKEND_DEVICE_FLAG, BACKEND_DEFAULT_DEVICE, BACKEND_EXTRACTOR_ALIASES,
  --vulkan-device arg, and devices dict; fix stale help text references.
- export_transnet_v2_placeholder.py / export_fastdvdnet_pre_placeholder.py:
  gate _export() call behind the --no-registry guard so --no-registry is a
  true dry-run that does not attempt writes to the default read-only path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): emit -32600 for batch requests, -32700 for parse errors on stdio transport

JSON-RPC 2.0 §6 requires servers that do not support batch requests to respond
with -32600 (Invalid Request), not -32700 (Parse error), when a JSON array
arrives on stdin — the payload is syntactically valid, it is the request type
that is unsupported. Previously _emit_parse_error always emitted -32700
regardless of whether the incoming line was malformed JSON or a valid-JSON-but-
array batch request.

The fix detects when the successfully-parsed value is a list and switches the
error code and message to -32600 / "Invalid Request" accordingly. All malformed-
JSON paths retain -32700. The _ParseErrorFilteredStdin wrapper that calls
_emit_parse_error is unchanged; all 407 existing tests continue to pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 12, 2026
…#877)

Dockerfile and Dockerfile.ffmpeg both declared ARG FFMPEG_TAG=n8.1 (the
bare minor-version tag) rather than n8.1.1 (the patch-level tag our
16-patch stack targets). The difference matters: n8.1 and n8.1.1 differ
by at least one commit ("Bump micro for 8.1.1"), and patch context is
validated against exact file content — any line-count divergence silently
shifts hunk offsets and can cascade into apply failures. We hit this
class of failure twice this session.

Also fix a stale context line in patch 0016 that was causing an apply
failure when the full series was replayed against a pristine n8.1.1
checkout. Hunk #1 of 0016-libvmaf-wire-score-fmt-on-all-vmaf-filters.patch
referenced "gpumask bit 1 disables" — the wording before patch 0014
(0014-libvmaf-expose-cpumask-and-gpumask) rewrote it to "gpumask bit 0
(value 1)". With the corrected context, all 16 patches apply cleanly via
git am --3way against a fresh n8.1.1 clone.

Series replay result: 16/16 patches apply cleanly against n8.1.1.

release/8.1 branch is 13 commits ahead of n8.1.1 tag; none of those
commits touch libavfilter/vf_libvmaf.c (our integration point), so no
action needed on the branch-vs-tag divergence.

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 27, 2026
…t + vmaf-tune stderr + eval shape guard (T-BUGHUNT-MCP-2026-06-27)

Bug-hunt sweep (mcp subsystem), 4 fixed / 2 already-fixed-and-skipped:

- mcp #1 (high): the Go cmd/vmafx-mcp streamable-HTTP transport had no auth,
  no body limit, and bound all interfaces — the Python ADR-0967 hardening was
  never ported. New cmd/vmafx-mcp/http_security.go adds a bearer-token
  middleware (VMAFX_MCP_HTTP_TOKEN, crypto/subtle constant-time compare,
  VMAFX_MCP_HTTP_NO_AUTH=1 opt-out, refuse-all 401 when neither is set), a
  4 MiB body limit (http.MaxBytesReader + Content-Length pre-flight -> 413),
  and a loopback-only default bind (VMAFX_MCP_HTTP_BIND, default 127.0.0.1,
  applied when mcp.http.addr has no host); wired into main.go runMCPTransport.
- mcp #3 (med): unify the score-precision default to "legacy" (%.6f, the
  C-CLI default per ADR-0119). The Python HTTP /v1/score path and the Go
  direct-cgo->subprocess fallback both defaulted to "17", diverging from the
  stdio path and the documented default.
- mcp #4 (low): add the pred/target shape-mismatch guard to the Go
  eval_model_on_split inline script (parity with Python _eval_model_on_split).
- mcp #5 (low): the three Go vmaf-tune wrappers now fold subprocess stderr
  into the error via a shared runVmafTune helper
  ("vmaf-tune <sub> exited <rc>: <stderr>"), instead of discarding it with
  exec.Output().

Skipped (already fixed in-tree): mcp #2 (subsample forwarding on
vmaf_score_encoded — scoreExtras.subsample) and mcp #6 (HTTP /v1/score strict
serializer — http_transport.py uses _dumps_strict).

Golden safety: no Netflix golden assertAlmostEqual value touched (MCP servers
only).

Reproducer:
  go test ./cmd/vmafx-mcp/        # TestSecurityMiddleware*, TestApplyBindHost,
                                  # TestRunVmafTune_*
  gosec ./cmd/vmafx-mcp/          # 0 issues
  PYTHONPATH=mcp-server/vmaf-mcp/src python -m pytest mcp-server/vmaf-mcp/tests -q
                                  # 451 passed, 2 skipped
                                  # (incl. test_score_precision_defaults_to_legacy)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 27, 2026
…t + vmaf-tune stderr + eval shape guard (T-BUGHUNT-MCP-2026-06-27)

Bug-hunt sweep (mcp subsystem), 4 fixed / 2 already-fixed-and-skipped:

- mcp #1 (high): the Go cmd/vmafx-mcp streamable-HTTP transport had no auth,
  no body limit, and bound all interfaces — the Python ADR-0967 hardening was
  never ported. New cmd/vmafx-mcp/http_security.go adds a bearer-token
  middleware (VMAFX_MCP_HTTP_TOKEN, crypto/subtle constant-time compare,
  VMAFX_MCP_HTTP_NO_AUTH=1 opt-out, refuse-all 401 when neither is set), a
  4 MiB body limit (http.MaxBytesReader + Content-Length pre-flight -> 413),
  and a loopback-only default bind (VMAFX_MCP_HTTP_BIND, default 127.0.0.1,
  applied when mcp.http.addr has no host); wired into main.go runMCPTransport.
- mcp #3 (med): unify the score-precision default to "legacy" (%.6f, the
  C-CLI default per ADR-0119). The Python HTTP /v1/score path and the Go
  direct-cgo->subprocess fallback both defaulted to "17", diverging from the
  stdio path and the documented default.
- mcp #4 (low): add the pred/target shape-mismatch guard to the Go
  eval_model_on_split inline script (parity with Python _eval_model_on_split).
- mcp #5 (low): the three Go vmaf-tune wrappers now fold subprocess stderr
  into the error via a shared runVmafTune helper
  ("vmaf-tune <sub> exited <rc>: <stderr>"), instead of discarding it with
  exec.Output().

Skipped (already fixed in-tree): mcp #2 (subsample forwarding on
vmaf_score_encoded — scoreExtras.subsample) and mcp #6 (HTTP /v1/score strict
serializer — http_transport.py uses _dumps_strict).

Golden safety: no Netflix golden assertAlmostEqual value touched (MCP servers
only).

Reproducer:
  go test ./cmd/vmafx-mcp/        # TestSecurityMiddleware*, TestApplyBindHost,
                                  # TestRunVmafTune_*
  gosec ./cmd/vmafx-mcp/          # 0 issues
  PYTHONPATH=mcp-server/vmaf-mcp/src python -m pytest mcp-server/vmaf-mcp/tests -q
                                  # 451 passed, 2 skipped
                                  # (incl. test_score_precision_defaults_to_legacy)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 27, 2026
…t + vmaf-tune stderr + eval shape guard (T-BUGHUNT-MCP-2026-06-27) (#1046)

Bug-hunt sweep (mcp subsystem), 4 fixed / 2 already-fixed-and-skipped:

- mcp #1 (high): the Go cmd/vmafx-mcp streamable-HTTP transport had no auth,
  no body limit, and bound all interfaces — the Python ADR-0967 hardening was
  never ported. New cmd/vmafx-mcp/http_security.go adds a bearer-token
  middleware (VMAFX_MCP_HTTP_TOKEN, crypto/subtle constant-time compare,
  VMAFX_MCP_HTTP_NO_AUTH=1 opt-out, refuse-all 401 when neither is set), a
  4 MiB body limit (http.MaxBytesReader + Content-Length pre-flight -> 413),
  and a loopback-only default bind (VMAFX_MCP_HTTP_BIND, default 127.0.0.1,
  applied when mcp.http.addr has no host); wired into main.go runMCPTransport.
- mcp #3 (med): unify the score-precision default to "legacy" (%.6f, the
  C-CLI default per ADR-0119). The Python HTTP /v1/score path and the Go
  direct-cgo->subprocess fallback both defaulted to "17", diverging from the
  stdio path and the documented default.
- mcp #4 (low): add the pred/target shape-mismatch guard to the Go
  eval_model_on_split inline script (parity with Python _eval_model_on_split).
- mcp #5 (low): the three Go vmaf-tune wrappers now fold subprocess stderr
  into the error via a shared runVmafTune helper
  ("vmaf-tune <sub> exited <rc>: <stderr>"), instead of discarding it with
  exec.Output().

Skipped (already fixed in-tree): mcp #2 (subsample forwarding on
vmaf_score_encoded — scoreExtras.subsample) and mcp #6 (HTTP /v1/score strict
serializer — http_transport.py uses _dumps_strict).

Golden safety: no Netflix golden assertAlmostEqual value touched (MCP servers
only).

Reproducer:
  go test ./cmd/vmafx-mcp/        # TestSecurityMiddleware*, TestApplyBindHost,
                                  # TestRunVmafTune_*
  gosec ./cmd/vmafx-mcp/          # 0 issues
  PYTHONPATH=mcp-server/vmaf-mcp/src python -m pytest mcp-server/vmaf-mcp/tests -q
                                  # 451 passed, 2 skipped
                                  # (incl. test_score_precision_defaults_to_legacy)

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 27, 2026
… golden drift

Two regressions rode into master via the 2026-06-27 admin-merge batch (per-PR
CI was bypassed). Both are required-check failures.

1. Sanitizers (TSan) link failure. The R2-9 OOM-injection test
   `test_gpu_dispatch_env_oom.cpp` replaces the global `operator new` /
   `operator delete`; under TSan/ASan/MSan the sanitizer runtimes interpose
   their own, so lld reports `duplicate symbol: operator new(unsigned long)`.
   Guard the override and self-skip the test under
   `__SANITIZE_THREAD__`/`__SANITIZE_ADDRESS__`/`__SANITIZE_MEMORY__` and the
   Clang `__has_feature` equivalents. The R2-9 slot-poisoning check still runs
   in every non-sanitized suite, so coverage is retained.

2. ARM golden drift. PR #1060's aarch64-only `-ffp-contract=off` guard on the
   scalar `adm_dwt2_s` / `adm_dwt2_lo_s` shifted the akiyo `disable_enhn_gain`
   ADM score on the ARM build matrix (88.030463 -> 88.030322), failing the
   `vmafexec_test.py` golden assertions. x86 D24 was unaffected (the guard was
   `ARCH_AARCH64`-gated), which is why it was not caught pre-merge. Per global
   rule #1 (immutable golden; fix the code, never the assertion) the scalar
   guard is reverted: `adm_tools.c` returns to its pre-#1060 FMA-default scalar
   arithmetic on aarch64 (byte-identical to 2d2f452^), restoring 88.030463.
   The now-stale `test_float_adm_simd.c` parity test and its meson
   registration are removed. The NEON-vs-scalar parity gap returns to the
   undispatched/FMA-free state of ADR-1057; T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06
   stays Open for a correct re-attempt (make the NEON side match the scalar FMA
   reference; validate against the full ARM quality suite, not only x86 golden).

Docs: ADR-1057 Update (2026-06-27) records the reverted follow-up and the
lesson; docs/state.md adds the Recently-closed row (and fixes a stale ADR-1057
link); changelog fragment replaced; arm64 AGENTS.md + research digest restored.

Verified: CPU build OK; fast suite 105/105; `test_gpu_dispatch_env_oom` builds
and passes (override active in the non-sanitized config);
`git diff 2d2f452^ -- core/src/feature/adm_tools.c` empty; clang-format,
assertion-density, check-copyright clean on touched files.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 27, 2026
… golden drift (#1063)

* fix(ci): repair master red — TSan operator-new dup + revert #1060 ARM golden drift

Two regressions rode into master via the 2026-06-27 admin-merge batch (per-PR
CI was bypassed). Both are required-check failures.

1. Sanitizers (TSan) link failure. The R2-9 OOM-injection test
   `test_gpu_dispatch_env_oom.cpp` replaces the global `operator new` /
   `operator delete`; under TSan/ASan/MSan the sanitizer runtimes interpose
   their own, so lld reports `duplicate symbol: operator new(unsigned long)`.
   Guard the override and self-skip the test under
   `__SANITIZE_THREAD__`/`__SANITIZE_ADDRESS__`/`__SANITIZE_MEMORY__` and the
   Clang `__has_feature` equivalents. The R2-9 slot-poisoning check still runs
   in every non-sanitized suite, so coverage is retained.

2. ARM golden drift. PR #1060's aarch64-only `-ffp-contract=off` guard on the
   scalar `adm_dwt2_s` / `adm_dwt2_lo_s` shifted the akiyo `disable_enhn_gain`
   ADM score on the ARM build matrix (88.030463 -> 88.030322), failing the
   `vmafexec_test.py` golden assertions. x86 D24 was unaffected (the guard was
   `ARCH_AARCH64`-gated), which is why it was not caught pre-merge. Per global
   rule #1 (immutable golden; fix the code, never the assertion) the scalar
   guard is reverted: `adm_tools.c` returns to its pre-#1060 FMA-default scalar
   arithmetic on aarch64 (byte-identical to 2d2f452^), restoring 88.030463.
   The now-stale `test_float_adm_simd.c` parity test and its meson
   registration are removed. The NEON-vs-scalar parity gap returns to the
   undispatched/FMA-free state of ADR-1057; T-NEON-FMA-FLOAT-ADM-DWT2-2026-06-06
   stays Open for a correct re-attempt (make the NEON side match the scalar FMA
   reference; validate against the full ARM quality suite, not only x86 golden).

Docs: ADR-1057 Update (2026-06-27) records the reverted follow-up and the
lesson; docs/state.md adds the Recently-closed row (and fixes a stale ADR-1057
link); changelog fragment replaced; arm64 AGENTS.md + research digest restored.

Verified: CPU build OK; fast suite 105/105; `test_gpu_dispatch_env_oom` builds
and passes (override active in the non-sanitized config);
`git diff 2d2f452^ -- core/src/feature/adm_tools.c` empty; clang-format,
assertion-density, check-copyright clean on touched files.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): Windows MSVC+CUDA D8021 — filter -Wno-write-strings via get_supported_arguments

The third required-check failure on master (pre-existing): the
`test_gpu_dispatch_env_oom` meson target passed the raw GCC/Clang flag
`-Wno-write-strings` unconditionally. MSVC `cl` parses `/W` as a warning-level
prefix, so it rejects the flag with
`D8021: invalid numeric argument '/Wno-write-strings'`, failing
`Build — Windows MSVC + CUDA (build only)`. Route the flag through
`cxx.get_supported_arguments('-Wno-write-strings')` so MSVC drops it while
GCC/Clang keep it (the flag is only needed to silence -Wwrite-strings from the
shared upstream `mu_assert` macro on the C++ TU).

A full grep confirms `-Wno-write-strings` was the only raw `-Wno-*` reaching the
compiler directly; `-ffp-contract=off` elsewhere is only a D9002 warning on MSVC
(ignored), and the lone `-Wl,-framework` is macOS-gated.

Verified: CPU build OK; fast suite 105/105; the flag still applies on Linux clang
(get_supported_arguments returns it). Completes the master-CI repair started in
this PR (Sanitizers-thread + D24 golden already green on the PR run).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): Windows MSVC C2440 — hold mu_assert messages in static char[] (supersede flag hack)

Follow-up to the get_supported_arguments(-Wno-write-strings) change: dropping the
flag on MSVC merely exposed the real, hard error it was masking —

  test_gpu_dispatch_env_oom.cpp(145): error C2440: 'return': cannot convert
  from 'const char [70]' to 'char *'

test.h's `mu_assert` does `return message;`. The harness is C-first, where a
string literal -> `char *` is legal; this is the ONLY test whose body is a C++
TU using mu_assert, and under MSVC /std:c++latest that conversion is a hard
C2440 (no flag silences an error). The other `.cpp` test executables either
build from `.c` bodies (test_dict/model/feature) or pull in library `.cpp`
sources with no mu_assert, so they are unaffected — this file is uniquely
exposed.

Fix: hold the two assert messages in `static char[]` buffers. A mutable array
decays to `char *` with no conversion (no C2440, no -Wwrite-strings), and static
storage keeps the returned pointer valid after the function returns. The
obsolete `-Wno-write-strings` cpp_args is removed. Compiles identically on
MSVC / GCC / Clang.

Verified: CPU build OK; fast suite green; test_gpu_dispatch_env_oom builds +
passes; clang-format clean. Completes the master-CI repair (all three required
failures: TSan dup-symbol, ARM golden drift, Windows MSVC C2440).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
lusoris added a commit that referenced this pull request Sep 26, 2026
vmaf_feature_extractor_context_close() unconditionally set is_closed=true
even when the underlying close() callback returned an error.  This broke
the retry contract end-to-end:

  1. A caller that catches a failed close() cannot retry: the next call to
     context_init() sees is_initialized=true (unchanged) and context_close()
     sees is_closed=true and returns 0 immediately — the extractor is locked
     in a half-torn-down state forever (blocker #1).

  2. After a partial GPU teardown the retained handles live in priv.  A
     subsequent context_destroy() frees priv unconditionally, releasing the
     memory while the handles still point into it (blocker #2).

Fix: only set is_closed when close() returns 0.  A non-zero return leaves
is_closed false so the caller can retry.  context_destroy() is unchanged —
callers remain responsible for calling close() before destroy().

Tests added to test_feature_extractor.c (device-free, CPU paths):
- test_close_retry_lifecycle: drives a synthetic extractor whose close()
  fails once then succeeds, verifying is_closed stays false on the first
  call and becomes true on the successful retry.

Closes: blocker #1 (retry lifecycle) and blocker #2 (destroy frees
retained handles after a failed close).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant