Skip to content

feat(cuda): runtime resolution-aware kernel variant dispatch (ADR-0753) - #91

Merged
lusoris merged 1 commit into
masterfrom
feat/cuda-resolution-dispatch-scaffold-20260529
May 29, 2026
Merged

lusoris merged 1 commit into
masterfrom
feat/cuda-resolution-dispatch-scaffold-20260529

Conversation

@lusoris

@lusoris lusoris commented May 29, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Introduces vmaf_cuda_workload_class(w, h) — a pure-C classifier mapping luma pixel count to WS_SMALL (< 720p) / WS_MEDIUM (720p–4K) / WS_LARGE (>= 4K).
  • Wires the first concrete consumer: adm_cm_device() in integer_adm_cuda.c selects adm_cm_line_kernel_8 (with __launch_bounds__(128,8)) at WS_MEDIUM and adm_cm_line_kernel_8_no_bounds otherwise — recovering the −9.3% 1080p gain without regressions at 576p or 4K.
  • Policy and design rationale documented in ADR-0753 + Research-0753.

Why DRAFT

Architecture review needed before retrofitting other kernels. The dispatch infrastructure is intentionally minimal; future consumers (filter1d WS_LARGE split, motion CPU fallback at WS_SMALL) follow the same pattern.

Policy table (ADR-0753)

Optimisation WS_SMALL WS_MEDIUM WS_LARGE
adm_cm __launch_bounds__(128,8) SKIP APPLY SKIP
filter1d __ldg + __launch_bounds SKIP APPLY APPLY
ms_ssim_decimate smem tiling SKIP SKIP SKIP

Reproducer / smoke test

# Build with CUDA
meson setup build -Denable_cuda=true && ninja -C build

# 576p — WS_SMALL path: should pick adm_cm_line_kernel_8_no_bounds
./build/tools/vmaf --feature adm --backend cuda \
  --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
  --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
  --width 576 --height 324 --pixel_format yuv420p --bitdepth 8

# 1080p — WS_MEDIUM path: should pick adm_cm_line_kernel_8
./build/tools/vmaf --feature adm --backend cuda \
  --reference python/test/resource/yuv/checkerboard_1920_1080_10_3_0_0.yuv \
  --distorted python/test/resource/yuv/checkerboard_1920_1080_10_3_0_1_0.yuv \
  --width 1920 --height 1080 --pixel_format yuv420p --bitdepth 10

# Parity gate (must pass at places=4)
make test-netflix-golden

Six deliverables (ADR-0108)

  • Research digest: docs/research/0753-cuda-resolution-aware-dispatch-design.md
  • Decision matrix: ADR-0753 ## Alternatives considered (compile-time, auto-tune, single-variant, per-SM-count)
  • AGENTS.md invariant: core/src/feature/cuda/AGENTS.md — "How to add a new resolution-aware variant"
  • Reproducer: build commands above
  • Changelog: changelog.d/added/cuda-resolution-aware-dispatch.md
  • Rebase notes: docs/rebase-notes.md entry for ADR-0753

Files touched

  • core/src/feature/cuda/resolution_dispatch.{h,c} — new, pure C
  • core/src/feature/cuda/integer_adm/adm_cm.cu — bounded + no-bounds macro variants
  • core/src/feature/cuda/integer_adm_cuda.c — include, struct field, init, dispatch
  • core/src/meson.build — register resolution_dispatch.c

🤖 Generated with Claude Code

lusoris added a commit that referenced this pull request May 29, 2026
…n dispatch (ADR-0753)

Extends PR #91's resolution-aware kernel dispatch scaffold to two additional
kernels per the ADR-0753 policy table. Three kernels now use the dispatch:

1. adm_cm_line_kernel_8 / _no_bounds (was already wired by PR #91)
2. filter1d_8_horizontal_kernel_2_17_9 / _no_bounds (new)
3. calculate_ssim_vert_combine / _no_bounds (new)

Policy (BOUNDED = __ldg + __launch_bounds, NO_BOUNDS = plain):
  - filter1d: BOUNDED at WS_MEDIUM + WS_LARGE; NO_BOUNDS at WS_SMALL
  - ssim_vert_combine: BOUNDED at WS_MEDIUM + WS_LARGE; NO_BOUNDS at WS_SMALL

Changes per file:
- integer_vif/filter1d.cu: FILTER1D_8_HORI_NO_BOUNDS macro + instantiation inside extern "C"
- integer_vif_cuda.c: VifStateCuda gains no_bounds function pointer; cuModuleGetFunction
  loads both; filter1d_8() branches on vmaf_cuda_workload_class(w,h)
- integer_ssim/ssim_score.cu: calculate_ssim_vert_combine_no_bounds inside extern "C";
  __ldg() loads retained in both variants
- integer_ssim_cuda.c: SsimStateCuda gains func_vert_no_bounds; submit_fex_cuda() branches
  on vmaf_cuda_workload_class(s->width, s->height)
- resolution_dispatch.h: policy table + kernel list comment updated
- docs/adr/0753: ssim_vert_combine row + three-consumer description
- docs/backends/cuda/overview.md: extended kernel dispatch table
- AGENTS.md: verified wirings table + rebase-sensitive invariants
- docs/rebase-notes.md + docs/state.md: extended scope entries

Correctness: CUDA vs CPU delta = 0.00e+00 on 576p pair at WS_SMALL (no_bounds
path). All 55 fast+gpu unit tests pass (including test_integer_vif_cpu_cuda_parity
and test_cuda_motion3_parity).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the feat/cuda-resolution-dispatch-scaffold-20260529 branch from 30cfaf5 to b8f2a79 Compare May 29, 2026 09:04
@lusoris

lusoris commented May 29, 2026

Copy link
Copy Markdown
Contributor Author

Extended to wire filter1d + ssim_vert_combine per the same policy. Now 3 kernels use the dispatch (adm_cm, filter1d, ssim_vert_combine).

…n dispatch (ADR-0753)

Extends PR #91's resolution-aware kernel dispatch scaffold to two additional
kernels per the ADR-0753 policy table. Three kernels now use the dispatch:

1. adm_cm_line_kernel_8 / _no_bounds (was already wired by PR #91)
2. filter1d_8_horizontal_kernel_2_17_9 / _no_bounds (new)
3. calculate_ssim_vert_combine / _no_bounds (new)

Policy (BOUNDED = __ldg + __launch_bounds, NO_BOUNDS = plain):
  - filter1d: BOUNDED at WS_MEDIUM + WS_LARGE; NO_BOUNDS at WS_SMALL
  - ssim_vert_combine: BOUNDED at WS_MEDIUM + WS_LARGE; NO_BOUNDS at WS_SMALL

Changes per file:
- integer_vif/filter1d.cu: FILTER1D_8_HORI_NO_BOUNDS macro + instantiation inside extern "C"
- integer_vif_cuda.c: VifStateCuda gains no_bounds function pointer; cuModuleGetFunction
  loads both; filter1d_8() branches on vmaf_cuda_workload_class(w,h)
- integer_ssim/ssim_score.cu: calculate_ssim_vert_combine_no_bounds inside extern "C";
  __ldg() loads retained in both variants
- integer_ssim_cuda.c: SsimStateCuda gains func_vert_no_bounds; submit_fex_cuda() branches
  on vmaf_cuda_workload_class(s->width, s->height)
- resolution_dispatch.h: policy table + kernel list comment updated
- docs/adr/0753: ssim_vert_combine row + three-consumer description
- docs/backends/cuda/overview.md: extended kernel dispatch table
- AGENTS.md: verified wirings table + rebase-sensitive invariants
- docs/rebase-notes.md + docs/state.md: extended scope entries

Correctness: CUDA vs CPU delta = 0.00e+00 on 576p pair at WS_SMALL (no_bounds
path). All 55 fast+gpu unit tests pass (including test_integer_vif_cpu_cuda_parity
and test_cuda_motion3_parity).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the feat/cuda-resolution-dispatch-scaffold-20260529 branch from b8f2a79 to 35a1fb6 Compare May 29, 2026 09:37
@lusoris
lusoris marked this pull request as ready for review May 29, 2026 09:37
@lusoris
lusoris merged commit fe5bcb2 into master May 29, 2026
40 of 62 checks passed
@lusoris
lusoris deleted the feat/cuda-resolution-dispatch-scaffold-20260529 branch May 29, 2026 09:37
@lusoris

lusoris commented May 29, 2026

Copy link
Copy Markdown
Contributor Author

Validation at 576p (WS_SMALL) — Research-0755 (PR #103):

Per-kernel nsys (CUPTI timestamps, 48 frames, src01 576×324):

Kernel Baseline avg NO_BOUNDS avg Delta
adm_cm_line_kernel_8 24.66 µs 21.90 µs −11.2%
filter1d_8_horizontal 18.69 µs 17.96 µs −3.9%
ssim_vert_combine 4.61 µs 4.71 µs +2.2% (noise)

Correctness: VMAF mean 76.6678 (both builds) — passes ADR-0214 places=2.

End-to-end wall time: Inconclusive. Five concurrent vmaf CUDA processes on the GPU during measurement produced ±18% run-to-run variance, masking any signal. An idle-GPU re-run is needed before quoting a wall-time number.

Verdict: NO_BOUNDS dispatch confirmed beneficial at SMALL for adm_cm (−11.2%) and filter1d_8_horizontal (−3.9%). Dispatch mechanism is correct. Recommend marking ready after idle-GPU end-to-end re-run.

Side note — master defect found: origin/master (70cb42a) has a committed merge conflict marker at core/src/feature/cuda/integer_vif_cuda.c line 336 (from commit 0c494cca05, post-merge-train sweep). This makes master CUDA-build-broken without a manual patch. Should be fixed as a hotfix before PR #91 merges.

Full digest: #103

lusoris added a commit that referenced this pull request May 29, 2026
…ch-0755)

nsys per-kernel profiling at 576x324 (WS_SMALL) comparing PR #91 dispatch
against master baseline:

- adm_cm_line_kernel_8_no_bounds: −11.2% vs bounded variant at 576p
- filter1d_8_horizontal_no_bounds: −3.9% vs bounded variant at 576p
- ssim_vert_combine_no_bounds: +2.2% (within noise)

End-to-end wall time inconclusive due to 5 concurrent vmaf CUDA processes on
the test GPU during measurement (+18.4% noise-dominated delta).

Correctness: VMAF 76.6678 mean (both builds), passes places=2 (ADR-0214).

Also flags pre-existing master defect: committed conflict marker in
core/src/feature/cuda/integer_vif_cuda.c line 336 (commit 0c494cc).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request May 29, 2026
…ch-0755) (#103)

nsys per-kernel profiling at 576x324 (WS_SMALL) comparing PR #91 dispatch
against master baseline:

- adm_cm_line_kernel_8_no_bounds: −11.2% vs bounded variant at 576p
- filter1d_8_horizontal_no_bounds: −3.9% vs bounded variant at 576p
- ssim_vert_combine_no_bounds: +2.2% (within noise)

End-to-end wall time inconclusive due to 5 concurrent vmaf CUDA processes on
the test GPU during measurement (+18.4% noise-dominated delta).

Correctness: VMAF 76.6678 mean (both builds), passes places=2 (ADR-0214).

Also flags pre-existing master defect: committed conflict marker in
core/src/feature/cuda/integer_vif_cuda.c line 336 (commit 0c494cc).

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
@lusoris
lusoris restored the feat/cuda-resolution-dispatch-scaffold-20260529 branch May 29, 2026 15:34
lusoris added a commit that referenced this pull request May 30, 2026
…tch) (#255)

The CUDA resolution-aware dispatch policy (ADR-0753) has been
implemented and merged via PR #91 (scaffold) + d6ef978 (filter1d
wiring). The ADR header still said Proposed; flipping to Accepted to
match deployed state.

Completes the residual ADR-state work from closed PR #214.

Co-authored-by: lusoris <lusoris@pm.me>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
Two Open rows in docs/state.md cited PRs that were CLOSED-not-merged
and flagged as follow-up by the PR #291 closing agent. Verified
against master tip bbcaa8d and master state of the underlying
issues:

1. T-CUDA-FILTER1D-RES-DISPATCH-CONFLICT-2026-05-29 — migrated from
   Open to Recently closed (superseded). The conflict markers only
   existed on the unmerged scaffold branch tip 35a1fb6. PR #91
   (merged 2026-05-29T09:37:48Z) landed ADR-0753 resolution-aware
   dispatch via the adm_cm_device() consumer without extending
   dispatch into filter1d_8(), so master never carried the
   build-failing scaffold variant. PR #214 (the planned
   conflict-marker cleanup) was CLOSED-not-merged 2026-05-30 when
   its base scaffold branch was abandoned. core/src/feature/cuda/
   integer_vif_cuda.c::filter1d_8() on master uses the clean
   unconditional cuLaunchKernel paths.

2. T-CPP23-READ-JSON-MODEL-PENDING-2026-05-29 — kept Open but the
   dead PR #215 citation removed. The C++23 Wave 8 conversion of
   core/src/read_json_model.c is still pending on master (still a
   .c source per core/src/meson.build:1578); PR #215 was
   CLOSED-not-merged 2026-05-30. Owner field rewritten to
   "Owner-driven; pending fresh PR per ADR-0846 Wave 8".

Scope-coordinated with DRAFT PR #291 (docs/state-md-drift-sync —
already migrates the 3 Vulkan rows + T-LEGACY-RUNNER-ANSNR-BROKEN
+ T-LEGACY-RUNNER-STUB-MISSING). Row 203
(T-LEGACY-RUNNER-STUB-MISSING-2026-05-29) cites closed PRs #213
and #181 as OPEN but is left untouched here since PR #291 already
rewrites it.

Net Open count -1; total T-row count unchanged (153).
No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- no digest needed: state.md hygiene
- no alternatives: only-one-way reconciliation
- no rebase-sensitive invariants
- Reproducer: `gh pr view 214 -R VMAFx/vmafx --json state,mergedAt`
  returns `{"state":"CLOSED","mergedAt":null}`;
  `grep -nE '^(<{7}|={7}|>{7})( |$)' core/src/feature/cuda/integer_vif_cuda.c`
  on master returns empty; `ls core/src/read_json_model.*` returns
  only `.c` and `.h`.
- Changelog: changelog.d/changed/state-md-closed-pr-row-sweep.md
- no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
Two Open rows in docs/state.md cited PRs that were CLOSED-not-merged
and flagged as follow-up by the PR #291 closing agent. Verified
against master tip bbcaa8d and master state of the underlying
issues:

1. T-CUDA-FILTER1D-RES-DISPATCH-CONFLICT-2026-05-29 — migrated from
   Open to Recently closed (superseded). The conflict markers only
   existed on the unmerged scaffold branch tip 35a1fb6. PR #91
   (merged 2026-05-29T09:37:48Z) landed ADR-0753 resolution-aware
   dispatch via the adm_cm_device() consumer without extending
   dispatch into filter1d_8(), so master never carried the
   build-failing scaffold variant. PR #214 (the planned
   conflict-marker cleanup) was CLOSED-not-merged 2026-05-30 when
   its base scaffold branch was abandoned. core/src/feature/cuda/
   integer_vif_cuda.c::filter1d_8() on master uses the clean
   unconditional cuLaunchKernel paths.

2. T-CPP23-READ-JSON-MODEL-PENDING-2026-05-29 — kept Open but the
   dead PR #215 citation removed. The C++23 Wave 8 conversion of
   core/src/read_json_model.c is still pending on master (still a
   .c source per core/src/meson.build:1578); PR #215 was
   CLOSED-not-merged 2026-05-30. Owner field rewritten to
   "Owner-driven; pending fresh PR per ADR-0846 Wave 8".

Scope-coordinated with DRAFT PR #291 (docs/state-md-drift-sync —
already migrates the 3 Vulkan rows + T-LEGACY-RUNNER-ANSNR-BROKEN
+ T-LEGACY-RUNNER-STUB-MISSING). Row 203
(T-LEGACY-RUNNER-STUB-MISSING-2026-05-29) cites closed PRs #213
and #181 as OPEN but is left untouched here since PR #291 already
rewrites it.

Net Open count -1; total T-row count unchanged (153).
No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- no digest needed: state.md hygiene
- no alternatives: only-one-way reconciliation
- no rebase-sensitive invariants
- Reproducer: `gh pr view 214 -R VMAFx/vmafx --json state,mergedAt`
  returns `{"state":"CLOSED","mergedAt":null}`;
  `grep -nE '^(<{7}|={7}|>{7})( |$)' core/src/feature/cuda/integer_vif_cuda.c`
  on master returns empty; `ls core/src/read_json_model.*` returns
  only `.c` and `.h`.
- Changelog: changelog.d/changed/state-md-closed-pr-row-sweep.md
- no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
Two Open rows in docs/state.md cited PRs that were CLOSED-not-merged
and flagged as follow-up by the PR #291 closing agent. Verified
against master tip bbcaa8d and master state of the underlying
issues:

1. T-CUDA-FILTER1D-RES-DISPATCH-CONFLICT-2026-05-29 — migrated from
   Open to Recently closed (superseded). The conflict markers only
   existed on the unmerged scaffold branch tip 35a1fb6. PR #91
   (merged 2026-05-29T09:37:48Z) landed ADR-0753 resolution-aware
   dispatch via the adm_cm_device() consumer without extending
   dispatch into filter1d_8(), so master never carried the
   build-failing scaffold variant. PR #214 (the planned
   conflict-marker cleanup) was CLOSED-not-merged 2026-05-30 when
   its base scaffold branch was abandoned. core/src/feature/cuda/
   integer_vif_cuda.c::filter1d_8() on master uses the clean
   unconditional cuLaunchKernel paths.

2. T-CPP23-READ-JSON-MODEL-PENDING-2026-05-29 — kept Open but the
   dead PR #215 citation removed. The C++23 Wave 8 conversion of
   core/src/read_json_model.c is still pending on master (still a
   .c source per core/src/meson.build:1578); PR #215 was
   CLOSED-not-merged 2026-05-30. Owner field rewritten to
   "Owner-driven; pending fresh PR per ADR-0846 Wave 8".

Scope-coordinated with DRAFT PR #291 (docs/state-md-drift-sync —
already migrates the 3 Vulkan rows + T-LEGACY-RUNNER-ANSNR-BROKEN
+ T-LEGACY-RUNNER-STUB-MISSING). Row 203
(T-LEGACY-RUNNER-STUB-MISSING-2026-05-29) cites closed PRs #213
and #181 as OPEN but is left untouched here since PR #291 already
rewrites it.

Net Open count -1; total T-row count unchanged (153).
No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- no digest needed: state.md hygiene
- no alternatives: only-one-way reconciliation
- no rebase-sensitive invariants
- Reproducer: `gh pr view 214 -R VMAFx/vmafx --json state,mergedAt`
  returns `{"state":"CLOSED","mergedAt":null}`;
  `grep -nE '^(<{7}|={7}|>{7})( |$)' core/src/feature/cuda/integer_vif_cuda.c`
  on master returns empty; `ls core/src/read_json_model.*` returns
  only `.c` and `.h`.
- Changelog: changelog.d/changed/state-md-closed-pr-row-sweep.md
- no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
lusoris added a commit that referenced this pull request Jun 3, 2026
…overage + #336 state.md row sweep + #383 README badges + #337 ADR-0865 ANSNR) (#529)

* docs(libvmaf): doxygen comments on 15 undocumented public C-API entry points

Round-2 follow-on to PR #302 (which closed five targeted gap-findings in
libvmaf.h / picture.h / dnn.h). This pass covers the public surfaces that
PR #302 left untouched, focusing on the headers the ffmpeg patch stack,
the upcoming Go/Rust bindings, and the embedded MCP server consume.

Entry points documented:

- feature.h (file was 100 % undocumented):
  - VmafFeatureDictionary (struct doc + ownership-transfer rules)
  - vmaf_feature_dictionary_set
  - vmaf_feature_dictionary_free

- model.h:
  - VmafModelFlags (enum + per-flag semantics)
  - VmafModelConfig (struct + per-field doc)
  - vmaf_model_load
  - vmaf_model_load_from_path
  - vmaf_model_feature_overload (incl. opts_dict ownership transfer)
  - vmaf_model_destroy (incl. do-not-destroy-after-collection-handoff)
  - VmafModelCollection (struct doc)
  - VmafModelCollectionScoreType (enum doc)
  - VmafModelCollectionScore (struct + per-field doc)
  - vmaf_model_collection_load
  - vmaf_model_collection_load_from_path
  - vmaf_model_collection_feature_overload
  - vmaf_model_collection_destroy

- dnn.h:
  - vmaf_dnn_session_close (pair-with-open contract)

Each block documents the negative-errno return convention, NULL-safety,
ownership-transfer semantics, and the destroy-pairing required to avoid
double-free of collection-owned sub-models. No semantic / no ABI change.

Cleanup pass (CLAUDE.md §12 r12 — touched-file lint-clean rule):
the three Netflix-copyright include guards (__VMAF_FEATURE_H__ /
__VMAF_MODEL_H__ / __VMAF_DNN_H__) trip clang-tidy's
bugprone-reserved-identifier check. Renaming them would diverge from
Netflix/vmaf master and break port-only upstream sync (CLAUDE.md §10),
so each #ifndef / #define gets an inline NOLINT citing the upstream-mirror
invariant — the exact pattern ADR-0278 endorses for load-bearing
upstream-parity identifiers.

Build + lint:
- meson setup build-cpu-doc core -Denable_cuda=false -Denable_sycl=false
- ninja -C build-cpu-doc (35 targets touched by header change; clean)
- clang-tidy -p build-cpu-doc on all 3 touched headers: 0 fork-local
  warnings (the 3 reserved-identifier warnings present on master are now
  NOLINT-cited; remaining warnings are in system headers and suppressed).
- pre-commit run --files <all 5 touched files>: green.

ADR-0108 deliverables:
- Research digest: no digest needed — trivial doc-only addition over
  Netflix-stable signatures already covered by the existing reference
  manual.
- Decision matrix: no alternatives needed — only-one-way fix; the
  ownership-transfer text matches the implementation in core/src/model.c
  and core/src/dict.c verbatim.
- AGENTS.md invariant note: no rebase-sensitive invariants — the doc
  text sits above unchanged upstream signatures; future merges from
  Netflix produce tractable 3-way merges. The NOLINT cites name the
  invariant they preserve (upstream-mirror include guards).
- Reproducer / smoke test: see PR body.
- Changelog fragment:
  changelog.d/added/libvmaf-public-header-doc-comments-round2.md.
- Rebase notes: docs/rebase-notes.md updated with a dedicated section.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* test(go): expand cmd/vmafx-{controller,server,mcp} coverage

Add unit tests for the lowest-coverage Go cmd/ subpackages identified by
the master-tip workflow audit (Section F). No behavior change.

Coverage deltas (against origin/master tip bbcaa8d):

  cmd/vmafx-controller        18.6% -> 32.4%  (+13.8 pp)
  cmd/vmafx-controller/nodes  80.7% -> 82.5%  (+1.8 pp)
  cmd/vmafx-server            27.5% -> 47.9%  (+20.4 pp)
  cmd/vmafx-mcp                3.5% -> 24.6%  (+21.1 pp)

New test files:

- cmd/vmafx-controller/main_extra_test.go
    405 method-not-allowed on /healthz, /readyz, /v1/score;
    400 invalid-JSON body; 500 scorer-error mapping via a stub vmaf
    binary; runHTTP graceful shutdown bounded by GracefulShutdownTimeout;
    envOr default+override; version().
- cmd/vmafx-controller/nodes/registry_edge_test.go
    Get(unknown), distinct-IDs-for-same-name contract pin,
    Heartbeat updates JobsRunning + advances LastHeartbeat
    (reaper eviction predicate), concurrent Register/Heartbeat under
    -race, defensive-copy assertion on All().
- cmd/vmafx-server/main_extra_test.go
    Same shape as the controller HTTP server tests; pins the PR #300
    bounded-timeout shutdown invariant.
- cmd/vmafx-mcp/impl_test.go
    All arg helpers (strArg, intArg, floatArg, boolArg, hasArg);
    every pure helper (classifySourceResolution, modelResolutionClass,
    resolutionMismatchWarning, inferBackendFromPayload,
    inferBackendFromSym, stripModelExt, toFFmpegPixfmt, pickWorstFrames,
    floatFromAny, roundF, truncate); representative handler error paths
    (handleProbeBackend missing/unknown backend, handleDescribeModel
    missing name, handleVmafScore invalid path / zero dimensions,
    handleCompareModels empty list). Pins the errorResult().IsError ==
    true invariant from project memory (MCP isError must be True so
    clients branch correctly).

Drive-by fix: .gitignore anchored the Go binary-ignore rules with a
leading slash so they only match the repo-root binaries, not any path
component sharing the name. Without this, untracked files like
cmd/vmafx-server/main_extra_test.go were silently ignored by `git add`.

Verification:

  go test -race -cover ./cmd/vmafx-controller/... \
                      ./cmd/vmafx-server/...     \
                      ./cmd/vmafx-mcp/...
  go vet  ./...

All green; no behavior change. The pre-existing
cmd/vmafx-operator/internal/controller failure (kubebuilder envtest
needs etcd binaries on PATH) is unrelated to this change and reproduces
on a clean origin/master checkout.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(state): reconcile 2 Open rows that cited CLOSED PRs (#214, #215)

Two Open rows in docs/state.md cited PRs that were CLOSED-not-merged
and flagged as follow-up by the PR #291 closing agent. Verified
against master tip bbcaa8d and master state of the underlying
issues:

1. T-CUDA-FILTER1D-RES-DISPATCH-CONFLICT-2026-05-29 — migrated from
   Open to Recently closed (superseded). The conflict markers only
   existed on the unmerged scaffold branch tip 35a1fb6. PR #91
   (merged 2026-05-29T09:37:48Z) landed ADR-0753 resolution-aware
   dispatch via the adm_cm_device() consumer without extending
   dispatch into filter1d_8(), so master never carried the
   build-failing scaffold variant. PR #214 (the planned
   conflict-marker cleanup) was CLOSED-not-merged 2026-05-30 when
   its base scaffold branch was abandoned. core/src/feature/cuda/
   integer_vif_cuda.c::filter1d_8() on master uses the clean
   unconditional cuLaunchKernel paths.

2. T-CPP23-READ-JSON-MODEL-PENDING-2026-05-29 — kept Open but the
   dead PR #215 citation removed. The C++23 Wave 8 conversion of
   core/src/read_json_model.c is still pending on master (still a
   .c source per core/src/meson.build:1578); PR #215 was
   CLOSED-not-merged 2026-05-30. Owner field rewritten to
   "Owner-driven; pending fresh PR per ADR-0846 Wave 8".

Scope-coordinated with DRAFT PR #291 (docs/state-md-drift-sync —
already migrates the 3 Vulkan rows + T-LEGACY-RUNNER-ANSNR-BROKEN
+ T-LEGACY-RUNNER-STUB-MISSING). Row 203
(T-LEGACY-RUNNER-STUB-MISSING-2026-05-29) cites closed PRs #213
and #181 as OPEN but is left untouched here since PR #291 already
rewrites it.

Net Open count -1; total T-row count unchanged (153).
No code changes — documentation cleanup only.

Deliverables (ADR-0108):
- no digest needed: state.md hygiene
- no alternatives: only-one-way reconciliation
- no rebase-sensitive invariants
- Reproducer: `gh pr view 214 -R VMAFx/vmafx --json state,mergedAt`
  returns `{"state":"CLOSED","mergedAt":null}`;
  `grep -nE '^(<{7}|={7}|>{7})( |$)' core/src/feature/cuda/integer_vif_cuda.c`
  on master returns empty; `ls core/src/read_json_model.*` returns
  only `.c` and `.h`.
- Changelog: changelog.d/changed/state-md-closed-pr-row-sweep.md
- no rebase impact: docs only

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(meta): README badge audit + Cargo / pyproject repo-metadata sweep

Audit + backfill the fork's repo-metadata surface so cargo / pip /
GitHub all advertise the canonical VMAFx/vmafx URLs:

- README.md: add Rust CI + Go CI workflow badges (both workflows ship
  on master but were not surfaced). All five pre-existing workflow
  badges already point at VMAFx/vmafx and reference real, active,
  master-green workflows; verified via the Actions API. License,
  Conventional Commits, OpenSSF Scorecard, ko-fi badges already
  present.
- Cargo.toml: add [workspace.package] with repository / homepage /
  documentation / license / authors. Both workspace members
  (bindings/rust/vmafx-sys, core/src/feature/rust/tad) switched to
  workspace-inherited metadata so URL drift is impossible across the
  Rust workspace. cargo metadata confirms both crates now expose the
  VMAFx/vmafx URLs.
- pyproject.toml (root, ai/, tools/vmaf-tune, tools/vmaf-roi-score,
  tools/ensemble-training-kit, dev-llm/, mcp-server/vmaf-mcp/): add
  [project.urls] with Homepage / Repository / Documentation / Issues /
  Changelog. All seven fork-authored projects now ship the same URL
  block; the package indexes (PyPI / internal) get a consistent
  repository link.
- deploy/helm/vmafx/Chart.yaml: already correct (home + sources
  already point at VMAFx/vmafx). No change needed.
- changelog.d/fixed/ + docs/rebase-notes.md: deliverables.

Coordination with PR #331 (rebrand sweep): #331 only touches the
line-1 copyright header of two of the seven pyproject files; this
PR adds a new [project.urls] block below — no merge conflict.

Deep-dive deliverables (ADR-0108):
- Research digest: no digest needed: trivial repo-metadata sweep.
- Decision matrix: no alternatives: only-one-way fix (URLs must
  match the rebrand target).
- AGENTS.md invariant: no rebase-sensitive invariants — fork-only
  metadata files, none mirror upstream Netflix.
- Reproducer: `cargo metadata --no-deps --format-version 1 | jq` +
  `python3 -c "import tomllib; tomllib.loads(open('pyproject.toml','rb').read().decode())"`.
- CHANGELOG fragment: changelog.d/fixed/readme-badges-metadata-audit.md.
- Rebase note: added to docs/rebase-notes.md (impact: none, fork-only).

* docs(adr): author ADR-0865 for ANSNR sunset (closes PR #38 ADR-0108 gap)

PR #38 (merged 2026-05-28) removed `float_ansnr` from the C backend
but cited `Parent ADR-0709` in its body — ADR-0709 is the Phase 4b
distributed-platform umbrella and contains zero ANSNR content.
PR #295 + PR #324 inherited the bad cite. No dedicated ANSNR-sunset
ADR existed in tree.

This change:
- Authors `docs/adr/0865-ansnr-sunset-pre-vmaf-metric-drop.md` as the
  missing parent ADR, back-dated to 2026-05-28 (PR #38 merge date) so
  the dependency chain (ADR-0865 -> PR #38 -> ADR-0749 Python sunset)
  is consistent.
- Documents the historical mis-cite in the new ADR's `## Notes`
  section so future readers landing on PR #38 can recover the trail.
  PR bodies on the remote are immutable merge-history and cannot be
  rewritten.
- Adds the index fragment + `_order.txt` row; regenerates
  `docs/adr/README.md` via `scripts/docs/concat-adr-index.sh`.
- Adds `docs/state.md` row (Updated note + Recently-closed entry).
- Adds `docs/rebase-notes.md` entry documenting the rebase invariant
  (upstream still ships `ansnr` extractors; fork must keep deleting).
- Adds `changelog.d/changed/ansnr-sunset-adr-authoring.md` fragment.

In-tree audit confirmed zero `ADR-0709` references mis-cite ANSNR —
all remaining tree-side `ADR-0709` cites correctly point at Phase 4b
distributed-platform content. No tree-side citation fix-up required.

Docs-only PR. No code changes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(bundle): add changelog fragment for doc-sweep bundle batch-1

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: Lusoris <lusoris@pm.me>
@lusoris
lusoris deleted the feat/cuda-resolution-dispatch-scaffold-20260529 branch June 4, 2026 08:32
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant