Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -1657,6 +1657,9 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
image.


- **The vendored Pelorus interop sources are re-vendored at the pelorus commit that clears their clang-tidy findings.** `scripts/sync-pelorus-interop.sh` pins `5f5614b0229d` (VMAFx/pelorus #78): the conformance test's long checks are split into helpers, blob headers are patched through `memcpy`, and each translation unit carries one cited `modernize-use-nullptr` block. No behaviour or ABI change (ABI 1.3).


- **Restore `adm_sum_cube_s_p3`, `adm_csf_den_scale_s_p3`, and `adm_cm_s_p3` fast-path
functions in `adm_tools.c` (ADR-0463 / BUG-048 B3).**
The specialized `adm_p_norm == 3.0` fast-paths eliminate all per-pixel `powf()`
Expand Down Expand Up @@ -3570,6 +3573,14 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
of [ADR-1317](docs/adr/1317-golden-gate-build-isolation.md).


- **Preserve explicit int8 paths in tiny-model DNN session loading.** When
`vmaf_dnn_session_open()` was called with an explicit `.int8.onnx` path,
`resolve_load_path()` lacked the `kInt8Suffix` early return present in
`dnn_attach_api.c`, causing it to append a redundant `.int8` suffix and derive
`<name>.int8.int8.onnx` before falling back to the fp32 path. The resolver now
checks `kInt8Suffix` upfront and preserves explicit int8 paths directly.


- Restored the section links that older pages and ADRs use into the CLI,
`vmaf_bench`, environment-variable and Getting started pages after their
rewrite (#1934, #1938): each former section name is a short heading that
Expand Down Expand Up @@ -4492,6 +4503,17 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
([licensing](docs/licensing.md#models)).


- **A model collection's score can be read more than once.**
`vmaf_score_at_index_model_collection()` failed with `-EINVAL` when the
frame had been scored before, and so did
`vmaf_score_pooled_model_collection()` over a range holding a frame already
scored per frame (the collector refused to write the members' scores of
that frame a second time, `feature "..." cannot be overwritten`). A frame
already predicted now returns its stored bootstrap scores, as a single
model's `vmaf_score_at_index()` does; the values are the first
prediction's, bit for bit.


- **The tiny-model registry validator no longer validates less when `jsonschema` is
missing.** `ai/scripts/validate_model_registry.py` used to fall back to a
four-field structural check and print `OK`; it now exits 2 and names the
Expand Down Expand Up @@ -4582,6 +4604,15 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
workflow, fails on the quiet form.


- **`float_vif` and SpEED refuse a prescaled plane past the `int` index of
their resampling and filter code.** `core/src/feature/vif_tools.c` indexes
a plane with `int`, so with `vif_prescale` or `speed_prescale` above about
1.414 at the 32768x32768 picture cap the index overflowed and the
resampling wrote outside the plane. `init()` now fails with `-EINVAL` and
names the plane size. 16K (15360x8640) and every smaller picture are
accepted at every prescale up to 4.0, as before, and no score changes.


- **The production CPU and MCP server images carry the licences of what they
contain, and publish the source their copyleft parts require.** Both images
have `/usr/local/share/vmafx/licenses/THIRD_PARTY_NOTICES.txt` with every
Expand Down
1 change: 1 addition & 0 deletions changelog.d/changed/pelorus-revendor-tidy-clean.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **The vendored Pelorus interop sources are re-vendored at the pelorus commit that clears their clang-tidy findings.** `scripts/sync-pelorus-interop.sh` pins `5f5614b0229d` (VMAFx/pelorus #78): the conformance test's long checks are split into helpers, blob headers are patched through `memcpy`, and each translation unit carries one cited `modernize-use-nullptr` block. No behaviour or ABI change (ABI 1.3).
6 changes: 6 additions & 0 deletions changelog.d/fixed/dnn-session-int8-explicit-path.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
- **Preserve explicit int8 paths in tiny-model DNN session loading.** When
`vmaf_dnn_session_open()` was called with an explicit `.int8.onnx` path,
`resolve_load_path()` lacked the `kInt8Suffix` early return present in
`dnn_attach_api.c`, causing it to append a redundant `.int8` suffix and derive
`<name>.int8.int8.onnx` before falling back to the fp32 path. The resolver now
checks `kInt8Suffix` upfront and preserves explicit int8 paths directly.
9 changes: 9 additions & 0 deletions changelog.d/fixed/model-collection-score-repeat.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **A model collection's score can be read more than once.**
`vmaf_score_at_index_model_collection()` failed with `-EINVAL` when the
frame had been scored before, and so did
`vmaf_score_pooled_model_collection()` over a range holding a frame already
scored per frame (the collector refused to write the members' scores of
that frame a second time, `feature "..." cannot be overwritten`). A frame
already predicted now returns its stored bootstrap scores, as a single
model's `vmaf_score_at_index()` does; the values are the first
prediction's, bit for bit.
7 changes: 7 additions & 0 deletions changelog.d/fixed/prescaled-plane-int-index-limit.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
- **`float_vif` and SpEED refuse a prescaled plane past the `int` index of
their resampling and filter code.** `core/src/feature/vif_tools.c` indexes
a plane with `int`, so with `vif_prescale` or `speed_prescale` above about
1.414 at the 32768x32768 picture cap the index overflowed and the
resampling wrote outside the plane. `init()` now fails with `-EINVAL` and
names the plane size. 16K (15360x8640) and every smaller picture are
accepted at every prescale up to 4.0, as before, and no score changes.
2 changes: 1 addition & 1 deletion core/include/libvmaf/pelorus/deband.h
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
*/

/*
* VENDORED FROM VMAFx/pelorus@42cb17106a2d3fae7790754f7cd8c6e1fbe6fa7f — DO NOT EDIT.
* VENDORED FROM VMAFx/pelorus@5f5614b0229d3461d5ef269bc75e9dd76497a6f1 — DO NOT EDIT.
* Append-only ABI; single
* source of truth is pelorus. Re-sync via scripts/sync-pelorus-interop.sh.
* See docs/adr/1113-vendor-pelorus-interop-abi.md.
Expand Down
2 changes: 1 addition & 1 deletion core/include/libvmaf/pelorus/denoise.h
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
*/

/*
* VENDORED FROM VMAFx/pelorus@42cb17106a2d3fae7790754f7cd8c6e1fbe6fa7f — DO NOT EDIT.
* VENDORED FROM VMAFx/pelorus@5f5614b0229d3461d5ef269bc75e9dd76497a6f1 — DO NOT EDIT.
* Append-only ABI; single
* source of truth is pelorus. Re-sync via scripts/sync-pelorus-interop.sh.
* See docs/adr/1113-vendor-pelorus-interop-abi.md.
Expand Down
2 changes: 1 addition & 1 deletion core/include/libvmaf/pelorus/interop.h
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
*/

/*
* VENDORED FROM VMAFx/pelorus@42cb17106a2d3fae7790754f7cd8c6e1fbe6fa7f — DO NOT EDIT.
* VENDORED FROM VMAFx/pelorus@5f5614b0229d3461d5ef269bc75e9dd76497a6f1 — DO NOT EDIT.
* Append-only ABI; single
* source of truth is pelorus. Re-sync via scripts/sync-pelorus-interop.sh.
* See docs/adr/1113-vendor-pelorus-interop-abi.md.
Expand Down
2 changes: 1 addition & 1 deletion core/include/libvmaf/pelorus/pelorus.h
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
*/

/*
* VENDORED FROM VMAFx/pelorus@42cb17106a2d3fae7790754f7cd8c6e1fbe6fa7f — DO NOT EDIT.
* VENDORED FROM VMAFx/pelorus@5f5614b0229d3461d5ef269bc75e9dd76497a6f1 — DO NOT EDIT.
* Append-only ABI; single
* source of truth is pelorus. Re-sync via scripts/sync-pelorus-interop.sh.
* See docs/adr/1113-vendor-pelorus-interop-abi.md.
Expand Down
4 changes: 4 additions & 0 deletions core/src/AGENTS.d/pooling-and-bootstrap.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,3 +49,7 @@ Do not restore translation-unit-local string literals: two paths would
again be able to publish different feature names. loops stay separate
because their callees and ownership contracts differ. fast source-contract
test is `core/test/test_bootstrap_name_contract.py`.

## Collection per-frame score reads stored values first (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05)

`vmaf_score_at_index_model_collection()` calls `read_predicted_collection_score()` before predicting: four named scores (`bootstrap_names.h` suffixes) already in collector -> return them. Prediction writes members' + named scores once per frame; collector refuses rewrite (`cannot be overwritten`), so without the read a second per-frame call, or `vmaf_score_pooled_model_collection()` (predicts every frame of its range) after a per-frame call, returned `-EINVAL`. Same rule as `vmaf_score_at_index()` for one model. Keep on upstream sync (upstream lacks it). Test: `core/test/test_model_collection_score_repeat.c`.
1 change: 1 addition & 0 deletions core/src/dnn/AGENTS.d/int8-quantization-and-scaling.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ invariant: Quantized int8 models redirect through fallback session opening and d

- **Sidecar `quant_mode` drives redirect**:
- entry points: `vmaf_use_tiny_model()` (`dnn_attach_api.c`), `vmaf_dnn_session_open()` (`dnn_api.c`).
- explicit `.int8.onnx` paths: if the caller passes a path ending in `.int8.onnx`, path resolution preserves it directly without appending a redundant `.int8` suffix (shared `kInt8Suffix` early return in both entry points).
- sidecar `quant_mode != VMAF_QUANT_FP32` -> load sibling `<basename>.int8.onnx` when present and valid; else fp32 baseline, logged at `VMAF_LOG_LEVEL_DEBUG` (ADR-1032).
- trigger 1: int8 file fails size cap or op allowlist -> each entry point's own path resolver.
- trigger 2: `vmaf_ort_open()` fails on int8 graph that passed those gates (ONNX Runtime build without kernel for quantised op; seen: `ConvInteger`) -> `vmaf_ort_open_with_fallback()` in `ort_backend.c`, only home. First attempt logs its `CreateSession` failure at DEBUG.
Expand Down
19 changes: 13 additions & 6 deletions core/src/dnn/dnn_api.c
Original file line number Diff line number Diff line change
Expand Up @@ -80,18 +80,25 @@ static int resolve_load_path(const VmafDnnSession *s, const char *onnx_path, siz
if (!s->has_sidecar || s->meta.quant_mode == VMAF_QUANT_FP32)
return 0;

static const char kInt8Suffix[] = ".int8.onnx";
static const char kOnnxSuffix[] = ".onnx";
const size_t plen = strlen(onnx_path);
const char *suffix = ".onnx";
const size_t suffix_len = 5u;
const size_t int8_len = sizeof(kInt8Suffix) - 1u;
const size_t onnx_len = sizeof(kOnnxSuffix) - 1u;

/* The caller already named the int8 graph — nothing to derive. */
if (plen >= int8_len && strcmp(onnx_path + plen - int8_len, kInt8Suffix) == 0)
return 0;

const size_t base_len =
(plen >= suffix_len && strcmp(onnx_path + plen - suffix_len, suffix) == 0) ?
plen - suffix_len :
(plen >= onnx_len && strcmp(onnx_path + plen - onnx_len, kOnnxSuffix) == 0) ?
plen - onnx_len :
plen;
if (base_len + sizeof(".int8.onnx") > int8_buf_sz)
if (base_len + sizeof(kInt8Suffix) > int8_buf_sz)
return -ENAMETOOLONG;

memcpy(int8_buf, onnx_path, base_len);
memcpy(int8_buf + base_len, ".int8.onnx", sizeof(".int8.onnx"));
memcpy(int8_buf + base_len, kInt8Suffix, sizeof(kInt8Suffix));

const int rc = vmaf_dnn_validate_onnx(int8_buf, max_bytes);
if (rc < 0) {
Expand Down
11 changes: 11 additions & 0 deletions core/src/feature/AGENTS.d/float-vif.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,3 +50,14 @@ invariant: float_vif lint decomposition, run-time Gaussian filter construction,
largest `((filter_width_s / 2) + 1) << s` over four-scale ladder, 16 at
default kernelscale — not from scale-0 filter alone. Do not replace it with
constant.

## Prescaled planes stay inside the int index of `vif_tools.c` (T-PRESCALED-PLANE-INT-INDEX-2026-10-05)

`vif_tools.c` indexes a plane as `y * stride + x` in `int`. `init()` refuses a
prescaled plane whose rows times stride (in samples) pass INT_MAX, through
`vif_plane_fits_int_index()` in `vif_tools.h`, before it allocates anything
(`init_scaled_plane()` holds the scaled-size checks).
`speed.c` and `speed_internal.c` call the same helper. Widening the index
instead would touch every function of the file; the plane such a prescale
needs is at least 8.6 GB per float buffer. `core/test/test_prescaled_plane_int_index.c`
holds the limit (16K accepted at prescale 4, the cap refused above 1.414).
8 changes: 8 additions & 0 deletions core/src/feature/AGENTS.d/speed-internal.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,3 +119,11 @@ all four `ALLOC_HOST` / `FREE_HOST` macro call sites in both TUs in
same PR. See `core/src/cuda/cuda_helper.cuh` for macro contract and
`core/src/cuda/picture_cuda.c` / `core/src/cuda/common.c` for
canonical usage of these members across codebase.

## The prescaled plane limit is part of the shared geometry

`speed_internal_init_dimensions()` refuses an allocation plane past the `int`
index of `vif_tools.c` (`vif_plane_fits_int_index()`), as `speed.c`'s own
`speed_init_dimensions()` does. Every SpEED device twin takes its geometry
from this function, so they refuse the same option combinations as the CPU
(T-PRESCALED-PLANE-INT-INDEX-2026-10-05).
41 changes: 30 additions & 11 deletions core/src/feature/float_vif.c
Original file line number Diff line number Diff line change
Expand Up @@ -276,6 +276,33 @@ static int alloc_buffers(VmafFeatureExtractor *fex, VifState *s, unsigned h)
return 0;
}

/* The prescaled plane compute_vif() works on: its size, which must clear the
* four-scale ladder minimum, and its strides; the plane must also fit the int
* index of vif_tools.c (T-PRESCALED-PLANE-INT-INDEX-2026-10-05). */
static int init_scaled_plane(VifState *s, unsigned w, unsigned h, int vif_min_dim)
{
s->scaled_w = (size_t)lround(w * s->vif_prescale);
s->scaled_h = (size_t)lround(h * s->vif_prescale);

if (s->scaled_w < (size_t)vif_min_dim || s->scaled_h < (size_t)vif_min_dim) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"float_vif requires scaled width >= %d and height >= %d for the "
"four-scale ladder (got %zux%zu)\n",
vif_min_dim, vif_min_dim, s->scaled_w, s->scaled_h);
return -EINVAL;
}
s->float_stride = ALIGN_CEIL(w * sizeof(float));
s->scaled_float_stride = ALIGN_CEIL(s->scaled_w * sizeof(float));
if (!vif_plane_fits_int_index(s->scaled_float_stride / sizeof(float), s->scaled_h)) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"float_vif: the prescaled plane (%zux%zu) has more samples than the "
"int index of vif_tools allows; lower vif_prescale\n",
s->scaled_w, s->scaled_h);
return -EINVAL;
}
return 0;
}

static int init(VmafFeatureExtractor *fex, enum VmafPixelFormat pix_fmt, unsigned bpc, unsigned w,
unsigned h)
{
Expand Down Expand Up @@ -314,18 +341,10 @@ static int init(VmafFeatureExtractor *fex, enum VmafPixelFormat pix_fmt, unsigne
return -EINVAL;
}

s->scaled_w = (size_t)lround(w * s->vif_prescale);
s->scaled_h = (size_t)lround(h * s->vif_prescale);

if (s->scaled_w < (size_t)vif_min_dim || s->scaled_h < (size_t)vif_min_dim) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"float_vif requires scaled width >= %d and height >= %d for the "
"four-scale ladder (got %zux%zu)\n",
vif_min_dim, vif_min_dim, s->scaled_w, s->scaled_h);
return -EINVAL;
const int plane_err = init_scaled_plane(s, w, h, vif_min_dim);
if (plane_err) {
return plane_err;
}
s->float_stride = ALIGN_CEIL(w * sizeof(float));
s->scaled_float_stride = ALIGN_CEIL(s->scaled_w * sizeof(float));

const int alloc_err = alloc_buffers(fex, s, h);
if (alloc_err) {
Expand Down
8 changes: 8 additions & 0 deletions core/src/feature/speed.c
Original file line number Diff line number Diff line change
Expand Up @@ -1306,6 +1306,14 @@ static int speed_init_dimensions(SpeedDimensions *dim, int w, int h, double spee
vmaf_log(VMAF_LOG_LEVEL_ERROR, "SpEED: image too small, operating width or height is 0\n");
return -EINVAL;
}
if (!vif_plane_fits_int_index(ALIGN_CEIL(dim->alloc_width * sizeof(float)) / sizeof(float),
dim->alloc_height)) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"SpEED: the prescaled plane (%zux%zu) has more samples than the int index "
"of vif_tools allows; lower speed_prescale\n",
dim->alloc_width, dim->alloc_height);
return -EINVAL;
}

dim->num_blocks_horizontal = dim->truncated_width / dim->block_size;
dim->num_blocks_vertical = dim->truncated_height / dim->block_size;
Expand Down
8 changes: 8 additions & 0 deletions core/src/feature/speed_internal.c
Original file line number Diff line number Diff line change
Expand Up @@ -92,6 +92,14 @@ int speed_internal_init_dimensions(SpeedInternalDimensions *dim, int w, int h, d
vmaf_log(VMAF_LOG_LEVEL_ERROR, "SpEED: image too small, operating width or height is 0\n");
return -EINVAL;
}
if (!vif_plane_fits_int_index(speed_internal_float_stride(dim->alloc_width) / sizeof(float),
dim->alloc_height)) {
vmaf_log(VMAF_LOG_LEVEL_ERROR,
"SpEED: the prescaled plane (%zux%zu) has more samples than the int index "
"of vif_tools allows; lower speed_prescale\n",
dim->alloc_width, dim->alloc_height);
return -EINVAL;
}
return 0;
}

Expand Down
13 changes: 13 additions & 0 deletions core/src/feature/vif_tools.h
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,9 @@
#ifndef VIF_TOOLS_H_
#define VIF_TOOLS_H_

#include <limits.h>
#include <stdbool.h>
#include <stddef.h>

/* The GPU backends that share vif_get_min_dim() are C++ (SYCL) and
* Objective-C++ (Metal) translation units, while vif_tools.c is C -- without
Expand Down Expand Up @@ -111,6 +113,17 @@ void vif_scale_frame_bilinear_precomputed_s(const float *src, float *dst, int sr
* VIF_LANCZOS4_TAPS * dst_len floats. */
void vif_scale_lanczos4_axis_weights(int src_len, int dst_len, float *weights);

/* The functions of vif_tools.c index a plane with int arithmetic,
* y * stride + x with the stride in elements. A plane of h rows of
* stride_elems elements fits that index when h * stride_elems <= INT_MAX.
* float_vif and SpEED refuse a prescaled plane that does not: from a prescale
* of about 1.4142 at the 32768 x 32768 picture cap; 16K fits at every prescale
* the options accept (T-PRESCALED-PLANE-INT-INDEX-2026-10-05). */
static inline bool vif_plane_fits_int_index(size_t stride_elems, size_t h)
{
return h == 0u || stride_elems <= (size_t)INT_MAX / h;
}

int vif_get_filter_size(int scale, float kernelscale);

/* Smallest frame dimension the four-scale VIF ladder can process without
Expand Down
8 changes: 8 additions & 0 deletions core/src/feature/x86/AGENTS.d/motion.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,11 @@ A change to the arithmetic goes into the helper that holds the statement, on
both files, and `core/test/test_motion_v2_simd.c` /
`core/test/test_motion_avx512_parity.c` compare the result with the scalar
reference. Keep every function at or under 60 lines (ADR-1142).

## `sad_avx512` takes the difference in unsigned lanes

`sad_avx512()` forms `|a - b|` as `max_epu16(a, b) - min_epu16(a, b)`. A
signed 16-bit subtraction followed by `abs_epi16` wraps for 16-bit samples
that differ by more than 32767 (65535 against 0 gave 1). The 16-bit cases of
`test_sad_avx512_*` in `core/test/test_motion_avx512_parity.c` fail on that
form (T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05).
10 changes: 6 additions & 4 deletions core/src/feature/x86/motion_avx512.c
Original file line number Diff line number Diff line change
Expand Up @@ -345,7 +345,10 @@ uint64_t motion_score_pipeline_8_avx512(const uint8_t *prev, ptrdiff_t prev_stri
*
* Operates on the luma plane (data[0]) only. Both pictures must have the
* same dimensions and bit-depth. Processes 32 uint16 samples per SIMD
* iteration using _mm512_abs_epi16 + widening accumulation.
* iteration: |a - b| = max(a, b) - min(a, b) in unsigned 16-bit lanes, then
* widening accumulation. A signed 16-bit difference wraps for 16-bit samples
* that differ by more than 32767 (65535 - 0 gave 1);
* T-SIMD-SAD-AVX512-INT16-DIFFERENCE-2026-10-05.
* ----------------------------------------------------------------------- */
void sad_avx512(VmafPicture *pic_a, VmafPicture *pic_b, uint64_t *sad_out)
{
Expand All @@ -370,9 +373,8 @@ void sad_avx512(VmafPicture *pic_a, VmafPicture *pic_b, uint64_t *sad_out)
for (; j + 32 <= w; j += 32) {
__m512i va = _mm512_loadu_si512((const __m512i *)(row_a + j));
__m512i vb = _mm512_loadu_si512((const __m512i *)(row_b + j));
/* Signed subtract, then abs -> |a[k]-b[k]| per int16 lane */
__m512i diff = _mm512_sub_epi16(va, vb);
__m512i abs_diff = _mm512_abs_epi16(diff);
/* |a[k]-b[k]| per uint16 lane, exact for every 16-bit sample */
__m512i abs_diff = _mm512_sub_epi16(_mm512_max_epu16(va, vb), _mm512_min_epu16(va, vb));
/* Widen uint16 -> uint32 in two halves and accumulate */
acc = _mm512_add_epi32(acc, _mm512_cvtepu16_epi32(_mm512_castsi512_si256(abs_diff)));
acc = _mm512_add_epi32(acc,
Expand Down
Loading
Loading