Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4492,6 +4492,17 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
([licensing](docs/licensing.md#models)).


- **A model collection's score can be read more than once.**
`vmaf_score_at_index_model_collection()` failed with `-EINVAL` when the
frame had been scored before, and so did
`vmaf_score_pooled_model_collection()` over a range holding a frame already
scored per frame (the collector refused to write the members' scores of
that frame a second time, `feature "..." cannot be overwritten`). A frame
already predicted now returns its stored bootstrap scores, as a single
model's `vmaf_score_at_index()` does; the values are the first
prediction's, bit for bit.


- **The tiny-model registry validator no longer validates less when `jsonschema` is
missing.** `ai/scripts/validate_model_registry.py` used to fall back to a
four-field structural check and print `OK`; it now exits 2 and names the
Expand Down
9 changes: 9 additions & 0 deletions changelog.d/fixed/model-collection-score-repeat.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
- **A model collection's score can be read more than once.**
`vmaf_score_at_index_model_collection()` failed with `-EINVAL` when the
frame had been scored before, and so did
`vmaf_score_pooled_model_collection()` over a range holding a frame already
scored per frame (the collector refused to write the members' scores of
that frame a second time, `feature "..." cannot be overwritten`). A frame
already predicted now returns its stored bootstrap scores, as a single
model's `vmaf_score_at_index()` does; the values are the first
prediction's, bit for bit.
4 changes: 4 additions & 0 deletions core/src/AGENTS.d/pooling-and-bootstrap.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,3 +49,7 @@ Do not restore translation-unit-local string literals: two paths would
again be able to publish different feature names. loops stay separate
because their callees and ownership contracts differ. fast source-contract
test is `core/test/test_bootstrap_name_contract.py`.

## Collection per-frame score reads stored values first (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05)

`vmaf_score_at_index_model_collection()` calls `read_predicted_collection_score()` before predicting: four named scores (`bootstrap_names.h` suffixes) already in collector -> return them. Prediction writes members' + named scores once per frame; collector refuses rewrite (`cannot be overwritten`), so without the read a second per-frame call, or `vmaf_score_pooled_model_collection()` (predicts every frame of its range) after a per-frame call, returned `-EINVAL`. Same rule as `vmaf_score_at_index()` for one model. Keep on upstream sync (upstream lacks it). Test: `core/test/test_model_collection_score_repeat.c`.
39 changes: 39 additions & 0 deletions core/src/libvmaf.c
Original file line number Diff line number Diff line change
Expand Up @@ -4347,6 +4347,37 @@ int vmaf_score_at_index(VmafContext *vmaf, VmafModel *model, double *score, unsi
return err;
}

/* The bootstrap score of a model collection already predicted at `index`:
* the four named scores its first prediction wrote into the collector
* (bootstrap_append_named_scores() in predict.c). Returns 0, or the
* collector's error for the first one missing. */
static int read_predicted_collection_score(VmafContext *vmaf,
const VmafModelCollection *model_collection,
VmafModelCollectionScore *score, unsigned index)
{
const size_t name_sz = BOOTSTRAP_NAME_BUF_SZ(model_collection->name);
char *name = (char *)calloc(1u, name_sz);
if (!name)
return -ENOMEM;
const char *const suffix[] = {BOOTSTRAP_SUFFIX_BAGGING, BOOTSTRAP_SUFFIX_STDDEV,
BOOTSTRAP_SUFFIX_CI_LO, BOOTSTRAP_SUFFIX_CI_HI};
double value[4] = {0.0, 0.0, 0.0, 0.0};
int err = 0;
for (unsigned i = 0; i < 4u && !err; i++) {
(void)snprintf(name, name_sz, "%s%s", model_collection->name, suffix[i]);
err = vmaf_feature_collector_get_score(vmaf->feature_collector, name, &value[i], index);
}
free(name);
if (err)
return err;
score->type = VMAF_MODEL_COLLECTION_SCORE_BOOTSTRAP;
score->bootstrap.bagging_score = value[0];
score->bootstrap.stddev = value[1];
score->bootstrap.ci.p95.lo = value[2];
score->bootstrap.ci.p95.hi = value[3];
return 0;
}

int vmaf_score_at_index_model_collection(VmafContext *vmaf, VmafModelCollection *model_collection,
VmafModelCollectionScore *score, unsigned index)
{
Expand All @@ -4357,6 +4388,14 @@ int vmaf_score_at_index_model_collection(VmafContext *vmaf, VmafModelCollection
if (!score)
return -EINVAL;

/* T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05: the prediction writes the
* members' and the four named scores of this frame into the collector,
* which refuses a second write. A frame already predicted (by an earlier
* per-frame call, or by the per-frame loop of
* vmaf_score_pooled_model_collection()) returns the stored values, as
* vmaf_score_at_index() does for a single model. */
if (read_predicted_collection_score(vmaf, model_collection, score, index) == 0)
return 0;
return vmaf_predict_score_at_index_model_collection(model_collection, vmaf->feature_collector,
index, score);
}
Expand Down
11 changes: 11 additions & 0 deletions core/test/meson.build
Original file line number Diff line number Diff line change
Expand Up @@ -949,6 +949,17 @@ test('test_hip_wavefront_reduce', test_hip_wavefront_reduce, suite : ['fast'])

# ADR-1188 — percentile temporal pooling (MEDIAN / PERC5 / PERC10 / PERC20).
# Pools imported per-frame scores, so it needs no YUV fixture and no model.
# T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05: a model collection's per-frame
# score read twice, and its pooled score after a per-frame read of a frame in
# the range, return the stored values instead of -EINVAL.
test_model_collection_score_repeat = executable('test_model_collection_score_repeat',
['test.c', 'test_model_collection_score_repeat.c'],
include_directories : [libvmaf_inc, test_inc],
link_with : get_option('default_library') == 'both' ? libvmaf.get_static_lib() : libvmaf,
dependencies : [math_lib, pthread_dependency, thread_lib, gpu_all_deps],
)
test('test_model_collection_score_repeat', test_model_collection_score_repeat, suite : ['fast'])

test_pool_percentile = executable('test_pool_percentile',
['test.c', 'test_pool_percentile.c'],
include_directories : [libvmaf_inc, test_inc],
Expand Down
175 changes: 175 additions & 0 deletions core/test/test_model_collection_score_repeat.c
Original file line number Diff line number Diff line change
@@ -0,0 +1,175 @@
/**
*
* Copyright 2026 Lusoris
*
* SPDX-License-Identifier: EUPL-1.2
*/

/*
* A model collection's score can be read more than once
* (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05).
*
* vmaf_score_at_index_model_collection() predicts every member model and
* writes the members' scores and four named bootstrap scores of the frame
* into the feature collector, which refuses a second write of a frame. On
* master before the fix a second per-frame call of a frame, or
* vmaf_score_pooled_model_collection() over a range holding a frame already
* scored per frame, failed with -EINVAL ("feature ... cannot be overwritten").
* A single model reads its stored score first (vmaf_score_at_index()); a
* collection now does the same.
*
* Failing first: without the fix, test_second_frame_score and
* test_pooled_after_frame_score fail (measured on master 782eba01f).
*/

#include <stdbool.h>
#include <stdint.h>
#include <string.h>

#include "libvmaf/libvmaf.h"
#include "libvmaf/model.h"
#include "mu_table.h"
#include "test.h"

/* NOLINTBEGIN(modernize-use-nullptr): C translation unit. The fork builds C as
* C23, where clang-tidy also proposes the `nullptr` keyword, but MSVC's
* documented /std:clatest C23 feature set does not include `nullptr` and the
* required Windows builds compile this TU with cl.exe (C2065). ADR-1138. */

enum { W = 176, H = 144, N_FRAMES = 5 };

typedef struct Session {
VmafContext *vmaf;
VmafModel *lead;
VmafModelCollection *collection;
} Session;

static int fill_picture(VmafPicture *pic, unsigned seed)
{
const int err = vmaf_picture_alloc(pic, VMAF_PIX_FMT_YUV420P, 8, W, H);
if (err)
return err;
for (unsigned p = 0; p < 3u; p++) {
uint8_t *row = pic->data[p];
for (unsigned y = 0; y < pic->h[p]; y++, row += pic->stride[p]) {
for (unsigned x = 0; x < pic->w[p]; x++)
row[x] = (uint8_t)((x * 7u + y * 13u + seed * 29u + p * 3u) & 0xffu);
}
}
return 0;
}

/* vmaf_b_v0.6.3 over N_FRAMES generated frames, flushed. */
static bool open_session(Session *s)
{
memset(s, 0, sizeof(*s));
VmafConfiguration cfg;
memset(&cfg, 0, sizeof(cfg));
VmafModelConfig mcfg = {0};
if (vmaf_init(&s->vmaf, cfg) ||
vmaf_model_collection_load(&s->lead, &s->collection, &mcfg, "vmaf_b_v0.6.3") ||
vmaf_use_features_from_model_collection(s->vmaf, s->collection))
return false;
for (unsigned i = 0; i < N_FRAMES; i++) {
VmafPicture ref;
VmafPicture dist;
if (fill_picture(&ref, i))
return false;
if (fill_picture(&dist, i + 50u)) {
(void)vmaf_picture_unref(&ref);
return false;
}
if (vmaf_read_pictures(s->vmaf, &ref, &dist, i))
return false;
}
return vmaf_read_pictures(s->vmaf, NULL, NULL, 0) == 0;
}

static bool close_session(Session *s)
{
const bool ok = vmaf_close(s->vmaf) == 0;
vmaf_model_destroy(s->lead);
vmaf_model_collection_destroy(s->collection);
return ok;
}

static bool same_bits(double a, double b)
{
uint64_t x = 0;
uint64_t y = 0;
memcpy(&x, &a, sizeof(x));
memcpy(&y, &b, sizeof(y));
return x == y;
}

static bool same_score(const VmafModelCollectionScore *a, const VmafModelCollectionScore *b)
{
return a->type == b->type &&
same_bits(a->bootstrap.bagging_score, b->bootstrap.bagging_score) &&
same_bits(a->bootstrap.stddev, b->bootstrap.stddev) &&
same_bits(a->bootstrap.ci.p95.lo, b->bootstrap.ci.p95.lo) &&
same_bits(a->bootstrap.ci.p95.hi, b->bootstrap.ci.p95.hi);
}

static char *test_second_frame_score(void)
{
Session s;
mu_assert("session", open_session(&s));
VmafModelCollectionScore first;
VmafModelCollectionScore second;
memset(&first, 0, sizeof(first));
memset(&second, 0, sizeof(second));
mu_assert("first", vmaf_score_at_index_model_collection(s.vmaf, s.collection, &first, 2) == 0);
mu_assert("second",
vmaf_score_at_index_model_collection(s.vmaf, s.collection, &second, 2) == 0);
mu_assert("the same score",
same_score(&first, &second) && first.type == VMAF_MODEL_COLLECTION_SCORE_BOOTSTRAP);
mu_assert("close", close_session(&s));
return NULL;
}

/* The pooled score of a fresh session, for comparison. */
static bool fresh_pooled(VmafModelCollectionScore *score)
{
Session s;
const bool ok = open_session(&s) &&
vmaf_score_pooled_model_collection(s.vmaf, s.collection, VMAF_POOL_METHOD_MEAN,
score, 0, N_FRAMES - 1) == 0;
return close_session(&s) && ok;
}

static char *test_pooled_after_frame_score(void)
{
VmafModelCollectionScore expected;
memset(&expected, 0, sizeof(expected));
mu_assert("fresh pooled score", fresh_pooled(&expected));
Session s;
mu_assert("session", open_session(&s));
VmafModelCollectionScore frame;
VmafModelCollectionScore pooled;
memset(&frame, 0, sizeof(frame));
memset(&pooled, 0, sizeof(pooled));
mu_assert("frame 2",
vmaf_score_at_index_model_collection(s.vmaf, s.collection, &frame, 2) == 0);
mu_assert("pooled over frame 2",
vmaf_score_pooled_model_collection(s.vmaf, s.collection, VMAF_POOL_METHOD_MEAN,
&pooled, 0, N_FRAMES - 1) == 0);
mu_assert("equals the fresh session's", same_score(&pooled, &expected));
mu_assert("pooled twice",
vmaf_score_pooled_model_collection(s.vmaf, s.collection, VMAF_POOL_METHOD_MEAN,
&pooled, 0, N_FRAMES - 1) == 0 &&
same_score(&pooled, &expected));
mu_assert("close", close_session(&s));
return NULL;
}

char *run_tests(void)
{
static const MuTest tests[] = {
MU_TEST(test_second_frame_score),
MU_TEST(test_pooled_after_frame_score),
};
return mu_run_table(tests, MU_TABLE_LEN(tests));
}

/* NOLINTEND(modernize-use-nullptr) */
14 changes: 14 additions & 0 deletions docs/rebase-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -61977,3 +61977,17 @@ No score, public API or FFmpeg patch impact.
transformers. The `vmaf-tune-train` test suite is removed from `.github/test-suites.json` and
`tests-and-quality-gates.yml`; a conflict there takes the side without it. No score, public C
API or FFmpeg patch impact.

## A model collection's per-frame score reads its stored values first

`fix/model-set-score-idempotent` (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05).

- `core/src/libvmaf.c` gains `read_predicted_collection_score()`, which
`vmaf_score_at_index_model_collection()` calls before
`vmaf_predict_score_at_index_model_collection()`: a frame whose four named
bootstrap scores are already in the collector returns them. Upstream
Netflix/vmaf predicts every time and has the same failure; an upstream
sync that touches this function keeps the read.
- New test `core/test/test_model_collection_score_repeat.c` and its block in
`core/test/meson.build`. No score or golden impact: a first prediction is
unchanged and a repeat returns its stored values.
1 change: 1 addition & 0 deletions docs/state.md
Original file line number Diff line number Diff line change
Expand Up @@ -1036,6 +1036,7 @@ landed fix yet._

## Recently closed

| **T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05** — a model collection scored per frame and then pooled over that frame failed (`-EINVAL`); no score was wrong. `vmaf_score_at_index_model_collection()` predicts every member model and writes the members' scores and four named bootstrap scores of the frame into the feature collector (`bootstrap_gather_scores()` / `bootstrap_append_named_scores()` in `core/src/predict.c`). `vmaf_score_pooled_model_collection()` predicts every frame of its range again, and the collector refuses a second write of a frame (`feature "vmaf_0001" cannot be overwritten at index 2`), so the pooled call, or a second per-frame call, returned `-EINVAL` after a per-frame call of a frame in the range. A single model reads its stored score first (`vmaf_score_at_index()`); a collection did not. Upstream Netflix/vmaf has the same code. Found by the VMAFx core API tests (RC4 WP2, `rc4/api-wp2-core`). | **FIXED on `fix/model-set-score-idempotent` (opened and closed by one PR).** `vmaf_score_at_index_model_collection()` returns the four stored named scores of a frame already predicted (`read_predicted_collection_score()` in `core/src/libvmaf.c`), bit for bit the first prediction's. `core/test/test_model_collection_score_repeat.c` fails on master without the fix (second per-frame call and pooled-after-per-frame both `-EINVAL`). | — | `fix/model-set-score-idempotent` | 2026-10-06 | fixed |
| **T-CUDA-FLOAT-MOTION-TILE-READ-BEFORE-PLANE-2026-10-05** — `float_motion_cuda` loads a 20x20 tile per 16x16 block with reflect-101 padding (`fm_mirror()` in `float_motion/float_motion_score.cu`) and did not clamp the reflected index. For a plane 3 to 9 samples wide or high, or 17, some padding loads reflect to a negative index (`fm_mirror(17, 5)` = -9): a read before the plane's row or before the plane. No output uses those tile cells, so no score changed, and `compute-sanitizer --tool memcheck` reported nothing on master at 5x5, 17x17 and 9x40 (the reads stay inside the device allocation). The HIP twin clamps (`fm_tile_index()`) | **FOUND by the RC3 integer-overflow audit of every accumulator (an aside of the CUDA rows: offset math) and FIXED on `fix/cuda-float-motion-tile-clamp` (opened and closed by one PR).** `fm_mirror()` returns `vmaf_cuda_tile_index(vmaf_cuda_reflect_101(idx, sup), sup)` from the shared `cuda/cuda_tile_index.h`, the form the CUDA `motion_v2` and `float_vif` kernels use; every index an output consumes is unchanged. `test_cuda_kernel_source_contract.py` requires the clamped form and reports master's (`test_unclamped_float_motion_tile_mirror_is_detected`). On `ryzen-4090-arc` (RTX 4090): `test_cuda_float_motion_parity{,_large}` and `test_cuda_exact_twins` pass, and `float_motion_cuda` equals `--backend cpu` on every output at 5x5, 17x17 and 9x40. | [ADR-1409](adr/1409-float-motion-twins-cpu-float-sum.md) | `fix/cuda-float-motion-tile-clamp` | 2026-10-05 | fixed |
| **T-GPU-PSNR-HVS-SCAN-32768-CHUNKS-2026-10-05** — the HIP and SYCL `psnr_hvs` twins stopped their prefix scan of the per-chunk term counts at 32,768 chunks of 256 blocks (`limit = num_chunks < 32768u ? num_chunks : 32768u` in `hvs_scan_prefix_hip()` and `launch_scan_prefix()`). Above 8,388,608 blocks the offsets of the later chunks were never written (the buffer is not cleared), so the compaction wrote those chunks' terms at uninitialised (or the previous frame's) offsets, out of the buffer's bounds, and the term total left them out. Read from source, not run past 16K. 16K in 4:4:4 needs 31,728 chunks (3.2 % under the cap); 16384x8640 in 4:4:4 already needs more, and the 32768 picture cap 256,779. The CUDA twin scans every chunk | **FOUND by the RC3 integer-overflow audit of every accumulator (HIP and SYCL rows, OVERFLOW@CAP-ONLY) and FIXED on `fix/psnr-hvs-gpu-scan-every-chunk` (opened and closed by one PR).** Both scans run to `num_chunks`, as the CUDA twin's does; the running offset and the term total stay `uint32_t` (64 terms x 65,735,283 blocks at the cap = 4.2e9 < 2^32). `test_psnr_hvs_gpu_scan_contract.py` holds all three scans to every chunk, reports the master form of the HIP and SYCL scans, and derives the chunk counts at 16K, at 16384x8640 and at the cap. No device run past 16K (device memory limits belong to a later candidate); on `ryzen-4090-arc` `test_hip_psnr_hvs_parity{,_large}` and `test_hip_exact_twins` pass on the gfx1036, `test_sycl_psnr_hvs_parity{,_large}`, `test_sycl_exact_twins` and `test_sycl_kernel_scratch` on the Arc A380. | [ADR-1401](adr/1401-psnr-hvs-sycl-hip-exact-twins.md) | `fix/psnr-hvs-gpu-scan-every-chunk` | 2026-10-05 | fixed |
| **T-PSNR-APSNR-CLIP-SSE-UINT64-WRAP-2026-10-05** — `apsnr_*` (the `psnr` extractor with `enable_apsnr`) summed each plane's SSE over the clip in a `uint64_t` on the CPU (`integer_psnr.c`) and in the CUDA, HIP, SYCL and Metal hosts. One frame's SSE is below 2^62, but at 16 bits with every sample at the maximum difference the clip sum wraps at frame 2072 of 1080p, 122 of 8K DCI and 33 of 16K (12 bits: 530,502 / 31,085 / 8,290; 10 bits: 8.5 million / 498,076 / 132,821), and `apsnr_*` comes out too high without a message. Upstream Netflix/vmaf has the same `uint64_t` sum | **FOUND by the RC3 integer-overflow audit of every accumulator (CPU and the CUDA / HIP rows; DEPENDS on the frame count) and FIXED on `fix/apsnr-clip-sse-128` (opened and closed by one PR).** `core/src/feature/psnr_score.h` holds the sum as `VmafPsnrClipSse` (two `uint64_t` words, `vmaf_psnr_clip_sse_add()` carries) and `vmaf_psnr_aggregate()` reads it as `(double)lo` while the high word is 0, so a clip that never reached 2^64 keeps its bits; the CPU extractor and the four device hosts use it. Evidence on `ryzen-4090-arc`: the new `test_psnr_apsnr_clip_sse_past_two_pow_64` case of `test_integer_psnr_coverage` scores 300 frames of 4096x4096 16-bit pictures at the maximum difference (sum 2.2e19) and holds `apsnr_y` at 0 dB; master's build reads 8.3 dB and fails it. On the Netflix 576x324 pair `apsnr_{y,cb,cr}` are identical to master's and identical on CPU, CUDA (RTX 4090), HIP (gfx1036) and SYCL (Arc A380); `test_{cuda,hip,sycl}_psnr_parity{,_large}` and `test_{cuda,hip,sycl}_exact_twins` pass; the GPU source contracts pin the helper. | [ADR-1193](adr/1193-psnr-uncapped-option.md) | `fix/apsnr-clip-sse-128` | 2026-10-05 | fixed |
Expand Down
Loading
Loading