Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 3 additions & 19 deletions .standards-baseline.json
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
{
"version": 1,
"generated_at": "2026-10-01T20:30:59Z",
"generated_at": "2026-10-01T20:44:38Z",
"repository": "VMAFx/vmafx",
"commit_sha": "4ae168bc06e2531cf7811c7e0ab7d3f20f790650",
"total_infractions": 365,
"commit_sha": "a5acac3477c451e562782ed4e6f7cf5842e5e2ef",
"total_infractions": 363,
"infractions": [
{
"rule_id": "HISS-04",
Expand Down Expand Up @@ -2102,14 +2102,6 @@
"message": "Function 'vmaf_sycl_import_va_surface' (301 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/src/sycl/dmabuf_import.cpp:366:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_barten_csf.c",
"line_number": 48,
"symbol": "test_barten_csf",
"message": "Function 'test_barten_csf' (87 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_barten_csf.c:48:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_pelorus_interop.c",
Expand Down Expand Up @@ -2142,14 +2134,6 @@
"message": "Function 'test_x265_csv_reader' (79 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_pelorus_interop.c:562:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_psnr_hvs_simd.c",
"line_number": 164,
"symbol": "ref_calc_psnrhvs",
"message": "Function 'ref_calc_psnrhvs' (115 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_psnr_hvs_simd.c:164:HISS-04"
},
{
"rule_id": "HISS-07",
"file_path": "dev/scripts/fetch-intel-neo.py",
Expand Down
45 changes: 45 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -660,6 +660,20 @@
`git rebase` conflicts in CI (ADR-1383).


- **Ten core C and C++ test files conform to clang-tidy and HISS standards (part 1).**
The first batch of CPU and core test sources (`test_barten_csf.c`,
`test_psnr_hvs_simd.c`, `test_propagate_metadata.c`, `test_thread_pool.c`,
`test_context.c`, `test_predict.c`, `test_ciede.c`,
`test_pic_preallocation.c`, `test_locale_handling.c`, and `test_dict.cpp`)
were refactored to zero clang-tidy findings in the CPU lane under
[ADR-1142](docs/adr/1142-clang-tidy-debt-ratchet.md). Functions exceeding
branch, nesting, and statement thresholds were split into modular helpers
satisfying HISS-04 / NASA JPL Rule 4, eliminating 2 recorded infractions from
`.standards-baseline.json`. C23 `nullptr` diagnostics in C files are scoped under
[ADR-1138](docs/adr/1138-c23-nullptr-msvc-compat.md), and `test_dict.cpp` uses
anonymous namespaces and standard `nullptr`. All tests continue to pass.


- **Nine HIP test files conform to clang-tidy and HISS standards (part 1).**
The first batch of HIP test sources (`test_hip_smoke.c`,
`test_hip_float_adm_parity.c`, `test_hip_motion3_parity.c`,
Expand Down Expand Up @@ -1015,6 +1029,37 @@
`float` arithmetic and the `5e-3` tolerance.


- **`float_adm_cuda` returns the CPU's scores bit for bit.** The CUDA twin
of `float_adm` was up to 1.3e-5 from the CPU extractor (`adm_scale0` at
3840x2160) and matched it on 144 of 791 measured scores. Nine things
differed: the association of the angle test's threshold (the 1.3e-5), the
order of the sums, CSF weights from a copied formula that rounded
differently, a true division where the CPU multiplies by a refined
reciprocal estimate, the order of the masking threshold's terms, `float`
constants and a `float` gain limit where the CPU uses `double`, and a
floor of the frame sums at `1e-2` where the CPU's is `1e-10`. The last one
could report `adm2 = 1` where the CPU reports 0, for content with almost no
reference detail scored without the noise floor. The twin now runs the
CPU's arithmetic in the CPU's types, adds each row on the device and the
rows on the host in the CPU's order, and takes the weights, the reduced
region, the pooling and the floor from the CPU's own routines
([ADR-1420](docs/adr/1420-cuda-float-adm-cpu-arithmetic.md)). The CPU's
division is built on the processor's `RCPSS` estimate, so the twin probes
that estimate when the extractor starts (about 10 ms) and evaluates it on
the device. Measured on an RTX 4090 at `--precision max`: every output of
every frame identical on the Netflix pair at 8, 10, 12 and 16 bits, both
1080p checkerboard pairs and BBB 3840x2160, also with `debug=true` and
with non-default `adm_enhn_gain_limit`, `adm_bypass_cm`,
`adm_noise_weight`, `adm_skip_aim_scale` and viewing geometry. The parity
gate compares this twin with tolerance 0. Not identical: `adm_p_norm`
other than 1 or 3, where the twin is within 1.1e-7 of the CPU (the two
`powf` implementations differ). A run of the twin alone takes 1.98 ms per
3840x2160 frame instead of 1.87 ms; its kernels take 1.11 ms instead of
0.76 ms. Stored `float_adm_cuda` outputs change in their low digits by at
most 1.3e-5. The SYCL, HIP and Metal twins still agree with the CPU to
four decimal places.


- **`float_motion_cuda` returns the CPU's scores bit for bit.** The CPU
`float_motion` extractor adds the absolute differences of a row into one
`float`, the row sums into another, and divides in `float`, so its score
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -184,7 +184,7 @@ generated agent context against `AGENTS.md`.
private scratch-link policy.

**Debt Baseline**: `.standards-baseline.json` anchors the debt ratchet at
365 recorded infractions; audit forbids growth.
363 recorded infractions; audit forbids growth.

[praetor-docs-badge]: https://github.com/vmafx/vmafx/actions/workflows/praetor-docs.yml/badge.svg
[praetor-docs-runs]: https://github.com/vmafx/vmafx/actions/workflows/praetor-docs.yml
Expand Down
12 changes: 12 additions & 0 deletions changelog.d/changed/std-core-tests-a.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
- **Ten core C and C++ test files conform to clang-tidy and HISS standards (part 1).**
The first batch of CPU and core test sources (`test_barten_csf.c`,
`test_psnr_hvs_simd.c`, `test_propagate_metadata.c`, `test_thread_pool.c`,
`test_context.c`, `test_predict.c`, `test_ciede.c`,
`test_pic_preallocation.c`, `test_locale_handling.c`, and `test_dict.cpp`)
were refactored to zero clang-tidy findings in the CPU lane under
[ADR-1142](docs/adr/1142-clang-tidy-debt-ratchet.md). Functions exceeding
branch, nesting, and statement thresholds were split into modular helpers
satisfying HISS-04 / NASA JPL Rule 4, eliminating 2 recorded infractions from
`.standards-baseline.json`. C23 `nullptr` diagnostics in C files are scoped under
[ADR-1138](docs/adr/1138-c23-nullptr-msvc-compat.md), and `test_dict.cpp` uses
anonymous namespaces and standard `nullptr`. All tests continue to pass.
29 changes: 29 additions & 0 deletions changelog.d/fixed/cuda-float-adm-cpu-arithmetic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
- **`float_adm_cuda` returns the CPU's scores bit for bit.** The CUDA twin
of `float_adm` was up to 1.3e-5 from the CPU extractor (`adm_scale0` at
3840x2160) and matched it on 144 of 791 measured scores. Nine things
differed: the association of the angle test's threshold (the 1.3e-5), the
order of the sums, CSF weights from a copied formula that rounded
differently, a true division where the CPU multiplies by a refined
reciprocal estimate, the order of the masking threshold's terms, `float`
constants and a `float` gain limit where the CPU uses `double`, and a
floor of the frame sums at `1e-2` where the CPU's is `1e-10`. The last one
could report `adm2 = 1` where the CPU reports 0, for content with almost no
reference detail scored without the noise floor. The twin now runs the
CPU's arithmetic in the CPU's types, adds each row on the device and the
rows on the host in the CPU's order, and takes the weights, the reduced
region, the pooling and the floor from the CPU's own routines
([ADR-1420](docs/adr/1420-cuda-float-adm-cpu-arithmetic.md)). The CPU's
division is built on the processor's `RCPSS` estimate, so the twin probes
that estimate when the extractor starts (about 10 ms) and evaluates it on
the device. Measured on an RTX 4090 at `--precision max`: every output of
every frame identical on the Netflix pair at 8, 10, 12 and 16 bits, both
1080p checkerboard pairs and BBB 3840x2160, also with `debug=true` and
with non-default `adm_enhn_gain_limit`, `adm_bypass_cm`,
`adm_noise_weight`, `adm_skip_aim_scale` and viewing geometry. The parity
gate compares this twin with tolerance 0. Not identical: `adm_p_norm`
other than 1 or 3, where the twin is within 1.1e-7 of the CPU (the two
`powf` implementations differ). A run of the twin alone takes 1.98 ms per
3840x2160 frame instead of 1.87 ms; its kernels take 1.11 ms instead of
0.76 ms. Stored `float_adm_cuda` outputs change in their low digits by at
most 1.3e-5. The SYCL, HIP and Metal twins still agree with the CPU to
four decimal places.
23 changes: 23 additions & 0 deletions core/src/feature/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -1008,6 +1008,29 @@ feature/
[Research-0024](../../../docs/research/0024-vif-upstream-divergence.md)
= history of the table era.

- **Float ADM reference exports for GPU twins** (ADR-1420,
[`adm_float_reference.h`](adm_float_reference.h)). `adm_tools.c`
exports `adm_border_s()`, `adm_csf_rfactor_s()`,
`adm_pool_bands_s()`, `adm_decouple_cos_1deg_sq_s()`,
`adm_divs_is_reciprocal_s()`, `adm_divs_reciprocal_estimate_s()`.
`float_adm_cuda.c` calls them instead of copying. Keep: the four
reductions (`adm_csf_den_scale_s[_p3]`, `adm_cm_s[_p3]`) end in
`adm_pool_bands_s()`; `rcp_s()` and the exported estimate share
`rcp_estimate_s()`. Twin copies drift: old CUDA copy of
`dwt_quant_step()` was 1-3 ulp off.
CPU float ADM = host-dependent: `DIVS()` on x86 (gcc / clang) builds
on the processor's `RCPSS` estimate, specified by error bound only
(`T-FLOAT-ADM-RECIPROCAL-ESTIMATE-HOST-DEPENDENT-2026-10-01`).
[`adm_reciprocal_model.{c,h}`](adm_reciprocal_model.h) = table model
of that estimate + host probe; device-compilable lookup. Twin types
that matter: gain limit `double`, `FLOAT_ONE_BY_30` / `_15` double
literals, threshold centre tap fifth, angle threshold
`(cos^2 * |o|^2) * |t|^2`, fp32 row + frame accumulators. Change any
-> change `cuda/float_adm/float_adm_device.h` same PR
(`test_float_adm_device_math` fails until it follows). SYCL / HIP /
Metal twins still old arithmetic:
`T-GPU-FLOAT-ADM-CPU-ARITHMETIC-2026-10-01`.

- **`compute_adm` signature stays on fork's parameter
list — Strategy E in Research-0024.** Netflix upstream
`4dcc2f7c` adds 12 new parameters (`luminance_level`,
Expand Down
68 changes: 68 additions & 0 deletions core/src/feature/adm_float_reference.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
/**
* Copyright 2016-2020 Netflix, Inc.
* Copyright 2026 Lusoris
* SPDX-License-Identifier: BSD-2-Clause-Patent
*
* The parts of the float ADM reference (adm_tools.c) that a GPU twin runs on
* the host instead of re-deriving: the reduced region, the CSF weights, the
* pooling of one scale's band accumulators, the decouple's angle threshold
* and the division the decouple uses (ADR-1420).
*
* A twin that computes any of these itself has a second implementation to
* keep in step with the reference. Calling them is what makes its result the
* reference's.
*/

#ifndef VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_
#define VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_

#include <stdbool.h>

#ifdef __cplusplus
extern "C" {
#endif

/* Half-open region [left, right) x [top, bottom) of one scale's bands. */
typedef struct AdmBorderS {
int left;
int top;
int right;
int bottom;
} AdmBorderS;

/* The region the reductions run over: `border_factor` of each frame edge is
* excluded. */
AdmBorderS adm_border_s(int w, int h, double border_factor);

/* CSF weights of DWT scale `scale`: rfactor[0..1] for the (h, v) bands,
* rfactor[2] for the (d) band. A negative adm_f1sN / adm_f2sN keeps the
* model's value. */
void adm_csf_rfactor_s(int scale, double adm_norm_view_dist, int adm_ref_display_height,
int adm_csf_mode, double luminance_level, double adm_csf_scale,
double adm_csf_diag_scale, double adm_f1s0, double adm_f1s1, double adm_f1s2,
double adm_f1s3, double adm_f2s0, double adm_f2s1, double adm_f2s2,
double adm_f2s3, float rfactor[3]);

/* One scale's value from its (h, v, d) accumulators over a region of
* region_w x region_h samples: each band's 1 / adm_p_norm root plus the noise
* floor, added in band order. */
float adm_pool_bands_s(const float accum[3], int region_w, int region_h, double adm_noise_weight,
double adm_p_norm);

/* cos(1 degree) squared as the fp32 the decouple's angle test compares with. */
float adm_decouple_cos_1deg_sq_s(void);

/* How the decouple divides in this build: true when it multiplies by a
* reciprocal refined from the processor's estimate (x86 with SSE2, gcc or
* clang), false when it is the IEEE quotient. */
bool adm_divs_is_reciprocal_s(void);

/* The estimate the reciprocal is refined from: the processor's RCPSS result
* where adm_divs_is_reciprocal_s(), the IEEE reciprocal otherwise. */
float adm_divs_reciprocal_estimate_s(float x);

#ifdef __cplusplus
}
#endif

#endif /* VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_ */
103 changes: 103 additions & 0 deletions core/src/feature/adm_reciprocal_model.c
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
/**
* Copyright 2026 Lusoris
* SPDX-License-Identifier: EUPL-1.2
*
* Host probe of the reciprocal estimate the float ADM decouple divides with
* (ADR-1420). See adm_reciprocal_model.h.
*/

#include <assert.h>
#include <stdbool.h>
#include <stdint.h>
#include <string.h>

#include "adm_float_reference.h"
#include "adm_reciprocal_model.h"

#define ADM_PROBE_MANTISSAS (1u << 23)
/* Mantissa stride of the per-exponent sweep: a prime, so the sampled
* mantissas do not line up with the table's buckets. */
#define ADM_PROBE_MANTISSA_STEP 4099u
#define ADM_PROBE_EXPONENTS 256u
#define ADM_PROBE_ONE_EXPONENT 127u

typedef uint32_t (*AdmReciprocalPredictor)(const uint32_t *table, uint32_t bits);

static uint32_t adm_bits_of(float x)
{
uint32_t bits;
memcpy(&bits, &x, sizeof(bits));
return bits;
}

static float adm_float_of(uint32_t bits)
{
float x;
memcpy(&x, &bits, sizeof(x));
return x;
}

static uint32_t adm_predict_ieee(const uint32_t *table, uint32_t bits)
{
(void)table;
return adm_bits_of(1.0f / adm_float_of(bits));
}

static bool adm_is_nan_bits(uint32_t bits)
{
return (bits & 0x7fffffffu) > 0x7f800000u;
}

/* The host's estimate of one input against a predictor. Two NaNs agree
* whatever their payloads: the decouple propagates a NaN, it never reads one. */
static bool adm_probe_input(AdmReciprocalPredictor predict, const uint32_t *table, uint32_t bits)
{
const uint32_t host = adm_bits_of(adm_divs_reciprocal_estimate_s(adm_float_of(bits)));
const uint32_t model = predict(table, bits);
return host == model || (adm_is_nan_bits(host) && adm_is_nan_bits(model));
}

/* Every mantissa at one exponent, then every exponent (zeros, denormals,
* infinities and NaNs included) and both signs at mantissas spread over the
* range and at the last one. */
static bool adm_probe_predictor(AdmReciprocalPredictor predict, const uint32_t *table)
{
bool agree = true;
for (uint32_t mantissa = 0u; mantissa < ADM_PROBE_MANTISSAS; mantissa++)
agree = adm_probe_input(predict, table, (ADM_PROBE_ONE_EXPONENT << 23) | mantissa) && agree;

for (uint32_t i = 0u; i < 2u * ADM_PROBE_EXPONENTS; i++) {
const uint32_t base = i << 23; /* sign and exponent */
for (uint32_t mantissa = 0u; mantissa < ADM_PROBE_MANTISSAS;
mantissa += ADM_PROBE_MANTISSA_STEP)
agree = adm_probe_input(predict, table, base | mantissa) && agree;
agree = adm_probe_input(predict, table, base | (ADM_PROBE_MANTISSAS - 1u)) && agree;
}
return agree;
}

void adm_reciprocal_model_probe(AdmReciprocalModel *model)
{
assert(model);
memset(model, 0, sizeof(*model));
model->division = ADM_DIVISION_IEEE;
model->reproduces_reference = true;
if (!adm_divs_is_reciprocal_s())
return;

const unsigned bucket_shift = 23u - ADM_RECIPROCAL_INDEX_BITS;
for (uint32_t i = 0u; i < ADM_RECIPROCAL_TABLE_SIZE; i++) {
const uint32_t in = (ADM_PROBE_ONE_EXPONENT << 23) | (i << bucket_shift);
model->table[i] = adm_bits_of(adm_divs_reciprocal_estimate_s(adm_float_of(in)));
}
if (adm_probe_predictor(adm_reciprocal_model_bits, model->table)) {
model->division = ADM_DIVISION_RECIPROCAL_TABLE;
return;
}

/* Not a table of the top mantissa bits. An emulator that computes the
* estimate as the IEEE reciprocal is still reproduced exactly; anything
* else is not. */
model->division = ADM_DIVISION_RECIPROCAL_IEEE;
model->reproduces_reference = adm_probe_predictor(adm_predict_ieee, model->table);
}
Loading
Loading