Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 3 additions & 19 deletions .standards-baseline.json
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
{
"version": 1,
"generated_at": "2026-10-01T20:30:59Z",
"generated_at": "2026-10-01T20:46:07Z",
"repository": "VMAFx/vmafx",
"commit_sha": "4ae168bc06e2531cf7811c7e0ab7d3f20f790650",
"total_infractions": 365,
"commit_sha": "a5acac3477c451e562782ed4e6f7cf5842e5e2ef",
"total_infractions": 363,
"infractions": [
{
"rule_id": "HISS-04",
Expand Down Expand Up @@ -2102,14 +2102,6 @@
"message": "Function 'vmaf_sycl_import_va_surface' (301 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/src/sycl/dmabuf_import.cpp:366:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_barten_csf.c",
"line_number": 48,
"symbol": "test_barten_csf",
"message": "Function 'test_barten_csf' (87 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_barten_csf.c:48:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_pelorus_interop.c",
Expand Down Expand Up @@ -2142,14 +2134,6 @@
"message": "Function 'test_x265_csv_reader' (79 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_pelorus_interop.c:562:HISS-04"
},
{
"rule_id": "HISS-04",
"file_path": "core/test/test_psnr_hvs_simd.c",
"line_number": 164,
"symbol": "ref_calc_psnrhvs",
"message": "Function 'ref_calc_psnrhvs' (115 LOC) exceeds HISS-04 / NASA Rule 4 limit of 60 LOC",
"fingerprint": "core/test/test_psnr_hvs_simd.c:164:HISS-04"
},
{
"rule_id": "HISS-07",
"file_path": "dev/scripts/fetch-intel-neo.py",
Expand Down
66 changes: 66 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -660,6 +660,20 @@
`git rebase` conflicts in CI (ADR-1383).


- **Ten core C and C++ test files conform to clang-tidy and HISS standards (part 1).**
The first batch of CPU and core test sources (`test_barten_csf.c`,
`test_psnr_hvs_simd.c`, `test_propagate_metadata.c`, `test_thread_pool.c`,
`test_context.c`, `test_predict.c`, `test_ciede.c`,
`test_pic_preallocation.c`, `test_locale_handling.c`, and `test_dict.cpp`)
were refactored to zero clang-tidy findings in the CPU lane under
[ADR-1142](docs/adr/1142-clang-tidy-debt-ratchet.md). Functions exceeding
branch, nesting, and statement thresholds were split into modular helpers
satisfying HISS-04 / NASA JPL Rule 4, eliminating 2 recorded infractions from
`.standards-baseline.json`. C23 `nullptr` diagnostics in C files are scoped under
[ADR-1138](docs/adr/1138-c23-nullptr-msvc-compat.md), and `test_dict.cpp` uses
anonymous namespaces and standard `nullptr`. All tests continue to pass.


- **Nine HIP test files conform to clang-tidy and HISS standards (part 1).**
The first batch of HIP test sources (`test_hip_smoke.c`,
`test_hip_float_adm_parity.c`, `test_hip_motion3_parity.c`,
Expand Down Expand Up @@ -1015,6 +1029,37 @@
`float` arithmetic and the `5e-3` tolerance.


- **`float_adm_cuda` returns the CPU's scores bit for bit.** The CUDA twin
of `float_adm` was up to 1.3e-5 from the CPU extractor (`adm_scale0` at
3840x2160) and matched it on 144 of 791 measured scores. Nine things
differed: the association of the angle test's threshold (the 1.3e-5), the
order of the sums, CSF weights from a copied formula that rounded
differently, a true division where the CPU multiplies by a refined
reciprocal estimate, the order of the masking threshold's terms, `float`
constants and a `float` gain limit where the CPU uses `double`, and a
floor of the frame sums at `1e-2` where the CPU's is `1e-10`. The last one
could report `adm2 = 1` where the CPU reports 0, for content with almost no
reference detail scored without the noise floor. The twin now runs the
CPU's arithmetic in the CPU's types, adds each row on the device and the
rows on the host in the CPU's order, and takes the weights, the reduced
region, the pooling and the floor from the CPU's own routines
([ADR-1420](docs/adr/1420-cuda-float-adm-cpu-arithmetic.md)). The CPU's
division is built on the processor's `RCPSS` estimate, so the twin probes
that estimate when the extractor starts (about 10 ms) and evaluates it on
the device. Measured on an RTX 4090 at `--precision max`: every output of
every frame identical on the Netflix pair at 8, 10, 12 and 16 bits, both
1080p checkerboard pairs and BBB 3840x2160, also with `debug=true` and
with non-default `adm_enhn_gain_limit`, `adm_bypass_cm`,
`adm_noise_weight`, `adm_skip_aim_scale` and viewing geometry. The parity
gate compares this twin with tolerance 0. Not identical: `adm_p_norm`
other than 1 or 3, where the twin is within 1.1e-7 of the CPU (the two
`powf` implementations differ). A run of the twin alone takes 1.98 ms per
3840x2160 frame instead of 1.87 ms; its kernels take 1.11 ms instead of
0.76 ms. Stored `float_adm_cuda` outputs change in their low digits by at
most 1.3e-5. The SYCL, HIP and Metal twins still agree with the CPU to
four decimal places.


- **`float_motion_cuda` returns the CPU's scores bit for bit.** The CPU
`float_motion` extractor adds the absolute differences of a row into one
`float`, the row sums into another, and divides in `float`, so its score
Expand Down Expand Up @@ -1162,6 +1207,27 @@
places or better.


- **`ssimulacra2_cuda` returns the CPU extractor's score bit for bit.** The
CUDA twin of `ssimulacra2` computed the CPU's per-pixel terms but added
them in a tree, where the CPU adds them one after the other into one
`double`; every add rounds, so the two ended a few units in the last place
apart. Measured on an RTX 4090 at `--precision max`, 8 of 113 frames
matched and the rest were up to 7.3e-11 away. The twin now forms the sums
of the CPU's loops on the device
([ADR-1433](docs/adr/1433-cuda-ssimulacra2-cpu-sum-order.md)): while a
running sum stays between two powers of two, adding a term moves it by a
whole number of steps, so the device adds those whole numbers per
1024-pixel chunk in parallel, one pass over the chunks puts them together,
and the few chunks in which the sum passes a power of two are added term by
term. All 113 frames are identical (Netflix 576x324 at 8, 10, 12 and 16
bits, both 1080p checkerboard pairs, BBB 3840x2160), and the parity gate
compares the cell at 0 instead of `5e-3`. The price is time: a 3840x2160
frame takes 15.6 ms instead of 7.8 ms and a 576x324 frame 1.7 ms instead of
0.4 ms; the CPU extractor takes 126 ms per 4K frame on sixteen threads.
Stored `ssimulacra2_cuda` scores change by up to 7.3e-11. The SYCL and HIP
twins keep their tree sums and the `5e-3` tolerance.


- **Several CUDA instances on one device no longer get wrong `vif` scores.**
`integer_vif_cuda` cleared its accumulators on its private stream while the
scale 0 kernels that add into them ran on the picture stream, with nothing
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -184,7 +184,7 @@ generated agent context against `AGENTS.md`.
private scratch-link policy.

**Debt Baseline**: `.standards-baseline.json` anchors the debt ratchet at
365 recorded infractions; audit forbids growth.
363 recorded infractions; audit forbids growth.

[praetor-docs-badge]: https://github.com/vmafx/vmafx/actions/workflows/praetor-docs.yml/badge.svg
[praetor-docs-runs]: https://github.com/vmafx/vmafx/actions/workflows/praetor-docs.yml
Expand Down
12 changes: 12 additions & 0 deletions changelog.d/changed/std-core-tests-a.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
- **Ten core C and C++ test files conform to clang-tidy and HISS standards (part 1).**
The first batch of CPU and core test sources (`test_barten_csf.c`,
`test_psnr_hvs_simd.c`, `test_propagate_metadata.c`, `test_thread_pool.c`,
`test_context.c`, `test_predict.c`, `test_ciede.c`,
`test_pic_preallocation.c`, `test_locale_handling.c`, and `test_dict.cpp`)
were refactored to zero clang-tidy findings in the CPU lane under
[ADR-1142](docs/adr/1142-clang-tidy-debt-ratchet.md). Functions exceeding
branch, nesting, and statement thresholds were split into modular helpers
satisfying HISS-04 / NASA JPL Rule 4, eliminating 2 recorded infractions from
`.standards-baseline.json`. C23 `nullptr` diagnostics in C files are scoped under
[ADR-1138](docs/adr/1138-c23-nullptr-msvc-compat.md), and `test_dict.cpp` uses
anonymous namespaces and standard `nullptr`. All tests continue to pass.
29 changes: 29 additions & 0 deletions changelog.d/fixed/cuda-float-adm-cpu-arithmetic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
- **`float_adm_cuda` returns the CPU's scores bit for bit.** The CUDA twin
of `float_adm` was up to 1.3e-5 from the CPU extractor (`adm_scale0` at
3840x2160) and matched it on 144 of 791 measured scores. Nine things
differed: the association of the angle test's threshold (the 1.3e-5), the
order of the sums, CSF weights from a copied formula that rounded
differently, a true division where the CPU multiplies by a refined
reciprocal estimate, the order of the masking threshold's terms, `float`
constants and a `float` gain limit where the CPU uses `double`, and a
floor of the frame sums at `1e-2` where the CPU's is `1e-10`. The last one
could report `adm2 = 1` where the CPU reports 0, for content with almost no
reference detail scored without the noise floor. The twin now runs the
CPU's arithmetic in the CPU's types, adds each row on the device and the
rows on the host in the CPU's order, and takes the weights, the reduced
region, the pooling and the floor from the CPU's own routines
([ADR-1420](docs/adr/1420-cuda-float-adm-cpu-arithmetic.md)). The CPU's
division is built on the processor's `RCPSS` estimate, so the twin probes
that estimate when the extractor starts (about 10 ms) and evaluates it on
the device. Measured on an RTX 4090 at `--precision max`: every output of
every frame identical on the Netflix pair at 8, 10, 12 and 16 bits, both
1080p checkerboard pairs and BBB 3840x2160, also with `debug=true` and
with non-default `adm_enhn_gain_limit`, `adm_bypass_cm`,
`adm_noise_weight`, `adm_skip_aim_scale` and viewing geometry. The parity
gate compares this twin with tolerance 0. Not identical: `adm_p_norm`
other than 1 or 3, where the twin is within 1.1e-7 of the CPU (the two
`powf` implementations differ). A run of the twin alone takes 1.98 ms per
3840x2160 frame instead of 1.87 ms; its kernels take 1.11 ms instead of
0.76 ms. Stored `float_adm_cuda` outputs change in their low digits by at
most 1.3e-5. The SYCL, HIP and Metal twins still agree with the CPU to
four decimal places.
19 changes: 19 additions & 0 deletions changelog.d/fixed/cuda-ssimulacra2-cpu-sum-order.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
- **`ssimulacra2_cuda` returns the CPU extractor's score bit for bit.** The
CUDA twin of `ssimulacra2` computed the CPU's per-pixel terms but added
them in a tree, where the CPU adds them one after the other into one
`double`; every add rounds, so the two ended a few units in the last place
apart. Measured on an RTX 4090 at `--precision max`, 8 of 113 frames
matched and the rest were up to 7.3e-11 away. The twin now forms the sums
of the CPU's loops on the device
([ADR-1433](docs/adr/1433-cuda-ssimulacra2-cpu-sum-order.md)): while a
running sum stays between two powers of two, adding a term moves it by a
whole number of steps, so the device adds those whole numbers per
1024-pixel chunk in parallel, one pass over the chunks puts them together,
and the few chunks in which the sum passes a power of two are added term by
term. All 113 frames are identical (Netflix 576x324 at 8, 10, 12 and 16
bits, both 1080p checkerboard pairs, BBB 3840x2160), and the parity gate
compares the cell at 0 instead of `5e-3`. The price is time: a 3840x2160
frame takes 15.6 ms instead of 7.8 ms and a 576x324 frame 1.7 ms instead of
0.4 ms; the CPU extractor takes 126 ms per 4K frame on sixteen threads.
Stored `ssimulacra2_cuda` scores change by up to 7.3e-11. The SYCL and HIP
twins keep their tree sums and the `5e-3` tolerance.
23 changes: 23 additions & 0 deletions core/src/feature/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -1008,6 +1008,29 @@ feature/
[Research-0024](../../../docs/research/0024-vif-upstream-divergence.md)
= history of the table era.

- **Float ADM reference exports for GPU twins** (ADR-1420,
[`adm_float_reference.h`](adm_float_reference.h)). `adm_tools.c`
exports `adm_border_s()`, `adm_csf_rfactor_s()`,
`adm_pool_bands_s()`, `adm_decouple_cos_1deg_sq_s()`,
`adm_divs_is_reciprocal_s()`, `adm_divs_reciprocal_estimate_s()`.
`float_adm_cuda.c` calls them instead of copying. Keep: the four
reductions (`adm_csf_den_scale_s[_p3]`, `adm_cm_s[_p3]`) end in
`adm_pool_bands_s()`; `rcp_s()` and the exported estimate share
`rcp_estimate_s()`. Twin copies drift: old CUDA copy of
`dwt_quant_step()` was 1-3 ulp off.
CPU float ADM = host-dependent: `DIVS()` on x86 (gcc / clang) builds
on the processor's `RCPSS` estimate, specified by error bound only
(`T-FLOAT-ADM-RECIPROCAL-ESTIMATE-HOST-DEPENDENT-2026-10-01`).
[`adm_reciprocal_model.{c,h}`](adm_reciprocal_model.h) = table model
of that estimate + host probe; device-compilable lookup. Twin types
that matter: gain limit `double`, `FLOAT_ONE_BY_30` / `_15` double
literals, threshold centre tap fifth, angle threshold
`(cos^2 * |o|^2) * |t|^2`, fp32 row + frame accumulators. Change any
-> change `cuda/float_adm/float_adm_device.h` same PR
(`test_float_adm_device_math` fails until it follows). SYCL / HIP /
Metal twins still old arithmetic:
`T-GPU-FLOAT-ADM-CPU-ARITHMETIC-2026-10-01`.

- **`compute_adm` signature stays on fork's parameter
list — Strategy E in Research-0024.** Netflix upstream
`4dcc2f7c` adds 12 new parameters (`luminance_level`,
Expand Down
68 changes: 68 additions & 0 deletions core/src/feature/adm_float_reference.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
/**
* Copyright 2016-2020 Netflix, Inc.
* Copyright 2026 Lusoris
* SPDX-License-Identifier: BSD-2-Clause-Patent
*
* The parts of the float ADM reference (adm_tools.c) that a GPU twin runs on
* the host instead of re-deriving: the reduced region, the CSF weights, the
* pooling of one scale's band accumulators, the decouple's angle threshold
* and the division the decouple uses (ADR-1420).
*
* A twin that computes any of these itself has a second implementation to
* keep in step with the reference. Calling them is what makes its result the
* reference's.
*/

#ifndef VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_
#define VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_

#include <stdbool.h>

#ifdef __cplusplus
extern "C" {
#endif

/* Half-open region [left, right) x [top, bottom) of one scale's bands. */
typedef struct AdmBorderS {
int left;
int top;
int right;
int bottom;
} AdmBorderS;

/* The region the reductions run over: `border_factor` of each frame edge is
* excluded. */
AdmBorderS adm_border_s(int w, int h, double border_factor);

/* CSF weights of DWT scale `scale`: rfactor[0..1] for the (h, v) bands,
* rfactor[2] for the (d) band. A negative adm_f1sN / adm_f2sN keeps the
* model's value. */
void adm_csf_rfactor_s(int scale, double adm_norm_view_dist, int adm_ref_display_height,
int adm_csf_mode, double luminance_level, double adm_csf_scale,
double adm_csf_diag_scale, double adm_f1s0, double adm_f1s1, double adm_f1s2,
double adm_f1s3, double adm_f2s0, double adm_f2s1, double adm_f2s2,
double adm_f2s3, float rfactor[3]);

/* One scale's value from its (h, v, d) accumulators over a region of
* region_w x region_h samples: each band's 1 / adm_p_norm root plus the noise
* floor, added in band order. */
float adm_pool_bands_s(const float accum[3], int region_w, int region_h, double adm_noise_weight,
double adm_p_norm);

/* cos(1 degree) squared as the fp32 the decouple's angle test compares with. */
float adm_decouple_cos_1deg_sq_s(void);

/* How the decouple divides in this build: true when it multiplies by a
* reciprocal refined from the processor's estimate (x86 with SSE2, gcc or
* clang), false when it is the IEEE quotient. */
bool adm_divs_is_reciprocal_s(void);

/* The estimate the reciprocal is refined from: the processor's RCPSS result
* where adm_divs_is_reciprocal_s(), the IEEE reciprocal otherwise. */
float adm_divs_reciprocal_estimate_s(float x);

#ifdef __cplusplus
}
#endif

#endif /* VMAF_SRC_FEATURE_ADM_FLOAT_REFERENCE_H_ */
Loading
Loading