Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/test-suites.json
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,7 @@
"scripts/ci/tests/test_helm_controller_auth.py",
"scripts/ci/tests/test_helm_controller_workload.py",
"scripts/ci/tests/test_helm_node_contract.py",
"scripts/ci/tests/test_helm_node_fuse_ebpf.py",
"scripts/ci/tests/test_helm_service_accounts.py"
],
"checks": ["helm lint + template"],
Expand Down
18 changes: 18 additions & 0 deletions .github/workflows/docker-publish-operator-node.yml
Original file line number Diff line number Diff line change
Expand Up @@ -1465,6 +1465,24 @@ jobs:
docker run --rm --entrypoint /usr/local/bin/ffmpeg "${image}" -version
# Remote storage inputs run rclone (pkg/storage, ADR-0719).
docker run --rm --entrypoint /usr/local/bin/rclone "${image}" version
# Mount mode (ADR-1593): as UID 65532, with only the bounding-set
# capabilities node.fuse grants, the setuid fusermount3 and the
# util-linux mount it runs mount an rclone remote; the file reads
# back through FUSE and unmounts. AppArmor (on the runner) denies
# mount under Docker's default profile.
mounted="vmafx-node-mount-smoke"
docker run --detach --name "${mounted}" --device /dev/fuse \
--cap-drop ALL --cap-add SYS_ADMIN --cap-add DAC_READ_SEARCH \
--security-opt apparmor=unconfined --read-only --tmpfs /tmp:mode=1777 \
-v "${frames}:/frames:ro" --entrypoint /usr/local/bin/rclone "${image}" \
rcd --rc-no-auth --rc-addr 127.0.0.1:5572
trap 'docker rm -f "${mounted}" > /dev/null' EXIT
docker exec "${mounted}" rclone mkdir /tmp/m
timeout 60 docker exec "${mounted}" \
rclone mount :local:/frames /tmp/m --daemon --vfs-cache-mode off
test "$(docker exec "${mounted}" rclone md5sum /tmp/m/ref.yuv | cut -d' ' -f1)" = \
"$(md5sum "${frames}/ref.yuv" | cut -d' ' -f1)"
docker exec "${mounted}" /usr/bin/fusermount3 -u /tmp/m

# -------------------------------------------------------------------------
# Summary gate — CI shows a single green/red status
Expand Down
6 changes: 6 additions & 0 deletions .github/workflows/helm-chart.yml
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,12 @@ jobs:
- name: Service accounts and tenant RBAC
run: python3 -B -m unittest discover -s scripts/ci/tests -p 'test_helm_service_accounts.py' -v

# ADR-1593: node.fuse (FUSE device resource, the setuid helper's
# capability bounding set) and node.ebpf (tracker env, UID 0 with BPF /
# PERFMON / SYS_ADMIN, tracefs), and the combinations the chart refuses.
- name: Node FUSE and eBPF renders
run: python3 -B -m unittest discover -s scripts/ci/tests -p 'test_helm_node_fuse_ebpf.py' -v

# ADR-1353: the server Deployment selected every Pod of the release, the
# operator's and node's included. Render each server workload with every
# other component on and require each selector to pick only its own Pods.
Expand Down
52 changes: 52 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,16 @@
controller workload").


- **Helm: `node.fuse` and `node.ebpf`
([ADR-1593](docs/adr/1593-helm-node-fuse-and-ebpf.md)).** `node.fuse` gives
the node pods `/dev/fuse` through a FUSE device plugin's resource and the
capability bounding set mount mode needs; `node.ebpf` turns on the eBPF
descriptor tracker (`VMAFX_EBPF_BYPASS`) with UID 0, `BPF`, `PERFMON` and
`SYS_ADMIN` and the host's tracefs read-only. `storage.mode: mount` without
`node.fuse` is now refused at render time; it used to deploy a node that
could not start.


- **`scripts/dev/hip_dispatch_drop_probe.hip` checks whether an AMD GPU runs
every command of a HIP stream.** Built with `hipcc`, it runs frames of one
memset, several small kernels and a readback on one stream and reports the
Expand Down Expand Up @@ -1307,6 +1317,16 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
[build flags](docs/development/build-flags.md#floating-point-contraction-is-off-everywhere).


- **Metal kernels compile without fast math or FP contraction**
(ADR-1498): every `.metal` file takes `-fno-fast-math -ffp-contract=off`,
so fp32 `+ - * /`, `sqrt` and `fma` are correctly rounded and no `a * b + c`
is fused behind the source's back, the policy the CUDA, HIP and SYCL
kernels already follow. The arithmetic of a ported Metal twin lives in a
header on `core/src/feature/metal/metal_portable.h` that also compiles on
the host, where a test holds it against the CPU extractor. Guide:
`docs/backends/metal/index.md`.


- **The tiny model cards quote the terms of the data each model was trained
on.** `fr_regressor_v1` to `v3`, `vmaf_tiny_v1` to `v4`, `nr_metric_v1`,
`learned_filter_v1`, `saliency_student_v1` and `v2` and `lpips_sq_v1` were
Expand Down Expand Up @@ -3733,6 +3753,28 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
stops `test_integer_vif_sv_sq`.


- **The Metal twins compute the CPU's scores by construction
([ADR-1498](docs/adr/1498-metal-twins-exact-designs.md)).** Each Metal
extractor now runs the design that makes its CUDA, HIP or SYCL twin return
the CPU's scores bit for bit, written in a header that also compiles on the
host, where a test holds it against the CPU extractor. Changes a Metal user
can see: `float_psnr_metal` and `float_moment_metal` are exact at 10 to 16
bits, `float_moment_metal` also on 16-bit frames whose sum passes 2^53
units; `integer_psnr_metal` no longer loses carries in its 64-bit error sum
and takes `enable_apsnr`; `integer_motion_metal` differences frames before
the blur, emits `motion_sad_score` and `motion3` and drops the
`motion_add_uv` option the CPU never had; `motion_v2_metal` and
`float_motion_metal` apply `motion_fps_weight` and `motion_max_val` per frame
and take every CPU option; `integer_vif_metal` hands frames below 16 pixels
to the CPU in a model run; `float_adm_metal` refuses frames below 17x17,
floors its sums as the CPU and takes `adm_f1s0` to `adm_f2s3`;
`float_vif_metal` runs every `vif_kernelscale`; `float_ssim_metal` gives an
identical flat frame a finite `enable_db` score; `integer_cambi_metal` takes
`src_width`, `src_height` and `full_ref`. No Apple device ran these builds
yet: each twin's state row closes when a report of the macOS tester bundle
shows its parity test passing. Guide: `docs/backends/metal/index.md`.


- **The licence notices name the right holders of the bundled models.** The
LPIPS-SqueezeNet weights are recorded as BSD-2-Clause (Zhang, Isola, Efros,
Shechtman, Wang) on torchvision's BSD-3-Clause SqueezeNet features, the
Expand Down Expand Up @@ -3789,6 +3831,16 @@ make `core/AGENTS.md` a generated index over `AGENTS.d/` topic pages ([ADR-1454]
`Windows MSVC+CUDA (full)` CI lane now runs `test_float_adm_x86`.


- **Mount mode works in the node image
([ADR-1593](docs/adr/1593-helm-node-fuse-and-ebpf.md)).** The published
`vmafx-node` image had no FUSE helper, so a node with
`VMAFX_STORAGE_MODE=mount` refused to start. The image now carries the
setuid `fusermount3` and the util-linux `mount` and `umount` it runs, listed
in the licence record with their Debian sources in the `-source` image. A
container needs `/dev/fuse` and the capabilities `SYS_ADMIN` and
`DAC_READ_SEARCH`; the node process itself keeps none.


- **vmafx-operator authenticates to the controller
([ADR-1569](docs/adr/1569-operator-controller-auth.md)).** The `VmafxJob`
reconciler polled `GetJob` without a token, so with the controller's auth
Expand Down
8 changes: 8 additions & 0 deletions changelog.d/added/helm-node-fuse-and-ebpf.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
- **Helm: `node.fuse` and `node.ebpf`
([ADR-1593](docs/adr/1593-helm-node-fuse-and-ebpf.md)).** `node.fuse` gives
the node pods `/dev/fuse` through a FUSE device plugin's resource and the
capability bounding set mount mode needs; `node.ebpf` turns on the eBPF
descriptor tracker (`VMAFX_EBPF_BYPASS`) with UID 0, `BPF`, `PERFMON` and
`SYS_ADMIN` and the host's tracefs read-only. `storage.mode: mount` without
`node.fuse` is now refused at render time; it used to deploy a node that
could not start.
8 changes: 8 additions & 0 deletions changelog.d/changed/metal-twins-exact.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
- **Metal kernels compile without fast math or FP contraction**
(ADR-1498): every `.metal` file takes `-fno-fast-math -ffp-contract=off`,
so fp32 `+ - * /`, `sqrt` and `fma` are correctly rounded and no `a * b + c`
is fused behind the source's back, the policy the CUDA, HIP and SYCL
kernels already follow. The arithmetic of a ported Metal twin lives in a
header on `core/src/feature/metal/metal_portable.h` that also compiles on
the host, where a test holds it against the CPU extractor. Guide:
`docs/backends/metal/index.md`.
20 changes: 20 additions & 0 deletions changelog.d/fixed/metal-twins-exact.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
- **The Metal twins compute the CPU's scores by construction
([ADR-1498](docs/adr/1498-metal-twins-exact-designs.md)).** Each Metal
extractor now runs the design that makes its CUDA, HIP or SYCL twin return
the CPU's scores bit for bit, written in a header that also compiles on the
host, where a test holds it against the CPU extractor. Changes a Metal user
can see: `float_psnr_metal` and `float_moment_metal` are exact at 10 to 16
bits, `float_moment_metal` also on 16-bit frames whose sum passes 2^53
units; `integer_psnr_metal` no longer loses carries in its 64-bit error sum
and takes `enable_apsnr`; `integer_motion_metal` differences frames before
the blur, emits `motion_sad_score` and `motion3` and drops the
`motion_add_uv` option the CPU never had; `motion_v2_metal` and
`float_motion_metal` apply `motion_fps_weight` and `motion_max_val` per frame
and take every CPU option; `integer_vif_metal` hands frames below 16 pixels
to the CPU in a model run; `float_adm_metal` refuses frames below 17x17,
floors its sums as the CPU and takes `adm_f1s0` to `adm_f2s3`;
`float_vif_metal` runs every `vif_kernelscale`; `float_ssim_metal` gives an
identical flat frame a finite `enable_db` score; `integer_cambi_metal` takes
`src_width`, `src_height` and `full_ref`. No Apple device ran these builds
yet: each twin's state row closes when a report of the macOS tester bundle
shows its parity test passing. Guide: `docs/backends/metal/index.md`.
8 changes: 8 additions & 0 deletions changelog.d/fixed/node-image-fuse-mount-mode.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
- **Mount mode works in the node image
([ADR-1593](docs/adr/1593-helm-node-fuse-and-ebpf.md)).** The published
`vmafx-node` image had no FUSE helper, so a node with
`VMAFX_STORAGE_MODE=mount` refused to start. The image now carries the
setuid `fusermount3` and the util-linux `mount` and `umount` it runs, listed
in the licence record with their Debian sources in the `-source` image. A
container needs `/dev/fuse` and the capabilities `SYS_ADMIN` and
`DAC_READ_SEARCH`; the node process itself keeps none.
22 changes: 11 additions & 11 deletions core/src/feature/adm_gain_limit.h
Original file line number Diff line number Diff line change
Expand Up @@ -18,9 +18,9 @@
* nothing else from floating point.
*
* adm_gain_limit_product() returns that value from 64-bit integer arithmetic,
* for a backend whose device code may not hold a double (SYCL, ADR-0220). It
* is not an approximation: for every limit the option admits and every int32
* sample it returns what the scalar's double product truncates to.
* for a backend whose device code may not hold a double (SYCL, ADR-0220; Metal,
* ADR-1498). It is not an approximation: for every limit the option admits and
* every int32 sample it returns what the scalar's double product truncates to.
*
* Derivation. Write gain = M * 2^-s with the 53-bit significand M in
* [2^52, 2^53), and a = |rst| <= 2^31. The exact product is P * 2^-s with
Expand All @@ -42,23 +42,23 @@

#ifndef LIBVMAF_FEATURE_ADM_GAIN_LIMIT_H_
#define LIBVMAF_FEATURE_ADM_GAIN_LIMIT_H_

#if !defined(__METAL_VERSION__)
#include <math.h>
#include <stdint.h>

#endif
/* The gain limit as significand * 2^-frac_bits, the significand in two halves
* so that a device without 128-bit integers can multiply by it. No typedef:
* the header is included from C and from C++ (SYCL) translation units, and
* `struct AdmGainLimit` reads the same in both. */
* the header is included from C, C++ (SYCL) and Metal Shading Language units
* (the includer defines UINT64_C there), and `struct AdmGainLimit` reads alike. */
struct AdmGainLimit {
uint32_t m_hi; /* significand bits 52..32 */
uint32_t m_lo; /* significand bits 31..0 */
int32_t frac_bits; /* 46 for a limit of 100, 52 for a limit of 1 */
};

/* Splits `gain` on the host. adm_gain_limit_product() needs
* 32 <= frac_bits <= 54, that is a limit in [0.25, 2^21); the option admits
* [1, 100]. */
#if !defined(__METAL_VERSION__) /* Metal has no double: the host splits. */
/* Splits `gain` on the host. adm_gain_limit_product() needs 32 <= frac_bits
* <= 54, that is a limit in [0.25, 2^21); the option admits [1, 100]. */
static inline struct AdmGainLimit adm_gain_limit_split(double gain)
{
int exponent = 0;
Expand All @@ -69,7 +69,7 @@ static inline struct AdmGainLimit adm_gain_limit_split(double gain)
.frac_bits = 53 - exponent};
return g;
}

#endif
/* (int64_t)((double)rst * gain): the double product of a sample and the gain
* limit, truncated toward zero. Integer arithmetic only. */
static inline int64_t adm_gain_limit_product(int32_t rst, struct AdmGainLimit g)
Expand Down
50 changes: 46 additions & 4 deletions core/src/feature/ciede_ff_math.h
Original file line number Diff line number Diff line change
Expand Up @@ -41,15 +41,26 @@
* make_constants(), which use fp64: they are host code, or evaluated by the
* compiler where a backend's host is C (both are constexpr). A translation
* unit that includes this header is compiled with contraction off.
*
* The Metal twin compiles the subset Metal Shading Language and host C++
* share (VMAF_FF_MSL_SUBSET, see ff_math.h; ADR-1498). Metal has no fp64
* type at all, so make_pair() and make_constants() do not exist in a Metal
* kernel: the twin's host evaluates them and hands the kernel the result.
* Every reference in Metal names its address space: the backend defines
* VMAF_FF_REF_SPACE, the space of the Constants and Tables the functions
* below take by reference, and this header spells the two parameter types
* with it. For the other backends the preprocessed header is what it was
* before the subset existed.
*/

#ifndef VMAF_FEATURE_CIEDE_FF_MATH_H_
#define VMAF_FEATURE_CIEDE_FF_MATH_H_

#if !defined(__METAL_VERSION__)
#include <cstddef>
#include <cstdint>
#include <numbers>

#endif
#include "ff_math.h"

#if !defined(VMAF_FF_LDEXP)
Expand All @@ -76,9 +87,9 @@ using vmaf_ffm_base::two_prod;
using vmaf_ffm_base::two_sum;

/* powf(25., 7): 25^7 = 6103515625 rounded to float. */
inline constexpr float kPowf25To7 = 6103515648.0f;
VMAF_FF_CONSTANT float kPowf25To7 = 6103515648.0f;
/* pow(25, 7), exactly: 6103515648 - 23. */
inline constexpr Ff kPow25To7 = {.hi = 6103515648.0f, .lo = -23.0f};
VMAF_FF_CONSTANT Ff kPow25To7 = VMAF_FF_PAIR_INIT(6103515648.0f, -23.0f);

/* ciede.c's fp64 constants as pairs. make_constants() builds them from the
* reference's own expressions, on the host or at compile time. */
Expand Down Expand Up @@ -128,6 +139,17 @@ struct Lab {
float b;
};

#if defined(VMAF_FF_MSL_SUBSET)
#if !defined(VMAF_FF_REF_SPACE)
#error "ciede_ff_math.h: VMAF_FF_MSL_SUBSET needs VMAF_FF_REF_SPACE (see the header comment)"
#endif
/* `const Constants &k` and `const Tables &tables` below name their address
* space, up to the #undef after pixel(). A macro is not expanded again inside
* its own expansion. */
#define Constants VMAF_FF_REF_SPACE Constants
#define Tables VMAF_FF_REF_SPACE Tables
#endif

/* pow(x, 2) for a float x: the exact fp64 square. */
VMAF_FF_INLINE Ff sq(float x)
{
Expand Down Expand Up @@ -183,9 +205,15 @@ VMAF_FF_INLINE Lab lab_color(float y, float u, float v, const Constants &k)

/* The three results are exact in fp64 before the rounding to float:
* products of a float with a small integer, sums of two floats. */
#if defined(VMAF_FF_MSL_SUBSET)
// NOLINTNEXTLINE(modernize-use-designated-initializers): MSL has none, ADR-1498
return {to_float(add_f(two_prod(116.0f, fy), -16.0f)),
to_float(mul_f(two_sum(fx, -fy), 500.0f)), to_float(mul_f(two_sum(fy, -fz), 200.0f))};
#else
return {.l = to_float(add_f(two_prod(116.0f, fy), -16.0f)),
.a = to_float(mul_f(two_sum(fx, -fy), 500.0f)),
.b = to_float(mul_f(two_sum(fy, -fz), 200.0f))};
#endif
}

/* ------------------------------------------------------------------ */
Expand Down Expand Up @@ -241,7 +269,7 @@ VMAF_FF_INLINE float upcase_h_bar_prime(float h_prime_1, float h_prime_2, const
/* get_upcase_t() */
VMAF_FF_INLINE float upcase_t(float h, const Constants &k, const Tables &tables)
{
const float *table = tables.sin_cos;
const VMAF_FF_TABLE_SPACE float *table = tables.sin_cos;
const Ff cos_1 = vmaf_ffm::sin_cos(sub(from_float(h), k.pi_over_6), table).cos;
const Ff cos_2 = vmaf_ffm::sin_cos(from_float(2.0f * h), table).cos;
const Ff cos_3 = vmaf_ffm::sin_cos(ff_add(two_prod(3.0f, h), k.pi_over_30), table).cos;
Expand Down Expand Up @@ -334,10 +362,17 @@ VMAF_FF_INLINE float pixel(Samples ref, Samples dis, const Constants &k, const T
return delta_e(c1, c2, k, tables);
}

#if defined(VMAF_FF_MSL_SUBSET)
#undef Constants
#undef Tables
#endif

/* ------------------------------------------------------------------ */
/* Host code, or the compiler's (constexpr) */
/* ------------------------------------------------------------------ */

#if !defined(__METAL_VERSION__)

/* An fp64 value as a pair: hi its nearest float, lo what remains. */
VMAF_FF_INLINE constexpr Ff make_pair(double value)
{
Expand All @@ -350,7 +385,12 @@ VMAF_FF_INLINE constexpr Ff make_pair(double value)
VMAF_FF_INLINE constexpr Constants make_constants(unsigned bpc)
{
const double pi = std::numbers::pi; /* the constant ciede.c names */
#if defined(VMAF_FF_MSL_SUBSET)
/* The same power of two, shifted unsigned. */
const double depth = (double)(1u << (bpc - 8u));
#else
const double depth = (double)(1 << (bpc - 8u));
#endif
Constants k = {};
k.luma_offset = (float)(16. * depth);
k.chroma_offset = (float)(128. * depth);
Expand Down Expand Up @@ -393,6 +433,8 @@ VMAF_FF_INLINE constexpr Constants make_constants(unsigned bpc)
return k;
}

#endif /* !__METAL_VERSION__ */

} // namespace vmaf_ciede_ff

#endif /* VMAF_FEATURE_CIEDE_FF_MATH_H_ */
6 changes: 3 additions & 3 deletions core/src/feature/common/convolution_internal.h
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ FORCE_INLINE float convolution_edge_s(bool horizontal, const float *filter, int
const float *src, int width, int height, int stride, int i,
int j)
{
int radius = filter_width / 2;
const int radius = filter_width / 2;

float accum = 0;
for (int k = 0; k < filter_width; ++k) {
Expand Down Expand Up @@ -104,7 +104,7 @@ FORCE_INLINE float convolution_edge_sq_s(bool horizontal, const float *filter, i
const float *src, int width, int height, int stride, int i,
int j)
{
int radius = filter_width / 2;
const int radius = filter_width / 2;

float accum = 0;
float src_val;
Expand Down Expand Up @@ -132,7 +132,7 @@ FORCE_INLINE float convolution_edge_xy_s(bool horizontal, const float *filter, i
const float *src1, const float *src2, int width,
int height, int stride1, int stride2, int i, int j)
{
int radius = filter_width / 2;
const int radius = filter_width / 2;

float accum = 0;
float src_val1;
Expand Down
Loading
Loading