Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 71 additions & 0 deletions docs/adr/0779-ebpf-fuse-bypass-rclone.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
<!-- markdownlint-disable MD013 MD041 MD060 -->
# ADR-0779: eBPF FUSE bypass for rclone zero-copy path in vmafx-node

- **Status**: Proposed
- **Date**: 2026-06-03
- **Deciders**: Lusoris
- **Tags**: `ci`, `go`, `ebpf`, `rclone`, `performance`, `security`, `supply-chain`

## Context

The vmafx-node Go worker fetches media clips from object storage via an rclone FUSE
mount. Each clip open under the FUSE mount incurs a FUSE kernel round-trip plus an
rclone network fetch even for clips that have been pulled to a local cache since the
previous run. At 1080p clip sizes this round-trip adds 370 ms p50 latency per clip
open (Research-0733). With a 150 k-clip dataset this dominates wall time.

A probe-only eBPF tracepoint program can track file-descriptor opens under the rclone
FUSE mount prefix without the overhead of a full FUSE intercept. By maintaining a BPF
hash map of bypass-eligible FDs and draining events via a ring buffer, the in-process
cache stays warm without polling, collapsing warm-cache clip-open latency to ~10 ms
(37× improvement). Research-0733 measured this on the fork's RTX 4090 dev machine.

The feature is gated behind `VMAFX_EBPF_BYPASS=1` (default off) and requires
Linux 5.15+ and `CAP_BPF`. A compile-time stub (`rclone_bypass_stub.go`) allows CI
builds without a BPF toolchain, preserving cross-platform compatibility.

## Decision

We will ship a probe-only eBPF tracepoint program (`rclone_bypass.bpf.c`) plus a Go
loader using `cilium/ebpf v0.21.0` in `cmd/vmafx-node/bpf/`. The feature is opt-in via
`VMAFX_EBPF_BYPASS=1`. BPF objects are generated artifacts regenerated with
`go generate ./cmd/vmafx-node/bpf/`. A stub auto-selects on non-Linux hosts or
when the BPF toolchain is absent.

## Alternatives considered

| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Full FUSE intercept (custom FUSE driver) | Complete control over all I/O paths | Requires kernel module; far higher complexity; breaks rclone compatibility | Too invasive; deployment requires root |
| Polling the rclone cache directory | Simple to implement; no kernel deps | High polling overhead; latency floor ~100 ms even at aggressive poll intervals; wastes CPU | Latency too high; CPU cost unacceptable at 150 k-clip scale |
| inotify watch on rclone cache dir | No BPF required; works on older kernels | inotify is per-inode; races between open and add; does not track FD → file mapping reliably | Race condition; inotify events arrive after the open, not before |
| Disable rclone FUSE; use direct S3 SDK | Eliminates FUSE overhead entirely | Requires S3 credentials in every worker pod; breaks the mount-abstraction that lets us swap object stores | Credential sprawl; mount abstraction is a design invariant (ADR-0709) |

## Consequences

- **Positive**: 37× warm-cache clip-open latency improvement (370 ms → 10 ms p50);
unlocks full 150 k-clip throughput on a single node without pre-staging;
opt-in design means existing deploys are unaffected.
- **Negative**: Requires Linux 5.15+, `CAP_BPF`, and `clang + bpf2go` for regeneration;
adds `cilium/ebpf v0.21.0` as a runtime dependency; widens the Linux-only surface.
- **Neutral / follow-ups**: BPF objects must be regenerated when the tracepoint
struct layout changes (`go generate ./cmd/vmafx-node/bpf/`); this is an
`AGENTS.md` invariant in `cmd/vmafx-node/bpf/`.

## Supply-chain impact

- **New dependencies**: `cilium/ebpf v0.21.0` (runtime, Apache-2.0,
`https://github.com/cilium/ebpf`).
- **Build-time fetches**: `go generate` invokes `bpf2go` (installed via Go toolchain);
pinned by `go.sum`.
- **Sigstore-signable**: Go module hash in `go.sum` provides integrity anchor.
- **CVE surface delta**: adds a BPF loader; BPF programs are kernel-verified and
run in read-only probe mode — no packet manipulation, no memory writes outside
the map.

## References

- Research-0733: 37× latency improvement measurement data.
- Open DRAFT PR: #137 (`feat(node): eBPF FUSE bypass for rclone`).
- ADR-0709: Phase 4b distributed platform (rclone mount-abstraction invariant).
- ADR-0719: vmafx-node rclone integration (the surface this optimizes).
64 changes: 64 additions & 0 deletions docs/adr/0783-k8s-e2e-integration-test-harness.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
<!-- markdownlint-disable MD013 MD041 MD060 -->
# ADR-0783: Kubernetes end-to-end integration test harness — kind + kuttl

- **Status**: Proposed
- **Date**: 2026-06-03
- **Deciders**: Lusoris
- **Tags**: `ci`, `testing`, `k8s`, `github`

## Context

The Phase 4b platform (ADR-0709) ships a Kubernetes Operator, Node worker, and sidecar
trainer that interact through CRDs and a gRPC control plane. Unit tests exercise
individual components in isolation, but there is no test that validates the full
controller → node → trainer loop on real Kubernetes. This gap means regressions in CRD
reconciliation, job dispatch, or cross-component communication go undetected until a
live deployment.

A lightweight end-to-end harness using kind (Kubernetes in Docker) eliminates the need
for a permanent cloud cluster while still exercising real Kubernetes APIs. kuttl
(KUbernetes Test TooL) provides a declarative YAML-based assertion layer that is
easier to maintain than raw Go integration tests.

The five test cases cover: CRD installation, VmafxJob pod lifecycle, VmafxNode
heartbeat, rclone-sourced scoring, and the sidecar trainer checkpoint flow.

A nightly CI workflow (`.github/workflows/e2e-k8s.yml`) runs the harness at 03:47 UTC
and is also opt-in on PRs via a `run-e2e-k8s` label. An 8-frame 64×64 YUV420p fixture
pair in `test/e2e/fixtures/` allows deterministic scoring without network access.

## Decision

We will ship `test/e2e/kind-cluster.sh` (idempotent kind cluster bootstrap with
real-NVIDIA or fake-GPU path via squat/k8s-fakedeviceplugin), `test/e2e/kuttl-tests/`
(five ordered kuttl test cases), and `.github/workflows/e2e-k8s.yml`. Documentation
lands in `docs/k8s/integration-tests.md`. This test surface is non-blocking on PRs
(opt-in label gate) until all five test cases pass reliably in CI.

## Alternatives considered

| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| kuttl (chosen) | Declarative YAML assertions; maintained by kube-burner community; no custom Go code | Requires kind + kubectl already present; sequential-only test ordering | Best balance of simplicity and real-k8s coverage |
| chainsaw (Kyverno's e2e tool) | Rich assertion DSL; supports parallel steps | Newer, smaller ecosystem; adds Kyverno dependency for a non-Kyverno project | Ecosystem risk; overkill for five sequential test cases |
| envtest (controller-runtime) | Pure Go; runs in-process; fast | Does not exercise Kubernetes networking, DNS, or admission controllers | In-process simulation misses the integration surface we need to test |
| Permanent cloud cluster (EKS / GKE) | Closest to production; tests real GPU scheduling | Cost; secret management; slow teardown; cluster drift | Cost prohibitive for nightly runs; kind achieves the same CRD/reconciliation coverage |

## Consequences

- **Positive**: Full controller → node → trainer loop is now automatically tested;
regressions in CRD reconciliation are caught before merge; local developers
can reproduce exactly with `bash test/e2e/kind-cluster.sh`.
- **Negative**: Nightly job adds ~15 min to CI wall time; fake-GPU path does not
exercise CUDA kernels (GPU scoring in test case 04 uses CPU fallback).
- **Neutral / follow-ups**: Test case 05 (sidecar-trainer) requires the operator
`currentSamples` increment logic to be implemented; the `required-aggregator.yml`
should mark `E2E — Kubernetes Integration` as non-blocking until all five cases pass.

## References

- Open DRAFT PR: #152 (`feat(ci): k8s e2e integration test harness — kind + kuttl`).
- ADR-0709: Phase 4b distributed platform.
- ADR-0711: vmafx-controller implementation.
- ADR-0713: vmafx-node implementation.
- ADR-0781: sidecar SGD-EMA online trainer.
77 changes: 77 additions & 0 deletions docs/adr/0815-operator-node-distroless-dockerfiles.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
<!-- markdownlint-disable MD013 MD041 MD060 -->
# ADR-0815: Distroless multi-arch Dockerfiles for vmafx-operator and vmafx-node + release CI

- **Status**: Proposed
- **Date**: 2026-06-03
- **Deciders**: Lusoris
- **Tags**: `ci`, `build`, `security`, `supply-chain`, `github`, `docker`

## Context

The vmafx-operator (Go, CGO_ENABLED=0) and vmafx-node CPU variant need production
container images for Kubernetes deployment. Without a published image, the k8s Operator
(ADR-0711) and Phase 4b platform (ADR-0709) cannot be deployed from the Helm chart
(ADR-0699). The node's CUDA image exists (`docker/Dockerfile.node`, ADR-0717) but is
not wired to any release CI workflow. The operator has no Dockerfile at all.

Using `gcr.io/distroless/static-debian12` as the base image minimises the CVE surface
(no shell, no package manager, no libc beyond musl stubs). Running as the distroless
`nonroot` user (uid 65532, aligned with ADR-0878) satisfies the Kubernetes
`pod-security.kubernetes.io/enforce=restricted` admission profile without extra pod
spec overrides.

Multi-arch (amd64 + arm64) via BuildKit native cross-compilation (CGO_ENABLED=0 makes
this straightforward for pure-Go binaries) avoids QEMU emulation and keeps build times
under two minutes.

## Decision

We will add `docker/Dockerfile.operator` (pure-Go operator binary into distroless,
runs as uid 65532, multi-arch amd64+arm64) and wire it alongside the existing
`docker/Dockerfile.node` (ADR-0717) into a new release CI workflow
(`.github/workflows/docker-publish-operator-node.yml`) that fires on `v*` tags and
`workflow_dispatch`. The workflow builds and pushes `ghcr.io/vmafx/vmafx-operator` and
`ghcr.io/vmafx/vmafx-node` (CPU variant), with cosign keyless signing and syft
CycloneDX SBOM attestation, mirroring the `docker-publish-production.yml` pattern
(ADR-0698).

## Alternatives considered

| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| `gcr.io/distroless/static-debian12` (chosen) | Minimal CVE surface; no shell; aligns with ADR-0878 uid 65532 | No debugging tools in container; must use ephemeral debug containers for prod troubleshooting | Best security posture; debug containers are the k8s-native solution |
| `gcr.io/distroless/cc-debian12` (C runtime variant) | Supports CGO-linked binaries | Larger image; CGO_ENABLED=0 makes this unnecessary for pure-Go operator | Unnecessary for a pure-Go binary |
| `alpine:3.19` | Has shell + apk for ad-hoc debugging; lighter than Debian | Not distroless; larger CVE surface; musl libc can cause subtle differences | CVE surface wider; distroless is the standard for production Go |
| Inline Dockerfile in workflow | Fewer files | Harder to test locally; no `docker build -f` reproducibility | Reproducibility is a first-class requirement |
| Separate Dockerfile per arch | Full control per arch | Doubles maintenance; BuildKit native cross-compile is equivalent | Unnecessary complexity |

## Consequences

- **Positive**: vmafx-operator and vmafx-node (CPU) are now release-published with
cosign attestation; Helm chart installs work without manual image builds;
multi-arch amd64+arm64 covers the arm node market without QEMU.
- **Negative**: Two new images add ~2 min to release CI; distroless debugging
requires ephemeral debug containers (`kubectl debug`).
- **Neutral / follow-ups**: Update `required-aggregator.yml` to include the new
workflow's smoke-test gate; update `docs/backends/operator.md` runbook with the
image reference and `docker pull` command.

## Supply-chain impact

- **New dependencies**: none at runtime (pure-Go, static binary).
- **Build-time fetches**: `gcr.io/distroless/static-debian12` base image pulled at
build time; pinned by digest in the Dockerfile.
- **Sigstore-signable**: cosign keyless signing via Sigstore OIDC is applied to both
images; syft CycloneDX SBOM is attached as an OCI attestation.
- **CVE surface delta**: narrows — distroless has no shell, no package manager,
no libc beyond the Go runtime's own calls.

## References

- Open DRAFT PR: #184 (`feat(docker): Dockerfile.operator + node publish CI — distroless multi-arch`).
- ADR-0698: `docker-publish-production.yml` pattern this mirrors.
- ADR-0699: Helm chart that consumes the published images.
- ADR-0709: Phase 4b distributed platform.
- ADR-0711: vmafx-operator implementation.
- ADR-0717: vmafx-node Dockerfile (existing, wired to release CI for the first time).
- ADR-0878: Trivy DS-0002 distroless `nonroot` uid 65532 baseline.
69 changes: 69 additions & 0 deletions docs/adr/0930-helm-networkpolicy-pss.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
<!-- markdownlint-disable MD013 MD041 MD060 -->
# ADR-0930: Helm chart NetworkPolicy default-deny + Pod Security Standards "restricted" baseline

- **Status**: Proposed
- **Date**: 2026-06-03
- **Deciders**: Lusoris
- **Tags**: `security`, `k8s`, `ci`, `build`, `github`

## Context

The vmafx Helm chart (`deploy/helm/vmafx/`, ADR-0699) shipped with minimal pod security
configuration. Default Kubernetes pod security contexts allow containers to run as root,
with writable root filesystems, and without seccomp profiles. This does not satisfy the
Kubernetes `pod-security.kubernetes.io/enforce=restricted` admission profile required
by most hardened cluster operators.

Additionally, the chart emitted no NetworkPolicy resources, leaving all cross-pod
traffic unrestricted. The Phase 4b platform (ADR-0709) introduces a controller → node
gRPC channel, a node → object-store HTTPS path, and operator → apiserver traffic, each
of which should be narrowly allowed with everything else denied by default.

The distroless `nonroot` uid used in the operator and node images (ADR-0878) is
65532, but the chart previously defaulted `podSecurityContext.runAsUser` to 65534
(the generic `nobody` uid). This drift caused file-ownership inconsistencies when
image-baked paths were accessed at runtime.

## Decision

We will update the Helm chart to: (1) align `podSecurityContext.runAsUser` to 65532
(distroless `nonroot`, per ADR-0878) across operator and node deployments; (2) set
`runAsNonRoot: true`, `allowPrivilegeEscalation: false`, `readOnlyRootFilesystem: true`,
and `seccompProfile.type: RuntimeDefault` as chart defaults, satisfying PSA
`restricted` out of the box; (3) ship `templates/networkpolicy.yaml` with a
default-deny baseline and narrow allow-rules (controller → node gRPC, node → object
store HTTPS, operator → apiserver, DNS, in-namespace HTTP ingress), gated by
`networkPolicy.enabled=false` so existing installs are unaffected by default.

## Alternatives considered

| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| NetworkPolicy enabled by default | Immediately hardens all installs | Breaks installs on clusters without a NetworkPolicy-aware CNI (Flannel, older EKS); migration friction | Opt-in via `networkPolicy.enabled=true` is safer for a Helm library chart |
| BYO NetworkPolicy (not in chart) | Maximum flexibility; no chart coupling | Every operator must maintain their own policies; no reference policy | Poor DX; reference policies reduce misconfiguration risk |
| UID 65534 (generic `nobody`) | Standard on non-distroless images | Mismatches distroless baked paths; causes file-ownership bugs at runtime | ADR-0878 already established 65532 as the canonical uid |
| Inline security blocks per workload template | Explicit per-template; no values indirection | Duplication; values override becomes impossible without template changes | Values-driven blocks allow cluster admins to override per-deploy |
| Pod Security Policy (deprecated) | Pre-1.25 clusters | Removed in Kubernetes 1.25; no-op on modern clusters | Deprecated; PSA labels are the successor |
| `seccompProfile.type: Localhost` (custom profile) | Maximum control over syscall allowlist | Requires profile distribution to each node; operational overhead | RuntimeDefault covers the common case; Localhost can be a follow-up |

## Consequences

- **Positive**: Chart default render satisfies PSA `restricted` without manual
overrides; NetworkPolicy opt-in provides a reference policy for hardened
deployments; uid/GID alignment with distroless images eliminates file-ownership
drift.
- **Negative**: `podSecurityContext.runAsUser` flip from 65534 → 65532 is a
potentially breaking change for installs that hard-coded the old value (migration
note in PR description and `NOTES.txt`).
- **Neutral / follow-ups**: When a dedicated `vmafx-controller` Service ships,
tighten the `controller-to-node` NetworkPolicy selector from "any pod in namespace"
to `component=controller`; track Kubernetes 1.31+ `appArmorProfile` for per-backend
AppArmor; tighten `nodeEgressObjectStore.cidrs` in production overlays.

## References

- Open DRAFT PR: #439 (`feat(helm): NetworkPolicy default-deny + Pod Security Standards "restricted" baseline`).
- ADR-0699: vmafx Helm chart foundation.
- ADR-0709: Phase 4b distributed platform (defines the traffic matrix).
- ADR-0719: vmafx-node rclone integration (node → object store HTTPS path).
- ADR-0878: Trivy DS-0002 distroless `nonroot` uid 65532 baseline.
4 changes: 4 additions & 0 deletions docs/adr/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -818,4 +818,8 @@ ADRs may exist there for local session continuity, but the tracked
| [ADR-0865](0865-ansnr-sunset-pre-vmaf-metric-drop.md) | Sunset ANSNR (pre-VMAF metric) — drop `ansnr` / `float_ansnr` feature extractors | Accepted | 2026-05-28 | metric, feature-extractor, breaking-change, cleanup, fork-local |
| [ADR-0984](0984-port-upstream-netflix-may-jun-2026.md) | Port Netflix upstream May–Jun 2026: direct-read default (e4b93c6ed), integer_motion_v2 rename (a4a1492d3), 2160p@1.5H CSF (c2155d6cd), ADM SIMD dispatch+fix (9a078011c), Speed_chroma covariance vectorization (30f472b14) | Accepted | 2026-06-01 | upstream-port, performance, cli, simd, csf, adm, speed, fork-local |
| [ADR-0774](0774-mcp-server-audit.md) | MCP server audit — fix stale `libvmaf/` path in `list_extractors`, wire `subsample` through to `--subsample` CLI flag, narrow `bitdepth` schema to `[8,10,12]`, remove dead `_run_benchmark` shadow, log swallowed exceptions in `_send_progress` and `_load_vlm`. | Accepted | 2026-05-29 | mcp, server, audit, bugfix, fork-local |
| [ADR-0779](0779-ebpf-fuse-bypass-rclone.md) | eBPF FUSE bypass for rclone zero-copy path in vmafx-node: probe-only tracepoint program + Go cilium/ebpf loader; 37× warm-cache latency improvement; opt-in via VMAFX_EBPF_BYPASS=1 | Proposed | 2026-06-03 | ci, go, ebpf, rclone, performance, security, supply-chain |
| [ADR-0783](0783-k8s-e2e-integration-test-harness.md) | Kubernetes end-to-end integration test harness — kind + kuttl: five test cases covering operator/node/trainer stack; nightly CI + opt-in PR label gate | Proposed | 2026-06-03 | ci, testing, k8s, github |
| [ADR-0815](0815-operator-node-distroless-dockerfiles.md) | Distroless multi-arch Dockerfiles for vmafx-operator and vmafx-node + release CI: gcr.io/distroless/static-debian12, uid 65532, amd64+arm64, cosign + syft SBOM | Proposed | 2026-06-03 | ci, build, security, supply-chain, github, docker |
| [ADR-0930](0930-helm-networkpolicy-pss.md) | Helm chart NetworkPolicy default-deny + Pod Security Standards "restricted" baseline: PSA-restricted pod contexts, uid 65532, opt-in NetworkPolicy bundle with narrow allow-rules | Proposed | 2026-06-03 | security, k8s, ci, build, github |
Comment on lines +821 to +824
| [ADR-0985](0985-sycl-parity-divergence-2026-06-03.md) | SYCL parity divergence investigation — float_ssim + ssimulacra2 on Arc A380: documents formula difference (CPU L×C×S vs GPU combined Eq.13), fp32 accumulation gap on fp64-less hardware, closes stale float_ansnr row; adds float_ssim_sycl parity test at places=3 | Proposed | 2026-06-03 | sycl, parity, ci, gpu, precision, arc, fork-local |
Loading