Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions changelog.d/added/helm-networkpolicy-pss.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
- **Helm chart NetworkPolicy + Pod Security Standards baseline (ADR-0930)** —
`deploy/helm/vmafx/` now renders pods that pass the Kubernetes
`pod-security.kubernetes.io/enforce=restricted` admission profile out
of the box: `runAsNonRoot`, distroless `nonroot` UID/GID `65532`
(aligned with ADR-0878), `readOnlyRootFilesystem`, dropped capabilities,
`seccompProfile.type=RuntimeDefault`, and `allowPrivilegeEscalation=false`
at both pod and container scope. A new opt-in
`templates/networkpolicy.yaml` (gated by `--set networkPolicy.enabled=true`)
emits a default-deny ingress + egress baseline plus narrow allow-rules
for in-namespace HTTP ingress, controller -> node gRPC, node -> object
store HTTPS (CIDR + `except` matrix), operator -> apiserver, and DNS
egress to CoreDNS. `operator-deployment.yaml` and
`tests/test-connection.yaml` now inherit `podSecurityContext` /
`securityContext` from `.Values` so chart-wide hardening changes can
no longer drift across templates. `NOTES.txt` and
`docs/development/k8s-deployment.md` document the namespace-labelling
command and the full NetworkPolicy matrix.

**Migration**: installs that hard-coded
`--set podSecurityContext.runAsUser=65534` should drop the override or
flip it to `65532` to keep file ownership consistent with the
distroless `nonroot` baked into every production image.
37 changes: 37 additions & 0 deletions deploy/helm/vmafx/templates/NOTES.txt
Original file line number Diff line number Diff line change
Expand Up @@ -62,4 +62,41 @@ GPU device-plugin is allocated. See docs/development/gpu-scheduling.md.
To verify device-plugin allocations on a node:
kubectl describe node <node> | grep -E "Capacity|Allocatable" -A 10

----------------------------------------------------------------------
Pod Security Standards — restricted profile (ADR-0930)
----------------------------------------------------------------------
This chart's pods satisfy the Kubernetes
"pod-security.kubernetes.io/enforce=restricted" admission profile:
runAsNonRoot, distroless UID 65532, readOnlyRootFilesystem, dropped
capabilities, seccompProfile=RuntimeDefault, allowPrivilegeEscalation=false.

To enforce the profile at the namespace level, label the install namespace:

kubectl label --overwrite namespace {{ .Release.Namespace }} \
pod-security.kubernetes.io/enforce=restricted \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/warn=restricted

----------------------------------------------------------------------
NetworkPolicy — default-deny posture (ADR-0930)
----------------------------------------------------------------------
{{- if .Values.networkPolicy.enabled }}
NetworkPolicies are ENABLED for this release. Every VMAFX pod runs under a
default-deny posture with explicit allow-rules for the documented
controller -> node, node -> object-store, operator -> apiserver, and DNS
egress flows. Verify the rules with:

kubectl get networkpolicy -n {{ .Release.Namespace }} -l app.kubernetes.io/instance={{ .Release.Name }}

Requires a NetworkPolicy-aware CNI (Cilium, Calico, kube-router, etc.).
{{- else }}
NetworkPolicies are DISABLED for this release. Opt in with:

helm upgrade {{ .Release.Name }} . --reuse-values --set networkPolicy.enabled=true

When enabled, the chart emits a default-deny ingress + egress baseline plus
explicit allow-rules for controller -> node, node -> object-store, operator
-> apiserver, and DNS. Requires a NetworkPolicy-aware CNI.
{{- end }}

Full operator guide: docs/development/k8s-deployment.md
250 changes: 250 additions & 0 deletions deploy/helm/vmafx/templates/networkpolicy.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,250 @@
{{- /*
SPDX-License-Identifier: BSD-3-Clause-Plus-Patent
Copyright 2026 Lusoris

deploy/helm/vmafx/templates/networkpolicy.yaml
Default-deny NetworkPolicies for the VMAFX release, plus the explicit
allow-rules that the documented controller/node/operator/DNS flows need.

Disabled by default (`networkPolicy.enabled=false`) — opt in with
`--set networkPolicy.enabled=true` once you have a NetworkPolicy-aware
CNI (Cilium, Calico, kube-router, etc.) installed in the cluster.

ADR-0930: Helm NetworkPolicy + Pod Security Standards baseline.
See docs/development/k8s-deployment.md#networkpolicy for the egress/ingress
matrix and the override knobs in values.yaml.
*/ -}}
{{- if .Values.networkPolicy.enabled }}
{{- $fullName := include "vmafx.fullname" . }}
{{- $selector := include "vmafx.selectorLabels" . }}
{{- $ns := .Release.Namespace }}
---
# ----------------------------------------------------------------------------
# Default-deny: drop every ingress and egress flow for VMAFX pods unless an
# explicit allow-rule below opens it. This is the safety net — even if a new
# workload is added to the chart without an accompanying allow-rule, it stays
# isolated by default.
# ----------------------------------------------------------------------------
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-default-deny
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
policyTypes:
- Ingress
- Egress
{{- if .Values.operator.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-operator-default-deny
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
app.kubernetes.io/component: operator
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
app.kubernetes.io/component: operator
policyTypes:
- Ingress
- Egress
{{- end }}
{{- if .Values.node.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-node-default-deny
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
app.kubernetes.io/component: node
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
app.kubernetes.io/component: node
policyTypes:
- Ingress
- Egress
{{- end }}

# ----------------------------------------------------------------------------
# Ingress for the scoring server (Deployment / StatefulSet workloads). All
# in-namespace clients may reach the HTTP scoring port; cross-namespace
# clients are out of scope (set up your own NetworkPolicy in their namespace).
# ----------------------------------------------------------------------------
{{- if or (eq .Values.workload "Deployment") (eq .Values.workload "StatefulSet") }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-allow-http-ingress
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
policyTypes:
- Ingress
ingress:
- from:
- podSelector: {}
ports:
- protocol: TCP
port: {{ .Values.service.targetPort }}
{{- end }}

# ----------------------------------------------------------------------------
# Allow rule: controller -> node (gRPC dispatch). Matches vmafx-controller
# pods (component=controller / selector match) reaching vmafx-node pods on
# the configured gRPC port.
# ----------------------------------------------------------------------------
{{- if and .Values.node.enabled .Values.networkPolicy.allow.controllerToNode.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-allow-controller-to-node
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
app.kubernetes.io/component: node
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
app.kubernetes.io/component: node
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
{{- $selector | nindent 14 }}
ports:
- protocol: TCP
port: {{ .Values.networkPolicy.allow.controllerToNode.nodePort }}
{{- end }}

# ----------------------------------------------------------------------------
# Allow rule: node -> object store (rclone egress). Rendered as a single
# egress policy with the CIDR/port matrix from values.yaml. Defaults to
# 0.0.0.0/0 minus RFC1918 — tighten to your bucket VPC CIDR in production.
# ----------------------------------------------------------------------------
{{- if and .Values.node.enabled .Values.networkPolicy.allow.nodeEgressObjectStore.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-allow-node-egress-object-store
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
app.kubernetes.io/component: node
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
app.kubernetes.io/component: node
policyTypes:
- Egress
egress:
- to:
{{- range .Values.networkPolicy.allow.nodeEgressObjectStore.cidrs }}
- ipBlock:
cidr: {{ . | quote }}
{{- with $.Values.networkPolicy.allow.nodeEgressObjectStore.except }}
except:
{{- range . }}
- {{ . | quote }}
{{- end }}
{{- end }}
{{- end }}
ports:
{{- range .Values.networkPolicy.allow.nodeEgressObjectStore.ports }}
- protocol: TCP
port: {{ . }}
{{- end }}
{{- end }}

# ----------------------------------------------------------------------------
# Allow rule: operator -> kube-apiserver (controller-runtime list/watch).
# The egress is namespace-anchored (kube-system / default Service IP) rather
# than CIDR-anchored, because the kube-apiserver Service IP is cluster-local.
# ----------------------------------------------------------------------------
{{- if and .Values.operator.enabled .Values.networkPolicy.allow.operatorToApiserver.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-allow-operator-to-apiserver
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
app.kubernetes.io/component: operator
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
app.kubernetes.io/component: operator
policyTypes:
- Egress
egress:
# Egress to kube-apiserver — most clusters expose it as a ClusterIP at
# `kubernetes.default.svc`; emit ipBlock 0.0.0.0/0 on the api-server
# ports because the Service IP isn't selectable by a NetworkPolicy peer.
# Operators can tighten this to their masters' CIDR when known.
- to:
- ipBlock:
cidr: "0.0.0.0/0"
ports:
{{- range .Values.networkPolicy.allow.operatorToApiserver.ports }}
- protocol: TCP
port: {{ . }}
{{- end }}
{{- end }}

# ----------------------------------------------------------------------------
# Allow rule: DNS egress. Every VMAFX pod must be able to resolve names
# via the cluster DNS service (CoreDNS / kube-dns in `kube-system`).
# ----------------------------------------------------------------------------
{{- if .Values.networkPolicy.allow.dns.enabled }}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: {{ $fullName }}-allow-dns-egress
namespace: {{ $ns }}
labels:
{{- include "vmafx.labels" . | nindent 4 }}
spec:
podSelector:
matchLabels:
{{- $selector | nindent 6 }}
policyTypes:
- Egress
egress:
- to:
- namespaceSelector:
{{- toYaml .Values.networkPolicy.allow.dns.namespaceSelector | nindent 12 }}
podSelector:
{{- toYaml .Values.networkPolicy.allow.dns.podSelector | nindent 12 }}
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
{{- end }}
{{- end }}
6 changes: 1 addition & 5 deletions deploy/helm/vmafx/templates/operator-deployment.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -74,10 +74,6 @@ spec:
resources:
{{- toYaml (.Values.operator.resources | default dict) | nindent 12 }}
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
{{- toYaml .Values.securityContext | nindent 12 }}
terminationGracePeriodSeconds: 10
{{- end }}
10 changes: 2 additions & 8 deletions deploy/helm/vmafx/templates/tests/test-connection.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,7 @@ metadata:
spec:
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 65534
runAsGroup: 65534
{{- toYaml .Values.podSecurityContext | nindent 4 }}
containers:
- name: wget
image: busybox:1.36
Expand All @@ -25,8 +23,4 @@ spec:
- --timeout=10
- "http://{{ include "vmafx.fullname" . }}.{{ .Release.Namespace }}.svc.cluster.local:{{ .Values.service.port }}/healthz"
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
{{- toYaml .Values.securityContext | nindent 8 }}
Loading
Loading