Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
94 changes: 94 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -198,6 +198,100 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
in S3 (CRD `spec.signingKeys[]` REQUIRED, CEL-gated non-empty);
controller-side auto-mint deferred to S7.

### S4 `phase2/inferencepolicy-reconciler` — full InferencePolicy reconciler + compile + helm CRD

This slice ships the controller side of the K8s primitive only. Per §3
non-compete, `InferencePolicy` is **not** a model-router — model
selection sits in Foundry; this is a sandbox-side budget / guardrail /
safety policy CR. Per user direction 2026-04-27, the runtime enforcement
substrate stays on Phase 1 (`inference-router::budget` env-fed token
tracker, Foundry-side Content Safety with flags reported to AGT
`BehaviorMonitor` via `safety::report_content_flags_to_agt`); the
informer that loads the compiled profile into `PolicyEnvelope` and the
optional upstream `BudgetTracker` port to AGT-Rust are both deferred to
S7. The compiled JSON ConfigMap is the hand-off contract.

#### Added

- **`InferencePolicy` CRD** — `controller/src/inference_policy.rs`,
group `azureclaw.azure.com`, version `v1alpha1`, namespaced,
shortname `ip`. Sub-types: `InferenceAppliesTo` (sandboxName,
sandboxMatchLabels, action), `TokenBudget` (perRequestTokens,
dailyTokens, monthlyTokens), `ContentSafetyFloor` (hate, selfHarm,
sexual, violence, requirePromptShields), `ModelPreference` (primary +
ordered fallback `Vec<ModelRef>`), `ModelRef` (provider, deployment).
Reuses `mcp_server::LocalObjectRef` for status pointer
(4th semantic client of that struct).
- **Pure compile module** — `controller/src/inference_policy_compile.rs`,
`compile_to_profile(&InferencePolicySpec) -> serde_json::Value` +
`version_hash(&Value) -> String` (sha256, first 16 bytes hex).
Deterministic; key-canonical via `serde_json::Value::Object`. Output
shape slots into `inference-router::policy_envelope::PolicyEntry::payload` —
no parallel hot-reload core.
- **Reconciler** — `controller/src/inference_policy_reconciler.rs`,
modeled directly on `a2a_agent_reconciler.rs` (S3). Field manager
`azureclaw-controller/inferencepolicy` (distinct per §10.4 #1);
finalizer `azureclaw.azure.com/inferencepolicy-cleanup`. Emits
`Ready`/`Progressing`/`Degraded` Conditions reusing
`status::conditions` helpers; preserves `lastTransitionTime` when
status doesn't flip. Closed-set `error_class`
(`kube_api`/`serde`) per §15.3.
- **Profile ConfigMap** — name `inferencepolicy-{name}-profile`, key
`profile.json`, annotated with version hash, labelled
`azureclaw.azure.com/artifact=inference-policy-profile` for the S7
router-side informer label selector.
- **CEL admission rules** (6) — `inference_policy_validations`:
`monthlyTokens >= dailyTokens`,
`monthlyTokens >= perRequestTokens`,
`contentSafety.{hate,selfHarm,sexual,violence}` ∈
`{Safe,Low,Medium,High}`,
`modelPreference.primary` non-empty provider+deployment,
`modelPreference.fallback[*]` non-empty provider+deployment,
`appliesTo.action` ∈ `{chat,responses,image,embeddings,*}`.
- **Helm CRD** — `deploy/helm/azureclaw/templates/crd-inferencepolicy.yaml`,
emitted by `helm_drift::tests::dump_inferencepolicy_crd_yaml` (env-gated)
and drift-checked by `helm_inferencepolicy_crd_matches_rust_schema`.
- **Audit doc** — `docs/security-audits/2026-04-27-phase2-inferencepolicy-reconciler.md`,
documenting the AGT boundary verification against
`agent-governance-toolkit` 3.3.0 on disk: AGT-Python has
`BudgetTracker`, AGT-Rust does not (yet); `cedar-policy` + `regorus`
available in AGT-Rust for future Content Safety floor encoding.
STRIDE coverage, OWASP A2A coverage, explicit out-of-scope list,
two sign-offs.

#### Tests

- `inference_policy_compile::tests` — 6 unit tests (empty/full
round-trip, determinism, version-hash change/stability, hex shape).
- `inference_policy_reconciler::tests` — 7 unit tests (rfc3339 shape,
error-class closed set, conditions on success/failure,
transition-time preservation, finalizer dns-subdomain, field-manager
distinctness from S1/S2/S3).
- `crd_validations::tests` — 5 new `inference_policy_*` tests
(non-empty rules, every-rule-has-message, after-injection count,
rule-mention invariants, serde round-trip).
- `helm_drift::tests` — 2 new (`dump_inferencepolicy_crd_yaml` env-gated
+ `helm_inferencepolicy_crd_matches_rust_schema`).
- Full controller suite: **218 passing** (was 193 after S3). Workspace
`cargo test`, `cargo fmt --all`, `cargo clippy --all-targets -D
warnings` — all green.

#### §14.6 impact

Strengthens column 7 (Foundry / M365 integration) of the competitive
matrix — the *primitive* lands in this slice; column-7 credibility
moves further when S7 wires the runtime consumers (token-budget swap,
floor compare-and-block, model-preference selection).

#### Notes

- AGT crate pin remains `agentmesh = "3.3.0"` from crates.io,
unmodified. `vendor/` directory untouched.
- Single new struct: none. `LocalObjectRef` semantically extended (4
clients now: signing/jwks, profile, agent-card, guardrail-profile).
- The runtime gate `inference-router::routes::inference_policy::check`
(Phase 1) is **not modified in this slice**.

## [Unreleased] — PR #44 `dev → main` uplift

This entry covers **186 commits** on `dev` since `main`, structured as Phase 0
Expand Down
133 changes: 133 additions & 0 deletions controller/src/crd_validations.rs
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ use k8s_openapi::apiextensions_apiserver::pkg::apis::apiextensions::v1::{
use kube::CustomResourceExt;

use crate::a2a_agent::A2AAgent;
use crate::inference_policy::InferencePolicy;
use crate::mcp_server::McpServer;
use crate::tool_policy::ToolPolicy;

Expand Down Expand Up @@ -221,6 +222,83 @@ pub fn a2a_agent_crd() -> CustomResourceDefinition {
.expect("kube-rs derive must produce a spec property on A2AAgent")
}

/// `InferencePolicy.spec` CEL rules. Phase 2 §8 entry 4 (S4).
///
/// Returns the validations injected on the `spec` schema node:
///
/// - `tokenBudget.monthlyTokens >= tokenBudget.dailyTokens` when both
/// are set (admission CEL — the runtime path also checks but
/// prevents the half-baked CR from ever landing).
/// - `tokenBudget.monthlyTokens >= tokenBudget.perRequestTokens` when
/// both are set (single request can't blow a monthly budget that
/// wouldn't accept it).
/// - `contentSafety.{hate,selfHarm,sexual,violence}` ∈ {`Safe`, `Low`,
/// `Medium`, `High`} when present — matches Microsoft Content Safety
/// `Microsoft.DefaultV2` severity levels exactly.
/// - `modelPreference.primary` requires non-empty `provider` and
/// `deployment` (any fallback entry too).
/// - `appliesTo.action` ∈ {`chat`, `responses`, `image`, `embeddings`,
/// `*`} — closed set matching `inference-router/src/routes/inference.rs`
/// call-site enumeration.
#[must_use]
pub fn inference_policy_validations() -> Vec<ValidationRule> {
let severities = "['Safe','Low','Medium','High']";
vec![
ValidationRule {
rule: "!has(self.tokenBudget) || !has(self.tokenBudget.monthlyTokens) || !has(self.tokenBudget.dailyTokens) || self.tokenBudget.monthlyTokens >= self.tokenBudget.dailyTokens".into(),
message: Some("spec.tokenBudget.monthlyTokens must be >= spec.tokenBudget.dailyTokens".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
ValidationRule {
rule: "!has(self.tokenBudget) || !has(self.tokenBudget.monthlyTokens) || !has(self.tokenBudget.perRequestTokens) || self.tokenBudget.monthlyTokens >= self.tokenBudget.perRequestTokens".into(),
message: Some("spec.tokenBudget.monthlyTokens must be >= spec.tokenBudget.perRequestTokens".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
ValidationRule {
rule: format!(
"!has(self.contentSafety) || (\
(!has(self.contentSafety.hate) || self.contentSafety.hate in {sev}) && \
(!has(self.contentSafety.selfHarm) || self.contentSafety.selfHarm in {sev}) && \
(!has(self.contentSafety.sexual) || self.contentSafety.sexual in {sev}) && \
(!has(self.contentSafety.violence) || self.contentSafety.violence in {sev}))",
sev = severities
),
message: Some("spec.contentSafety severities must be one of: Safe, Low, Medium, High".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
ValidationRule {
rule: "!has(self.modelPreference) || (size(self.modelPreference.primary.provider) > 0 && size(self.modelPreference.primary.deployment) > 0)".into(),
message: Some("spec.modelPreference.primary requires non-empty provider and deployment".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
ValidationRule {
rule: "!has(self.modelPreference) || self.modelPreference.fallback.all(f, size(f.provider) > 0 && size(f.deployment) > 0)".into(),
message: Some("spec.modelPreference.fallback[*] requires non-empty provider and deployment".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
ValidationRule {
rule: "!has(self.appliesTo.action) || self.appliesTo.action in ['chat','responses','image','embeddings','*']".into(),
message: Some("spec.appliesTo.action must be one of: chat, responses, image, embeddings, *".into()),
reason: Some("FieldValueInvalid".into()),
..ValidationRule::default()
},
]
}

/// `InferencePolicy` CRD with [`inference_policy_validations`] injected.
///
/// Panics only if kube-rs ever produces a CRD whose `spec` is missing.
#[must_use]
pub fn inference_policy_crd() -> CustomResourceDefinition {
inject_spec_validations(InferencePolicy::crd(), inference_policy_validations())
.expect("kube-rs derive must produce a spec property on InferencePolicy")
}

#[cfg(test)]
mod tests {
use super::*;
Expand Down Expand Up @@ -378,6 +456,61 @@ mod tests {
assert!(y.contains("signingKeys"));
}

#[test]
fn inference_policy_validations_are_non_empty() {
assert!(!inference_policy_validations().is_empty());
}

#[test]
fn every_inference_policy_rule_has_message_and_rule() {
for rule in inference_policy_validations() {
assert!(!rule.rule.is_empty(), "rule body must not be empty");
let msg = rule.message.as_deref().unwrap_or("");
assert!(!msg.is_empty(), "rule '{}' missing message", rule.rule);
}
}

#[test]
fn inference_policy_crd_has_spec_validations_after_injection() {
let crd = inference_policy_crd();
let v = spec_validations(&crd);
assert_eq!(v.len(), inference_policy_validations().len());
}

#[test]
fn inference_policy_rules_mention_token_budget_and_severity_invariants() {
let rules: Vec<String> = inference_policy_validations()
.into_iter()
.map(|r| r.rule)
.collect();
assert!(
rules
.iter()
.any(|r| r.contains("monthlyTokens") && r.contains("dailyTokens")),
"must enforce monthlyTokens >= dailyTokens; got rules: {rules:?}"
);
assert!(
rules
.iter()
.any(|r| r.contains("Safe") && r.contains("High")),
"must enforce content-safety severity closed set; got rules: {rules:?}"
);
assert!(
rules
.iter()
.any(|r| r.contains("appliesTo.action") && r.contains("chat")),
"must enforce appliesTo.action closed set; got rules: {rules:?}"
);
}

#[test]
fn inference_policy_crd_is_serde_round_trippable() {
let crd = inference_policy_crd();
let y = serde_yaml::to_string(&crd).expect("serializes");
assert!(y.contains("x-kubernetes-validations"));
assert!(y.contains("monthlyTokens"));
}

#[test]
fn injection_returns_none_when_spec_property_missing() {
// Build a CRD with an empty schema tree to confirm the helper
Expand Down
34 changes: 33 additions & 1 deletion controller/src/helm_drift.rs
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,9 @@
#![allow(dead_code)]

#[cfg(test)]
use crate::crd_validations::{a2a_agent_crd, mcp_server_crd, tool_policy_crd};
use crate::crd_validations::{
a2a_agent_crd, inference_policy_crd, mcp_server_crd, tool_policy_crd,
};

const MCP_HELM_CRD_PATH: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
Expand All @@ -45,6 +47,11 @@ const A2AAGENT_HELM_CRD_PATH: &str = concat!(
"/../deploy/helm/azureclaw/templates/crd-a2aagent.yaml"
);

const INFERENCEPOLICY_HELM_CRD_PATH: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../deploy/helm/azureclaw/templates/crd-inferencepolicy.yaml"
);

/// Strip non-schema fields that legitimately differ between the Rust
/// `CustomResource::crd()` output and the helm template (helm labels,
/// status block, metadata.creationTimestamp, etc.). The comparison key
Expand Down Expand Up @@ -158,4 +165,29 @@ mod tests {
serde_json::to_value(a2a_agent_crd()).expect("rust crd serializes to JSON");
assert_helm_matches_rust(A2AAGENT_HELM_CRD_PATH, rust_crd_value, "a2aagent");
}

/// One-shot dumper for the inferencepolicy CRD. Run via:
///
/// DUMP_INFERENCEPOLICY_CRD_YAML=1 cargo test --bin azureclaw-controller \
/// helm_drift::tests::dump_inferencepolicy_crd_yaml -- --nocapture
#[test]
fn dump_inferencepolicy_crd_yaml() {
if std::env::var("DUMP_INFERENCEPOLICY_CRD_YAML").is_err() {
return;
}
let crd = inference_policy_crd();
let yaml = serde_yaml::to_string(&crd).expect("serialize crd to YAML");
println!("---\n{yaml}");
}

#[test]
fn helm_inferencepolicy_crd_matches_rust_schema() {
let rust_crd_value =
serde_json::to_value(inference_policy_crd()).expect("rust crd serializes to JSON");
assert_helm_matches_rust(
INFERENCEPOLICY_HELM_CRD_PATH,
rust_crd_value,
"inferencepolicy",
);
}
}
Loading
Loading