Repository navigation
Conversation
@microsoft-github-policy-service agree |
053dd90 to
f9b9ea4
Compare
Pal Lakatos-Toth (pallakatos)
left a comment
There was a problem hiding this comment.
Thanks for the PR! Reviewed all 5 files — here's the breakdown:
✅ Accept: sandbox-images/openclaw/entrypoint.sh
The temp-dir pre-creation and router log redirect fix look solid. Correct UID handling, mode 700, and the sh -c wrapper for the redirect is the right approach. Happy to merge this part.
❌ Reject: inference-router/src/proxy.rs
This conflicts with changes already on main:
- We upgraded
prometheus0.13 → 0.14 on main, which requireswith_label_valuesto use consistent types. Your diff reverts those back to&strliterals — this will fail to compile against main. - The embedding model routing fix (extracting model from request body) — we already landed an equivalent fix in
routes.rsfor the AKS/Workload-Identity path. Your fix covers the API-key path, which is good in principle, but the larger streaming refactor (removinginject_stream_usageand the token-metrics wrapper) is a separate concern and needs its own discussion. - The streaming doc-comment changes are fine but tangled with the above.
Suggestion: Split this into two PRs — (a) the API-key embedding routing fix rebased on current main, and (b) the streaming simplification as a separate discussion.
❌ Skip: cli/src/plugin.ts (Bing discovery)
We already have type + properties.category discovery on main, and it works on our dev tenant. The metadata.type fallback is harmless but not needed right now. If you're seeing a tenant where the existing discovery fails, let's discuss — otherwise let's skip this.
❌ Skip: cli/skills/foundry-web-search/SKILL.md
Agree with the intent (removing hardcoded connection ID), but this is a one-liner we can fold into a future commit. Not worth the merge overhead on its own.
❌ Skip: policy-engine/profiles/seccomp/azureclaw-strict.json
Verified this is purely formatting (identical syscall set). Reformatting is fine but adds noise to the diff for no functional benefit.
TL;DR: Please split out just the entrypoint.sh changes into a clean PR and I'll merge immediately. The proxy.rs changes need a rebase on main (prometheus 0.14 broke the API) and should be split from the streaming refactor.
…irect
Two bugs in entrypoint.sh that prevent the sandbox from starting correctly:
1. OpenClaw temp dirs missing: OpenClaw requires /tmp/openclaw-{UID} dirs
(owned by user, mode 700) but --tmpfs /tmp starts empty. Every openclaw
command crashes with 'Unable to create fallback OpenClaw temp dir'.
Fix: create dirs for current user, and for sandbox/router UIDs when
running as root (dev mode). Safe no-op in AKS (non-root) mode.
2. Router log redirect runs as root: the shell redirect in
'runuser -u router -- azureclaw-inference-router > /tmp/inference-router.log'
executes before the user switch, creating a root-owned log file.
Fix: wrap in sh -c so redirect runs as the router user, and rm -f
stale log files from previous runs.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- plugin.ts: Match Bing connections by metadata.type='bing_grounding' (Foundry API returns type='ApiKey', not 'GroundingWithBingSearch') - SKILL.md: Use placeholder instead of hardcoded connection ID - seccomp: Allow inotify syscalls (inotify_init, inotify_init1, inotify_add_watch, inotify_rm_watch) — fixes EPERM crash when OpenClaw watches MEMORY.md for persistent memory Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
dc2901d to
04dc9f5
Compare
|
Thanks woddll! Cherry-picked the entrypoint fix (commit 4712609 → 80fcf4f on main) — temp dir pre-creation and router log redirect are now on main. The Closing since the entrypoint changes are merged. 🎉 |
) * phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only) S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting. Replaces the legacy `spec.openclaw` field with a discriminated union `spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`. The `kind` discriminator selects which sibling struct is required; the others must be absent. Mutual exclusion enforced at admission via Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules `(self.kind == 'X') == has(self.x)`); controller-side defensive guard will land alongside the reconciler dispatch in a follow-up commit. Pre-release simplification: in-place v1alpha1 schema edit. No v1alpha2 cut, no conversion webhook (no installed base, per plan.md S10 + S13). One PR = one breaking change. What this commit delivers ------------------------- - `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum, `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`, `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`, `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent default per rubber-duck #9). `ClawSandboxSpec.runtime` is required on the wire; `Default` retained for test ergonomics, returns `OpenClaw` with empty config. - `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid wiping a populated value via merge patch). - `controller/src/crd.rs`: 8 new tests — PascalCase wire-format guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw, per-variant round-trip, BYO contractVersion required-not-default, serializer omits absent variants, runtimeKind status absence. - `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips from `["openclaw", ...]` to `["runtime", ...]`. New `runtime` block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains `runtimeKind`. `Runtime` printer column added. - `controller/src/reconciler/mod.rs`: minimal call-site fix — `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the build green. Full dispatch refactor (`RuntimeDeploymentPlan` per rubber-duck #2/#3) lands in step 2. What is NOT yet wired (intentional, follow-ups) ----------------------------------------------- - Reconciler dispatch per `runtime.kind` (single-seam plan struct). - `RuntimeReady` Condition machinery (folded into `build_running_status_patch` + `running_status_matches` per rubber-duck #1 to avoid status-merge churn). - OpenAI Agents / MAF deployment SKIP (must NOT silently use `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded + AdapterMissing). - `validate_runtime_shape` controller-side guard. - Examples / fixtures / CLI templates / convert / from_kagent migration. - CHANGELOG.md, audit doc. Verification ------------ - `cargo test --package azureclaw-controller`: 284/284 pass (8 new RuntimeSpec tests included). - `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean. - `cargo fmt --all`: applied. - Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime block, byo.required = [image, contractVersion], runtimeKind status field present, Runtime printer column added. NOTE: This commit alone is NOT mergeable on its own. Without the fixture/CLI/example migrations + reconciler dispatch, every existing `spec.openclaw` manifest in-tree would fail admission. Branch `phase2-multi-runtime-crd` will accumulate the remaining steps before the PR opens. Refs: plan.md S10.A1; rubber-duck critique applied (status merge risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL shape #6, contractVersion required #9, container name stays 'openclaw' #4). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures Completes the in-place v1alpha1 schema migration for the multi-runtime CRD spine. CRD types + Helm schema + reconciler reader landed in d11d41d; this finishes the long tail of CRD-emitting / CRD-reading sites the user called out ("every aspect of the code — including the cloud offload"). Cloud offload (the user's explicit ask): - controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw). - controller/src/reconciler/mod.rs:796: comment text aligned to new path. CLI emitters: - add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}. - convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape. - migrate.ts: --image help text references spec.runtime.openclaw.image. - migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning message aligned. - handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model. Fixtures + examples (8 yaml files): - examples/{basic,confidential,telegram}-agent/clawsandbox.yaml - examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml - tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied fixture left unchanged — orthogonal to this migration; static scanner is the active enforcement, not CRD validation.) Tests updated to assert the new shape: - add.test.ts: 4 assertions - convert.test.ts: 6 assertion blocks + 1 multi-container test - from_kagent.test.ts: 4 assertions Verification: - cargo test --package azureclaw-controller: 284/284 pass - cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean - cli npm test: 435/435 pass + 2 skipped - cli npm run typecheck: clean - grep confirms no remaining spec.openclaw emission/read sites; only intentional docstring/comment references documenting the legacy shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: runtime-aware status surface + AdapterMissing dispatch guard Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior two commits (d11d41d CRD spine; 3202d13 emitter migration) by wiring the runtime kind through the status surface and refusing to deploy a Pod for runtime kinds whose adapter has not yet shipped. Status surface - New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in controller/src/status/conditions.rs. - `build_running_status_patch` / `running_status_matches` / `build_overlay_status_patch` / `overlay_status_matches` take `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and append `RuntimeReady` to the conditions array (True/Reconciled on the running path, False/OverlayMode on overlay). Stamping inside the existing patch (rather than via a separate patch_status) avoids the merge-patch array overwrite that would erase the new Condition and re-introduce the resourceVersion-bump reconcile storm — see plan S10.A1 rubber-duck #1. - New `build_runtime_unsupported_status_patch` / `runtime_unsupported_status_matches` / `stamp_runtime_unsupported` helper trio mirrors the existing degraded_* trio. Stamps Degraded=True + Ready=False + RuntimeReady=False, all Reason=AdapterMissing. Reconciler dispatch - controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a static-str discriminator and explicitly skips namespace/SA/Deployment creation when the kind is not OpenClaw. Stamps AdapterMissing and returns Action::requeue(300s) BEFORE any K8s-resource builder is invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1 rubber-duck #2). - Status-patch call sites at :1481-1498 thread the runtime_kind_str. Tests - 5 new tests for the AdapterMissing helper trio (stamp shape, status-missing/runtime-mismatch idempotency rejection, settled-status match, transition-time preservation across repeat patches). - 4 existing tests updated for new conditions array shape + runtimeKind field (Running: 2 conds, Overlay: 4 conds). - 289/289 controller tests pass (was 284); cargo clippy clean; CLI 435/435 still green. Docs - CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking- change marker spanning the three commits. - docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full audit doc with threat model (silent fallthrough, status churn, CEL-disabled, BYO contract bypass, convert hard-fail), existing- implementation survey, wire-format invariants, test matrix, and S10.A2-A5 deferral list. Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam, per-variant image/entrypoint/env/agentCode resolution, BYO contract verifier, validate_runtime_shape defensive guard. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic) Locks the CRD wire shape now for three additional declared-roadmap runtimes so adding their adapters in a later slice is not a breaking schema change. The CRD becomes a public roadmap signal: customers can pin `spec.runtime.kind` today and know the schema won't shift under them. Tiering: Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3), MicrosoftAgentFramework (S10.A4) Tier 2 (placeholder, Phase 3+ adapters): SemanticKernel, LangGraph, Anthropic BYO (warn-only contract verifier in S10.A2) Schema changes - crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs `SemanticKernelConfig` (language: python|dotnet|java), `LangGraphConfig` (language: python|typescript), `AnthropicConfig` (pythonVersion). All three carry the universal agentCode + entrypoint + extraEnv shape — same as OpenAIAgents/MAF. - helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property blocks; kind enum extended in both spec + status surfaces; nested AgentCodeRef exactly-one CEL on every variant that carries code. - reconciler: runtime_kind_str match extended; AdapterMissing message enumerates Tier-2 placeholders. Behavior - All three Tier-2 kinds short-circuit through the existing `stamp_runtime_unsupported` path: Degraded + Ready=False + RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths; pure schema scaffold. Tests: 4 new round-trip tests (one per variant + a defaults check that SkLanguage and LangGraphLanguage default to python). 293/293 controller tests pass (was 289). Clippy clean. Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL matrix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…emo-clawshield (#225) * examples: lethal-trifecta-demo — reproduces Claude Cowork attack on AKS A reproducible launch-day demo anchored on three real, recent agentic-AI exploits: - Claude Cowork file-exfiltration (PromptArmor, Jan 2026) - Google Antigravity .env exfiltration (PromptArmor, Nov 2025) - EchoLeak / M365 Copilot (CVE-2025-32711, Jun 2025) All three exploit the lethal trifecta (Simon Willison): private data + untrusted content + exfil channel. The demo deploys two side-by-side namespaces — vanilla OpenClaw with a domain-only egress allowlist vs. a full AzureClaw stack — and shows six independent AzureClaw layers each catching the attack alone: 1. Inline Content Safety (Foundry DefaultV2 prompt-shield) 2. ToolPolicy URL+method allowlist (not just domain) 3. ClawIdentity strips attacker-controlled bearer 4. Egress-guard (UID 1000 iptables) 5. Token budget cap 6. AGT BehaviorMonitor auto-quarantine + tamper-evident audit chain Files: - README.md — threat model, citations, quick run - WALKTHROUGH.md — 7-min timed live/recorded script - bait/poisoned-skill.md — the 1pt-font injection (markdown form) - scenarios/00-namespaces.yaml - scenarios/01-naked-claw.yaml — vanilla Pod, falls to attack - scenarios/02-azureclaw-sandbox.yaml — full ClawSandbox CRD - scenarios/03-bait-server.yaml - scripts/{deploy,run-attack,verify-defense,teardown}.sh Leaves examples/demo-clawshield in place for now — that demo covers multi-tenancy / Kata isolation, which is a different story. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(examples): real READMEs for basic-agent / confidential-agent / demo-clawshield Three top-level entries in examples/README.md were either missing a README entirely (basic-agent, confidential-agent) or shipping a 14-line shell-comment stub with promised section headings and zero content (demo-clawshield). The stub was discoverable from the GitHub deep-link `#3-networkpolicy-default-deny-egress` and bounced to nothing. This adds proper READMEs: - examples/basic-agent/README.md — what it ships, default posture table, deploy + customize + cleanup, links to confidential-agent and lethal-trifecta-demo for variants - examples/confidential-agent/README.md — explicitly documents how it differs from basic-agent (single `isolation: confidential` field), Kata add-on prereq, runtimeClassName verification one-liner, links to blueprints/02-enterprise-self-hosted - examples/demo-clawshield/README.md — full content replacing the stub: what each YAML does, layer-per-phase mapping table, the this-vs-lethal-trifecta-demo orientation paragraph, pointer to docs/internal/DEMO.md for the 30-min timed walkthrough Cross-links between the four attack/security examples (basic-agent, confidential-agent, demo-clawshield, lethal-trifecta-demo) so users can navigate between them. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…hments, sibling trust race, final-deliverable rule Live multi-agent demo (analyst→viz→writer fan-out) surfaced four distinct breakages that each silently degraded the run while the parent agent still self-reported success. 1. Image generation 404 (router URL prefix regression) Foundry's account-scoped /openai/v1/images/generations endpoint does NOT accept the /api/projects/<project>/ URL prefix that chat-completions tolerates. Commit c13f302 unified everything through that prefix. In dev (raw Azure OpenAI account, no project path) it works; in AKS prod against a Foundry project endpoint the upstream returns a fast 404 and image_generation falls back to written descriptions. Fix: strip /api/projects/<name>/ from the upstream endpoint inside the images_generations handler before forwarding. Add strip_project_prefix helper + four unit tests (project-stripped, trailing-slash variant, AOAI passthrough, no-prefix passthrough). 2. foundry_code_execute drops container files The Responses API tool only walked output[].type=='message' for text. matplotlib PNGs / CSVs generated by code_interpreter live inside Foundry's per-run container and are referenced via container_file_citation annotations or { type: 'image' } entries in code_interpreter_call.outputs. They were never downloaded, so the demo's bar chart silently degraded to ASCII. Fix: collect every (container_id, file_id, filename) reference from both shapes, GET each via the new /openai/containers/... router route (added to foundry_standalone_routes), and write the bytes to /sandbox/.openclaw/workspace/. Append the local paths to the tool result so downstream tools (mesh_transfer_file, file_write) can ship them. Adds routerCallBinary helper for binary downloads through the router. 3. Sibling KNOCK race in parallel fan-out AGT_TRUSTED_PEERS is baked at spawn time and consumed once at sub-agent boot. When the parent spawns analyst → viz → writer in sequence, only writer (last) sees all siblings. analyst's parentTrustedAmids only contains parent — so when viz or writer later try to KNOCK analyst, the trust score is 0 + 0 = 0 and the KNOCK is rejected at threshold 500. The demo logs confirm only 1 of 3 sibling pairs ever opened a session. Fix: after every successful spawn, the parent broadcasts a peers_update message containing the new sibling's AMID to every already-running sibling. Each sub-agent now records the parent's AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and handles peers_update only from that AMID, extending its parentTrustedAmids set at runtime. 4. Sub-agent system prompt missing FINAL DELIVERABLE rule Sub-agents were told they could mesh_transfer_file artifacts to peers, but nothing forced them to mesh_transfer_file the FINAL artifact back to the parent before returning a summary. The writer's executive_brief.md sat in its local /sandbox forever while the parent reported success. Fix: append a hard rule to the sub-agent system prompt requiring mesh_transfer_file(to_agent='parent', ...) as the last action before any "task complete" reply, with one call per output file. Tests - inference-router: 643 lib tests pass (4 new strip_project_prefix tests). - inference-router: cargo clippy --all-targets clean. - runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only pre-existing warnings. - cli: 553 tests pass. Deployment - For #2: rebuild + push inference-router image. - For #1, #3, #4: rebuild + push sandbox image. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Per docs/implementation-plan.md §5.4. Behavioral conformance corpus — the net that catches 'endpoint returned 200 but skipped the crypto step' bugs. tests/conformance/ layout - package.json / tsconfig.json / vitest.config.ts — own workspace, fork-isolated pool so ratchet state doesn't leak across specs. - README.md — corpus table (9 corpora across Phases 0, 1, 3), provider-axis rule (plan §11.3), invariant discipline. - fixtures/README.md — vendored-source policy per principle §0.2 #10. - harness/index.ts — empty scaffold; helpers land per-PR. Phase 0 specs (all 'it.todo', 59 invariants total — vitest reports 59 todo / 0 pass / 0 fail, no false-green): - signal-x3dh.spec.ts (14) — key-exchange shape, symmetric ratchet invariants, base64 input hygiene (vendor patches #3, #4). - signal-knock.spec.ts (15) — KNOCK happy-path, trust threshold, relay disruption, wire-shape parity (vendor patches #5, #7, #8, #1/#2 timestamps, vendored<->AGT byte-identical KNOCK). - signal-negative.spec.ts (13) — ciphertext integrity, replay, session clobber (vendor patch #10), DoS surfaces. - sandbox-isolation.spec.ts (17) — seccomp, Landlock, egress-guard, router-as-only-network-path. Guarded by CONFORMANCE_E2E=1 (requires Kind; compat suite Kind harness wires it in Phase 1). Principle §0.2 #8 ('solid, not look-alike'): it.todo is a documented pending assertion, not a silently-passing no-op. Each spec's top comment cites the vendor patch or production bug it exists to prevent recurring. Principle §0.2 #10: every invariant that references an external protocol cites its upstream source (libsignal, RFC3339 'Z' suffix, Signal Double Ratchet spec) either inline or in the README. ci/no-stubs.sh already allow-lists tests/ subtrees; gate remains PASS. No new dependencies outside the pinned vitest / typescript / @types dev-deps already used by tests/compat/. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… validator
Implements plan §1.3 (Outage semantics) as a pure, deterministic decision
function. No I/O, no AGT import surface, clock injected.
inference-router/src/providers/outage.rs (new):
- OutageMode enum: Strict | CachedRead | DegradedDev
* serde camelCase; FromStr accepts camel/kebab/snake; Display; Default = Strict
- OutageConfig { mode, cached_ttl }
* validate_for_env(is_dev_env) rejects DegradedDev in prod
* rejects cached_ttl = 0 on CachedRead
* enforces MAX_CACHED_TTL = 15 min
- CachedDecision<T> with is_expired(ttl, now) — backwards-clock = expired
- OutageAction<T>: Deny { mode } | ServeCached { verdict } | AllowWithWarning
- decide_outage(config, cached, now) — pure, test-friendly
- 19 unit tests covering all three modes, cache freshness, clock skew,
serde round-trip, env validation, TTL bounds
inference-router/src/providers/mod.rs:
- remove placeholder OutageMode stub; re-export real types
controller/src/providers/mod.rs:
- mirrored OutageMode (from_spec, is_dev_only, validate_for_env)
- OutageModeError::DegradedDevInProd
- 4 new unit tests
docs/security-audits/2026-04-24-phase1-outage-semantics.md:
- STRIDE, principle-mapping, re-audit triggers; both sign-offs present.
No call-site in the router consumes decide_outage yet — that lands with
the first AGT provider. Landing the pure semantics first locks the
decision rules before any provider wiring pressures them.
Verification:
- cargo test --all: 106+155+15+26+3 = 305 passed (was 286, +19)
- cargo clippy --all-targets --all-features -- -D warnings: clean
- six CI gates PASS on the branch tip
Plan refs: §0.2 #1/#2/#3/#4/#5/#8/#9/#10 | §1.3 | §1.4
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
) * phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only) S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting. Replaces the legacy `spec.openclaw` field with a discriminated union `spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`. The `kind` discriminator selects which sibling struct is required; the others must be absent. Mutual exclusion enforced at admission via Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules `(self.kind == 'X') == has(self.x)`); controller-side defensive guard will land alongside the reconciler dispatch in a follow-up commit. Pre-release simplification: in-place v1alpha1 schema edit. No v1alpha2 cut, no conversion webhook (no installed base, per plan.md S10 + S13). One PR = one breaking change. What this commit delivers ------------------------- - `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum, `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`, `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`, `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent default per rubber-duck #9). `ClawSandboxSpec.runtime` is required on the wire; `Default` retained for test ergonomics, returns `OpenClaw` with empty config. - `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid wiping a populated value via merge patch). - `controller/src/crd.rs`: 8 new tests — PascalCase wire-format guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw, per-variant round-trip, BYO contractVersion required-not-default, serializer omits absent variants, runtimeKind status absence. - `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips from `["openclaw", ...]` to `["runtime", ...]`. New `runtime` block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains `runtimeKind`. `Runtime` printer column added. - `controller/src/reconciler/mod.rs`: minimal call-site fix — `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the build green. Full dispatch refactor (`RuntimeDeploymentPlan` per rubber-duck #2/#3) lands in step 2. What is NOT yet wired (intentional, follow-ups) ----------------------------------------------- - Reconciler dispatch per `runtime.kind` (single-seam plan struct). - `RuntimeReady` Condition machinery (folded into `build_running_status_patch` + `running_status_matches` per rubber-duck #1 to avoid status-merge churn). - OpenAI Agents / MAF deployment SKIP (must NOT silently use `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded + AdapterMissing). - `validate_runtime_shape` controller-side guard. - Examples / fixtures / CLI templates / convert / from_kagent migration. - CHANGELOG.md, audit doc. Verification ------------ - `cargo test --package azureclaw-controller`: 284/284 pass (8 new RuntimeSpec tests included). - `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean. - `cargo fmt --all`: applied. - Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime block, byo.required = [image, contractVersion], runtimeKind status field present, Runtime printer column added. NOTE: This commit alone is NOT mergeable on its own. Without the fixture/CLI/example migrations + reconciler dispatch, every existing `spec.openclaw` manifest in-tree would fail admission. Branch `phase2-multi-runtime-crd` will accumulate the remaining steps before the PR opens. Refs: plan.md S10.A1; rubber-duck critique applied (status merge risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL shape #6, contractVersion required #9, container name stays 'openclaw' #4). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures Completes the in-place v1alpha1 schema migration for the multi-runtime CRD spine. CRD types + Helm schema + reconciler reader landed in 98c629d; this finishes the long tail of CRD-emitting / CRD-reading sites the user called out ("every aspect of the code — including the cloud offload"). Cloud offload (the user's explicit ask): - controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw). - controller/src/reconciler/mod.rs:796: comment text aligned to new path. CLI emitters: - add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}. - convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape. - migrate.ts: --image help text references spec.runtime.openclaw.image. - migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning message aligned. - handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model. Fixtures + examples (8 yaml files): - examples/{basic,confidential,telegram}-agent/clawsandbox.yaml - examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml - tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied fixture left unchanged — orthogonal to this migration; static scanner is the active enforcement, not CRD validation.) Tests updated to assert the new shape: - add.test.ts: 4 assertions - convert.test.ts: 6 assertion blocks + 1 multi-container test - from_kagent.test.ts: 4 assertions Verification: - cargo test --package azureclaw-controller: 284/284 pass - cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean - cli npm test: 435/435 pass + 2 skipped - cli npm run typecheck: clean - grep confirms no remaining spec.openclaw emission/read sites; only intentional docstring/comment references documenting the legacy shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: runtime-aware status surface + AdapterMissing dispatch guard Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior two commits (98c629d CRD spine; 56ad076 emitter migration) by wiring the runtime kind through the status surface and refusing to deploy a Pod for runtime kinds whose adapter has not yet shipped. Status surface - New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in controller/src/status/conditions.rs. - `build_running_status_patch` / `running_status_matches` / `build_overlay_status_patch` / `overlay_status_matches` take `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and append `RuntimeReady` to the conditions array (True/Reconciled on the running path, False/OverlayMode on overlay). Stamping inside the existing patch (rather than via a separate patch_status) avoids the merge-patch array overwrite that would erase the new Condition and re-introduce the resourceVersion-bump reconcile storm — see plan S10.A1 rubber-duck #1. - New `build_runtime_unsupported_status_patch` / `runtime_unsupported_status_matches` / `stamp_runtime_unsupported` helper trio mirrors the existing degraded_* trio. Stamps Degraded=True + Ready=False + RuntimeReady=False, all Reason=AdapterMissing. Reconciler dispatch - controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a static-str discriminator and explicitly skips namespace/SA/Deployment creation when the kind is not OpenClaw. Stamps AdapterMissing and returns Action::requeue(300s) BEFORE any K8s-resource builder is invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1 rubber-duck #2). - Status-patch call sites at :1481-1498 thread the runtime_kind_str. Tests - 5 new tests for the AdapterMissing helper trio (stamp shape, status-missing/runtime-mismatch idempotency rejection, settled-status match, transition-time preservation across repeat patches). - 4 existing tests updated for new conditions array shape + runtimeKind field (Running: 2 conds, Overlay: 4 conds). - 289/289 controller tests pass (was 284); cargo clippy clean; CLI 435/435 still green. Docs - CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking- change marker spanning the three commits. - docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full audit doc with threat model (silent fallthrough, status churn, CEL-disabled, BYO contract bypass, convert hard-fail), existing- implementation survey, wire-format invariants, test matrix, and S10.A2-A5 deferral list. Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam, per-variant image/entrypoint/env/agentCode resolution, BYO contract verifier, validate_runtime_shape defensive guard. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic) Locks the CRD wire shape now for three additional declared-roadmap runtimes so adding their adapters in a later slice is not a breaking schema change. The CRD becomes a public roadmap signal: customers can pin `spec.runtime.kind` today and know the schema won't shift under them. Tiering: Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3), MicrosoftAgentFramework (S10.A4) Tier 2 (placeholder, Phase 3+ adapters): SemanticKernel, LangGraph, Anthropic BYO (warn-only contract verifier in S10.A2) Schema changes - crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs `SemanticKernelConfig` (language: python|dotnet|java), `LangGraphConfig` (language: python|typescript), `AnthropicConfig` (pythonVersion). All three carry the universal agentCode + entrypoint + extraEnv shape — same as OpenAIAgents/MAF. - helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property blocks; kind enum extended in both spec + status surfaces; nested AgentCodeRef exactly-one CEL on every variant that carries code. - reconciler: runtime_kind_str match extended; AdapterMissing message enumerates Tier-2 placeholders. Behavior - All three Tier-2 kinds short-circuit through the existing `stamp_runtime_unsupported` path: Degraded + Ready=False + RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths; pure schema scaffold. Tests: 4 new round-trip tests (one per variant + a defaults check that SkLanguage and LangGraphLanguage default to python). 293/293 controller tests pass (was 289). Clippy clean. Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL matrix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…emo-clawshield (#225) * examples: lethal-trifecta-demo — reproduces Claude Cowork attack on AKS A reproducible launch-day demo anchored on three real, recent agentic-AI exploits: - Claude Cowork file-exfiltration (PromptArmor, Jan 2026) - Google Antigravity .env exfiltration (PromptArmor, Nov 2025) - EchoLeak / M365 Copilot (CVE-2025-32711, Jun 2025) All three exploit the lethal trifecta (Simon Willison): private data + untrusted content + exfil channel. The demo deploys two side-by-side namespaces — vanilla OpenClaw with a domain-only egress allowlist vs. a full AzureClaw stack — and shows six independent AzureClaw layers each catching the attack alone: 1. Inline Content Safety (Foundry DefaultV2 prompt-shield) 2. ToolPolicy URL+method allowlist (not just domain) 3. ClawIdentity strips attacker-controlled bearer 4. Egress-guard (UID 1000 iptables) 5. Token budget cap 6. AGT BehaviorMonitor auto-quarantine + tamper-evident audit chain Files: - README.md — threat model, citations, quick run - WALKTHROUGH.md — 7-min timed live/recorded script - bait/poisoned-skill.md — the 1pt-font injection (markdown form) - scenarios/00-namespaces.yaml - scenarios/01-naked-claw.yaml — vanilla Pod, falls to attack - scenarios/02-azureclaw-sandbox.yaml — full ClawSandbox CRD - scenarios/03-bait-server.yaml - scripts/{deploy,run-attack,verify-defense,teardown}.sh Leaves examples/demo-clawshield in place for now — that demo covers multi-tenancy / Kata isolation, which is a different story. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(examples): real READMEs for basic-agent / confidential-agent / demo-clawshield Three top-level entries in examples/README.md were either missing a README entirely (basic-agent, confidential-agent) or shipping a 14-line shell-comment stub with promised section headings and zero content (demo-clawshield). The stub was discoverable from the GitHub deep-link `#3-networkpolicy-default-deny-egress` and bounced to nothing. This adds proper READMEs: - examples/basic-agent/README.md — what it ships, default posture table, deploy + customize + cleanup, links to confidential-agent and lethal-trifecta-demo for variants - examples/confidential-agent/README.md — explicitly documents how it differs from basic-agent (single `isolation: confidential` field), Kata add-on prereq, runtimeClassName verification one-liner, links to blueprints/02-enterprise-self-hosted - examples/demo-clawshield/README.md — full content replacing the stub: what each YAML does, layer-per-phase mapping table, the this-vs-lethal-trifecta-demo orientation paragraph, pointer to docs/internal/DEMO.md for the 30-min timed walkthrough Cross-links between the four attack/security examples (basic-agent, confidential-agent, demo-clawshield, lethal-trifecta-demo) so users can navigate between them. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…hments, sibling trust race, final-deliverable rule Live multi-agent demo (analyst→viz→writer fan-out) surfaced four distinct breakages that each silently degraded the run while the parent agent still self-reported success. 1. Image generation 404 (router URL prefix regression) Foundry's account-scoped /openai/v1/images/generations endpoint does NOT accept the /api/projects/<project>/ URL prefix that chat-completions tolerates. Commit 582eb39 unified everything through that prefix. In dev (raw Azure OpenAI account, no project path) it works; in AKS prod against a Foundry project endpoint the upstream returns a fast 404 and image_generation falls back to written descriptions. Fix: strip /api/projects/<name>/ from the upstream endpoint inside the images_generations handler before forwarding. Add strip_project_prefix helper + four unit tests (project-stripped, trailing-slash variant, AOAI passthrough, no-prefix passthrough). 2. foundry_code_execute drops container files The Responses API tool only walked output[].type=='message' for text. matplotlib PNGs / CSVs generated by code_interpreter live inside Foundry's per-run container and are referenced via container_file_citation annotations or { type: 'image' } entries in code_interpreter_call.outputs. They were never downloaded, so the demo's bar chart silently degraded to ASCII. Fix: collect every (container_id, file_id, filename) reference from both shapes, GET each via the new /openai/containers/... router route (added to foundry_standalone_routes), and write the bytes to /sandbox/.openclaw/workspace/. Append the local paths to the tool result so downstream tools (mesh_transfer_file, file_write) can ship them. Adds routerCallBinary helper for binary downloads through the router. 3. Sibling KNOCK race in parallel fan-out AGT_TRUSTED_PEERS is baked at spawn time and consumed once at sub-agent boot. When the parent spawns analyst → viz → writer in sequence, only writer (last) sees all siblings. analyst's parentTrustedAmids only contains parent — so when viz or writer later try to KNOCK analyst, the trust score is 0 + 0 = 0 and the KNOCK is rejected at threshold 500. The demo logs confirm only 1 of 3 sibling pairs ever opened a session. Fix: after every successful spawn, the parent broadcasts a peers_update message containing the new sibling's AMID to every already-running sibling. Each sub-agent now records the parent's AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and handles peers_update only from that AMID, extending its parentTrustedAmids set at runtime. 4. Sub-agent system prompt missing FINAL DELIVERABLE rule Sub-agents were told they could mesh_transfer_file artifacts to peers, but nothing forced them to mesh_transfer_file the FINAL artifact back to the parent before returning a summary. The writer's executive_brief.md sat in its local /sandbox forever while the parent reported success. Fix: append a hard rule to the sub-agent system prompt requiring mesh_transfer_file(to_agent='parent', ...) as the last action before any "task complete" reply, with one call per output file. Tests - inference-router: 643 lib tests pass (4 new strip_project_prefix tests). - inference-router: cargo clippy --all-targets clean. - runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only pre-existing warnings. - cli: 553 tests pass. Deployment - For #2: rebuild + push inference-router image. - For #1, #3, #4: rebuild + push sandbox image. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
) (#292) * Slice 4d.1 — mcpServerRefs plural CRD field + admission CEL Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the existing singular mcpServerRef, which is now deprecated and honored as a length-1 alias via effective_mcp_server_refs(). Controller-side changes: - New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET in status/conditions.rs::reason. - Reconciler mirror loop refactored to iterate effective_mcp_server_refs(). Singular-field use emits tracing::warn with McpSingularDeprecated. len > 1 short-circuits via degrade! macro until Slice 4d.2 wires per-server addressing — principles §3 honest 'not-yet-enforced' signal. - 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty + plural-wins-when-both-set. Controller suite: 555 passing. Admission CEL on deploy/helm/azureclaw/templates/crd.yaml: - Mutex: singular and plural cannot both be set. - maxItems: 8 (router-side scheme is sized for this). - Per-name uniqueness across mcpServerRefs. Out of scope for 4d.1 (queued for 4d.2): - Per-server jwks-{name}.json / tools-{name}.json file scheme. - Router-side McpServerRegistry + namespaced tool dispatch. - Stale-file sweep (DoD #6). - e2e fixture with ≥ 3 servers (DoD #1). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Slice 4d.2 — per-server McpServer mounts + router discovery (DoD #1 + #6) Closes Slice 4 DoD #1 (≥3 plural McpServers reachable e2e) at the mount-and-discovery layer, and DoD #6 (stale-file sweep) via reconciler-driven volume rebuild. Multi-JWKS OAuth + namespaced tool dispatch (DoD #3) follow in Slice 4d.3. Controller: - reconciler/mod.rs: replace len>1 short-circuit with full iteration over effective_mcp_server_refs(). Per-name volumes (mcp-jwks-<name>, mcp-signing-<name>) mounted at /etc/azureclaw/mcp/<name>/ and /etc/azureclaw/mcp-signing/<name>/. - First entry (idx 0) keeps legacy MCP_JWKS_PATH + MCP_SIGNING_KEY_DIR env vars for backwards compat with current single-JWKS OAuth path. - New MCP_JWKS_DIR=/etc/azureclaw/mcp env set once per pod. - governance_mounts.rs: new inject_container_env helper for idempotent env-var injection without a volume mount. - Removed obsolete PLURAL_MCP_SERVERS_UNSUPPORTED_YET reason constant. - CRD doc comment updated for 4d.2 mount layout + 4d.3 forward-ref. Router: - New mcp/registry.rs with scan() + discover_from_env(). At startup, reads MCP_JWKS_DIR, enumerates subdirs with parseable jwks.json, emits tracing::info!(servers=?, count=N) per discovery + warn! per skipped candidate. - Empty/missing dir handled gracefully (sandbox with zero mcpServerRefs is a valid steady state). - 7 registry unit tests + 4 inject_container_env unit tests. Stale-file sweep (DoD #6): reconcile rebuilds pod-spec from current refs; removed refs disappear via SSA. No explicit sweep code needed. Verified: 559 controller + 804 router tests pass, clippy -D warnings clean across workspace, cargo fmt --check clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
— OAuth half) (#293) Closes the OAuth half of Slice 4 DoD #3. Each McpServer now has its tokens validated against its own JWKS with a per-issuer audience and scopes override — all driven by a meta.json file the controller writes alongside jwks.json in the per-server ConfigMap. Controller-side producer: - mcp_server_reconciler.rs: ensure_jwks_configmap now writes a second key, meta.json, alongside jwks.json. Wire shape (camelCase): { issuer, audience?, scopes[] }. Optional/empty fields skipped for forward-compat. - New McpServerMeta struct + McpServerMeta::from_spec helper. Router-side consumer: - DiscoveredMcpServerMeta in mcp::registry, loaded adjacent to jwks.json during scan(). Missing meta is silently back-compat; malformed or empty-issuer meta is recorded in skipped (principles §3 — no silent failure). - OAuthVerifierConfig gains per_issuer_audience + per_issuer_scopes. verify_access_token now uses the per-issuer entry keyed off the token's iss claim, falling back to the global expected_audience / required_scopes when none is set. iss is already pinned per-issuer via JWKS lookup, so this lifts the audience/scope checks onto the same secure key. - New OAuthVerifierConfig::from_registry constructor: returns Ok(None) when no servers carry meta (caller falls back to legacy MCP_JWKS_PATH); hard-errors on conflicting JWKS for the same issuer. - main.rs build_mcp_router: prefers the registry-first multi-issuer path when MCP_PRODUCTION_MODE=true and the registry is non-empty; legacy single-JWKS path remains as fallback. Tests: - 4 new registry tests (meta load happy, empty-issuer skip, malformed parse skip, no-meta back-compat). - 6 new oauth tests (per-issuer audience accept across two servers, per-issuer audience cross-token rejection, per-issuer scopes apply per token, from_registry None on empty, from_registry happy with two issuers, from_registry conflict rejection). - 820 router lib tests pass (was 810), 559 controller bin tests pass. - cargo clippy --all-targets -- -D warnings clean. - cargo fmt --check clean. Out of scope (4d.4 will land): - Namespaced tool dispatch (/mcp/<server>/...) — waiting on a real upstream MCP forwarder backend so dispatch isn't scaffolding. Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Closes Slice 4 DoD #3 dispatch half: /mcp now forwards tool calls to upstream McpServer.spec.url instead of serving only the in-tree EchoDispatcher. Producer (controller): - McpServerMeta gains url + allowed_tools fields, written into meta.json alongside the OAuth metadata from Slice 4d.3. Consumer (router): - New mcp::forwarder module with RouterToolDispatcher implementing AsyncToolDispatcher. - Startup discovery POSTs tools/list to each upstream, filters through allowed_tools ("*" exposes all; explicit list selects named subset; empty fails closed), prefixes each tool with snake_case(server_name). - Servers with empty URL, empty allow-list, outbound OAuth requirement (deferred to 4d.5), unreachable upstream, or HTTP errors are recorded in skipped[] with operator-actionable reason and excluded from the catalog. Router still starts. - Dispatch splits name on first '.', looks up server, forwards tools/call to upstream URL. JSON-RPC errors collapse to is_error:true; 5xx surfaces as DispatchError::ExecutionFailed; bad names surface as DispatchError::UnknownTool. - McpRouteState gains with_tools() builder; build_mcp_router (now async) swaps the dispatcher when registry is non-empty. Scope: - Unauthenticated upstreams only. Outbound OAuth is Slice 4d.5 (requires per-server credential source). Servers requiring bearer credentials are deliberately skipped per §5 (no scaffolding — only ship the consumer we can drive end-to-end). - Single tools/list page (multi-page support deferred until a real upstream hits the cap). Test coverage: 15 new unit tests covering every skip reason, both allow-list semantics, namespaced dispatch (single + multi server), upstream JSON-RPC errors, HTTP 5xx, and unknown-tool branches. Backward compatibility: registries with no meta.json or no url fall through to the existing EchoDispatcher unchanged. Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replaces ClawSandbox.spec.networkPolicy.learnEgress: bool with egressMode: Strict | Learn across controller, router, CLI, Headlamp, helm CRD, and docs. Default = Learn. The Approval variant is deferred to Slice 5c (no consumer yet per principles §5). No back-compat shim: repo has no live deployments yet, so the legacy field was removed cleanly. Slice 5b DoD #3 closed. - controller/src/crd.rs: new EgressMode enum with #[derive(Default)] and #[default] on Learn. Replaces NetworkPolicyConfig.learn_egress. - controller/src/reconciler/mod.rs: emits EGRESS_MODE=strict|learn (no back-compat EGRESS_LEARN_MODE). - inference-router/src/routes/mod.rs: consumes EGRESS_MODE, defaults to Learn when unset. - inference-router/src/spawn/mod.rs: sub-agent reader/producer migrated. Internal SpawnRequest.learn_egress retained (not wire). - deploy/helm/azureclaw/templates/crd.yaml: learnEgress removed from schema; egressMode enum-validated. - cli/src/commands/{add,policy,handoff,handoff/helpers,up/sandbox_bringup, dev/local-k8s,operator/actions}.ts: every CR producer/reader migrated. - tools/headlamp-plugin/src/index.tsx: all 4 egress readers consume egressMode directly. - docs/use-cases.md, docs/egress-proxy.md: samples updated. Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…atency overclaims Second pass of the OSS-readiness audit, focused on the architecture diagrams (the first pass missed factual errors here). architecture-diagrams.md: - Diagram #3: audit record is hash-chained, append-only (NOT signed today — cryptographic signing of the chain head is on the roadmap, see security.md). Tamper-detection vs tamper-proof is a real distinction. - Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list), prose says nine + mentions ClawPairing as the controller-internal 10th. - Diagram #7: CRD relationship arrows were wrong against the actual Rust structs. Corrected: - policyRef → spec.governance.toolPolicyRef - mcpRefs → spec.governance.mcpServerRefs - inferenceRef → spec.inferenceRef (top-level, was already correct) - memoryRef → spec.memoryRef (top-level, was already correct) - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy. - CE 'sandboxRef' → spec.targetSandboxRef. - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected to every sandbox by the controller (no ref). Noted in prose. - EgressApproval added (it was missing from the diagram entirely); links to ClawSandbox via spec.sandbox (string name, not ref object). security.md: - Headline #1 'agent does not see Azure credentials. Period.' — softened with a dev-mode footnote matching the README hero. In azureclaw dev, agent+router share a container with separate UIDs but a kernel-level container escape defeats the boundary; the hard guarantee is the AKS path. Anchor points at architecture.md#two-modes. - Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation latency on the router hot path.' Microsecond was an overclaim without a benchmark to back it. architecture.md: - Controller row in the components table: 'watches the eight peer CRDs' → 'nine peer CRDs (plus controller-internal ClawPairing)'. All 9 mermaid blocks pass a bracket-balance sanity check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…#325) * docs: OSS-readiness pass — strip internal jargon + correct overclaims Doc + comment-only audit across user-facing docs ahead of OSS launch. No code paths touched. Factual corrections (the heaviest changes): - docs/security.md headline guarantee #4: drop 'signed by the router' claim. The audit log is hash-chained and detects modification but is not signed today; signing is on the v1.1 roadmap. - docs/security.md Layer 6 inference-safety table: * Content Safety: 'Always on, server-side' → 'Always on for Foundry-provider requests; Copilot/GitHub-Models providers do not return prompt_filter_results'. * Token budget: was 'not yet aggregated'; the router DOES aggregate per-tenant daily and monthly UTC counters with on-disk persistence (inference-router/src/budget.rs:200-260). * New 'Operator escape hatches' subsection documents the two env-var knobs (AZURECLAW_SUPPRESS_CONTENT_FLAGS, AZURECLAW_CONTENT_FLAG_MIN_SEVERITY) honestly. - README hero: 'agent never sees an Azure key' softened to call out that this is the AKS guarantee; dev mode co-locates agent + router in one container. - README/architecture: '31 commands' → '30+ commands'; remove unverified '18 Foundry API groups' count. - docs/architecture.md design goal #1 + #4: explicit dev-vs-prod scoping; 'same code path' → 'same data-path code with documented AZURECLAW_DEV_MODE branches'. - docs/architecture/a2a-gateway.md: port 8445 is config-locked but the mTLS listener itself is still being wired; operators should set A2A_GATEWAY_UPSTREAM_URL explicitly until the listener is GA. - mesh-plugin/src/agt-identity.ts comment: replace 'encrypted at rest with per-host KEK' claim with honest 'chmod 0600 is the real boundary' note (mirrors PR #324 identity-store fix). - .github/copilot-instructions.md: '__AGT_INITIALIZED env guard' → 'Symbol.for(agt-mesh-client)' (matches current code in runtimes/openclaw/src/index.ts:458-480). Internal-jargon strip in user-facing docs: - docs/api/lifecycle.md: remove 'Slice 4', 'Slice 4d.3/4d.4', 'Slice 0', 'Slice 1c', 'Slice 2a/2b/2c/2d.1', 'Slice 2d.2', 'Slice 3a' references — replaced with descriptive prose. - docs/api/conditions.md: drop 'Slice 1c invariant' and 'principles.md §3' references. - docs/architecture/agt-boundary.md, docs/security-mcp-top10.md, docs/cli-reference.md: drop 'Phase 5.2' shibboleth — keep the facts (vendored AgentMesh fork was retired upstream). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: deep-dive diagram audit — fix CRD field-name labels + signing/latency overclaims Second pass of the OSS-readiness audit, focused on the architecture diagrams (the first pass missed factual errors here). architecture-diagrams.md: - Diagram #3: audit record is hash-chained, append-only (NOT signed today — cryptographic signing of the chain head is on the roadmap, see security.md). Tamper-detection vs tamper-proof is a real distinction. - Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list), prose says nine + mentions ClawPairing as the controller-internal 10th. - Diagram #7: CRD relationship arrows were wrong against the actual Rust structs. Corrected: - policyRef → spec.governance.toolPolicyRef - mcpRefs → spec.governance.mcpServerRefs - inferenceRef → spec.inferenceRef (top-level, was already correct) - memoryRef → spec.memoryRef (top-level, was already correct) - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy. - CE 'sandboxRef' → spec.targetSandboxRef. - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected to every sandbox by the controller (no ref). Noted in prose. - EgressApproval added (it was missing from the diagram entirely); links to ClawSandbox via spec.sandbox (string name, not ref object). security.md: - Headline #1 'agent does not see Azure credentials. Period.' — softened with a dev-mode footnote matching the README hero. In azureclaw dev, agent+router share a container with separate UIDs but a kernel-level container escape defeats the boundary; the hard guarantee is the AKS path. Anchor points at architecture.md#two-modes. - Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation latency on the router hot path.' Microsecond was an overclaim without a benchmark to back it. architecture.md: - Controller row in the components table: 'watches the eight peer CRDs' → 'nine peer CRDs (plus controller-internal ClawPairing)'. All 9 mermaid blocks pass a bracket-balance sanity check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…diagrams (#327) Link each provider bullet in 'Pluggable inference backend' to the relevant architecture chapter and diagram: - GitHub Copilot → architecture.md#dev-mode + diagram #1 (dev mode pod) - Foundry / Azure OpenAI → architecture.md#prod-mode + diagrams #2 (prod mode pod) and #3 (data path) - GitHub Models → security.md (Content Safety caveat) + architecture.md data-path section (provider routing notes) Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Closes the rubber-duck research findings from the Phase 3 critique and Microsoft's Entra Agent ID design-patterns audit. CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM): - KarsSandbox.spec.meshAuth.customSecurityAttributes: BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which attributes (from a tenant-declared set) the controller should PATCH onto each per-sandbox agent identity. - AgentIdentityClient::patch_custom_security_attributes Graph client method. Constructs the documented CustomSecurityAttribute- Value envelope with the required @odata.type per attribute, inferred from the JSON value shape. - odata_type_for_value helper: maps serde_json::Value → '#String' | '#Int32' | '#Boolean' | '#Collection($T)' and rejects floats, nulls, nested objects, mixed-type arrays, and empty arrays with clear error messages BEFORE the call goes out. - Wired into ensure_agent_identity_for_sandbox: PATCH runs on every reconcile (idempotent on Graph). Failures surface as ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing silent missing-attribute drift. SCALE-OUT INVARIANT (rubber-duck #4): - Documented in agent_id_provisioning.rs module doc: the agent identity is keyed on KarsSandbox.metadata.uid (and cluster UID), with NO per-pod / per-replica / per-ordinal dimension. All replicas of one KarsSandbox share ONE agent identity. - New test tag_layout_excludes_per_pod_attributes pins the tag layout — any future PR that adds a 'kars-pod-' / 'kars-replica-' / 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix breaks the test. - Visibility change: AgentIdentityClient::tags_for is now pub(crate) so the cross-module invariant test can call it. BOOTSTRAP SCRIPTS (deploy/bicep/standalone/): - custom-security-attributes.sh: declares the recommended AgentGovernance set with 4 attributes — AgentClassification (Standard|Restricted|Confidential), DataSensitivity (Public|Internal|Confidential), ProductOwner, ManagedBy. Idempotent via az rest against the Graph beta endpoint. (The Microsoft.Graph Bicep extension does not yet ship typed attributeSets / customSecurityAttributeDefinitions resources, so this is a shell script that operators run once per tenant.) - conditional-access-baseline.sh: applies Microsoft's policy-autonomous-agents template — blocks sign-ins where the risk level meets the configured threshold (default: high), targeted via the ManagedBy=kars-controller attribute filter. Defaults to report-only state for safe rollout. Idempotent via upsert. FOUNDRY RBAC (research R5): - foundry-rbac.bicep grants Azure AI User on a Foundry resource to the BLUEPRINT SP (not per-agent). All derived agent identities inherit access — eliminates per-agent role-assignment churn. Supports RG-scoped (default) and resource-scoped assignment via foundryResourceName parameter. az bicep build clean, no warnings. DOCS: - docs/architecture/entra-agent-id/05-security-alignment.md: 6.8KB operator-facing runbook covering the bootstrap order, KarsSandbox YAML example, failure modes table, scale-out invariant rationale, Foundry RBAC inheritance. TESTS: - Controller: 802/802 (+13 new — 12 odata_type_for_value variants covering supported + rejected shapes, 1 scale-out invariant). - Bicep: foundry-rbac.bicep builds clean. JSON compiled. - Bash: both .sh scripts pass 'bash -n' syntax check. - CLI: 786/786 (unchanged). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs(entra-agent-id): import POC findings as architecture reference
Captures the end-to-end token-acquisition flow validated on real AKS
in the Microsoft tenant during the POC phase.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): scaffold Entra Agent ID auth machinery
Adds controller-side primitives for operating kars sandboxes as
per-sandbox Entra Agent Identities. Foundation only — pod-spec
integration lands in a follow-up commit.
New modules:
- auth_config.rs: KarsAuthConfig CRD (cluster-scoped singleton).
Tenant + blueprint + controller MI anchors plus optional
serviceManagementReference for Microsoft-style tenants.
- auth_config_reconciler.rs: materialises kars-auth-sidecar-env
ConfigMap with AzureAd__ and DownstreamApis__ env vars. Idle when
the CRD is absent (cluster stays in anonymous tier).
- agent_identity.rs: Graph client for per-sandbox agent identity SP
lifecycle. Implements IMDS to MI to blueprint to Graph chain
proven during POC.
- sidecar_injection.rs: pure-function pod-spec helpers (sidecar
container shape, pinned-identity env vars, egress-guard iptables).
CRD additions in crd.rs:
- KarsSandbox.spec.meshAuth.mode (auto/agent-id/anonymous)
- KarsSandbox.status.agentIdentity
786 controller tests pass, 14 new tests added.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli): auto-provision Entra Agent ID trust during kars up
Makes Entra Agent ID setup invisible to end users: when `kars up`
runs and the cluster does not yet have a `KarsAuthConfig/default`
resource, the new `mesh/agent_id_setup.ts` module idempotently
provisions the tenant trust anchor end-to-end:
1. Blueprint app via Graph (POST /v1.0/applications/ with
@odata.type=#Microsoft.Graph.AgentIdentityBlueprint),
including the optional serviceManagementReference required by
Microsoft-style enterprise tenants.
2. Blueprint service principal (visible in the Entra Agents portal).
3. Controller managed identity in the customer's subscription.
4. MI-as-FIC on the blueprint (issuer=login.microsoftonline.com),
the anti-loop-safe credential path proven by the POC.
5. KarsAuthConfig CR written to the cluster.
The step is non-fatal: if the user lacks `Agent ID Developer`,
`kars up` continues and the cluster runs in anonymous tier until
the role is granted and `kars mesh setup-trust` is rerun. This
matches the three-tier fallback model documented in
docs/architecture/entra-agent-id/.
New `--service-tree <guid>` flag on `kars up` (and KARS_SERVICE_TREE
env var) propagates the ServiceTree GUID to the blueprint creation.
No hardcoding — tenants that do not require it leave it empty.
Tests: 6 new unit tests for agent_id_setup (idempotence detection,
dry-run, env-var threading, error propagation). 775 CLI tests pass
(2 pre-existing skipped). typecheck + lint clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs+preflight: user-facing Entra Agent ID guidance
Adds the user-facing documentation layer that the previous controller
and CLI commits were missing. Every doc that mentioned the old
`api://agentmesh` flow now points to the new per-sandbox Entra Agent
ID model.
New docs:
- docs/agent-identity.md — day-1 and day-2 user guide. Prerequisites,
walk-through of what `kars up` does for auth, sub-agent semantics,
troubleshooting, teardown.
Updated docs:
- README.md — top-line mention of per-sandbox Entra Agent ID in the
`kars up` summary, linking to the new guide.
- docs/getting-started.md — Step 2.1 calls out the `Agent ID
Developer` role prerequisite; Step 2.2 expands the bring-up list to
include the preflight role check and the Entra trust provisioning.
- docs/permissions.md — rewrites the "Tenant-level (Entra ID)
considerations" section to describe the new model. Replaces
`api://agentmesh` failure rows with Entra-Agent-ID-specific ones
(CredentialInvalidLifetimeAsPerAppPolicy,
InvalidFederatedIdentityCredentialValue, missing sidecar).
- docs/cli-reference.md — documents the new `--service-tree` flag
+ Microsoft-corp example.
- docs/SUMMARY.md — wires agent-identity.md into the mdbook index.
New preflight check:
- cli/src/preflight.ts — calls `checkAgentIdRole` from agent_id_setup.
Warns (not blocks) when the signed-in user lacks the `Agent ID
Developer` directory role. The existing `api://agentmesh` warning
line is removed.
- cli/src/commands/mesh/agent_id_setup.ts — new exports:
- `checkAgentIdRole` returns hasRole/inconclusive/message via
Graph `/me/transitiveMemberOf`. Matches by role template id
(stable) AND display name (forward-compatible).
- `detectExistingBlueprint` returns whether the configured
blueprint already exists in Graph.
- `AgentIdSetupOptions.blueprintName` added so multi-cluster
deployments can share a tenant-wide blueprint.
Tests: 781 CLI tests pass (+6 new for checkAgentIdRole / blueprint
detect). Typecheck + lint clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(cli): UX polish — from-scratch resets context, surface CA block
Two papercuts surfaced when the user ran `kars up --from-scratch` on
a tenant that already had a previous deployment:
1. `--from-scratch` cleared the resume state but NOT the cached
deployment context, so the "fresh" run silently reused the prior
region/RG/Foundry endpoint instead of re-prompting. Fixed in
`cli/src/commands/up/preflight.ts` — when `fromScratch=true`, skip
the `loadContext()` prefill entirely and print a clear
"ignoring any cached deployment context" line. Adds the
`fromScratch?` field to `UpOptionsForPreflight`.
2. The Entra Agent ID preflight check correctly soft-failed on the
well-known AADSTS530084 Conditional Access token-binding block
(common in Microsoft-corporate tenants), but the warning was
generic and gave the user no actionable next step. Now detects
AADSTS530084 (and the related AADSTS65001/65002 missing-consent
codes) specifically and surfaces the exact `az login --scope
https://graph.microsoft.com//.default` workaround inline.
Docs: adds an `#az-cli-ca-block` anchor section to
`docs/agent-identity.md` so the inline preflight message can link
directly to the troubleshooting paragraph.
Tests: 2 new test cases in agent_id_setup.test.ts pinning the
AADSTS530084 + AADSTS65001 detection paths. Total CLI tests:
783 pass (+2 vs prior commit).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(controller): correct system namespace kars-system (was azureclaw-system)
Caught while auditing residual `azureclaw-` references after the
Azure/kars rebrand. The auth-config reconciler materialised the
sidecar env ConfigMap into `azureclaw-system`, but the actual
controller + helm chart deploy everything into `kars-system`. With
the wrong namespace the ConfigMap was unreachable from sandbox pods
even when the rest of the wiring was correct.
Two-line fix:
- controller/src/auth_config_reconciler.rs:61 — namespace constant.
- controller/src/auth_config.rs:43 — doc comment.
The rest of the controller already uses `kars-system` consistently
(pairing_reconciler, trust_graph_reconciler, signer_policy,
egress_approval_reconciler). This brings the new modules in line.
786 controller tests still pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli): kars mesh setup-trust --mode agent-id (standalone retry)
The Entra Agent ID auto-provisioning that `kars up` runs is now
also reachable as a standalone command — necessary for retrying
after a transient failure (e.g. AADSTS530084 Conditional Access
block on Microsoft Graph) without re-running every other `kars up`
phase.
cli/src/commands/mesh/setup-trust.ts grows a `--mode <agent-id|legacy>`
flag, defaulting to `agent-id`. Agent-id mode forwards to the same
`ensureAgentIdTrust` helper used by `kars up`. Legacy mode preserves
the original api://agentmesh app-registration flow for installations
that haven't migrated yet (slated for removal once all consumers
have switched).
Also surfaces:
- --service-tree <guid> for Microsoft-style tenants
- --cluster-name / --resource-group / --region for controller MI
override
- Clear AADSTS530084 + "Agent ID Developer missing" remediation
messages on failure.
Bonus: cli/src/commands/up.ts agentmesh-image import block from a
prior in-flight edit — imports agentmesh-relay-agt + agentmesh-
registry-agt from the public source ACR in --build mode so the
agentmesh deploy step does not block on ImagePullBackOff. (Future
improvement: build them locally from .agt-sdk when the AGT SDK
tarball is present.)
Tests + typecheck clean: 783 CLI tests pass, build green.
* chore(helm): install KarsAuthConfig CRD
Adds the helm template for the new cluster-scoped singleton CRD
introduced in c9ce68f. Without this, `helm upgrade` does not install
the CRD and `kars mesh setup-trust --mode agent-id` fails with
"the server doesn't have a resource type karsauthconfig" when it
tries to `kubectl apply` the CR.
The schema uses `x-kubernetes-preserve-unknown-fields: true` on
spec/status with required-field validation for the top-level
tenant / agentId / controller blocks. The canonical type definitions
remain in controller/src/auth_config.rs; the controller validates
spec shape on reconcile. A future PR will add karsauthconfig to the
existing helm-vs-Rust drift test in controller/src/helm_drift.rs.
helm lint clean. YAML parses.
* feat(cli+bicep): auto-fallback to Bicep when az CLI Graph is CA-blocked
End-to-end UX for Microsoft-corporate-style tenants where the Azure
CLI cannot acquire a Microsoft Graph token (AADSTS530084) — the same
block we hit repeatedly during the POC phase.
What changes:
- deploy/bicep/agent-id-trust.bicep — sub-scope Bicep template that
provisions everything the imperative path does: blueprint app + SP
(tagged EntraAgentId), controller MI, MI-as-FIC on the blueprint
using login.microsoftonline.com (universally allow-listed). Goes
through ARM + the Microsoft.Graph extension, which has its own auth
path and is not subject to az CLI CA token-binding policy.
- deploy/bicep/modules/controller-mi.bicep — RG-scope module for the
controller MI (Bicep needs RG scope for UAMI, parent runs at sub
scope to create the RG).
- deploy/bicep/bicepconfig.json — enables the Microsoft.Graph
extension (preview).
- cli/src/commands/mesh/agent_id_setup_bicep.ts — driver that runs
`az deployment sub create` against the template, parses outputs,
and writes the KarsAuthConfig CR.
- cli/src/commands/mesh/agent_id_setup.ts — adds
ensureAgentIdTrustAutoFallback() which tries the fast Graph REST
path first and transparently switches to the Bicep path on
AADSTS530084. Other Graph errors propagate unchanged (don't mask
real permission failures with a Bicep retry).
- cli/src/commands/up.ts — uses the auto-fallback wrapper so
`kars up` "just works" in any tenant.
- cli/src/commands/mesh/setup-trust.ts — same auto-fallback for
`--mode agent-id`, and a new `--mode bicep` for users who want to
skip the CLI attempt entirely.
- Also makes blueprintName configurable so multi-cluster deployments
in the same tenant share one blueprint by default.
Tests: 783 CLI pass, typecheck clean, helm-lint + bicep-build clean.
Operational note (not code): the AGT registry/relay images the user
manually built+pushed during this session were from a feature branch
(copilot/secure-mcp-agent-governance) that adds proof-of-possession
to /v1/agents — incompatible with SDK 3.7.0 which doesn't sign yet.
Rebuilding from tag v3.7.0 + restarting the deployments resolves the
422 flood. Documented for future kars releases to build from a known
AGT tag rather than HEAD.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(bicep): clean error reporting + linter-clean issuer URL
Three small fixes shaken out by the user's first end-to-end run of
`kars mesh setup-trust --mode bicep` in the Microsoft corporate
tenant:
1. deploy/bicep/agent-id-trust.bicep — the MI-as-FIC issuer was
hardcoded to `login.microsoftonline.com`. Bicep linter emits
`no-hardcoded-env-urls` as a warning on stderr, which the CLI
driver was mistaking for a deployment failure. Switched to
`environment().authentication.loginEndpoint` so the template is
cloud-portable AND the linter is silent.
2. cli/src/commands/mesh/agent_id_setup_bicep.ts — error capture
reworked to use execa's merged `all` stream so the real ARM
deployment error surfaces instead of being masked by stderr-only
linter warnings. Lines starting with `WARNING:` are stripped from
the error summary.
3. Same module — the final `kubectl apply KarsAuthConfig` is now
correctly recognised as a separate step from the Bicep deployment.
When the CRD isn't installed (e.g. older controller helm release),
the user sees:
"Bicep deployment succeeded, but KarsAuthConfig CRD is not
installed. All Entra resources are already created — just
install the CRD and re-run."
plus the exact `helm upgrade` command. The function returns the
Bicep result so callers know the Entra side is complete.
Also briefly probed the Microsoft.Graph Bicep extension's beta
channel (1.0.0) to see if it supports the `@odata.type` discriminator
for typed AgentIdentityBlueprint. It does not (`BCP037 The property
"@odata.type" is not allowed on objects of type
"Microsoft.Graph/applications"`). The Bicep path therefore creates a
regular Application tagged `EntraAgentId` — functional for runtime
but not visible under the Entra Agents portal page. Documented as a
limitation; users who need portal visibility for the typed form must
use Graph Explorer or PowerShell (both have different first-party
auth paths that aren't CA-blocked).
Bicep build clean, typecheck clean, 783 tests pass.
* fix(cli+docs): switch Graph calls from /v1.0 to /beta
User reported that the typed POST creating an
`#Microsoft.Graph.AgentIdentityBlueprint` works in Graph Explorer
against the /beta endpoint — and that's what makes the resulting
app visible in the Entra Agents portal. The /v1.0 endpoint accepts
the same body but does NOT route through the typed-resource
discriminator on the server, so the app ends up listed only under
App registrations.
Updates `cli/src/commands/mesh/agent_id_setup.ts` to use
/beta/applications, /beta/servicePrincipals, /beta/users, /beta/me,
and /beta/applications/{id}/federatedIdentityCredentials in every
Graph REST call. The `@odata.type` body field stays
`#Microsoft.Graph.AgentIdentityBlueprint` — only the URL prefix
changes.
docs/permissions.md updated to match (the manual escape-hatch
snippet showing `az rest --method POST --url
https://graph.microsoft.com/v1.0/applications/...` now reads `/beta/`).
Note: this only helps when the imperative Graph REST path runs
successfully. In tenants where the az CLI is Conditional-Access-
blocked from Graph (AADSTS530084), the CLI auto-falls back to the
Bicep ARM path. The Bicep Microsoft.Graph extension does NOT
support the @odata.type discriminator at v1.0 or beta (1.0.0), so
the Bicep-created app is functional but stays in the tag-based
detection mode (App registrations only, not Agents portal). This
is a current limitation of the extension, documented in
agent_id_setup_bicep.ts.
Tests: 783 CLI pass, typecheck clean.
* feat(cli): device-code re-login + exact-match Graph body for typed blueprint
Two coordinated changes to make the imperative Graph REST path
succeed in tenants with Conditional Access token-binding policy
(AADSTS530084) on the az CLI's first-party app:
1. Match Graph Explorer's exact working body shape:
- @odata.type value WITHOUT the leading `#` ("Microsoft.Graph.
AgentIdentityBlueprint", not "#Microsoft.Graph.…"). Both forms
are accepted by /v1.0; only the unprefixed form survives the
/beta route reliably per user testing.
- sponsors@odata.bind / owners@odata.bind URLs use /v1.0/users/.
Graph rejects /beta/users/ as an @odata.bind target.
- POST URL stays /beta/applications/ — the typed
AgentIdentityBlueprint resource only surfaces in the Entra
Agents portal page when created via the /beta route.
2. Auto device-code re-login on AADSTS530084:
- When `az rest` returns AADSTS530084, the helper does a one-shot
`az login --use-device-code --scope https://graph.microsoft.com//.default`
which goes through a different OAuth flow that often bypasses
the token-binding CA policy applied to the default interactive
flow.
- User is prompted in the terminal: "visit https://microsoft.com/
devicelogin, paste this code". After login, the original Graph
call is retried once. If it still fails, the
ensureAgentIdTrustAutoFallback wrapper proceeds to the Bicep
ARM path (untyped but functional).
Updates existing test fixture for checkAgentIdRole so the
device-code retry path is mocked. 783 CLI tests pass.
Together these should let Microsoft-corporate-tenant users get the
typed AgentIdentityBlueprint that surfaces in the Entra Agents
portal page, without needing to fall back to Bicep (which can only
produce the tag-based form).
* docs(agent-identity): document AADSTS530033 + Graph Explorer workaround
User hit the next layer of CA policy after device-code re-login:
`AADSTS530033` — "device must be Intune-managed" — applies to
both interactive and device-code flows of the Microsoft Azure CLI
first-party app. Bicep ARM path keeps working (different auth
surface), but it produces an untyped blueprint that only shows
under App registrations, not the Entra Agents portal page.
Updates the existing AADSTS troubleshooting section with:
- Table of auto-handled error codes + what the kars CLI does
- Explicit "even Bicep cannot produce the typed form" path:
Graph Explorer PATCH to upgrade the untyped blueprint app to
the typed AgentIdentityBlueprint discriminator in place
- Note that runtime is unaffected by typed-vs-untyped — only
portal categorisation differs
- Long-term Intune-enrolment recommendation
* fix(cli): kars mesh setup-trust short-circuits when already provisioned
Two bugs surfaced when the user re-ran setup-trust after the
initial successful Bicep + Graph Explorer flow:
1. `karsAuthConfigExists()` checked for "karsauthconfig/default" in
kubectl's `-o name` output, but newer clusters return the
fully-qualified form `karsauthconfig.kars.azure.com/default`.
The bug caused the existence check to always return false, so
the wrapper always tried the Graph REST path even when the trust
was already in place. Fix: match the invariant `/default`
suffix.
2. `kars mesh setup-trust --mode agent-id|bicep` did not consult
`karsAuthConfigExists()` at all, so every re-run triggered the
full provisioning attempt — which, in CA-blocked tenants, kicks
off a device-code login prompt the user has no reason to deal
with when nothing needs to be done. Fix: short-circuit at the
top of both modes when the CR already exists, with a clear
message and the `kubectl delete + retry` escape hatch for users
who genuinely want to re-provision.
Tests: 784 CLI pass (+1 new for the FQ-name kubectl output).
* fix(bicep-fallback): print full Graph Explorer PATCH for portal visibility
The Bicep `Microsoft.Graph` extension cannot set `@odata.type`, so the
Bicep-created blueprint is a plain `Application` that the kars runtime
uses fine but that the Entra portal's Agents page does not show. The
existing docs explained the workaround as a one-field PATCH that
sets `@odata.type` only — that does upgrade the type but the entry
*still* stays hidden in the portal because the Agents-page filter
also requires `sponsors` and `owners` to be set.
Changes:
- agent_id_setup_bicep.ts: at the end of every Bicep success path
(and also on the CRD-missing soft-failure branch), print a fully
populated Graph Explorer PATCH body the user can paste verbatim.
The body includes `@odata.type` + `sponsors@odata.bind` +
`owners@odata.bind` so the resulting typed blueprint actually
shows up under Entra portal → Identity → Agents.
We try `az ad signed-in-user show --query id -o tsv` to auto-fill
the user OID. In the very tenants where this matters (Microsoft
corp Macs without Intune enrollment) that command also fails with
AADSTS530084 — handled silently with a `<YOUR_USER_OID>` placeholder
and a one-line hint on where to find the OID in the Entra portal.
- docs/agent-identity.md: replace the misleading one-field PATCH
example with the full body, including notes about: no `#` prefix on
`@odata.type` in the request body, `/v1.0/users/` required in
`@odata.bind` even when the parent URL is `/beta/`, and a pointer
to the CLI's auto-generated copy-paste body.
Tests: 786 CLI pass (incl. 15 agent_id_setup tests).
* docs+cli: full delete+recreate runbook for typed blueprint (SP+FIC included)
A user hit the case where the in-place @odata.type PATCH was rejected
and they recreated the blueprint via Graph Explorer's POST /applications.
The recreated app then lacked an SP (so it couldn't receive RBAC role
assignments and stayed invisible in the Agents portal) and lacked the
MI-as-FIC (so the controller couldn't mint child identities).
Bicep creates all three resources (app + SP + FIC) but the Bicep
Microsoft.Graph extension cannot set the @odata.type discriminator
needed for a typed agentIdentityBlueprint — so Graph-Explorer recovery
remains the only path in CA-blocked tenants, and it must do all three
steps explicitly.
- agent_id_setup_bicep.ts: extend printPortalVisibilityHint() to print
the full 5-step recovery runbook (POST app, POST SP, POST FIC, kubectl
patch CR, optional DELETE old app) underneath the simpler in-place
PATCH path that's still tried first.
- docs/agent-identity.md: replace the one-line "delete + recreate"
hand-wave with the full POST sequence with all body shapes, plus the
kubectl jsonpath one-liner to extract tenantId + MI principalId from
the cluster.
Tests: 786 CLI pass.
* fix(controller): own egress-guard script in one place + fix iptables rule order
The original `agent_id_egress_rules()` shipped on `feat/entra-agent-id`
contained a security bug: every rule used `-A OUTPUT` (append). The
pre-existing baseline egress-guard script (currently emitted as a
`concat!` literal in `reconciler/mod.rs`) starts with
`-A OUTPUT --uid-owner 1000 -o lo -j ACCEPT`, so any later appended
rule blocking UID 1000 → 127.0.0.1:8080 would NEVER fire — the
loopback-allow would match first. This would have silently broken the
"agent cannot impersonate the router and mint downstream tokens"
boundary as soon as sidecar injection is wired up.
Rubber-duck critique caught this (finding #5). Fix:
- Refactor `agent_id_egress_rules()` to emit two correctly-positioned
rules. The sidecar-block uses `-I OUTPUT 1` (insert at chain head)
so it runs BEFORE the baseline loopback-allow. The router-IMDS
block can stay `-A` because no prior `--uid-owner 1001` rule exists.
- Drop the redundant UID 1000 → IMDS rule (catch-all `DROP UID 1000`
already covers it).
- Drop the explicit UID 1002 → IMDS ACCEPT (OUTPUT chain default is
ACCEPT and no UID 1002 restriction exists).
- New `build_egress_guard_command(agent_id_mode: bool)` composes the
full shell script: agent-id rules first (so `-I` semantics + script
text order both keep the security boundary), then the seven
baseline iptables lines, then a mode-specific echo. `&&`-chained so
any iptables failure aborts init-container startup (partial policy
is worse than no policy).
- Reconciler/mod.rs replaces the `concat!` literal with a call to the
new helper (currently always passing `false`; agent-id pass that
flips to `true` lands in a follow-up commit on the same branch).
Tests:
- New `egress_rules_use_insert_before_baseline_loopback_allow` — pins
the `-I OUTPUT 1` semantics. Direct regression for the security bug.
- New `egress_guard_command_legacy_mode_matches_existing_behaviour` —
byte-for-byte pin on the seven historical iptables lines so the
refactor is provably no-op for non-agent-id sandboxes.
- New `egress_guard_command_agent_id_mode_prepends_security_rules` —
asserts the sidecar REJECT appears BEFORE the loopback ACCEPT in
the script text (defence-in-depth even if a reader misreads -I).
- New `egress_guard_command_is_shell_safe_chained` — every step
starts with `iptables ` or `echo `.
789/789 controller tests pass.
* feat(controller): per-sandbox agent identity provisioning + sidecar injection
The end-to-end controller-side of agent-id mode. New module
`agent_id_provisioning` orchestrates the full flow that the rubber-duck
critique identified as the missing glue: resolve mesh-auth mode →
provision (or recover) the per-sandbox Entra Agent Identity via Graph
→ patch sandbox status → materialise the per-namespace sidecar env
ConfigMap → return a Ready outcome that the sandbox reconciler uses
to inject the sidecar container, pin the router env, and flip the
egress-guard into agent-id mode.
Architecture: model (D) from the critique — the controller provisions
the per-sandbox identity BEFORE pod creation, status-pins the appId,
and the router uses that pinned appId in every sidecar request. No
per-sandbox Secret; no sandbox-side workload identity binding.
## What this commit adds
### New module: `controller/src/agent_id_provisioning.rs`
- `ProvisionerCache` — process-wide cache of `AgentIdentityClient`s
keyed by blueprint client ID. Shares token caches and connection
pools across concurrent sandbox reconciles.
- `resolve_mesh_auth_mode` — pure-function 3-way resolution (Auto →
AgentId or Anonymous based on KarsAuthConfig readiness; explicit
AgentId without ready config surfaces a distinct `AuthConfigNotReady`
reason so operators can distinguish "tenant not set up" from
"auto-fallback to anonymous").
- `load_auth_config` — singleton fetch with explicit `Ok(None)` for
the 404 (anonymous-tier fallback) vs. `Err` for transient failures.
- `ensure_agent_identity_for_sandbox` — idempotent orchestration with
the three-step recovery flow the critique specified:
1. If `status.agentIdentity` recorded, GET via Graph; reuse on
200, reprovision on 404, requeue on 5xx.
2. If status empty, `list_cluster_agent_identities` filtered by
the `kars-sandbox-uid:<uid>` tag. Catches the
"Graph create succeeded but controller crashed before status
patch" crash window.
3. Otherwise create new, patch status, return.
- `materialise_sidecar_configmap` — copies the rendered sidecar env
into the sandbox namespace (`envFrom` cannot cross namespaces — the
critique caught this). Owned by the KarsSandbox via
`ownerReferences` so K8s garbage-collects on sandbox deletion.
### Modified: `controller/src/reconciler/mod.rs`
- `Context` gains `cluster_uid` (read from `kube-system` ns metadata
at startup — canonical "this cluster" identifier in K8s; falls back
to `KARS_CLUSTER_UID` env or generated string with a warning) and
`agent_id_cache: Arc<ProvisionerCache>`.
- `reconcile` calls `ensure_agent_identity_for_sandbox` BEFORE pod-spec
assembly. Match on outcome: `Skipped` → legacy path; `Ready` →
capture identity for downstream injection; `Failed` → patch
Degraded status with `AgentIdentityProvisioningFailed` reason and
requeue. No silent fallback to legacy on AgentId failure — explicit
user intent must not be downgraded.
- Pod-spec assembly:
- Egress-guard `command` flips to `build_egress_guard_command(true)`
when `agent_id_active.is_some()` — adds the security-critical
`-I OUTPUT 1` REJECT rule for UID 1000 → sidecar.
- Router `env` gains `PINNED_AGENT_IDENTITY_APP_ID` (the
per-sandbox appId) + `AUTH_SIDECAR_URL` (loopback :8080).
- Sidecar container is appended to `containers` after the
`runtimeClassName` block. Image pinned to the GA Microsoft
distroless build (overrideable via `KARS_SIDECAR_IMAGE`).
### Modified: `controller/src/auth_config_reconciler.rs`
- After successful ConfigMap apply, `patch_ready_status` patches
`phase=PHASE_READY` + `SidecarConfigMaterialized=True` condition.
This is the real readiness signal that the sandbox reconciler's
`resolve_mesh_auth_mode` gates on (per critique #7). Best-effort:
status-patch failure is logged but doesn't fail the reconcile.
- `build_condition_blueprint_ready` no longer `#[allow(dead_code)]` —
it's now consumed by `patch_ready_status`. Type renamed from
`BlueprintReady` (Graph-call check) to `SidecarConfigMaterialized`
(controller-side materialisation check) since the BlueprintReady
signal will come from a separate cluster-health probe.
## Test status
- 795/795 controller tests pass (+6 new for the new module).
- phase_taxonomy_guard pre-existing integration test passes (the
status patch uses `PHASE_READY` constant, not a string literal).
- Sidecar-injection tests still all pass (the iptables refactor from
the previous commit holds).
## Not yet in this commit (deferred to follow-up commits/PRs)
- Inference-router `sidecar_client.rs` that actually consumes the
pinned env vars (todays AgentIdentity unused on the router side).
- CLI changes: `kars up` VMSS-MI assignment; `kars mesh setup-trust
verify` end-to-end audit; Foundry RBAC assignment print.
- `agent_identity_reaper` for cleaning up orphan SPs.
- Sponsor user object IDs in KarsAuthConfig spec (currently empty
array passed; works in tenants where sponsors aren't required).
* fix(e2e): make Entra Agent ID chain work end-to-end on real AKS
Concludes a long live-debug session against kars-aks where we walked
the full controller → sidecar → router auth chain and fixed every
blocker until the sidecar successfully mints tokens from Entra and
relays them to the router. The remaining gate at end-of-chain is the
Microsoft corp tenant's Conditional Access policy on agent identities
(AADSTS53003 with a `capolids` claims challenge) — a tenant
configuration issue, not a code issue. The architecture is proven
correct.
- **Drop the GET-by-id verify on recorded identity.** Graph's
`GET /servicePrincipals/{id}` has multi-second eventual-consistency
after creation; treating a 404 as "stale, reprovision" creates a
runaway-creation loop (live cluster produced 70+ duplicate SPs per
minute before this fix). The reaper (separate PR) handles
out-of-band deletes via tag-scoped listing.
- **Idempotency guard on `patch_sandbox_status`.** Skip the SSA patch
when the recorded status already matches — same `lastTransitionTime`
drift pattern that bit the auth-config reconciler.
- **Drop ownerReference on the per-namespace sidecar CM.** KarsSandbox
lives in `kars-system`; the per-sandbox CM lives in `kars-<name>`.
K8s rejects cross-namespace owner refs with `OwnerRefInvalidNamespace`
and GC's the CM seconds after creation (debugged via the kubelet
events stream). Cleanup happens via namespace deletion instead.
- **Add apiVersion+kind to status patch body.** SSA patches without
the top-level type meta return `BadRequest: invalid object type:
/, Kind=`.
- **IMDS-first, WI fallback.** Workload-Identity-derived tokens cannot
be used as FIC assertions (Entra anti-loop AADSTS700231). The
controller MUST acquire its MI token via IMDS — which requires the
MI assigned to the AKS node-pool VMSS (`kars up` does this).
- **Permissive Graph list parser.** The agent identity list response
uses `agentAppId`, not `appId`. The strict-schema deserialiser
failed silently on every list call, so the tag-based recovery path
never found anything and the controller created a fresh SP every
reconcile. Fallback parser handles both field names.
- **Patch `phase=Ready`** so `resolve_mesh_auth_mode` has a real
signal to gate on (rubber-duck critique #7). No-op when the
observed status already matches the desired — avoids the SSA
`lastTransitionTime` reconcile loop.
- **`AUTH_SIDECAR_URL=http://localhost:8080`** not `http://127.0.0.1:8080`.
The Microsoft Entra SDK sidecar's HostFiltering middleware only
allows `Host: localhost`; calls with `Host: 127.0.0.1` are rejected
with `400 Bad Request - Invalid Hostname`. /etc/hosts in every pod
maps `localhost` → `127.0.0.1` so the loopback semantics are
identical.
- **`httpHeaders: Host=localhost:8080`** on the readiness/liveness
probes. The kubelet sends the pod IP as the default Host header,
which the sidecar rejects with the same 400. Without this override,
every sidecar stays ready=false even when fully functional.
- **emptyDir at /app/keys.** The sidecar's ASP.NET Data Protection
writes encryption keys to `/app/keys`; the distroless image has
/app owned by root and UID 1002 cannot mkdir there. Per-pod
emptyDir gives a writable scratch.
- Append the keys volume to `pod_spec.volumes` whenever the sidecar
is injected. Without this the container immediately crashes on
startup with `Access to the path '/app/keys' is denied`.
The Microsoft Entra SDK sidecar's response shape varies across
builds. We've observed three forms in the wild:
1. `{"AuthorizationHeader": "Bearer xxx", "ExpiresIn": 3600}` —
documented contract.
2. `"Bearer xxx"` — JSON-quoted string only.
3. `Bearer xxx` — plain text, no JSON.
The strict parser broke on shape 2/3 with `parse auth-sidecar JSON
response`, silently disabling agent-id auth for the entire pod. New
`parse_sidecar_body()` handles all three lenient + 5 dedicated tests.
Without this rule the controller's IMDS call times out (CILIUM
default-deny drops 169.254.169.254). IMDS is per-node link-local —
not externally routable, no exfiltration risk.
- KarsAuthConfig.spec.agentId.sponsorUserObjectIds — required in
tenants where Graph rejects agent identity creation without a
sponsor (live error: `No sponsor specified`).
- KarsSandbox.status.agentIdentity — appId/objectId/displayName/
createdAt. Without this in the CRD schema, SSA patches fail with
`field not declared in schema`.
- Grant the controller `get/list/watch/create/update/patch/delete`
on `karsauthconfigs` (and `/status`). Without this the auth-config
reconciler stays dormant ("KarsAuthConfig CRD not installed").
- 795/795 controller tests pass (+ 0 net change; the existing
sidecar_injection tests covered the URL constants we updated).
- 890/890 inference-router lib tests pass (+5 new
`parse_sidecar_body*` regression tests).
- agent_identity_reaper module (cleans Graph orphans by tag — needed
long-term to handle out-of-band deletes; for now duplicates require
a manual cleanup script).
- CLI flow: `kars up` should run `az identity federated-credential
create` for the controller SA so the WI fallback works in clusters
without VMSS-attached MI.
- Per-sandbox Foundry RBAC automation — verify already prints the
exact `az role assignment` commands; landing this would auto-grant.
```
✅ KarsAuthConfig provisioning (Bicep + Graph Explorer fallback)
✅ Controller MI assigned to AKS VMSSes
✅ Controller IMDS → MI → blueprint Graph token chain working
✅ Per-sandbox agent identity creation via Graph (with sponsors)
✅ Sandbox status.agentIdentity correctly recorded
✅ Per-namespace sidecar env ConfigMap materialized
✅ Sandbox pod contains 3 containers (openclaw + router + auth-sidecar)
✅ auth-sidecar healthy + listening on localhost:8080
✅ Router detects sidecar mode and routes auth through it (fail-closed)
✅ Sidecar mints Entra tokens for the pinned agent identity
✅ Sidecar relays token (or Entra error) back to router
⚠️ Foundry call returns 401 AADSTS53003 — Microsoft corp tenant's
Conditional Access policy blocks agent identities. This is the
same family of CA wall the user has been hitting all session for
their human account. Tenant configuration issue; the chain itself
is provably correct.
```
* fix(sidecar): bind on all interfaces so kubelet probes can reach /healthz
The Microsoft Entra SDK sidecar was bound to `127.0.0.1:8080` only,
which meant the kubelet's readiness/liveness probes (which connect to
the pod IP, not loopback) got `connection refused` at the TCP layer
before any HTTP exchange could happen. Overriding the probe's `Host`
header is useless when the connect itself fails.
Fix: bind on all interfaces (`http://+:8080`). Three defence-in-depth
layers keep cross-pod traffic out of the sidecar:
1. Sidecar's HostFiltering middleware accepts only `Host: localhost`;
anything else gets 400 Bad Request.
2. The sandbox NetworkPolicy has no ingress rule allowing port 8080,
so cross-pod packets are dropped before reaching the sidecar.
3. Egress-guard iptables rule REJECTs UID 1000 → 127.0.0.1:8080,
blocking in-pod privilege escalation from the agent container.
Verified live on kars-aks (sandbox kars-testrun):
- auth-sidecar: ready=true, restarts=0
- kubelet probe now succeeds via pod-IP connect + Host-header
override
- cross-pod connections are still rejected (no NP ingress rule)
This was the last container-level blocker keeping the chain from
running end-to-end. With this fix the sidecar mints real Entra tokens
and the router relays them to Foundry (where the only remaining gate
is the per-agent-identity Azure AI User role assignment).
* refactor(controller): pivot to shared-sidecar architecture (Phase 0)
This commit prepares the branch for the shared-sidecar redesign:
deletes the per-pod sidecar injection plumbing while preserving the
controller-side provisioning machinery, the Graph client, and all
hard-won bug fixes.
## What's removed
- `controller/src/sidecar_injection.rs` — per-pod container spec
helpers (build_sidecar_container, build_router_pinned_identity_env,
build_router_sidecar_url_env, egress-guard agent-id-mode rules).
In the new design the auth-sidecar runs ONCE per cluster as a
Helm-managed Deployment in `kars-system`; per-pod injection is no
longer needed.
- `cli/src/commands/mesh/vmss_mi_assign.ts` — VMSS managed-identity
assignment driver. Not needed when the sidecar runs in kars-system
with its own Workload Identity binding.
- `materialise_sidecar_configmap` per-namespace mirror in
`agent_id_provisioning.rs`. The shared sidecar consumes a single
`kars-system`-scoped ConfigMap managed by `auth_config_reconciler`.
- The agent-id-mode iptables rules in the egress-guard init container
(UID 1000 → REJECT 127.0.0.1:8080, UID 1001 → REJECT IMDS). Trust
boundary for cross-pod sidecar access moves to a NetworkPolicy on
the sidecar's namespace (added in a follow-up commit).
## What's preserved
- `controller/src/agent_identity.rs` — Graph client. Same provisioning
endpoints, same auth flow. With or without per-pod sidecar.
- `controller/src/agent_id_provisioning.rs` — provisioning loop with
idempotency, tag-based recovery from crashes, permissive Graph list
parser, status patch idempotency. Hard-won lessons preserved.
- `controller/src/auth_config.rs` + `auth_config_reconciler.rs` —
KarsAuthConfig CRD + its reconciler. Still cluster-level singleton
config; the sidecar just consumes a kars-system-scoped CM now.
- `inference-router/src/sidecar_client.rs` — URL-agnostic client.
The new design points it at the cluster Service DNS instead of
loopback; no code change needed.
- All Helm CRD templates, RBAC additions, IMDS NetworkPolicy egress
rule, the Bicep blueprint provisioning template, and the
Graph Explorer recovery runbook docs.
## What's reset to baseline behaviour
- `reconciler/mod.rs` egress-guard: back to the seven-line baseline
iptables script (the same one shipped on `main`). No agent-id-mode
branching; the security boundary for cross-pod sidecar access is
the sidecar-namespace NetworkPolicy, not pod-local iptables.
- Inference-router env injection: still pushes
`PINNED_AGENT_IDENTITY_APP_ID` + `AUTH_SIDECAR_URL` when an agent
identity is provisioned, but `AUTH_SIDECAR_URL` now points at
`http://entra-auth-sidecar.kars-system.svc:5000` instead of
loopback. Helm chart for the sidecar Deployment lands in the next
commit.
## Why this redesign
Microsoft's own Agent ID design-pattern docs and the Auth Sidecar
API support `?AgentIdentity=<appId>` so one sidecar instance can
mint tokens for any of the blueprint's children. We had been running
one sidecar per pod (N × 128MB), iptables-isolating cross-UID access
to localhost:8080. Switching to one shared sidecar (kars-system,
2 replicas, NetworkPolicy-gated) saves ~5x memory at 10 sandboxes,
eliminates the per-pod ASPNETCORE bind + probe gymnastics, and aligns
with the Microsoft.Identity.Web library's centralized MSAL cache.
Independent research review confirmed the per-pod model is one of
two documented patterns but is more resource-intensive; the
`?AgentIdentity=<appId>` shared-sidecar pattern is explicitly
documented as a valid alternative and the SDK's `AzureAd__ClientId`
takes the blueprint appId, making each child mintable on demand.
## Tests
785/785 controller tests pass. Inference-router tests untouched.
* feat(helm): shared entra-auth-sidecar Deployment (Phase 1)
One auth-sidecar Deployment per cluster (2 replicas, HA) serves all
sandboxes via the Microsoft AuthorizationHeaderUnauthenticated
?AgentIdentity=<appId> query parameter — replacing the per-pod
sidecar injection model from earlier iterations.
Templates added (all gated on `entraSidecar.enabled`, default false):
- auth-sidecar-serviceaccount.yaml - WI-annotated SA
- auth-sidecar-deployment.yaml - distroless sidecar, /healthz probes
- auth-sidecar-service.yaml - ClusterIP :5000
- auth-sidecar-networkpolicy.yaml - ingress only from sandbox routers
values.yaml gets the `entraSidecar:` config block populated by
`kars up` from the KarsAuthConfig CR.
Trust boundary rationale and credential-supply patterns (WI for OSS,
IMDS-MI for corp tenant) explained inline in each template.
Resource footprint: ~160MB cluster-total vs ~1.28GB at 10 sandboxes
under the per-pod model.
Tests: helm lint clean, helm template renders 4 resources when
enabled, 0 when disabled. Controller 785/785 unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): allow sandbox→kars-system:5000 egress for shared sidecar (Phase 2)
The shared entra-auth-sidecar lives in the kars-system namespace and
is reached by every sandbox's inference-router over TCP 5000. The
sandbox-policy NetworkPolicy now includes an explicit egress rule
selecting namespaces labeled
app.kubernetes.io/name=kars
app.kubernetes.io/component=system
on port 5000.
Also fixed the auth-sidecar's own ingress NetworkPolicy: the selector
now correctly targets pods labeled kars.azure.com/component=sandbox
(the pod label) rather than the non-existent
kars.azure.com/component=inference-router (which is a container, not
a pod, and so was never matchable).
Trust boundary is now two-sided:
- Sandbox-side: egress allows reaching only the kars-system namespace
on port 5000, where the sidecar Service is the only listener.
- Sidecar-side: ingress allows only pods in sandbox-labeled namespaces.
When entraSidecar.enabled=false at the Helm level, no auth-sidecar
exists and the egress rule is a harmless no-op (no destination to
reach).
Tests: controller 785/785, helm lint clean, helm template renders
correctly with enabled=true and enabled=false.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(router): tid+principal+aud+exp pinning on sidecar tokens (Phase 3)
Wire the previously-orphaned sidecar_client module into the
inference-router and harden it against token substitution attacks.
CONTROLLER (reconciler/mod.rs):
- Extends agent_id_active tuple to carry tenant_id alongside agent
identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via
ProvisioningOutcome::Ready.auth_spec.
- Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode
sandboxes. Without it, router-side tid pinning is disabled and
warns at boot (insecure, dev-only).
ROUTER (sidecar_client.rs, auth.rs, lib.rs):
- pub mod sidecar_client; — fix the orphaned-module bug.
- WorkloadIdentityAuth now consults SidecarClient first; sidecar mode
is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves
per-sandbox audit attribution in downstream Azure RBAC.
- from_env() now Result<Option<Self>>. Partial config (only URL or
only PINNED set) returns Err; WorkloadIdentityAuth::new() panics →
AKS surfaces as CrashLoopBackoff. No more silent fallback to a
different identity model.
- HARD-fail validate_token_claims on EVERY check (rubber-duck #1):
* tid mismatch / missing (cross-tenant guard)
* appid/azp mismatch when present (audit attribution guard)
— if both present, BOTH must match
* aud mismatch against per-service expected set (cache-poisoning
guard) — Foundry, Graph, OpenAI, Management, Search mapped
* exp in past or within 60s skew (stale-token guard)
* exp missing when tid-pinning enabled (unbounded-lifetime guard)
- Returns a TTL cap = exp - now - 60s; caller takes
min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching.
- JWT decoder tightened: requires EXACTLY 3 segments (rejects
2-segment unsecured JWS and 4-segment JWE).
- aud claim normalized: accepts both single-string (Entra default)
and array (RFC 7519); rejects out-of-spec shapes (number, bool).
- Soft WARN when both appid and azp are absent (some MSI tokens
legitimately omit both).
- New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning
state for operator visibility.
TESTS:
- 39 sidecar_client unit + wiremock integration tests, covering all
HARD-fail branches, the cache-skip on validation failure, the TTL
cap math, the partial-config Err path, JWT decoder edge cases
(2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud,
out-of-spec aud).
- Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client.
Rubber-duck review caught: (1) appid/azp should be hard not soft, (2)
partial env config must fail closed, (3) 3-segment requirement, (4)
aud validation against requested resource, (5) exp validation + TTL
cap. All five addressed in this commit.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli,bicep,controller): dual credential mode auto-detect (Phase 4)
Adds first-class support for two auth-sidecar credential supply
patterns and auto-detects which one works in the current Entra tenant.
PATTERN A (ManagedIdentityImds, default):
- Sidecar auths via SignedAssertionFromManagedIdentity against the
controller MI's IMDS endpoint. Corp-tenant safe; required when
the tenant's FIC issuer-allowlist policy blocks AKS OIDC
(Microsoft-corporate: InvalidFederatedIdentityCredentialValue).
- Bicep creates: blueprint + SP + controller MI + MI-as-FIC.
PATTERN B (WorkloadIdentity):
- Sidecar auths via SignedAssertionFilePath against the projected
K8s SA token. No per-cluster MI, no VMSS identity assignment.
- Bicep creates: blueprint + SP + SA-as-FIC pointing at the AKS
cluster's OIDC issuer URL.
CRD (controller/src/auth_config.rs):
- Added `controller.credentialMode` enum (ManagedIdentityImds default).
- Made `controller.managedIdentity{Client,Resource,Principal}Id`
Optional (required only in MI mode).
- Added is_valid_for_mode() validator.
CONTROLLER (auth_config_reconciler.rs):
- render_sidecar_env branches on credentialMode and emits the
correct AzureAd__ClientCredentials__0__SourceType. MI mode
emits ManagedIdentityClientId; WI mode emits
SignedAssertionFileDiskPath=/var/run/secrets/azure/tokens/azure-identity-token.
- Spec validation runs BEFORE ConfigMap materialisation. Invalid
specs (MI mode + empty clientId) surface as phase=Degraded
with InvalidCredentialMode condition — refuses to propagate
the misconfiguration to running sandboxes.
- New patch_degraded_status helper.
BICEP (deploy/bicep/agent-id-trust.bicep):
- New credentialMode parameter (allowed: ManagedIdentityImds |
WorkloadIdentity).
- Conditional resources: controller MI + MI-as-FIC only in
Pattern A; SA-as-FIC only in Pattern B.
- New aksOidcIssuerUrl parameter (required for Pattern B).
- New outputs: credentialMode, aksOidcIssuerUrl.
- az bicep build clean, no warnings.
CLI:
- ensureAgentIdTrust now accepts credentialMode (auto | WorkloadIdentity
| ManagedIdentityImds), aksClusterName, aksClusterResourceGroup.
- AUTO MODE: tries Pattern B first by discovering the AKS OIDC
issuer URL via az aks show, then creating an SA-as-FIC. On
InvalidFederatedIdentityCredentialValue (corp tenant signature),
falls back to creating the controller MI + MI-as-FIC.
- New TenantRejectedAksOidcIssuer sentinel for clean fallback
orchestration.
- New discoverAksOidcIssuerUrl helper — accepts explicit args or
walks kubeconfig + az aks list to guess.
- writeKarsAuthConfig strips MI fields when credentialMode is
WorkloadIdentity (matches the controller's Optional schema).
- setup-trust gains --credential-mode, --aks-cluster-name,
--aks-cluster-resource-group, --aks-oidc-issuer-url flags.
- Bicep wrapper threads credentialMode + aksOidcIssuerUrl through
to ARM.
TESTS:
- Controller: 789/789 (4 new — WI mode render, MI empty-field
rendering, is_valid_for_mode in MI and WI modes).
- Router: 1014/1014 (unchanged).
- CLI: 786/786 (2 new — credentialMode propagation through
dry-run with default + explicit).
- Bicep: az bicep build clean.
- Helm lint + template clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller,bicep,docs): security alignment (Phase 5)
Closes the rubber-duck research findings from the Phase 3 critique
and Microsoft's Entra Agent ID design-patterns audit.
CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM):
- KarsSandbox.spec.meshAuth.customSecurityAttributes:
BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which
attributes (from a tenant-declared set) the controller should
PATCH onto each per-sandbox agent identity.
- AgentIdentityClient::patch_custom_security_attributes Graph
client method. Constructs the documented CustomSecurityAttribute-
Value envelope with the required @odata.type per attribute,
inferred from the JSON value shape.
- odata_type_for_value helper: maps serde_json::Value →
'#String' | '#Int32' | '#Boolean' | '#Collection($T)' and
rejects floats, nulls, nested objects, mixed-type arrays, and
empty arrays with clear error messages BEFORE the call goes out.
- Wired into ensure_agent_identity_for_sandbox: PATCH runs on
every reconcile (idempotent on Graph). Failures surface as
ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing
silent missing-attribute drift.
SCALE-OUT INVARIANT (rubber-duck #4):
- Documented in agent_id_provisioning.rs module doc: the agent
identity is keyed on KarsSandbox.metadata.uid (and cluster UID),
with NO per-pod / per-replica / per-ordinal dimension. All
replicas of one KarsSandbox share ONE agent identity.
- New test tag_layout_excludes_per_pod_attributes pins the tag
layout — any future PR that adds a 'kars-pod-' / 'kars-replica-'
/ 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix
breaks the test.
- Visibility change: AgentIdentityClient::tags_for is now
pub(crate) so the cross-module invariant test can call it.
BOOTSTRAP SCRIPTS (deploy/bicep/standalone/):
- custom-security-attributes.sh: declares the recommended
AgentGovernance set with 4 attributes — AgentClassification
(Standard|Restricted|Confidential), DataSensitivity
(Public|Internal|Confidential), ProductOwner, ManagedBy.
Idempotent via az rest against the Graph beta endpoint.
(The Microsoft.Graph Bicep extension does not yet ship typed
attributeSets / customSecurityAttributeDefinitions resources,
so this is a shell script that operators run once per tenant.)
- conditional-access-baseline.sh: applies Microsoft's
policy-autonomous-agents template — blocks sign-ins where the
risk level meets the configured threshold (default: high),
targeted via the ManagedBy=kars-controller attribute filter.
Defaults to report-only state for safe rollout. Idempotent via
upsert.
FOUNDRY RBAC (research R5):
- foundry-rbac.bicep grants Azure AI User on a Foundry resource
to the BLUEPRINT SP (not per-agent). All derived agent
identities inherit access — eliminates per-agent role-assignment
churn. Supports RG-scoped (default) and resource-scoped
assignment via foundryResourceName parameter. az bicep build
clean, no warnings.
DOCS:
- docs/architecture/entra-agent-id/05-security-alignment.md:
6.8KB operator-facing runbook covering the bootstrap order,
KarsSandbox YAML example, failure modes table, scale-out
invariant rationale, Foundry RBAC inheritance.
TESTS:
- Controller: 802/802 (+13 new — 12 odata_type_for_value variants
covering supported + rejected shapes, 1 scale-out invariant).
- Bicep: foundry-rbac.bicep builds clean. JSON compiled.
- Bash: both .sh scripts pass 'bash -n' syntax check.
- CLI: 786/786 (unchanged).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(phase-7): live deploy bug fixes + Foundry RBAC fallback (Phase 7)
End-to-end live validation of the Phase 0-5 shared-sidecar architecture
against real Microsoft corp tenant (kars-aks). Three discoveries
addressed in this commit:
1) ASP.NET Core HostFiltering rejects Service-DNS callers (router fix)
- Microsoft Entra SDK auth-sidecar's HostFiltering middleware rejects
non-localhost Host headers regardless of AllowedHosts=* env.
- inference-router/src/sidecar_client.rs: always send
Host: localhost:5000 on sidecar requests.
- deploy/helm/kars/templates/auth-sidecar-deployment.yaml: set
AllowedHosts=* env for completeness (sidecar config source doesn't
honour it for dynamic-binding path).
- NetworkPolicy remains the real ingress boundary; bypassing host
filter does not weaken security.
2) Azure RBAC inheritance correction (sandbox_bringup.ts)
- Phase 5 originally assumed Foundry RBAC would inherit blueprint ->
derived agent identity SPs.
- Microsoft docs (concept-agent-id-design-patterns) clarify that
inheritable permissions = Microsoft Graph delegated permissions only.
Azure RBAC is per-principal.
- Confirmed empirically live: granting on blueprint SP alone did NOT
unblock chat; direct grant on each agent identity SP did.
- cli/src/commands/up/sandbox_bringup.ts: inline Bicep now grants
Cognitive Services OpenAI User on the blueprint SP as a
break-glass / fallback (kept for the case where the controller
can't grant per-agent). Inheritable-permissions language removed.
- Operator-facing error message documents the per-agent grant path
when ARM deployment fails for permission reasons.
- Phase 5b (future PR): controller must assign Azure RBAC per agent.
3) Multi-agent exec-brief demo: chain proven multi-agent
- Demo applies CRDs, brings up parent execbrief + 3 sub-agents.
- Each sub-agent gets its own typed Entra agent identity SP.
- Verified live (with all 4 SPs granted Cognitive Services OpenAI
User AND Azure AI User on the Foundry account via az rest):
* All 4 sandboxes booted in fail-closed sidecar mode
* 65 successful Foundry 200s across all 4 sandboxes in 10 min
* ZERO PermissionDenied responses after roles in place
* Mesh routing: parent dispatched to analyst, replied
* E2E encrypted relay: file transfer (analyst.json) ACKed
* AGT trust scoring: +0.8 reputation submitted, accepted=true
* NetworkPolicy: 0 egress denials, 0 ingress drops
- Demo's verify.json shows 3/9 checks pass — passing checks are the
kars-runtime mechanisms (sub-agents active, egress 0 denials,
telegram skipped). 6 failing checks ALL trace to one root cause:
Foundry project's Bing Grounding connection is not configured
(foundry_web_search requires bing project_connection_id). This is
orthogonal to our work; the auth chain is proven.
TESTS:
- inference-router sidecar_client: 39/39 pass.
- CLI: 786 pass + 2 skipped (no regressions).
- Bicep: existing modules build clean.
- Helm: lint clean, template renders correctly.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs(entra-agent-id): top-level README + migration guide (Phase 8)
Consolidates the entra-agent-id architecture documentation into a
navigable index and adds a migration guide for operators upgrading
from earlier per-pod-sidecar branch heads.
- docs/architecture/entra-agent-id/README.md (new top-level index):
* TL;DR architecture diagram
* Pattern A/B selection table
* Phase ledger linking each commit to its scope
* Key files map
* Live validation snapshot (verified on kars-aks 2026-05-28)
- docs/architecture/entra-agent-id/00-poc-archive.md:
* Previous POC README, renamed to archive. The POC scaffolding
drove early design but is no longer the canonical reference.
- docs/architecture/entra-agent-id/04-migration-guide.md (new):
* Step-by-step upgrade path from per-pod-sidecar branch heads
* Helm upgrade command with the entraSidecar values
* Per-agent role grant runbook (az rest workaround for the
CA-blocked az role assignment create path)
* Rollback procedure
* Known caveats (HostFiltering, MSAL cache per replica,
Bing Grounding orthogonal issue)
The numbered ordering reflects deployment lifecycle:
00 — historical POC (archived)
01 — runtime token flow
02 — alternative ACI flow (reference)
03 — original POC findings
04 — migration guide (this commit)
05 — security alignment (Phase 5)
Closes Phase 8 of the feat/entra-agent-id PR.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): per-agent ARM RBAC + agent identity cleanup (Phase 5b)
Eliminates the manual az role assignment runbook step from the Phase 5
migration guide by having the controller assign and revoke Azure RBAC
roles per agent identity automatically. Closes the orphan SP risk by
adding agent-identity deprovision to the KarsSandbox deletion
finalizer.
CRD (auth_config.rs):
- New `KarsAuthConfig.spec.foundryRbac: Vec<FoundryRbacAssignment>`.
- Each assignment carries an ARM `scope` and a list of built-in role
definition GUIDs to PUT against every per-sandbox agent identity SP
at provisioning time.
- Empty list (default) preserves the manual-grant workflow for
backward compatibility with existing kars deployments.
ARM REST extensions (agent_identity.rs):
- `arm_token()` — MI token for management.azure.com audience.
- `assign_role_to_agent_identity()` — PUT roleAssignment with a
deterministic UUIDv4 derived from (scope, principal, role), so
repeated PUTs are idempotent on Azure's side. Treats 200/201 as
success and 409 RoleAssignmentExists as success.
- `delete_role_assignments_for_principal()` — GET assignments
filtered by `principalId` (NOT combined with `atScope()` — Azure
REST returns 400 UnsupportedQuery), narrows to the requested scope
client-side, then DELETEs each. Used by the deletion finalizer.
- New helpers: `deterministic_assignment_guid()` (SHA-256 → UUIDv4)
+ `extract_subscription_id()` (scope parser).
- 9 new unit tests cover GUID stability, case-insensitivity, UUID
format, scope parsing happy-path + rejected forms.
Provisioning (agent_id_provisioning.rs):
- After Graph create (step 3c): iterate `spec.foundry_rbac` and assign
every (scope, role) tuple to the new identity. WARN-and-continue on
failure so the sandbox still boots; ARM RBAC converges on retry.
- Early-return path (recorded identity): RE-ASSERT the same role
assignments so existing sandboxes pick up retroactive RBAC config
the first time the operator adds `foundryRbac` to KarsAuthConfig.
Idempotent via the deterministic GUID.
- Logs `Phase 5b reconcile: re-asserting ARM role assignments` with
`foundry_rbac_entries` count for visibility.
Cleanup (agent_id_provisioning.rs + reconciler/mod.rs):
- New `cleanup_agent_identity_for_sandbox()` orchestrates the two
Azure-side teardown steps on sandbox delete:
1. DELETE role assignments held by the agent identity SP at every
configured `foundryRbac` scope.
2. Graph DELETE the agent identity SP itself.
- Hooked into the KarsSandbox deletion finalizer in
reconciler/mod.rs. Best-effort: failures logged as WARN but do not
block finalizer removal, so the K8s CR doesn't get stuck Terminating
when Azure is degraded. The existing orphan reaper backstops what
this path misses.
VERIFIED LIVE on kars-aks:
- All 5 existing sandboxes' agent identities (execbrief + 3 sub-agents
+ kars-testrun) now have exactly 2 role assignments each on the
Foundry account (Cognitive Services OpenAI User + Azure AI User),
re-asserted by the new code on the early-return path.
- Deleting `kars-testrun` sandbox: controller logged
"deleted role assignments for principal" → 2 deleted
"agent identity SP deprovisioned"
Verified Foundry now reports 0 role assignments for the deleted
appId — full cleanup successful.
Tests: 811/811 controller (+9 new for the GUID + scope helpers).
Build clean, helm template + lint clean.
One-time operator prerequisite: the controller's identity (the MI
whose IMDS token the controller uses for ARM calls — typically the
AKS kubelet/agentpool MI) must hold
`Microsoft.Authorization/roleAssignments/write` at every scope listed
in `foundryRbac`. The kars-recommended way is to grant Role Based
Access Control Administrator on the Foundry RG. This is intentionally
a one-time operator action documented in the migration guide rather
than auto-granted, since it crosses cluster/customer trust boundaries.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): scaffold MeshAuthBackend CRD field for Phase 6 design
Adds KarsAuthConfig.spec.meshAuthBackend (enum: Anonymous default,
EntraAgentIdentity opt-in) and optional meshAuthAudience override for
the next milestone — verified Entra-signed AGT mesh peer authentication.
Defaults preserve full backward compatibility (every existing cluster
keeps registering anonymously with no behaviour change). Operators on
clusters that have completed sidecar-based entrypoint mint + relay JWKS
verification (the next-PR work) flip the field to EntraAgentIdentity.
Includes a design doc (docs/architecture/entra-agent-id/06-mesh-trust-design.md)
capturing the three independent pieces that must land for end-to-end
enforcement (CRD scaffold here; sandbox entrypoint via sidecar; AGT
relay JWKS verification — last piece is upstream-coordination work).
Unit tests pin: default is Anonymous (backward compat), both variants
deserialize, unknown variants are rejected (no silent fallback).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(router,sandbox): /v1/mesh-token endpoint + Entra-signed mesh peer path
Phase 6.b — sandbox-side path for verified AGT mesh trust. When
KarsAuthConfig.spec.meshAuthBackend=EntraAgentIdentity, the controller
sets MESH_AUTH_BACKEND on the router, the router exposes
GET /v1/mesh-token, and entrypoint.sh acquires a verified-tier agent
identity token via the shared auth-sidecar instead of doing the legacy
direct Workload-Identity → Entra exchange.
Router changes:
- New inference-router/src/routes/mesh_token.rs route gated on
MESH_AUTH_BACKEND; returns 404 when disabled, 503 if sidecar not
configured, 502 on sidecar call failure, 200 with {access_token,
token_type, expires_in} on success
- Extended sidecar_client::resource_to_service_name to map
api://agentmesh* → AgentMesh service key
- Optional MESH_AUTH_AUDIENCE env override (defaults to
api://agentmesh/.default)
- 7 new mesh_token tests + agent-mesh resource-mapping test, all
serialised on a process-wide ENV_LOCK mutex to avoid env races
Controller changes:
- reconciler/mod.rs injects MESH_AUTH_BACKEND + MESH_AUTH_AUDIENCE env
on the sandbox's router container when the CRD field is set
- auth_config_reconciler.rs auto-emits the DownstreamApis__AgentMesh__*
cluster onto the shared sidecar when meshAuthBackend=EntraAgentIdentity
(skipped if operator has supplied an explicit AgentMesh entry)
- 4 new reconciler tests pin: anonymous default emits nothing, entra
variant emits expected env, custom audience honoured, operator-
supplied entry wins
Sandbox changes:
- entrypoint.sh: new branch BEFORE the legacy WI-exchange one. When
MESH_AUTH_BACKEND=EntraAgentIdentity, curl localhost:8443/v1/mesh-token,
export AGT_OAUTH_TOKEN on success, force AGT_TRUST_THRESHOLD=0 on
failure (anonymous-tier fail-open contract matches existing logic)
- Backward compat: when MESH_AUTH_BACKEND is unset (default), the
entrypoint follows the exact same path as before — no behaviour
change for existing clusters
Test results:
- 819 controller tests pass (was 815, +4)
- 930 router tests pass (was 923, +7)
- entrypoint.sh sh…
…cal-k8s, docker (#368) * docs(security-validation): cross-platform validation report — AKS, local-k8s, docker Comprehensive security & runtime validation across all three deploy modes, mapped to the 9-layer model in docs/security.md. Per-platform live evidence collected: - CRD presence + InferencePolicyCompiled/ToolPolicyCompiled/EgressAllowlistCompiled digests - Pod securityContext (UID, readOnlyRootFilesystem, capabilities, seccomp) - Workload Identity / Entra Agent ID federated token plumbing (AKS only) - entra-auth-sidecar token issuance + per-sandbox pinned_agent_id - Output authenticity (real Foundry calls, real URLs verified via HTTP 200) - AGT mesh KNOCK + E2E channel establishment per agent - Native AGT governance modules (PolicyEngine, AuditLogger, etc.) - NetworkPolicy enforcement + egress-guard caps - 9-check verify run on every platform Verified all three platforms 9/9 PASS with current main checks.py. Findings (no fixes applied — tracked for separate PRs): #1 HIGH (docker macOS) — UID 1000 reads Foundry API key (Docker Desktop UID virtualization) #2 HIGH (local-k8s) — API key + GitHub token in plaintext pod env (should be secretKeyRef) #3 MEDIUM (docker) — NET_ADMIN on persistent parent container (vs init-only in K8s mode) #4 MEDIUM (AKS, local-k8s) — blocklist-refresh CronJob failing every 6h (VAP collision) #5 MEDIUM (all) — audit JSONL not persisted (RO root FS, no emptyDir mount) #6 LOW (AKS) — AllowlistVerified=False on execbrief (inline endpoints, no cosign attestation) #7 LOW (AKS) — TrustGraph router-side enforcement not active (documented roadmap) All findings include a remediation plan in §10. None block the AKS verified-tier security story documented in docs/security.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> * docs(security-validation): add env-var inventory addendum for AKS containers Adds detailed per-container env-var analysis answering the question: 'do AKS containers have more env variables than they should?' Per-container inventories with full categorization: - openclaw container: 42 vars in 8 categories - inference-router container: 54 vars (router-only paths/toggles) Three additional findings on env-var hygiene: #8 LOW — OPENCLAW_GATEWAY_TOKEN exposed via env (should be file mount) #9 LOW — enableServiceLinks=true leaks internal cluster IPs (16 env vars) #10 LOW — possibly-redundant Foundry/Mesh-auth env vars on openclaw Headline confirmations: ✅ NO AZURE_OPENAI_API_KEY on either AKS container ✅ NO COPILOT_GITHUB_TOKEN on either AKS container ✅ Federated identity token mounted RO, never as env value ✅ Auth mode 'shared entra-auth-sidecar fail-closed, no WI/IMDS/API-key fallback' Comparison table across all 3 platforms included. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> --------- Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Summary
Two bugs in
entrypoint.shthat prevent the sandbox from starting correctly in local dev mode (azureclaw dev).Bug 1: OpenClaw temp dirs missing on startup
OpenClaw requires
/tmp/openclaw-{UID}directories to exist, be owned by the current user, and have mode 700. Since the container uses--tmpfs /tmp(starts empty), these dirs never exist. Everyopenclawcommand (gateway, tui, agent) crashes:Fix: Pre-create temp dirs in entrypoint after the
IS_ROOTdetection. Creates dir for the current user unconditionally, plus sandbox (1000) and router (1001) dirs when running as root (dev mode). Safe no-op in AKS non-root mode.Bug 2: Router log redirect runs as wrong user
The router startup line:
The shell redirect
>executes as root (beforerunuserswitches to UID 1001). The log file gets created owned by root, and the router process (UID 1001) can't write to it. Router silently fails to start.Fix: Wrap in
sh -cso the redirect happens after the user switch. Alsorm -fstale log files from previous runs.Testing
Validated on WSL2 with Docker Desktop:
azureclaw dev --build --no-agt→ clean start, no manual intervention