Skip to content

fix(entrypoint): pre-create OpenClaw temp dirs and fix router log redirect - #3

Closed
woddll wants to merge 2 commits into
Azure:mainfrom
woddll:fix/entrypoint-openclaw-tmpdir-and-router-log
Closed

woddll wants to merge 2 commits into
Azure:mainfrom
woddll:fix/entrypoint-openclaw-tmpdir-and-router-log

Conversation

@woddll

@woddll woddll commented Mar 24, 2026

Copy link
Copy Markdown

Summary

Two bugs in entrypoint.sh that prevent the sandbox from starting correctly in local dev mode (azureclaw dev).

Bug 1: OpenClaw temp dirs missing on startup

OpenClaw requires /tmp/openclaw-{UID} directories to exist, be owned by the current user, and have mode 700. Since the container uses --tmpfs /tmp (starts empty), these dirs never exist. Every openclaw command (gateway, tui, agent) crashes:

Error: Unable to create fallback OpenClaw temp dir: /tmp/openclaw-0

Fix: Pre-create temp dirs in entrypoint after the IS_ROOT detection. Creates dir for the current user unconditionally, plus sandbox (1000) and router (1001) dirs when running as root (dev mode). Safe no-op in AKS non-root mode.

Bug 2: Router log redirect runs as wrong user

The router startup line:

$AS_ROUTER azureclaw-inference-router > /tmp/inference-router.log 2>&1 &

The shell redirect > executes as root (before runuser switches to UID 1001). The log file gets created owned by root, and the router process (UID 1001) can't write to it. Router silently fails to start.

Fix: Wrap in sh -c so the redirect happens after the user switch. Also rm -f stale log files from previous runs.

Testing

Validated on WSL2 with Docker Desktop:

  • azureclaw dev --build --no-agt → clean start, no manual intervention
  • Temp dirs created with correct ownership and mode 700
  • Router starts successfully, logs writing correctly
  • End-to-end chat working (gateway → router → Azure OpenAI → response)

@woddll

woddll commented Mar 24, 2026

Copy link
Copy Markdown
Author

woddll please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.

@microsoft-github-policy-service agree [company="{your company}"]

Options:

  • (default - no company specified) I have sole ownership of intellectual property rights to my Submissions and I am not making Submissions in the course of work for my employer.
@microsoft-github-policy-service agree
  • (when company given) I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.
@microsoft-github-policy-service agree company="Microsoft"

Contributor License Agreement

Contribution License Agreement

This Contribution License Agreement (“Agreement”) is agreed to by the party signing below (“You”), and conveys certain license rights to Microsoft Corporation and its affiliates (“Microsoft”) for Your contributions to Microsoft open source projects. This Agreement is effective as of the latest signature date below.

  1. Definitions.
    “Code” means the computer software code, whether in human-readable or machine-executable form,
    that is delivered by You to Microsoft under this Agreement.
    “Project” means any of the projects owned or managed by Microsoft and offered under a license
    approved by the Open Source Initiative (www.opensource.org).
    “Submit” is the act of uploading, submitting, transmitting, or distributing code or other content to any
    Project, including but not limited to communication on electronic mailing lists, source code control
    systems, and issue tracking systems that are managed by, or on behalf of, the Project for the purpose of
    discussing and improving that Project, but excluding communication that is conspicuously marked or
    otherwise designated in writing by You as “Not a Submission.”
    “Submission” means the Code and any other copyrightable material Submitted by You, including any
    associated comments and documentation.
  2. Your Submission. You must agree to the terms of this Agreement before making a Submission to any
    Project. This Agreement covers any and all Submissions that You, now or in the future (except as
    described in Section 4 below), Submit to any Project.
  3. Originality of Work. You represent that each of Your Submissions is entirely Your original work.
    Should You wish to Submit materials that are not Your original work, You may Submit them separately
    to the Project if You (a) retain all copyright and license information that was in the materials as You
    received them, (b) in the description accompanying Your Submission, include the phrase “Submission
    containing materials of a third party:” followed by the names of the third party and any licenses or other
    restrictions of which You are aware, and (c) follow any other instructions in the Project’s written
    guidelines concerning Submissions.
  4. Your Employer. References to “employer” in this Agreement include Your employer or anyone else
    for whom You are acting in making Your Submission, e.g. as a contractor, vendor, or agent. If Your
    Submission is made in the course of Your work for an employer or Your employer has intellectual
    property rights in Your Submission by contract or applicable law, You must secure permission from Your
    employer to make the Submission before signing this Agreement. In that case, the term “You” in this
    Agreement will refer to You and the employer collectively. If You change employers in the future and
    desire to Submit additional Submissions for the new employer, then You agree to sign a new Agreement
    and secure permission from the new employer before Submitting those Submissions.
  5. Licenses.
  • Copyright License. You grant Microsoft, and those who receive the Submission directly or
    indirectly from Microsoft, a perpetual, worldwide, non-exclusive, royalty-free, irrevocable license in the
    Submission to reproduce, prepare derivative works of, publicly display, publicly perform, and distribute
    the Submission and such derivative works, and to sublicense any or all of the foregoing rights to third
    parties.
  • Patent License. You grant Microsoft, and those who receive the Submission directly or
    indirectly from Microsoft, a perpetual, worldwide, non-exclusive, royalty-free, irrevocable license under
    Your patent claims that are necessarily infringed by the Submission or the combination of the
    Submission with the Project to which it was Submitted to make, have made, use, offer to sell, sell and
    import or otherwise dispose of the Submission alone or with the Project.
  • Other Rights Reserved. Each party reserves all rights not expressly granted in this Agreement.
    No additional licenses or rights whatsoever (including, without limitation, any implied licenses) are
    granted by implication, exhaustion, estoppel or otherwise.
  1. Representations and Warranties. You represent that You are legally entitled to grant the above
    licenses. You represent that each of Your Submissions is entirely Your original work (except as You may
    have disclosed under Section 3). You represent that You have secured permission from Your employer to
    make the Submission in cases where Your Submission is made in the course of Your work for Your
    employer or Your employer has intellectual property rights in Your Submission by contract or applicable
    law. If You are signing this Agreement on behalf of Your employer, You represent and warrant that You
    have the necessary authority to bind the listed employer to the obligations contained in this Agreement.
    You are not expected to provide support for Your Submission, unless You choose to do so. UNLESS
    REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING, AND EXCEPT FOR THE WARRANTIES
    EXPRESSLY STATED IN SECTIONS 3, 4, AND 6, THE SUBMISSION PROVIDED UNDER THIS AGREEMENT IS
    PROVIDED WITHOUT WARRANTY OF ANY KIND, INCLUDING, BUT NOT LIMITED TO, ANY WARRANTY OF
    NONINFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
  2. Notice to Microsoft. You agree to notify Microsoft in writing of any facts or circumstances of which
    You later become aware that would make Your representations in this Agreement inaccurate in any
    respect.
  3. Information about Submissions. You agree that contributions to Projects and information about
    contributions may be maintained indefinitely and disclosed publicly, including Your name and other
    information that You submit with Your Submission.
  4. Governing Law/Jurisdiction. This Agreement is governed by the laws of the State of Washington, and
    the parties consent to exclusive jurisdiction and venue in the federal courts sitting in King County,
    Washington, unless no federal subject matter jurisdiction exists, in which case the parties consent to
    exclusive jurisdiction and venue in the Superior Court of King County, Washington. The parties waive all
    defenses of lack of personal jurisdiction and forum non-conveniens.
  5. Entire Agreement/Assignment. This Agreement is the entire agreement between the parties, and
    supersedes any and all prior agreements, understandings or communications, written or oral, between
    the parties relating to the subject matter hereof. This Agreement may be assigned by Microsoft.

@microsoft-github-policy-service agree

@woddll
woddll force-pushed the fix/entrypoint-openclaw-tmpdir-and-router-log branch from 053dd90 to f9b9ea4 Compare March 25, 2026 20:58

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR! Reviewed all 5 files — here's the breakdown:

✅ Accept: sandbox-images/openclaw/entrypoint.sh

The temp-dir pre-creation and router log redirect fix look solid. Correct UID handling, mode 700, and the sh -c wrapper for the redirect is the right approach. Happy to merge this part.

❌ Reject: inference-router/src/proxy.rs

This conflicts with changes already on main:

  1. We upgraded prometheus 0.13 → 0.14 on main, which requires with_label_values to use consistent types. Your diff reverts those back to &str literals — this will fail to compile against main.
  2. The embedding model routing fix (extracting model from request body) — we already landed an equivalent fix in routes.rs for the AKS/Workload-Identity path. Your fix covers the API-key path, which is good in principle, but the larger streaming refactor (removing inject_stream_usage and the token-metrics wrapper) is a separate concern and needs its own discussion.
  3. The streaming doc-comment changes are fine but tangled with the above.

Suggestion: Split this into two PRs — (a) the API-key embedding routing fix rebased on current main, and (b) the streaming simplification as a separate discussion.

❌ Skip: cli/src/plugin.ts (Bing discovery)

We already have type + properties.category discovery on main, and it works on our dev tenant. The metadata.type fallback is harmless but not needed right now. If you're seeing a tenant where the existing discovery fails, let's discuss — otherwise let's skip this.

❌ Skip: cli/skills/foundry-web-search/SKILL.md

Agree with the intent (removing hardcoded connection ID), but this is a one-liner we can fold into a future commit. Not worth the merge overhead on its own.

❌ Skip: policy-engine/profiles/seccomp/azureclaw-strict.json

Verified this is purely formatting (identical syscall set). Reformatting is fine but adds noise to the diff for no functional benefit.


TL;DR: Please split out just the entrypoint.sh changes into a clean PR and I'll merge immediately. The proxy.rs changes need a rebase on main (prometheus 0.14 broke the API) and should be split from the streaming refactor.

donghapaek and others added 2 commits March 26, 2026 10:31
…irect

Two bugs in entrypoint.sh that prevent the sandbox from starting correctly:

1. OpenClaw temp dirs missing: OpenClaw requires /tmp/openclaw-{UID} dirs
   (owned by user, mode 700) but --tmpfs /tmp starts empty. Every openclaw
   command crashes with 'Unable to create fallback OpenClaw temp dir'.
   Fix: create dirs for current user, and for sandbox/router UIDs when
   running as root (dev mode). Safe no-op in AKS (non-root) mode.

2. Router log redirect runs as root: the shell redirect in
   'runuser -u router -- azureclaw-inference-router > /tmp/inference-router.log'
   executes before the user switch, creating a root-owned log file.
   Fix: wrap in sh -c so redirect runs as the router user, and rm -f
   stale log files from previous runs.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- plugin.ts: Match Bing connections by metadata.type='bing_grounding'
  (Foundry API returns type='ApiKey', not 'GroundingWithBingSearch')
- SKILL.md: Use placeholder instead of hardcoded connection ID
- seccomp: Allow inotify syscalls (inotify_init, inotify_init1,
  inotify_add_watch, inotify_rm_watch) — fixes EPERM crash when
  OpenClaw watches MEMORY.md for persistent memory

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@woddll
woddll force-pushed the fix/entrypoint-openclaw-tmpdir-and-router-log branch from dc2901d to 04dc9f5 Compare March 26, 2026 17:33
@pallakatos

Copy link
Copy Markdown
Collaborator

Thanks woddll! Cherry-picked the entrypoint fix (commit 4712609 → 80fcf4f on main) — temp dir pre-creation and router log redirect are now on main.

The plugin.ts Bing discovery fallback (metadata?.type) isn't needed right now (we already have type + properties.category discovery working). If you hit a tenant where it fails, please open a focused PR and we'll add it.

Closing since the entrypoint changes are merged. 🎉

Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Apr 28, 2026
)

* phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only)

S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting.

Replaces the legacy `spec.openclaw` field with a discriminated union
`spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`.
The `kind` discriminator selects which sibling struct is required;
the others must be absent. Mutual exclusion enforced at admission via
Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules
`(self.kind == 'X') == has(self.x)`); controller-side defensive guard
will land alongside the reconciler dispatch in a follow-up commit.

Pre-release simplification: in-place v1alpha1 schema edit. No
v1alpha2 cut, no conversion webhook (no installed base, per plan.md
S10 + S13). One PR = one breaking change.

What this commit delivers
-------------------------
- `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum,
  `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`,
  `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`,
  `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent
  default per rubber-duck #9). `ClawSandboxSpec.runtime` is required
  on the wire; `Default` retained for test ergonomics, returns
  `OpenClaw` with empty config.
- `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option
  field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid
  wiping a populated value via merge patch).
- `controller/src/crd.rs`: 8 new tests — PascalCase wire-format
  guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw,
  per-variant round-trip, BYO contractVersion required-not-default,
  serializer omits absent variants, runtimeKind status absence.
- `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips
  from `["openclaw", ...]` to `["runtime", ...]`. New `runtime`
  block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner
  CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains
  `runtimeKind`. `Runtime` printer column added.
- `controller/src/reconciler/mod.rs`: minimal call-site fix —
  `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the
  build green. Full dispatch refactor (`RuntimeDeploymentPlan` per
  rubber-duck #2/#3) lands in step 2.

What is NOT yet wired (intentional, follow-ups)
-----------------------------------------------
- Reconciler dispatch per `runtime.kind` (single-seam plan struct).
- `RuntimeReady` Condition machinery (folded into
  `build_running_status_patch` + `running_status_matches` per
  rubber-duck #1 to avoid status-merge churn).
- OpenAI Agents / MAF deployment SKIP (must NOT silently use
  `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded +
  AdapterMissing).
- `validate_runtime_shape` controller-side guard.
- Examples / fixtures / CLI templates / convert / from_kagent migration.
- CHANGELOG.md, audit doc.

Verification
------------
- `cargo test --package azureclaw-controller`: 284/284 pass
  (8 new RuntimeSpec tests included).
- `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean.
- `cargo fmt --all`: applied.
- Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime
  block, byo.required = [image, contractVersion], runtimeKind status
  field present, Runtime printer column added.

NOTE: This commit alone is NOT mergeable on its own. Without the
fixture/CLI/example migrations + reconciler dispatch, every existing
`spec.openclaw` manifest in-tree would fail admission. Branch
`phase2-multi-runtime-crd` will accumulate the remaining steps before
the PR opens.

Refs: plan.md S10.A1; rubber-duck critique applied (status merge
risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL
shape #6, contractVersion required #9, container name stays
'openclaw' #4).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures

Completes the in-place v1alpha1 schema migration for the multi-runtime CRD
spine. CRD types + Helm schema + reconciler reader landed in d11d41d; this
finishes the long tail of CRD-emitting / CRD-reading sites the user called
out ("every aspect of the code — including the cloud offload").

Cloud offload (the user's explicit ask):
- controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with
  spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate
  spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw).
- controller/src/reconciler/mod.rs:796: comment text aligned to new path.

CLI emitters:
- add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}.
- convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails
  with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape
  for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape.
- migrate.ts: --image help text references spec.runtime.openclaw.image.
- migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning
  message aligned.
- handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model.

Fixtures + examples (8 yaml files):
- examples/{basic,confidential,telegram}-agent/clawsandbox.yaml
- examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml
- tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml
  (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied
  fixture left unchanged — orthogonal to this migration; static scanner is the
  active enforcement, not CRD validation.)

Tests updated to assert the new shape:
- add.test.ts: 4 assertions
- convert.test.ts: 6 assertion blocks + 1 multi-container test
- from_kagent.test.ts: 4 assertions

Verification:
- cargo test --package azureclaw-controller: 284/284 pass
- cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean
- cli npm test: 435/435 pass + 2 skipped
- cli npm run typecheck: clean
- grep confirms no remaining spec.openclaw emission/read sites; only intentional
  docstring/comment references documenting the legacy shape.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: runtime-aware status surface + AdapterMissing dispatch guard

Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior
two commits (d11d41d CRD spine; 3202d13 emitter migration) by wiring the
runtime kind through the status surface and refusing to deploy a Pod for
runtime kinds whose adapter has not yet shipped.

Status surface
- New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in
  controller/src/status/conditions.rs.
- `build_running_status_patch` / `running_status_matches` /
  `build_overlay_status_patch` / `overlay_status_matches` take
  `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and
  append `RuntimeReady` to the conditions array (True/Reconciled on
  the running path, False/OverlayMode on overlay). Stamping inside the
  existing patch (rather than via a separate patch_status) avoids the
  merge-patch array overwrite that would erase the new Condition and
  re-introduce the resourceVersion-bump reconcile storm — see plan
  S10.A1 rubber-duck #1.
- New `build_runtime_unsupported_status_patch` /
  `runtime_unsupported_status_matches` / `stamp_runtime_unsupported`
  helper trio mirrors the existing degraded_* trio. Stamps Degraded=True
  + Ready=False + RuntimeReady=False, all Reason=AdapterMissing.

Reconciler dispatch
- controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a
  static-str discriminator and explicitly skips namespace/SA/Deployment
  creation when the kind is not OpenClaw. Stamps AdapterMissing and
  returns Action::requeue(300s) BEFORE any K8s-resource builder is
  invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1
  rubber-duck #2).
- Status-patch call sites at :1481-1498 thread the runtime_kind_str.

Tests
- 5 new tests for the AdapterMissing helper trio (stamp shape,
  status-missing/runtime-mismatch idempotency rejection, settled-status
  match, transition-time preservation across repeat patches).
- 4 existing tests updated for new conditions array shape + runtimeKind
  field (Running: 2 conds, Overlay: 4 conds).
- 289/289 controller tests pass (was 284); cargo clippy clean; CLI
  435/435 still green.

Docs
- CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking-
  change marker spanning the three commits.
- docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full
  audit doc with threat model (silent fallthrough, status churn,
  CEL-disabled, BYO contract bypass, convert hard-fail), existing-
  implementation survey, wire-format invariants, test matrix, and
  S10.A2-A5 deferral list.

Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam,
per-variant image/entrypoint/env/agentCode resolution, BYO contract
verifier, validate_runtime_shape defensive guard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic)

Locks the CRD wire shape now for three additional declared-roadmap
runtimes so adding their adapters in a later slice is not a breaking
schema change. The CRD becomes a public roadmap signal: customers can
pin `spec.runtime.kind` today and know the schema won't shift under
them.

Tiering:
  Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3),
                             MicrosoftAgentFramework (S10.A4)
  Tier 2 (placeholder, Phase 3+ adapters):
                             SemanticKernel, LangGraph, Anthropic
  BYO (warn-only contract verifier in S10.A2)

Schema changes
- crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs
  `SemanticKernelConfig` (language: python|dotnet|java),
  `LangGraphConfig` (language: python|typescript), `AnthropicConfig`
  (pythonVersion). All three carry the universal agentCode + entrypoint
  + extraEnv shape — same as OpenAIAgents/MAF.
- helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property
  blocks; kind enum extended in both spec + status surfaces; nested
  AgentCodeRef exactly-one CEL on every variant that carries code.
- reconciler: runtime_kind_str match extended; AdapterMissing message
  enumerates Tier-2 placeholders.

Behavior
- All three Tier-2 kinds short-circuit through the existing
  `stamp_runtime_unsupported` path: Degraded + Ready=False +
  RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths;
  pure schema scaffold.

Tests: 4 new round-trip tests (one per variant + a defaults check that
SkLanguage and LangGraphLanguage default to python). 293/293 controller
tests pass (was 289). Clippy clean.

Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md
updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL
matrix.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 5, 2026
…emo-clawshield (#225)

* examples: lethal-trifecta-demo — reproduces Claude Cowork attack on AKS

A reproducible launch-day demo anchored on three real, recent
agentic-AI exploits:

- Claude Cowork file-exfiltration (PromptArmor, Jan 2026)
- Google Antigravity .env exfiltration (PromptArmor, Nov 2025)
- EchoLeak / M365 Copilot (CVE-2025-32711, Jun 2025)

All three exploit the lethal trifecta (Simon Willison): private data +
untrusted content + exfil channel. The demo deploys two side-by-side
namespaces — vanilla OpenClaw with a domain-only egress allowlist vs.
a full AzureClaw stack — and shows six independent AzureClaw layers
each catching the attack alone:

  1. Inline Content Safety (Foundry DefaultV2 prompt-shield)
  2. ToolPolicy URL+method allowlist (not just domain)
  3. ClawIdentity strips attacker-controlled bearer
  4. Egress-guard (UID 1000 iptables)
  5. Token budget cap
  6. AGT BehaviorMonitor auto-quarantine
  + tamper-evident audit chain

Files:
- README.md           — threat model, citations, quick run
- WALKTHROUGH.md      — 7-min timed live/recorded script
- bait/poisoned-skill.md         — the 1pt-font injection (markdown form)
- scenarios/00-namespaces.yaml
- scenarios/01-naked-claw.yaml   — vanilla Pod, falls to attack
- scenarios/02-azureclaw-sandbox.yaml — full ClawSandbox CRD
- scenarios/03-bait-server.yaml
- scripts/{deploy,run-attack,verify-defense,teardown}.sh

Leaves examples/demo-clawshield in place for now — that demo covers
multi-tenancy / Kata isolation, which is a different story.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(examples): real READMEs for basic-agent / confidential-agent / demo-clawshield

Three top-level entries in examples/README.md were either missing a
README entirely (basic-agent, confidential-agent) or shipping a 14-line
shell-comment stub with promised section headings and zero content
(demo-clawshield). The stub was discoverable from the GitHub deep-link
`#3-networkpolicy-default-deny-egress` and bounced to nothing.

This adds proper READMEs:

- examples/basic-agent/README.md — what it ships, default posture table,
  deploy + customize + cleanup, links to confidential-agent and
  lethal-trifecta-demo for variants

- examples/confidential-agent/README.md — explicitly documents how it
  differs from basic-agent (single `isolation: confidential` field),
  Kata add-on prereq, runtimeClassName verification one-liner, links to
  blueprints/02-enterprise-self-hosted

- examples/demo-clawshield/README.md — full content replacing the stub:
  what each YAML does, layer-per-phase mapping table, the
  this-vs-lethal-trifecta-demo orientation paragraph, pointer to
  docs/internal/DEMO.md for the 30-min timed walkthrough

Cross-links between the four attack/security examples (basic-agent,
confidential-agent, demo-clawshield, lethal-trifecta-demo) so users
can navigate between them.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 5, 2026
…hments, sibling trust race, final-deliverable rule

Live multi-agent demo (analyst→viz→writer fan-out) surfaced four
distinct breakages that each silently degraded the run while the
parent agent still self-reported success.

1. Image generation 404 (router URL prefix regression)

   Foundry's account-scoped /openai/v1/images/generations endpoint
   does NOT accept the /api/projects/<project>/ URL prefix that
   chat-completions tolerates. Commit c13f302 unified everything
   through that prefix. In dev (raw Azure OpenAI account, no project
   path) it works; in AKS prod against a Foundry project endpoint
   the upstream returns a fast 404 and image_generation falls back
   to written descriptions.

   Fix: strip /api/projects/<name>/ from the upstream endpoint
   inside the images_generations handler before forwarding. Add
   strip_project_prefix helper + four unit tests (project-stripped,
   trailing-slash variant, AOAI passthrough, no-prefix passthrough).

2. foundry_code_execute drops container files

   The Responses API tool only walked output[].type=='message' for
   text. matplotlib PNGs / CSVs generated by code_interpreter live
   inside Foundry's per-run container and are referenced via
   container_file_citation annotations or { type: 'image' } entries
   in code_interpreter_call.outputs. They were never downloaded, so
   the demo's bar chart silently degraded to ASCII.

   Fix: collect every (container_id, file_id, filename) reference
   from both shapes, GET each via the new /openai/containers/...
   router route (added to foundry_standalone_routes), and write the
   bytes to /sandbox/.openclaw/workspace/. Append the local paths
   to the tool result so downstream tools (mesh_transfer_file,
   file_write) can ship them. Adds routerCallBinary helper for
   binary downloads through the router.

3. Sibling KNOCK race in parallel fan-out

   AGT_TRUSTED_PEERS is baked at spawn time and consumed once at
   sub-agent boot. When the parent spawns analyst → viz → writer in
   sequence, only writer (last) sees all siblings. analyst's
   parentTrustedAmids only contains parent — so when viz or writer
   later try to KNOCK analyst, the trust score is 0 + 0 = 0 and
   the KNOCK is rejected at threshold 500. The demo logs confirm
   only 1 of 3 sibling pairs ever opened a session.

   Fix: after every successful spawn, the parent broadcasts a
   peers_update message containing the new sibling's AMID to every
   already-running sibling. Each sub-agent now records the parent's
   AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and
   handles peers_update only from that AMID, extending its
   parentTrustedAmids set at runtime.

4. Sub-agent system prompt missing FINAL DELIVERABLE rule

   Sub-agents were told they could mesh_transfer_file artifacts to
   peers, but nothing forced them to mesh_transfer_file the FINAL
   artifact back to the parent before returning a summary. The
   writer's executive_brief.md sat in its local /sandbox forever
   while the parent reported success.

   Fix: append a hard rule to the sub-agent system prompt requiring
   mesh_transfer_file(to_agent='parent', ...) as the last action
   before any "task complete" reply, with one call per output file.

Tests
- inference-router: 643 lib tests pass (4 new strip_project_prefix
  tests).
- inference-router: cargo clippy --all-targets clean.
- runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only
  pre-existing warnings.
- cli: 553 tests pass.

Deployment
- For #2: rebuild + push inference-router image.
- For #1, #3, #4: rebuild + push sandbox image.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
Per docs/implementation-plan.md §5.4. Behavioral conformance corpus —
the net that catches 'endpoint returned 200 but skipped the crypto
step' bugs.

tests/conformance/ layout
- package.json / tsconfig.json / vitest.config.ts — own workspace,
  fork-isolated pool so ratchet state doesn't leak across specs.
- README.md — corpus table (9 corpora across Phases 0, 1, 3),
  provider-axis rule (plan §11.3), invariant discipline.
- fixtures/README.md — vendored-source policy per principle §0.2 #10.
- harness/index.ts — empty scaffold; helpers land per-PR.

Phase 0 specs (all 'it.todo', 59 invariants total — vitest reports
59 todo / 0 pass / 0 fail, no false-green):
- signal-x3dh.spec.ts (14) — key-exchange shape, symmetric ratchet
  invariants, base64 input hygiene (vendor patches #3, #4).
- signal-knock.spec.ts (15) — KNOCK happy-path, trust threshold, relay
  disruption, wire-shape parity (vendor patches #5, #7, #8, #1/#2
  timestamps, vendored<->AGT byte-identical KNOCK).
- signal-negative.spec.ts (13) — ciphertext integrity, replay, session
  clobber (vendor patch #10), DoS surfaces.
- sandbox-isolation.spec.ts (17) — seccomp, Landlock, egress-guard,
  router-as-only-network-path. Guarded by CONFORMANCE_E2E=1 (requires
  Kind; compat suite Kind harness wires it in Phase 1).

Principle §0.2 #8 ('solid, not look-alike'): it.todo is a documented
pending assertion, not a silently-passing no-op. Each spec's top
comment cites the vendor patch or production bug it exists to prevent
recurring.

Principle §0.2 #10: every invariant that references an external
protocol cites its upstream source (libsignal, RFC3339 'Z' suffix,
Signal Double Ratchet spec) either inline or in the README.

ci/no-stubs.sh already allow-lists tests/ subtrees; gate remains PASS.
No new dependencies outside the pinned vitest / typescript / @types
dev-deps already used by tests/compat/.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
… validator

Implements plan §1.3 (Outage semantics) as a pure, deterministic decision
function. No I/O, no AGT import surface, clock injected.

inference-router/src/providers/outage.rs (new):
 - OutageMode enum: Strict | CachedRead | DegradedDev
   * serde camelCase; FromStr accepts camel/kebab/snake; Display; Default = Strict
 - OutageConfig { mode, cached_ttl }
   * validate_for_env(is_dev_env) rejects DegradedDev in prod
   * rejects cached_ttl = 0 on CachedRead
   * enforces MAX_CACHED_TTL = 15 min
 - CachedDecision<T> with is_expired(ttl, now) — backwards-clock = expired
 - OutageAction<T>: Deny { mode } | ServeCached { verdict } | AllowWithWarning
 - decide_outage(config, cached, now) — pure, test-friendly
 - 19 unit tests covering all three modes, cache freshness, clock skew,
   serde round-trip, env validation, TTL bounds

inference-router/src/providers/mod.rs:
 - remove placeholder OutageMode stub; re-export real types

controller/src/providers/mod.rs:
 - mirrored OutageMode (from_spec, is_dev_only, validate_for_env)
 - OutageModeError::DegradedDevInProd
 - 4 new unit tests

docs/security-audits/2026-04-24-phase1-outage-semantics.md:
 - STRIDE, principle-mapping, re-audit triggers; both sign-offs present.

No call-site in the router consumes decide_outage yet — that lands with
the first AGT provider. Landing the pure semantics first locks the
decision rules before any provider wiring pressures them.

Verification:
 - cargo test --all: 106+155+15+26+3 = 305 passed (was 286, +19)
 - cargo clippy --all-targets --all-features -- -D warnings: clean
 - six CI gates PASS on the branch tip

Plan refs: §0.2 #1/#2/#3/#4/#5/#8/#9/#10 | §1.3 | §1.4

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
)

* phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only)

S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting.

Replaces the legacy `spec.openclaw` field with a discriminated union
`spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`.
The `kind` discriminator selects which sibling struct is required;
the others must be absent. Mutual exclusion enforced at admission via
Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules
`(self.kind == 'X') == has(self.x)`); controller-side defensive guard
will land alongside the reconciler dispatch in a follow-up commit.

Pre-release simplification: in-place v1alpha1 schema edit. No
v1alpha2 cut, no conversion webhook (no installed base, per plan.md
S10 + S13). One PR = one breaking change.

What this commit delivers
-------------------------
- `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum,
  `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`,
  `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`,
  `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent
  default per rubber-duck #9). `ClawSandboxSpec.runtime` is required
  on the wire; `Default` retained for test ergonomics, returns
  `OpenClaw` with empty config.
- `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option
  field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid
  wiping a populated value via merge patch).
- `controller/src/crd.rs`: 8 new tests — PascalCase wire-format
  guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw,
  per-variant round-trip, BYO contractVersion required-not-default,
  serializer omits absent variants, runtimeKind status absence.
- `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips
  from `["openclaw", ...]` to `["runtime", ...]`. New `runtime`
  block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner
  CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains
  `runtimeKind`. `Runtime` printer column added.
- `controller/src/reconciler/mod.rs`: minimal call-site fix —
  `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the
  build green. Full dispatch refactor (`RuntimeDeploymentPlan` per
  rubber-duck #2/#3) lands in step 2.

What is NOT yet wired (intentional, follow-ups)
-----------------------------------------------
- Reconciler dispatch per `runtime.kind` (single-seam plan struct).
- `RuntimeReady` Condition machinery (folded into
  `build_running_status_patch` + `running_status_matches` per
  rubber-duck #1 to avoid status-merge churn).
- OpenAI Agents / MAF deployment SKIP (must NOT silently use
  `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded +
  AdapterMissing).
- `validate_runtime_shape` controller-side guard.
- Examples / fixtures / CLI templates / convert / from_kagent migration.
- CHANGELOG.md, audit doc.

Verification
------------
- `cargo test --package azureclaw-controller`: 284/284 pass
  (8 new RuntimeSpec tests included).
- `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean.
- `cargo fmt --all`: applied.
- Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime
  block, byo.required = [image, contractVersion], runtimeKind status
  field present, Runtime printer column added.

NOTE: This commit alone is NOT mergeable on its own. Without the
fixture/CLI/example migrations + reconciler dispatch, every existing
`spec.openclaw` manifest in-tree would fail admission. Branch
`phase2-multi-runtime-crd` will accumulate the remaining steps before
the PR opens.

Refs: plan.md S10.A1; rubber-duck critique applied (status merge
risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL
shape #6, contractVersion required #9, container name stays
'openclaw' #4).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures

Completes the in-place v1alpha1 schema migration for the multi-runtime CRD
spine. CRD types + Helm schema + reconciler reader landed in 98c629d; this
finishes the long tail of CRD-emitting / CRD-reading sites the user called
out ("every aspect of the code — including the cloud offload").

Cloud offload (the user's explicit ask):
- controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with
  spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate
  spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw).
- controller/src/reconciler/mod.rs:796: comment text aligned to new path.

CLI emitters:
- add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}.
- convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails
  with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape
  for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape.
- migrate.ts: --image help text references spec.runtime.openclaw.image.
- migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning
  message aligned.
- handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model.

Fixtures + examples (8 yaml files):
- examples/{basic,confidential,telegram}-agent/clawsandbox.yaml
- examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml
- tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml
  (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied
  fixture left unchanged — orthogonal to this migration; static scanner is the
  active enforcement, not CRD validation.)

Tests updated to assert the new shape:
- add.test.ts: 4 assertions
- convert.test.ts: 6 assertion blocks + 1 multi-container test
- from_kagent.test.ts: 4 assertions

Verification:
- cargo test --package azureclaw-controller: 284/284 pass
- cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean
- cli npm test: 435/435 pass + 2 skipped
- cli npm run typecheck: clean
- grep confirms no remaining spec.openclaw emission/read sites; only intentional
  docstring/comment references documenting the legacy shape.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: runtime-aware status surface + AdapterMissing dispatch guard

Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior
two commits (98c629d CRD spine; 56ad076 emitter migration) by wiring the
runtime kind through the status surface and refusing to deploy a Pod for
runtime kinds whose adapter has not yet shipped.

Status surface
- New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in
  controller/src/status/conditions.rs.
- `build_running_status_patch` / `running_status_matches` /
  `build_overlay_status_patch` / `overlay_status_matches` take
  `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and
  append `RuntimeReady` to the conditions array (True/Reconciled on
  the running path, False/OverlayMode on overlay). Stamping inside the
  existing patch (rather than via a separate patch_status) avoids the
  merge-patch array overwrite that would erase the new Condition and
  re-introduce the resourceVersion-bump reconcile storm — see plan
  S10.A1 rubber-duck #1.
- New `build_runtime_unsupported_status_patch` /
  `runtime_unsupported_status_matches` / `stamp_runtime_unsupported`
  helper trio mirrors the existing degraded_* trio. Stamps Degraded=True
  + Ready=False + RuntimeReady=False, all Reason=AdapterMissing.

Reconciler dispatch
- controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a
  static-str discriminator and explicitly skips namespace/SA/Deployment
  creation when the kind is not OpenClaw. Stamps AdapterMissing and
  returns Action::requeue(300s) BEFORE any K8s-resource builder is
  invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1
  rubber-duck #2).
- Status-patch call sites at :1481-1498 thread the runtime_kind_str.

Tests
- 5 new tests for the AdapterMissing helper trio (stamp shape,
  status-missing/runtime-mismatch idempotency rejection, settled-status
  match, transition-time preservation across repeat patches).
- 4 existing tests updated for new conditions array shape + runtimeKind
  field (Running: 2 conds, Overlay: 4 conds).
- 289/289 controller tests pass (was 284); cargo clippy clean; CLI
  435/435 still green.

Docs
- CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking-
  change marker spanning the three commits.
- docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full
  audit doc with threat model (silent fallthrough, status churn,
  CEL-disabled, BYO contract bypass, convert hard-fail), existing-
  implementation survey, wire-format invariants, test matrix, and
  S10.A2-A5 deferral list.

Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam,
per-variant image/entrypoint/env/agentCode resolution, BYO contract
verifier, validate_runtime_shape defensive guard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic)

Locks the CRD wire shape now for three additional declared-roadmap
runtimes so adding their adapters in a later slice is not a breaking
schema change. The CRD becomes a public roadmap signal: customers can
pin `spec.runtime.kind` today and know the schema won't shift under
them.

Tiering:
  Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3),
                             MicrosoftAgentFramework (S10.A4)
  Tier 2 (placeholder, Phase 3+ adapters):
                             SemanticKernel, LangGraph, Anthropic
  BYO (warn-only contract verifier in S10.A2)

Schema changes
- crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs
  `SemanticKernelConfig` (language: python|dotnet|java),
  `LangGraphConfig` (language: python|typescript), `AnthropicConfig`
  (pythonVersion). All three carry the universal agentCode + entrypoint
  + extraEnv shape — same as OpenAIAgents/MAF.
- helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property
  blocks; kind enum extended in both spec + status surfaces; nested
  AgentCodeRef exactly-one CEL on every variant that carries code.
- reconciler: runtime_kind_str match extended; AdapterMissing message
  enumerates Tier-2 placeholders.

Behavior
- All three Tier-2 kinds short-circuit through the existing
  `stamp_runtime_unsupported` path: Degraded + Ready=False +
  RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths;
  pure schema scaffold.

Tests: 4 new round-trip tests (one per variant + a defaults check that
SkLanguage and LangGraphLanguage default to python). 293/293 controller
tests pass (was 289). Clippy clean.

Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md
updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL
matrix.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
…emo-clawshield (#225)

* examples: lethal-trifecta-demo — reproduces Claude Cowork attack on AKS

A reproducible launch-day demo anchored on three real, recent
agentic-AI exploits:

- Claude Cowork file-exfiltration (PromptArmor, Jan 2026)
- Google Antigravity .env exfiltration (PromptArmor, Nov 2025)
- EchoLeak / M365 Copilot (CVE-2025-32711, Jun 2025)

All three exploit the lethal trifecta (Simon Willison): private data +
untrusted content + exfil channel. The demo deploys two side-by-side
namespaces — vanilla OpenClaw with a domain-only egress allowlist vs.
a full AzureClaw stack — and shows six independent AzureClaw layers
each catching the attack alone:

  1. Inline Content Safety (Foundry DefaultV2 prompt-shield)
  2. ToolPolicy URL+method allowlist (not just domain)
  3. ClawIdentity strips attacker-controlled bearer
  4. Egress-guard (UID 1000 iptables)
  5. Token budget cap
  6. AGT BehaviorMonitor auto-quarantine
  + tamper-evident audit chain

Files:
- README.md           — threat model, citations, quick run
- WALKTHROUGH.md      — 7-min timed live/recorded script
- bait/poisoned-skill.md         — the 1pt-font injection (markdown form)
- scenarios/00-namespaces.yaml
- scenarios/01-naked-claw.yaml   — vanilla Pod, falls to attack
- scenarios/02-azureclaw-sandbox.yaml — full ClawSandbox CRD
- scenarios/03-bait-server.yaml
- scripts/{deploy,run-attack,verify-defense,teardown}.sh

Leaves examples/demo-clawshield in place for now — that demo covers
multi-tenancy / Kata isolation, which is a different story.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(examples): real READMEs for basic-agent / confidential-agent / demo-clawshield

Three top-level entries in examples/README.md were either missing a
README entirely (basic-agent, confidential-agent) or shipping a 14-line
shell-comment stub with promised section headings and zero content
(demo-clawshield). The stub was discoverable from the GitHub deep-link
`#3-networkpolicy-default-deny-egress` and bounced to nothing.

This adds proper READMEs:

- examples/basic-agent/README.md — what it ships, default posture table,
  deploy + customize + cleanup, links to confidential-agent and
  lethal-trifecta-demo for variants

- examples/confidential-agent/README.md — explicitly documents how it
  differs from basic-agent (single `isolation: confidential` field),
  Kata add-on prereq, runtimeClassName verification one-liner, links to
  blueprints/02-enterprise-self-hosted

- examples/demo-clawshield/README.md — full content replacing the stub:
  what each YAML does, layer-per-phase mapping table, the
  this-vs-lethal-trifecta-demo orientation paragraph, pointer to
  docs/internal/DEMO.md for the 30-min timed walkthrough

Cross-links between the four attack/security examples (basic-agent,
confidential-agent, demo-clawshield, lethal-trifecta-demo) so users
can navigate between them.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
…hments, sibling trust race, final-deliverable rule

Live multi-agent demo (analyst→viz→writer fan-out) surfaced four
distinct breakages that each silently degraded the run while the
parent agent still self-reported success.

1. Image generation 404 (router URL prefix regression)

   Foundry's account-scoped /openai/v1/images/generations endpoint
   does NOT accept the /api/projects/<project>/ URL prefix that
   chat-completions tolerates. Commit 582eb39 unified everything
   through that prefix. In dev (raw Azure OpenAI account, no project
   path) it works; in AKS prod against a Foundry project endpoint
   the upstream returns a fast 404 and image_generation falls back
   to written descriptions.

   Fix: strip /api/projects/<name>/ from the upstream endpoint
   inside the images_generations handler before forwarding. Add
   strip_project_prefix helper + four unit tests (project-stripped,
   trailing-slash variant, AOAI passthrough, no-prefix passthrough).

2. foundry_code_execute drops container files

   The Responses API tool only walked output[].type=='message' for
   text. matplotlib PNGs / CSVs generated by code_interpreter live
   inside Foundry's per-run container and are referenced via
   container_file_citation annotations or { type: 'image' } entries
   in code_interpreter_call.outputs. They were never downloaded, so
   the demo's bar chart silently degraded to ASCII.

   Fix: collect every (container_id, file_id, filename) reference
   from both shapes, GET each via the new /openai/containers/...
   router route (added to foundry_standalone_routes), and write the
   bytes to /sandbox/.openclaw/workspace/. Append the local paths
   to the tool result so downstream tools (mesh_transfer_file,
   file_write) can ship them. Adds routerCallBinary helper for
   binary downloads through the router.

3. Sibling KNOCK race in parallel fan-out

   AGT_TRUSTED_PEERS is baked at spawn time and consumed once at
   sub-agent boot. When the parent spawns analyst → viz → writer in
   sequence, only writer (last) sees all siblings. analyst's
   parentTrustedAmids only contains parent — so when viz or writer
   later try to KNOCK analyst, the trust score is 0 + 0 = 0 and
   the KNOCK is rejected at threshold 500. The demo logs confirm
   only 1 of 3 sibling pairs ever opened a session.

   Fix: after every successful spawn, the parent broadcasts a
   peers_update message containing the new sibling's AMID to every
   already-running sibling. Each sub-agent now records the parent's
   AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and
   handles peers_update only from that AMID, extending its
   parentTrustedAmids set at runtime.

4. Sub-agent system prompt missing FINAL DELIVERABLE rule

   Sub-agents were told they could mesh_transfer_file artifacts to
   peers, but nothing forced them to mesh_transfer_file the FINAL
   artifact back to the parent before returning a summary. The
   writer's executive_brief.md sat in its local /sandbox forever
   while the parent reported success.

   Fix: append a hard rule to the sub-agent system prompt requiring
   mesh_transfer_file(to_agent='parent', ...) as the last action
   before any "task complete" reply, with one call per output file.

Tests
- inference-router: 643 lib tests pass (4 new strip_project_prefix
  tests).
- inference-router: cargo clippy --all-targets clean.
- runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only
  pre-existing warnings.
- cli: 553 tests pass.

Deployment
- For #2: rebuild + push inference-router image.
- For #1, #3, #4: rebuild + push sandbox image.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
) (#292)

* Slice 4d.1 — mcpServerRefs plural CRD field + admission CEL

Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds
GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the
existing singular mcpServerRef, which is now deprecated and honored
as a length-1 alias via effective_mcp_server_refs().

Controller-side changes:
- New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET
  in status/conditions.rs::reason.
- Reconciler mirror loop refactored to iterate effective_mcp_server_refs().
  Singular-field use emits tracing::warn with McpSingularDeprecated.
  len > 1 short-circuits via degrade! macro until Slice 4d.2 wires
  per-server addressing — principles §3 honest 'not-yet-enforced' signal.
- 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty
  + plural-wins-when-both-set. Controller suite: 555 passing.

Admission CEL on deploy/helm/azureclaw/templates/crd.yaml:
- Mutex: singular and plural cannot both be set.
- maxItems: 8 (router-side scheme is sized for this).
- Per-name uniqueness across mcpServerRefs.

Out of scope for 4d.1 (queued for 4d.2):
- Per-server jwks-{name}.json / tools-{name}.json file scheme.
- Router-side McpServerRegistry + namespaced tool dispatch.
- Stale-file sweep (DoD #6).
- e2e fixture with ≥ 3 servers (DoD #1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Slice 4d.2 — per-server McpServer mounts + router discovery (DoD #1 + #6)

Closes Slice 4 DoD #1 (≥3 plural McpServers reachable e2e) at the
mount-and-discovery layer, and DoD #6 (stale-file sweep) via
reconciler-driven volume rebuild. Multi-JWKS OAuth + namespaced tool
dispatch (DoD #3) follow in Slice 4d.3.

Controller:
- reconciler/mod.rs: replace len>1 short-circuit with full iteration
  over effective_mcp_server_refs(). Per-name volumes (mcp-jwks-<name>,
  mcp-signing-<name>) mounted at /etc/azureclaw/mcp/<name>/ and
  /etc/azureclaw/mcp-signing/<name>/.
- First entry (idx 0) keeps legacy MCP_JWKS_PATH + MCP_SIGNING_KEY_DIR
  env vars for backwards compat with current single-JWKS OAuth path.
- New MCP_JWKS_DIR=/etc/azureclaw/mcp env set once per pod.
- governance_mounts.rs: new inject_container_env helper for idempotent
  env-var injection without a volume mount.
- Removed obsolete PLURAL_MCP_SERVERS_UNSUPPORTED_YET reason constant.
- CRD doc comment updated for 4d.2 mount layout + 4d.3 forward-ref.

Router:
- New mcp/registry.rs with scan() + discover_from_env(). At startup,
  reads MCP_JWKS_DIR, enumerates subdirs with parseable jwks.json,
  emits tracing::info!(servers=?, count=N) per discovery + warn!
  per skipped candidate.
- Empty/missing dir handled gracefully (sandbox with zero
  mcpServerRefs is a valid steady state).
- 7 registry unit tests + 4 inject_container_env unit tests.

Stale-file sweep (DoD #6): reconcile rebuilds pod-spec from current
refs; removed refs disappear via SSA. No explicit sweep code needed.

Verified: 559 controller + 804 router tests pass, clippy -D warnings
clean across workspace, cargo fmt --check clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
 — OAuth half) (#293)

Closes the OAuth half of Slice 4 DoD #3. Each McpServer now has its
tokens validated against its own JWKS with a per-issuer audience and
scopes override — all driven by a meta.json file the controller
writes alongside jwks.json in the per-server ConfigMap.

Controller-side producer:
- mcp_server_reconciler.rs: ensure_jwks_configmap now writes a second
  key, meta.json, alongside jwks.json. Wire shape (camelCase):
  { issuer, audience?, scopes[] }. Optional/empty fields skipped for
  forward-compat.
- New McpServerMeta struct + McpServerMeta::from_spec helper.

Router-side consumer:
- DiscoveredMcpServerMeta in mcp::registry, loaded adjacent to
  jwks.json during scan(). Missing meta is silently back-compat;
  malformed or empty-issuer meta is recorded in skipped (principles
  §3 — no silent failure).
- OAuthVerifierConfig gains per_issuer_audience + per_issuer_scopes.
  verify_access_token now uses the per-issuer entry keyed off the
  token's iss claim, falling back to the global expected_audience /
  required_scopes when none is set. iss is already pinned per-issuer
  via JWKS lookup, so this lifts the audience/scope checks onto the
  same secure key.
- New OAuthVerifierConfig::from_registry constructor: returns
  Ok(None) when no servers carry meta (caller falls back to legacy
  MCP_JWKS_PATH); hard-errors on conflicting JWKS for the same issuer.
- main.rs build_mcp_router: prefers the registry-first multi-issuer
  path when MCP_PRODUCTION_MODE=true and the registry is non-empty;
  legacy single-JWKS path remains as fallback.

Tests:
- 4 new registry tests (meta load happy, empty-issuer skip, malformed
  parse skip, no-meta back-compat).
- 6 new oauth tests (per-issuer audience accept across two servers,
  per-issuer audience cross-token rejection, per-issuer scopes apply
  per token, from_registry None on empty, from_registry happy with
  two issuers, from_registry conflict rejection).
- 820 router lib tests pass (was 810), 559 controller bin tests pass.
- cargo clippy --all-targets -- -D warnings clean.
- cargo fmt --check clean.

Out of scope (4d.4 will land):
- Namespaced tool dispatch (/mcp/<server>/...) — waiting on a real
  upstream MCP forwarder backend so dispatch isn't scaffolding.

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
Closes Slice 4 DoD #3 dispatch half: /mcp now forwards tool calls
to upstream McpServer.spec.url instead of serving only the in-tree
EchoDispatcher.

Producer (controller):
- McpServerMeta gains url + allowed_tools fields, written into
  meta.json alongside the OAuth metadata from Slice 4d.3.

Consumer (router):
- New mcp::forwarder module with RouterToolDispatcher implementing
  AsyncToolDispatcher.
- Startup discovery POSTs tools/list to each upstream, filters
  through allowed_tools ("*" exposes all; explicit list selects
  named subset; empty fails closed), prefixes each tool with
  snake_case(server_name).
- Servers with empty URL, empty allow-list, outbound OAuth
  requirement (deferred to 4d.5), unreachable upstream, or HTTP
  errors are recorded in skipped[] with operator-actionable reason
  and excluded from the catalog. Router still starts.
- Dispatch splits name on first '.', looks up server, forwards
  tools/call to upstream URL. JSON-RPC errors collapse to
  is_error:true; 5xx surfaces as DispatchError::ExecutionFailed;
  bad names surface as DispatchError::UnknownTool.
- McpRouteState gains with_tools() builder; build_mcp_router (now
  async) swaps the dispatcher when registry is non-empty.

Scope:
- Unauthenticated upstreams only. Outbound OAuth is Slice 4d.5
  (requires per-server credential source). Servers requiring
  bearer credentials are deliberately skipped per §5 (no
  scaffolding — only ship the consumer we can drive end-to-end).
- Single tools/list page (multi-page support deferred until a
  real upstream hits the cap).

Test coverage: 15 new unit tests covering every skip reason, both
allow-list semantics, namespaced dispatch (single + multi server),
upstream JSON-RPC errors, HTTP 5xx, and unknown-tool branches.

Backward compatibility: registries with no meta.json or no url
fall through to the existing EchoDispatcher unchanged.

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
Replaces ClawSandbox.spec.networkPolicy.learnEgress: bool with
egressMode: Strict | Learn across controller, router, CLI, Headlamp,
helm CRD, and docs. Default = Learn. The Approval variant is deferred
to Slice 5c (no consumer yet per principles §5).

No back-compat shim: repo has no live deployments yet, so the legacy
field was removed cleanly. Slice 5b DoD #3 closed.

- controller/src/crd.rs: new EgressMode enum with #[derive(Default)]
  and #[default] on Learn. Replaces NetworkPolicyConfig.learn_egress.
- controller/src/reconciler/mod.rs: emits EGRESS_MODE=strict|learn
  (no back-compat EGRESS_LEARN_MODE).
- inference-router/src/routes/mod.rs: consumes EGRESS_MODE, defaults
  to Learn when unset.
- inference-router/src/spawn/mod.rs: sub-agent reader/producer
  migrated. Internal SpawnRequest.learn_egress retained (not wire).
- deploy/helm/azureclaw/templates/crd.yaml: learnEgress removed
  from schema; egressMode enum-validated.
- cli/src/commands/{add,policy,handoff,handoff/helpers,up/sandbox_bringup,
  dev/local-k8s,operator/actions}.ts: every CR producer/reader migrated.
- tools/headlamp-plugin/src/index.tsx: all 4 egress readers consume
  egressMode directly.
- docs/use-cases.md, docs/egress-proxy.md: samples updated.

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 16, 2026
…atency overclaims

Second pass of the OSS-readiness audit, focused on the architecture
diagrams (the first pass missed factual errors here).

architecture-diagrams.md:
- Diagram #3: audit record is hash-chained, append-only (NOT signed today —
  cryptographic signing of the chain head is on the roadmap, see
  security.md). Tamper-detection vs tamper-proof is a real distinction.
- Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list),
  prose says nine + mentions ClawPairing as the controller-internal 10th.
- Diagram #7: CRD relationship arrows were wrong against the actual Rust
  structs. Corrected:
    - policyRef → spec.governance.toolPolicyRef
    - mcpRefs → spec.governance.mcpServerRefs
    - inferenceRef → spec.inferenceRef (top-level, was already correct)
    - memoryRef → spec.memoryRef (top-level, was already correct)
    - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real
      link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy.
    - CE 'sandboxRef' → spec.targetSandboxRef.
    - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected
      to every sandbox by the controller (no ref). Noted in prose.
    - EgressApproval added (it was missing from the diagram entirely);
      links to ClawSandbox via spec.sandbox (string name, not ref object).

security.md:
- Headline #1 'agent does not see Azure credentials. Period.' — softened
  with a dev-mode footnote matching the README hero. In azureclaw dev,
  agent+router share a container with separate UIDs but a kernel-level
  container escape defeats the boundary; the hard guarantee is the AKS
  path. Anchor points at architecture.md#two-modes.
- Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation
  latency on the router hot path.' Microsecond was an overclaim without
  a benchmark to back it.

architecture.md:
- Controller row in the components table: 'watches the eight peer CRDs'
  → 'nine peer CRDs (plus controller-internal ClawPairing)'.

All 9 mermaid blocks pass a bracket-balance sanity check.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 16, 2026
…#325)

* docs: OSS-readiness pass — strip internal jargon + correct overclaims

Doc + comment-only audit across user-facing docs ahead of OSS launch.
No code paths touched.

Factual corrections (the heaviest changes):

- docs/security.md headline guarantee #4: drop 'signed by the router'
  claim. The audit log is hash-chained and detects modification but
  is not signed today; signing is on the v1.1 roadmap.
- docs/security.md Layer 6 inference-safety table:
  * Content Safety: 'Always on, server-side' → 'Always on for
    Foundry-provider requests; Copilot/GitHub-Models providers do
    not return prompt_filter_results'.
  * Token budget: was 'not yet aggregated'; the router DOES aggregate
    per-tenant daily and monthly UTC counters with on-disk persistence
    (inference-router/src/budget.rs:200-260).
  * New 'Operator escape hatches' subsection documents the two env-var
    knobs (AZURECLAW_SUPPRESS_CONTENT_FLAGS,
    AZURECLAW_CONTENT_FLAG_MIN_SEVERITY) honestly.
- README hero: 'agent never sees an Azure key' softened to call out
  that this is the AKS guarantee; dev mode co-locates agent + router
  in one container.
- README/architecture: '31 commands' → '30+ commands'; remove
  unverified '18 Foundry API groups' count.
- docs/architecture.md design goal #1 + #4: explicit dev-vs-prod
  scoping; 'same code path' → 'same data-path code with documented
  AZURECLAW_DEV_MODE branches'.
- docs/architecture/a2a-gateway.md: port 8445 is config-locked but the
  mTLS listener itself is still being wired; operators should set
  A2A_GATEWAY_UPSTREAM_URL explicitly until the listener is GA.
- mesh-plugin/src/agt-identity.ts comment: replace 'encrypted at rest
  with per-host KEK' claim with honest 'chmod 0600 is the real
  boundary' note (mirrors PR #324 identity-store fix).
- .github/copilot-instructions.md: '__AGT_INITIALIZED env guard' →
  'Symbol.for(agt-mesh-client)' (matches current code in
  runtimes/openclaw/src/index.ts:458-480).

Internal-jargon strip in user-facing docs:

- docs/api/lifecycle.md: remove 'Slice 4', 'Slice 4d.3/4d.4',
  'Slice 0', 'Slice 1c', 'Slice 2a/2b/2c/2d.1', 'Slice 2d.2',
  'Slice 3a' references — replaced with descriptive prose.
- docs/api/conditions.md: drop 'Slice 1c invariant' and 'principles.md
  §3' references.
- docs/architecture/agt-boundary.md, docs/security-mcp-top10.md,
  docs/cli-reference.md: drop 'Phase 5.2' shibboleth — keep the
  facts (vendored AgentMesh fork was retired upstream).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: deep-dive diagram audit — fix CRD field-name labels + signing/latency overclaims

Second pass of the OSS-readiness audit, focused on the architecture
diagrams (the first pass missed factual errors here).

architecture-diagrams.md:
- Diagram #3: audit record is hash-chained, append-only (NOT signed today —
  cryptographic signing of the chain head is on the roadmap, see
  security.md). Tamper-detection vs tamper-proof is a real distinction.
- Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list),
  prose says nine + mentions ClawPairing as the controller-internal 10th.
- Diagram #7: CRD relationship arrows were wrong against the actual Rust
  structs. Corrected:
    - policyRef → spec.governance.toolPolicyRef
    - mcpRefs → spec.governance.mcpServerRefs
    - inferenceRef → spec.inferenceRef (top-level, was already correct)
    - memoryRef → spec.memoryRef (top-level, was already correct)
    - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real
      link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy.
    - CE 'sandboxRef' → spec.targetSandboxRef.
    - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected
      to every sandbox by the controller (no ref). Noted in prose.
    - EgressApproval added (it was missing from the diagram entirely);
      links to ClawSandbox via spec.sandbox (string name, not ref object).

security.md:
- Headline #1 'agent does not see Azure credentials. Period.' — softened
  with a dev-mode footnote matching the README hero. In azureclaw dev,
  agent+router share a container with separate UIDs but a kernel-level
  container escape defeats the boundary; the hard guarantee is the AKS
  path. Anchor points at architecture.md#two-modes.
- Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation
  latency on the router hot path.' Microsecond was an overclaim without
  a benchmark to back it.

architecture.md:
- Controller row in the components table: 'watches the eight peer CRDs'
  → 'nine peer CRDs (plus controller-internal ClawPairing)'.

All 9 mermaid blocks pass a bracket-balance sanity check.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 16, 2026
…diagrams (#327)

Link each provider bullet in 'Pluggable inference backend' to the
relevant architecture chapter and diagram:

- GitHub Copilot → architecture.md#dev-mode + diagram #1 (dev mode pod)
- Foundry / Azure OpenAI → architecture.md#prod-mode + diagrams #2 (prod
  mode pod) and #3 (data path)
- GitHub Models → security.md (Content Safety caveat) + architecture.md
  data-path section (provider routing notes)

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 28, 2026
Closes the rubber-duck research findings from the Phase 3 critique
and Microsoft's Entra Agent ID design-patterns audit.

CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM):
- KarsSandbox.spec.meshAuth.customSecurityAttributes:
  BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which
  attributes (from a tenant-declared set) the controller should
  PATCH onto each per-sandbox agent identity.
- AgentIdentityClient::patch_custom_security_attributes Graph
  client method. Constructs the documented CustomSecurityAttribute-
  Value envelope with the required @odata.type per attribute,
  inferred from the JSON value shape.
- odata_type_for_value helper: maps serde_json::Value →
  '#String' | '#Int32' | '#Boolean' | '#Collection($T)' and
  rejects floats, nulls, nested objects, mixed-type arrays, and
  empty arrays with clear error messages BEFORE the call goes out.
- Wired into ensure_agent_identity_for_sandbox: PATCH runs on
  every reconcile (idempotent on Graph). Failures surface as
  ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing
  silent missing-attribute drift.

SCALE-OUT INVARIANT (rubber-duck #4):
- Documented in agent_id_provisioning.rs module doc: the agent
  identity is keyed on KarsSandbox.metadata.uid (and cluster UID),
  with NO per-pod / per-replica / per-ordinal dimension. All
  replicas of one KarsSandbox share ONE agent identity.
- New test tag_layout_excludes_per_pod_attributes pins the tag
  layout — any future PR that adds a 'kars-pod-' / 'kars-replica-'
  / 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix
  breaks the test.
- Visibility change: AgentIdentityClient::tags_for is now
  pub(crate) so the cross-module invariant test can call it.

BOOTSTRAP SCRIPTS (deploy/bicep/standalone/):
- custom-security-attributes.sh: declares the recommended
  AgentGovernance set with 4 attributes — AgentClassification
  (Standard|Restricted|Confidential), DataSensitivity
  (Public|Internal|Confidential), ProductOwner, ManagedBy.
  Idempotent via az rest against the Graph beta endpoint.
  (The Microsoft.Graph Bicep extension does not yet ship typed
  attributeSets / customSecurityAttributeDefinitions resources,
  so this is a shell script that operators run once per tenant.)
- conditional-access-baseline.sh: applies Microsoft's
  policy-autonomous-agents template — blocks sign-ins where the
  risk level meets the configured threshold (default: high),
  targeted via the ManagedBy=kars-controller attribute filter.
  Defaults to report-only state for safe rollout. Idempotent via
  upsert.

FOUNDRY RBAC (research R5):
- foundry-rbac.bicep grants Azure AI User on a Foundry resource
  to the BLUEPRINT SP (not per-agent). All derived agent
  identities inherit access — eliminates per-agent role-assignment
  churn. Supports RG-scoped (default) and resource-scoped
  assignment via foundryResourceName parameter. az bicep build
  clean, no warnings.

DOCS:
- docs/architecture/entra-agent-id/05-security-alignment.md:
  6.8KB operator-facing runbook covering the bootstrap order,
  KarsSandbox YAML example, failure modes table, scale-out
  invariant rationale, Foundry RBAC inheritance.

TESTS:
- Controller: 802/802 (+13 new — 12 odata_type_for_value variants
  covering supported + rejected shapes, 1 scale-out invariant).
- Bicep: foundry-rbac.bicep builds clean. JSON compiled.
- Bash: both .sh scripts pass 'bash -n' syntax check.
- CLI: 786/786 (unchanged).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 29, 2026
* docs(entra-agent-id): import POC findings as architecture reference

Captures the end-to-end token-acquisition flow validated on real AKS
in the Microsoft tenant during the POC phase.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): scaffold Entra Agent ID auth machinery

Adds controller-side primitives for operating kars sandboxes as
per-sandbox Entra Agent Identities. Foundation only — pod-spec
integration lands in a follow-up commit.

New modules:
- auth_config.rs: KarsAuthConfig CRD (cluster-scoped singleton).
  Tenant + blueprint + controller MI anchors plus optional
  serviceManagementReference for Microsoft-style tenants.
- auth_config_reconciler.rs: materialises kars-auth-sidecar-env
  ConfigMap with AzureAd__ and DownstreamApis__ env vars. Idle when
  the CRD is absent (cluster stays in anonymous tier).
- agent_identity.rs: Graph client for per-sandbox agent identity SP
  lifecycle. Implements IMDS to MI to blueprint to Graph chain
  proven during POC.
- sidecar_injection.rs: pure-function pod-spec helpers (sidecar
  container shape, pinned-identity env vars, egress-guard iptables).

CRD additions in crd.rs:
- KarsSandbox.spec.meshAuth.mode (auto/agent-id/anonymous)
- KarsSandbox.status.agentIdentity

786 controller tests pass, 14 new tests added.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli): auto-provision Entra Agent ID trust during kars up

Makes Entra Agent ID setup invisible to end users: when `kars up`
runs and the cluster does not yet have a `KarsAuthConfig/default`
resource, the new `mesh/agent_id_setup.ts` module idempotently
provisions the tenant trust anchor end-to-end:

  1. Blueprint app via Graph (POST /v1.0/applications/ with
     @odata.type=#Microsoft.Graph.AgentIdentityBlueprint),
     including the optional serviceManagementReference required by
     Microsoft-style enterprise tenants.
  2. Blueprint service principal (visible in the Entra Agents portal).
  3. Controller managed identity in the customer's subscription.
  4. MI-as-FIC on the blueprint (issuer=login.microsoftonline.com),
     the anti-loop-safe credential path proven by the POC.
  5. KarsAuthConfig CR written to the cluster.

The step is non-fatal: if the user lacks `Agent ID Developer`,
`kars up` continues and the cluster runs in anonymous tier until
the role is granted and `kars mesh setup-trust` is rerun. This
matches the three-tier fallback model documented in
docs/architecture/entra-agent-id/.

New `--service-tree <guid>` flag on `kars up` (and KARS_SERVICE_TREE
env var) propagates the ServiceTree GUID to the blueprint creation.
No hardcoding — tenants that do not require it leave it empty.

Tests: 6 new unit tests for agent_id_setup (idempotence detection,
dry-run, env-var threading, error propagation). 775 CLI tests pass
(2 pre-existing skipped). typecheck + lint clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs+preflight: user-facing Entra Agent ID guidance

Adds the user-facing documentation layer that the previous controller
and CLI commits were missing. Every doc that mentioned the old
`api://agentmesh` flow now points to the new per-sandbox Entra Agent
ID model.

New docs:
- docs/agent-identity.md — day-1 and day-2 user guide. Prerequisites,
  walk-through of what `kars up` does for auth, sub-agent semantics,
  troubleshooting, teardown.

Updated docs:
- README.md — top-line mention of per-sandbox Entra Agent ID in the
  `kars up` summary, linking to the new guide.
- docs/getting-started.md — Step 2.1 calls out the `Agent ID
  Developer` role prerequisite; Step 2.2 expands the bring-up list to
  include the preflight role check and the Entra trust provisioning.
- docs/permissions.md — rewrites the "Tenant-level (Entra ID)
  considerations" section to describe the new model. Replaces
  `api://agentmesh` failure rows with Entra-Agent-ID-specific ones
  (CredentialInvalidLifetimeAsPerAppPolicy,
  InvalidFederatedIdentityCredentialValue, missing sidecar).
- docs/cli-reference.md — documents the new `--service-tree` flag
  + Microsoft-corp example.
- docs/SUMMARY.md — wires agent-identity.md into the mdbook index.

New preflight check:
- cli/src/preflight.ts — calls `checkAgentIdRole` from agent_id_setup.
  Warns (not blocks) when the signed-in user lacks the `Agent ID
  Developer` directory role. The existing `api://agentmesh` warning
  line is removed.
- cli/src/commands/mesh/agent_id_setup.ts — new exports:
  - `checkAgentIdRole` returns hasRole/inconclusive/message via
    Graph `/me/transitiveMemberOf`. Matches by role template id
    (stable) AND display name (forward-compatible).
  - `detectExistingBlueprint` returns whether the configured
    blueprint already exists in Graph.
  - `AgentIdSetupOptions.blueprintName` added so multi-cluster
    deployments can share a tenant-wide blueprint.

Tests: 781 CLI tests pass (+6 new for checkAgentIdRole / blueprint
detect). Typecheck + lint clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(cli): UX polish — from-scratch resets context, surface CA block

Two papercuts surfaced when the user ran `kars up --from-scratch` on
a tenant that already had a previous deployment:

1. `--from-scratch` cleared the resume state but NOT the cached
   deployment context, so the "fresh" run silently reused the prior
   region/RG/Foundry endpoint instead of re-prompting. Fixed in
   `cli/src/commands/up/preflight.ts` — when `fromScratch=true`, skip
   the `loadContext()` prefill entirely and print a clear
   "ignoring any cached deployment context" line. Adds the
   `fromScratch?` field to `UpOptionsForPreflight`.

2. The Entra Agent ID preflight check correctly soft-failed on the
   well-known AADSTS530084 Conditional Access token-binding block
   (common in Microsoft-corporate tenants), but the warning was
   generic and gave the user no actionable next step. Now detects
   AADSTS530084 (and the related AADSTS65001/65002 missing-consent
   codes) specifically and surfaces the exact `az login --scope
   https://graph.microsoft.com//.default` workaround inline.

Docs: adds an `#az-cli-ca-block` anchor section to
`docs/agent-identity.md` so the inline preflight message can link
directly to the troubleshooting paragraph.

Tests: 2 new test cases in agent_id_setup.test.ts pinning the
AADSTS530084 + AADSTS65001 detection paths. Total CLI tests:
783 pass (+2 vs prior commit).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): correct system namespace kars-system (was azureclaw-system)

Caught while auditing residual `azureclaw-` references after the
Azure/kars rebrand. The auth-config reconciler materialised the
sidecar env ConfigMap into `azureclaw-system`, but the actual
controller + helm chart deploy everything into `kars-system`. With
the wrong namespace the ConfigMap was unreachable from sandbox pods
even when the rest of the wiring was correct.

Two-line fix:
- controller/src/auth_config_reconciler.rs:61 — namespace constant.
- controller/src/auth_config.rs:43 — doc comment.

The rest of the controller already uses `kars-system` consistently
(pairing_reconciler, trust_graph_reconciler, signer_policy,
egress_approval_reconciler). This brings the new modules in line.

786 controller tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli): kars mesh setup-trust --mode agent-id (standalone retry)

The Entra Agent ID auto-provisioning that `kars up` runs is now
also reachable as a standalone command — necessary for retrying
after a transient failure (e.g. AADSTS530084 Conditional Access
block on Microsoft Graph) without re-running every other `kars up`
phase.

cli/src/commands/mesh/setup-trust.ts grows a `--mode <agent-id|legacy>`
flag, defaulting to `agent-id`. Agent-id mode forwards to the same
`ensureAgentIdTrust` helper used by `kars up`. Legacy mode preserves
the original api://agentmesh app-registration flow for installations
that haven't migrated yet (slated for removal once all consumers
have switched).

Also surfaces:
- --service-tree <guid> for Microsoft-style tenants
- --cluster-name / --resource-group / --region for controller MI
  override
- Clear AADSTS530084 + "Agent ID Developer missing" remediation
  messages on failure.

Bonus: cli/src/commands/up.ts agentmesh-image import block from a
prior in-flight edit — imports agentmesh-relay-agt + agentmesh-
registry-agt from the public source ACR in --build mode so the
agentmesh deploy step does not block on ImagePullBackOff. (Future
improvement: build them locally from .agt-sdk when the AGT SDK
tarball is present.)

Tests + typecheck clean: 783 CLI tests pass, build green.

* chore(helm): install KarsAuthConfig CRD

Adds the helm template for the new cluster-scoped singleton CRD
introduced in c9ce68f. Without this, `helm upgrade` does not install
the CRD and `kars mesh setup-trust --mode agent-id` fails with
"the server doesn't have a resource type karsauthconfig" when it
tries to `kubectl apply` the CR.

The schema uses `x-kubernetes-preserve-unknown-fields: true` on
spec/status with required-field validation for the top-level
tenant / agentId / controller blocks. The canonical type definitions
remain in controller/src/auth_config.rs; the controller validates
spec shape on reconcile. A future PR will add karsauthconfig to the
existing helm-vs-Rust drift test in controller/src/helm_drift.rs.

helm lint clean. YAML parses.

* feat(cli+bicep): auto-fallback to Bicep when az CLI Graph is CA-blocked

End-to-end UX for Microsoft-corporate-style tenants where the Azure
CLI cannot acquire a Microsoft Graph token (AADSTS530084) — the same
block we hit repeatedly during the POC phase.

What changes:
- deploy/bicep/agent-id-trust.bicep — sub-scope Bicep template that
  provisions everything the imperative path does: blueprint app + SP
  (tagged EntraAgentId), controller MI, MI-as-FIC on the blueprint
  using login.microsoftonline.com (universally allow-listed). Goes
  through ARM + the Microsoft.Graph extension, which has its own auth
  path and is not subject to az CLI CA token-binding policy.
- deploy/bicep/modules/controller-mi.bicep — RG-scope module for the
  controller MI (Bicep needs RG scope for UAMI, parent runs at sub
  scope to create the RG).
- deploy/bicep/bicepconfig.json — enables the Microsoft.Graph
  extension (preview).
- cli/src/commands/mesh/agent_id_setup_bicep.ts — driver that runs
  `az deployment sub create` against the template, parses outputs,
  and writes the KarsAuthConfig CR.
- cli/src/commands/mesh/agent_id_setup.ts — adds
  ensureAgentIdTrustAutoFallback() which tries the fast Graph REST
  path first and transparently switches to the Bicep path on
  AADSTS530084. Other Graph errors propagate unchanged (don't mask
  real permission failures with a Bicep retry).
- cli/src/commands/up.ts — uses the auto-fallback wrapper so
  `kars up` "just works" in any tenant.
- cli/src/commands/mesh/setup-trust.ts — same auto-fallback for
  `--mode agent-id`, and a new `--mode bicep` for users who want to
  skip the CLI attempt entirely.
- Also makes blueprintName configurable so multi-cluster deployments
  in the same tenant share one blueprint by default.

Tests: 783 CLI pass, typecheck clean, helm-lint + bicep-build clean.

Operational note (not code): the AGT registry/relay images the user
manually built+pushed during this session were from a feature branch
(copilot/secure-mcp-agent-governance) that adds proof-of-possession
to /v1/agents — incompatible with SDK 3.7.0 which doesn't sign yet.
Rebuilding from tag v3.7.0 + restarting the deployments resolves the
422 flood. Documented for future kars releases to build from a known
AGT tag rather than HEAD.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(bicep): clean error reporting + linter-clean issuer URL

Three small fixes shaken out by the user's first end-to-end run of
`kars mesh setup-trust --mode bicep` in the Microsoft corporate
tenant:

1. deploy/bicep/agent-id-trust.bicep — the MI-as-FIC issuer was
   hardcoded to `login.microsoftonline.com`. Bicep linter emits
   `no-hardcoded-env-urls` as a warning on stderr, which the CLI
   driver was mistaking for a deployment failure. Switched to
   `environment().authentication.loginEndpoint` so the template is
   cloud-portable AND the linter is silent.

2. cli/src/commands/mesh/agent_id_setup_bicep.ts — error capture
   reworked to use execa's merged `all` stream so the real ARM
   deployment error surfaces instead of being masked by stderr-only
   linter warnings. Lines starting with `WARNING:` are stripped from
   the error summary.

3. Same module — the final `kubectl apply KarsAuthConfig` is now
   correctly recognised as a separate step from the Bicep deployment.
   When the CRD isn't installed (e.g. older controller helm release),
   the user sees:
     "Bicep deployment succeeded, but KarsAuthConfig CRD is not
      installed. All Entra resources are already created — just
      install the CRD and re-run."
   plus the exact `helm upgrade` command. The function returns the
   Bicep result so callers know the Entra side is complete.

Also briefly probed the Microsoft.Graph Bicep extension's beta
channel (1.0.0) to see if it supports the `@odata.type` discriminator
for typed AgentIdentityBlueprint. It does not (`BCP037 The property
"@odata.type" is not allowed on objects of type
"Microsoft.Graph/applications"`). The Bicep path therefore creates a
regular Application tagged `EntraAgentId` — functional for runtime
but not visible under the Entra Agents portal page. Documented as a
limitation; users who need portal visibility for the typed form must
use Graph Explorer or PowerShell (both have different first-party
auth paths that aren't CA-blocked).

Bicep build clean, typecheck clean, 783 tests pass.

* fix(cli+docs): switch Graph calls from /v1.0 to /beta

User reported that the typed POST creating an
`#Microsoft.Graph.AgentIdentityBlueprint` works in Graph Explorer
against the /beta endpoint — and that's what makes the resulting
app visible in the Entra Agents portal. The /v1.0 endpoint accepts
the same body but does NOT route through the typed-resource
discriminator on the server, so the app ends up listed only under
App registrations.

Updates `cli/src/commands/mesh/agent_id_setup.ts` to use
/beta/applications, /beta/servicePrincipals, /beta/users, /beta/me,
and /beta/applications/{id}/federatedIdentityCredentials in every
Graph REST call. The `@odata.type` body field stays
`#Microsoft.Graph.AgentIdentityBlueprint` — only the URL prefix
changes.

docs/permissions.md updated to match (the manual escape-hatch
snippet showing `az rest --method POST --url
https://graph.microsoft.com/v1.0/applications/...` now reads `/beta/`).

Note: this only helps when the imperative Graph REST path runs
successfully. In tenants where the az CLI is Conditional-Access-
blocked from Graph (AADSTS530084), the CLI auto-falls back to the
Bicep ARM path. The Bicep Microsoft.Graph extension does NOT
support the @odata.type discriminator at v1.0 or beta (1.0.0), so
the Bicep-created app is functional but stays in the tag-based
detection mode (App registrations only, not Agents portal). This
is a current limitation of the extension, documented in
agent_id_setup_bicep.ts.

Tests: 783 CLI pass, typecheck clean.

* feat(cli): device-code re-login + exact-match Graph body for typed blueprint

Two coordinated changes to make the imperative Graph REST path
succeed in tenants with Conditional Access token-binding policy
(AADSTS530084) on the az CLI's first-party app:

1. Match Graph Explorer's exact working body shape:
   - @odata.type value WITHOUT the leading `#` ("Microsoft.Graph.
     AgentIdentityBlueprint", not "#Microsoft.Graph.…"). Both forms
     are accepted by /v1.0; only the unprefixed form survives the
     /beta route reliably per user testing.
   - sponsors@odata.bind / owners@odata.bind URLs use /v1.0/users/.
     Graph rejects /beta/users/ as an @odata.bind target.
   - POST URL stays /beta/applications/ — the typed
     AgentIdentityBlueprint resource only surfaces in the Entra
     Agents portal page when created via the /beta route.

2. Auto device-code re-login on AADSTS530084:
   - When `az rest` returns AADSTS530084, the helper does a one-shot
     `az login --use-device-code --scope https://graph.microsoft.com//.default`
     which goes through a different OAuth flow that often bypasses
     the token-binding CA policy applied to the default interactive
     flow.
   - User is prompted in the terminal: "visit https://microsoft.com/
     devicelogin, paste this code". After login, the original Graph
     call is retried once. If it still fails, the
     ensureAgentIdTrustAutoFallback wrapper proceeds to the Bicep
     ARM path (untyped but functional).

Updates existing test fixture for checkAgentIdRole so the
device-code retry path is mocked. 783 CLI tests pass.

Together these should let Microsoft-corporate-tenant users get the
typed AgentIdentityBlueprint that surfaces in the Entra Agents
portal page, without needing to fall back to Bicep (which can only
produce the tag-based form).

* docs(agent-identity): document AADSTS530033 + Graph Explorer workaround

User hit the next layer of CA policy after device-code re-login:
`AADSTS530033` — "device must be Intune-managed" — applies to
both interactive and device-code flows of the Microsoft Azure CLI
first-party app. Bicep ARM path keeps working (different auth
surface), but it produces an untyped blueprint that only shows
under App registrations, not the Entra Agents portal page.

Updates the existing AADSTS troubleshooting section with:
- Table of auto-handled error codes + what the kars CLI does
- Explicit "even Bicep cannot produce the typed form" path:
  Graph Explorer PATCH to upgrade the untyped blueprint app to
  the typed AgentIdentityBlueprint discriminator in place
- Note that runtime is unaffected by typed-vs-untyped — only
  portal categorisation differs
- Long-term Intune-enrolment recommendation

* fix(cli): kars mesh setup-trust short-circuits when already provisioned

Two bugs surfaced when the user re-ran setup-trust after the
initial successful Bicep + Graph Explorer flow:

1. `karsAuthConfigExists()` checked for "karsauthconfig/default" in
   kubectl's `-o name` output, but newer clusters return the
   fully-qualified form `karsauthconfig.kars.azure.com/default`.
   The bug caused the existence check to always return false, so
   the wrapper always tried the Graph REST path even when the trust
   was already in place. Fix: match the invariant `/default`
   suffix.

2. `kars mesh setup-trust --mode agent-id|bicep` did not consult
   `karsAuthConfigExists()` at all, so every re-run triggered the
   full provisioning attempt — which, in CA-blocked tenants, kicks
   off a device-code login prompt the user has no reason to deal
   with when nothing needs to be done. Fix: short-circuit at the
   top of both modes when the CR already exists, with a clear
   message and the `kubectl delete + retry` escape hatch for users
   who genuinely want to re-provision.

Tests: 784 CLI pass (+1 new for the FQ-name kubectl output).

* fix(bicep-fallback): print full Graph Explorer PATCH for portal visibility

The Bicep `Microsoft.Graph` extension cannot set `@odata.type`, so the
Bicep-created blueprint is a plain `Application` that the kars runtime
uses fine but that the Entra portal's Agents page does not show. The
existing docs explained the workaround as a one-field PATCH that
sets `@odata.type` only — that does upgrade the type but the entry
*still* stays hidden in the portal because the Agents-page filter
also requires `sponsors` and `owners` to be set.

Changes:

- agent_id_setup_bicep.ts: at the end of every Bicep success path
  (and also on the CRD-missing soft-failure branch), print a fully
  populated Graph Explorer PATCH body the user can paste verbatim.
  The body includes `@odata.type` + `sponsors@odata.bind` +
  `owners@odata.bind` so the resulting typed blueprint actually
  shows up under Entra portal → Identity → Agents.

  We try `az ad signed-in-user show --query id -o tsv` to auto-fill
  the user OID. In the very tenants where this matters (Microsoft
  corp Macs without Intune enrollment) that command also fails with
  AADSTS530084 — handled silently with a `<YOUR_USER_OID>` placeholder
  and a one-line hint on where to find the OID in the Entra portal.

- docs/agent-identity.md: replace the misleading one-field PATCH
  example with the full body, including notes about: no `#` prefix on
  `@odata.type` in the request body, `/v1.0/users/` required in
  `@odata.bind` even when the parent URL is `/beta/`, and a pointer
  to the CLI's auto-generated copy-paste body.

Tests: 786 CLI pass (incl. 15 agent_id_setup tests).

* docs+cli: full delete+recreate runbook for typed blueprint (SP+FIC included)

A user hit the case where the in-place @odata.type PATCH was rejected
and they recreated the blueprint via Graph Explorer's POST /applications.
The recreated app then lacked an SP (so it couldn't receive RBAC role
assignments and stayed invisible in the Agents portal) and lacked the
MI-as-FIC (so the controller couldn't mint child identities).

Bicep creates all three resources (app + SP + FIC) but the Bicep
Microsoft.Graph extension cannot set the @odata.type discriminator
needed for a typed agentIdentityBlueprint — so Graph-Explorer recovery
remains the only path in CA-blocked tenants, and it must do all three
steps explicitly.

- agent_id_setup_bicep.ts: extend printPortalVisibilityHint() to print
  the full 5-step recovery runbook (POST app, POST SP, POST FIC, kubectl
  patch CR, optional DELETE old app) underneath the simpler in-place
  PATCH path that's still tried first.

- docs/agent-identity.md: replace the one-line "delete + recreate"
  hand-wave with the full POST sequence with all body shapes, plus the
  kubectl jsonpath one-liner to extract tenantId + MI principalId from
  the cluster.

Tests: 786 CLI pass.

* fix(controller): own egress-guard script in one place + fix iptables rule order

The original `agent_id_egress_rules()` shipped on `feat/entra-agent-id`
contained a security bug: every rule used `-A OUTPUT` (append). The
pre-existing baseline egress-guard script (currently emitted as a
`concat!` literal in `reconciler/mod.rs`) starts with
`-A OUTPUT --uid-owner 1000 -o lo -j ACCEPT`, so any later appended
rule blocking UID 1000 → 127.0.0.1:8080 would NEVER fire — the
loopback-allow would match first. This would have silently broken the
"agent cannot impersonate the router and mint downstream tokens"
boundary as soon as sidecar injection is wired up.

Rubber-duck critique caught this (finding #5). Fix:

- Refactor `agent_id_egress_rules()` to emit two correctly-positioned
  rules. The sidecar-block uses `-I OUTPUT 1` (insert at chain head)
  so it runs BEFORE the baseline loopback-allow. The router-IMDS
  block can stay `-A` because no prior `--uid-owner 1001` rule exists.
- Drop the redundant UID 1000 → IMDS rule (catch-all `DROP UID 1000`
  already covers it).
- Drop the explicit UID 1002 → IMDS ACCEPT (OUTPUT chain default is
  ACCEPT and no UID 1002 restriction exists).
- New `build_egress_guard_command(agent_id_mode: bool)` composes the
  full shell script: agent-id rules first (so `-I` semantics + script
  text order both keep the security boundary), then the seven
  baseline iptables lines, then a mode-specific echo. `&&`-chained so
  any iptables failure aborts init-container startup (partial policy
  is worse than no policy).
- Reconciler/mod.rs replaces the `concat!` literal with a call to the
  new helper (currently always passing `false`; agent-id pass that
  flips to `true` lands in a follow-up commit on the same branch).

Tests:
- New `egress_rules_use_insert_before_baseline_loopback_allow` — pins
  the `-I OUTPUT 1` semantics. Direct regression for the security bug.
- New `egress_guard_command_legacy_mode_matches_existing_behaviour` —
  byte-for-byte pin on the seven historical iptables lines so the
  refactor is provably no-op for non-agent-id sandboxes.
- New `egress_guard_command_agent_id_mode_prepends_security_rules` —
  asserts the sidecar REJECT appears BEFORE the loopback ACCEPT in
  the script text (defence-in-depth even if a reader misreads -I).
- New `egress_guard_command_is_shell_safe_chained` — every step
  starts with `iptables ` or `echo `.

789/789 controller tests pass.

* feat(controller): per-sandbox agent identity provisioning + sidecar injection

The end-to-end controller-side of agent-id mode. New module
`agent_id_provisioning` orchestrates the full flow that the rubber-duck
critique identified as the missing glue: resolve mesh-auth mode →
provision (or recover) the per-sandbox Entra Agent Identity via Graph
→ patch sandbox status → materialise the per-namespace sidecar env
ConfigMap → return a Ready outcome that the sandbox reconciler uses
to inject the sidecar container, pin the router env, and flip the
egress-guard into agent-id mode.

Architecture: model (D) from the critique — the controller provisions
the per-sandbox identity BEFORE pod creation, status-pins the appId,
and the router uses that pinned appId in every sidecar request. No
per-sandbox Secret; no sandbox-side workload identity binding.

## What this commit adds

### New module: `controller/src/agent_id_provisioning.rs`

- `ProvisionerCache` — process-wide cache of `AgentIdentityClient`s
  keyed by blueprint client ID. Shares token caches and connection
  pools across concurrent sandbox reconciles.
- `resolve_mesh_auth_mode` — pure-function 3-way resolution (Auto →
  AgentId or Anonymous based on KarsAuthConfig readiness; explicit
  AgentId without ready config surfaces a distinct `AuthConfigNotReady`
  reason so operators can distinguish "tenant not set up" from
  "auto-fallback to anonymous").
- `load_auth_config` — singleton fetch with explicit `Ok(None)` for
  the 404 (anonymous-tier fallback) vs. `Err` for transient failures.
- `ensure_agent_identity_for_sandbox` — idempotent orchestration with
  the three-step recovery flow the critique specified:
    1. If `status.agentIdentity` recorded, GET via Graph; reuse on
       200, reprovision on 404, requeue on 5xx.
    2. If status empty, `list_cluster_agent_identities` filtered by
       the `kars-sandbox-uid:<uid>` tag. Catches the
       "Graph create succeeded but controller crashed before status
       patch" crash window.
    3. Otherwise create new, patch status, return.
- `materialise_sidecar_configmap` — copies the rendered sidecar env
  into the sandbox namespace (`envFrom` cannot cross namespaces — the
  critique caught this). Owned by the KarsSandbox via
  `ownerReferences` so K8s garbage-collects on sandbox deletion.

### Modified: `controller/src/reconciler/mod.rs`

- `Context` gains `cluster_uid` (read from `kube-system` ns metadata
  at startup — canonical "this cluster" identifier in K8s; falls back
  to `KARS_CLUSTER_UID` env or generated string with a warning) and
  `agent_id_cache: Arc<ProvisionerCache>`.
- `reconcile` calls `ensure_agent_identity_for_sandbox` BEFORE pod-spec
  assembly. Match on outcome: `Skipped` → legacy path; `Ready` →
  capture identity for downstream injection; `Failed` → patch
  Degraded status with `AgentIdentityProvisioningFailed` reason and
  requeue. No silent fallback to legacy on AgentId failure — explicit
  user intent must not be downgraded.
- Pod-spec assembly:
    - Egress-guard `command` flips to `build_egress_guard_command(true)`
      when `agent_id_active.is_some()` — adds the security-critical
      `-I OUTPUT 1` REJECT rule for UID 1000 → sidecar.
    - Router `env` gains `PINNED_AGENT_IDENTITY_APP_ID` (the
      per-sandbox appId) + `AUTH_SIDECAR_URL` (loopback :8080).
    - Sidecar container is appended to `containers` after the
      `runtimeClassName` block. Image pinned to the GA Microsoft
      distroless build (overrideable via `KARS_SIDECAR_IMAGE`).

### Modified: `controller/src/auth_config_reconciler.rs`

- After successful ConfigMap apply, `patch_ready_status` patches
  `phase=PHASE_READY` + `SidecarConfigMaterialized=True` condition.
  This is the real readiness signal that the sandbox reconciler's
  `resolve_mesh_auth_mode` gates on (per critique #7). Best-effort:
  status-patch failure is logged but doesn't fail the reconcile.
- `build_condition_blueprint_ready` no longer `#[allow(dead_code)]` —
  it's now consumed by `patch_ready_status`. Type renamed from
  `BlueprintReady` (Graph-call check) to `SidecarConfigMaterialized`
  (controller-side materialisation check) since the BlueprintReady
  signal will come from a separate cluster-health probe.

## Test status

- 795/795 controller tests pass (+6 new for the new module).
- phase_taxonomy_guard pre-existing integration test passes (the
  status patch uses `PHASE_READY` constant, not a string literal).
- Sidecar-injection tests still all pass (the iptables refactor from
  the previous commit holds).

## Not yet in this commit (deferred to follow-up commits/PRs)

- Inference-router `sidecar_client.rs` that actually consumes the
  pinned env vars (todays AgentIdentity unused on the router side).
- CLI changes: `kars up` VMSS-MI assignment; `kars mesh setup-trust
  verify` end-to-end audit; Foundry RBAC assignment print.
- `agent_identity_reaper` for cleaning up orphan SPs.
- Sponsor user object IDs in KarsAuthConfig spec (currently empty
  array passed; works in tenants where sponsors aren't required).

* fix(e2e): make Entra Agent ID chain work end-to-end on real AKS

Concludes a long live-debug session against kars-aks where we walked
the full controller → sidecar → router auth chain and fixed every
blocker until the sidecar successfully mints tokens from Entra and
relays them to the router. The remaining gate at end-of-chain is the
Microsoft corp tenant's Conditional Access policy on agent identities
(AADSTS53003 with a `capolids` claims challenge) — a tenant
configuration issue, not a code issue. The architecture is proven
correct.

- **Drop the GET-by-id verify on recorded identity.** Graph's
  `GET /servicePrincipals/{id}` has multi-second eventual-consistency
  after creation; treating a 404 as "stale, reprovision" creates a
  runaway-creation loop (live cluster produced 70+ duplicate SPs per
  minute before this fix). The reaper (separate PR) handles
  out-of-band deletes via tag-scoped listing.
- **Idempotency guard on `patch_sandbox_status`.** Skip the SSA patch
  when the recorded status already matches — same `lastTransitionTime`
  drift pattern that bit the auth-config reconciler.
- **Drop ownerReference on the per-namespace sidecar CM.** KarsSandbox
  lives in `kars-system`; the per-sandbox CM lives in `kars-<name>`.
  K8s rejects cross-namespace owner refs with `OwnerRefInvalidNamespace`
  and GC's the CM seconds after creation (debugged via the kubelet
  events stream). Cleanup happens via namespace deletion instead.
- **Add apiVersion+kind to status patch body.** SSA patches without
  the top-level type meta return `BadRequest: invalid object type:
  /, Kind=`.

- **IMDS-first, WI fallback.** Workload-Identity-derived tokens cannot
  be used as FIC assertions (Entra anti-loop AADSTS700231). The
  controller MUST acquire its MI token via IMDS — which requires the
  MI assigned to the AKS node-pool VMSS (`kars up` does this).
- **Permissive Graph list parser.** The agent identity list response
  uses `agentAppId`, not `appId`. The strict-schema deserialiser
  failed silently on every list call, so the tag-based recovery path
  never found anything and the controller created a fresh SP every
  reconcile. Fallback parser handles both field names.

- **Patch `phase=Ready`** so `resolve_mesh_auth_mode` has a real
  signal to gate on (rubber-duck critique #7). No-op when the
  observed status already matches the desired — avoids the SSA
  `lastTransitionTime` reconcile loop.

- **`AUTH_SIDECAR_URL=http://localhost:8080`** not `http://127.0.0.1:8080`.
  The Microsoft Entra SDK sidecar's HostFiltering middleware only
  allows `Host: localhost`; calls with `Host: 127.0.0.1` are rejected
  with `400 Bad Request - Invalid Hostname`. /etc/hosts in every pod
  maps `localhost` → `127.0.0.1` so the loopback semantics are
  identical.
- **`httpHeaders: Host=localhost:8080`** on the readiness/liveness
  probes. The kubelet sends the pod IP as the default Host header,
  which the sidecar rejects with the same 400. Without this override,
  every sidecar stays ready=false even when fully functional.
- **emptyDir at /app/keys.** The sidecar's ASP.NET Data Protection
  writes encryption keys to `/app/keys`; the distroless image has
  /app owned by root and UID 1002 cannot mkdir there. Per-pod
  emptyDir gives a writable scratch.

- Append the keys volume to `pod_spec.volumes` whenever the sidecar
  is injected. Without this the container immediately crashes on
  startup with `Access to the path '/app/keys' is denied`.

The Microsoft Entra SDK sidecar's response shape varies across
builds. We've observed three forms in the wild:

1. `{"AuthorizationHeader": "Bearer xxx", "ExpiresIn": 3600}` —
   documented contract.
2. `"Bearer xxx"` — JSON-quoted string only.
3. `Bearer xxx` — plain text, no JSON.

The strict parser broke on shape 2/3 with `parse auth-sidecar JSON
response`, silently disabling agent-id auth for the entire pod. New
`parse_sidecar_body()` handles all three lenient + 5 dedicated tests.

Without this rule the controller's IMDS call times out (CILIUM
default-deny drops 169.254.169.254). IMDS is per-node link-local —
not externally routable, no exfiltration risk.

- KarsAuthConfig.spec.agentId.sponsorUserObjectIds — required in
  tenants where Graph rejects agent identity creation without a
  sponsor (live error: `No sponsor specified`).
- KarsSandbox.status.agentIdentity — appId/objectId/displayName/
  createdAt. Without this in the CRD schema, SSA patches fail with
  `field not declared in schema`.

- Grant the controller `get/list/watch/create/update/patch/delete`
  on `karsauthconfigs` (and `/status`). Without this the auth-config
  reconciler stays dormant ("KarsAuthConfig CRD not installed").

- 795/795 controller tests pass (+ 0 net change; the existing
  sidecar_injection tests covered the URL constants we updated).
- 890/890 inference-router lib tests pass (+5 new
  `parse_sidecar_body*` regression tests).

- agent_identity_reaper module (cleans Graph orphans by tag — needed
  long-term to handle out-of-band deletes; for now duplicates require
  a manual cleanup script).
- CLI flow: `kars up` should run `az identity federated-credential
  create` for the controller SA so the WI fallback works in clusters
  without VMSS-attached MI.
- Per-sandbox Foundry RBAC automation — verify already prints the
  exact `az role assignment` commands; landing this would auto-grant.

```
✅ KarsAuthConfig provisioning (Bicep + Graph Explorer fallback)
✅ Controller MI assigned to AKS VMSSes
✅ Controller IMDS → MI → blueprint Graph token chain working
✅ Per-sandbox agent identity creation via Graph (with sponsors)
✅ Sandbox status.agentIdentity correctly recorded
✅ Per-namespace sidecar env ConfigMap materialized
✅ Sandbox pod contains 3 containers (openclaw + router + auth-sidecar)
✅ auth-sidecar healthy + listening on localhost:8080
✅ Router detects sidecar mode and routes auth through it (fail-closed)
✅ Sidecar mints Entra tokens for the pinned agent identity
✅ Sidecar relays token (or Entra error) back to router

⚠️ Foundry call returns 401 AADSTS53003 — Microsoft corp tenant's
   Conditional Access policy blocks agent identities. This is the
   same family of CA wall the user has been hitting all session for
   their human account. Tenant configuration issue; the chain itself
   is provably correct.
```

* fix(sidecar): bind on all interfaces so kubelet probes can reach /healthz

The Microsoft Entra SDK sidecar was bound to `127.0.0.1:8080` only,
which meant the kubelet's readiness/liveness probes (which connect to
the pod IP, not loopback) got `connection refused` at the TCP layer
before any HTTP exchange could happen. Overriding the probe's `Host`
header is useless when the connect itself fails.

Fix: bind on all interfaces (`http://+:8080`). Three defence-in-depth
layers keep cross-pod traffic out of the sidecar:

  1. Sidecar's HostFiltering middleware accepts only `Host: localhost`;
     anything else gets 400 Bad Request.
  2. The sandbox NetworkPolicy has no ingress rule allowing port 8080,
     so cross-pod packets are dropped before reaching the sidecar.
  3. Egress-guard iptables rule REJECTs UID 1000 → 127.0.0.1:8080,
     blocking in-pod privilege escalation from the agent container.

Verified live on kars-aks (sandbox kars-testrun):
  - auth-sidecar: ready=true, restarts=0
  - kubelet probe now succeeds via pod-IP connect + Host-header
    override
  - cross-pod connections are still rejected (no NP ingress rule)

This was the last container-level blocker keeping the chain from
running end-to-end. With this fix the sidecar mints real Entra tokens
and the router relays them to Foundry (where the only remaining gate
is the per-agent-identity Azure AI User role assignment).

* refactor(controller): pivot to shared-sidecar architecture (Phase 0)

This commit prepares the branch for the shared-sidecar redesign:
deletes the per-pod sidecar injection plumbing while preserving the
controller-side provisioning machinery, the Graph client, and all
hard-won bug fixes.

## What's removed

- `controller/src/sidecar_injection.rs` — per-pod container spec
  helpers (build_sidecar_container, build_router_pinned_identity_env,
  build_router_sidecar_url_env, egress-guard agent-id-mode rules).
  In the new design the auth-sidecar runs ONCE per cluster as a
  Helm-managed Deployment in `kars-system`; per-pod injection is no
  longer needed.
- `cli/src/commands/mesh/vmss_mi_assign.ts` — VMSS managed-identity
  assignment driver. Not needed when the sidecar runs in kars-system
  with its own Workload Identity binding.
- `materialise_sidecar_configmap` per-namespace mirror in
  `agent_id_provisioning.rs`. The shared sidecar consumes a single
  `kars-system`-scoped ConfigMap managed by `auth_config_reconciler`.
- The agent-id-mode iptables rules in the egress-guard init container
  (UID 1000 → REJECT 127.0.0.1:8080, UID 1001 → REJECT IMDS). Trust
  boundary for cross-pod sidecar access moves to a NetworkPolicy on
  the sidecar's namespace (added in a follow-up commit).

## What's preserved

- `controller/src/agent_identity.rs` — Graph client. Same provisioning
  endpoints, same auth flow. With or without per-pod sidecar.
- `controller/src/agent_id_provisioning.rs` — provisioning loop with
  idempotency, tag-based recovery from crashes, permissive Graph list
  parser, status patch idempotency. Hard-won lessons preserved.
- `controller/src/auth_config.rs` + `auth_config_reconciler.rs` —
  KarsAuthConfig CRD + its reconciler. Still cluster-level singleton
  config; the sidecar just consumes a kars-system-scoped CM now.
- `inference-router/src/sidecar_client.rs` — URL-agnostic client.
  The new design points it at the cluster Service DNS instead of
  loopback; no code change needed.
- All Helm CRD templates, RBAC additions, IMDS NetworkPolicy egress
  rule, the Bicep blueprint provisioning template, and the
  Graph Explorer recovery runbook docs.

## What's reset to baseline behaviour

- `reconciler/mod.rs` egress-guard: back to the seven-line baseline
  iptables script (the same one shipped on `main`). No agent-id-mode
  branching; the security boundary for cross-pod sidecar access is
  the sidecar-namespace NetworkPolicy, not pod-local iptables.
- Inference-router env injection: still pushes
  `PINNED_AGENT_IDENTITY_APP_ID` + `AUTH_SIDECAR_URL` when an agent
  identity is provisioned, but `AUTH_SIDECAR_URL` now points at
  `http://entra-auth-sidecar.kars-system.svc:5000` instead of
  loopback. Helm chart for the sidecar Deployment lands in the next
  commit.

## Why this redesign

Microsoft's own Agent ID design-pattern docs and the Auth Sidecar
API support `?AgentIdentity=<appId>` so one sidecar instance can
mint tokens for any of the blueprint's children. We had been running
one sidecar per pod (N × 128MB), iptables-isolating cross-UID access
to localhost:8080. Switching to one shared sidecar (kars-system,
2 replicas, NetworkPolicy-gated) saves ~5x memory at 10 sandboxes,
eliminates the per-pod ASPNETCORE bind + probe gymnastics, and aligns
with the Microsoft.Identity.Web library's centralized MSAL cache.

Independent research review confirmed the per-pod model is one of
two documented patterns but is more resource-intensive; the
`?AgentIdentity=<appId>` shared-sidecar pattern is explicitly
documented as a valid alternative and the SDK's `AzureAd__ClientId`
takes the blueprint appId, making each child mintable on demand.

## Tests

785/785 controller tests pass. Inference-router tests untouched.

* feat(helm): shared entra-auth-sidecar Deployment (Phase 1)

One auth-sidecar Deployment per cluster (2 replicas, HA) serves all
sandboxes via the Microsoft AuthorizationHeaderUnauthenticated
?AgentIdentity=<appId> query parameter — replacing the per-pod
sidecar injection model from earlier iterations.

Templates added (all gated on `entraSidecar.enabled`, default false):
- auth-sidecar-serviceaccount.yaml  - WI-annotated SA
- auth-sidecar-deployment.yaml      - distroless sidecar, /healthz probes
- auth-sidecar-service.yaml         - ClusterIP :5000
- auth-sidecar-networkpolicy.yaml   - ingress only from sandbox routers

values.yaml gets the `entraSidecar:` config block populated by
`kars up` from the KarsAuthConfig CR.

Trust boundary rationale and credential-supply patterns (WI for OSS,
IMDS-MI for corp tenant) explained inline in each template.

Resource footprint: ~160MB cluster-total vs ~1.28GB at 10 sandboxes
under the per-pod model.

Tests: helm lint clean, helm template renders 4 resources when
enabled, 0 when disabled. Controller 785/785 unchanged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): allow sandbox→kars-system:5000 egress for shared sidecar (Phase 2)

The shared entra-auth-sidecar lives in the kars-system namespace and
is reached by every sandbox's inference-router over TCP 5000. The
sandbox-policy NetworkPolicy now includes an explicit egress rule
selecting namespaces labeled
  app.kubernetes.io/name=kars
  app.kubernetes.io/component=system
on port 5000.

Also fixed the auth-sidecar's own ingress NetworkPolicy: the selector
now correctly targets pods labeled kars.azure.com/component=sandbox
(the pod label) rather than the non-existent
kars.azure.com/component=inference-router (which is a container, not
a pod, and so was never matchable).

Trust boundary is now two-sided:
- Sandbox-side: egress allows reaching only the kars-system namespace
  on port 5000, where the sidecar Service is the only listener.
- Sidecar-side: ingress allows only pods in sandbox-labeled namespaces.

When entraSidecar.enabled=false at the Helm level, no auth-sidecar
exists and the egress rule is a harmless no-op (no destination to
reach).

Tests: controller 785/785, helm lint clean, helm template renders
correctly with enabled=true and enabled=false.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(router): tid+principal+aud+exp pinning on sidecar tokens (Phase 3)

Wire the previously-orphaned sidecar_client module into the
inference-router and harden it against token substitution attacks.

CONTROLLER (reconciler/mod.rs):
- Extends agent_id_active tuple to carry tenant_id alongside agent
  identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via
  ProvisioningOutcome::Ready.auth_spec.
- Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode
  sandboxes. Without it, router-side tid pinning is disabled and
  warns at boot (insecure, dev-only).

ROUTER (sidecar_client.rs, auth.rs, lib.rs):
- pub mod sidecar_client; — fix the orphaned-module bug.
- WorkloadIdentityAuth now consults SidecarClient first; sidecar mode
  is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves
  per-sandbox audit attribution in downstream Azure RBAC.
- from_env() now Result<Option<Self>>. Partial config (only URL or
  only PINNED set) returns Err; WorkloadIdentityAuth::new() panics →
  AKS surfaces as CrashLoopBackoff. No more silent fallback to a
  different identity model.
- HARD-fail validate_token_claims on EVERY check (rubber-duck #1):
    * tid mismatch / missing (cross-tenant guard)
    * appid/azp mismatch when present (audit attribution guard)
      — if both present, BOTH must match
    * aud mismatch against per-service expected set (cache-poisoning
      guard) — Foundry, Graph, OpenAI, Management, Search mapped
    * exp in past or within 60s skew (stale-token guard)
    * exp missing when tid-pinning enabled (unbounded-lifetime guard)
- Returns a TTL cap = exp - now - 60s; caller takes
  min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching.
- JWT decoder tightened: requires EXACTLY 3 segments (rejects
  2-segment unsecured JWS and 4-segment JWE).
- aud claim normalized: accepts both single-string (Entra default)
  and array (RFC 7519); rejects out-of-spec shapes (number, bool).
- Soft WARN when both appid and azp are absent (some MSI tokens
  legitimately omit both).
- New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning
  state for operator visibility.

TESTS:
- 39 sidecar_client unit + wiremock integration tests, covering all
  HARD-fail branches, the cache-skip on validation failure, the TTL
  cap math, the partial-config Err path, JWT decoder edge cases
  (2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud,
  out-of-spec aud).
- Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client.

Rubber-duck review caught: (1) appid/azp should be hard not soft, (2)
partial env config must fail closed, (3) 3-segment requirement, (4)
aud validation against requested resource, (5) exp validation + TTL
cap. All five addressed in this commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli,bicep,controller): dual credential mode auto-detect (Phase 4)

Adds first-class support for two auth-sidecar credential supply
patterns and auto-detects which one works in the current Entra tenant.

PATTERN A (ManagedIdentityImds, default):
- Sidecar auths via SignedAssertionFromManagedIdentity against the
  controller MI's IMDS endpoint. Corp-tenant safe; required when
  the tenant's FIC issuer-allowlist policy blocks AKS OIDC
  (Microsoft-corporate: InvalidFederatedIdentityCredentialValue).
- Bicep creates: blueprint + SP + controller MI + MI-as-FIC.

PATTERN B (WorkloadIdentity):
- Sidecar auths via SignedAssertionFilePath against the projected
  K8s SA token. No per-cluster MI, no VMSS identity assignment.
- Bicep creates: blueprint + SP + SA-as-FIC pointing at the AKS
  cluster's OIDC issuer URL.

CRD (controller/src/auth_config.rs):
- Added `controller.credentialMode` enum (ManagedIdentityImds default).
- Made `controller.managedIdentity{Client,Resource,Principal}Id`
  Optional (required only in MI mode).
- Added is_valid_for_mode() validator.

CONTROLLER (auth_config_reconciler.rs):
- render_sidecar_env branches on credentialMode and emits the
  correct AzureAd__ClientCredentials__0__SourceType. MI mode
  emits ManagedIdentityClientId; WI mode emits
  SignedAssertionFileDiskPath=/var/run/secrets/azure/tokens/azure-identity-token.
- Spec validation runs BEFORE ConfigMap materialisation. Invalid
  specs (MI mode + empty clientId) surface as phase=Degraded
  with InvalidCredentialMode condition — refuses to propagate
  the misconfiguration to running sandboxes.
- New patch_degraded_status helper.

BICEP (deploy/bicep/agent-id-trust.bicep):
- New credentialMode parameter (allowed: ManagedIdentityImds |
  WorkloadIdentity).
- Conditional resources: controller MI + MI-as-FIC only in
  Pattern A; SA-as-FIC only in Pattern B.
- New aksOidcIssuerUrl parameter (required for Pattern B).
- New outputs: credentialMode, aksOidcIssuerUrl.
- az bicep build clean, no warnings.

CLI:
- ensureAgentIdTrust now accepts credentialMode (auto | WorkloadIdentity
  | ManagedIdentityImds), aksClusterName, aksClusterResourceGroup.
- AUTO MODE: tries Pattern B first by discovering the AKS OIDC
  issuer URL via az aks show, then creating an SA-as-FIC. On
  InvalidFederatedIdentityCredentialValue (corp tenant signature),
  falls back to creating the controller MI + MI-as-FIC.
- New TenantRejectedAksOidcIssuer sentinel for clean fallback
  orchestration.
- New discoverAksOidcIssuerUrl helper — accepts explicit args or
  walks kubeconfig + az aks list to guess.
- writeKarsAuthConfig strips MI fields when credentialMode is
  WorkloadIdentity (matches the controller's Optional schema).
- setup-trust gains --credential-mode, --aks-cluster-name,
  --aks-cluster-resource-group, --aks-oidc-issuer-url flags.
- Bicep wrapper threads credentialMode + aksOidcIssuerUrl through
  to ARM.

TESTS:
- Controller: 789/789 (4 new — WI mode render, MI empty-field
  rendering, is_valid_for_mode in MI and WI modes).
- Router: 1014/1014 (unchanged).
- CLI: 786/786 (2 new — credentialMode propagation through
  dry-run with default + explicit).
- Bicep: az bicep build clean.
- Helm lint + template clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller,bicep,docs): security alignment (Phase 5)

Closes the rubber-duck research findings from the Phase 3 critique
and Microsoft's Entra Agent ID design-patterns audit.

CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM):
- KarsSandbox.spec.meshAuth.customSecurityAttributes:
  BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which
  attributes (from a tenant-declared set) the controller should
  PATCH onto each per-sandbox agent identity.
- AgentIdentityClient::patch_custom_security_attributes Graph
  client method. Constructs the documented CustomSecurityAttribute-
  Value envelope with the required @odata.type per attribute,
  inferred from the JSON value shape.
- odata_type_for_value helper: maps serde_json::Value →
  '#String' | '#Int32' | '#Boolean' | '#Collection($T)' and
  rejects floats, nulls, nested objects, mixed-type arrays, and
  empty arrays with clear error messages BEFORE the call goes out.
- Wired into ensure_agent_identity_for_sandbox: PATCH runs on
  every reconcile (idempotent on Graph). Failures surface as
  ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing
  silent missing-attribute drift.

SCALE-OUT INVARIANT (rubber-duck #4):
- Documented in agent_id_provisioning.rs module doc: the agent
  identity is keyed on KarsSandbox.metadata.uid (and cluster UID),
  with NO per-pod / per-replica / per-ordinal dimension. All
  replicas of one KarsSandbox share ONE agent identity.
- New test tag_layout_excludes_per_pod_attributes pins the tag
  layout — any future PR that adds a 'kars-pod-' / 'kars-replica-'
  / 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix
  breaks the test.
- Visibility change: AgentIdentityClient::tags_for is now
  pub(crate) so the cross-module invariant test can call it.

BOOTSTRAP SCRIPTS (deploy/bicep/standalone/):
- custom-security-attributes.sh: declares the recommended
  AgentGovernance set with 4 attributes — AgentClassification
  (Standard|Restricted|Confidential), DataSensitivity
  (Public|Internal|Confidential), ProductOwner, ManagedBy.
  Idempotent via az rest against the Graph beta endpoint.
  (The Microsoft.Graph Bicep extension does not yet ship typed
  attributeSets / customSecurityAttributeDefinitions resources,
  so this is a shell script that operators run once per tenant.)
- conditional-access-baseline.sh: applies Microsoft's
  policy-autonomous-agents template — blocks sign-ins where the
  risk level meets the configured threshold (default: high),
  targeted via the ManagedBy=kars-controller attribute filter.
  Defaults to report-only state for safe rollout. Idempotent via
  upsert.

FOUNDRY RBAC (research R5):
- foundry-rbac.bicep grants Azure AI User on a Foundry resource
  to the BLUEPRINT SP (not per-agent). All derived agent
  identities inherit access — eliminates per-agent role-assignment
  churn. Supports RG-scoped (default) and resource-scoped
  assignment via foundryResourceName parameter. az bicep build
  clean, no warnings.

DOCS:
- docs/architecture/entra-agent-id/05-security-alignment.md:
  6.8KB operator-facing runbook covering the bootstrap order,
  KarsSandbox YAML example, failure modes table, scale-out
  invariant rationale, Foundry RBAC inheritance.

TESTS:
- Controller: 802/802 (+13 new — 12 odata_type_for_value variants
  covering supported + rejected shapes, 1 scale-out invariant).
- Bicep: foundry-rbac.bicep builds clean. JSON compiled.
- Bash: both .sh scripts pass 'bash -n' syntax check.
- CLI: 786/786 (unchanged).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(phase-7): live deploy bug fixes + Foundry RBAC fallback (Phase 7)

End-to-end live validation of the Phase 0-5 shared-sidecar architecture
against real Microsoft corp tenant (kars-aks). Three discoveries
addressed in this commit:

1) ASP.NET Core HostFiltering rejects Service-DNS callers (router fix)
   - Microsoft Entra SDK auth-sidecar's HostFiltering middleware rejects
     non-localhost Host headers regardless of AllowedHosts=* env.
   - inference-router/src/sidecar_client.rs: always send
     Host: localhost:5000 on sidecar requests.
   - deploy/helm/kars/templates/auth-sidecar-deployment.yaml: set
     AllowedHosts=* env for completeness (sidecar config source doesn't
     honour it for dynamic-binding path).
   - NetworkPolicy remains the real ingress boundary; bypassing host
     filter does not weaken security.

2) Azure RBAC inheritance correction (sandbox_bringup.ts)
   - Phase 5 originally assumed Foundry RBAC would inherit blueprint ->
     derived agent identity SPs.
   - Microsoft docs (concept-agent-id-design-patterns) clarify that
     inheritable permissions = Microsoft Graph delegated permissions only.
     Azure RBAC is per-principal.
   - Confirmed empirically live: granting on blueprint SP alone did NOT
     unblock chat; direct grant on each agent identity SP did.
   - cli/src/commands/up/sandbox_bringup.ts: inline Bicep now grants
     Cognitive Services OpenAI User on the blueprint SP as a
     break-glass / fallback (kept for the case where the controller
     can't grant per-agent). Inheritable-permissions language removed.
   - Operator-facing error message documents the per-agent grant path
     when ARM deployment fails for permission reasons.
   - Phase 5b (future PR): controller must assign Azure RBAC per agent.

3) Multi-agent exec-brief demo: chain proven multi-agent
   - Demo applies CRDs, brings up parent execbrief + 3 sub-agents.
   - Each sub-agent gets its own typed Entra agent identity SP.
   - Verified live (with all 4 SPs granted Cognitive Services OpenAI
     User AND Azure AI User on the Foundry account via az rest):
     * All 4 sandboxes booted in fail-closed sidecar mode
     * 65 successful Foundry 200s across all 4 sandboxes in 10 min
     * ZERO PermissionDenied responses after roles in place
     * Mesh routing: parent dispatched to analyst, replied
     * E2E encrypted relay: file transfer (analyst.json) ACKed
     * AGT trust scoring: +0.8 reputation submitted, accepted=true
     * NetworkPolicy: 0 egress denials, 0 ingress drops
   - Demo's verify.json shows 3/9 checks pass — passing checks are the
     kars-runtime mechanisms (sub-agents active, egress 0 denials,
     telegram skipped). 6 failing checks ALL trace to one root cause:
     Foundry project's Bing Grounding connection is not configured
     (foundry_web_search requires bing project_connection_id). This is
     orthogonal to our work; the auth chain is proven.

TESTS:
- inference-router sidecar_client: 39/39 pass.
- CLI: 786 pass + 2 skipped (no regressions).
- Bicep: existing modules build clean.
- Helm: lint clean, template renders correctly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(entra-agent-id): top-level README + migration guide (Phase 8)

Consolidates the entra-agent-id architecture documentation into a
navigable index and adds a migration guide for operators upgrading
from earlier per-pod-sidecar branch heads.

- docs/architecture/entra-agent-id/README.md (new top-level index):
  * TL;DR architecture diagram
  * Pattern A/B selection table
  * Phase ledger linking each commit to its scope
  * Key files map
  * Live validation snapshot (verified on kars-aks 2026-05-28)
- docs/architecture/entra-agent-id/00-poc-archive.md:
  * Previous POC README, renamed to archive. The POC scaffolding
    drove early design but is no longer the canonical reference.
- docs/architecture/entra-agent-id/04-migration-guide.md (new):
  * Step-by-step upgrade path from per-pod-sidecar branch heads
  * Helm upgrade command with the entraSidecar values
  * Per-agent role grant runbook (az rest workaround for the
    CA-blocked az role assignment create path)
  * Rollback procedure
  * Known caveats (HostFiltering, MSAL cache per replica,
    Bing Grounding orthogonal issue)

The numbered ordering reflects deployment lifecycle:
  00 — historical POC (archived)
  01 — runtime token flow
  02 — alternative ACI flow (reference)
  03 — original POC findings
  04 — migration guide (this commit)
  05 — security alignment (Phase 5)

Closes Phase 8 of the feat/entra-agent-id PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): per-agent ARM RBAC + agent identity cleanup (Phase 5b)

Eliminates the manual az role assignment runbook step from the Phase 5
migration guide by having the controller assign and revoke Azure RBAC
roles per agent identity automatically. Closes the orphan SP risk by
adding agent-identity deprovision to the KarsSandbox deletion
finalizer.

CRD (auth_config.rs):
- New `KarsAuthConfig.spec.foundryRbac: Vec<FoundryRbacAssignment>`.
- Each assignment carries an ARM `scope` and a list of built-in role
  definition GUIDs to PUT against every per-sandbox agent identity SP
  at provisioning time.
- Empty list (default) preserves the manual-grant workflow for
  backward compatibility with existing kars deployments.

ARM REST extensions (agent_identity.rs):
- `arm_token()` — MI token for management.azure.com audience.
- `assign_role_to_agent_identity()` — PUT roleAssignment with a
  deterministic UUIDv4 derived from (scope, principal, role), so
  repeated PUTs are idempotent on Azure's side. Treats 200/201 as
  success and 409 RoleAssignmentExists as success.
- `delete_role_assignments_for_principal()` — GET assignments
  filtered by `principalId` (NOT combined with `atScope()` — Azure
  REST returns 400 UnsupportedQuery), narrows to the requested scope
  client-side, then DELETEs each. Used by the deletion finalizer.
- New helpers: `deterministic_assignment_guid()` (SHA-256 → UUIDv4)
  + `extract_subscription_id()` (scope parser).
- 9 new unit tests cover GUID stability, case-insensitivity, UUID
  format, scope parsing happy-path + rejected forms.

Provisioning (agent_id_provisioning.rs):
- After Graph create (step 3c): iterate `spec.foundry_rbac` and assign
  every (scope, role) tuple to the new identity. WARN-and-continue on
  failure so the sandbox still boots; ARM RBAC converges on retry.
- Early-return path (recorded identity): RE-ASSERT the same role
  assignments so existing sandboxes pick up retroactive RBAC config
  the first time the operator adds `foundryRbac` to KarsAuthConfig.
  Idempotent via the deterministic GUID.
- Logs `Phase 5b reconcile: re-asserting ARM role assignments` with
  `foundry_rbac_entries` count for visibility.

Cleanup (agent_id_provisioning.rs + reconciler/mod.rs):
- New `cleanup_agent_identity_for_sandbox()` orchestrates the two
  Azure-side teardown steps on sandbox delete:
    1. DELETE role assignments held by the agent identity SP at every
       configured `foundryRbac` scope.
    2. Graph DELETE the agent identity SP itself.
- Hooked into the KarsSandbox deletion finalizer in
  reconciler/mod.rs. Best-effort: failures logged as WARN but do not
  block finalizer removal, so the K8s CR doesn't get stuck Terminating
  when Azure is degraded. The existing orphan reaper backstops what
  this path misses.

VERIFIED LIVE on kars-aks:
- All 5 existing sandboxes' agent identities (execbrief + 3 sub-agents
  + kars-testrun) now have exactly 2 role assignments each on the
  Foundry account (Cognitive Services OpenAI User + Azure AI User),
  re-asserted by the new code on the early-return path.
- Deleting `kars-testrun` sandbox: controller logged
    "deleted role assignments for principal"  →  2 deleted
    "agent identity SP deprovisioned"
  Verified Foundry now reports 0 role assignments for the deleted
  appId — full cleanup successful.

Tests: 811/811 controller (+9 new for the GUID + scope helpers).
Build clean, helm template + lint clean.

One-time operator prerequisite: the controller's identity (the MI
whose IMDS token the controller uses for ARM calls — typically the
AKS kubelet/agentpool MI) must hold
`Microsoft.Authorization/roleAssignments/write` at every scope listed
in `foundryRbac`. The kars-recommended way is to grant Role Based
Access Control Administrator on the Foundry RG. This is intentionally
a one-time operator action documented in the migration guide rather
than auto-granted, since it crosses cluster/customer trust boundaries.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): scaffold MeshAuthBackend CRD field for Phase 6 design

Adds KarsAuthConfig.spec.meshAuthBackend (enum: Anonymous default,
EntraAgentIdentity opt-in) and optional meshAuthAudience override for
the next milestone — verified Entra-signed AGT mesh peer authentication.

Defaults preserve full backward compatibility (every existing cluster
keeps registering anonymously with no behaviour change). Operators on
clusters that have completed sidecar-based entrypoint mint + relay JWKS
verification (the next-PR work) flip the field to EntraAgentIdentity.

Includes a design doc (docs/architecture/entra-agent-id/06-mesh-trust-design.md)
capturing the three independent pieces that must land for end-to-end
enforcement (CRD scaffold here; sandbox entrypoint via sidecar; AGT
relay JWKS verification — last piece is upstream-coordination work).

Unit tests pin: default is Anonymous (backward compat), both variants
deserialize, unknown variants are rejected (no silent fallback).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(router,sandbox): /v1/mesh-token endpoint + Entra-signed mesh peer path

Phase 6.b — sandbox-side path for verified AGT mesh trust. When
KarsAuthConfig.spec.meshAuthBackend=EntraAgentIdentity, the controller
sets MESH_AUTH_BACKEND on the router, the router exposes
GET /v1/mesh-token, and entrypoint.sh acquires a verified-tier agent
identity token via the shared auth-sidecar instead of doing the legacy
direct Workload-Identity → Entra exchange.

Router changes:
- New inference-router/src/routes/mesh_token.rs route gated on
  MESH_AUTH_BACKEND; returns 404 when disabled, 503 if sidecar not
  configured, 502 on sidecar call failure, 200 with {access_token,
  token_type, expires_in} on success
- Extended sidecar_client::resource_to_service_name to map
  api://agentmesh* → AgentMesh service key
- Optional MESH_AUTH_AUDIENCE env override (defaults to
  api://agentmesh/.default)
- 7 new mesh_token tests + agent-mesh resource-mapping test, all
  serialised on a process-wide ENV_LOCK mutex to avoid env races

Controller changes:
- reconciler/mod.rs injects MESH_AUTH_BACKEND + MESH_AUTH_AUDIENCE env
  on the sandbox's router container when the CRD field is set
- auth_config_reconciler.rs auto-emits the DownstreamApis__AgentMesh__*
  cluster onto the shared sidecar when meshAuthBackend=EntraAgentIdentity
  (skipped if operator has supplied an explicit AgentMesh entry)
- 4 new reconciler tests pin: anonymous default emits nothing, entra
  variant emits expected env, custom audience honoured, operator-
  supplied entry wins

Sandbox changes:
- entrypoint.sh: new branch BEFORE the legacy WI-exchange one. When
  MESH_AUTH_BACKEND=EntraAgentIdentity, curl localhost:8443/v1/mesh-token,
  export AGT_OAUTH_TOKEN on success, force AGT_TRUST_THRESHOLD=0 on
  failure (anonymous-tier fail-open contract matches existing logic)
- Backward compat: when MESH_AUTH_BACKEND is unset (default), the
  entrypoint follows the exact same path as before — no behaviour
  change for existing clusters

Test results:
- 819 controller tests pass (was 815, +4)
- 930 router tests pass (was 923, +7)
- entrypoint.sh sh…
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 31, 2026
…cal-k8s, docker (#368)

* docs(security-validation): cross-platform validation report — AKS, local-k8s, docker

Comprehensive security & runtime validation across all three deploy
modes, mapped to the 9-layer model in docs/security.md.

Per-platform live evidence collected:
- CRD presence + InferencePolicyCompiled/ToolPolicyCompiled/EgressAllowlistCompiled digests
- Pod securityContext (UID, readOnlyRootFilesystem, capabilities, seccomp)
- Workload Identity / Entra Agent ID federated token plumbing (AKS only)
- entra-auth-sidecar token issuance + per-sandbox pinned_agent_id
- Output authenticity (real Foundry calls, real URLs verified via HTTP 200)
- AGT mesh KNOCK + E2E channel establishment per agent
- Native AGT governance modules (PolicyEngine, AuditLogger, etc.)
- NetworkPolicy enforcement + egress-guard caps
- 9-check verify run on every platform

Verified all three platforms 9/9 PASS with current main checks.py.

Findings (no fixes applied — tracked for separate PRs):
  #1 HIGH (docker macOS) — UID 1000 reads Foundry API key (Docker Desktop UID virtualization)
  #2 HIGH (local-k8s) — API key + GitHub token in plaintext pod env (should be secretKeyRef)
  #3 MEDIUM (docker) — NET_ADMIN on persistent parent container (vs init-only in K8s mode)
  #4 MEDIUM (AKS, local-k8s) — blocklist-refresh CronJob failing every 6h (VAP collision)
  #5 MEDIUM (all) — audit JSONL not persisted (RO root FS, no emptyDir mount)
  #6 LOW (AKS) — AllowlistVerified=False on execbrief (inline endpoints, no cosign attestation)
  #7 LOW (AKS) — TrustGraph router-side enforcement not active (documented roadmap)

All findings include a remediation plan in §10. None block the AKS
verified-tier security story documented in docs/security.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>

* docs(security-validation): add env-var inventory addendum for AKS containers

Adds detailed per-container env-var analysis answering the question:
'do AKS containers have more env variables than they should?'

Per-container inventories with full categorization:
- openclaw container: 42 vars in 8 categories
- inference-router container: 54 vars (router-only paths/toggles)

Three additional findings on env-var hygiene:
  #8 LOW — OPENCLAW_GATEWAY_TOKEN exposed via env (should be file mount)
  #9 LOW — enableServiceLinks=true leaks internal cluster IPs (16 env vars)
  #10 LOW — possibly-redundant Foundry/Mesh-auth env vars on openclaw

Headline confirmations:
  ✅ NO AZURE_OPENAI_API_KEY on either AKS container
  ✅ NO COPILOT_GITHUB_TOKEN on either AKS container
  ✅ Federated identity token mounted RO, never as env value
  ✅ Auth mode 'shared entra-auth-sidecar fail-closed, no WI/IMDS/API-key fallback'

Comparison table across all 3 platforms included.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>

---------

Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants