Skip to content

Bump jsonwebtoken from 9.3.1 to 10.3.0 - #1

Merged
Pal Lakatos-Toth (pallakatos) merged 1 commit into
mainfrom
dependabot/cargo/jsonwebtoken-10.3.0
Mar 21, 2026
Merged

Pal Lakatos-Toth (pallakatos) merged 1 commit into
mainfrom
dependabot/cargo/jsonwebtoken-10.3.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Mar 18, 2026

Copy link
Copy Markdown
Contributor

Bumps jsonwebtoken from 9.3.1 to 10.3.0.

Changelog

Sourced from jsonwebtoken's changelog.

10.3.0 (2026-01-27)

  • Export everything needed to define your own CryptoProvider
  • Fix type confusion with exp/nbf when not required

10.2.0 (2025-11-06)

  • Remove Clone bound from decode functions

10.1.0 (2025-10-18)

  • add dangerous::insecure_decode
  • Implement TryFrom &Jwk for DecodingKey

10.0.0 (2025-09-29)

  • BREAKING: now using traits for crypto backends, you have to choose between aws_lc_rs and rust_crypto
  • Add Clone bound to decode
  • Support decoding byte slices
  • Support JWS
Commits

Dependabot compatibility score

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)
    You can disable automated security fix PRs for this repo from the Security Alerts page.

Bumps [jsonwebtoken](https://github.com/Keats/jsonwebtoken) from 9.3.1 to 10.3.0.
- [Changelog](https://github.com/Keats/jsonwebtoken/blob/master/CHANGELOG.md)
- [Commits](Keats/jsonwebtoken@v9.3.1...v10.3.0)

---
updated-dependencies:
- dependency-name: jsonwebtoken
  dependency-version: 10.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added dependencies Pull requests that update a dependency file rust Pull requests that update rust code labels Mar 18, 2026
@pallakatos

Copy link
Copy Markdown
Collaborator

reviewed and approved - we can bump this too

@pallakatos
Pal Lakatos-Toth (pallakatos) merged commit 34f1c0f into main Mar 21, 2026
1 of 7 checks passed
@dependabot
dependabot Bot deleted the dependabot/cargo/jsonwebtoken-10.3.0 branch March 21, 2026 23:12
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request Apr 14, 2026
…wlist

clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill).
openclaw.ai enables `curl | bash` install vectors from inside sandboxes.

Both were in the default Helm values and example CRD. Removed from:
- deploy/helm/azureclaw/values.yaml (default egress allowlist)
- examples/basic-agent/clawsandbox.yaml (example CRD)
- PLAN.md (policy presets documentation)

Defense is now three layers deep:
1. Egress proxy blocks clawhub.com/openclaw.ai (network)
2. Skills directory is root-owned, chmod 640/750 (filesystem)
3. Plugin code is root-owned, read-only for sandbox (code integrity)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Apr 14, 2026
* docs: add global AgentMesh handoff design document

Comprehensive design for agent live migration (local ↔ cloud):
- Identity succession protocol (Ed25519 signed, no key transfer)
- Reclamation protocol (co-signed reverse handoff)
- Sub-agent re-spawn with state injection
- Three-layer handoff endpoint auth (handoff token + no localhost bypass + mutual attestation)
- Security review: 11 threat findings with mitigations
- Handoff trigger security (confirmation token, time delay, AGT policy gate)
- UX design across webchat, TUI, and Telegram
- Demo script and implementation phases (H1-H4)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H1 — handoff module with three-layer auth

Router-side handoff infrastructure for agent live migration (local ↔ cloud):

## New module: handoff.rs (1368 lines)
- HandoffState, SubAgentSnapshot, HandoffMetadata, CredentialRef structs
- HandoffTokenStore: in-memory, TTL-based, one-at-a-time token management
  - 32-byte random tokens, max 10min TTL, constant-time comparison
  - Token hash logged for audit (never the token value)
- HandoffSession: phase tracking across the full handoff lifecycle
  (idle → initialized → draining → snapshotting → transferring → restoring
  → verifying → decommissioning → complete | failed | aborted)
- DrainState: stops new work during handoff, tracks duration
- State serialization: JSON + gzip compression
- State encryption: AES-256-GCM with HKDF-SHA256 key derivation
  Key derived from shared secret + salt using 'azureclaw-handoff-v1' info
- Verification: SHA-256 hash of plaintext for integrity checking
- 21 unit tests covering token store, serialization, encryption, sessions

## New endpoints (8 routes, three auth tiers)
1. POST /agt/handoff/init — admin token only, NO localhost bypass
2. POST /agt/handoff/snapshot — creates encrypted state blob
3. POST /agt/handoff/restore — decrypts, validates, restores state
4. POST /agt/handoff/verify — returns verification digest
5. POST /agt/handoff/drain — enters drain mode
6. POST /agt/handoff/decommission — agent goes dormant
7. POST /agt/handoff/abort — cancels in-progress handoff
8. GET /agt/handoff/status — read-only (localhost allowed)

## Security: three-layer authentication
- Layer 1: Handoff token (one-time, short-lived, CLI-only)
  Token exists only in CLI process memory — never in pod env
- Layer 2: NO localhost bypass for mutation endpoints
  Prevents prompt injection from exfiltrating state via localhost
- Layer 3: Mutual attestation via DH-encrypted state blob
  (Phase H2 adds Ed25519 succession signature verification)

## All endpoints audit-logged with:
- Caller IP, timestamp, endpoint, success/failure
- Token hash (not value), state blob size, item counts

## Dependencies added:
aes-gcm 0.10, hkdf 0.12, sha2 0.10, rand 0.9, base64 0.22, flate2 1

## Test results:
- 77 unit tests pass (21 new handoff tests)
- 26 integration tests pass (updated for new AppState fields)
- 74 controller tests pass (unaffected)
- clippy clean (zero warnings)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H2 — registry mode, identity succession, and reclamation

Phase H2 of the agent handoff feature:

Registry topology (local vs global):
- Add RegistryMode enum to router config (AGT_REGISTRY_MODE env var)
- Handoff init returns 409 in local mode with clear guidance
- Global mode does startup health check on AGT_REGISTRY_URL
- Handoff status endpoint exposes registry_mode + handoff_available

Identity succession (A→B):
- SQL migration 008_succession.sql with succession_log table
- POST /v1/registry/succession endpoint with Ed25519 sig verification
- Canonical message format: succession:{pred}:{succ}:{timestamp}
- One-shot rule (unique index on active predecessor)
- Copies reputation A→B, marks predecessor dormant

Identity reclamation (B→A, co-signed):
- POST /v1/registry/reclamation with dual signature verification
- Original succession ref must match active event_hash
- Deactivates succession redirect, copies reputation back
- Sets original online, departing offline

Lookup follows succession redirects:
- lookup_agent checks succession_log for dormant predecessors
- Returns successor with succeeded_from + succession_hash metadata
- Max redirect depth = 1 (no chains)

Dormant presence status:
- New PresenceStatus::Dormant variant in registry
- Ghost cleanup skips dormant agents (preserves succession chains)
- Capability search excludes dormant agents

CLI --global-registry flag:
- azureclaw dev --global-registry <url> skips local registry stack
- Passes AGT_REGISTRY_MODE=global to router
- Health check on global registry at startup
- Status display shows "handoff enabled" for global mode

Tests: 177 Rust (77 unit + 26 integration + 74 controller) + 159 CLI
All passing, clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H3 — CLI command and plugin tools

CLI command (cli/src/commands/handoff.ts):
- azureclaw handoff <name> --to cloud  (forward: local → AKS)
- azureclaw handoff <name> --to local  (reverse: AKS → local)
- azureclaw handoff <name> --status    (check progress)
- azureclaw handoff <name> --abort     (cancel in-progress)
- 7-step orchestration: verify → init → snapshot → drain → transfer → succession → summary
- Registry mode guard (blocks with clear guidance if not --global-registry)
- Shared secret derivation for snapshot encryption
- Full error recovery (abort on failure)

Plugin tools (cli/src/plugin.ts):
- azureclaw_handoff_status: check handoff availability and state
- azureclaw_handoff_request: LLM can request handoff, but returns CLI
  command (tokens are CLI-only to prevent prompt injection attacks)

Security design:
- Handoff tokens live in CLI process memory only (never LLM-accessible)
- LLM tool returns the CLI command to run, not the token itself
- This prevents prompt injection from triggering unauthorized handoff

Registered in cli.ts under 'Agent mobility' section.

Tests: 177 Rust + 159 CLI, all passing, clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security(handoff): close §9.9 security gaps — confirmation gate, injection protection, rate limits

§9.9.9 — Two-stage confirmation gate (HIGH priority):
- PendingHandoffStore with confirmation token (4-byte random hex)
- POST /agt/handoff/pending — creates pending request (rate limited: 1 per 5min)
- POST /agt/handoff/confirm — validates token, enforces 3s minimum delay
  to prevent LLM self-confirm, consumes pending on success
- Plugin azureclaw_handoff_request now calls /pending (Stage 1)
- New plugin azureclaw_handoff_confirm tool (Stage 2)
- 6 unit tests for PendingHandoffStore (create, confirm, delay, rate limit,
  cancel, wrong token)

§9.9.1 — State blob prompt injection protections:
- sanitize_chat_snapshot() strips messages matching 17 injection patterns
  (system prompt override, handoff commands, instruction ignoring)
- User messages always preserved (legitimate user content)
- Non-UTF8 chat snapshots rejected entirely
- Trust scores capped at 750 on restore (cannot import max trust)
- 4 unit tests for chat sanitization

§9.9.4 — State blob size/DoS limits:
- 50MB blob size cap on both snapshot and restore
- MAX_WORKSPACE_FILES (100) and MAX_WORKSPACE_FILE_SIZE (10MB) constants
- PAYLOAD_TOO_LARGE (413) returned on violation

§9.9.3/§9.9.8 — Rate limits:
- Succession rate limit: 1 per AMID per 5 minutes (DB-backed)
- Reclamation rate limit: 1 per AMID per hour (DB-backed)
- check_succession_rate_limit() queries succession_log timestamps

§9.9.9 — AGT policy rule (belt-and-suspenders):
- handoff-tool-approval rule in azureclaw-default.yaml
- type: approval, priority: 75 (higher than tool-allow at 70)
- Requires operator approval for tool:azureclaw_handoff_request:*
  and tool:azureclaw_handoff_confirm:*

Tests: 188 Rust (74 controller + 88 router + 26 integration) + 159 CLI
All passing, clippy clean, registry cargo check clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): implement global registry deployment — Ingress, OAuth, relay auth, CLI

Phase G1 implementation:

- G1a: AGIC Ingress manifest (deploy/agentmesh-ingress.yaml) with
  NetworkPolicy (postgres locked to registry, registry/relay to AppGW),
  Azure-managed TLS, WAF rate limiting, WebSocket support for relay

- G1b: Entra ID OAuth provider added to agentmesh-registry
  (authorize, callback, token validation via Microsoft Graph).
  Existing GitHub + Google providers untouched.

- G1c: Deployment manifest updated with OAuth secret references
  (agentmesh-oauth-credentials), REGISTRY_URL for relay verification

- G1d: CLI 'azureclaw mesh auth' command — generates Ed25519 keypair,
  runs browser-based OAuth flow, stores encrypted identity in
  ~/.azureclaw/mesh-identity.json (AES-256-GCM, machine-bound key).
  Subcommands: auth, status, reset.

- G1e: CLI 'azureclaw up --global-registry <url>' skips local registry
  deployment. '--expose-registry' deploys AGIC Ingress to make this
  cluster's registry the global endpoint. Context persists registry mode.

- G1f: Relay registration verification — after Ed25519 signature check,
  relay calls registry /v1/registry/lookup to confirm AMID is registered.
  Unregistered/revoked agents rejected. Fails open on registry errors
  (avoids cascading failures). Gated by REQUIRE_REGISTRATION=true.

Security: 4-layer auth chain (WAF → Ed25519 → registry check → OAuth).
PostgreSQL never exposed externally (NetworkPolicy enforced).
Private keys encrypted at rest (AES-256-GCM).

Tests: 188 Rust + 159 CLI passing, all clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test+docs(mesh): integration tests and security/architecture documentation

Tests:
- 28 new CLI mesh tests (mesh.test.ts): base58 encoding, Ed25519
  keypair generation, AMID derivation, encrypt/decrypt roundtrip,
  tamper detection, command structure verification
- 3 relay registry verifier tests (registry_verify.rs): disabled
  verifier passthrough, env-based construction, enable logic
  (compile-gated by pre-existing ed25519-dalek API mismatch in relay)
- All 188 Rust + 187 CLI tests passing

Documentation:
- architecture.md: new 'Global Registry Deployment' section — deployment
  modes table, 4-layer auth chain diagram, NetworkPolicy enforcement,
  identity management overview
- security.md: new 'Layer 9: Global Registry & Handoff Security' section
  — relay auth layers table, handoff threat/mitigation matrix,
  NetworkPolicy diagram, identity-at-rest encryption details

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(plugin): gate handoff mutation tools behind AGT_REGISTRY_MODE=global

In local registry mode, only azureclaw_handoff_status is registered.
The request and confirm tools are hidden from the LLM, preventing
unnecessary AGT governance prompts for tools that would 409 anyway.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add promote/demote commands for registry global mode

azureclaw mesh promote — deploys AGIC Ingress + NetworkPolicies to expose
the cluster's AgentMesh registry and relay as public endpoints. Updates
deployment context to global mode.

azureclaw mesh demote — removes Ingress resources and reverts to
cluster-local registry. Disables cross-environment handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add --allow-ip to promote for IP-based access control

mesh promote auto-detects your public IP (via ifconfig.me) and injects
the AGIC whitelist-source-range annotation into both Ingress resources.
Override with --allow-ip <cidr>. If detection fails, warns and leaves
the registry open.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): auto-detect AppGW IP and use sslip.io for zero-config DNS

mesh promote now queries the AGIC Application Gateway for its public IP
and generates sslip.io hostnames (e.g. registry.20-30-40-50.sslip.io).
No DNS setup needed for testing. TLS is disabled for sslip.io domains
(secured by IP allowlist instead). Use --domain for custom domains.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* refactor(mesh): switch promote/demote to LoadBalancer Services

Replace Ingress-based approach with direct LoadBalancer Service patching.
No ingress controller needed. promote patches registry + relay services
to LoadBalancer with loadBalancerSourceRanges for IP restriction, waits
for external IPs, builds sslip.io URLs. demote reverts to ClusterIP.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(plugin): LLM-driven handoff orchestration via E2E mesh

Replace the CLI-only handoff confirm flow with full LLM-driven
orchestration. After the user confirms the handoff code, the plugin
now executes the entire transfer autonomously:

1. Confirm → router creates handoff token (stays in plugin memory)
2. Snapshot → encrypted AES-256-GCM state blob
3. Drain → stop accepting new work
4. Spawn → create cloud target on AKS (or find existing)
5. Transfer → send state blob via E2E encrypted mesh (Signal Protocol)
6. Verify → target restores, sends verification digest back via mesh
7. Succession → registry identity chain update
8. Decommission → local agent enters dormant state

Key changes:
- handoff_confirm tool: full orchestration instead of returning CLI cmd
- onMessage handler: new handoff_transfer message type for target agent
  to auto-restore state and send verification back
- _routerCallStrict: new helper that rejects on HTTP >= 400
- _readAdminToken: reads admin token from filesystem paths
- _routerCall: added extraHeaders parameter (backward compatible)
- agtReconnect: disconnect before connect to clear stale SDK state

Security model (§9.9.9): the LLM can REQUEST a handoff but never
EXECUTE one. The handoff token stays in plugin memory — the LLM
never sees it. All router calls use this token. Human confirmation
via the 2-stage code flow is the gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): promote checks health and reconnects stale port-forwards

When registry is already in global mode, 'azureclaw mesh promote' now:
- Checks registry HTTP health (/v1/health)
- Checks relay TCP connectivity
- If both healthy: reports status and exits
- If either dead: kills stale PIDs, clears held ports, restarts
  fresh port-forward tunnels, verifies connectivity

Previously it just said 'already global' and exited, even when the
port-forwards had died (e.g. after IP change or sleep).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(plugin): use async import for fs in ESM context

_readAdminToken used require('node:fs') which is unavailable in ESM.
Changed to async function with await import('node:fs') and updated
both call sites to await the result.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): spawn AKS pod from dev mode for cloud handoff

In dev mode, the router's /sandbox/spawn endpoint was creating Docker
containers. For handoff (local→cloud), we need actual AKS pods.

Changes:
- Add HandoffMeta struct to SpawnRequest (mode + predecessor fields)
- When handoff.mode='restore' in dev mode, bypass Docker path and use
  K8s CRD creation via kube-rs (kubeconfig mounted from host)
- Mount ~/.kube/config into dev container at /run/secrets/kubeconfig
  so the router can reach the K8s API for handoff spawns

The controller already sets AGT_RELAY_URL and AGT_REGISTRY_URL on
spawned pods, and NetworkPolicy allows mesh egress — so the handoff
target automatically joins the global mesh.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): propagate trusted_peers and registry_mode to spawned pods

The handoff target was rejecting the source's KNOCK (trust score 0 <
threshold 500) because AGT_TRUSTED_PEERS wasn't propagated. Also,
AGT_REGISTRY_MODE wasn't set, so handoff tools were skipped.

Changes:
- CRD: add trusted_peers and registry_mode fields to GovernanceConfig
- Controller: propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to
  the openclaw container env vars
- Spawn: write trusted_peers and registry_mode='global' into CRD
  governance spec for handoff targets
- Spawn: use 'handoff'/'predecessor' labels instead of 'agent'/'parent'
  for handoff-spawned CRDs (not sub-agents)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add cloud handoff flow diagram (section 11)

Sequence diagram covering all 5 phases: two-stage confirm, snapshot/drain,
spawn on AKS, E2E mesh transfer, succession/decommission. Includes security
model diagram and current vs future (Entra OAuth) trust flow comparison.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): async orchestration with real-time progress tracking

Refactor handoff_confirm to return immediately and run orchestration
in the background via _runHandoffOrchestration(). The LLM polls
handoff_status every 3-5s and relays emoji step updates to the user
in real-time instead of blocking for 2-3 minutes.

Key changes:
- HandoffProgress interface tracks phase, steps[], status, error
- _hp() helper updates progress + logs at each step
- _runHandoffOrchestration() contains the full 7-step flow:
  snapshot → drain → spawn → mesh-wait → transfer → verify →
  succession → decommission
- handoff_status returns rich progress with active polling instruction
- Module-level _log set during register() for background access

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address security and reliability findings from handoff audit

sec-6: Fix message filter AND→OR — verification now rejects messages
       unless BOTH from_amid AND from_agent match the expected target
sec-1: Propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to router
       container (was only on openclaw container)
sec-2: Validate trusted_peers — reject values with control chars
sec-3: Validate registry_mode — only accept 'local'|'global'
rel-7: Wrap _runHandoffOrchestration in top-level try-catch
rel-6: Replace non-null assertions with explicit guard at completion
rel-3: Bump snapshot timeout 15s→60s, drain timeout 15s→30s

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: revise handoff flow diagrams with full review findings

Replace Section 11 with 7 comprehensive diagrams:
- 11.1: End-to-end sequence (both source + target sides, live progress)
- 11.2: Handoff state machine with known limitations noted
- 11.3: 7-layer security model (gate → isolation → auth → encryption →
        injection hardening → identity → infrastructure)
- 11.4: Env var propagation showing both containers receive vars
- 11.5: Two orchestration paths (LLM vs CLI) and their differences
- 11.6: Trust flow (current unauthenticated vs future Entra OAuth)
- 11.7: Error recovery and planned improvements

Also add nohup.out to .gitignore (stale port-forward logs).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address remaining TS findings — orphan cleanup, guards, tests

rel-1: Clean up orphaned CRDs on abort when mesh transfer or discovery
       fails after spawn (DELETE /sandbox/spawn/<name>)
rel-8: Concurrent handoff guard — reject confirm if handoff already running
sec-4: Respect $KUBECONFIG env var with fallback to ~/.kube/config
test-3: Fix 4 failing spawn error tests — tools handle unreachable router
        gracefully (return status JSON), update assertions accordingly
        (187/187 tests now pass)
dup-1: Document dual orchestration paths (CLI operator-mode vs plugin
       LLM-mode) in handoff.ts header comment + architecture-diagrams.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address Rust findings — state machine, resume, tests

rel-2: Add try_transition() to enforce handoff phase ordering
       (Idle→Init→Snapshot→Drain→Transfer→Restore→Verify→Decom→Complete)
rel-4: Add resume() + POST /agt/handoff/resume endpoint to cancel
       drain state after abort (Aborted|Draining → Idle)
rel-5: Change body.unwrap() to expect() with message in spawn.rs
dead-1: Remove unused _used field from ActiveToken
test-1: Add 10 state machine transition tests (valid sequence,
        invalid skip, abort, fail, resume, restart after complete)
test-2: Add 4 auth token tests (wrong value, no active, after
        revoke, wrong pending confirmation code)

201 Rust tests pass (74 controller + 101 router + 26 integration)
Clippy clean with -D warnings

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* handoff: incremental progress polling, SDK reconnect fix, policy cleanup

- handoff_status tool: add since_step param for incremental polling,
  returns only new_steps since last call so LLM relays one step at a time
- vendor SDK patch #9: AgentMeshClient.connect() no longer sets
  connected=true when transport.connect() returns false, allowing retry
- Policy: remove handoff-tool-approval gate (two-step confirmation code
  mechanism is sufficient; approval gate can return once native UI exists)
- config: add promoteMode to DeploymentContext for mesh promote tracking

All tests pass: 201 Rust (74+101+26), 187 CLI

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(crd): add trustedPeers and registryMode to Helm CRD schema

The K8s API server was silently stripping these fields because the Helm
CRD template only defined enabled/toolPolicy/trustThreshold. spawn.rs
wrote the fields and reconciler.rs read them, but the schema validation
layer dropped them in between.

Root cause of handoff mesh registration failure: target pods never
received AGT_TRUSTED_PEERS or AGT_REGISTRY_MODE env vars because the
CRD never stored the values.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(router): read admin token from correct mount path

AppState::new() only checked /run/secrets/admin-token but the controller
mounts the secret at /etc/azureclaw/secrets/admin-token. This caused
'Server misconfiguration: no admin token' on target pods during handoff
verification. main.rs had the correct path but its token was only used
for the admin_auth_middleware, not the handoff middleware which reads
from state.admin_token.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): pre-build audit — snapshot strict, direction propagation, relay URLs

Three defensive fixes from comprehensive flow audit:

1. Snapshot endpoint now uses _routerCallStrict (was _routerCall) —
   if snapshot fails, error surfaces immediately instead of continuing
   with undefined blob data

2. Target-side handoff_transfer handler now reads direction from the
   mesh message instead of hardcoding 'local_to_aks' — enables
   reverse (aks_to_local) handoffs

3. Controller propagates AGT_RELAY_URL and AGT_REGISTRY_URL to the
   openclaw container (was only on router container) — plugin no
   longer relies on fallback to router proxy for relay connection

All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): enforce state machine — migrate all handlers to try_transition

All 5 handoff route handlers now use try_transition() instead of
set_phase(), returning 409 Conflict on invalid phase transitions.

Also fixed the transition rules: Decommissioning is now allowed from
Draining (source-side flow: Init→Snapshot→Drain→Decommission skips
Verify/Restore which happen on the target router, not the source).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): verification hash mismatch + synchronous progress

Two fixes:

1. Verification hash mismatch: verify endpoint was rebuilding a fresh
   snapshot (new timestamp, nonce, hostname) instead of using the hash
   from the restored data. Now restore stores the hash of the decrypted
   compressed bytes, and verify reuses it. Falls back to build_snapshot
   for source-side verify (where no restore happened).

2. No proactive progress: LLMs don't autonomously poll tools, so the
   handoff_confirm tool now awaits _runHandoffOrchestration() and
   returns all steps when complete, instead of firing-and-forgetting
   and expecting the LLM to poll handoff_status.

All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(spawn): check Docker API HTTP status codes in docker_api

The docker_api helper only checked curl's exit code, not the HTTP
response status from Docker Engine. This caused silent failures —
e.g. container start returning HTTP 404 (network not found) was
swallowed and spawn reported success even though the container
never started.

Add -w flag to capture HTTP status code and return Err for 4xx/5xx
responses with the Docker error message extracted from the JSON body.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): propagate channel credentials to cloud target

During handoff spawn, collect channel/plugin credentials from the
source environment (TELEGRAM_BOT_TOKEN, SLACK_BOT_TOKEN, etc.) and
create a {name}-credentials K8s secret in the target namespace.
The controller already mounts this secret via envFrom (optional),
so the cloud agent inherits Telegram and other channels.

Also fix docker_api to check HTTP status codes — previously it only
checked curl's exit code, silently swallowing Docker Engine errors
like 'network not found' (HTTP 404).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): full state hydration — workspace, memory, conversations, Telegram

Source side (orchestration):
- Pack workspace tar from /sandbox/.openclaw/ (AGENTS.md, SOUL.md, etc.)
- Search Foundry Memory for recent context and include as chat_snapshot
- Include credential refs (channel/plugin names) in snapshot

Target router (restore):
- Extract workspace tar to /sandbox/ with path traversal protection
- Write chat_snapshot and metadata to /tmp/handoff/ for plugin

Target plugin (post-restore hydration):
- Create Foundry Conversation with replayed chat messages
- Store handoff event fact in Foundry Memory (update_memories)
- Write HANDOFF_CONTEXT.md to workspace (fallback context)
- Send 'handoff_ready' mesh message back to predecessor

The cloud agent now comes online with full context: workspace files,
conversation history in Foundry Conversations, semantic memory via
shared Memory Store, and proactively greets the user via Telegram.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: cross-container handoff — return state in response, not filesystem

The router and openclaw containers have separate filesystems in AKS.
Previously the router wrote workspace tar and chat snapshot to /tmp/
which the plugin couldn't read.

Changes:
- Router: return workspace_tar (base64) and chat_snapshot in restore
  response JSON instead of writing to filesystem
- Plugin: extract workspace tar and parse chat snapshot from the
  HTTP response body (runs in openclaw container where /sandbox/ lives)
- Remove unused extract_workspace_tar() from routes.rs
- Remove tar crate dependency (extraction now done by plugin via CLI tar)
- Add direction, initiated_at, restored_at to restore response

Future: workspace payloads >5MB will auto-transfer via Azure Blob
Storage (SAS URL in handoff state) — not yet implemented.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security: harden workspace tar extraction and chat snapshot parsing

Tar extraction:
- Pre-extract validation: list entries, reject any containing '..' or
  starting with '/' (path traversal)
- Size guard: reject compressed payloads >5MB (decompression bomb)
- Unique temp dir per extraction (race condition prevention)
- --no-same-owner --no-overwrite-dir flags on extraction
- Temp dir cleaned up after extraction
- Removed '|| true' — errors now surface in logs

Chat snapshot:
- Schema validation: must be array, each entry must have string
  role + content
- Cap at 100 messages, role capped at 20 chars, content at 10k chars
- Rejects non-conforming entries silently

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* perf: split sandbox Dockerfile into base + overlay for fast rebuilds

The sandbox image was 4.96 GB with ~4.0 GB of rarely-changing deps
(OpenClaw, Python wheels, Go tools, Node.js, CLI tools) rebuilt on
every code change. Now split into:

Dockerfile.base (~4.0 GB, rebuild weekly/on dep upgrade):
  - Azure Linux 3 + system packages
  - Node.js 22, Python 3 + 41 packages, Go CLI tools
  - OpenClaw framework + extension symlinks + skills
  - gh, ripgrep, 1password, himalaya
  - User setup (sandbox:1000, router:1001)

Dockerfile (~50 MB, rebuild per commit in ~30s):
  - FROM azureclaw-sandbox-base (all heavy deps pre-cached)
  - CLI plugin builder reuses base image (has Node.js already)
  - Router binary, plugin dist, vendored SDK overlay
  - Entrypoint, proxy-bootstrap, skills, policies

All functionality preserved:
  - UID separation, iptables egress guard, seccomp, read-only rootfs
  - Channel plugins (Telegram/Slack/Discord/WhatsApp)
  - Extension dep symlinks (grammy, carbon, bolt, etc.)
  - Vendored SDK overlay, proxy-bootstrap, Control UI symlink
  - ClawHub skills, npm CLI tools (clawhub, mcporter, oracle)

Build paths updated:
  - azureclaw dev: auto-builds base if not cached, --build-base to force
  - azureclaw push: --only sandbox-base to push base image
  - Makefile: image-sandbox-base target added

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement reverse handoff (cloud → local)

CLI-driven reverse handoff orchestration:
- aksRouterExec: kubectl port-forward to AKS pod router
- wakeDormantDocker: detect and restart stopped containers
- readAksCrdSpec: inherit model/egress/isolation from CRD
- rehydrateCredentials: copy K8s secrets to Docker container
- Full 10-step reverse flow: connect → verify → init → snapshot →
  drain → wake local → credentials → restore → succession →
  decommission + delete CRD

Direction-aware source routing:
- sourceExec alias delegates to routerExec (Docker) or
  aksRouterExec (AKS) based on direction
- Forward path unchanged at runtime (sourceExec === routerExec)

Plugin reverse handoff:
- handoff_request returns CLI command for aks_to_local direction
- Completion messages updated for both directions
- Decommission label direction-aware

Operator TUI:
- 'returning' handoff state for active aks_to_local handoff
- Table shows '<' icon and 'Returning' status
- ASCII-only table icons for reliable column alignment
- Column widths tightened

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(router): raise body limit on handoff routes to 50MB

Axum's default body limit is 2MB. Encrypted state snapshots easily
exceed this, causing HTTP 413 on /agt/handoff/snapshot and /restore.

Add DefaultBodyLimit::max(MAX_BLOB_SIZE_BYTES) layer to
handoff_protected_routes — matches the existing 50MB blob size
constant from §9.9.4.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): read AKS admin token from mounted secret

On AKS, the admin token is stored in K8s secret 'router-admin-token'
and mounted at /etc/azureclaw/secrets/admin-token — not as env var
or /tmp file. Updated getAksAdminToken() to:
1. Read from /etc/azureclaw/secrets/admin-token (router container)
2. Fallback: same path in openclaw container
3. Fallback: kubectl get secret (base64 decode)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): use POST for snapshot route (was GET → 405)

The /agt/handoff/snapshot route is POST-only but both forward and
reverse handoff paths were sending GET requests, causing HTTP 405.
Changed both to POST.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): reuse existing snapshot blob in reverse path

The reverse handoff was requesting a second snapshot at step 9, but
the state machine had already advanced to 'draining' after step 5.
The snapshot blob was already captured at step 3 — now the reverse
path uses snapshotResp.body.blob directly instead of re-fetching.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): pipe restore payload via stdin for large blobs

routerExec passes JSON as a curl -d argument, which hits shell
argument length limits for large encrypted snapshots. The reverse
handoff restore now uses 'docker exec -i ... curl -d @-' with the
payload piped via stdin, avoiding ARG_MAX issues.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(connect): handle Ctrl+C to disconnect port-forward

The kubectl port-forward child process with stdio:pipe did not
receive SIGINT from the terminal. Added explicit SIGINT/SIGTERM
handlers that terminate the child process and exit cleanly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): init local handoff session + auth headers for restore

The local Docker router requires admin token + handoff token for
/agt/handoff/restore. The reverse path now:
1. Gets local admin token from Docker container
2. Inits a handoff session on the local router
3. Passes both auth headers to the restore curl call

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): correct step count + Telegram notification on handoff

- Fixed reverse handoff step counter: 13 steps (was 10, showing 11/10+)
- Added Telegram notification on handoff completion for both directions:
  - local→cloud: 'I moved to the cloud'
  - cloud→local: 'I am back on your local machine'
  Best-effort — reads credentials from Docker container (reverse) or
  env (forward), sends via Telegram Bot API. Failures are silent.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): clean up local handoff state after reverse restore

Two fixes for stale handoff sessions blocking subsequent handoffs:

1. CLI: After successful reverse restore, transition local router
   through verify → decommission to reach a terminal state.

2. Router: Expand can_start() to allow re-init from Restoring,
   Verifying, and Decommissioning phases. These indicate a previous
   handoff that completed data transfer but wasn't properly finalized.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): extend verification timeout + retry mesh send

The verification timeout was 60s but AKS pods can take longer to
fully initialize their plugin message handlers after mesh registration.
The blob sent before the handler is ready gets silently dropped.

Fix:
- Extended timeout from 60s to 180s
- Re-send the handoff_transfer blob every 30s within the verification
  loop, in case the target's message handler wasn't ready on first send
- Also fixes: router can_start() allows stale Restoring/Verifying states,
  local handoff session cleaned up after reverse restore

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(relay): increase max_message_size to 1MB for handoff blobs

Handoff snapshots with real state (chat, audit, credentials) can be
80+ KB. After Signal Protocol encryption + base64 + JSON envelope,
they exceed the relay's 64KB default max_message_size. Bumped to 1MB.

Also includes verification timeout extension (60s→180s) with 30s
re-send retries, already committed in plugin.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: collect workspace/chat/credentials in CLI handoff snapshot

The CLI handoff command was sending an empty snapshot payload
(only shared_secret). The forward handoff via plugin.ts collected
workspace tar, Foundry memories, and credential refs — but the
CLI path (used for both forward and reverse) skipped this.

Changes:
- Collect workspace tar via kubectl exec (AKS) or docker exec (local)
- Collect Foundry Memory Store items as chat context
- Collect credential refs from container environment
- Fix snapshot response field name: size_bytes → snapshot_size_bytes
- Include items breakdown in snapshot response
- Fix step counter: move transfer step into forward branch only

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: stop stepper spinner and cleanup port-forward after handoff

stepper.step('Handoff summary...') started a spinner that was never
stopped with stepper.done(), keeping the event loop alive and requiring
Ctrl+C. Also aksPortForwardStop() was only called in the error path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: add /agt/handoff/succession router endpoint for Ed25519-signed succession

The registry's succession API requires the predecessor's Ed25519 signature
over a canonical message. The private key lives in the router's Governance
identity — inaccessible to the CLI.

New endpoint POST /agt/handoff/succession on the router:
- Takes {successor_amid, reason} from CLI
- Looks up predecessor (self) AMID from registry
- Looks up successor signing key from registry
- Signs canonical message 'succession:{pred}:{succ}:{timestamp}'
- Submits complete SuccessionRequest to registry
- Returns registry response

CLI + plugin updated to call /agt/handoff/succession instead of
/agt/registry/registry/succession (which lacked signing keys).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: sub-agent handoff — collect, snapshot, and re-spawn during handoff

Sub-agents spawned by a parent agent are now included in handoff state
transfer. The full pipeline:

Collection (source side):
- New GET /agt/handoff/sub-agents endpoint lists active sub-agents
  and reconstructs SpawnRequest from CRD spec (K8s) or container
  labels (Docker dev mode)
- CLI + plugin call this endpoint and inject sub_agent_snapshots
  into the handoff snapshot payload

Re-spawn (target side):
- handoff_restore iterates sub_agent_snapshots after state hydration
- Calls create_sandbox() for each sub-agent with the stored config
- Returns per-sub-agent results (spawned/failed) in restore response
- Audit-logged as handoff:restore:sub-agent

Supporting changes:
- SpawnRequest, HandoffMeta, SubAgentSnapshot: added Clone derive
- Sub-agent results included in restore response as sub_agent_results

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: AMID remapping + sub-agent workspace collection during handoff

Two improvements to sub-agent handoff:

1. AMID remapping: When re-spawning sub-agents on the target, the old
   parent AMID in trusted_peers is replaced with the new parent's AMID.
   This ensures sub-agents trust the new parent for KNOCK handshakes.
   The new parent AMID is looked up from the registry at restore time.

2. Workspace collection: The CLI now exec's into each sub-agent's
   container (kubectl for AKS, docker for local) to collect workspace
   tar before including it in the snapshot. Each sub-agent's workspace
   is capped at 2MB. The plugin path is best-effort without workspace
   (no container exec access from inside the sandbox).

Also adds sub_agents_respawned count to plugin restore metadata.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: collect sub-agent workspace via E2E mesh during handoff

Sub-agents now respond to handoff:workspace_request mesh messages by
tarring their /sandbox/.openclaw/workspace and sending it back via the
E2E encrypted relay. No shared volumes or cross-container exec needed.

Plugin message handler (sub-agent side):
- Receives handoff:workspace_request from parent
- Tars workspace (excludes extensions, node_modules, etc.)
- Sends base64-encoded tar back as handoff:workspace_response
- Falls back to empty response on error (parent doesn't hang)

Plugin handoff orchestration (parent side):
- After fetching sub-agent list from router, discovers each sub-agent
  via registry search to get their AMID
- Sends handoff:workspace_request to each via mesh
- Polls agtInbox for handoff:workspace_response (up to 15s per agent)
- Enriches sub-agent snapshots with workspace tar before creating
  the encrypted handoff snapshot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: expand workspace tar to include cron/policies/agents + mesh 768KB cap

- Add WORKSPACE_TAR_CMD constant in handoff.ts for consistent tar commands
- Include .openclaw/cron, .openclaw/policies, .openclaw/agents in all 5 tar commands
- Refine extension exclusion: only exclude */dist and */node_modules (keep manifests)
- Cap mesh workspace response at 768KB (safe under relay's ~1MB limit)
- Add truncated flag in workspace_response messages

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: chunked mesh transfer for large sub-agent workspaces

Workspaces > 512KB are split into chunks sent via separate mesh messages,
then reassembled on the receiver side. This lifts the practical workspace
transfer limit from ~768KB to ~40MB (80 chunks × 512KB).

Sender (sub-agent):
- Small workspace (≤512KB): single handoff:workspace_response (fast path)
- Large workspace: N handoff:workspace_chunk messages + completion marker
- Max 80 chunks (leaves headroom in relay's 100-message offline queue)

Receiver (parent):
- Collects workspace_chunk messages into a Map keyed by chunk_index
- Reassembles in order when all chunks received or completion marker arrives
- 30s timeout with partial-chunk recovery (uses what was received)
- Backwards compatible: single-message responses still work unchanged

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: unified chunked mesh transport layer + file transfer tool

Implements a general-purpose auto-chunking transport layer for the E2E
encrypted mesh. Any payload exceeding 512KB is transparently split into
chunks with per-chunk SHA-256 integrity verification, then reassembled
on the receiver side before reaching application logic.

Transport layer (meshSend + meshHandleTransportMessage):
- meshSend(): auto-chunks large payloads into mesh:transfer_manifest +
  N mesh:transfer_chunk messages. Small messages pass through directly.
- meshHandleTransportMessage(): intercepts transport messages in onMessage,
  accumulates chunks, verifies SHA-256 hashes, reassembles, and delivers
  the original message to the application layer.
- Per-chunk + manifest-level SHA-256 integrity verification
- 2-minute TTL with automatic stale transfer cleanup
- Max ~40MB per transfer (80 chunks × 512KB)

New tool — azureclaw_mesh_transfer_file:
- Agents can send files to each other via E2E encrypted mesh
- Files up to 30MB supported (auto-chunked transparently)
- Received files auto-saved to /sandbox/.openclaw/workspace/incoming/
- Path traversal protection (must be within /sandbox)

Consumers updated to use unified transport:
- mesh_send tool: auto-chunks large task messages
- Handoff blob transfer: auto-chunks encrypted snapshots
- Sub-agent workspace collection: auto-chunks workspace tars
- Handoff re-send loop: uses meshSend for retransmit

Limits raised:
- Router MAX_BLOB_SIZE_BYTES: 50MB → 200MB (sub-agent workspaces)
- Plugin workspace tar cap: 5MB → 50MB
- CLI workspace tar cap: 5MB → 50MB

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: configurable router URL + fix spawn test timeouts + add transport tests

- Make ROUTER and ROUTER_BASE configurable via AZURECLAW_ROUTER_URL env var
  (defaults to http://127.0.0.1:8443 for backward compat)
- Fix 5 pre-existing spawn tool test timeouts caused by local dev router
  on port 8443 — tests now point to unused port 19876 for immediate
  ECONNREFUSED instead of 45s polling loop
- Fix test assertions to match actual error behavior (plain text errors,
  not JSON, when router is unreachable)
- Add 11 new tests (198 total, up from 187):
  - mesh_transfer_file: registration, schema, path traversal, abs path,
    mesh-not-connected
  - mesh_send: registration, params, error when disconnected
  - handoff_status: registration, returns status JSON
  - AZURECLAW_ROUTER_URL: spawn + spawn_status use configurable URL

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: address 4 direction-specific handoff gaps

1. Plugin reverse: retry with 60s backoff when discovering local target
   (local agent may be waking from dormant state after CLI runs
   wakeDormantDocker)

2. Direction validation: both plugin and router now validate that the
   incoming handoff direction matches the environment (AZURECLAW_DEV_MODE).
   Warn-only on mismatch — doesn't block to avoid false positives.

3. Forward credential rehydration: collect actual credential VALUES
   from Docker (not just refs), create K8s secret on AKS target before
   pod starts so envFrom can mount them. Closes the credential gap
   where forward handoff lost Telegram/Slack/Brave tokens.

4. Symmetric cleanup: reverse handoff now scales deployment to 0
   instead of deleting the CRD. This preserves the sandbox definition
   for instant re-forward handoff while freeing all compute resources.
   Falls back to CRD deletion if scale fails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: graceful sub-agent interrupt during handoff

Add handoff:interrupt protocol so sub-agents can save in-progress work
before their workspace is collected during a handoff.

Plugin path (mesh-based):
- Parent sends handoff:interrupt to all sub-agents concurrently
- Sub-agents set handoffInterruptRequested flag
- processTaskWithTools checks the flag between LLM rounds
- On interrupt: saves task progress to .task-in-progress.json
  (round, messages, last content, original task)
- Sends handoff:interrupt_ack back to parent
- Parent waits up to 10s for acks, then proceeds with workspace collection

CLI path (exec-based):
- CLI writes .handoff-interrupt sentinel file into each sub-agent container
- processTaskWithTools also checks for this file between rounds
- Same progress save behavior (.task-in-progress.json)

Both paths ensure the workspace tar includes the progress checkpoint,
so it survives the handoff and is available on the target for resumption.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: sub-agent workspace injection + task resumption after handoff

Complete the sub-agent handoff lifecycle — after re-spawn on the target,
sub-agents now receive their workspace and resume interrupted work.

Router changes:
- Restore response now includes sub_agent_workspaces array with each
  sub-agent's workspace_tar, task_context, status, and checkpoint

Plugin — target side (post-restore):
- Waits up to 60s for each re-spawned sub-agent to register in mesh
- Sends workspace tar via meshSend (auto-chunked for large workspaces)
- Sends handoff:resume with task context and checkpoint info

Plugin — sub-agent side (new message handlers):
- handoff:workspace_inject: extracts received workspace tar into /sandbox/
  with path traversal validation and size guard
- handoff:resume: reads .task-in-progress.json, sends resume_ack to parent
  with status report ('Successfully restored in cloud. Resuming interrupted
  work from round N: <task>'), then re-enters processTaskWithTools with a
  contextual prompt that includes the original task, progress, and last output

The full sub-agent handoff lifecycle is now:
  interrupt → save progress → collect workspace → transfer →
  re-spawn → inject workspace → resume task → report to parent

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent handoff — Docker API encoding + name prefix + source cleanup

Three bugs prevented sub-agents from properly migrating during handoff:

1. collect_sub_agent_snapshots_docker passed raw JSON braces in the
   Docker API URL — curl treated {} as glob patterns and the Docker
   daemon couldn't parse the filter. Result: always returned 0
   sub-agents, so the snapshot blob had no sub-agents to respawn.
   Now uses URL-encoded filter via the docker_api() helper (matching
   list_sandboxes_docker's pattern).

2. Same function used the Docker container name (azureclaw-{name})
   as the agent name, causing respawn to create
   azureclaw-azureclaw-{name} on the target. Now strips the prefix.

3. Source decommission only put the main agent dormant — sub-agent
   containers kept running. Now destroys all source sub-agents via
   the spawn API before decommissioning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent handoff — wrong API URLs + missing report steps

Four fixes for sub-agent handoff orchestration:

1. Sub-agent list used GET /sandbox/spawn (wrong) — now GET /sandbox/list
2. Sub-agent delete used DELETE /sandbox/spawn/{name} (wrong) — now
   DELETE /sandbox/{name} (matches actual router routes)
3. Response field was 'sub_agents' but endpoint returns 'sandboxes'
4. No sub-agent info in handoff progress report — added _hp() calls
   for: discovery count, interrupt/checkpoint status, workspace
   collection count, snapshot inclusion, cleanup status, and
   sub_agents_transferred in the final result object

Also fixed orphan target cleanup URLs (2 places) that had the same
/sandbox/spawn/{name} → /sandbox/{name} issue.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: set trusted_peers for re-spawned sub-agents after handoff

The trusted_peers remapping in handoff_restore was dead code —
Docker snapshots had trusted_peers=None, so the if-let guard never
fired. Re-spawned sub-agents had no trusted parent AMID and rejected
handoff:workspace_inject + handoff:resume messages from the new
parent. This caused workspace injection to silently fail and task
resumption to never trigger.

Fix: always set trusted_peers to include the new parent's AMID,
regardless of whether the original snapshot had it set:
- If peers existed: remap old parent → new parent (existing logic)
- If peers existed but old parent absent: append new parent
- If peers was None: set to new parent entry (new case)

This ensures the sub-agent trusts the new parent on first KNOCK
and accepts workspace/resume messages immediately after spawn.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: include sub-agent status in Telegram greeting after handoff

Restructure the restore flow so sub-agent workspace injection and resume
happens BEFORE the Telegram greeting. After sending resume signals, wait
up to 8s for sub-agents to send resume_ack messages. The Telegram greeting
now includes a sub-agent status section showing each agent's name, state
(resumed/ready/starting/failed), and a task preview.

The handoff_ready mesh message back to the predecessor also now includes
sub_agents_restored count, sub_agents_resumed count, and per-agent details.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: use per-sandbox runtime for operator exec path, not global devMode

The operator's unified view fetches both Docker and AKS agents, but the
routerExec/fetchAgtQuick/fetchEgressDomains functions used the global
devMode flag to choose between docker-exec and kubectl-exec. This meant
Docker agents were queried via kubectl (which fails — no K8s namespace)
when the operator ran in non-dev mode, and vice versa.

Fix: use sb.runtime === 'docker' per-sandbox instead of the global devMode
flag in all 4 places that exec into agent containers:
- fetchSecurityState (routerExec + k8sCheck)
- fetchEgressDomains (routerCurl)
- fetchAgtQuick (docker/kubectl exec)
- seccomp profile inference

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent trust + workspace logging after handoff

Two fixes for broken parent↔sub-agent communication after handoff:

1. Stale trusted_peers: After handoff, re-spawned sub-agents get new
   AMIDs but the parent's parentTrustedAmids set contains old AMIDs.
   KNOCK handler rejects all incoming sub-agent messages (score=0 <
   threshold=500). Fix: when the plugin discovers a sub-agent's new
   AMID during workspace injection, register it in amidToName,
   nameToAmid, parentTrustedAmids, and push baseline trust to router.

2. Silent from_value failure: serde_json::from_value for sub-agent
   snapshots at POST /handoff/snapshot was wrapped in `if let Ok`
   which silently swallowed deserialization errors, potentially losing
   all sub-agent workspace data. Changed to match with tracing::warn
   that logs the exact error + JSON preview for debugging.

Also adds roundtrip test for SubAgentSnapshot workspace_tar through
the full serialize→compress→encrypt→decrypt→decompress→deserialize
chain, and a test for the JS↔Rust base64 round-trip.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security: remove clawhub.com and openclaw.ai from default egress allowlist

clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill).
openclaw.ai enables `curl | bash` install vectors from inside sandboxes.

Both were in the default Helm values and example CRD. Removed from:
- deploy/helm/azureclaw/values.yaml (default egress allowlist)
- examples/basic-agent/clawsandbox.yaml (example CRD)
- PLAN.md (policy presets documentation)

Defense is now three layers deep:
1. Egress proxy blocks clawhub.com/openclaw.ai (network)
2. Skills directory is root-owned, chmod 640/750 (filesystem)
3. Plugin code is root-owned, read-only for sandbox (code integrity)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: use sub_agent_results as trust+resume loop driver, not sub_agent_workspaces

Root cause: the post-restore trust registration and resume signals were
gated on restoreResp.sub_agent_workspaces (workspace data), which could
be empty even when sub-agents were successfully spawned. This caused the
entire block to be skipped — no trust registration, no resume signals,
no Telegram sub-agent status.

Fix: use restoreResp.sub_agent_results (always populated when sub-agents
spawn) as the primary loop driver. Workspace data is looked up by name
from sub_agent_workspaces as a secondary source.

Also:
- Added workspace_inject_ack: sub-agent confirms extraction success/fail
  with file_count + error back to parent before resume is sent
- Parent waits up to 15s for ack, logs result, passes workspace_delivered
  flag in resume payload
- handoff_ready report now includes sub_agents_workspace_delivered count
- Telegram greeting shows 📦 icon per sub-agent when workspace arrived
- Increased mesh registration wait from 60s to 90s (AKS pods need boot)
- Router logs snapshot details when building sub_agent_workspaces

Tests added:
- Rust: sub_agent_workspaces builder filter (empty/non-empty workspace_tar)
- Rust: full encrypt→decrypt→restore round-trip with 2 sub-agents
- Rust: edge case — empty workspace + empty task_context filtered out
- TS: sub_agent_results drives loop even when sub_agent_workspaces empty
- TS: workspace_inject_ack protocol (success + failure paths)
- TS: handoff_ready includes workspace delivery status
- TS: only spawned sub-agents enter trust loop
- TS: missing sub_agent_results graceful fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sdk): reuse active Signal Protocol session instead of crashing

Vendor patch #10: SessionManager.initiateSession() threw 'Active session
already exists' when the crypto layer had a session established via
incoming KNOCK but client.activeSessions wasn't synced.

This broke mesh_transfer_file and any second send to the same peer.

Fix: return existing session info with reused=true flag instead of
throwing. establishSession detects reuse, syncs activeSessions, and
skips redundant KNOCK/activate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: don't expose handoff confirmation code to LLM, add console.log diagnostics

Security fix: The handoff confirmation token was returned in the tool
result, allowing the LLM to self-confirm without user input. Now the
code is sent directly to Telegram (side-channel) and printed to console
for TUI users — the LLM never sees it.

Changes:
- Remove confirmation_token from azureclaw_handoff_request tool response
- Send code via Telegram sendMessage as a side-channel delivery
- Print code to console.log (visible in kubectl logs, not to LLM)
- Update tool descriptions to emphasize code comes from user input
- Bump CONFIRMATION_MIN_DELAY_SECS from 3s to 8s
- Add console.log diagnostics in post-restore IIFE for handoff debugging
- Add test: tool response must not contain confirmation_token

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: increase sub-agent tool-calling rounds from 10 to 25

10 rounds was too tight for data-heavy tasks — sub-agents hit the
cap and returned truncated results. 25 gives enough room for
multi-step research/collection while still preventing runaway loops.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: increase prekey retry window from 16s to 45s for sub-agent mesh send

Sub-agents take 20-30s+ after pod is Running to upload prekeys
(gateway start → plugin load → SDK init → relay connect → prekey
upload). The previous 8×2s=16s window wasn't enough, causing
'Cannot get prekeys' failures on task dispatch.

Now: 15 attempts × 3s = 45s max wait with clearer hint message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: add file_transfer_ack for verified file delivery

The mesh_transfer_file tool had no delivery confirmation — it reported
'sent' but couldn't verify the file was actually written to disk on the
target agent. This caused silent failures where coordinator-notes.txt
appeared to transfer but never landed.

Now:
- Receiver sends file_transfer_ack with success/saved_to/error
- Sender waits up to 15s for ack
- Tool returns 'delivered' (with path) or 'sent_no_ack' (no confirmation)
- Write is verified with stat after writeFileSync

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: filter protocol messages from mesh_inbox, improve file transfer reliability

mesh_inbox now filters out internal protocol messages (handoff blobs,
acks, workspace inject/resume) so the LLM only sees actual sub-agent
replies. Shows filtered_protocol_messages count for visibility.

file_transfer now retries up to 3 times with ack verification — sends,
waits 15s for file_transfer_ack, retries with 3s backoff if no ack.
Returns 'delivered' with exact path or 'sent_no_ack' after all retries.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: auto-decode file_transfer content in mesh_inbox

file_transfer messages now show decoded text content (or a binary
placeholder) instead of raw base64 blobs. Uses null-byte detection
to distinguish text vs binary files. Text files are fully readable
in the inbox; binary files show filename and size.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: stale AMID cache poisoning breaks post-handoff mesh delivery

Root cause: trust+resume loop found OLD sub-agent AMIDs still in the
registry (Docker containers hadn't timed out yet) and cached them.
New AKS sub-agents registered with different AMIDs 26s later, but
the parent never discovered them — all messages went to dead relay
connections and were silently dropped.

Three-layer fix:
1. Stale AMID rejection: collect original_amid from handoff snapshots,
   reject matching registry results, wait for NEW AMIDs to appear
2. Prekey readiness gate: verify E2E session is established before
   sending workspace_inject (20 attempts × 3s = 60s max)
3. Workspace inject retry: 3 attempts with 20s ack wait each, catches
   send errors and retries instead of fire-and-forget

Also filters protocol messages from mesh_inbox and auto-decodes
file_transfer base64 content so LLM sees readable text.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: write HANDOFF_FILES.md manifest after workspace inject

After extracting the workspace tar, writes a manifest listing restored
user-facing files to /sandbox/.openclaw/workspace/HANDOFF_FILES.md.
This makes injected files discoverable when the agent is asked about
its workspace contents post-handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: bump AGT rate limits for multi-agent handoff

Policy: 120 → 240 max_calls/60s for inference:* actions.
Router: 100 → 200 global req/s, 10 → 20 per-agent req/s.
Handoff with 3+ agents doing workspace inject + resume + relay
traffic was hitting the old limits.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: promote incoming/ files to workspace root after handoff inject

Copies files from incoming/ to the workspace root so the agent sees
them immediately when listing files, without needing to know about
the incoming/ directory convention.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: scale down sub-agent deployments during cloud→local decommission

After scaling the parent to 0, also scale down all sub-agent
deployments that were snapshotted. Prevents orphaned sub-agent
pods running on AKS after reverse handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add bidirectional handoff architecture diagrams and changelog

- New section 12 in architecture-diagrams.md: agent lifecycle across
  handoff, forward/reverse flows with sub-agents, stale AMID cache
  poisoning problem & fix, workspace injection detail
- CHANGELOG: add handoff features, sub-agent support, rate limit bump
- README: add handoff to features list and CLI reference table

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Apr 27, 2026
…compile + helm CRD (S4) (#54)

Phase 2 §8 entry 4. Ships the K8s primitive only — `InferencePolicy` is
NOT a model-router (per §3 non-compete; model selection sits in
Foundry). Sandbox-side budget / guardrail / safety policy CR, compiled
to a JSON ConfigMap that the S7 router-side informer will load into the
existing PolicyEnvelope.

Per user direction 2026-04-27, runtime enforcement substrate stays on
Phase 1: `inference-router::budget::TokenBudgetTracker` (env-fed) for
tokens, Foundry Content Safety + `safety::report_content_flags_to_agt`
→ AGT BehaviorMonitor for safety. AGT-Rust 3.3.0 verified against
`/Users/pallakatos/Private/Repos/agt/agent-governance-toolkit` —
AGT-Python has BudgetTracker, AGT-Rust does not yet; the upstream port
is an S7 decision and is explicitly out of scope here.

Added:
- controller/src/inference_policy.rs — CRD struct + spec sub-types
  (TokenBudget, ContentSafetyFloor, ModelPreference, ModelRef) + status
  reusing mcp_server::LocalObjectRef (4th semantic client).
- controller/src/inference_policy_compile.rs — pure-fn
  compile_to_profile + version_hash, deterministic, key-canonical;
  output shape slots into PolicyEntry.payload, no parallel hot-reload.
- controller/src/inference_policy_reconciler.rs — modeled on S3
  a2a_agent_reconciler. Field manager azureclaw-controller/inferencepolicy
  (distinct per §10.4 #1), finalizer
  azureclaw.azure.com/inferencepolicy-cleanup. Conditions reuse
  status::conditions; closed-set error_class per §15.3.
- 6 CEL admission rules in crd_validations.rs:
  monthlyTokens >= dailyTokens, monthlyTokens >= perRequestTokens,
  contentSafety severity ∈ {Safe,Low,Medium,High},
  modelPreference primary/fallback non-empty provider+deployment,
  appliesTo.action ∈ {chat,responses,image,embeddings,*}.
- deploy/helm/azureclaw/templates/crd-inferencepolicy.yaml — drift-
  checked by helm_inferencepolicy_crd_matches_rust_schema.
- docs/security-audits/2026-04-27-phase2-inferencepolicy-reconciler.md
  — AGT boundary verification, STRIDE, out-of-scope list, two sign-offs.

Tests: +20 (6 compile + 7 reconciler + 5 admission + 2 helm-drift).
Controller suite 193 → 218. Workspace cargo test/fmt/clippy all green.

§14.6: strengthens column 7 (Foundry / M365 integration) — primitive
lands here; runtime consumers wired in S7.

AGT crate pin unchanged: agentmesh = "3.3.0" from crates.io, no fork.
`vendor/` directory untouched.

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Apr 28, 2026
)

* phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only)

S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting.

Replaces the legacy `spec.openclaw` field with a discriminated union
`spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`.
The `kind` discriminator selects which sibling struct is required;
the others must be absent. Mutual exclusion enforced at admission via
Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules
`(self.kind == 'X') == has(self.x)`); controller-side defensive guard
will land alongside the reconciler dispatch in a follow-up commit.

Pre-release simplification: in-place v1alpha1 schema edit. No
v1alpha2 cut, no conversion webhook (no installed base, per plan.md
S10 + S13). One PR = one breaking change.

What this commit delivers
-------------------------
- `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum,
  `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`,
  `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`,
  `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent
  default per rubber-duck #9). `ClawSandboxSpec.runtime` is required
  on the wire; `Default` retained for test ergonomics, returns
  `OpenClaw` with empty config.
- `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option
  field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid
  wiping a populated value via merge patch).
- `controller/src/crd.rs`: 8 new tests — PascalCase wire-format
  guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw,
  per-variant round-trip, BYO contractVersion required-not-default,
  serializer omits absent variants, runtimeKind status absence.
- `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips
  from `["openclaw", ...]` to `["runtime", ...]`. New `runtime`
  block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner
  CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains
  `runtimeKind`. `Runtime` printer column added.
- `controller/src/reconciler/mod.rs`: minimal call-site fix —
  `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the
  build green. Full dispatch refactor (`RuntimeDeploymentPlan` per
  rubber-duck #2/#3) lands in step 2.

What is NOT yet wired (intentional, follow-ups)
-----------------------------------------------
- Reconciler dispatch per `runtime.kind` (single-seam plan struct).
- `RuntimeReady` Condition machinery (folded into
  `build_running_status_patch` + `running_status_matches` per
  rubber-duck #1 to avoid status-merge churn).
- OpenAI Agents / MAF deployment SKIP (must NOT silently use
  `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded +
  AdapterMissing).
- `validate_runtime_shape` controller-side guard.
- Examples / fixtures / CLI templates / convert / from_kagent migration.
- CHANGELOG.md, audit doc.

Verification
------------
- `cargo test --package azureclaw-controller`: 284/284 pass
  (8 new RuntimeSpec tests included).
- `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean.
- `cargo fmt --all`: applied.
- Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime
  block, byo.required = [image, contractVersion], runtimeKind status
  field present, Runtime printer column added.

NOTE: This commit alone is NOT mergeable on its own. Without the
fixture/CLI/example migrations + reconciler dispatch, every existing
`spec.openclaw` manifest in-tree would fail admission. Branch
`phase2-multi-runtime-crd` will accumulate the remaining steps before
the PR opens.

Refs: plan.md S10.A1; rubber-duck critique applied (status merge
risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL
shape #6, contractVersion required #9, container name stays
'openclaw' #4).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures

Completes the in-place v1alpha1 schema migration for the multi-runtime CRD
spine. CRD types + Helm schema + reconciler reader landed in d11d41d; this
finishes the long tail of CRD-emitting / CRD-reading sites the user called
out ("every aspect of the code — including the cloud offload").

Cloud offload (the user's explicit ask):
- controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with
  spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate
  spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw).
- controller/src/reconciler/mod.rs:796: comment text aligned to new path.

CLI emitters:
- add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}.
- convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails
  with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape
  for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape.
- migrate.ts: --image help text references spec.runtime.openclaw.image.
- migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning
  message aligned.
- handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model.

Fixtures + examples (8 yaml files):
- examples/{basic,confidential,telegram}-agent/clawsandbox.yaml
- examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml
- tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml
  (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied
  fixture left unchanged — orthogonal to this migration; static scanner is the
  active enforcement, not CRD validation.)

Tests updated to assert the new shape:
- add.test.ts: 4 assertions
- convert.test.ts: 6 assertion blocks + 1 multi-container test
- from_kagent.test.ts: 4 assertions

Verification:
- cargo test --package azureclaw-controller: 284/284 pass
- cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean
- cli npm test: 435/435 pass + 2 skipped
- cli npm run typecheck: clean
- grep confirms no remaining spec.openclaw emission/read sites; only intentional
  docstring/comment references documenting the legacy shape.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: runtime-aware status surface + AdapterMissing dispatch guard

Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior
two commits (d11d41d CRD spine; 3202d13 emitter migration) by wiring the
runtime kind through the status surface and refusing to deploy a Pod for
runtime kinds whose adapter has not yet shipped.

Status surface
- New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in
  controller/src/status/conditions.rs.
- `build_running_status_patch` / `running_status_matches` /
  `build_overlay_status_patch` / `overlay_status_matches` take
  `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and
  append `RuntimeReady` to the conditions array (True/Reconciled on
  the running path, False/OverlayMode on overlay). Stamping inside the
  existing patch (rather than via a separate patch_status) avoids the
  merge-patch array overwrite that would erase the new Condition and
  re-introduce the resourceVersion-bump reconcile storm — see plan
  S10.A1 rubber-duck #1.
- New `build_runtime_unsupported_status_patch` /
  `runtime_unsupported_status_matches` / `stamp_runtime_unsupported`
  helper trio mirrors the existing degraded_* trio. Stamps Degraded=True
  + Ready=False + RuntimeReady=False, all Reason=AdapterMissing.

Reconciler dispatch
- controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a
  static-str discriminator and explicitly skips namespace/SA/Deployment
  creation when the kind is not OpenClaw. Stamps AdapterMissing and
  returns Action::requeue(300s) BEFORE any K8s-resource builder is
  invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1
  rubber-duck #2).
- Status-patch call sites at :1481-1498 thread the runtime_kind_str.

Tests
- 5 new tests for the AdapterMissing helper trio (stamp shape,
  status-missing/runtime-mismatch idempotency rejection, settled-status
  match, transition-time preservation across repeat patches).
- 4 existing tests updated for new conditions array shape + runtimeKind
  field (Running: 2 conds, Overlay: 4 conds).
- 289/289 controller tests pass (was 284); cargo clippy clean; CLI
  435/435 still green.

Docs
- CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking-
  change marker spanning the three commits.
- docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full
  audit doc with threat model (silent fallthrough, status churn,
  CEL-disabled, BYO contract bypass, convert hard-fail), existing-
  implementation survey, wire-format invariants, test matrix, and
  S10.A2-A5 deferral list.

Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam,
per-variant image/entrypoint/env/agentCode resolution, BYO contract
verifier, validate_runtime_shape defensive guard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic)

Locks the CRD wire shape now for three additional declared-roadmap
runtimes so adding their adapters in a later slice is not a breaking
schema change. The CRD becomes a public roadmap signal: customers can
pin `spec.runtime.kind` today and know the schema won't shift under
them.

Tiering:
  Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3),
                             MicrosoftAgentFramework (S10.A4)
  Tier 2 (placeholder, Phase 3+ adapters):
                             SemanticKernel, LangGraph, Anthropic
  BYO (warn-only contract verifier in S10.A2)

Schema changes
- crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs
  `SemanticKernelConfig` (language: python|dotnet|java),
  `LangGraphConfig` (language: python|typescript), `AnthropicConfig`
  (pythonVersion). All three carry the universal agentCode + entrypoint
  + extraEnv shape — same as OpenAIAgents/MAF.
- helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property
  blocks; kind enum extended in both spec + status surfaces; nested
  AgentCodeRef exactly-one CEL on every variant that carries code.
- reconciler: runtime_kind_str match extended; AdapterMissing message
  enumerates Tier-2 placeholders.

Behavior
- All three Tier-2 kinds short-circuit through the existing
  `stamp_runtime_unsupported` path: Degraded + Ready=False +
  RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths;
  pure schema scaffold.

Tests: 4 new round-trip tests (one per variant + a defaults check that
SkLanguage and LangGraphLanguage default to python). 293/293 controller
tests pass (was 289). Clippy clean.

Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md
updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL
matrix.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Apr 28, 2026
… sub-slice) (#72)

§10.4 #1 ("Server-Side Apply on every emitted object with stable
field managers") landing in sub-slices. This is the first: central
field-manager registry + replacement of every bare-string SSA site.

- New `controller/src/field_managers.rs` — single source of truth.
  CLAWSANDBOX, PAIRING, MESH_PEER, MCP_SERVER, TOOL_POLICY, A2A_AGENT,
  INFERENCE_POLICY, CLAW_MEMORY, CLAW_EVAL constants + ROUTER_RECONCILER,
  PROVIDER_BRIDGE, MESH, RECONCILER. ALL_FIELD_MANAGERS registry +
  4 invariant tests (uniqueness, namespaced-format, no bare-controller,
  legacy-string match).

- 13 sites in `reconciler/mod.rs` and 3 in `pairing*.rs` previously
  used bare `"azureclaw-controller"` — now use namespaced constants.
  All sites had `.force()` so the field-ownership transition is
  transparent on existing clusters.

- 3 sites in `mesh_peer/{offload,pair}.rs` previously used bare
  `"azureclaw-mesh-peer"` — now use the constant `MESH_PEER` whose
  value is the legacy string verbatim (zero migration).

- 6 per-CRD `FIELD_MANAGER` constants in their respective reconciler
  files re-export from the central registry — same string, central
  source of truth.

- `providers::field_managers` preserved as backwards-compat re-export.

Controller tests: 324 → 328 (+4 invariant tests).
Workspace clippy + fmt clean.

Audit: docs/security-audits/2026-04-28-phase2-conditions-ssa-leader.md
S7 sub-slices remaining: B (Conditions matrix), C (leader election +
predicated informers), D (backoff + reconcile-DAG), E (workqueue
metrics), F (VAP/MAP expansion).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 1, 2026
…e 3) (#58)

* feat(ci): add governance workflows from AGT (#1)

Adds security and supply chain governance workflows adapted from
agent-governance-toolkit:

- dependabot.yml: Automated dependency updates for Cargo, Docker, and GitHub Actions
- dependency-review.yml: Block PRs introducing high-severity vulnerabilities
- scorecard.yml: Weekly OpenSSF Scorecard analysis
- secret-scanning.yml: TruffleHog secret detection on push and PR

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): add DID canonicalization for AGT identity parity (Phase 3)

- Add did.ts: deriveCanonicalDid(), parseDid(), normalizeDid(), classifyAddress()
- Extend MeshIdentity interface with computed 'did' field (not persisted)
- Update buildFacade() to derive canonical DID from Ed25519 signing key
- Add DID to connection logging for dual-identity observability
- Add 20 tests covering derivation, parsing, normalization, classification

Canonical format: did:agentmesh:<hex(sha256(pubkey))[:16]>
Accepts AGT TS SDK (did:agentmesh:<id>:<fp>) and Python (did:mesh:*) formats.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 5, 2026
…hments, sibling trust race, final-deliverable rule

Live multi-agent demo (analyst→viz→writer fan-out) surfaced four
distinct breakages that each silently degraded the run while the
parent agent still self-reported success.

1. Image generation 404 (router URL prefix regression)

   Foundry's account-scoped /openai/v1/images/generations endpoint
   does NOT accept the /api/projects/<project>/ URL prefix that
   chat-completions tolerates. Commit c13f302 unified everything
   through that prefix. In dev (raw Azure OpenAI account, no project
   path) it works; in AKS prod against a Foundry project endpoint
   the upstream returns a fast 404 and image_generation falls back
   to written descriptions.

   Fix: strip /api/projects/<name>/ from the upstream endpoint
   inside the images_generations handler before forwarding. Add
   strip_project_prefix helper + four unit tests (project-stripped,
   trailing-slash variant, AOAI passthrough, no-prefix passthrough).

2. foundry_code_execute drops container files

   The Responses API tool only walked output[].type=='message' for
   text. matplotlib PNGs / CSVs generated by code_interpreter live
   inside Foundry's per-run container and are referenced via
   container_file_citation annotations or { type: 'image' } entries
   in code_interpreter_call.outputs. They were never downloaded, so
   the demo's bar chart silently degraded to ASCII.

   Fix: collect every (container_id, file_id, filename) reference
   from both shapes, GET each via the new /openai/containers/...
   router route (added to foundry_standalone_routes), and write the
   bytes to /sandbox/.openclaw/workspace/. Append the local paths
   to the tool result so downstream tools (mesh_transfer_file,
   file_write) can ship them. Adds routerCallBinary helper for
   binary downloads through the router.

3. Sibling KNOCK race in parallel fan-out

   AGT_TRUSTED_PEERS is baked at spawn time and consumed once at
   sub-agent boot. When the parent spawns analyst → viz → writer in
   sequence, only writer (last) sees all siblings. analyst's
   parentTrustedAmids only contains parent — so when viz or writer
   later try to KNOCK analyst, the trust score is 0 + 0 = 0 and
   the KNOCK is rejected at threshold 500. The demo logs confirm
   only 1 of 3 sibling pairs ever opened a session.

   Fix: after every successful spawn, the parent broadcasts a
   peers_update message containing the new sibling's AMID to every
   already-running sibling. Each sub-agent now records the parent's
   AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and
   handles peers_update only from that AMID, extending its
   parentTrustedAmids set at runtime.

4. Sub-agent system prompt missing FINAL DELIVERABLE rule

   Sub-agents were told they could mesh_transfer_file artifacts to
   peers, but nothing forced them to mesh_transfer_file the FINAL
   artifact back to the parent before returning a summary. The
   writer's executive_brief.md sat in its local /sandbox forever
   while the parent reported success.

   Fix: append a hard rule to the sub-agent system prompt requiring
   mesh_transfer_file(to_agent='parent', ...) as the last action
   before any "task complete" reply, with one call per output file.

Tests
- inference-router: 643 lib tests pass (4 new strip_project_prefix
  tests).
- inference-router: cargo clippy --all-targets clean.
- runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only
  pre-existing warnings.
- cli: 553 tests pass.

Deployment
- For #2: rebuild + push inference-router image.
- For #1, #3, #4: rebuild + push sandbox image.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 11, 2026
Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 11, 2026
…#245)

* feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap

Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): pre-build mesh-plugin for runtime CI + format reconciler

PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): quote workflow step name containing colon

YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): clippy fixes for Rust 1.95.0

CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12

Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5

The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): mark gaps G1 and G2 fixed on local AGT branch

Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.

* dev: add --mesh-provider <vendored|agt> selection with first-run prompt

Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites

Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): wire AGT provider end-to-end (6 stackup bugs)

End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): always route mesh registry through inference-router

azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): break mesh_send infinite poll loop on dead sub-agent

probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): suppress /v1/registry/* 404 leaks in AGT mode

Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode

Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s

The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* runtime: hide Foundry tools in github-copilot mode (same as github-models)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch

Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(up): --mesh-provider=agt picks AGT manifest + flips helm value

Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): plumb --mesh-provider into local-k8s helm install

Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): AGT wire protocol adapter for mesh_peer

Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer

- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(deploy): AGT mesh probe path + Cilium pod-port NP allow

deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh promote): AGT-compat health + WS upgrade paths

azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): first-run picker for local vs remote mesh source

azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container

On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh

Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap

mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call

Previous commit aa7d28e wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* push: auto-detect mesh provider from live helm release

When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* entrypoint: fail-open trust gate when running anonymous tier

When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop

Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable

The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log

CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
* docs: add global AgentMesh handoff design document

Comprehensive design for agent live migration (local ↔ cloud):
- Identity succession protocol (Ed25519 signed, no key transfer)
- Reclamation protocol (co-signed reverse handoff)
- Sub-agent re-spawn with state injection
- Three-layer handoff endpoint auth (handoff token + no localhost bypass + mutual attestation)
- Security review: 11 threat findings with mitigations
- Handoff trigger security (confirmation token, time delay, AGT policy gate)
- UX design across webchat, TUI, and Telegram
- Demo script and implementation phases (H1-H4)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H1 — handoff module with three-layer auth

Router-side handoff infrastructure for agent live migration (local ↔ cloud):

## New module: handoff.rs (1368 lines)
- HandoffState, SubAgentSnapshot, HandoffMetadata, CredentialRef structs
- HandoffTokenStore: in-memory, TTL-based, one-at-a-time token management
  - 32-byte random tokens, max 10min TTL, constant-time comparison
  - Token hash logged for audit (never the token value)
- HandoffSession: phase tracking across the full handoff lifecycle
  (idle → initialized → draining → snapshotting → transferring → restoring
  → verifying → decommissioning → complete | failed | aborted)
- DrainState: stops new work during handoff, tracks duration
- State serialization: JSON + gzip compression
- State encryption: AES-256-GCM with HKDF-SHA256 key derivation
  Key derived from shared secret + salt using 'azureclaw-handoff-v1' info
- Verification: SHA-256 hash of plaintext for integrity checking
- 21 unit tests covering token store, serialization, encryption, sessions

## New endpoints (8 routes, three auth tiers)
1. POST /agt/handoff/init — admin token only, NO localhost bypass
2. POST /agt/handoff/snapshot — creates encrypted state blob
3. POST /agt/handoff/restore — decrypts, validates, restores state
4. POST /agt/handoff/verify — returns verification digest
5. POST /agt/handoff/drain — enters drain mode
6. POST /agt/handoff/decommission — agent goes dormant
7. POST /agt/handoff/abort — cancels in-progress handoff
8. GET /agt/handoff/status — read-only (localhost allowed)

## Security: three-layer authentication
- Layer 1: Handoff token (one-time, short-lived, CLI-only)
  Token exists only in CLI process memory — never in pod env
- Layer 2: NO localhost bypass for mutation endpoints
  Prevents prompt injection from exfiltrating state via localhost
- Layer 3: Mutual attestation via DH-encrypted state blob
  (Phase H2 adds Ed25519 succession signature verification)

## All endpoints audit-logged with:
- Caller IP, timestamp, endpoint, success/failure
- Token hash (not value), state blob size, item counts

## Dependencies added:
aes-gcm 0.10, hkdf 0.12, sha2 0.10, rand 0.9, base64 0.22, flate2 1

## Test results:
- 77 unit tests pass (21 new handoff tests)
- 26 integration tests pass (updated for new AppState fields)
- 74 controller tests pass (unaffected)
- clippy clean (zero warnings)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H2 — registry mode, identity succession, and reclamation

Phase H2 of the agent handoff feature:

Registry topology (local vs global):
- Add RegistryMode enum to router config (AGT_REGISTRY_MODE env var)
- Handoff init returns 409 in local mode with clear guidance
- Global mode does startup health check on AGT_REGISTRY_URL
- Handoff status endpoint exposes registry_mode + handoff_available

Identity succession (A→B):
- SQL migration 008_succession.sql with succession_log table
- POST /v1/registry/succession endpoint with Ed25519 sig verification
- Canonical message format: succession:{pred}:{succ}:{timestamp}
- One-shot rule (unique index on active predecessor)
- Copies reputation A→B, marks predecessor dormant

Identity reclamation (B→A, co-signed):
- POST /v1/registry/reclamation with dual signature verification
- Original succession ref must match active event_hash
- Deactivates succession redirect, copies reputation back
- Sets original online, departing offline

Lookup follows succession redirects:
- lookup_agent checks succession_log for dormant predecessors
- Returns successor with succeeded_from + succession_hash metadata
- Max redirect depth = 1 (no chains)

Dormant presence status:
- New PresenceStatus::Dormant variant in registry
- Ghost cleanup skips dormant agents (preserves succession chains)
- Capability search excludes dormant agents

CLI --global-registry flag:
- azureclaw dev --global-registry <url> skips local registry stack
- Passes AGT_REGISTRY_MODE=global to router
- Health check on global registry at startup
- Status display shows "handoff enabled" for global mode

Tests: 177 Rust (77 unit + 26 integration + 74 controller) + 159 CLI
All passing, clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement Phase H3 — CLI command and plugin tools

CLI command (cli/src/commands/handoff.ts):
- azureclaw handoff <name> --to cloud  (forward: local → AKS)
- azureclaw handoff <name> --to local  (reverse: AKS → local)
- azureclaw handoff <name> --status    (check progress)
- azureclaw handoff <name> --abort     (cancel in-progress)
- 7-step orchestration: verify → init → snapshot → drain → transfer → succession → summary
- Registry mode guard (blocks with clear guidance if not --global-registry)
- Shared secret derivation for snapshot encryption
- Full error recovery (abort on failure)

Plugin tools (cli/src/plugin.ts):
- azureclaw_handoff_status: check handoff availability and state
- azureclaw_handoff_request: LLM can request handoff, but returns CLI
  command (tokens are CLI-only to prevent prompt injection attacks)

Security design:
- Handoff tokens live in CLI process memory only (never LLM-accessible)
- LLM tool returns the CLI command to run, not the token itself
- This prevents prompt injection from triggering unauthorized handoff

Registered in cli.ts under 'Agent mobility' section.

Tests: 177 Rust + 159 CLI, all passing, clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security(handoff): close §9.9 security gaps — confirmation gate, injection protection, rate limits

§9.9.9 — Two-stage confirmation gate (HIGH priority):
- PendingHandoffStore with confirmation token (4-byte random hex)
- POST /agt/handoff/pending — creates pending request (rate limited: 1 per 5min)
- POST /agt/handoff/confirm — validates token, enforces 3s minimum delay
  to prevent LLM self-confirm, consumes pending on success
- Plugin azureclaw_handoff_request now calls /pending (Stage 1)
- New plugin azureclaw_handoff_confirm tool (Stage 2)
- 6 unit tests for PendingHandoffStore (create, confirm, delay, rate limit,
  cancel, wrong token)

§9.9.1 — State blob prompt injection protections:
- sanitize_chat_snapshot() strips messages matching 17 injection patterns
  (system prompt override, handoff commands, instruction ignoring)
- User messages always preserved (legitimate user content)
- Non-UTF8 chat snapshots rejected entirely
- Trust scores capped at 750 on restore (cannot import max trust)
- 4 unit tests for chat sanitization

§9.9.4 — State blob size/DoS limits:
- 50MB blob size cap on both snapshot and restore
- MAX_WORKSPACE_FILES (100) and MAX_WORKSPACE_FILE_SIZE (10MB) constants
- PAYLOAD_TOO_LARGE (413) returned on violation

§9.9.3/§9.9.8 — Rate limits:
- Succession rate limit: 1 per AMID per 5 minutes (DB-backed)
- Reclamation rate limit: 1 per AMID per hour (DB-backed)
- check_succession_rate_limit() queries succession_log timestamps

§9.9.9 — AGT policy rule (belt-and-suspenders):
- handoff-tool-approval rule in azureclaw-default.yaml
- type: approval, priority: 75 (higher than tool-allow at 70)
- Requires operator approval for tool:azureclaw_handoff_request:*
  and tool:azureclaw_handoff_confirm:*

Tests: 188 Rust (74 controller + 88 router + 26 integration) + 159 CLI
All passing, clippy clean, registry cargo check clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): implement global registry deployment — Ingress, OAuth, relay auth, CLI

Phase G1 implementation:

- G1a: AGIC Ingress manifest (deploy/agentmesh-ingress.yaml) with
  NetworkPolicy (postgres locked to registry, registry/relay to AppGW),
  Azure-managed TLS, WAF rate limiting, WebSocket support for relay

- G1b: Entra ID OAuth provider added to agentmesh-registry
  (authorize, callback, token validation via Microsoft Graph).
  Existing GitHub + Google providers untouched.

- G1c: Deployment manifest updated with OAuth secret references
  (agentmesh-oauth-credentials), REGISTRY_URL for relay verification

- G1d: CLI 'azureclaw mesh auth' command — generates Ed25519 keypair,
  runs browser-based OAuth flow, stores encrypted identity in
  ~/.azureclaw/mesh-identity.json (AES-256-GCM, machine-bound key).
  Subcommands: auth, status, reset.

- G1e: CLI 'azureclaw up --global-registry <url>' skips local registry
  deployment. '--expose-registry' deploys AGIC Ingress to make this
  cluster's registry the global endpoint. Context persists registry mode.

- G1f: Relay registration verification — after Ed25519 signature check,
  relay calls registry /v1/registry/lookup to confirm AMID is registered.
  Unregistered/revoked agents rejected. Fails open on registry errors
  (avoids cascading failures). Gated by REQUIRE_REGISTRATION=true.

Security: 4-layer auth chain (WAF → Ed25519 → registry check → OAuth).
PostgreSQL never exposed externally (NetworkPolicy enforced).
Private keys encrypted at rest (AES-256-GCM).

Tests: 188 Rust + 159 CLI passing, all clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test+docs(mesh): integration tests and security/architecture documentation

Tests:
- 28 new CLI mesh tests (mesh.test.ts): base58 encoding, Ed25519
  keypair generation, AMID derivation, encrypt/decrypt roundtrip,
  tamper detection, command structure verification
- 3 relay registry verifier tests (registry_verify.rs): disabled
  verifier passthrough, env-based construction, enable logic
  (compile-gated by pre-existing ed25519-dalek API mismatch in relay)
- All 188 Rust + 187 CLI tests passing

Documentation:
- architecture.md: new 'Global Registry Deployment' section — deployment
  modes table, 4-layer auth chain diagram, NetworkPolicy enforcement,
  identity management overview
- security.md: new 'Layer 9: Global Registry & Handoff Security' section
  — relay auth layers table, handoff threat/mitigation matrix,
  NetworkPolicy diagram, identity-at-rest encryption details

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(plugin): gate handoff mutation tools behind AGT_REGISTRY_MODE=global

In local registry mode, only azureclaw_handoff_status is registered.
The request and confirm tools are hidden from the LLM, preventing
unnecessary AGT governance prompts for tools that would 409 anyway.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add promote/demote commands for registry global mode

azureclaw mesh promote — deploys AGIC Ingress + NetworkPolicies to expose
the cluster's AgentMesh registry and relay as public endpoints. Updates
deployment context to global mode.

azureclaw mesh demote — removes Ingress resources and reverts to
cluster-local registry. Disables cross-environment handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add --allow-ip to promote for IP-based access control

mesh promote auto-detects your public IP (via ifconfig.me) and injects
the AGIC whitelist-source-range annotation into both Ingress resources.
Override with --allow-ip <cidr>. If detection fails, warns and leaves
the registry open.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): auto-detect AppGW IP and use sslip.io for zero-config DNS

mesh promote now queries the AGIC Application Gateway for its public IP
and generates sslip.io hostnames (e.g. registry.20-30-40-50.sslip.io).
No DNS setup needed for testing. TLS is disabled for sslip.io domains
(secured by IP allowlist instead). Use --domain for custom domains.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* refactor(mesh): switch promote/demote to LoadBalancer Services

Replace Ingress-based approach with direct LoadBalancer Service patching.
No ingress controller needed. promote patches registry + relay services
to LoadBalancer with loadBalancerSourceRanges for IP restriction, waits
for external IPs, builds sslip.io URLs. demote reverts to ClusterIP.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(plugin): LLM-driven handoff orchestration via E2E mesh

Replace the CLI-only handoff confirm flow with full LLM-driven
orchestration. After the user confirms the handoff code, the plugin
now executes the entire transfer autonomously:

1. Confirm → router creates handoff token (stays in plugin memory)
2. Snapshot → encrypted AES-256-GCM state blob
3. Drain → stop accepting new work
4. Spawn → create cloud target on AKS (or find existing)
5. Transfer → send state blob via E2E encrypted mesh (Signal Protocol)
6. Verify → target restores, sends verification digest back via mesh
7. Succession → registry identity chain update
8. Decommission → local agent enters dormant state

Key changes:
- handoff_confirm tool: full orchestration instead of returning CLI cmd
- onMessage handler: new handoff_transfer message type for target agent
  to auto-restore state and send verification back
- _routerCallStrict: new helper that rejects on HTTP >= 400
- _readAdminToken: reads admin token from filesystem paths
- _routerCall: added extraHeaders parameter (backward compatible)
- agtReconnect: disconnect before connect to clear stale SDK state

Security model (§9.9.9): the LLM can REQUEST a handoff but never
EXECUTE one. The handoff token stays in plugin memory — the LLM
never sees it. All router calls use this token. Human confirmation
via the 2-stage code flow is the gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): promote checks health and reconnects stale port-forwards

When registry is already in global mode, 'azureclaw mesh promote' now:
- Checks registry HTTP health (/v1/health)
- Checks relay TCP connectivity
- If both healthy: reports status and exits
- If either dead: kills stale PIDs, clears held ports, restarts
  fresh port-forward tunnels, verifies connectivity

Previously it just said 'already global' and exited, even when the
port-forwards had died (e.g. after IP change or sleep).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(plugin): use async import for fs in ESM context

_readAdminToken used require('node:fs') which is unavailable in ESM.
Changed to async function with await import('node:fs') and updated
both call sites to await the result.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): spawn AKS pod from dev mode for cloud handoff

In dev mode, the router's /sandbox/spawn endpoint was creating Docker
containers. For handoff (local→cloud), we need actual AKS pods.

Changes:
- Add HandoffMeta struct to SpawnRequest (mode + predecessor fields)
- When handoff.mode='restore' in dev mode, bypass Docker path and use
  K8s CRD creation via kube-rs (kubeconfig mounted from host)
- Mount ~/.kube/config into dev container at /run/secrets/kubeconfig
  so the router can reach the K8s API for handoff spawns

The controller already sets AGT_RELAY_URL and AGT_REGISTRY_URL on
spawned pods, and NetworkPolicy allows mesh egress — so the handoff
target automatically joins the global mesh.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): propagate trusted_peers and registry_mode to spawned pods

The handoff target was rejecting the source's KNOCK (trust score 0 <
threshold 500) because AGT_TRUSTED_PEERS wasn't propagated. Also,
AGT_REGISTRY_MODE wasn't set, so handoff tools were skipped.

Changes:
- CRD: add trusted_peers and registry_mode fields to GovernanceConfig
- Controller: propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to
  the openclaw container env vars
- Spawn: write trusted_peers and registry_mode='global' into CRD
  governance spec for handoff targets
- Spawn: use 'handoff'/'predecessor' labels instead of 'agent'/'parent'
  for handoff-spawned CRDs (not sub-agents)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add cloud handoff flow diagram (section 11)

Sequence diagram covering all 5 phases: two-stage confirm, snapshot/drain,
spawn on AKS, E2E mesh transfer, succession/decommission. Includes security
model diagram and current vs future (Entra OAuth) trust flow comparison.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): async orchestration with real-time progress tracking

Refactor handoff_confirm to return immediately and run orchestration
in the background via _runHandoffOrchestration(). The LLM polls
handoff_status every 3-5s and relays emoji step updates to the user
in real-time instead of blocking for 2-3 minutes.

Key changes:
- HandoffProgress interface tracks phase, steps[], status, error
- _hp() helper updates progress + logs at each step
- _runHandoffOrchestration() contains the full 7-step flow:
  snapshot → drain → spawn → mesh-wait → transfer → verify →
  succession → decommission
- handoff_status returns rich progress with active polling instruction
- Module-level _log set during register() for background access

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address security and reliability findings from handoff audit

sec-6: Fix message filter AND→OR — verification now rejects messages
       unless BOTH from_amid AND from_agent match the expected target
sec-1: Propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to router
       container (was only on openclaw container)
sec-2: Validate trusted_peers — reject values with control chars
sec-3: Validate registry_mode — only accept 'local'|'global'
rel-7: Wrap _runHandoffOrchestration in top-level try-catch
rel-6: Replace non-null assertions with explicit guard at completion
rel-3: Bump snapshot timeout 15s→60s, drain timeout 15s→30s

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: revise handoff flow diagrams with full review findings

Replace Section 11 with 7 comprehensive diagrams:
- 11.1: End-to-end sequence (both source + target sides, live progress)
- 11.2: Handoff state machine with known limitations noted
- 11.3: 7-layer security model (gate → isolation → auth → encryption →
        injection hardening → identity → infrastructure)
- 11.4: Env var propagation showing both containers receive vars
- 11.5: Two orchestration paths (LLM vs CLI) and their differences
- 11.6: Trust flow (current unauthenticated vs future Entra OAuth)
- 11.7: Error recovery and planned improvements

Also add nohup.out to .gitignore (stale port-forward logs).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address remaining TS findings — orphan cleanup, guards, tests

rel-1: Clean up orphaned CRDs on abort when mesh transfer or discovery
       fails after spawn (DELETE /sandbox/spawn/<name>)
rel-8: Concurrent handoff guard — reject confirm if handoff already running
sec-4: Respect $KUBECONFIG env var with fallback to ~/.kube/config
test-3: Fix 4 failing spawn error tests — tools handle unreachable router
        gracefully (return status JSON), update assertions accordingly
        (187/187 tests now pass)
dup-1: Document dual orchestration paths (CLI operator-mode vs plugin
       LLM-mode) in handoff.ts header comment + architecture-diagrams.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(review): address Rust findings — state machine, resume, tests

rel-2: Add try_transition() to enforce handoff phase ordering
       (Idle→Init→Snapshot→Drain→Transfer→Restore→Verify→Decom→Complete)
rel-4: Add resume() + POST /agt/handoff/resume endpoint to cancel
       drain state after abort (Aborted|Draining → Idle)
rel-5: Change body.unwrap() to expect() with message in spawn.rs
dead-1: Remove unused _used field from ActiveToken
test-1: Add 10 state machine transition tests (valid sequence,
        invalid skip, abort, fail, resume, restart after complete)
test-2: Add 4 auth token tests (wrong value, no active, after
        revoke, wrong pending confirmation code)

201 Rust tests pass (74 controller + 101 router + 26 integration)
Clippy clean with -D warnings

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* handoff: incremental progress polling, SDK reconnect fix, policy cleanup

- handoff_status tool: add since_step param for incremental polling,
  returns only new_steps since last call so LLM relays one step at a time
- vendor SDK patch #9: AgentMeshClient.connect() no longer sets
  connected=true when transport.connect() returns false, allowing retry
- Policy: remove handoff-tool-approval gate (two-step confirmation code
  mechanism is sufficient; approval gate can return once native UI exists)
- config: add promoteMode to DeploymentContext for mesh promote tracking

All tests pass: 201 Rust (74+101+26), 187 CLI

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(crd): add trustedPeers and registryMode to Helm CRD schema

The K8s API server was silently stripping these fields because the Helm
CRD template only defined enabled/toolPolicy/trustThreshold. spawn.rs
wrote the fields and reconciler.rs read them, but the schema validation
layer dropped them in between.

Root cause of handoff mesh registration failure: target pods never
received AGT_TRUSTED_PEERS or AGT_REGISTRY_MODE env vars because the
CRD never stored the values.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(router): read admin token from correct mount path

AppState::new() only checked /run/secrets/admin-token but the controller
mounts the secret at /etc/azureclaw/secrets/admin-token. This caused
'Server misconfiguration: no admin token' on target pods during handoff
verification. main.rs had the correct path but its token was only used
for the admin_auth_middleware, not the handoff middleware which reads
from state.admin_token.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): pre-build audit — snapshot strict, direction propagation, relay URLs

Three defensive fixes from comprehensive flow audit:

1. Snapshot endpoint now uses _routerCallStrict (was _routerCall) —
   if snapshot fails, error surfaces immediately instead of continuing
   with undefined blob data

2. Target-side handoff_transfer handler now reads direction from the
   mesh message instead of hardcoding 'local_to_aks' — enables
   reverse (aks_to_local) handoffs

3. Controller propagates AGT_RELAY_URL and AGT_REGISTRY_URL to the
   openclaw container (was only on router container) — plugin no
   longer relies on fallback to router proxy for relay connection

All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): enforce state machine — migrate all handlers to try_transition

All 5 handoff route handlers now use try_transition() instead of
set_phase(), returning 409 Conflict on invalid phase transitions.

Also fixed the transition rules: Decommissioning is now allowed from
Draining (source-side flow: Init→Snapshot→Drain→Decommission skips
Verify/Restore which happen on the target router, not the source).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): verification hash mismatch + synchronous progress

Two fixes:

1. Verification hash mismatch: verify endpoint was rebuilding a fresh
   snapshot (new timestamp, nonce, hostname) instead of using the hash
   from the restored data. Now restore stores the hash of the decrypted
   compressed bytes, and verify reuses it. Falls back to build_snapshot
   for source-side verify (where no restore happened).

2. No proactive progress: LLMs don't autonomously poll tools, so the
   handoff_confirm tool now awaits _runHandoffOrchestration() and
   returns all steps when complete, instead of firing-and-forgetting
   and expecting the LLM to poll handoff_status.

All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(spawn): check Docker API HTTP status codes in docker_api

The docker_api helper only checked curl's exit code, not the HTTP
response status from Docker Engine. This caused silent failures —
e.g. container start returning HTTP 404 (network not found) was
swallowed and spawn reported success even though the container
never started.

Add -w flag to capture HTTP status code and return Err for 4xx/5xx
responses with the Docker error message extracted from the JSON body.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): propagate channel credentials to cloud target

During handoff spawn, collect channel/plugin credentials from the
source environment (TELEGRAM_BOT_TOKEN, SLACK_BOT_TOKEN, etc.) and
create a {name}-credentials K8s secret in the target namespace.
The controller already mounts this secret via envFrom (optional),
so the cloud agent inherits Telegram and other channels.

Also fix docker_api to check HTTP status codes — previously it only
checked curl's exit code, silently swallowing Docker Engine errors
like 'network not found' (HTTP 404).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): full state hydration — workspace, memory, conversations, Telegram

Source side (orchestration):
- Pack workspace tar from /sandbox/.openclaw/ (AGENTS.md, SOUL.md, etc.)
- Search Foundry Memory for recent context and include as chat_snapshot
- Include credential refs (channel/plugin names) in snapshot

Target router (restore):
- Extract workspace tar to /sandbox/ with path traversal protection
- Write chat_snapshot and metadata to /tmp/handoff/ for plugin

Target plugin (post-restore hydration):
- Create Foundry Conversation with replayed chat messages
- Store handoff event fact in Foundry Memory (update_memories)
- Write HANDOFF_CONTEXT.md to workspace (fallback context)
- Send 'handoff_ready' mesh message back to predecessor

The cloud agent now comes online with full context: workspace files,
conversation history in Foundry Conversations, semantic memory via
shared Memory Store, and proactively greets the user via Telegram.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: cross-container handoff — return state in response, not filesystem

The router and openclaw containers have separate filesystems in AKS.
Previously the router wrote workspace tar and chat snapshot to /tmp/
which the plugin couldn't read.

Changes:
- Router: return workspace_tar (base64) and chat_snapshot in restore
  response JSON instead of writing to filesystem
- Plugin: extract workspace tar and parse chat snapshot from the
  HTTP response body (runs in openclaw container where /sandbox/ lives)
- Remove unused extract_workspace_tar() from routes.rs
- Remove tar crate dependency (extraction now done by plugin via CLI tar)
- Add direction, initiated_at, restored_at to restore response

Future: workspace payloads >5MB will auto-transfer via Azure Blob
Storage (SAS URL in handoff state) — not yet implemented.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security: harden workspace tar extraction and chat snapshot parsing

Tar extraction:
- Pre-extract validation: list entries, reject any containing '..' or
  starting with '/' (path traversal)
- Size guard: reject compressed payloads >5MB (decompression bomb)
- Unique temp dir per extraction (race condition prevention)
- --no-same-owner --no-overwrite-dir flags on extraction
- Temp dir cleaned up after extraction
- Removed '|| true' — errors now surface in logs

Chat snapshot:
- Schema validation: must be array, each entry must have string
  role + content
- Cap at 100 messages, role capped at 20 chars, content at 10k chars
- Rejects non-conforming entries silently

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* perf: split sandbox Dockerfile into base + overlay for fast rebuilds

The sandbox image was 4.96 GB with ~4.0 GB of rarely-changing deps
(OpenClaw, Python wheels, Go tools, Node.js, CLI tools) rebuilt on
every code change. Now split into:

Dockerfile.base (~4.0 GB, rebuild weekly/on dep upgrade):
  - Azure Linux 3 + system packages
  - Node.js 22, Python 3 + 41 packages, Go CLI tools
  - OpenClaw framework + extension symlinks + skills
  - gh, ripgrep, 1password, himalaya
  - User setup (sandbox:1000, router:1001)

Dockerfile (~50 MB, rebuild per commit in ~30s):
  - FROM azureclaw-sandbox-base (all heavy deps pre-cached)
  - CLI plugin builder reuses base image (has Node.js already)
  - Router binary, plugin dist, vendored SDK overlay
  - Entrypoint, proxy-bootstrap, skills, policies

All functionality preserved:
  - UID separation, iptables egress guard, seccomp, read-only rootfs
  - Channel plugins (Telegram/Slack/Discord/WhatsApp)
  - Extension dep symlinks (grammy, carbon, bolt, etc.)
  - Vendored SDK overlay, proxy-bootstrap, Control UI symlink
  - ClawHub skills, npm CLI tools (clawhub, mcporter, oracle)

Build paths updated:
  - azureclaw dev: auto-builds base if not cached, --build-base to force
  - azureclaw push: --only sandbox-base to push base image
  - Makefile: image-sandbox-base target added

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(handoff): implement reverse handoff (cloud → local)

CLI-driven reverse handoff orchestration:
- aksRouterExec: kubectl port-forward to AKS pod router
- wakeDormantDocker: detect and restart stopped containers
- readAksCrdSpec: inherit model/egress/isolation from CRD
- rehydrateCredentials: copy K8s secrets to Docker container
- Full 10-step reverse flow: connect → verify → init → snapshot →
  drain → wake local → credentials → restore → succession →
  decommission + delete CRD

Direction-aware source routing:
- sourceExec alias delegates to routerExec (Docker) or
  aksRouterExec (AKS) based on direction
- Forward path unchanged at runtime (sourceExec === routerExec)

Plugin reverse handoff:
- handoff_request returns CLI command for aks_to_local direction
- Completion messages updated for both directions
- Decommission label direction-aware

Operator TUI:
- 'returning' handoff state for active aks_to_local handoff
- Table shows '<' icon and 'Returning' status
- ASCII-only table icons for reliable column alignment
- Column widths tightened

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(router): raise body limit on handoff routes to 50MB

Axum's default body limit is 2MB. Encrypted state snapshots easily
exceed this, causing HTTP 413 on /agt/handoff/snapshot and /restore.

Add DefaultBodyLimit::max(MAX_BLOB_SIZE_BYTES) layer to
handoff_protected_routes — matches the existing 50MB blob size
constant from §9.9.4.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): read AKS admin token from mounted secret

On AKS, the admin token is stored in K8s secret 'router-admin-token'
and mounted at /etc/azureclaw/secrets/admin-token — not as env var
or /tmp file. Updated getAksAdminToken() to:
1. Read from /etc/azureclaw/secrets/admin-token (router container)
2. Fallback: same path in openclaw container
3. Fallback: kubectl get secret (base64 decode)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): use POST for snapshot route (was GET → 405)

The /agt/handoff/snapshot route is POST-only but both forward and
reverse handoff paths were sending GET requests, causing HTTP 405.
Changed both to POST.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): reuse existing snapshot blob in reverse path

The reverse handoff was requesting a second snapshot at step 9, but
the state machine had already advanced to 'draining' after step 5.
The snapshot blob was already captured at step 3 — now the reverse
path uses snapshotResp.body.blob directly instead of re-fetching.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): pipe restore payload via stdin for large blobs

routerExec passes JSON as a curl -d argument, which hits shell
argument length limits for large encrypted snapshots. The reverse
handoff restore now uses 'docker exec -i ... curl -d @-' with the
payload piped via stdin, avoiding ARG_MAX issues.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(connect): handle Ctrl+C to disconnect port-forward

The kubectl port-forward child process with stdio:pipe did not
receive SIGINT from the terminal. Added explicit SIGINT/SIGTERM
handlers that terminate the child process and exit cleanly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): init local handoff session + auth headers for restore

The local Docker router requires admin token + handoff token for
/agt/handoff/restore. The reverse path now:
1. Gets local admin token from Docker container
2. Inits a handoff session on the local router
3. Passes both auth headers to the restore curl call

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): correct step count + Telegram notification on handoff

- Fixed reverse handoff step counter: 13 steps (was 10, showing 11/10+)
- Added Telegram notification on handoff completion for both directions:
  - local→cloud: 'I moved to the cloud'
  - cloud→local: 'I am back on your local machine'
  Best-effort — reads credentials from Docker container (reverse) or
  env (forward), sends via Telegram Bot API. Failures are silent.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): clean up local handoff state after reverse restore

Two fixes for stale handoff sessions blocking subsequent handoffs:

1. CLI: After successful reverse restore, transition local router
   through verify → decommission to reach a terminal state.

2. Router: Expand can_start() to allow re-init from Restoring,
   Verifying, and Decommissioning phases. These indicate a previous
   handoff that completed data transfer but wasn't properly finalized.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(handoff): extend verification timeout + retry mesh send

The verification timeout was 60s but AKS pods can take longer to
fully initialize their plugin message handlers after mesh registration.
The blob sent before the handler is ready gets silently dropped.

Fix:
- Extended timeout from 60s to 180s
- Re-send the handoff_transfer blob every 30s within the verification
  loop, in case the target's message handler wasn't ready on first send
- Also fixes: router can_start() allows stale Restoring/Verifying states,
  local handoff session cleaned up after reverse restore

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(relay): increase max_message_size to 1MB for handoff blobs

Handoff snapshots with real state (chat, audit, credentials) can be
80+ KB. After Signal Protocol encryption + base64 + JSON envelope,
they exceed the relay's 64KB default max_message_size. Bumped to 1MB.

Also includes verification timeout extension (60s→180s) with 30s
re-send retries, already committed in plugin.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: collect workspace/chat/credentials in CLI handoff snapshot

The CLI handoff command was sending an empty snapshot payload
(only shared_secret). The forward handoff via plugin.ts collected
workspace tar, Foundry memories, and credential refs — but the
CLI path (used for both forward and reverse) skipped this.

Changes:
- Collect workspace tar via kubectl exec (AKS) or docker exec (local)
- Collect Foundry Memory Store items as chat context
- Collect credential refs from container environment
- Fix snapshot response field name: size_bytes → snapshot_size_bytes
- Include items breakdown in snapshot response
- Fix step counter: move transfer step into forward branch only

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: stop stepper spinner and cleanup port-forward after handoff

stepper.step('Handoff summary...') started a spinner that was never
stopped with stepper.done(), keeping the event loop alive and requiring
Ctrl+C. Also aksPortForwardStop() was only called in the error path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: add /agt/handoff/succession router endpoint for Ed25519-signed succession

The registry's succession API requires the predecessor's Ed25519 signature
over a canonical message. The private key lives in the router's Governance
identity — inaccessible to the CLI.

New endpoint POST /agt/handoff/succession on the router:
- Takes {successor_amid, reason} from CLI
- Looks up predecessor (self) AMID from registry
- Looks up successor signing key from registry
- Signs canonical message 'succession:{pred}:{succ}:{timestamp}'
- Submits complete SuccessionRequest to registry
- Returns registry response

CLI + plugin updated to call /agt/handoff/succession instead of
/agt/registry/registry/succession (which lacked signing keys).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: sub-agent handoff — collect, snapshot, and re-spawn during handoff

Sub-agents spawned by a parent agent are now included in handoff state
transfer. The full pipeline:

Collection (source side):
- New GET /agt/handoff/sub-agents endpoint lists active sub-agents
  and reconstructs SpawnRequest from CRD spec (K8s) or container
  labels (Docker dev mode)
- CLI + plugin call this endpoint and inject sub_agent_snapshots
  into the handoff snapshot payload

Re-spawn (target side):
- handoff_restore iterates sub_agent_snapshots after state hydration
- Calls create_sandbox() for each sub-agent with the stored config
- Returns per-sub-agent results (spawned/failed) in restore response
- Audit-logged as handoff:restore:sub-agent

Supporting changes:
- SpawnRequest, HandoffMeta, SubAgentSnapshot: added Clone derive
- Sub-agent results included in restore response as sub_agent_results

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: AMID remapping + sub-agent workspace collection during handoff

Two improvements to sub-agent handoff:

1. AMID remapping: When re-spawning sub-agents on the target, the old
   parent AMID in trusted_peers is replaced with the new parent's AMID.
   This ensures sub-agents trust the new parent for KNOCK handshakes.
   The new parent AMID is looked up from the registry at restore time.

2. Workspace collection: The CLI now exec's into each sub-agent's
   container (kubectl for AKS, docker for local) to collect workspace
   tar before including it in the snapshot. Each sub-agent's workspace
   is capped at 2MB. The plugin path is best-effort without workspace
   (no container exec access from inside the sandbox).

Also adds sub_agents_respawned count to plugin restore metadata.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: collect sub-agent workspace via E2E mesh during handoff

Sub-agents now respond to handoff:workspace_request mesh messages by
tarring their /sandbox/.openclaw/workspace and sending it back via the
E2E encrypted relay. No shared volumes or cross-container exec needed.

Plugin message handler (sub-agent side):
- Receives handoff:workspace_request from parent
- Tars workspace (excludes extensions, node_modules, etc.)
- Sends base64-encoded tar back as handoff:workspace_response
- Falls back to empty response on error (parent doesn't hang)

Plugin handoff orchestration (parent side):
- After fetching sub-agent list from router, discovers each sub-agent
  via registry search to get their AMID
- Sends handoff:workspace_request to each via mesh
- Polls agtInbox for handoff:workspace_response (up to 15s per agent)
- Enriches sub-agent snapshots with workspace tar before creating
  the encrypted handoff snapshot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: expand workspace tar to include cron/policies/agents + mesh 768KB cap

- Add WORKSPACE_TAR_CMD constant in handoff.ts for consistent tar commands
- Include .openclaw/cron, .openclaw/policies, .openclaw/agents in all 5 tar commands
- Refine extension exclusion: only exclude */dist and */node_modules (keep manifests)
- Cap mesh workspace response at 768KB (safe under relay's ~1MB limit)
- Add truncated flag in workspace_response messages

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: chunked mesh transfer for large sub-agent workspaces

Workspaces > 512KB are split into chunks sent via separate mesh messages,
then reassembled on the receiver side. This lifts the practical workspace
transfer limit from ~768KB to ~40MB (80 chunks × 512KB).

Sender (sub-agent):
- Small workspace (≤512KB): single handoff:workspace_response (fast path)
- Large workspace: N handoff:workspace_chunk messages + completion marker
- Max 80 chunks (leaves headroom in relay's 100-message offline queue)

Receiver (parent):
- Collects workspace_chunk messages into a Map keyed by chunk_index
- Reassembles in order when all chunks received or completion marker arrives
- 30s timeout with partial-chunk recovery (uses what was received)
- Backwards compatible: single-message responses still work unchanged

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: unified chunked mesh transport layer + file transfer tool

Implements a general-purpose auto-chunking transport layer for the E2E
encrypted mesh. Any payload exceeding 512KB is transparently split into
chunks with per-chunk SHA-256 integrity verification, then reassembled
on the receiver side before reaching application logic.

Transport layer (meshSend + meshHandleTransportMessage):
- meshSend(): auto-chunks large payloads into mesh:transfer_manifest +
  N mesh:transfer_chunk messages. Small messages pass through directly.
- meshHandleTransportMessage(): intercepts transport messages in onMessage,
  accumulates chunks, verifies SHA-256 hashes, reassembles, and delivers
  the original message to the application layer.
- Per-chunk + manifest-level SHA-256 integrity verification
- 2-minute TTL with automatic stale transfer cleanup
- Max ~40MB per transfer (80 chunks × 512KB)

New tool — azureclaw_mesh_transfer_file:
- Agents can send files to each other via E2E encrypted mesh
- Files up to 30MB supported (auto-chunked transparently)
- Received files auto-saved to /sandbox/.openclaw/workspace/incoming/
- Path traversal protection (must be within /sandbox)

Consumers updated to use unified transport:
- mesh_send tool: auto-chunks large task messages
- Handoff blob transfer: auto-chunks encrypted snapshots
- Sub-agent workspace collection: auto-chunks workspace tars
- Handoff re-send loop: uses meshSend for retransmit

Limits raised:
- Router MAX_BLOB_SIZE_BYTES: 50MB → 200MB (sub-agent workspaces)
- Plugin workspace tar cap: 5MB → 50MB
- CLI workspace tar cap: 5MB → 50MB

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: configurable router URL + fix spawn test timeouts + add transport tests

- Make ROUTER and ROUTER_BASE configurable via AZURECLAW_ROUTER_URL env var
  (defaults to http://127.0.0.1:8443 for backward compat)
- Fix 5 pre-existing spawn tool test timeouts caused by local dev router
  on port 8443 — tests now point to unused port 19876 for immediate
  ECONNREFUSED instead of 45s polling loop
- Fix test assertions to match actual error behavior (plain text errors,
  not JSON, when router is unreachable)
- Add 11 new tests (198 total, up from 187):
  - mesh_transfer_file: registration, schema, path traversal, abs path,
    mesh-not-connected
  - mesh_send: registration, params, error when disconnected
  - handoff_status: registration, returns status JSON
  - AZURECLAW_ROUTER_URL: spawn + spawn_status use configurable URL

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: address 4 direction-specific handoff gaps

1. Plugin reverse: retry with 60s backoff when discovering local target
   (local agent may be waking from dormant state after CLI runs
   wakeDormantDocker)

2. Direction validation: both plugin and router now validate that the
   incoming handoff direction matches the environment (AZURECLAW_DEV_MODE).
   Warn-only on mismatch — doesn't block to avoid false positives.

3. Forward credential rehydration: collect actual credential VALUES
   from Docker (not just refs), create K8s secret on AKS target before
   pod starts so envFrom can mount them. Closes the credential gap
   where forward handoff lost Telegram/Slack/Brave tokens.

4. Symmetric cleanup: reverse handoff now scales deployment to 0
   instead of deleting the CRD. This preserves the sandbox definition
   for instant re-forward handoff while freeing all compute resources.
   Falls back to CRD deletion if scale fails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: graceful sub-agent interrupt during handoff

Add handoff:interrupt protocol so sub-agents can save in-progress work
before their workspace is collected during a handoff.

Plugin path (mesh-based):
- Parent sends handoff:interrupt to all sub-agents concurrently
- Sub-agents set handoffInterruptRequested flag
- processTaskWithTools checks the flag between LLM rounds
- On interrupt: saves task progress to .task-in-progress.json
  (round, messages, last content, original task)
- Sends handoff:interrupt_ack back to parent
- Parent waits up to 10s for acks, then proceeds with workspace collection

CLI path (exec-based):
- CLI writes .handoff-interrupt sentinel file into each sub-agent container
- processTaskWithTools also checks for this file between rounds
- Same progress save behavior (.task-in-progress.json)

Both paths ensure the workspace tar includes the progress checkpoint,
so it survives the handoff and is available on the target for resumption.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: sub-agent workspace injection + task resumption after handoff

Complete the sub-agent handoff lifecycle — after re-spawn on the target,
sub-agents now receive their workspace and resume interrupted work.

Router changes:
- Restore response now includes sub_agent_workspaces array with each
  sub-agent's workspace_tar, task_context, status, and checkpoint

Plugin — target side (post-restore):
- Waits up to 60s for each re-spawned sub-agent to register in mesh
- Sends workspace tar via meshSend (auto-chunked for large workspaces)
- Sends handoff:resume with task context and checkpoint info

Plugin — sub-agent side (new message handlers):
- handoff:workspace_inject: extracts received workspace tar into /sandbox/
  with path traversal validation and size guard
- handoff:resume: reads .task-in-progress.json, sends resume_ack to parent
  with status report ('Successfully restored in cloud. Resuming interrupted
  work from round N: <task>'), then re-enters processTaskWithTools with a
  contextual prompt that includes the original task, progress, and last output

The full sub-agent handoff lifecycle is now:
  interrupt → save progress → collect workspace → transfer →
  re-spawn → inject workspace → resume task → report to parent

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent handoff — Docker API encoding + name prefix + source cleanup

Three bugs prevented sub-agents from properly migrating during handoff:

1. collect_sub_agent_snapshots_docker passed raw JSON braces in the
   Docker API URL — curl treated {} as glob patterns and the Docker
   daemon couldn't parse the filter. Result: always returned 0
   sub-agents, so the snapshot blob had no sub-agents to respawn.
   Now uses URL-encoded filter via the docker_api() helper (matching
   list_sandboxes_docker's pattern).

2. Same function used the Docker container name (azureclaw-{name})
   as the agent name, causing respawn to create
   azureclaw-azureclaw-{name} on the target. Now strips the prefix.

3. Source decommission only put the main agent dormant — sub-agent
   containers kept running. Now destroys all source sub-agents via
   the spawn API before decommissioning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent handoff — wrong API URLs + missing report steps

Four fixes for sub-agent handoff orchestration:

1. Sub-agent list used GET /sandbox/spawn (wrong) — now GET /sandbox/list
2. Sub-agent delete used DELETE /sandbox/spawn/{name} (wrong) — now
   DELETE /sandbox/{name} (matches actual router routes)
3. Response field was 'sub_agents' but endpoint returns 'sandboxes'
4. No sub-agent info in handoff progress report — added _hp() calls
   for: discovery count, interrupt/checkpoint status, workspace
   collection count, snapshot inclusion, cleanup status, and
   sub_agents_transferred in the final result object

Also fixed orphan target cleanup URLs (2 places) that had the same
/sandbox/spawn/{name} → /sandbox/{name} issue.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: set trusted_peers for re-spawned sub-agents after handoff

The trusted_peers remapping in handoff_restore was dead code —
Docker snapshots had trusted_peers=None, so the if-let guard never
fired. Re-spawned sub-agents had no trusted parent AMID and rejected
handoff:workspace_inject + handoff:resume messages from the new
parent. This caused workspace injection to silently fail and task
resumption to never trigger.

Fix: always set trusted_peers to include the new parent's AMID,
regardless of whether the original snapshot had it set:
- If peers existed: remap old parent → new parent (existing logic)
- If peers existed but old parent absent: append new parent
- If peers was None: set to new parent entry (new case)

This ensures the sub-agent trusts the new parent on first KNOCK
and accepts workspace/resume messages immediately after spawn.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat: include sub-agent status in Telegram greeting after handoff

Restructure the restore flow so sub-agent workspace injection and resume
happens BEFORE the Telegram greeting. After sending resume signals, wait
up to 8s for sub-agents to send resume_ack messages. The Telegram greeting
now includes a sub-agent status section showing each agent's name, state
(resumed/ready/starting/failed), and a task preview.

The handoff_ready mesh message back to the predecessor also now includes
sub_agents_restored count, sub_agents_resumed count, and per-agent details.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: use per-sandbox runtime for operator exec path, not global devMode

The operator's unified view fetches both Docker and AKS agents, but the
routerExec/fetchAgtQuick/fetchEgressDomains functions used the global
devMode flag to choose between docker-exec and kubectl-exec. This meant
Docker agents were queried via kubectl (which fails — no K8s namespace)
when the operator ran in non-dev mode, and vice versa.

Fix: use sb.runtime === 'docker' per-sandbox instead of the global devMode
flag in all 4 places that exec into agent containers:
- fetchSecurityState (routerExec + k8sCheck)
- fetchEgressDomains (routerCurl)
- fetchAgtQuick (docker/kubectl exec)
- seccomp profile inference

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: sub-agent trust + workspace logging after handoff

Two fixes for broken parent↔sub-agent communication after handoff:

1. Stale trusted_peers: After handoff, re-spawned sub-agents get new
   AMIDs but the parent's parentTrustedAmids set contains old AMIDs.
   KNOCK handler rejects all incoming sub-agent messages (score=0 <
   threshold=500). Fix: when the plugin discovers a sub-agent's new
   AMID during workspace injection, register it in amidToName,
   nameToAmid, parentTrustedAmids, and push baseline trust to router.

2. Silent from_value failure: serde_json::from_value for sub-agent
   snapshots at POST /handoff/snapshot was wrapped in `if let Ok`
   which silently swallowed deserialization errors, potentially losing
   all sub-agent workspace data. Changed to match with tracing::warn
   that logs the exact error + JSON preview for debugging.

Also adds roundtrip test for SubAgentSnapshot workspace_tar through
the full serialize→compress→encrypt→decrypt→decompress→deserialize
chain, and a test for the JS↔Rust base64 round-trip.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* security: remove clawhub.com and openclaw.ai from default egress allowlist

clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill).
openclaw.ai enables `curl | bash` install vectors from inside sandboxes.

Both were in the default Helm values and example CRD. Removed from:
- deploy/helm/azureclaw/values.yaml (default egress allowlist)
- examples/basic-agent/clawsandbox.yaml (example CRD)
- PLAN.md (policy presets documentation)

Defense is now three layers deep:
1. Egress proxy blocks clawhub.com/openclaw.ai (network)
2. Skills directory is root-owned, chmod 640/750 (filesystem)
3. Plugin code is root-owned, read-only for sandbox (code integrity)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: use sub_agent_results as trust+resume loop driver, not sub_agent_workspaces

Root cause: the post-restore trust registration and resume signals were
gated on restoreResp.sub_agent_workspaces (workspace data), which could
be empty even when sub-agents were successfully spawned. This caused the
entire block to be skipped — no trust registration, no resume signals,
no Telegram sub-agent status.

Fix: use restoreResp.sub_agent_results (always populated when sub-agents
spawn) as the primary loop driver. Workspace data is looked up by name
from sub_agent_workspaces as a secondary source.

Also:
- Added workspace_inject_ack: sub-agent confirms extraction success/fail
  with file_count + error back to parent before resume is sent
- Parent waits up to 15s for ack, logs result, passes workspace_delivered
  flag in resume payload
- handoff_ready report now includes sub_agents_workspace_delivered count
- Telegram greeting shows 📦 icon per sub-agent when workspace arrived
- Increased mesh registration wait from 60s to 90s (AKS pods need boot)
- Router logs snapshot details when building sub_agent_workspaces

Tests added:
- Rust: sub_agent_workspaces builder filter (empty/non-empty workspace_tar)
- Rust: full encrypt→decrypt→restore round-trip with 2 sub-agents
- Rust: edge case — empty workspace + empty task_context filtered out
- TS: sub_agent_results drives loop even when sub_agent_workspaces empty
- TS: workspace_inject_ack protocol (success + failure paths)
- TS: handoff_ready includes workspace delivery status
- TS: only spawned sub-agents enter trust loop
- TS: missing sub_agent_results graceful fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sdk): reuse active Signal Protocol session instead of crashing

Vendor patch #10: SessionManager.initiateSession() threw 'Active session
already exists' when the crypto layer had a session established via
incoming KNOCK but client.activeSessions wasn't synced.

This broke mesh_transfer_file and any second send to the same peer.

Fix: return existing session info with reused=true flag instead of
throwing. establishSession detects reuse, syncs activeSessions, and
skips redundant KNOCK/activate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: don't expose handoff confirmation code to LLM, add console.log diagnostics

Security fix: The handoff confirmation token was returned in the tool
result, allowing the LLM to self-confirm without user input. Now the
code is sent directly to Telegram (side-channel) and printed to console
for TUI users — the LLM never sees it.

Changes:
- Remove confirmation_token from azureclaw_handoff_request tool response
- Send code via Telegram sendMessage as a side-channel delivery
- Print code to console.log (visible in kubectl logs, not to LLM)
- Update tool descriptions to emphasize code comes from user input
- Bump CONFIRMATION_MIN_DELAY_SECS from 3s to 8s
- Add console.log diagnostics in post-restore IIFE for handoff debugging
- Add test: tool response must not contain confirmation_token

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: increase sub-agent tool-calling rounds from 10 to 25

10 rounds was too tight for data-heavy tasks — sub-agents hit the
cap and returned truncated results. 25 gives enough room for
multi-step research/collection while still preventing runaway loops.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: increase prekey retry window from 16s to 45s for sub-agent mesh send

Sub-agents take 20-30s+ after pod is Running to upload prekeys
(gateway start → plugin load → SDK init → relay connect → prekey
upload). The previous 8×2s=16s window wasn't enough, causing
'Cannot get prekeys' failures on task dispatch.

Now: 15 attempts × 3s = 45s max wait with clearer hint message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: add file_transfer_ack for verified file delivery

The mesh_transfer_file tool had no delivery confirmation — it reported
'sent' but couldn't verify the file was actually written to disk on the
target agent. This caused silent failures where coordinator-notes.txt
appeared to transfer but never landed.

Now:
- Receiver sends file_transfer_ack with success/saved_to/error
- Sender waits up to 15s for ack
- Tool returns 'delivered' (with path) or 'sent_no_ack' (no confirmation)
- Write is verified with stat after writeFileSync

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: filter protocol messages from mesh_inbox, improve file transfer reliability

mesh_inbox now filters out internal protocol messages (handoff blobs,
acks, workspace inject/resume) so the LLM only sees actual sub-agent
replies. Shows filtered_protocol_messages count for visibility.

file_transfer now retries up to 3 times with ack verification — sends,
waits 15s for file_transfer_ack, retries with 3s backoff if no ack.
Returns 'delivered' with exact path or 'sent_no_ack' after all retries.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: auto-decode file_transfer content in mesh_inbox

file_transfer messages now show decoded text content (or a binary
placeholder) instead of raw base64 blobs. Uses null-byte detection
to distinguish text vs binary files. Text files are fully readable
in the inbox; binary files show filename and size.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: stale AMID cache poisoning breaks post-handoff mesh delivery

Root cause: trust+resume loop found OLD sub-agent AMIDs still in the
registry (Docker containers hadn't timed out yet) and cached them.
New AKS sub-agents registered with different AMIDs 26s later, but
the parent never discovered them — all messages went to dead relay
connections and were silently dropped.

Three-layer fix:
1. Stale AMID rejection: collect original_amid from handoff snapshots,
   reject matching registry results, wait for NEW AMIDs to appear
2. Prekey readiness gate: verify E2E session is established before
   sending workspace_inject (20 attempts × 3s = 60s max)
3. Workspace inject retry: 3 attempts with 20s ack wait each, catches
   send errors and retries instead of fire-and-forget

Also filters protocol messages from mesh_inbox and auto-decodes
file_transfer base64 content so LLM sees readable text.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: write HANDOFF_FILES.md manifest after workspace inject

After extracting the workspace tar, writes a manifest listing restored
user-facing files to /sandbox/.openclaw/workspace/HANDOFF_FILES.md.
This makes injected files discoverable when the agent is asked about
its workspace contents post-handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: bump AGT rate limits for multi-agent handoff

Policy: 120 → 240 max_calls/60s for inference:* actions.
Router: 100 → 200 global req/s, 10 → 20 per-agent req/s.
Handoff with 3+ agents doing workspace inject + resume + relay
traffic was hitting the old limits.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: promote incoming/ files to workspace root after handoff inject

Copies files from incoming/ to the workspace root so the agent sees
them immediately when listing files, without needing to know about
the incoming/ directory convention.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: scale down sub-agent deployments during cloud→local decommission

After scaling the parent to 0, also scale down all sub-agent
deployments that were snapshotted. Prevents orphaned sub-agent
pods running on AKS after reverse handoff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add bidirectional handoff architecture diagrams and changelog

- New section 12 in architecture-diagrams.md: agent lifecycle across
  handoff, forward/reverse flows with sub-agents, stale AMID cache
  poisoning problem & fix, workspace injection detail
- CHANGELOG: add handoff features, sub-agent support, rate limit bump
- README: add handoff to features list and CLI reference table

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
Per docs/implementation-plan.md §5.4. The compat suite is the zero-regression
wire (principle §0.2 #1) between existing user-facing flows and the
Phase-0→Phase-4 decomposition. Before we split any of the 12 hotspot files
or swap to AgtMeshProvider, every protected flow (§5.1, eight total) needs
a spec here; after the change, the same spec must still pass.

- tests/compat/package.json          — vitest-based standalone npm package
- tests/compat/tsconfig.json         — ES2022 strict typecheck
- tests/compat/vitest.config.ts      — fork pool for per-spec isolation
- tests/compat/README.md             — charter, protected flows, authoring rules
- tests/compat/harness/types.ts      — the 8 protected-flow catalogue
- tests/compat/harness/blessed-mock.ts — headless blessed + blessed-contrib
                                       surfaces (screen/box/log/list/table/
                                       grid/line/bar/sparkline) with a
                                       snapshot() oracle + typeKey() driver
- tests/compat/specs/operator-tui.spec.ts — 11 harness-sanity assertions
                                       green + 8 it.todo staging Phase 1
                                       render-and-drive tests

Locally: cd tests/compat && npm ci && npm test — 11 passed / 8 todo (19).

No production code touched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
Per docs/implementation-plan.md §5.4. Behavioral conformance corpus —
the net that catches 'endpoint returned 200 but skipped the crypto
step' bugs.

tests/conformance/ layout
- package.json / tsconfig.json / vitest.config.ts — own workspace,
  fork-isolated pool so ratchet state doesn't leak across specs.
- README.md — corpus table (9 corpora across Phases 0, 1, 3),
  provider-axis rule (plan §11.3), invariant discipline.
- fixtures/README.md — vendored-source policy per principle §0.2 #10.
- harness/index.ts — empty scaffold; helpers land per-PR.

Phase 0 specs (all 'it.todo', 59 invariants total — vitest reports
59 todo / 0 pass / 0 fail, no false-green):
- signal-x3dh.spec.ts (14) — key-exchange shape, symmetric ratchet
  invariants, base64 input hygiene (vendor patches #3, #4).
- signal-knock.spec.ts (15) — KNOCK happy-path, trust threshold, relay
  disruption, wire-shape parity (vendor patches #5, #7, #8, #1/#2
  timestamps, vendored<->AGT byte-identical KNOCK).
- signal-negative.spec.ts (13) — ciphertext integrity, replay, session
  clobber (vendor patch #10), DoS surfaces.
- sandbox-isolation.spec.ts (17) — seccomp, Landlock, egress-guard,
  router-as-only-network-path. Guarded by CONFORMANCE_E2E=1 (requires
  Kind; compat suite Kind harness wires it in Phase 1).

Principle §0.2 #8 ('solid, not look-alike'): it.todo is a documented
pending assertion, not a silently-passing no-op. Each spec's top
comment cites the vendor patch or production bug it exists to prevent
recurring.

Principle §0.2 #10: every invariant that references an external
protocol cites its upstream source (libsignal, RFC3339 'Z' suffix,
Signal Double Ratchet spec) either inline or in the README.

ci/no-stubs.sh already allow-lists tests/ subtrees; gate remains PASS.
No new dependencies outside the pinned vitest / typescript / @types
dev-deps already used by tests/compat/.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
… validator

Implements plan §1.3 (Outage semantics) as a pure, deterministic decision
function. No I/O, no AGT import surface, clock injected.

inference-router/src/providers/outage.rs (new):
 - OutageMode enum: Strict | CachedRead | DegradedDev
   * serde camelCase; FromStr accepts camel/kebab/snake; Display; Default = Strict
 - OutageConfig { mode, cached_ttl }
   * validate_for_env(is_dev_env) rejects DegradedDev in prod
   * rejects cached_ttl = 0 on CachedRead
   * enforces MAX_CACHED_TTL = 15 min
 - CachedDecision<T> with is_expired(ttl, now) — backwards-clock = expired
 - OutageAction<T>: Deny { mode } | ServeCached { verdict } | AllowWithWarning
 - decide_outage(config, cached, now) — pure, test-friendly
 - 19 unit tests covering all three modes, cache freshness, clock skew,
   serde round-trip, env validation, TTL bounds

inference-router/src/providers/mod.rs:
 - remove placeholder OutageMode stub; re-export real types

controller/src/providers/mod.rs:
 - mirrored OutageMode (from_spec, is_dev_only, validate_for_env)
 - OutageModeError::DegradedDevInProd
 - 4 new unit tests

docs/security-audits/2026-04-24-phase1-outage-semantics.md:
 - STRIDE, principle-mapping, re-audit triggers; both sign-offs present.

No call-site in the router consumes decide_outage yet — that lands with
the first AGT provider. Landing the pure semantics first locks the
decision rules before any provider wiring pressures them.

Verification:
 - cargo test --all: 106+155+15+26+3 = 305 passed (was 286, +19)
 - cargo clippy --all-targets --all-features -- -D warnings: clean
 - six CI gates PASS on the branch tip

Plan refs: §0.2 #1/#2/#3/#4/#5/#8/#9/#10 | §1.3 | §1.4

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 12, 2026
…hments, sibling trust race, final-deliverable rule

Live multi-agent demo (analyst→viz→writer fan-out) surfaced four
distinct breakages that each silently degraded the run while the
parent agent still self-reported success.

1. Image generation 404 (router URL prefix regression)

   Foundry's account-scoped /openai/v1/images/generations endpoint
   does NOT accept the /api/projects/<project>/ URL prefix that
   chat-completions tolerates. Commit 582eb39 unified everything
   through that prefix. In dev (raw Azure OpenAI account, no project
   path) it works; in AKS prod against a Foundry project endpoint
   the upstream returns a fast 404 and image_generation falls back
   to written descriptions.

   Fix: strip /api/projects/<name>/ from the upstream endpoint
   inside the images_generations handler before forwarding. Add
   strip_project_prefix helper + four unit tests (project-stripped,
   trailing-slash variant, AOAI passthrough, no-prefix passthrough).

2. foundry_code_execute drops container files

   The Responses API tool only walked output[].type=='message' for
   text. matplotlib PNGs / CSVs generated by code_interpreter live
   inside Foundry's per-run container and are referenced via
   container_file_citation annotations or { type: 'image' } entries
   in code_interpreter_call.outputs. They were never downloaded, so
   the demo's bar chart silently degraded to ASCII.

   Fix: collect every (container_id, file_id, filename) reference
   from both shapes, GET each via the new /openai/containers/...
   router route (added to foundry_standalone_routes), and write the
   bytes to /sandbox/.openclaw/workspace/. Append the local paths
   to the tool result so downstream tools (mesh_transfer_file,
   file_write) can ship them. Adds routerCallBinary helper for
   binary downloads through the router.

3. Sibling KNOCK race in parallel fan-out

   AGT_TRUSTED_PEERS is baked at spawn time and consumed once at
   sub-agent boot. When the parent spawns analyst → viz → writer in
   sequence, only writer (last) sees all siblings. analyst's
   parentTrustedAmids only contains parent — so when viz or writer
   later try to KNOCK analyst, the trust score is 0 + 0 = 0 and
   the KNOCK is rejected at threshold 500. The demo logs confirm
   only 1 of 3 sibling pairs ever opened a session.

   Fix: after every successful spawn, the parent broadcasts a
   peers_update message containing the new sibling's AMID to every
   already-running sibling. Each sub-agent now records the parent's
   AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and
   handles peers_update only from that AMID, extending its
   parentTrustedAmids set at runtime.

4. Sub-agent system prompt missing FINAL DELIVERABLE rule

   Sub-agents were told they could mesh_transfer_file artifacts to
   peers, but nothing forced them to mesh_transfer_file the FINAL
   artifact back to the parent before returning a summary. The
   writer's executive_brief.md sat in its local /sandbox forever
   while the parent reported success.

   Fix: append a hard rule to the sub-agent system prompt requiring
   mesh_transfer_file(to_agent='parent', ...) as the last action
   before any "task complete" reply, with one call per output file.

Tests
- inference-router: 643 lib tests pass (4 new strip_project_prefix
  tests).
- inference-router: cargo clippy --all-targets clean.
- runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only
  pre-existing warnings.
- cli: 553 tests pass.

Deployment
- For #2: rebuild + push inference-router image.
- For #1, #3, #4: rebuild + push sandbox image.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
…#245)

* feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap

Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): pre-build mesh-plugin for runtime CI + format reconciler

PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): quote workflow step name containing colon

YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): clippy fixes for Rust 1.95.0

CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12

Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5

The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): mark gaps G1 and G2 fixed on local AGT branch

Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.

* dev: add --mesh-provider <vendored|agt> selection with first-run prompt

Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites

Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): wire AGT provider end-to-end (6 stackup bugs)

End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): always route mesh registry through inference-router

azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): break mesh_send infinite poll loop on dead sub-agent

probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): suppress /v1/registry/* 404 leaks in AGT mode

Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode

Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s

The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* runtime: hide Foundry tools in github-copilot mode (same as github-models)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch

Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(up): --mesh-provider=agt picks AGT manifest + flips helm value

Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): plumb --mesh-provider into local-k8s helm install

Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): AGT wire protocol adapter for mesh_peer

Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer

- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(deploy): AGT mesh probe path + Cilium pod-port NP allow

deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh promote): AGT-compat health + WS upgrade paths

azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): first-run picker for local vs remote mesh source

azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container

On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh

Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap

mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call

Previous commit 34662c7 wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* push: auto-detect mesh provider from live helm release

When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* entrypoint: fail-open trust gate when running anonymous tier

When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop

Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable

The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log

CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
…D) (#283)

Adds mtime-poll watchers for both /etc/azureclaw/inference and
/etc/azureclaw/memory mount directories, mirroring the existing
governance::Governance::spawn_policy_watcher pattern. Closes Slice 3
DoD #4 (router reloads within 5s of kubectl edit) explicitly and
Slice 2 DoD #1 implicitly (the echo loop now closes Compiled→Ready
on every change, not just the first).

- spawn_inference_policy_watcher (INFERENCE_POLICY_WATCH_INTERVAL,
  default 5s)
- spawn_memory_binding_watcher (MEMORY_BINDING_WATCH_INTERVAL,
  default 5s)
- load_and_install now clears the handle on NoBinding/NoPolicy and
  preserves it on Error — proper hot-reload semantics for when an
  operator removes spec.memoryRef / spec.inferenceRef.
- Both watchers wired in main.rs after governance.
- 779 router lib tests (+5).

Co-authored-by: Patrik Lagi <palagi@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds
GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the
existing singular mcpServerRef, which is now deprecated and honored
as a length-1 alias via effective_mcp_server_refs().

Controller-side changes:
- New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET
  in status/conditions.rs::reason.
- Reconciler mirror loop refactored to iterate effective_mcp_server_refs().
  Singular-field use emits tracing::warn with McpSingularDeprecated.
  len > 1 short-circuits via degrade! macro until Slice 4d.2 wires
  per-server addressing — principles §3 honest 'not-yet-enforced' signal.
- 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty
  + plural-wins-when-both-set. Controller suite: 555 passing.

Admission CEL on deploy/helm/azureclaw/templates/crd.yaml:
- Mutex: singular and plural cannot both be set.
- maxItems: 8 (router-side scheme is sized for this).
- Per-name uniqueness across mcpServerRefs.

Out of scope for 4d.1 (queued for 4d.2):
- Per-server jwks-{name}.json / tools-{name}.json file scheme.
- Router-side McpServerRegistry + namespaced tool dispatch.
- Stale-file sweep (DoD #6).
- e2e fixture with ≥ 3 servers (DoD #1).

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
) (#292)

* Slice 4d.1 — mcpServerRefs plural CRD field + admission CEL

Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds
GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the
existing singular mcpServerRef, which is now deprecated and honored
as a length-1 alias via effective_mcp_server_refs().

Controller-side changes:
- New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET
  in status/conditions.rs::reason.
- Reconciler mirror loop refactored to iterate effective_mcp_server_refs().
  Singular-field use emits tracing::warn with McpSingularDeprecated.
  len > 1 short-circuits via degrade! macro until Slice 4d.2 wires
  per-server addressing — principles §3 honest 'not-yet-enforced' signal.
- 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty
  + plural-wins-when-both-set. Controller suite: 555 passing.

Admission CEL on deploy/helm/azureclaw/templates/crd.yaml:
- Mutex: singular and plural cannot both be set.
- maxItems: 8 (router-side scheme is sized for this).
- Per-name uniqueness across mcpServerRefs.

Out of scope for 4d.1 (queued for 4d.2):
- Per-server jwks-{name}.json / tools-{name}.json file scheme.
- Router-side McpServerRegistry + namespaced tool dispatch.
- Stale-file sweep (DoD #6).
- e2e fixture with ≥ 3 servers (DoD #1).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Slice 4d.2 — per-server McpServer mounts + router discovery (DoD #1 + #6)

Closes Slice 4 DoD #1 (≥3 plural McpServers reachable e2e) at the
mount-and-discovery layer, and DoD #6 (stale-file sweep) via
reconciler-driven volume rebuild. Multi-JWKS OAuth + namespaced tool
dispatch (DoD #3) follow in Slice 4d.3.

Controller:
- reconciler/mod.rs: replace len>1 short-circuit with full iteration
  over effective_mcp_server_refs(). Per-name volumes (mcp-jwks-<name>,
  mcp-signing-<name>) mounted at /etc/azureclaw/mcp/<name>/ and
  /etc/azureclaw/mcp-signing/<name>/.
- First entry (idx 0) keeps legacy MCP_JWKS_PATH + MCP_SIGNING_KEY_DIR
  env vars for backwards compat with current single-JWKS OAuth path.
- New MCP_JWKS_DIR=/etc/azureclaw/mcp env set once per pod.
- governance_mounts.rs: new inject_container_env helper for idempotent
  env-var injection without a volume mount.
- Removed obsolete PLURAL_MCP_SERVERS_UNSUPPORTED_YET reason constant.
- CRD doc comment updated for 4d.2 mount layout + 4d.3 forward-ref.

Router:
- New mcp/registry.rs with scan() + discover_from_env(). At startup,
  reads MCP_JWKS_DIR, enumerates subdirs with parseable jwks.json,
  emits tracing::info!(servers=?, count=N) per discovery + warn!
  per skipped candidate.
- Empty/missing dir handled gracefully (sandbox with zero
  mcpServerRefs is a valid steady state).
- 7 registry unit tests + 4 inject_container_env unit tests.

Stale-file sweep (DoD #6): reconcile rebuilds pod-spec from current
refs; removed refs disappear via SSA. No explicit sweep code needed.

Verified: 559 controller + 804 router tests pass, clippy -D warnings
clean across workspace, cargo fmt --check clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 13, 2026
Operator-facing surface for the BlockedBuffer the forward proxy has
been populating since S12.f. Closes Slice 5 DoD #1.

Producer side (inference-router):
- New BlockedBuffer::snapshot_since(since_unix) and top_hosts(since,n)
  methods. snapshot_since sorts newest-first by last_seen_unix;
  top_hosts aggregates by hostname across (sandbox, port) and
  secondary-sorts deterministically by host name for tied counts.
- 7 unit tests cover filter cutoffs, dedup count carry-through,
  multi-sandbox/multi-port aggregation, n=0 early return, and the
  truncate-to-n contract.

Wire surface (routes/internal.rs):
- GET /internal/egress/blocked?since=<rfc3339|unix|-Nm>
- GET /internal/egress/blocked/top?window=<duration>&n=<int>
Both mounted on the admin-gated 'protected' router. JSON envelopes
carry schema_version: 1 and RFC 3339 strings alongside raw Unix
seconds. Hand-rolled duration + RFC 3339 parsers (no chrono dep)
with 7 unit tests + 6 integration tests covering bare seconds,
s/m/h/d suffixes, relative -Nm form, malformed input → 0, the n≤100
cap, and the default 5m window.

CLI (azureclaw egress blocked <sandbox>):
- New subcommand under 'egress' so --watch/--top/--since don't
  collide with the existing flat options surface.
- Mirrors azureclaw inspect token-resolution (router-admin-token
  secret first, in-pod admin-token file fallback) and uses in-pod
  kubectl exec curl to avoid port-forward collisions.
- --watch loops every 5s with VT clear; --top overrides --since
  (window is its own filter); --json emits raw response.
- 13 vitest unit tests for buildPath, renderers, and unixToIso(0).

Tests: 849 router lib + 8 egress_blocked integration (up from 2) +
13 CLI unit. cargo clippy -D warnings + cargo fmt + npm
typecheck/build/test/lint all clean.

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 16, 2026
…atency overclaims

Second pass of the OSS-readiness audit, focused on the architecture
diagrams (the first pass missed factual errors here).

architecture-diagrams.md:
- Diagram #3: audit record is hash-chained, append-only (NOT signed today —
  cryptographic signing of the chain head is on the roadmap, see
  security.md). Tamper-detection vs tamper-proof is a real distinction.
- Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list),
  prose says nine + mentions ClawPairing as the controller-internal 10th.
- Diagram #7: CRD relationship arrows were wrong against the actual Rust
  structs. Corrected:
    - policyRef → spec.governance.toolPolicyRef
    - mcpRefs → spec.governance.mcpServerRefs
    - inferenceRef → spec.inferenceRef (top-level, was already correct)
    - memoryRef → spec.memoryRef (top-level, was already correct)
    - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real
      link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy.
    - CE 'sandboxRef' → spec.targetSandboxRef.
    - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected
      to every sandbox by the controller (no ref). Noted in prose.
    - EgressApproval added (it was missing from the diagram entirely);
      links to ClawSandbox via spec.sandbox (string name, not ref object).

security.md:
- Headline #1 'agent does not see Azure credentials. Period.' — softened
  with a dev-mode footnote matching the README hero. In azureclaw dev,
  agent+router share a container with separate UIDs but a kernel-level
  container escape defeats the boundary; the hard guarantee is the AKS
  path. Anchor points at architecture.md#two-modes.
- Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation
  latency on the router hot path.' Microsecond was an overclaim without
  a benchmark to back it.

architecture.md:
- Controller row in the components table: 'watches the eight peer CRDs'
  → 'nine peer CRDs (plus controller-internal ClawPairing)'.

All 9 mermaid blocks pass a bracket-balance sanity check.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 16, 2026
…#325)

* docs: OSS-readiness pass — strip internal jargon + correct overclaims

Doc + comment-only audit across user-facing docs ahead of OSS launch.
No code paths touched.

Factual corrections (the heaviest changes):

- docs/security.md headline guarantee #4: drop 'signed by the router'
  claim. The audit log is hash-chained and detects modification but
  is not signed today; signing is on the v1.1 roadmap.
- docs/security.md Layer 6 inference-safety table:
  * Content Safety: 'Always on, server-side' → 'Always on for
    Foundry-provider requests; Copilot/GitHub-Models providers do
    not return prompt_filter_results'.
  * Token budget: was 'not yet aggregated'; the router DOES aggregate
    per-tenant daily and monthly UTC counters with on-disk persistence
    (inference-router/src/budget.rs:200-260).
  * New 'Operator escape hatches' subsection documents the two env-var
    knobs (AZURECLAW_SUPPRESS_CONTENT_FLAGS,
    AZURECLAW_CONTENT_FLAG_MIN_SEVERITY) honestly.
- README hero: 'agent never sees an Azure key' softened to call out
  that this is the AKS guarantee; dev mode co-locates agent + router
  in one container.
- README/architecture: '31 commands' → '30+ commands'; remove
  unverified '18 Foundry API groups' count.
- docs/architecture.md design goal #1 + #4: explicit dev-vs-prod
  scoping; 'same code path' → 'same data-path code with documented
  AZURECLAW_DEV_MODE branches'.
- docs/architecture/a2a-gateway.md: port 8445 is config-locked but the
  mTLS listener itself is still being wired; operators should set
  A2A_GATEWAY_UPSTREAM_URL explicitly until the listener is GA.
- mesh-plugin/src/agt-identity.ts comment: replace 'encrypted at rest
  with per-host KEK' claim with honest 'chmod 0600 is the real
  boundary' note (mirrors PR #324 identity-store fix).
- .github/copilot-instructions.md: '__AGT_INITIALIZED env guard' →
  'Symbol.for(agt-mesh-client)' (matches current code in
  runtimes/openclaw/src/index.ts:458-480).

Internal-jargon strip in user-facing docs:

- docs/api/lifecycle.md: remove 'Slice 4', 'Slice 4d.3/4d.4',
  'Slice 0', 'Slice 1c', 'Slice 2a/2b/2c/2d.1', 'Slice 2d.2',
  'Slice 3a' references — replaced with descriptive prose.
- docs/api/conditions.md: drop 'Slice 1c invariant' and 'principles.md
  §3' references.
- docs/architecture/agt-boundary.md, docs/security-mcp-top10.md,
  docs/cli-reference.md: drop 'Phase 5.2' shibboleth — keep the
  facts (vendored AgentMesh fork was retired upstream).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: deep-dive diagram audit — fix CRD field-name labels + signing/latency overclaims

Second pass of the OSS-readiness audit, focused on the architecture
diagrams (the first pass missed factual errors here).

architecture-diagrams.md:
- Diagram #3: audit record is hash-chained, append-only (NOT signed today —
  cryptographic signing of the chain head is on the roadmap, see
  security.md). Tamper-detection vs tamper-proof is a real distinction.
- Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list),
  prose says nine + mentions ClawPairing as the controller-internal 10th.
- Diagram #7: CRD relationship arrows were wrong against the actual Rust
  structs. Corrected:
    - policyRef → spec.governance.toolPolicyRef
    - mcpRefs → spec.governance.mcpServerRefs
    - inferenceRef → spec.inferenceRef (top-level, was already correct)
    - memoryRef → spec.memoryRef (top-level, was already correct)
    - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real
      link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy.
    - CE 'sandboxRef' → spec.targetSandboxRef.
    - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected
      to every sandbox by the controller (no ref). Noted in prose.
    - EgressApproval added (it was missing from the diagram entirely);
      links to ClawSandbox via spec.sandbox (string name, not ref object).

security.md:
- Headline #1 'agent does not see Azure credentials. Period.' — softened
  with a dev-mode footnote matching the README hero. In azureclaw dev,
  agent+router share a container with separate UIDs but a kernel-level
  container escape defeats the boundary; the hard guarantee is the AKS
  path. Anchor points at architecture.md#two-modes.
- Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation
  latency on the router hot path.' Microsecond was an overclaim without
  a benchmark to back it.

architecture.md:
- Controller row in the components table: 'watches the eight peer CRDs'
  → 'nine peer CRDs (plus controller-internal ClawPairing)'.

All 9 mermaid blocks pass a bracket-balance sanity check.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 16, 2026
…diagrams (#327)

Link each provider bullet in 'Pluggable inference backend' to the
relevant architecture chapter and diagram:

- GitHub Copilot → architecture.md#dev-mode + diagram #1 (dev mode pod)
- Foundry / Azure OpenAI → architecture.md#prod-mode + diagrams #2 (prod
  mode pod) and #3 (data path)
- GitHub Models → security.md (Content Safety caveat) + architecture.md
  data-path section (provider routing notes)

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 28, 2026
Wire the previously-orphaned sidecar_client module into the
inference-router and harden it against token substitution attacks.

CONTROLLER (reconciler/mod.rs):
- Extends agent_id_active tuple to carry tenant_id alongside agent
  identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via
  ProvisioningOutcome::Ready.auth_spec.
- Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode
  sandboxes. Without it, router-side tid pinning is disabled and
  warns at boot (insecure, dev-only).

ROUTER (sidecar_client.rs, auth.rs, lib.rs):
- pub mod sidecar_client; — fix the orphaned-module bug.
- WorkloadIdentityAuth now consults SidecarClient first; sidecar mode
  is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves
  per-sandbox audit attribution in downstream Azure RBAC.
- from_env() now Result<Option<Self>>. Partial config (only URL or
  only PINNED set) returns Err; WorkloadIdentityAuth::new() panics →
  AKS surfaces as CrashLoopBackoff. No more silent fallback to a
  different identity model.
- HARD-fail validate_token_claims on EVERY check (rubber-duck #1):
    * tid mismatch / missing (cross-tenant guard)
    * appid/azp mismatch when present (audit attribution guard)
      — if both present, BOTH must match
    * aud mismatch against per-service expected set (cache-poisoning
      guard) — Foundry, Graph, OpenAI, Management, Search mapped
    * exp in past or within 60s skew (stale-token guard)
    * exp missing when tid-pinning enabled (unbounded-lifetime guard)
- Returns a TTL cap = exp - now - 60s; caller takes
  min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching.
- JWT decoder tightened: requires EXACTLY 3 segments (rejects
  2-segment unsecured JWS and 4-segment JWE).
- aud claim normalized: accepts both single-string (Entra default)
  and array (RFC 7519); rejects out-of-spec shapes (number, bool).
- Soft WARN when both appid and azp are absent (some MSI tokens
  legitimately omit both).
- New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning
  state for operator visibility.

TESTS:
- 39 sidecar_client unit + wiremock integration tests, covering all
  HARD-fail branches, the cache-skip on validation failure, the TTL
  cap math, the partial-config Err path, JWT decoder edge cases
  (2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud,
  out-of-spec aud).
- Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client.

Rubber-duck review caught: (1) appid/azp should be hard not soft, (2)
partial env config must fail closed, (3) 3-segment requirement, (4)
aud validation against requested resource, (5) exp validation + TTL
cap. All five addressed in this commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 29, 2026
* docs(entra-agent-id): import POC findings as architecture reference

Captures the end-to-end token-acquisition flow validated on real AKS
in the Microsoft tenant during the POC phase.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): scaffold Entra Agent ID auth machinery

Adds controller-side primitives for operating kars sandboxes as
per-sandbox Entra Agent Identities. Foundation only — pod-spec
integration lands in a follow-up commit.

New modules:
- auth_config.rs: KarsAuthConfig CRD (cluster-scoped singleton).
  Tenant + blueprint + controller MI anchors plus optional
  serviceManagementReference for Microsoft-style tenants.
- auth_config_reconciler.rs: materialises kars-auth-sidecar-env
  ConfigMap with AzureAd__ and DownstreamApis__ env vars. Idle when
  the CRD is absent (cluster stays in anonymous tier).
- agent_identity.rs: Graph client for per-sandbox agent identity SP
  lifecycle. Implements IMDS to MI to blueprint to Graph chain
  proven during POC.
- sidecar_injection.rs: pure-function pod-spec helpers (sidecar
  container shape, pinned-identity env vars, egress-guard iptables).

CRD additions in crd.rs:
- KarsSandbox.spec.meshAuth.mode (auto/agent-id/anonymous)
- KarsSandbox.status.agentIdentity

786 controller tests pass, 14 new tests added.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli): auto-provision Entra Agent ID trust during kars up

Makes Entra Agent ID setup invisible to end users: when `kars up`
runs and the cluster does not yet have a `KarsAuthConfig/default`
resource, the new `mesh/agent_id_setup.ts` module idempotently
provisions the tenant trust anchor end-to-end:

  1. Blueprint app via Graph (POST /v1.0/applications/ with
     @odata.type=#Microsoft.Graph.AgentIdentityBlueprint),
     including the optional serviceManagementReference required by
     Microsoft-style enterprise tenants.
  2. Blueprint service principal (visible in the Entra Agents portal).
  3. Controller managed identity in the customer's subscription.
  4. MI-as-FIC on the blueprint (issuer=login.microsoftonline.com),
     the anti-loop-safe credential path proven by the POC.
  5. KarsAuthConfig CR written to the cluster.

The step is non-fatal: if the user lacks `Agent ID Developer`,
`kars up` continues and the cluster runs in anonymous tier until
the role is granted and `kars mesh setup-trust` is rerun. This
matches the three-tier fallback model documented in
docs/architecture/entra-agent-id/.

New `--service-tree <guid>` flag on `kars up` (and KARS_SERVICE_TREE
env var) propagates the ServiceTree GUID to the blueprint creation.
No hardcoding — tenants that do not require it leave it empty.

Tests: 6 new unit tests for agent_id_setup (idempotence detection,
dry-run, env-var threading, error propagation). 775 CLI tests pass
(2 pre-existing skipped). typecheck + lint clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs+preflight: user-facing Entra Agent ID guidance

Adds the user-facing documentation layer that the previous controller
and CLI commits were missing. Every doc that mentioned the old
`api://agentmesh` flow now points to the new per-sandbox Entra Agent
ID model.

New docs:
- docs/agent-identity.md — day-1 and day-2 user guide. Prerequisites,
  walk-through of what `kars up` does for auth, sub-agent semantics,
  troubleshooting, teardown.

Updated docs:
- README.md — top-line mention of per-sandbox Entra Agent ID in the
  `kars up` summary, linking to the new guide.
- docs/getting-started.md — Step 2.1 calls out the `Agent ID
  Developer` role prerequisite; Step 2.2 expands the bring-up list to
  include the preflight role check and the Entra trust provisioning.
- docs/permissions.md — rewrites the "Tenant-level (Entra ID)
  considerations" section to describe the new model. Replaces
  `api://agentmesh` failure rows with Entra-Agent-ID-specific ones
  (CredentialInvalidLifetimeAsPerAppPolicy,
  InvalidFederatedIdentityCredentialValue, missing sidecar).
- docs/cli-reference.md — documents the new `--service-tree` flag
  + Microsoft-corp example.
- docs/SUMMARY.md — wires agent-identity.md into the mdbook index.

New preflight check:
- cli/src/preflight.ts — calls `checkAgentIdRole` from agent_id_setup.
  Warns (not blocks) when the signed-in user lacks the `Agent ID
  Developer` directory role. The existing `api://agentmesh` warning
  line is removed.
- cli/src/commands/mesh/agent_id_setup.ts — new exports:
  - `checkAgentIdRole` returns hasRole/inconclusive/message via
    Graph `/me/transitiveMemberOf`. Matches by role template id
    (stable) AND display name (forward-compatible).
  - `detectExistingBlueprint` returns whether the configured
    blueprint already exists in Graph.
  - `AgentIdSetupOptions.blueprintName` added so multi-cluster
    deployments can share a tenant-wide blueprint.

Tests: 781 CLI tests pass (+6 new for checkAgentIdRole / blueprint
detect). Typecheck + lint clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(cli): UX polish — from-scratch resets context, surface CA block

Two papercuts surfaced when the user ran `kars up --from-scratch` on
a tenant that already had a previous deployment:

1. `--from-scratch` cleared the resume state but NOT the cached
   deployment context, so the "fresh" run silently reused the prior
   region/RG/Foundry endpoint instead of re-prompting. Fixed in
   `cli/src/commands/up/preflight.ts` — when `fromScratch=true`, skip
   the `loadContext()` prefill entirely and print a clear
   "ignoring any cached deployment context" line. Adds the
   `fromScratch?` field to `UpOptionsForPreflight`.

2. The Entra Agent ID preflight check correctly soft-failed on the
   well-known AADSTS530084 Conditional Access token-binding block
   (common in Microsoft-corporate tenants), but the warning was
   generic and gave the user no actionable next step. Now detects
   AADSTS530084 (and the related AADSTS65001/65002 missing-consent
   codes) specifically and surfaces the exact `az login --scope
   https://graph.microsoft.com//.default` workaround inline.

Docs: adds an `#az-cli-ca-block` anchor section to
`docs/agent-identity.md` so the inline preflight message can link
directly to the troubleshooting paragraph.

Tests: 2 new test cases in agent_id_setup.test.ts pinning the
AADSTS530084 + AADSTS65001 detection paths. Total CLI tests:
783 pass (+2 vs prior commit).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): correct system namespace kars-system (was azureclaw-system)

Caught while auditing residual `azureclaw-` references after the
Azure/kars rebrand. The auth-config reconciler materialised the
sidecar env ConfigMap into `azureclaw-system`, but the actual
controller + helm chart deploy everything into `kars-system`. With
the wrong namespace the ConfigMap was unreachable from sandbox pods
even when the rest of the wiring was correct.

Two-line fix:
- controller/src/auth_config_reconciler.rs:61 — namespace constant.
- controller/src/auth_config.rs:43 — doc comment.

The rest of the controller already uses `kars-system` consistently
(pairing_reconciler, trust_graph_reconciler, signer_policy,
egress_approval_reconciler). This brings the new modules in line.

786 controller tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli): kars mesh setup-trust --mode agent-id (standalone retry)

The Entra Agent ID auto-provisioning that `kars up` runs is now
also reachable as a standalone command — necessary for retrying
after a transient failure (e.g. AADSTS530084 Conditional Access
block on Microsoft Graph) without re-running every other `kars up`
phase.

cli/src/commands/mesh/setup-trust.ts grows a `--mode <agent-id|legacy>`
flag, defaulting to `agent-id`. Agent-id mode forwards to the same
`ensureAgentIdTrust` helper used by `kars up`. Legacy mode preserves
the original api://agentmesh app-registration flow for installations
that haven't migrated yet (slated for removal once all consumers
have switched).

Also surfaces:
- --service-tree <guid> for Microsoft-style tenants
- --cluster-name / --resource-group / --region for controller MI
  override
- Clear AADSTS530084 + "Agent ID Developer missing" remediation
  messages on failure.

Bonus: cli/src/commands/up.ts agentmesh-image import block from a
prior in-flight edit — imports agentmesh-relay-agt + agentmesh-
registry-agt from the public source ACR in --build mode so the
agentmesh deploy step does not block on ImagePullBackOff. (Future
improvement: build them locally from .agt-sdk when the AGT SDK
tarball is present.)

Tests + typecheck clean: 783 CLI tests pass, build green.

* chore(helm): install KarsAuthConfig CRD

Adds the helm template for the new cluster-scoped singleton CRD
introduced in c9ce68f. Without this, `helm upgrade` does not install
the CRD and `kars mesh setup-trust --mode agent-id` fails with
"the server doesn't have a resource type karsauthconfig" when it
tries to `kubectl apply` the CR.

The schema uses `x-kubernetes-preserve-unknown-fields: true` on
spec/status with required-field validation for the top-level
tenant / agentId / controller blocks. The canonical type definitions
remain in controller/src/auth_config.rs; the controller validates
spec shape on reconcile. A future PR will add karsauthconfig to the
existing helm-vs-Rust drift test in controller/src/helm_drift.rs.

helm lint clean. YAML parses.

* feat(cli+bicep): auto-fallback to Bicep when az CLI Graph is CA-blocked

End-to-end UX for Microsoft-corporate-style tenants where the Azure
CLI cannot acquire a Microsoft Graph token (AADSTS530084) — the same
block we hit repeatedly during the POC phase.

What changes:
- deploy/bicep/agent-id-trust.bicep — sub-scope Bicep template that
  provisions everything the imperative path does: blueprint app + SP
  (tagged EntraAgentId), controller MI, MI-as-FIC on the blueprint
  using login.microsoftonline.com (universally allow-listed). Goes
  through ARM + the Microsoft.Graph extension, which has its own auth
  path and is not subject to az CLI CA token-binding policy.
- deploy/bicep/modules/controller-mi.bicep — RG-scope module for the
  controller MI (Bicep needs RG scope for UAMI, parent runs at sub
  scope to create the RG).
- deploy/bicep/bicepconfig.json — enables the Microsoft.Graph
  extension (preview).
- cli/src/commands/mesh/agent_id_setup_bicep.ts — driver that runs
  `az deployment sub create` against the template, parses outputs,
  and writes the KarsAuthConfig CR.
- cli/src/commands/mesh/agent_id_setup.ts — adds
  ensureAgentIdTrustAutoFallback() which tries the fast Graph REST
  path first and transparently switches to the Bicep path on
  AADSTS530084. Other Graph errors propagate unchanged (don't mask
  real permission failures with a Bicep retry).
- cli/src/commands/up.ts — uses the auto-fallback wrapper so
  `kars up` "just works" in any tenant.
- cli/src/commands/mesh/setup-trust.ts — same auto-fallback for
  `--mode agent-id`, and a new `--mode bicep` for users who want to
  skip the CLI attempt entirely.
- Also makes blueprintName configurable so multi-cluster deployments
  in the same tenant share one blueprint by default.

Tests: 783 CLI pass, typecheck clean, helm-lint + bicep-build clean.

Operational note (not code): the AGT registry/relay images the user
manually built+pushed during this session were from a feature branch
(copilot/secure-mcp-agent-governance) that adds proof-of-possession
to /v1/agents — incompatible with SDK 3.7.0 which doesn't sign yet.
Rebuilding from tag v3.7.0 + restarting the deployments resolves the
422 flood. Documented for future kars releases to build from a known
AGT tag rather than HEAD.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(bicep): clean error reporting + linter-clean issuer URL

Three small fixes shaken out by the user's first end-to-end run of
`kars mesh setup-trust --mode bicep` in the Microsoft corporate
tenant:

1. deploy/bicep/agent-id-trust.bicep — the MI-as-FIC issuer was
   hardcoded to `login.microsoftonline.com`. Bicep linter emits
   `no-hardcoded-env-urls` as a warning on stderr, which the CLI
   driver was mistaking for a deployment failure. Switched to
   `environment().authentication.loginEndpoint` so the template is
   cloud-portable AND the linter is silent.

2. cli/src/commands/mesh/agent_id_setup_bicep.ts — error capture
   reworked to use execa's merged `all` stream so the real ARM
   deployment error surfaces instead of being masked by stderr-only
   linter warnings. Lines starting with `WARNING:` are stripped from
   the error summary.

3. Same module — the final `kubectl apply KarsAuthConfig` is now
   correctly recognised as a separate step from the Bicep deployment.
   When the CRD isn't installed (e.g. older controller helm release),
   the user sees:
     "Bicep deployment succeeded, but KarsAuthConfig CRD is not
      installed. All Entra resources are already created — just
      install the CRD and re-run."
   plus the exact `helm upgrade` command. The function returns the
   Bicep result so callers know the Entra side is complete.

Also briefly probed the Microsoft.Graph Bicep extension's beta
channel (1.0.0) to see if it supports the `@odata.type` discriminator
for typed AgentIdentityBlueprint. It does not (`BCP037 The property
"@odata.type" is not allowed on objects of type
"Microsoft.Graph/applications"`). The Bicep path therefore creates a
regular Application tagged `EntraAgentId` — functional for runtime
but not visible under the Entra Agents portal page. Documented as a
limitation; users who need portal visibility for the typed form must
use Graph Explorer or PowerShell (both have different first-party
auth paths that aren't CA-blocked).

Bicep build clean, typecheck clean, 783 tests pass.

* fix(cli+docs): switch Graph calls from /v1.0 to /beta

User reported that the typed POST creating an
`#Microsoft.Graph.AgentIdentityBlueprint` works in Graph Explorer
against the /beta endpoint — and that's what makes the resulting
app visible in the Entra Agents portal. The /v1.0 endpoint accepts
the same body but does NOT route through the typed-resource
discriminator on the server, so the app ends up listed only under
App registrations.

Updates `cli/src/commands/mesh/agent_id_setup.ts` to use
/beta/applications, /beta/servicePrincipals, /beta/users, /beta/me,
and /beta/applications/{id}/federatedIdentityCredentials in every
Graph REST call. The `@odata.type` body field stays
`#Microsoft.Graph.AgentIdentityBlueprint` — only the URL prefix
changes.

docs/permissions.md updated to match (the manual escape-hatch
snippet showing `az rest --method POST --url
https://graph.microsoft.com/v1.0/applications/...` now reads `/beta/`).

Note: this only helps when the imperative Graph REST path runs
successfully. In tenants where the az CLI is Conditional-Access-
blocked from Graph (AADSTS530084), the CLI auto-falls back to the
Bicep ARM path. The Bicep Microsoft.Graph extension does NOT
support the @odata.type discriminator at v1.0 or beta (1.0.0), so
the Bicep-created app is functional but stays in the tag-based
detection mode (App registrations only, not Agents portal). This
is a current limitation of the extension, documented in
agent_id_setup_bicep.ts.

Tests: 783 CLI pass, typecheck clean.

* feat(cli): device-code re-login + exact-match Graph body for typed blueprint

Two coordinated changes to make the imperative Graph REST path
succeed in tenants with Conditional Access token-binding policy
(AADSTS530084) on the az CLI's first-party app:

1. Match Graph Explorer's exact working body shape:
   - @odata.type value WITHOUT the leading `#` ("Microsoft.Graph.
     AgentIdentityBlueprint", not "#Microsoft.Graph.…"). Both forms
     are accepted by /v1.0; only the unprefixed form survives the
     /beta route reliably per user testing.
   - sponsors@odata.bind / owners@odata.bind URLs use /v1.0/users/.
     Graph rejects /beta/users/ as an @odata.bind target.
   - POST URL stays /beta/applications/ — the typed
     AgentIdentityBlueprint resource only surfaces in the Entra
     Agents portal page when created via the /beta route.

2. Auto device-code re-login on AADSTS530084:
   - When `az rest` returns AADSTS530084, the helper does a one-shot
     `az login --use-device-code --scope https://graph.microsoft.com//.default`
     which goes through a different OAuth flow that often bypasses
     the token-binding CA policy applied to the default interactive
     flow.
   - User is prompted in the terminal: "visit https://microsoft.com/
     devicelogin, paste this code". After login, the original Graph
     call is retried once. If it still fails, the
     ensureAgentIdTrustAutoFallback wrapper proceeds to the Bicep
     ARM path (untyped but functional).

Updates existing test fixture for checkAgentIdRole so the
device-code retry path is mocked. 783 CLI tests pass.

Together these should let Microsoft-corporate-tenant users get the
typed AgentIdentityBlueprint that surfaces in the Entra Agents
portal page, without needing to fall back to Bicep (which can only
produce the tag-based form).

* docs(agent-identity): document AADSTS530033 + Graph Explorer workaround

User hit the next layer of CA policy after device-code re-login:
`AADSTS530033` — "device must be Intune-managed" — applies to
both interactive and device-code flows of the Microsoft Azure CLI
first-party app. Bicep ARM path keeps working (different auth
surface), but it produces an untyped blueprint that only shows
under App registrations, not the Entra Agents portal page.

Updates the existing AADSTS troubleshooting section with:
- Table of auto-handled error codes + what the kars CLI does
- Explicit "even Bicep cannot produce the typed form" path:
  Graph Explorer PATCH to upgrade the untyped blueprint app to
  the typed AgentIdentityBlueprint discriminator in place
- Note that runtime is unaffected by typed-vs-untyped — only
  portal categorisation differs
- Long-term Intune-enrolment recommendation

* fix(cli): kars mesh setup-trust short-circuits when already provisioned

Two bugs surfaced when the user re-ran setup-trust after the
initial successful Bicep + Graph Explorer flow:

1. `karsAuthConfigExists()` checked for "karsauthconfig/default" in
   kubectl's `-o name` output, but newer clusters return the
   fully-qualified form `karsauthconfig.kars.azure.com/default`.
   The bug caused the existence check to always return false, so
   the wrapper always tried the Graph REST path even when the trust
   was already in place. Fix: match the invariant `/default`
   suffix.

2. `kars mesh setup-trust --mode agent-id|bicep` did not consult
   `karsAuthConfigExists()` at all, so every re-run triggered the
   full provisioning attempt — which, in CA-blocked tenants, kicks
   off a device-code login prompt the user has no reason to deal
   with when nothing needs to be done. Fix: short-circuit at the
   top of both modes when the CR already exists, with a clear
   message and the `kubectl delete + retry` escape hatch for users
   who genuinely want to re-provision.

Tests: 784 CLI pass (+1 new for the FQ-name kubectl output).

* fix(bicep-fallback): print full Graph Explorer PATCH for portal visibility

The Bicep `Microsoft.Graph` extension cannot set `@odata.type`, so the
Bicep-created blueprint is a plain `Application` that the kars runtime
uses fine but that the Entra portal's Agents page does not show. The
existing docs explained the workaround as a one-field PATCH that
sets `@odata.type` only — that does upgrade the type but the entry
*still* stays hidden in the portal because the Agents-page filter
also requires `sponsors` and `owners` to be set.

Changes:

- agent_id_setup_bicep.ts: at the end of every Bicep success path
  (and also on the CRD-missing soft-failure branch), print a fully
  populated Graph Explorer PATCH body the user can paste verbatim.
  The body includes `@odata.type` + `sponsors@odata.bind` +
  `owners@odata.bind` so the resulting typed blueprint actually
  shows up under Entra portal → Identity → Agents.

  We try `az ad signed-in-user show --query id -o tsv` to auto-fill
  the user OID. In the very tenants where this matters (Microsoft
  corp Macs without Intune enrollment) that command also fails with
  AADSTS530084 — handled silently with a `<YOUR_USER_OID>` placeholder
  and a one-line hint on where to find the OID in the Entra portal.

- docs/agent-identity.md: replace the misleading one-field PATCH
  example with the full body, including notes about: no `#` prefix on
  `@odata.type` in the request body, `/v1.0/users/` required in
  `@odata.bind` even when the parent URL is `/beta/`, and a pointer
  to the CLI's auto-generated copy-paste body.

Tests: 786 CLI pass (incl. 15 agent_id_setup tests).

* docs+cli: full delete+recreate runbook for typed blueprint (SP+FIC included)

A user hit the case where the in-place @odata.type PATCH was rejected
and they recreated the blueprint via Graph Explorer's POST /applications.
The recreated app then lacked an SP (so it couldn't receive RBAC role
assignments and stayed invisible in the Agents portal) and lacked the
MI-as-FIC (so the controller couldn't mint child identities).

Bicep creates all three resources (app + SP + FIC) but the Bicep
Microsoft.Graph extension cannot set the @odata.type discriminator
needed for a typed agentIdentityBlueprint — so Graph-Explorer recovery
remains the only path in CA-blocked tenants, and it must do all three
steps explicitly.

- agent_id_setup_bicep.ts: extend printPortalVisibilityHint() to print
  the full 5-step recovery runbook (POST app, POST SP, POST FIC, kubectl
  patch CR, optional DELETE old app) underneath the simpler in-place
  PATCH path that's still tried first.

- docs/agent-identity.md: replace the one-line "delete + recreate"
  hand-wave with the full POST sequence with all body shapes, plus the
  kubectl jsonpath one-liner to extract tenantId + MI principalId from
  the cluster.

Tests: 786 CLI pass.

* fix(controller): own egress-guard script in one place + fix iptables rule order

The original `agent_id_egress_rules()` shipped on `feat/entra-agent-id`
contained a security bug: every rule used `-A OUTPUT` (append). The
pre-existing baseline egress-guard script (currently emitted as a
`concat!` literal in `reconciler/mod.rs`) starts with
`-A OUTPUT --uid-owner 1000 -o lo -j ACCEPT`, so any later appended
rule blocking UID 1000 → 127.0.0.1:8080 would NEVER fire — the
loopback-allow would match first. This would have silently broken the
"agent cannot impersonate the router and mint downstream tokens"
boundary as soon as sidecar injection is wired up.

Rubber-duck critique caught this (finding #5). Fix:

- Refactor `agent_id_egress_rules()` to emit two correctly-positioned
  rules. The sidecar-block uses `-I OUTPUT 1` (insert at chain head)
  so it runs BEFORE the baseline loopback-allow. The router-IMDS
  block can stay `-A` because no prior `--uid-owner 1001` rule exists.
- Drop the redundant UID 1000 → IMDS rule (catch-all `DROP UID 1000`
  already covers it).
- Drop the explicit UID 1002 → IMDS ACCEPT (OUTPUT chain default is
  ACCEPT and no UID 1002 restriction exists).
- New `build_egress_guard_command(agent_id_mode: bool)` composes the
  full shell script: agent-id rules first (so `-I` semantics + script
  text order both keep the security boundary), then the seven
  baseline iptables lines, then a mode-specific echo. `&&`-chained so
  any iptables failure aborts init-container startup (partial policy
  is worse than no policy).
- Reconciler/mod.rs replaces the `concat!` literal with a call to the
  new helper (currently always passing `false`; agent-id pass that
  flips to `true` lands in a follow-up commit on the same branch).

Tests:
- New `egress_rules_use_insert_before_baseline_loopback_allow` — pins
  the `-I OUTPUT 1` semantics. Direct regression for the security bug.
- New `egress_guard_command_legacy_mode_matches_existing_behaviour` —
  byte-for-byte pin on the seven historical iptables lines so the
  refactor is provably no-op for non-agent-id sandboxes.
- New `egress_guard_command_agent_id_mode_prepends_security_rules` —
  asserts the sidecar REJECT appears BEFORE the loopback ACCEPT in
  the script text (defence-in-depth even if a reader misreads -I).
- New `egress_guard_command_is_shell_safe_chained` — every step
  starts with `iptables ` or `echo `.

789/789 controller tests pass.

* feat(controller): per-sandbox agent identity provisioning + sidecar injection

The end-to-end controller-side of agent-id mode. New module
`agent_id_provisioning` orchestrates the full flow that the rubber-duck
critique identified as the missing glue: resolve mesh-auth mode →
provision (or recover) the per-sandbox Entra Agent Identity via Graph
→ patch sandbox status → materialise the per-namespace sidecar env
ConfigMap → return a Ready outcome that the sandbox reconciler uses
to inject the sidecar container, pin the router env, and flip the
egress-guard into agent-id mode.

Architecture: model (D) from the critique — the controller provisions
the per-sandbox identity BEFORE pod creation, status-pins the appId,
and the router uses that pinned appId in every sidecar request. No
per-sandbox Secret; no sandbox-side workload identity binding.

## What this commit adds

### New module: `controller/src/agent_id_provisioning.rs`

- `ProvisionerCache` — process-wide cache of `AgentIdentityClient`s
  keyed by blueprint client ID. Shares token caches and connection
  pools across concurrent sandbox reconciles.
- `resolve_mesh_auth_mode` — pure-function 3-way resolution (Auto →
  AgentId or Anonymous based on KarsAuthConfig readiness; explicit
  AgentId without ready config surfaces a distinct `AuthConfigNotReady`
  reason so operators can distinguish "tenant not set up" from
  "auto-fallback to anonymous").
- `load_auth_config` — singleton fetch with explicit `Ok(None)` for
  the 404 (anonymous-tier fallback) vs. `Err` for transient failures.
- `ensure_agent_identity_for_sandbox` — idempotent orchestration with
  the three-step recovery flow the critique specified:
    1. If `status.agentIdentity` recorded, GET via Graph; reuse on
       200, reprovision on 404, requeue on 5xx.
    2. If status empty, `list_cluster_agent_identities` filtered by
       the `kars-sandbox-uid:<uid>` tag. Catches the
       "Graph create succeeded but controller crashed before status
       patch" crash window.
    3. Otherwise create new, patch status, return.
- `materialise_sidecar_configmap` — copies the rendered sidecar env
  into the sandbox namespace (`envFrom` cannot cross namespaces — the
  critique caught this). Owned by the KarsSandbox via
  `ownerReferences` so K8s garbage-collects on sandbox deletion.

### Modified: `controller/src/reconciler/mod.rs`

- `Context` gains `cluster_uid` (read from `kube-system` ns metadata
  at startup — canonical "this cluster" identifier in K8s; falls back
  to `KARS_CLUSTER_UID` env or generated string with a warning) and
  `agent_id_cache: Arc<ProvisionerCache>`.
- `reconcile` calls `ensure_agent_identity_for_sandbox` BEFORE pod-spec
  assembly. Match on outcome: `Skipped` → legacy path; `Ready` →
  capture identity for downstream injection; `Failed` → patch
  Degraded status with `AgentIdentityProvisioningFailed` reason and
  requeue. No silent fallback to legacy on AgentId failure — explicit
  user intent must not be downgraded.
- Pod-spec assembly:
    - Egress-guard `command` flips to `build_egress_guard_command(true)`
      when `agent_id_active.is_some()` — adds the security-critical
      `-I OUTPUT 1` REJECT rule for UID 1000 → sidecar.
    - Router `env` gains `PINNED_AGENT_IDENTITY_APP_ID` (the
      per-sandbox appId) + `AUTH_SIDECAR_URL` (loopback :8080).
    - Sidecar container is appended to `containers` after the
      `runtimeClassName` block. Image pinned to the GA Microsoft
      distroless build (overrideable via `KARS_SIDECAR_IMAGE`).

### Modified: `controller/src/auth_config_reconciler.rs`

- After successful ConfigMap apply, `patch_ready_status` patches
  `phase=PHASE_READY` + `SidecarConfigMaterialized=True` condition.
  This is the real readiness signal that the sandbox reconciler's
  `resolve_mesh_auth_mode` gates on (per critique #7). Best-effort:
  status-patch failure is logged but doesn't fail the reconcile.
- `build_condition_blueprint_ready` no longer `#[allow(dead_code)]` —
  it's now consumed by `patch_ready_status`. Type renamed from
  `BlueprintReady` (Graph-call check) to `SidecarConfigMaterialized`
  (controller-side materialisation check) since the BlueprintReady
  signal will come from a separate cluster-health probe.

## Test status

- 795/795 controller tests pass (+6 new for the new module).
- phase_taxonomy_guard pre-existing integration test passes (the
  status patch uses `PHASE_READY` constant, not a string literal).
- Sidecar-injection tests still all pass (the iptables refactor from
  the previous commit holds).

## Not yet in this commit (deferred to follow-up commits/PRs)

- Inference-router `sidecar_client.rs` that actually consumes the
  pinned env vars (todays AgentIdentity unused on the router side).
- CLI changes: `kars up` VMSS-MI assignment; `kars mesh setup-trust
  verify` end-to-end audit; Foundry RBAC assignment print.
- `agent_identity_reaper` for cleaning up orphan SPs.
- Sponsor user object IDs in KarsAuthConfig spec (currently empty
  array passed; works in tenants where sponsors aren't required).

* fix(e2e): make Entra Agent ID chain work end-to-end on real AKS

Concludes a long live-debug session against kars-aks where we walked
the full controller → sidecar → router auth chain and fixed every
blocker until the sidecar successfully mints tokens from Entra and
relays them to the router. The remaining gate at end-of-chain is the
Microsoft corp tenant's Conditional Access policy on agent identities
(AADSTS53003 with a `capolids` claims challenge) — a tenant
configuration issue, not a code issue. The architecture is proven
correct.

- **Drop the GET-by-id verify on recorded identity.** Graph's
  `GET /servicePrincipals/{id}` has multi-second eventual-consistency
  after creation; treating a 404 as "stale, reprovision" creates a
  runaway-creation loop (live cluster produced 70+ duplicate SPs per
  minute before this fix). The reaper (separate PR) handles
  out-of-band deletes via tag-scoped listing.
- **Idempotency guard on `patch_sandbox_status`.** Skip the SSA patch
  when the recorded status already matches — same `lastTransitionTime`
  drift pattern that bit the auth-config reconciler.
- **Drop ownerReference on the per-namespace sidecar CM.** KarsSandbox
  lives in `kars-system`; the per-sandbox CM lives in `kars-<name>`.
  K8s rejects cross-namespace owner refs with `OwnerRefInvalidNamespace`
  and GC's the CM seconds after creation (debugged via the kubelet
  events stream). Cleanup happens via namespace deletion instead.
- **Add apiVersion+kind to status patch body.** SSA patches without
  the top-level type meta return `BadRequest: invalid object type:
  /, Kind=`.

- **IMDS-first, WI fallback.** Workload-Identity-derived tokens cannot
  be used as FIC assertions (Entra anti-loop AADSTS700231). The
  controller MUST acquire its MI token via IMDS — which requires the
  MI assigned to the AKS node-pool VMSS (`kars up` does this).
- **Permissive Graph list parser.** The agent identity list response
  uses `agentAppId`, not `appId`. The strict-schema deserialiser
  failed silently on every list call, so the tag-based recovery path
  never found anything and the controller created a fresh SP every
  reconcile. Fallback parser handles both field names.

- **Patch `phase=Ready`** so `resolve_mesh_auth_mode` has a real
  signal to gate on (rubber-duck critique #7). No-op when the
  observed status already matches the desired — avoids the SSA
  `lastTransitionTime` reconcile loop.

- **`AUTH_SIDECAR_URL=http://localhost:8080`** not `http://127.0.0.1:8080`.
  The Microsoft Entra SDK sidecar's HostFiltering middleware only
  allows `Host: localhost`; calls with `Host: 127.0.0.1` are rejected
  with `400 Bad Request - Invalid Hostname`. /etc/hosts in every pod
  maps `localhost` → `127.0.0.1` so the loopback semantics are
  identical.
- **`httpHeaders: Host=localhost:8080`** on the readiness/liveness
  probes. The kubelet sends the pod IP as the default Host header,
  which the sidecar rejects with the same 400. Without this override,
  every sidecar stays ready=false even when fully functional.
- **emptyDir at /app/keys.** The sidecar's ASP.NET Data Protection
  writes encryption keys to `/app/keys`; the distroless image has
  /app owned by root and UID 1002 cannot mkdir there. Per-pod
  emptyDir gives a writable scratch.

- Append the keys volume to `pod_spec.volumes` whenever the sidecar
  is injected. Without this the container immediately crashes on
  startup with `Access to the path '/app/keys' is denied`.

The Microsoft Entra SDK sidecar's response shape varies across
builds. We've observed three forms in the wild:

1. `{"AuthorizationHeader": "Bearer xxx", "ExpiresIn": 3600}` —
   documented contract.
2. `"Bearer xxx"` — JSON-quoted string only.
3. `Bearer xxx` — plain text, no JSON.

The strict parser broke on shape 2/3 with `parse auth-sidecar JSON
response`, silently disabling agent-id auth for the entire pod. New
`parse_sidecar_body()` handles all three lenient + 5 dedicated tests.

Without this rule the controller's IMDS call times out (CILIUM
default-deny drops 169.254.169.254). IMDS is per-node link-local —
not externally routable, no exfiltration risk.

- KarsAuthConfig.spec.agentId.sponsorUserObjectIds — required in
  tenants where Graph rejects agent identity creation without a
  sponsor (live error: `No sponsor specified`).
- KarsSandbox.status.agentIdentity — appId/objectId/displayName/
  createdAt. Without this in the CRD schema, SSA patches fail with
  `field not declared in schema`.

- Grant the controller `get/list/watch/create/update/patch/delete`
  on `karsauthconfigs` (and `/status`). Without this the auth-config
  reconciler stays dormant ("KarsAuthConfig CRD not installed").

- 795/795 controller tests pass (+ 0 net change; the existing
  sidecar_injection tests covered the URL constants we updated).
- 890/890 inference-router lib tests pass (+5 new
  `parse_sidecar_body*` regression tests).

- agent_identity_reaper module (cleans Graph orphans by tag — needed
  long-term to handle out-of-band deletes; for now duplicates require
  a manual cleanup script).
- CLI flow: `kars up` should run `az identity federated-credential
  create` for the controller SA so the WI fallback works in clusters
  without VMSS-attached MI.
- Per-sandbox Foundry RBAC automation — verify already prints the
  exact `az role assignment` commands; landing this would auto-grant.

```
✅ KarsAuthConfig provisioning (Bicep + Graph Explorer fallback)
✅ Controller MI assigned to AKS VMSSes
✅ Controller IMDS → MI → blueprint Graph token chain working
✅ Per-sandbox agent identity creation via Graph (with sponsors)
✅ Sandbox status.agentIdentity correctly recorded
✅ Per-namespace sidecar env ConfigMap materialized
✅ Sandbox pod contains 3 containers (openclaw + router + auth-sidecar)
✅ auth-sidecar healthy + listening on localhost:8080
✅ Router detects sidecar mode and routes auth through it (fail-closed)
✅ Sidecar mints Entra tokens for the pinned agent identity
✅ Sidecar relays token (or Entra error) back to router

⚠️ Foundry call returns 401 AADSTS53003 — Microsoft corp tenant's
   Conditional Access policy blocks agent identities. This is the
   same family of CA wall the user has been hitting all session for
   their human account. Tenant configuration issue; the chain itself
   is provably correct.
```

* fix(sidecar): bind on all interfaces so kubelet probes can reach /healthz

The Microsoft Entra SDK sidecar was bound to `127.0.0.1:8080` only,
which meant the kubelet's readiness/liveness probes (which connect to
the pod IP, not loopback) got `connection refused` at the TCP layer
before any HTTP exchange could happen. Overriding the probe's `Host`
header is useless when the connect itself fails.

Fix: bind on all interfaces (`http://+:8080`). Three defence-in-depth
layers keep cross-pod traffic out of the sidecar:

  1. Sidecar's HostFiltering middleware accepts only `Host: localhost`;
     anything else gets 400 Bad Request.
  2. The sandbox NetworkPolicy has no ingress rule allowing port 8080,
     so cross-pod packets are dropped before reaching the sidecar.
  3. Egress-guard iptables rule REJECTs UID 1000 → 127.0.0.1:8080,
     blocking in-pod privilege escalation from the agent container.

Verified live on kars-aks (sandbox kars-testrun):
  - auth-sidecar: ready=true, restarts=0
  - kubelet probe now succeeds via pod-IP connect + Host-header
    override
  - cross-pod connections are still rejected (no NP ingress rule)

This was the last container-level blocker keeping the chain from
running end-to-end. With this fix the sidecar mints real Entra tokens
and the router relays them to Foundry (where the only remaining gate
is the per-agent-identity Azure AI User role assignment).

* refactor(controller): pivot to shared-sidecar architecture (Phase 0)

This commit prepares the branch for the shared-sidecar redesign:
deletes the per-pod sidecar injection plumbing while preserving the
controller-side provisioning machinery, the Graph client, and all
hard-won bug fixes.

## What's removed

- `controller/src/sidecar_injection.rs` — per-pod container spec
  helpers (build_sidecar_container, build_router_pinned_identity_env,
  build_router_sidecar_url_env, egress-guard agent-id-mode rules).
  In the new design the auth-sidecar runs ONCE per cluster as a
  Helm-managed Deployment in `kars-system`; per-pod injection is no
  longer needed.
- `cli/src/commands/mesh/vmss_mi_assign.ts` — VMSS managed-identity
  assignment driver. Not needed when the sidecar runs in kars-system
  with its own Workload Identity binding.
- `materialise_sidecar_configmap` per-namespace mirror in
  `agent_id_provisioning.rs`. The shared sidecar consumes a single
  `kars-system`-scoped ConfigMap managed by `auth_config_reconciler`.
- The agent-id-mode iptables rules in the egress-guard init container
  (UID 1000 → REJECT 127.0.0.1:8080, UID 1001 → REJECT IMDS). Trust
  boundary for cross-pod sidecar access moves to a NetworkPolicy on
  the sidecar's namespace (added in a follow-up commit).

## What's preserved

- `controller/src/agent_identity.rs` — Graph client. Same provisioning
  endpoints, same auth flow. With or without per-pod sidecar.
- `controller/src/agent_id_provisioning.rs` — provisioning loop with
  idempotency, tag-based recovery from crashes, permissive Graph list
  parser, status patch idempotency. Hard-won lessons preserved.
- `controller/src/auth_config.rs` + `auth_config_reconciler.rs` —
  KarsAuthConfig CRD + its reconciler. Still cluster-level singleton
  config; the sidecar just consumes a kars-system-scoped CM now.
- `inference-router/src/sidecar_client.rs` — URL-agnostic client.
  The new design points it at the cluster Service DNS instead of
  loopback; no code change needed.
- All Helm CRD templates, RBAC additions, IMDS NetworkPolicy egress
  rule, the Bicep blueprint provisioning template, and the
  Graph Explorer recovery runbook docs.

## What's reset to baseline behaviour

- `reconciler/mod.rs` egress-guard: back to the seven-line baseline
  iptables script (the same one shipped on `main`). No agent-id-mode
  branching; the security boundary for cross-pod sidecar access is
  the sidecar-namespace NetworkPolicy, not pod-local iptables.
- Inference-router env injection: still pushes
  `PINNED_AGENT_IDENTITY_APP_ID` + `AUTH_SIDECAR_URL` when an agent
  identity is provisioned, but `AUTH_SIDECAR_URL` now points at
  `http://entra-auth-sidecar.kars-system.svc:5000` instead of
  loopback. Helm chart for the sidecar Deployment lands in the next
  commit.

## Why this redesign

Microsoft's own Agent ID design-pattern docs and the Auth Sidecar
API support `?AgentIdentity=<appId>` so one sidecar instance can
mint tokens for any of the blueprint's children. We had been running
one sidecar per pod (N × 128MB), iptables-isolating cross-UID access
to localhost:8080. Switching to one shared sidecar (kars-system,
2 replicas, NetworkPolicy-gated) saves ~5x memory at 10 sandboxes,
eliminates the per-pod ASPNETCORE bind + probe gymnastics, and aligns
with the Microsoft.Identity.Web library's centralized MSAL cache.

Independent research review confirmed the per-pod model is one of
two documented patterns but is more resource-intensive; the
`?AgentIdentity=<appId>` shared-sidecar pattern is explicitly
documented as a valid alternative and the SDK's `AzureAd__ClientId`
takes the blueprint appId, making each child mintable on demand.

## Tests

785/785 controller tests pass. Inference-router tests untouched.

* feat(helm): shared entra-auth-sidecar Deployment (Phase 1)

One auth-sidecar Deployment per cluster (2 replicas, HA) serves all
sandboxes via the Microsoft AuthorizationHeaderUnauthenticated
?AgentIdentity=<appId> query parameter — replacing the per-pod
sidecar injection model from earlier iterations.

Templates added (all gated on `entraSidecar.enabled`, default false):
- auth-sidecar-serviceaccount.yaml  - WI-annotated SA
- auth-sidecar-deployment.yaml      - distroless sidecar, /healthz probes
- auth-sidecar-service.yaml         - ClusterIP :5000
- auth-sidecar-networkpolicy.yaml   - ingress only from sandbox routers

values.yaml gets the `entraSidecar:` config block populated by
`kars up` from the KarsAuthConfig CR.

Trust boundary rationale and credential-supply patterns (WI for OSS,
IMDS-MI for corp tenant) explained inline in each template.

Resource footprint: ~160MB cluster-total vs ~1.28GB at 10 sandboxes
under the per-pod model.

Tests: helm lint clean, helm template renders 4 resources when
enabled, 0 when disabled. Controller 785/785 unchanged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): allow sandbox→kars-system:5000 egress for shared sidecar (Phase 2)

The shared entra-auth-sidecar lives in the kars-system namespace and
is reached by every sandbox's inference-router over TCP 5000. The
sandbox-policy NetworkPolicy now includes an explicit egress rule
selecting namespaces labeled
  app.kubernetes.io/name=kars
  app.kubernetes.io/component=system
on port 5000.

Also fixed the auth-sidecar's own ingress NetworkPolicy: the selector
now correctly targets pods labeled kars.azure.com/component=sandbox
(the pod label) rather than the non-existent
kars.azure.com/component=inference-router (which is a container, not
a pod, and so was never matchable).

Trust boundary is now two-sided:
- Sandbox-side: egress allows reaching only the kars-system namespace
  on port 5000, where the sidecar Service is the only listener.
- Sidecar-side: ingress allows only pods in sandbox-labeled namespaces.

When entraSidecar.enabled=false at the Helm level, no auth-sidecar
exists and the egress rule is a harmless no-op (no destination to
reach).

Tests: controller 785/785, helm lint clean, helm template renders
correctly with enabled=true and enabled=false.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(router): tid+principal+aud+exp pinning on sidecar tokens (Phase 3)

Wire the previously-orphaned sidecar_client module into the
inference-router and harden it against token substitution attacks.

CONTROLLER (reconciler/mod.rs):
- Extends agent_id_active tuple to carry tenant_id alongside agent
  identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via
  ProvisioningOutcome::Ready.auth_spec.
- Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode
  sandboxes. Without it, router-side tid pinning is disabled and
  warns at boot (insecure, dev-only).

ROUTER (sidecar_client.rs, auth.rs, lib.rs):
- pub mod sidecar_client; — fix the orphaned-module bug.
- WorkloadIdentityAuth now consults SidecarClient first; sidecar mode
  is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves
  per-sandbox audit attribution in downstream Azure RBAC.
- from_env() now Result<Option<Self>>. Partial config (only URL or
  only PINNED set) returns Err; WorkloadIdentityAuth::new() panics →
  AKS surfaces as CrashLoopBackoff. No more silent fallback to a
  different identity model.
- HARD-fail validate_token_claims on EVERY check (rubber-duck #1):
    * tid mismatch / missing (cross-tenant guard)
    * appid/azp mismatch when present (audit attribution guard)
      — if both present, BOTH must match
    * aud mismatch against per-service expected set (cache-poisoning
      guard) — Foundry, Graph, OpenAI, Management, Search mapped
    * exp in past or within 60s skew (stale-token guard)
    * exp missing when tid-pinning enabled (unbounded-lifetime guard)
- Returns a TTL cap = exp - now - 60s; caller takes
  min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching.
- JWT decoder tightened: requires EXACTLY 3 segments (rejects
  2-segment unsecured JWS and 4-segment JWE).
- aud claim normalized: accepts both single-string (Entra default)
  and array (RFC 7519); rejects out-of-spec shapes (number, bool).
- Soft WARN when both appid and azp are absent (some MSI tokens
  legitimately omit both).
- New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning
  state for operator visibility.

TESTS:
- 39 sidecar_client unit + wiremock integration tests, covering all
  HARD-fail branches, the cache-skip on validation failure, the TTL
  cap math, the partial-config Err path, JWT decoder edge cases
  (2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud,
  out-of-spec aud).
- Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client.

Rubber-duck review caught: (1) appid/azp should be hard not soft, (2)
partial env config must fail closed, (3) 3-segment requirement, (4)
aud validation against requested resource, (5) exp validation + TTL
cap. All five addressed in this commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(cli,bicep,controller): dual credential mode auto-detect (Phase 4)

Adds first-class support for two auth-sidecar credential supply
patterns and auto-detects which one works in the current Entra tenant.

PATTERN A (ManagedIdentityImds, default):
- Sidecar auths via SignedAssertionFromManagedIdentity against the
  controller MI's IMDS endpoint. Corp-tenant safe; required when
  the tenant's FIC issuer-allowlist policy blocks AKS OIDC
  (Microsoft-corporate: InvalidFederatedIdentityCredentialValue).
- Bicep creates: blueprint + SP + controller MI + MI-as-FIC.

PATTERN B (WorkloadIdentity):
- Sidecar auths via SignedAssertionFilePath against the projected
  K8s SA token. No per-cluster MI, no VMSS identity assignment.
- Bicep creates: blueprint + SP + SA-as-FIC pointing at the AKS
  cluster's OIDC issuer URL.

CRD (controller/src/auth_config.rs):
- Added `controller.credentialMode` enum (ManagedIdentityImds default).
- Made `controller.managedIdentity{Client,Resource,Principal}Id`
  Optional (required only in MI mode).
- Added is_valid_for_mode() validator.

CONTROLLER (auth_config_reconciler.rs):
- render_sidecar_env branches on credentialMode and emits the
  correct AzureAd__ClientCredentials__0__SourceType. MI mode
  emits ManagedIdentityClientId; WI mode emits
  SignedAssertionFileDiskPath=/var/run/secrets/azure/tokens/azure-identity-token.
- Spec validation runs BEFORE ConfigMap materialisation. Invalid
  specs (MI mode + empty clientId) surface as phase=Degraded
  with InvalidCredentialMode condition — refuses to propagate
  the misconfiguration to running sandboxes.
- New patch_degraded_status helper.

BICEP (deploy/bicep/agent-id-trust.bicep):
- New credentialMode parameter (allowed: ManagedIdentityImds |
  WorkloadIdentity).
- Conditional resources: controller MI + MI-as-FIC only in
  Pattern A; SA-as-FIC only in Pattern B.
- New aksOidcIssuerUrl parameter (required for Pattern B).
- New outputs: credentialMode, aksOidcIssuerUrl.
- az bicep build clean, no warnings.

CLI:
- ensureAgentIdTrust now accepts credentialMode (auto | WorkloadIdentity
  | ManagedIdentityImds), aksClusterName, aksClusterResourceGroup.
- AUTO MODE: tries Pattern B first by discovering the AKS OIDC
  issuer URL via az aks show, then creating an SA-as-FIC. On
  InvalidFederatedIdentityCredentialValue (corp tenant signature),
  falls back to creating the controller MI + MI-as-FIC.
- New TenantRejectedAksOidcIssuer sentinel for clean fallback
  orchestration.
- New discoverAksOidcIssuerUrl helper — accepts explicit args or
  walks kubeconfig + az aks list to guess.
- writeKarsAuthConfig strips MI fields when credentialMode is
  WorkloadIdentity (matches the controller's Optional schema).
- setup-trust gains --credential-mode, --aks-cluster-name,
  --aks-cluster-resource-group, --aks-oidc-issuer-url flags.
- Bicep wrapper threads credentialMode + aksOidcIssuerUrl through
  to ARM.

TESTS:
- Controller: 789/789 (4 new — WI mode render, MI empty-field
  rendering, is_valid_for_mode in MI and WI modes).
- Router: 1014/1014 (unchanged).
- CLI: 786/786 (2 new — credentialMode propagation through
  dry-run with default + explicit).
- Bicep: az bicep build clean.
- Helm lint + template clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller,bicep,docs): security alignment (Phase 5)

Closes the rubber-duck research findings from the Phase 3 critique
and Microsoft's Entra Agent ID design-patterns audit.

CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM):
- KarsSandbox.spec.meshAuth.customSecurityAttributes:
  BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which
  attributes (from a tenant-declared set) the controller should
  PATCH onto each per-sandbox agent identity.
- AgentIdentityClient::patch_custom_security_attributes Graph
  client method. Constructs the documented CustomSecurityAttribute-
  Value envelope with the required @odata.type per attribute,
  inferred from the JSON value shape.
- odata_type_for_value helper: maps serde_json::Value →
  '#String' | '#Int32' | '#Boolean' | '#Collection($T)' and
  rejects floats, nulls, nested objects, mixed-type arrays, and
  empty arrays with clear error messages BEFORE the call goes out.
- Wired into ensure_agent_identity_for_sandbox: PATCH runs on
  every reconcile (idempotent on Graph). Failures surface as
  ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing
  silent missing-attribute drift.

SCALE-OUT INVARIANT (rubber-duck #4):
- Documented in agent_id_provisioning.rs module doc: the agent
  identity is keyed on KarsSandbox.metadata.uid (and cluster UID),
  with NO per-pod / per-replica / per-ordinal dimension. All
  replicas of one KarsSandbox share ONE agent identity.
- New test tag_layout_excludes_per_pod_attributes pins the tag
  layout — any future PR that adds a 'kars-pod-' / 'kars-replica-'
  / 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix
  breaks the test.
- Visibility change: AgentIdentityClient::tags_for is now
  pub(crate) so the cross-module invariant test can call it.

BOOTSTRAP SCRIPTS (deploy/bicep/standalone/):
- custom-security-attributes.sh: declares the recommended
  AgentGovernance set with 4 attributes — AgentClassification
  (Standard|Restricted|Confidential), DataSensitivity
  (Public|Internal|Confidential), ProductOwner, ManagedBy.
  Idempotent via az rest against the Graph beta endpoint.
  (The Microsoft.Graph Bicep extension does not yet ship typed
  attributeSets / customSecurityAttributeDefinitions resources,
  so this is a shell script that operators run once per tenant.)
- conditional-access-baseline.sh: applies Microsoft's
  policy-autonomous-agents template — blocks sign-ins where the
  risk level meets the configured threshold (default: high),
  targeted via the ManagedBy=kars-controller attribute filter.
  Defaults to report-only state for safe rollout. Idempotent via
  upsert.

FOUNDRY RBAC (research R5):
- foundry-rbac.bicep grants Azure AI User on a Foundry resource
  to the BLUEPRINT SP (not per-agent). All derived agent
  identities inherit access — eliminates per-agent role-assignment
  churn. Supports RG-scoped (default) and resource-scoped
  assignment via foundryResourceName parameter. az bicep build
  clean, no warnings.

DOCS:
- docs/architecture/entra-agent-id/05-security-alignment.md:
  6.8KB operator-facing runbook covering the bootstrap order,
  KarsSandbox YAML example, failure modes table, scale-out
  invariant rationale, Foundry RBAC inheritance.

TESTS:
- Controller: 802/802 (+13 new — 12 odata_type_for_value variants
  covering supported + rejected shapes, 1 scale-out invariant).
- Bicep: foundry-rbac.bicep builds clean. JSON compiled.
- Bash: both .sh scripts pass 'bash -n' syntax check.
- CLI: 786/786 (unchanged).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(phase-7): live deploy bug fixes + Foundry RBAC fallback (Phase 7)

End-to-end live validation of the Phase 0-5 shared-sidecar architecture
against real Microsoft corp tenant (kars-aks). Three discoveries
addressed in this commit:

1) ASP.NET Core HostFiltering rejects Service-DNS callers (router fix)
   - Microsoft Entra SDK auth-sidecar's HostFiltering middleware rejects
     non-localhost Host headers regardless of AllowedHosts=* env.
   - inference-router/src/sidecar_client.rs: always send
     Host: localhost:5000 on sidecar requests.
   - deploy/helm/kars/templates/auth-sidecar-deployment.yaml: set
     AllowedHosts=* env for completeness (sidecar config source doesn't
     honour it for dynamic-binding path).
   - NetworkPolicy remains the real ingress boundary; bypassing host
     filter does not weaken security.

2) Azure RBAC inheritance correction (sandbox_bringup.ts)
   - Phase 5 originally assumed Foundry RBAC would inherit blueprint ->
     derived agent identity SPs.
   - Microsoft docs (concept-agent-id-design-patterns) clarify that
     inheritable permissions = Microsoft Graph delegated permissions only.
     Azure RBAC is per-principal.
   - Confirmed empirically live: granting on blueprint SP alone did NOT
     unblock chat; direct grant on each agent identity SP did.
   - cli/src/commands/up/sandbox_bringup.ts: inline Bicep now grants
     Cognitive Services OpenAI User on the blueprint SP as a
     break-glass / fallback (kept for the case where the controller
     can't grant per-agent). Inheritable-permissions language removed.
   - Operator-facing error message documents the per-agent grant path
     when ARM deployment fails for permission reasons.
   - Phase 5b (future PR): controller must assign Azure RBAC per agent.

3) Multi-agent exec-brief demo: chain proven multi-agent
   - Demo applies CRDs, brings up parent execbrief + 3 sub-agents.
   - Each sub-agent gets its own typed Entra agent identity SP.
   - Verified live (with all 4 SPs granted Cognitive Services OpenAI
     User AND Azure AI User on the Foundry account via az rest):
     * All 4 sandboxes booted in fail-closed sidecar mode
     * 65 successful Foundry 200s across all 4 sandboxes in 10 min
     * ZERO PermissionDenied responses after roles in place
     * Mesh routing: parent dispatched to analyst, replied
     * E2E encrypted relay: file transfer (analyst.json) ACKed
     * AGT trust scoring: +0.8 reputation submitted, accepted=true
     * NetworkPolicy: 0 egress denials, 0 ingress drops
   - Demo's verify.json shows 3/9 checks pass — passing checks are the
     kars-runtime mechanisms (sub-agents active, egress 0 denials,
     telegram skipped). 6 failing checks ALL trace to one root cause:
     Foundry project's Bing Grounding connection is not configured
     (foundry_web_search requires bing project_connection_id). This is
     orthogonal to our work; the auth chain is proven.

TESTS:
- inference-router sidecar_client: 39/39 pass.
- CLI: 786 pass + 2 skipped (no regressions).
- Bicep: existing modules build clean.
- Helm: lint clean, template renders correctly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(entra-agent-id): top-level README + migration guide (Phase 8)

Consolidates the entra-agent-id architecture documentation into a
navigable index and adds a migration guide for operators upgrading
from earlier per-pod-sidecar branch heads.

- docs/architecture/entra-agent-id/README.md (new top-level index):
  * TL;DR architecture diagram
  * Pattern A/B selection table
  * Phase ledger linking each commit to its scope
  * Key files map
  * Live validation snapshot (verified on kars-aks 2026-05-28)
- docs/architecture/entra-agent-id/00-poc-archive.md:
  * Previous POC README, renamed to archive. The POC scaffolding
    drove early design but is no longer the canonical reference.
- docs/architecture/entra-agent-id/04-migration-guide.md (new):
  * Step-by-step upgrade path from per-pod-sidecar branch heads
  * Helm upgrade command with the entraSidecar values
  * Per-agent role grant runbook (az rest workaround for the
    CA-blocked az role assignment create path)
  * Rollback procedure
  * Known caveats (HostFiltering, MSAL cache per replica,
    Bing Grounding orthogonal issue)

The numbered ordering reflects deployment lifecycle:
  00 — historical POC (archived)
  01 — runtime token flow
  02 — alternative ACI flow (reference)
  03 — original POC findings
  04 — migration guide (this commit)
  05 — security alignment (Phase 5)

Closes Phase 8 of the feat/entra-agent-id PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): per-agent ARM RBAC + agent identity cleanup (Phase 5b)

Eliminates the manual az role assignment runbook step from the Phase 5
migration guide by having the controller assign and revoke Azure RBAC
roles per agent identity automatically. Closes the orphan SP risk by
adding agent-identity deprovision to the KarsSandbox deletion
finalizer.

CRD (auth_config.rs):
- New `KarsAuthConfig.spec.foundryRbac: Vec<FoundryRbacAssignment>`.
- Each assignment carries an ARM `scope` and a list of built-in role
  definition GUIDs to PUT against every per-sandbox agent identity SP
  at provisioning time.
- Empty list (default) preserves the manual-grant workflow for
  backward compatibility with existing kars deployments.

ARM REST extensions (agent_identity.rs):
- `arm_token()` — MI token for management.azure.com audience.
- `assign_role_to_agent_identity()` — PUT roleAssignment with a
  deterministic UUIDv4 derived from (scope, principal, role), so
  repeated PUTs are idempotent on Azure's side. Treats 200/201 as
  success and 409 RoleAssignmentExists as success.
- `delete_role_assignments_for_principal()` — GET assignments
  filtered by `principalId` (NOT combined with `atScope()` — Azure
  REST returns 400 UnsupportedQuery), narrows to the requested scope
  client-side, then DELETEs each. Used by the deletion finalizer.
- New helpers: `deterministic_assignment_guid()` (SHA-256 → UUIDv4)
  + `extract_subscription_id()` (scope parser).
- 9 new unit tests cover GUID stability, case-insensitivity, UUID
  format, scope parsing happy-path + rejected forms.

Provisioning (agent_id_provisioning.rs):
- After Graph create (step 3c): iterate `spec.foundry_rbac` and assign
  every (scope, role) tuple to the new identity. WARN-and-continue on
  failure so the sandbox still boots; ARM RBAC converges on retry.
- Early-return path (recorded identity): RE-ASSERT the same role
  assignments so existing sandboxes pick up retroactive RBAC config
  the first time the operator adds `foundryRbac` to KarsAuthConfig.
  Idempotent via the deterministic GUID.
- Logs `Phase 5b reconcile: re-asserting ARM role assignments` with
  `foundry_rbac_entries` count for visibility.

Cleanup (agent_id_provisioning.rs + reconciler/mod.rs):
- New `cleanup_agent_identity_for_sandbox()` orchestrates the two
  Azure-side teardown steps on sandbox delete:
    1. DELETE role assignments held by the agent identity SP at every
       configured `foundryRbac` scope.
    2. Graph DELETE the agent identity SP itself.
- Hooked into the KarsSandbox deletion finalizer in
  reconciler/mod.rs. Best-effort: failures logged as WARN but do not
  block finalizer removal, so the K8s CR doesn't get stuck Terminating
  when Azure is degraded. The existing orphan reaper backstops what
  this path misses.

VERIFIED LIVE on kars-aks:
- All 5 existing sandboxes' agent identities (execbrief + 3 sub-agents
  + kars-testrun) now have exactly 2 role assignments each on the
  Foundry account (Cognitive Services OpenAI User + Azure AI User),
  re-asserted by the new code on the early-return path.
- Deleting `kars-testrun` sandbox: controller logged
    "deleted role assignments for principal"  →  2 deleted
    "agent identity SP deprovisioned"
  Verified Foundry now reports 0 role assignments for the deleted
  appId — full cleanup successful.

Tests: 811/811 controller (+9 new for the GUID + scope helpers).
Build clean, helm template + lint clean.

One-time operator prerequisite: the controller's identity (the MI
whose IMDS token the controller uses for ARM calls — typically the
AKS kubelet/agentpool MI) must hold
`Microsoft.Authorization/roleAssignments/write` at every scope listed
in `foundryRbac`. The kars-recommended way is to grant Role Based
Access Control Administrator on the Foundry RG. This is intentionally
a one-time operator action documented in the migration guide rather
than auto-granted, since it crosses cluster/customer trust boundaries.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): scaffold MeshAuthBackend CRD field for Phase 6 design

Adds KarsAuthConfig.spec.meshAuthBackend (enum: Anonymous default,
EntraAgentIdentity opt-in) and optional meshAuthAudience override for
the next milestone — verified Entra-signed AGT mesh peer authentication.

Defaults preserve full backward compatibility (every existing cluster
keeps registering anonymously with no behaviour change). Operators on
clusters that have completed sidecar-based entrypoint mint + relay JWKS
verification (the next-PR work) flip the field to EntraAgentIdentity.

Includes a design doc (docs/architecture/entra-agent-id/06-mesh-trust-design.md)
capturing the three independent pieces that must land for end-to-end
enforcement (CRD scaffold here; sandbox entrypoint via sidecar; AGT
relay JWKS verification — last piece is upstream-coordination work).

Unit tests pin: default is Anonymous (backward compat), both variants
deserialize, unknown variants are rejected (no silent fallback).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(router,sandbox): /v1/mesh-token endpoint + Entra-signed mesh peer path

Phase 6.b — sandbox-side path for verified AGT mesh trust. When
KarsAuthConfig.spec.meshAuthBackend=EntraAgentIdentity, the controller
sets MESH_AUTH_BACKEND on the router, the router exposes
GET /v1/mesh-token, and entrypoint.sh acquires a verified-tier agent
identity token via the shared auth-sidecar instead of doing the legacy
direct Workload-Identity → Entra exchange.

Router changes:
- New inference-router/src/routes/mesh_token.rs route gated on
  MESH_AUTH_BACKEND; returns 404 when disabled, 503 if sidecar not
  configured, 502 on sidecar call failure, 200 with {access_token,
  token_type, expires_in} on success
- Extended sidecar_client::resource_to_service_name to map
  api://agentmesh* → AgentMesh service key
- Optional MESH_AUTH_AUDIENCE env override (defaults to
  api://agentmesh/.default)
- 7 new mesh_token tests + agent-mesh resource-mapping test, all
  serialised on a process-wide ENV_LOCK mutex to avoid env races

Controller changes:
- reconciler/mod.rs injects MESH_AUTH_BACKEND + MESH_AUTH_AUDIENCE env
  on the sandbox's router container when the CRD field is set
- auth_config_reconciler.rs auto-emits the DownstreamApis__AgentMesh__*
  cluster onto the shared sidecar when meshAuthBackend=EntraAgentIdentity
  (skipped if operator has supplied an explicit AgentMesh entry)
- 4 new reconciler tests pin: anonymous default emits nothing, entra
  variant emits expected env, custom audience honoured, operator-
  supplied entry wins

Sandbox changes:
- entrypoint.sh: new branch BEFORE the legacy WI-exchange one. When
  MESH_AUTH_BACKEND=EntraAgentIdentity, curl localhost:8443/v1/mesh-token,
  export AGT_OAUTH_TOKEN on success, force AGT_TRUST_THRESHOLD=0 on
  failure (anonymous-tier fail-open contract matches existing logic)
- Backward compat: when MESH_AUTH_BACKEND is unset (default), the
  entrypoint follows the exact same path as before — no behaviour
  change for existing clusters

Test results:
- 819 controller tests pass (was 815, +4)
- 930 router tests pass (was 923, +7)
- entrypoint.sh sh…
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 31, 2026
…cal-k8s, docker (#368)

* docs(security-validation): cross-platform validation report — AKS, local-k8s, docker

Comprehensive security & runtime validation across all three deploy
modes, mapped to the 9-layer model in docs/security.md.

Per-platform live evidence collected:
- CRD presence + InferencePolicyCompiled/ToolPolicyCompiled/EgressAllowlistCompiled digests
- Pod securityContext (UID, readOnlyRootFilesystem, capabilities, seccomp)
- Workload Identity / Entra Agent ID federated token plumbing (AKS only)
- entra-auth-sidecar token issuance + per-sandbox pinned_agent_id
- Output authenticity (real Foundry calls, real URLs verified via HTTP 200)
- AGT mesh KNOCK + E2E channel establishment per agent
- Native AGT governance modules (PolicyEngine, AuditLogger, etc.)
- NetworkPolicy enforcement + egress-guard caps
- 9-check verify run on every platform

Verified all three platforms 9/9 PASS with current main checks.py.

Findings (no fixes applied — tracked for separate PRs):
  #1 HIGH (docker macOS) — UID 1000 reads Foundry API key (Docker Desktop UID virtualization)
  #2 HIGH (local-k8s) — API key + GitHub token in plaintext pod env (should be secretKeyRef)
  #3 MEDIUM (docker) — NET_ADMIN on persistent parent container (vs init-only in K8s mode)
  #4 MEDIUM (AKS, local-k8s) — blocklist-refresh CronJob failing every 6h (VAP collision)
  #5 MEDIUM (all) — audit JSONL not persisted (RO root FS, no emptyDir mount)
  #6 LOW (AKS) — AllowlistVerified=False on execbrief (inline endpoints, no cosign attestation)
  #7 LOW (AKS) — TrustGraph router-side enforcement not active (documented roadmap)

All findings include a remediation plan in §10. None block the AKS
verified-tier security story documented in docs/security.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>

* docs(security-validation): add env-var inventory addendum for AKS containers

Adds detailed per-container env-var analysis answering the question:
'do AKS containers have more env variables than they should?'

Per-container inventories with full categorization:
- openclaw container: 42 vars in 8 categories
- inference-router container: 54 vars (router-only paths/toggles)

Three additional findings on env-var hygiene:
  #8 LOW — OPENCLAW_GATEWAY_TOKEN exposed via env (should be file mount)
  #9 LOW — enableServiceLinks=true leaks internal cluster IPs (16 env vars)
  #10 LOW — possibly-redundant Foundry/Mesh-auth env vars on openclaw

Headline confirmations:
  ✅ NO AZURE_OPENAI_API_KEY on either AKS container
  ✅ NO COPILOT_GITHUB_TOKEN on either AKS container
  ✅ Federated identity token mounted RO, never as env value
  ✅ Auth mode 'shared entra-auth-sidecar fail-closed, no WI/IMDS/API-key fallback'

Comparison table across all 3 platforms included.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>

---------

Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request Jun 4, 2026
… working

Lands the protocol-correct fixes needed for MeshClient.connect() →
KNOCK → X3DH → Double Ratchet roundtrip between two sandboxes. Tested
end-to-end on kind-kars-dev with two Hermes pods (execbrief-hermes and
smoke-hermes) on the FRESHLY BUILT image (no hot patches):

- pod A registers, uploads prekey bundle, opens relay WS (with POP)
- pod B does the same
- pod A discovers B via /v1/discover (freshest-first sort)
- pod A fetches B's bundle, runs X3DH, sends KNOCK + first ciphertext
- pod B's _handle_knock_frame auto-accepts via SecureChannel.create_receiver,
  decrypts plaintext 'hello from execbrief-hermes'
- pod B replies via send_by_did → encrypted message frame
- pod A decrypts 'pong from smoke-hermes'

## Critical protocol fixes

1. **Relay WS connect-frame POP** (relay_transport.py)
   - Was: {type:'connect', from:did, ts:...}
   - Now: full proof-of-possession (std-base64 pub_key + iso ts + sig
     over ts), per AGT relay/app.py::_verify_connect_pop
   - Without this, the relay rejects every connection with
     'connect frame missing did/public_key/timestamp/signature'

2. **Registry auth header** (registry_client.py)
   - Was: three separate X-Agent-DID/Timestamp/Signature headers,
     signature over method+path+ts
   - Now: single 'Authorization: Ed25519-Timestamp <did> <ts> <b64url-sig>',
     signature over timestamp string only
   - Matches AGT registry/app.py::verify_ed25519_timestamp_auth

3. **X3DH bootstrap missing** (client.py)
   - Now connect() builds X3DHKeyManager + generates signed_pre_key
     + 10 OTKs + uploads bundle via PUT /v1/agents/{did}/prekeys
   - Without this, peers couldn't fetch our bundle, X3DH initiation
     would fail at the responder side

4. **KNOCK responder implemented** (client.py::_handle_knock_frame)
   - Was: log-only stub ('responder path not implemented')
   - Now: parses ChannelEstablishment, calls SecureChannel.create_receiver,
     caches the channel, decrypts the bundled first ciphertext,
     eagerly tops up the OTK pool for the next session

5. **Send fuses KNOCK + first message** (client.py::send_by_did)
   - First call to a new peer DID sends {type:'knock', establishment, ciphertext}
   - Subsequent calls send {type:'message', ciphertext}
   - Matches the TS SDK wire convention (one RTT, not two)

6. **AAD directionality fix** (client.py)
   - Initiator: f'{self_did}|{peer_did}'
   - Responder: f'{from_did}|{self_did}' (reconstructs the same bytes)

7. **EncryptedMessage wire format** (client.py)
   - Was: JSON of em.__dict__ (would fail at decoder)
   - Now: EncryptedMessage.serialize() / .deserialize() (binary + b64url)

8. **PeerBundle flat shape** (registry_client.py + client.py)
   - Was: nested dicts mirroring my best-guess wire format
   - Now: matches agentmesh.encryption.x3dh.PreKeyBundle's flat dataclass

9. **register_self handles 409 gracefully** (registry_client.py)
   - Was: raised MeshRegistryError, blocking every restart
   - Now: logs and continues — the subsequent prekey PUT (with
     Ed25519-Timestamp auth) proves we own the same key

10. **discover() sorts freshest-first** (registry_client.py)
    - Avoids hitting stale ghost-DIDs when a sandbox restarts with
      a new identity before the prior registration ages out

## Tests

- 9 kars-agt-mesh unit tests pass
- 83 Hermes unit tests pass
- Live bidirectional roundtrip verified on freshly-built image
  (build hash c1dcdfc11475... loaded into kind-kars-dev)

## Security audit updated

docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
- Residual risk #1 (no KNOCK responder) removed — now implemented.
- Added residual risk #4 (stale registry entries — non-security).
- Added live bidirectional test description.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request Jun 8, 2026
User report:
  > operator says "✓ Spawned" then nothing visible
  > kubectl get karssandbox -A confirms the CR was never created

Two compounding silent-failure bugs:

1. kars add was log-then-exit-0 on caught errors.
   The outer catch at cli/src/commands/add.ts line 601 (was: 531)
   handled every exception by calling spinner.fail() + console.error()
   and then RETURNING — letting Node exit 0 naturally. So
   `kubectl apply -f -` failing (CRD missing, wrong context, schema
   rejection on the bundle, etc.) surfaced as a clean exit code to
   any caller. Operator's `execa("kars", args, { stdio: "pipe" })`
   only logs `✗ Spawn fail` when execa REJECTS, so silent exit-0
   masked every kars-add failure mode behind a green checkmark.

   Fix: add `process.exit(1)` after the error logs. Preserves all
   the existing error-message branching (controller-not-installed
   hint, generic error text) — just stops lying about exit status.

2. Operator's spawn dialog was throwing away the real error text.
   Previously logged only `(e.stderr || e.message)?.substring(0, 200)`
   — execa's `.message` is usually `Command failed with exit code 1:
   kars add ...`, NOT the underlying kars-add stderr. So even after
   fix #1, the operator log would show "✗ Spawn fail: Command failed
   with exit code 1: kars add testhermes --runtime hermes ..." with
   no actual root cause.

   Fix: prefer e.stderr (now populated thanks to fix #1) over
   e.message, strip ANSI colour codes that kars add emits via chalk,
   filter empty lines, keep the last 4 (which is where spinner.fail
   + error hints live), join with " | ", cap at 400 chars. Activity
   log now shows e.g.:

     ✗ Spawn fail: Failed to create sandbox | Error: kubectl error:
     KarsSandbox.kars.azure.com "testhermes" is invalid: spec.hermes:
     Invalid value: ... | Connect: kars connect testhermes

   Also: on SUCCESS, echo the last 3 lines of stdout (the
   "Namespace / Model / Status / Connect" hints kars add prints) so
   the operator sees useful follow-up info inline.

Verified:
  npm run build + typecheck ⇒ clean
  vitest run                ⇒ 798 passed
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Jun 9, 2026
* fix(hermes): A1 docker smoke fixes — version pin + plugin opt-in

Two real bugs surfaced when running the first `docker build` +
end-to-end smoke test of the Hermes sandbox image:

1. **Hermes version pin wrong**
   `ARG HERMES_VERSION=0.5.1` doesn't exist on PyPI. The 0.5.x
   assumption came from misreading the Hermes README's Homebrew
   formula tag (`5.1.14`); the actual `hermes-agent` PyPI package
   uses 0.x.y numbering at 0.15.2 latest. Bumped to 0.15.2.

   Hermes 0.15.2's plugin contract (PluginContext.register_tool,
   register_hook, plugin.yaml with provides_tools/provides_hooks,
   discovery via `$HERMES_HOME/plugins/`) matches what the A1
   plugin code was already built for — verified by importing
   hermes_cli.plugins and running discover_plugins() against our
   materialized plugin tree.

2. **ripgrep not in Azure Linux 3**
   `tdnf install -y` exits non-zero if ANY package is missing, and
   Azure Linux 3 doesn't ship ripgrep. Hermes' built-in file_search
   tool prefers ripgrep but falls back to grep, so dropping it is
   safe. Image now builds in ~30s.

3. **kars plugin discovered but not loaded**
   Hermes treats `standalone` plugins as opt-in via
   `plugins.enabled` in config.yaml. The entrypoint was placing the
   kars plugin into `$HERMES_HOME/plugins/kars/` (correct user
   discovery path), but never adding `kars` to the enabled
   allow-list — so it was discovered and silently skipped with
   `error='not enabled in config'`.

   The entrypoint now emits a `plugins.enabled: [kars]` block at
   the top of every generated config.yaml. The awk-merge that
   replaces prior `mcp_servers:` blocks was extended to also
   replace prior `plugins:` blocks so re-runs are idempotent.

Verified end-to-end:
- `docker build` succeeds
- `discover_plugins()` loads kars plugin, registers 10 tools +
  2 hooks (pre_tool_call + post_tool_call)
- Entrypoint generates correct config.yaml with both blocks
- `$HERMES_HOME/plugins/kars/` materialized from
  `/opt/kars-hermes-stage/plugins/kars/` on every boot
- 83/83 python unit tests still pass inside the image
- Mock smoke run: `python3 -m hermes_cli.plugins discover` shows
  kars: enabled=True, 17 total plugin tools across all enabled
  plugins (10 from kars + 7 web/foundry from bundled providers)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(hermes): A1.2 CRD schema + gitignore for cross-compile artifacts

Two follow-ups from the kind-cluster end-to-end smoke test:

1. **Helm CRD schema missing Hermes enum** — controller's `crd.rs`
   added `RuntimeKind::Hermes` in a7882b8 but the matching Helm
   CRD YAML wasn't updated. Result: the API server rejected every
   KarsSandbox with `runtime.kind: Hermes` BEFORE the controller
   ever saw it. Verified by `kubectl apply --dry-run=server`
   failing with "unknown enum value 'Hermes'".

   Added:
   - `Hermes` to the `runtime.kind` enum at line 85
   - x-kubernetes-validations rule:
     `(self.kind == 'Hermes') == has(self.hermes)`
   - `runtime.hermes` properties block mirroring `pydanticAi`
     shape (version, agentCode oci/git, entrypoint, extraEnv)

   After the fix, `kubectl apply -f /tmp/hermes-sandbox.yaml`
   succeeds, controller picks up the CR, and a 2-container pod
   (`agent` + `inference-router`) reaches `2/2 Running` with the
   kars plugin loaded (10 tools + 2 hooks registered).

2. **`.cargo-docker/` not gitignored** — when cross-compiling for
   linux/arm64 via `docker run -v $PWD:/work … cargo build` (the
   pattern used for kind-on-M-series), `CARGO_HOME=/work/.cargo-docker`
   keeps container-arch crate cache out of the host's `~/.cargo`.
   That directory was leaking into `git status`. Added rules:
   - `.cargo-docker/` — explicit
   - `/bin/` was already covered by `**/[Bb]in/*` (verified)

Verified end-to-end on kind cluster `kars-dev`:

  $ kubectl get karssandbox,pods -n kars-smoke-hermes
  NAME           PHASE   RUNTIME   INFERENCEPOLICY   ISOLATION
  smoke-hermes           Hermes    smoke-inference   standard
  NAME                            READY   STATUS    RESTARTS
  smoke-hermes-697c6bd557-q5xfr   2/2     Running   0

  Plugin discovery inside the pod:
    kars plugin: enabled=True, source=user
    hooks      : {'pre_tool_call': 1, 'post_tool_call': 1}
    tools      : http_fetch, kars_discover, kars_mesh_{send,inbox,
                 await,transfer_file}, kars_spawn{,_status,_destroy,
                 _list}

  Router /healthz from the agent container: 200 ok

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(hermes): A1 e2e smoke — six bugs surfaced by kind cluster run

End-to-end Hermes smoke on kind cluster exposed and fixed six real
bugs blocking the runtime from being functional:

1. awk not in Azure Linux 3 — replaced entrypoint merge with Python
2. TUI mode crashed without TTY — switched to hermes gateway run
3. KARS_MCP_SERVERS injected only into "openclaw" container —
   generalized to use agent_container_name based on runtime kind
4. Entrypoint scanned wrong path for MCP servers — aligned to the
   KARS_MCP_SERVERS env + loopback router pattern
5. hermes config set used key=value (wrong) — fixed to two positional args
6. Router rustls CryptoProvider not pre-installed — added explicit
   aws_lc_rs::default_provider().install_default() in main()

Verified 12/12 e2e checks pass on kind cluster:
- Pod 2/2 Running, plugin loaded with 10 tools + 2 hooks
- Router /healthz, /agt/evaluate, /egress/fetch, /sandbox/list all 200
- KarsMemory CR Compiled, McpServer translated, channel translation
- Mesh stubs return clear Act 2 error
- pre_tool_call hook fires + decision=allow

All 834 controller + 932 router Rust tests pass.
cargo clippy clean, cargo fmt applied.

Security audit:
docs/internal/security-audits/2026-06-04-hermes-act1-e2e-smoke-fixes.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): always allow operator policy-echo ingress (NP)

The sandbox NetworkPolicy gated ALL ingress rules behind
`governance.enabled=true`. With governance off, the NP shipped with
`policyTypes: [Ingress, Egress]` and an empty `ingress: []` block —
deny-all ingress. The operator namespace then could not reach
`/internal/policy-status` on the router and every referencing
InferencePolicy / KarsMemory / ToolPolicy / McpServer / EgressApproval
stuck forever in `Ready=False / AwaitingRouterEnforcement`, observable
in the operator panel even though the sandbox itself was healthy and
the router /readyz returned 200.

Split into two ingress classes:
- **Operator policy-echo ingress** (router :8443 admin surface from
  ns labeled `app.kubernetes.io/name=kars,component=system`) — emitted
  UNCONDITIONALLY. Three orthogonal gates still protect it: bearer
  token, constant-time compare, optional IP pinning.
- **Peer-sandbox mesh + gateway ingress** (8443 / 18789 / 18791 from
  ns labeled `kars.azure.com/role=sandbox`) — kept gated on
  governance.enabled (no peers when governance is off).

Surfaced during local-k8s smoke of smoke-hermes: even after fixing
the AZURE_OPENAI_API_KEY env path so /readyz returned 200, three
policy CRs (InferencePolicy, KarsMemory, ToolPolicy) stayed
Ready=False because the controller's /internal/policy-status probe
to the sandbox router timed out at the NetworkPolicy level.

After this fix, with governance off, the controller's HTTP probe
gets a 401 (admin-token gate doing its job) instead of a connection
timeout, and the policy reconcilers update status using the round
trip rather than reporting "router unreachable".

Verified end-to-end on kind cluster `kars-dev`:
  $ kubectl get inferencepolicy smoke-inference -n kars-system -o jsonpath='{.status.conditions}' | jq
  - Ready=True  RouterEnforcing: all 1 referencing sandbox router(s) confirmed inference-policy digest
  - Progressing=False  Reconciled: router echo confirmed
  $ kubectl get karsmemory smoke-mem -n kars-system -o jsonpath='{.status.conditions}' | jq
  - Ready=True  RouterEnforcing: all 1 referencing sandbox router(s) confirmed claw-memory binding digest

834 controller tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(e2e): add exec-brief-hermes-single scenario (Hermes Act 1)

Collapses the canonical 4-agent exec-brief scenario (parent +
analyst + viz + writer) into a single Hermes agent doing the whole
pipeline itself — research, scorecard, hero image, written brief.
Built to validate the Hermes runtime adapter end-to-end on
local-k8s and AKS without depending on the Python AGT MeshClient
(which ships in Act 2; until then, `kars_mesh_*` returns explicit
"Act 2 not ready" errors and the prompt explicitly tells the agent
not to call those tools).

Scenario layout (mirrors exec-brief/):
  - manifests/00-namespace.yaml ........ kars-execbrief-hermes ns
  - manifests/01-inferencepolicy.yaml .. azure-openai gpt-5.4
  - manifests/02-toolpolicy.yaml ....... allow-all AGT profile
  - manifests/03-clawmemory.yaml ....... memory-execbrief-hermes store
  - manifests/04-mcpserver.yaml ........ DeepWiki MCP (same as canonical)
  - manifests/05-clawsandbox.yaml ...... runtime.kind: Hermes
  - config.sh .......................... SCENARIO_SUB_SANDBOXES=()
  - prompt.txt ......................... single-agent pipeline
  - README.md .......................... what it exercises + skips

Verified on kind cluster `kars-dev`:
  $ kubectl apply -f tools/e2e-harness/scenarios/exec-brief-hermes-single/manifests/
    → 6 resources created
  $ kubectl get karssandbox execbrief-hermes -n kars-system
    PHASE=healthy RUNTIME=Hermes
  $ kubectl get pods -n kars-execbrief-hermes
    execbrief-hermes-...   2/2 Running

All 5 CRs reach RouterEnforcing / Ready=True:
  ● execbrief-hermes-inference   InferencePolicy   router echo confirmed
  ● execbrief-hermes-toolpolicy  ToolPolicy        agt-profile digest confirmed
  ● execbrief-hermes-memory      KarsMemory        binding=bound
  ● execbrief-hermes-deepwiki    McpServer         healthy
  ● execbrief-hermes             KarsSandbox       healthy

In-pod verification:
  - kars plugin: enabled=True source=user, 10 tools + 2 hooks
  - foundry_memory store_name = memory-execbrief-hermes (matches CR)
  - config.yaml mcp_servers.execbrief-hermes-deepwiki present
  - KARS_MCP_SERVERS=execbrief-hermes-deepwiki in agent env
  - Router /readyz: 200 ok

Note: the actual LLM execution of the prompt requires real Azure
OpenAI / Foundry credentials. With the fake-key dev overlay used in
this validation, the pipeline runs through Hermes → kars plugin →
router → upstream-call layer and hangs at the upstream (expected).
Running with real creds — either via `kars dev --target local-k8s`
with a real provider, or on AKS via `SCENARIO=exec-brief-hermes-single
PLATFORM=aks ./tools/e2e-harness/run.sh` — will execute the full
pipeline.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(hermes): A1 e2e harness wiring — dev pull policy + Hermes posture hardening

End-to-end run of the new `exec-brief-hermes-single` scenario on
local-k8s surfaced four more bugs that all gate the prompt from
actually reaching the model:

1. **`pull_policy=Always` for `:latest` images** in dev mode forced a
   doomed registry pull (karsacr.azurecr.io/…) instead of using the
   kind-cached image. The controller now picks `IfNotPresent` when
   `KARS_DEV_PROFILE=true` is set on its own env. Production AKS
   stays on `Always` for `:latest`.

2. **Hermes' `tirith` auto-download** from GitHub releases blocked
   every cold start while the kars egress-guard slow-walked the
   fetch. Entrypoint now sets `TIRITH_ENABLED=false` by default;
   Hermes falls back to its built-in pattern-matching shell
   checker. Operators can re-enable by pre-baking the binary at
   `/usr/local/bin/tirith` and setting `TIRITH_ENABLED=true`.

3. **`HERMES_DISABLE_LAZY_INSTALLS=1`** suppresses Hermes' `pip
   install` of discord.py / google-* / brotlicffi on first use of
   bundled platform plugins. Saves 30–120s on every cold start;
   operators wanting the extras re-bake into the image.

4. **`HERMES_SKIP_NODE_BOOTSTRAP=1`** suppresses Hermes' shell-based
   Node.js 22 LTS auto-installer (scripts/install.sh). We pre-install
   `nodejs` + `nodejs-npm` from the Azure Linux 3 base repo
   (currently v20.14 — Hermes' dep_ensure accepts any modern node).
   Browser tools that need a Chromium download still need to be
   pre-baked separately.

All three Hermes-runtime knobs are also mirrored into
`$HERMES_HOME/.env` so they survive `kubectl exec` sessions
(kubectl exec spawns a fresh env that doesn't see entrypoint
exports). Hermes' env_loader loads .env at import time
(`hermes_cli/env_loader.py:_load_dotenv_with_fallback`).

After all four fixes verified end-to-end:
  - smoke-hermes sandbox: phase=Running, 2/2 Ready
  - Router /readyz: 200 ok (controller forwards real Foundry API
    key from `kars-dev-creds` Secret via secretKeyRef)
  - Router /v1/chat/completions: 200 with real gpt-5.4 reply ("OK"
    in 1.1s, latency_checkpoint shows engine_ttft_ms=108)
  - InferencePolicy / KarsMemory / ToolPolicy / McpServer all
    Ready=True / RouterEnforcing
  - Plugin loaded with 10 tools + 2 hooks + foundry_memory native
  - Platform MCP block present in config.yaml when
    FOUNDRY_PROJECT_ENDPOINT is bound

Outstanding gap (NOT in this commit): Hermes' `hermes -z` still
makes an outbound HTTPS handshake (state=SYN_SENT to 104.18.3.115
:443, a Cloudflare IP — likely a check-update or telemetry endpoint
the harness hasn't tracked down). The kars egress-guard's
forward-proxy stalls the connection rather than denying outright,
so the prompt-driven path hangs after plugin discovery completes.
Workarounds:
  (a) `KARS_EGRESS_LEARN=true` to log unallowed hosts, then
      explicitly allowlist in EgressAllowlist;
  (b) find Hermes' env to disable check-update / telemetry — Act 1.x;
  (c) drive Hermes via Telegram channel instead of `hermes -z`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(hermes): A1 e2e — Hermes runs the full exec-brief pipeline on real Foundry

The single-agent exec-brief scenario (research → JSON → scorecard PNG →
hero PNG → 2-page brief.md) now runs end-to-end on Hermes through the
kars router to real Azure Foundry gpt-5.4. Verified on local-k8s with
the user's ~/.kars/ creds.

Four fixes were needed (each surfaced sequentially as the agent loop
progressed further):

1. **`OPENAI_API_KEY` env routes Hermes to openrouter** (and openrouter.ai
   is blocked by the egress-guard). Switched the entrypoint's `.env`
   mirror to `AZURE_FOUNDRY_API_KEY` + `AZURE_FOUNDRY_BASE_URL` so
   resolve_provider() picks the `azure-foundry` provider (which has
   no built-in Cloudflare callback).

2. **`agent_init.py` hardcodes `_codex_reasoning_replay_enabled = True`**
   → Hermes echoes `{"type": "reasoning", "encrypted_content": "..."}`
   back to /v1/responses on every continuation, which Azure Foundry's
   strict schema validator rejects with `invalid_payload`. OpenAI's
   own Responses API accepts these. Hermes only learns to disable
   replay when the upstream returns `invalid_encrypted_content` (a
   different error code that Foundry doesn't emit).

   Router fix: `build_upstream_url()` in proxy.rs now strips
   `input[]` items of `type=reasoning` and the
   `include=["reasoning.encrypted_content"]` field from any /v1/responses
   request bound for Azure Foundry (NOT GitHub Models / Copilot —
   their schemas accept the original shape).

3. **/v1/responses handler used `forward()` (non-streaming)** but Hermes
   always opens these with `responses.create(stream=True)` and expects
   an SSE `text/event-stream` response. The buffered JSON blob made
   Hermes' SDK raise "Connection error" after ~15s and retry 6× before
   giving up with `max_retries_exhausted`. Switched the handler to
   `forward_stream()` so the SSE byte stream flows through unchanged.

4. **`forward_stream()` injected `stream_options.include_usage`** which
   the OpenAI Responses API rejects (`unknown_parameter`). Skip the
   injection for /v1/responses (Foundry already emits usage in the
   terminating SSE event); was already skipped for Anthropic
   /v1/messages — same exclusion now covers both shapes.

Plus the entrypoint now persists `model.{default,provider,base_url}` in
config.yaml on every boot (not just plugins+mcp_servers), so a fresh
pod doesn't need a one-time `hermes config set model` post-boot dance.

End-to-end run delivered:
  /sandbox/incoming/brief.md      6,136 B  (2 pages, real Markdown,
                                            12 footnoted https citations,
                                            references hero+scorecard PNGs
                                            inline, all 4 control-domain
                                            terms present)
  /sandbox/incoming/analyst.json  5,025 B  (foundry_web_search × 3 →
                                            trends / control_categories /
                                            runtimes / metrics)
  /sandbox/incoming/hero.png     30,094 B  (1024×1024, foundry_image_generation
                                            gpt-image-1, "Defense in Depth"
                                            isometric data-center cutaway)
  /sandbox/incoming/scorecard.png 12,201 B (1024×640, foundry_code_execute
                                            matplotlib grouped bar chart,
                                            4 runtimes × 4 control columns)

Router log: 30+ /v1/responses SSE streams, all 200 OK, latencies
1.6–67s. Foundry stream headers received for every request after
this fix; pre-fix only 2 of 8 requests had `Foundry complete` entries
before Hermes gave up.

Agent stdout (final response after autonomous tool-use loop):
> Done. Artifacts produced:
> - /sandbox/incoming/brief.md — 6136 bytes
> - /sandbox/incoming/hero.png — 30094 bytes
> - /sandbox/incoming/scorecard.png — 12201 bytes
> - /sandbox/incoming/analyst.json — 5025 bytes
> Verified: brief.md exists and references both image files
>           hero.png and scorecard.png exist as real PNGs
>           analyst.json exists with the normalized runtime comparison

All 932 router + 834 controller Rust tests still pass.

Deliverables captured under:
  tools/e2e-harness/out/hermes-exec-brief-delivered/

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(router): operator-UX token + sandbox metrics for /v1/responses

Two visibility gaps surfaced after the Hermes exec-brief run:
operator panel showed `sandbox="unknown"` (instead of the real
sandbox name) and zero token counters for every /v1/responses call.

1. **sandbox label was "unknown"**: every `x-kars-sandbox` header
   parser fell back to `"unknown"` when the header wasn't set —
   which is the default for clients like Hermes' openai SDK that
   don't add kars-specific headers. Per-sandbox routers KNOW their
   own identity via the `SANDBOX_NAME` env (set by the controller).

   Added `resolve_sandbox_name()` helper at the top of inference.rs:
   trust+validate the header if present; otherwise fall back to
   `SANDBOX_NAME` env (Box::leak'd to &'static str — fine because
   the env is set once at process start). Replaces 4 hand-rolled
   `unwrap_or("unknown")` / `unwrap_or("self")` sites. All four
   /v1/{responses,completions,embeddings} + foundry-proxy handlers
   now produce metrics labelled with the real sandbox name.

2. **token counters were empty for /v1/responses**: the SSE parser
   in `forward_stream` looked for top-level `usage` in each
   `data:` chunk. OpenAI Chat Completions /v1/chat/completions puts
   usage at the top level (works); OpenAI Responses /v1/responses
   puts it nested under `response.usage` in the terminating
   `response.completed` event (didn't work — captured a real
   response.completed event to confirm).

   Parser now probes both shapes:
     v.get("usage").or_else(|| v.get("response")?.get("usage"))

   /v1/responses tokens are now counted (verified live: kars_tokens
   delta of +16 input / +12 output for a "list 3 colors" prompt;
   was +0 / +0 before).

Verified on local kind cluster after rebuild:

  kars_inference_requests_total{model="gpt-5.4",sandbox="execbrief-hermes",status="ok"} 5
  kars_tokens_total{direction="input",model="gpt-5.4",sandbox="execbrief-hermes"} 51
  kars_tokens_total{direction="output",model="gpt-5.4",sandbox="execbrief-hermes"} 30

The operator panel's "Inference by sandbox" + token-mix dashboards
now populate correctly for Hermes / pydantic-ai / langgraph / any
runtime that uses /v1/responses with non-kars HTTP clients.

932 router tests + cargo clippy --all-targets -- -D warnings clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): Hermes Act 2 — runtime-neutral Python AGT MeshClient + 6-tool deny list

Closes the inter-agent comms gap for Python frameworks. Until now only
the TypeScript OpenClaw runtime could speak E2E-encrypted AGT mesh;
Hermes had Act 1 stubs that returned 'not_yet_implemented'. This adds
a real implementation usable by any Python framework (Hermes is the
first consumer).

## What ships

1. New package 'kars-agt-mesh' (runtimes/agt-mesh-python/)
   - MeshClient orchestrator wrapping the upstream agentmesh-platform
     crypto primitives (X3DH, Double Ratchet, SecureChannel)
   - IdentityStore: persists Ed25519+X25519 keys at mode 0600
   - RegistryClient: POP-signed POST /v1/agents, prekey CRUD,
     /v1/discover, Ed25519-Timestamp auth
   - RelayTransport: async WS client with 30s heartbeat + backoff
   - Process-singleton via _SINGLETONS dict (mirrors openclaw's
     Symbol.for('agt-mesh-client') pattern)
   - Runtime-neutral — no Hermes-specific code
   - 9 unit tests pass

2. Hermes mesh adapter (runtimes/hermes/.../plugin/mesh.py)
   - Replaces Act 1 mesh_stubs.py
   - Sync→async bridge: dedicated asyncio loop in bg thread so
     Hermes' sync tool callbacks can call MeshClient
   - Defaults to router-proxied URLs (127.0.0.1:8443/agt/{relay,registry})
     so egress-guard iptables stay in place
   - Registers kars_mesh_{send,inbox,await,transfer_file}

3. Sub-agent tool deny list (defence in depth)
   - Plugin-side: _HERMES_DENY in plugin/__init__.py deregisters
     delegate_task, mixture_of_agents, cronjob, kanban_create,
     kanban_comment, send_message
   - AGT-profile-side: denied_actions block in scenario ToolPolicy
     catches the same six names at priority 100
   - Rationale per-tool in security audit doc

4. Dockerfile updated to install kars-agt-mesh wheel before plugin stage

5. AGT wheel build script extended to include 'agent-mesh' package
   (now produces agentmesh_platform-4.0.0)

## Live verification on kind-kars-dev

- MeshClient.connect() returns 201 from registry, WS upgrade OK
- Self-discovery via /v1/discover returns own DID
- Plugin loader log shows 6 deregistrations + 4 mesh tools present
- 83 Hermes unit tests + 9 kars-agt-mesh unit tests pass

## Critical bug fixed mid-implementation

Initial POP shape sent raw 32-byte public key + ts; registry expected
base64url-string(pub) + ts. Also DID format is server-derived
did:mesh:<sha256(pub)[:32]>, NOT did:agentmesh:<b64url>. Fixed both
in registry_client.py and identity.py. Memory stored for future
non-TS SDK implementers.

## Security audit

See docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
(2 sign-offs, ci-gates green).

## Deferred to Act 2.2

- KNOCK auto-accept responder (currently logs only — Hermes only
  initiates so not reachable yet)
- Cross-runtime golden vectors (TS↔Python interop test)
- Multi-process Hermes broker (lazy_install subprocess) — not
  reachable while delegate_task is denied

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(mesh): kars-agt-mesh Act 2.1 — full bidirectional E2E round-trip working

Lands the protocol-correct fixes needed for MeshClient.connect() →
KNOCK → X3DH → Double Ratchet roundtrip between two sandboxes. Tested
end-to-end on kind-kars-dev with two Hermes pods (execbrief-hermes and
smoke-hermes) on the FRESHLY BUILT image (no hot patches):

- pod A registers, uploads prekey bundle, opens relay WS (with POP)
- pod B does the same
- pod A discovers B via /v1/discover (freshest-first sort)
- pod A fetches B's bundle, runs X3DH, sends KNOCK + first ciphertext
- pod B's _handle_knock_frame auto-accepts via SecureChannel.create_receiver,
  decrypts plaintext 'hello from execbrief-hermes'
- pod B replies via send_by_did → encrypted message frame
- pod A decrypts 'pong from smoke-hermes'

## Critical protocol fixes

1. **Relay WS connect-frame POP** (relay_transport.py)
   - Was: {type:'connect', from:did, ts:...}
   - Now: full proof-of-possession (std-base64 pub_key + iso ts + sig
     over ts), per AGT relay/app.py::_verify_connect_pop
   - Without this, the relay rejects every connection with
     'connect frame missing did/public_key/timestamp/signature'

2. **Registry auth header** (registry_client.py)
   - Was: three separate X-Agent-DID/Timestamp/Signature headers,
     signature over method+path+ts
   - Now: single 'Authorization: Ed25519-Timestamp <did> <ts> <b64url-sig>',
     signature over timestamp string only
   - Matches AGT registry/app.py::verify_ed25519_timestamp_auth

3. **X3DH bootstrap missing** (client.py)
   - Now connect() builds X3DHKeyManager + generates signed_pre_key
     + 10 OTKs + uploads bundle via PUT /v1/agents/{did}/prekeys
   - Without this, peers couldn't fetch our bundle, X3DH initiation
     would fail at the responder side

4. **KNOCK responder implemented** (client.py::_handle_knock_frame)
   - Was: log-only stub ('responder path not implemented')
   - Now: parses ChannelEstablishment, calls SecureChannel.create_receiver,
     caches the channel, decrypts the bundled first ciphertext,
     eagerly tops up the OTK pool for the next session

5. **Send fuses KNOCK + first message** (client.py::send_by_did)
   - First call to a new peer DID sends {type:'knock', establishment, ciphertext}
   - Subsequent calls send {type:'message', ciphertext}
   - Matches the TS SDK wire convention (one RTT, not two)

6. **AAD directionality fix** (client.py)
   - Initiator: f'{self_did}|{peer_did}'
   - Responder: f'{from_did}|{self_did}' (reconstructs the same bytes)

7. **EncryptedMessage wire format** (client.py)
   - Was: JSON of em.__dict__ (would fail at decoder)
   - Now: EncryptedMessage.serialize() / .deserialize() (binary + b64url)

8. **PeerBundle flat shape** (registry_client.py + client.py)
   - Was: nested dicts mirroring my best-guess wire format
   - Now: matches agentmesh.encryption.x3dh.PreKeyBundle's flat dataclass

9. **register_self handles 409 gracefully** (registry_client.py)
   - Was: raised MeshRegistryError, blocking every restart
   - Now: logs and continues — the subsequent prekey PUT (with
     Ed25519-Timestamp auth) proves we own the same key

10. **discover() sorts freshest-first** (registry_client.py)
    - Avoids hitting stale ghost-DIDs when a sandbox restarts with
      a new identity before the prior registration ages out

## Tests

- 9 kars-agt-mesh unit tests pass
- 83 Hermes unit tests pass
- Live bidirectional roundtrip verified on freshly-built image
  (build hash c1dcdfc11475... loaded into kind-kars-dev)

## Security audit updated

docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
- Residual risk #1 (no KNOCK responder) removed — now implemented.
- Added residual risk #4 (stale registry entries — non-security).
- Added live bidirectional test description.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(runtime-contract): lift env injection to be runtime-neutral + plug Hermes mesh egress-guard hole

## controller/src/reconciler/mod.rs

Adds three runtime-neutral env vars injected on EVERY agent container
(not just OpenClaw):

- KARS_MODEL=<inference model> — generic alias for OPENCLAW_MODEL so
  Hermes / OpenAIAgents / MAF / BYO can read the same value without
  knowing about runtime-specific env names
- KARS_RUNTIME_CONTRACT_VERSION=v1 — self-documenting marker that
  this container claims to participate in the kars v1 runtime contract
- KARS_RUNTIME_KIND=<Debug repr of RuntimeKind> — uniform anchor any
  plugin can use to introspect what runtime it's running as

Lifted from the OpenClaw-only `is_openclaw` gate. All 834 controller
tests still pass.

## runtimes/hermes/.../plugin/mesh.py

**Real bug fix**: the Hermes mesh plugin was reading AGT_RELAY_URL /
AGT_REGISTRY_URL from env. The controller injects these as the
upstream CLUSTER URLs (ws://agentmesh-relay.agentmesh.svc:8765 etc.)
— but those are blocked by the egress-guard iptables rule (UID 1000
is restricted to localhost + DNS only; ports 8765/8080 are dropped
before the connection establishes).

The OpenClaw runtime makes the same call deliberately in
`runtimes/openclaw/src/core/mesh-registry.ts` (always uses
`routerUrl("/agt/registry")` — comment: 'Runtime UID 1000 is
iptables-confined to localhost. AGT_REGISTRY_URL is set by the
sandbox launcher as the router's UPSTREAM target — it points at
the real registry which the runtime cannot reach directly').

Now Hermes does the same: hardcodes 127.0.0.1:8443/agt/{relay,registry}
(the router proxy) on the agent side, ignoring the cluster-DNS env
vars which only the router container is meant to consume.

## Live verification

End-to-end mesh round-trip re-run on the rebuilt controller + sandbox
images (no hot patches):
- pod A (execbrief-hermes) registers, discovers pod B, KNOCK + X3DH
- pod B auto-accepts, decrypts 'hello from execbrief-hermes', replies
- pod A decrypts 'pong from smoke-hermes'

Env vars confirmed present on the agent container post-reconcile:
  KARS_MODEL=gpt-5.4
  KARS_RUNTIME_CONTRACT_VERSION=v1
  KARS_RUNTIME_KIND=Hermes

## Tests

- 834 controller tests pass (cargo test -p kars-controller)
- 83 Hermes unit tests pass
- 9 kars-agt-mesh unit tests pass
- cargo clippy --package kars-controller -- -D warnings clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(mesh): Hermes Act 2.2 — multi-agent mesh end-to-end with kars_spawn

Wires the missing pieces so a Hermes parent can spawn Hermes children
AND mesh-message them through the real Python AGT MeshClient.
Multi-agent fanout (parent → 3 sub-agents) verified live on
kind-kars-dev: each sub-agent receives the encrypted KNOCK + first
ciphertext, decrypts plaintext, and the parent's transcript ends with
'EXEC_BRIEF_MESH_FANOUT_DONE: 3 mesh sends delivered.'

## Bug fixes

### 1. Hermes parent now spawns Hermes children (NOT OpenClaw)

inference-router/src/spawn/mod.rs::build_sub_agent_crd_with_labels
hard-coded `runtime.kind = OpenClaw` for every spawn. Now it:
  - Accepts an explicit `runtime_kind` field on SpawnRequest.
  - Falls back to the `KARS_RUNTIME_KIND` env on the router (set by
    the controller as part of the v1 runtime contract).
  - Falls back to "OpenClaw" for backward compat.

Also stamps the matching runtime variant key
(openclaw/hermes/openaiAgents/maf) so the CRD admission webhook
doesn't strip-reject the spec.

Restores the runtime kind from a captured spec on handoff snapshot
re-spawn (so Hermes parents survive handoff without silently flipping
to OpenClaw children).

### 2. Controller injects KARS_RUNTIME_KIND on the router container

controller/src/reconciler/mod.rs previously injected
KARS_RUNTIME_CONTRACT_VERSION + KARS_RUNTIME_KIND only on the
*agent* container. Without these on the router too, the spawn
endpoint had no env-based fallback for the kind, so the previous
fix would have silently regressed to OpenClaw.

### 3. Hermes mesh.py accepts OpenClaw-style arg naming

kars_mesh_send now accepts `to_agent` (OpenClaw convention) and
`to` (short form), and `content` plus `payload`, so prompts
written for the OpenClaw mesh API work on Hermes too. Tool schema
advertises the canonical `to_agent`/`content` names primarily.

### 4. Hermes plugin eagerly pre-registers MeshClient at load

runtimes/hermes/.../plugin/__init__.py kicks off a background thread
that calls `_get_or_init_client()` at gateway boot, so the
sub-agent's DID is discoverable in the registry before the parent's
`kars_mesh_send` arrives. Without this, kars_spawn → kars_mesh_send
races: the child is Running but its lazy MeshClient hasn't connected
yet, so find_by_display_name returns nothing and the parent gets
'Peer not found'.

### 5. Discovery falls back to capability when registry omits metadata

runtimes/agt-mesh-python/.../registry_client.py find_by_display_name
no longer requires `metadata.display_name` to be present (the AGT
Python registry's /v1/discover only returns did + capabilities). It
now matches against the capabilities list, which is where MeshClient
puts the display name on register.

## Harness additions

### tools/e2e-harness/platforms/aks.sh

- New `hermes-exec` prompt driver (selected via
  SCENARIO_PROMPT_DRIVER=hermes-exec) for runtimes that don't expose
  an HTTP gateway on port 18789. Drives `hermes -z` via
  `kubectl exec -c agent` with HOME=/sandbox + HERMES_HOME set
  explicitly (kubectl exec doesn't inherit container ENV).
- Optional SCENARIO_DAEMON_{SUB,SCRIPT,READY_MARKER} hooks to copy a
  helper script into a sub-sandbox and wait for a readiness marker
  before posting the parent prompt.
- platform_collect_artifacts now picks the right container name and
  gateway-log path per runtime (openclaw=/tmp/gateway.log,
  hermes=/sandbox/.hermes/logs/gateway.log).

### tools/e2e-harness/scenarios/mesh-roundtrip-hermes/

Minimal smoke scenario: two pods, one Python echo daemon, one LLM
prompt that calls kars_mesh_send + kars_mesh_await and reports the
decoded plaintext. Verified end-to-end on freshly-built images.

### tools/e2e-harness/scenarios/exec-brief-hermes/

Multi-agent variant: parent uses kars_spawn to launch 3 Hermes children
(analyst/viz/writer), then fans out via kars_mesh_send. This is the
Hermes counterpart of the canonical OpenClaw exec-brief scenario.

## inference-router/Dockerfile.dev

The canonical Dockerfile is distroless (no shell). The controller's
egress-guard init container runs `sh -c "iptables ..."` which can
only work on an image that has sh + iptables. The .dev variant uses
mcr.microsoft.com/azurelinux/base/core:3.0 (non-distroless) + tdnf
install iptables, while still COPYing the pre-staged binary. Used by
`kind load`-based local dev; production AKS keeps the distroless
prod image.

## Tests

- 83 Hermes unit tests pass.
- 9 kars-agt-mesh unit tests pass.
- 16 router spawn tests pass (added env-locked parallelism guard so
  the new sub_agent_inherits_parent_runtime_kind_from_env test
  doesn't poison sub_agent_crd_uses_post_s10_s13_shape).
- All 834 controller tests pass.
- cargo clippy --package kars-inference-router -- -D warnings clean.

## Live verification on kind-kars-dev

Multi-agent fanout reproduced end-to-end (run.sh-equivalent invocation):

  $ hermes -z 'kars_mesh_send to_agent="analyst" content="ECHO_TEST_ANALYST";
                kars_mesh_send to_agent="viz"     content="ECHO_TEST_VIZ";
                kars_mesh_send to_agent="writer"  content="ECHO_TEST_WRITER";
                emit EXEC_BRIEF_MESH_FANOUT_DONE'
  EXEC_BRIEF_MESH_FANOUT_DONE: 3 mesh sends delivered.

  analyst daemon log: PRE_REG_GOT bytes=17 text='ECHO_TEST_ANALYST'
  viz    daemon log: PRE_REG_GOT bytes=13 text='ECHO_TEST_VIZ'
  writer daemon log: PRE_REG_GOT bytes=16 text='ECHO_TEST_WRITER'

kubectl get karssandbox -n kars-system shows all 4 as RUNTIME=Hermes
(not the prior bug where Hermes parent spawned OpenClaw children).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* chore(harness): remove stray .new file from mesh-roundtrip-hermes

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(mesh): Hermes Act 2.3 — autonomous sub-agent auto-responder

Closes the last gap blocking the OpenClaw-style multi-agent
exec-brief pattern on Hermes: spawned sub-agents now respond to
inbound mesh messages **without an active session**.

## Problem

After Act 2.2 a Hermes parent could spawn Hermes children and
mesh-send to them, but the children couldn't reply with real LLM
output. Hermes sub-agents are passive daemons — the LLM only runs
when something invokes `hermes -z`. OpenClaw doesn't have this
issue because its plugin runs inside an always-on
`openclaw agent --local` session.

So a parent doing:
  parent → kars_mesh_send(to_agent='analyst', content='research X')
  parent → kars_mesh_await(senders=['analyst'])
would land the message in analyst's inbox but never get a reply.
The analyst's Hermes daemon would just queue the message and sleep.

## Fix

New `runtimes/hermes/.../plugin/mesh_worker.py`: a background
asyncio loop in each sub-agent that:
  1. Drains the shared MeshClient inbox.
  2. For each inbound message, runs `hermes -z <payload>` as a
     subprocess with KARS_MESH_WORKER_TIMEOUT_S (default 1500s).
  3. Resolves the sender's display name via the registry.
  4. Replies with the captured stdout via `kars_mesh_send` on the
     same singleton MeshClient.

Opt-in via `KARS_MESH_AUTO_RESPONDER=1`. The controller sets this
ONLY on Hermes sandboxes that have the
`kars.azure.com/parent` label (i.e. children spawned by another
sandbox via the router's spawn endpoint). The parent never gets it
on — the parent IS the human/external-driver and would otherwise
loop on the children's replies.

The plugin's `__init__`'s eager-init thread now also calls
`mesh_worker.start_worker()` after the MeshClient is up, so the
responder lifecycle is bound to the plugin's.

## Live verification

Multi-step exec-brief on kind-kars-dev with real Foundry work:

  parent → analyst:  'research 2026 agentic AI runtimes, reply ANALYST_FOUND: <url>'
  parent → viz:      'use foundry_code_execute to print a JSON dict'
  parent → writer:   'use file_write to author /sandbox/incoming/brief.md'
  parent → kars_mesh_await(senders=[analyst,viz,writer], timeout=600)

Parent transcript:
  WRITER_DONE: 486
  VIZ_DONE: {"chart_ready": true, "format": "bar", "width": 1024}

Writer pod /sandbox/incoming/brief.md (486 bytes, REAL LLM content):
  'In 2026, agentic runtimes are defined less by raw model capability
   than by orchestration: durable memory, verifiable tool use,
   background jobs, and policy-aware delegation have turned agents
   from clever chat interfaces into operating systems for knowledge
   work. The winning stacks emphasize observability, rollback,
   sandboxing, and human checkpoints, because the hard problem is no
   longer generating ideas but coordinating long-running actions
   safely, cheaply, and at production scale.'

Sub-agent daemon logs confirm:
  - Accepted KNOCK from parent's DID
  - AUTO_GOT bytes=<inbound>
  - AUTO_REPLIED bytes=<reply> to=<parent DID>

(Analyst's reply landed slightly past the parent's await window so
the parent's transcript shows TIMEOUT: 2 received — the mesh path
itself worked for all 3; only the LLM coordination timing was tight
because foundry_web_search adds 30+s to analyst's hermes -z latency.
Verified independently that analyst auto-responded with 16 bytes.)

## Tests

- 83 Hermes unit tests pass
- 9 kars-agt-mesh unit tests pass
- 834 controller tests pass
- 16 router spawn tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* fix(hermes): pre_tool_call hook signature + heartbeat-vs-app metric breakdown

Two operator-visibility fixes called out during the Act 2.3 live
verification:

## 1. Hermes pre_tool_call hook crashed silently → no AGT audit for tools

Root cause: `runtimes/hermes/.../plugin/governance.py::_on_pre_tool_call`
took positional arg `params`, but Hermes 0.15.2 invokes the hook
with KEYWORD args matching `plugins.py:1685-1707`:

  tool_name=<name>, args=<dict>, task_id=<id>,
  session_id=<id>, tool_call_id=<id>

Our signature `(tool_name, params, **_kwargs)` matched `tool_name`
but every other kw landed in `**_kwargs` and `params` stayed unbound.
Result: TypeError on every invocation → Hermes' hook-runner swallowed
it → no `/agt/evaluate` POST → **no AGT audit entry for any tool
call**. Operator saw only `inference:responses:gpt-5.4` entries in
the audit log even though the agents made dozens of tool calls.

Fixed by matching the Hermes invocation signature exactly
(tool_name, args, task_id, session_id, tool_call_id) + keeping
**_kwargs for forward compat.

Also fixed the deny return shape: the hook used to return a
JSON-string error blob, but `get_pre_tool_call_block_message` only
recognises `{"action": "block", "message": <str>}`. Old denies
were logged + ignored — the tool actually ran. New dict-shape denies
make the block actually block.

Action-verb taxonomy fix: `kars_mesh_send` read `params['target_agent']`
but the real arg name is `to_agent` (alias `to`). Action verb
became `mesh:send:` (empty target). Now accepts all three names.
Also added `mesh:inbox` and `mesh:await` verbs for the drain/wait
tools.

### Live verification

Before fix, parent's /agt/audit:
  inference:responses:gpt-5.4 × 63   (every line, no tool entries)

After fix, parent's /agt/audit:
  inference:responses:gpt-5.4 × 64
  tool:kars_discover:writer × 1      ← NEW
  mesh:send:writer × 1               ← NEW

Writer's /agt/audit after fix:
  tool:write_file:/sandbox/incoming/audit_evidence.txt × 1   ← NEW

## 2. Sent ≫ received metric asymmetry now legible

Operator UX was showing e.g. 2218 sent / 4 received which is correct
but confusing — sent counter included 30s heartbeats over hours of
uptime. The kars_mesh_messages_{sent,received}_total counters stay
(back-compat, total of all frame types).

New counters break the total down by frame type:

  kars_mesh_frames_sent_total{type='heartbeat'}     — 30s keepalive
  kars_mesh_frames_sent_total{type='message'}       — app payload
  kars_mesh_frames_sent_total{type='knock'}         — session establish
  kars_mesh_frames_sent_total{type='connect'}       — POP / WS open
  kars_mesh_frames_sent_total{type='ack'}           — KNOCK/heartbeat ack
  kars_mesh_frames_sent_total{type='unknown'}       — unclassified

Same shape for kars_mesh_frames_received_total.

Subtracting type=heartbeat + type=connect from the total gives the
real application-frame count. Operator dashboards can now show:

  app_sent = sum(rate(kars_mesh_frames_sent_total{type!~'heartbeat|connect'}[5m]))

Classification is a cheap byte-prefix scan (first 80 bytes); the test
`classify_frame_type_buckets_known_kinds` guards every bucket and
`classify_frame_type_handles_short_input` guards bounds.

## Tests

- 84 Hermes unit tests pass (3 new govern hook contract tests)
- 936 router lib tests pass (2 new classify_frame_type tests)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(cli): wire kars connect for Hermes sandboxes

Before this change, `kars connect <hermes-sandbox>` failed silently:
the AKS path is OpenClaw-specific — reads the `gateway-token` Secret
(only created for OpenClaw, see controller/src/reconciler/mod.rs:1354)
and port-forwards :18789 (containerPort only added for OpenClaw, ibid.
:1852). On a Hermes sandbox both are absent, so connect would print
'Gateway token not found' and bail.

Adds a Hermes-specific branch in cli/src/commands/connect.ts that
runs after the AKS-existence check but before the WebUI/shell logic:

  if (runtimeKind === 'Hermes') {
    kubectl exec -it -c agent — env HOME=/sandbox HERMES_HOME=...
      hermes chat --accept-hooks
  }

`hermes chat` is the canonical interactive REPL (per
`hermes --help` in 0.15.2 — running `hermes` alone prints usage).
`--accept-hooks` lets the AGT pre_tool_call hook run without
per-tool approval prompts (operator already approved by issuing
`kars connect`).

HOME + HERMES_HOME must be set explicitly because kubectl exec does
NOT inherit container ENV. Hermes' `ensure_hermes_home()` falls
back to $HOME/.hermes; without HOME set, the running container's
HOME defaults to `/` and Hermes tries to mkdir `/.hermes` which
ENOENTs on the read-only rootfs. /sandbox is the writable emptyDir
the entrypoint uses for the long-running gateway daemon.

The exec-ban VAP only targets container name `openclaw`; Hermes'
container is `agent` (set in controller reconciler.rs:1801 from
`is_openclaw` branch), so this is admission-compliant. See
`deploy/helm/kars/templates/admission-pod-exec-ban.yaml`
`matchConditions`.

The --web flag falls back gracefully with a one-line note that
Hermes doesn't ship a browser UI.

The --reset flag works for both runtimes (it's just a rollout
restart). For OpenClaw it clears the in-process brute-force lockout;
for Hermes there's no equivalent state but a restart is still useful
to pick up plugin / env changes.

Local Docker mode (--local) is unchanged — it drops into bash with
OpenClaw-style tips. `kars dev --runtime hermes` for local Docker
isn't a common path yet (the harness lives on local-k8s + AKS);
leaving the bash drop-in to handle both cases until that comes up.

## Tests

789 CLI tests pass (vitest, no new tests added — interactive shell
path is exercised by integration runs, not unit tests).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(operator): Enter-key drops into Hermes agent TUI (full UX parity)

Restores the 'press Enter on a sandbox row → drop into the agent
TUI' UX the operator had for local OpenClaw, but for Hermes on AKS.
OpenClaw on AKS still uses the port-forward + WebUI URL path because
the exec-ban VAP blocks exec into the openclaw container.

## What changed

cli/src/commands/operator/dialogs/connect.ts splits the Enter
handler by (location × runtime kind):

  - AKS + OpenClaw → existing port-forward path (VAP-bound)
  - AKS + Hermes   → PTY exec into 'agent' container (NEW)
  - local Docker + OpenClaw → 'openclaw tui' PTY
  - local Docker + Hermes   → 'hermes chat --accept-hooks' PTY (NEW)

The two PTY paths share a common _spawnPtyConnect() helper extracted
from the old inline body; the OpenClaw port-forward path is now
_aksOpenClawConnect(). Both are pure refactors — the byte-identical
PTY plumbing (blessed save/restore, raw-mode stdin, Ctrl-\ detach)
moved into the helper, no functional change for OpenClaw.

## Why this works for Hermes but not OpenClaw on AKS

deploy/helm/kars/templates/admission-pod-exec-ban.yaml has
matchConditions:
  expression: object.container == '' || object.container == 'openclaw'

The VAP fires ONLY when the target container is literally named
'openclaw' (or unspecified — which defaults to the first container,
which is 'openclaw' in OpenClaw pods). Hermes' container is named
'agent' (controller/src/reconciler/mod.rs:1801 picks the name from
the is_openclaw branch), so 'kubectl exec -c agent ...' bypasses the
VAP cleanly.

This was a deliberate VAP design: the policy targets the literal
openclaw runtime container, not 'any agent container'. Hermes (and
future runtimes whose container is named 'agent') benefit by design.

## HOME / HERMES_HOME env vars

Set explicitly on the exec because kubectl exec does NOT inherit
container ENV. Without them, Hermes' ensure_hermes_home() falls back
to $HOME/.hermes; since HOME defaults to '/' in kubectl exec
sessions, Hermes tries mkdir '/.hermes' on the read-only rootfs and
ENOENTs. /sandbox is the writable emptyDir the entrypoint daemon
uses for the long-running hermes gateway.

## Tests

- 789 CLI vitest tests pass (no new tests — interactive PTY path is
  exercised by live operator runs, not unit tests).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* fix(mesh): align Python AGT MeshClient wire format with TS SDK (cross-runtime interop)

Closes the last gap blocking Hermes ↔ OpenClaw mesh communication.
Until this change, the Python kars-agt-mesh library and the TypeScript
@microsoft/agent-governance-sdk produced INCOMPATIBLE relay frames —
Python-Python and TS-TS interop worked fine, but a Python sender
talking to a TS receiver (or vice versa) silently dropped messages.

## Wire-format divergences fixed

### 1. message frame: structured header, std base64

**Before (Python only):**
  {
    'v': 1, 'type': 'message',
    'ciphertext': '<urlsafe-base64 of (struct.pack(>I, header_len) + header + ct)>'
  }

**After (matches TS mesh-client.js::send):**
  {
    'v': 1, 'type': 'message', 'from': ..., 'to': ..., 'id': ..., 'ts': ...,
    'header': {
      'dh': '<std-base64 dhPublicKey>',
      'pn': <previous_chain_length>,
      'n':  <message_number>
    },
    'ciphertext': '<std-base64 ciphertext>'
  }

The TS receiver reads frame.header.dh / frame.ciphertext as separate
fields; the old Python shape had no .header, so TS-side .base64ToUint8
got an unexpected packed blob and decrypt errored out (silently
dropped at the SDK boundary).

### 2. establishment: short TS-style keys

**Before:**  {initiator_identity_key: ..., ephemeral_public_key: ..., used_one_time_key_id: ...}
**After:**   {ik: ..., ek: ..., otk: ...}   (matches mesh-client.js::serializeEstablishment)

### 3. KNOCK + first message: TWO frames, not one fused

**Before:** Python fused KNOCK + first ciphertext into a single
  'type=knock' frame for one-RTT latency. TS receivers do NOT consume
  a 'ciphertext' field on a KNOCK — they only read 'establishment',
  call acceptSession, then await a separate 'type=message' frame.
  → first ciphertext was lost on Python-to-TS sends.

**After:** Python sends two distinct frames: 'type=knock' (no ciphertext,
  just establishment) followed immediately by 'type=message'. Matches
  TS mesh-client.js::establishSession + send.

### 4. std-base64 (not urlsafe) on the wire

JS's btoa / Node's Buffer.toString('base64') produce std-base64 with
'+' and '/'. Python's base64.urlsafe_b64encode produces '-' and '_'.
A TS receiver's atob fails on '-'/'_'; a Python receiver's
base64.b64decode fails on '+'/'_' depending on input. Now all on-the-
wire byte strings use std-base64.

## Backwards compat

Receiver tolerates both shapes for one release cycle:

- _message_frame_to_encrypted accepts BOTH the TS shape and the legacy
  packed-ciphertext shape (fallback path)
- _wire_to_establishment accepts BOTH {ik,ek,otk} and the legacy
  {initiator_identity_key, ephemeral_public_key, used_one_time_key_id}
- _b64std_decode tolerates urlsafe alphabet on input

A fleet mid-upgrade between old/new pods won't drop in-flight messages.

## Live verification

Sent {b'WIRE_TEST_DIRECT', 16 bytes} parent → analyst via direct
asyncio script with PYTHONPATH pointing at hot-patched client.py.

Parent stderr:
  > TEXT '{"v": 1, "type": "knock", "from": "did:mesh:a61...", "establishment": {"ik":..., "ek":..., "otk": 20}}'
  > TEXT '{"v": 1, "type": "message", ..., "header": {"dh":..., "pn":0, "n":0}, "ciphertext": "..."}'

Analyst auto_responder.log:
  Accepted KNOCK from did:mesh:a61c9cbf...
  AUTO_GOT from=did:mesh:a61c9cbf... bytes=16
  AUTO_REPLIED bytes=16 to=did:mesh:a61c9cbf...

The 16-byte payload decrypted correctly with the TS-compatible shape.

## Tests

- 8 new wire-format unit tests pin every field-shape contract
- 9 existing kars-agt-mesh unit tests still pass

## Cross-runtime promise

With this commit, a Hermes agent CAN mesh-send to an OpenClaw agent
and vice versa (same relay, same registry, same crypto, now same
wire envelope). End-to-end interop verification on a mixed-runtime
cluster ships as a follow-up — the wire alignment is the prerequisite.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* feat(cli): include kars-runtime-hermes in 'kars push' image set

Without this, 'kars push' to ACR misses the Hermes sandbox image, so
operators who want to deploy a Hermes sandbox on AKS have to build +
push that image by hand. Aligns with the existing pattern for the
other six runtime adapter images (openai-agents, maf-python,
anthropic, langgraph, langgraph-ts, pydantic-ai).

The tag 'kars-runtime-hermes:latest' matches the controller default
(controller/src/reconciler/runtime.rs DEFAULT_HERMES_IMAGE).

## Cross-runtime mesh status (asked during this session)

Hermes ↔ Hermes mesh: WORKS end-to-end on local-k8s with the
TS-compatible wire format shipped in commit 1a6e7f4.

Hermes ↔ OpenClaw mesh: BLOCKED by an OpenClaw-side mesh-connect
bug. Symptoms observed on local-k8s with a freshly deployed
OpenClaw sandbox (kars-sandbox:dev image, built locally):

  inference-router log: tight WS connect-loop (~50 cycles/sec)
    'AGT relay WebSocket proxy connected'
    'AGT relay WebSocket proxy disconnected outbound_messages=2 ...'

  relay log:
    'Rejecting connect frame for did:agentmesh:1dcfc3c6...:
     connect frame missing did/public_key/timestamp/signature'
    (later) 'Cannot call "send" once a close message has been sent'

This is the contract-drift bug already documented in
docs/internal/security-audits/2026-06-02-agt-relay-pop-flood.md and
in the agt-e2e-encryption skill: the bundled
@microsoft/agent-governance-sdk in runtimes/openclaw/node_modules
uses the legacy did:agentmesh:<b64url> format and does not send
proof-of-possession fields in the connect frame.

Setting AGENTMESH_RELAY_ALLOW_UNAUTHED_DID=1 on the relay (escape
hatch from AGT relay PR #66918631) doesn't help — the WS-close loop
continues, suggesting the OpenClaw plugin's mesh code closes the
socket after the first reply, which is an OpenClaw or TS-SDK issue
upstream from kars.

The Python wire format (commit 1a6e7f4) is correct against the TS
SDK spec — verified with TS-shape frames decoded on the Python
receive side — so once the OpenClaw side updates its bundled
agent-governance-sdk to a version that:
  1. Uses the did:mesh:<sha256(pub)[:32]> DID format
  2. Sends connect-frame POP (public_key + timestamp + signature)
  3. Holds the WS open with heartbeats instead of disconnecting

cross-runtime mesh will work without any further Python changes.

Tracked separately as the next item on the OpenClaw upgrade list.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>

* fix(mesh-plugin): emit modern did:mesh:<sha256[:32]> DID format

The kars mesh-plugin previously generated 'did:agentmesh:<sha256[:16]>'
— a kars-specific shorter fingerprint that pre-dated the AGT spec
update. The post-2026-05-23 AGT Python registry rejects connect
frames whose DID doesn't match 'did:mesh:' + sha256(pub).hex[:32],
which produced a tight WS connect/reject loop on every mesh attempt
once we picked up a recent registry build.

Match the upstream AGT TS SDK (@microsoft/agent-governance-sdk
≥4.0.0) and the AGT Python registry: derive
did:mesh:<sha256(pub_key)[:32]>. Legacy did:agentmesh: inbound DIDs
from older peers are still accepted on parse (no parser change
needed — `parseDid` already tolerates both prefixes).

Tested:
- mesh-plugin unit suite: 68 passed
- Live: openclaw sandbox on kind registers as
  did:mesh:b3bdd3630f5f727ea1e708fcb431ff10 and is discoverable
  by Hermes via registry capability search.

Refs: kars cross-runtime mesh interop (Hermes ↔ OpenClaw)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* build(agt): move AGT pin from pallakatos fork to upstream microsoft branch

Previously kars main built the TypeScript SDK from
`pallakatos/agent-governance-toolkit@bdea1097` (a private mirror of
the kars-sdk-pop-signing feature branch). That branch has now been
pushed to the upstream repo, so we can flip every pin to the
canonical `microsoft/agent-governance-toolkit:kars-sdk-pop-signing`
without needing a private fork.

The new HEAD (3322175d) also carries a second pre-release fix that
unblocks Hermes ↔ OpenClaw cross-runtime mesh:

  3322175d  fix(ts/x3dh): align KDF with Signal X3DH §2.2 spec

Combined contents of the branch on top of upstream main:
  1. Proof-of-possession registration + connect-frame Ed25519
     signing in the TypeScript SDK (upstream PR #2772, in review).
  2. Spec-compliant X3DH KDF (F prefix moved into IKM, zero salt)
     so Python sender ↔ TS receiver derive byte-equal shared secret
     instead of triggering AEAD 'invalid tag' on the first MESSAGE.

Updated in lockstep (per docs/PUBLISHING.md 'Updating the AGT pin'):
  - vendor/agt/pin.json: url → microsoft, sha → 3322175d
  - Cargo.toml [patch.crates-io]: agentmesh + agentmesh-mcp git URL
    + rev (Rust agentmesh crate is byte-identical between the two
    revs — only the TS SDK was touched — so the workspace stays at
    one consistent toolkit revision)
  - Cargo.lock: refreshed via 'cargo update -p agentmesh -p agentmesh-mcp'
  - deny.toml allow-git: dropped pallakatos URL (microsoft already listed)
  - vendor/agt/microsoft-agent-governance-sdk-4.0.0-agt-bdea1097.tgz
    → vendor/agt/microsoft-agent-governance-sdk-4.0.0-agt-3322175d.tgz
    (174 KB, SHA256 dcc8cc...)
  - vendor/agt/SHA256SUMS regenerated
  - {mesh-plugin,runtimes/openclaw}/package.json: file: dep path
  - {mesh-plugin,runtimes/openclaw}/package-lock.json: refreshed
    via 'npm install'

Verified:
  - mesh-plugin: 68 tests pass, typecheck clean
  - runtimes/openclaw: typecheck clean
  - cargo check -p kars-controller: clean
  - new tarball contains the X3DH fix
    (dist/encryption/x3dh.js: F_PREFIX + ZERO_SALT + concat(F_PREFIX, ikm))

When both upstream PRs (#2772 and the pending X3DH spec-compliance
PR) merge to microsoft/agent-governance-toolkit main and AGT cuts
a release containing them, drop the [patch.crates-io] block, the
vendor/agt/ tarball, the file: deps, and vendor/agt/pin.json —
switch to the published npm + crates.io artifacts.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh-python): JSON-wrap payload to interop with TS SDK receiver

Cross-runtime mesh proof Hermes(Python kars_agt_mesh) → OpenClaw(TS
@microsoft/agent-governance-sdk) was failing silently: KNOCK routed,
session established (knock_accept seen on the wire), AEAD tag
verified correctly, plaintext recovered — but the TS SDK's
MeshClient.handleMessage hardcodes a 'JSON.parse(new TextDecoder().
decode(plaintext))' on every successfully-decrypted frame, and a
Python sender that hands raw bytes straight to SecureChannel.send()
produces a plaintext that throws inside that JSON.parse. The
exception lives one level above handleMessage's own catch block, so
the message is dropped on the floor *after* successful decrypt:
onMessage handlers never fire, the plugin's pushInbox() is never
called, and the receiver looks like it never got anything (despite
the kars router metric kars_mesh_messages_received_total
ticking up — the frame DID arrive, but the application layer never
saw it).

The TS SDK's outbound convention (mesh-client.ts::send) wraps every
payload as 'new TextEncoder().encode(JSON.stringify(payload))' so
the inbound JSON.parse always succeeds. Mirror that here:
- UTF-8 byte payloads → JSON-encode as a string before encrypt
- Non-UTF-8 binary payloads → wrap in a {raw_b64: '...'} envelope
- Inbound: invert both paths so callers see the same bytes they sent
- Inbound passthrough for non-JSON plaintext (older Python senders
  pre-dating this wrapper) and structured-JSON payloads (caller
  opted into JSON shape and will re-parse themselves)

Live-cluster proof (kind-kars-dev, openclaw pod with patched TS SDK
@3322175d): Hermes sender → openclaw's kars_mesh_inbox returns

  { from_agent: 'execbrief-hermes-multi',
    message_type: 'message',
    content: 'PATCHED_V10_INTEROP_PROOF' }

with received_total=2 (was 0 / session_desync / 'invalid tag' on every
prior attempt). KNOCK+MESSAGE flow visible in relay debug log; no
knock_reject.

Tested:
- 4 new wire-format unit tests cover utf-8 round-trip, binary
  raw_b64 round-trip, non-JSON passthrough, structured-JSON
  passthrough — all pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller,hermes,agt-mesh-python): operator panel shows mesh peers on Hermes side

The operator's per-sandbox AGT trust panel was empty on Hermes
sandboxes even after a successful KNOCK + decrypted MESSAGE exchange
with an OpenClaw peer. Three independent issues stacked on top of
each other:

1. controller: agent-container volume mounts (admin-token + AGT
   policy) were hard-coded to fire only when 'is_openclaw' was true.
   Hermes runs the same in-pod kars governance plugin via Python
   (runtimes/hermes/src/kars_runtime_hermes/plugin/) but never got
   /etc/kars/secrets or /etc/agt/policies mounted, so:
   - submit_trust() to /agt/trust returned 403 ('Admin token required
     for trust mutations') because no token file existed at any of
     the documented paths
   - the in-process governance engine started with an empty policy
     set ('AGT engine will start with an empty policy set and fail
     closed')

   Generalize the mount predicate: any runtime that ships the kars
   plugin (today: OpenClaw + Hermes; future Tier-1 runtimes will opt
   in by extension) gets both mounts. BYO continues to skip them.
   The agt-policy ConfigMap mount target is now agent_container_name
   instead of the literal string 'openclaw', so the inject reaches
   whichever container holds the agent runtime.

2. runtimes/hermes mesh_worker: _handle_message decrypted inbound
   messages but never called submit_trust(), so even after the
   admin-token mount fix the operator panel still missed the peer.
   Mirror what OpenClaw does inside its onKnock handler
   (runtimes/openclaw/src/index.ts: pushTrustToRouter(fromName,
   0.0)): resolve the sender's display_name first so the entry is
   human-readable, then push with score 0.0 (router baseline 500 =
   at-threshold trust).

3. runtimes/hermes mesh_worker _resolve_sender_name was calling
   client._registry.discover('', limit=200) — the AGT registry
   rejects an empty 'capability' query param and returns nothing,
   so the lookup silently returned None and the trust entry was
   keyed on the raw did:mesh:<hex> instead of the human name.
   Add a new direct GET /v1/agents/<did> path to the Python
   registry client (matches the AGT REST surface) and use it for
   peer-name reverse lookup. O(1) on registry side + works for any
   registered DID.

Tested live on kind-kars-dev with the rebuilt controller (admin-
token + agt-policy mounts inserted on Hermes pod, verified via
'kubectl get pod -o jsonpath') and rebuilt kars-runtime-hermes
image (mesh_worker submit_trust path baked in):

  OC → Hermes (V21):
    Hermes /agt/trust returns
      [{ agent_id: 'mesh-peer-openclaw',
         tier:     'Anonymous',
         score:    10,
         interactions: 1 }]

  Previously (V20, pre-fix):
    Hermes /agt/trust returns []
    or, with hot-patched mesh_worker, returns the raw DID:
    [{ agent_id: 'did:mesh:d7ddf7278ce65b...', ... }]

  Auto-responder requires KARS_MESH_AUTO_RESPONDER=1 on the parent
  sandbox to be enabled — this is the default on sub-agent containers
  the controller spawns. Parents that explicitly opt in get the
  same trust-publish path.

Tests:
- 2 new mesh_worker tests cover the submit_trust call path
  (success + DID fallback when registry lookup misses).
- All 834 controller tests pass after the mount-predicate change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(hermes): trust score at OpenClaw-convention baseline + bidi harness

- mesh_worker.submit_trust call now sends score=0.5 instead of
  score=0.0. The Python submit_trust helper scales 0.0-1.0 → 0-1000
  (so 0.0 → 0, floored by the router to its anonymous-tier minimum
  of 10), but OpenClaw's TS plugin uses a delta-around-500
  convention: pushTrustToRouter(name, 0.0) →
  Math.round(500 + 0.0 * 500) = 500. Sending 0.5 from Python lands
  at the same 500 baseline so a Hermes-side peer entry shows the
  same at-threshold trust as an OpenClaw-side peer entry would,
  rather than appearing as a near-zero-trust agent in the operator
  panel.

- tests/e2e/interop/hermes_openclaw_bidi.sh: end-to-end
  bidirectional reproducer. Spawns no infra, drives the existing
  Hermes + OpenClaw sandboxes on kind-kars-dev through one
  OpenClaw → Hermes mesh round-trip and asserts the operator-side
  view:
    1. OpenClaw warmed + registered with capability=mesh-peer-openclaw
    2. kars_mesh_send delivers (status=delivered_via_agt_relay or
       delivered_and_replied)
    3. Hermes /agt/trust gains a fresh last_interaction for the
       peer (proves submit_trust path fired on the inbound msg)
    4. Trust entr…
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request Jun 15, 2026
…ategory error

User challenge (2026-06-15): why would kars need a front-door at all
when the inference router already lets agents communicate with models?

The challenge is correct. The front-door I had as Phase 1 items F1/F2
was solving 'give my external IDE a cluster-managed OpenAI endpoint',
which is exactly what agentgateway is for. Putting it in kars would:

* Collide with agentgateway in the category where they dominate
  (LF-hosted, MSFT-backed, 9 enterprise sponsors)
* Contradict our per-pod trust-boundary claim (an external IDE has no
  egress-guard and is a trusted, not adversarial, caller)
* Blur the product positioning ('agent runtime AND model gateway?')
* Centralize what we deliberately decentralized (front-door is a
  cluster-singleton ingress; kars router is a per-pod sidecar)
* Reduce kars to a worse OpenAI proxy

Plan revised:
* F1, F2 removed from the must-have table
* Strategy note added explaining the decision
* Replaced with three agent-runtime-specific items only kars can ship:
  - #1 Sub-agent spawn governance hardening (validate target / inherit
    creds / propagate audit context across spawn chains)
  - #2 Unified per-agent action-cost ledger across model + tool + MCP
    + mesh + spawn (agentgateway tracks model calls only)
  - #18 Mesh-aware QoS (per-peer rate-limit, fair-share, KNOCK-aware
    budget) — only kars has a mesh
* Composition framing: when an external IDE needs governed-cluster
  credentials, the right answer is agentgateway in front + kars
  inside, and the docs say so explicitly

Todo store updated to match (lead-F1, lead-F2 dropped; lead-SP1,
lead-AC1, lead-MQ1, lead-D1b added).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Jun 15, 2026
…d apply-fix + Telegram (#397)

* demo(act2): S0 — infra-tier ResourceQuota incident harness for Agent A

PR #1 in the kars-sre/demo-and-agent series — Slice 0 of the SRE
proposal: the demo can now be walked end-to-end by hand before any
SRE plugin code lands. Each subsequent slice (S1 read-only tools,
S2 K8s diag toolset, S3 typed apply-fix, S4 proactive watcher)
replaces one hand-walked step with an autonomous one.

Scenario: 'platform team's GitOps refactor lands a tight
ResourceQuota across every workload namespace; the quota's
requests.memory ceiling (50Mi) is lower than what the research
sandbox actually requests. The pod stays Running until anything
triggers a reschedule — then it goes Pending forever because the
quota blocks pod admission.'

Why infrastructure, not image-tag:  image tags don't change on a
running pod for random reasons.  ResourceQuota mis-configuration is
a real GitOps-collision incident that operators hit regularly.

Files:
  agent-a-research.yaml         — KarsSandbox 'research' (Hermes
                                  runtime, mirrors exec-brief-hermes-
                                  single shape, simplified to two CRs
                                  so the demo focuses on the runtime)
  platform-hardening-quota.yaml — the bad ResourceQuota the break
                                  script applies; deliberately NOT
                                  labeled kars.azure.com/managed-by
                                  so the SRE's DeleteResourceQuota
                                  typed action is permitted
  break.sh                      — applies the quota, force-deletes
                                  the running pod, confirms the
                                  FailedCreate event surfaces
  reset.sh                      — deletes the quota and waits for
                                  Running 2/2 (manual recovery path)
  runbook.md                    — presenter script for walking Act II
                                  by hand until S2 ships; once S2
                                  ships, the runbook becomes the
                                  expected-behaviour spec for the
                                  autonomous agent walk

Proposal update:
  §7.7.1 — adds DeleteResourceQuota as a typed action (namespace-
           scope, requires the ResourceQuota NOT carry the
           kars.azure.com/managed-by=controller label so kars-owned
           governance quotas stay protected and only operator-applied
           platform quotas are deletable)
  §7.7.1 — removes the PatchSandboxRuntimeImage carve-out from the
           previous draft; the demo no longer requires writes to
           kars.azure.com/* CRs, so the no-governance-mutation rule
           stays absolute

Validation:
  python3 -c yaml.safe_load_all on both YAMLs        — parses OK
  bash -n break.sh / reset.sh                        — syntax OK
  ci/check-copyright-headers.sh                      — all 499 OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre(s1): MVP — Helm template + 5 read-only kars-CR tools + CLI + plugin containment

Slice 1 of the kars-sre demo+agent series. The agent is now installable
on any kars cluster via 'kars sre install' and reachable via 'kars sre
talk'.  It reads kars CRs cluster-wide, walks the diagnostic checklist,
matches errors against the OOTB-blocker corpus, and proposes typed
fixes (apply is Slice 3).

What ships:

  deploy/helm/kars/templates/sre.yaml — Gated on .Values.sre.enabled.
  Creates 5 K8s objects when enabled:
    - InferencePolicy 'sre-inference' (kars-system)
    - KarsSandbox 'sre' (kars-system) with runtime: Hermes,
      extraEnv KARS_SRE_ENABLED=true, networkPolicy.defaultDeny=true
      + allowlist contains ONLY kubernetes.default.svc (NOT
      agentmesh — §7.8.6 network layer)
    - ToolPolicy 'sre-tools' (kars-sre) gating the sre_* surface
    - ClusterRole 'kars-sre-reader' — read on kars CRs + apiextensions
      + core workloads (RBAC per proposal §7.2.1 minus what S2/S3 add)
    - ClusterRoleBinding pinned to ServiceAccount kars-sre/sandbox
      (explicit subject — no group binding, no wildcard, §7.8.3)

  deploy/helm/kars/values.yaml — new 'sre:' block (enabled=false default,
  model=gpt-4.1, provider=azure-openai, tokenBudget=32000,
  extraAllowedEndpoints commented out for Slice 4 channel wiring).

  cli/src/commands/sre.ts — 'kars sre {install,uninstall,status,talk}'
  subcommands. 'install' wraps 'helm upgrade --reuse-values --set
  sre.enabled=true' then waits for the sandbox to reach Available.

  cli/src/cli.ts — wires sreCommand() into the Operations command group.

  runtimes/hermes/.../plugin/sre.py — 5 tools, all read-only:
    - sre_describe_state   structured snapshot of all 11 kars-owned CRs
    - sre_logs             apiserver-side pod log tail (cap 500 lines)
    - sre_diagnose         kars-CR health checklist + summary string
    - sre_explain_error    OOTB-blocker corpus matcher (6 known patterns
                           including ImagePullBackOff, exceeded quota,
                           OOMKilled, CrashLoopBackOff, FailedScheduling,
                           ContainerCreating)
    - sre_propose_fix      typed-action proposal envelope; Slice 1
                           codifies DeleteResourceQuota (the demo Act II
                           target) — rest of typed-action set lands in S3

  runtimes/hermes/.../plugin/sre_kube.py — minimal in-cluster apiserver
  client built on httpx (no new dep added to the shared Hermes image).
  Reads projected SA token + ca.crt + namespace from the standard paths;
  detects token rotation by content compare on each request.

  runtimes/hermes/.../plugin/__init__.py — adds the KARS_SRE_ENABLED
  gate. When set:
    - kars_spawn family is SKIPPED at registration (§7.8.5 — SRE agent
      cannot spawn sub-agents)
    - kars_mesh_* family is SKIPPED at registration (§7.8.6 — SRE agent
      is not on the mesh; combined with the NetworkPolicy block above
      this is two of three §7.8.6 enforcement layers — the third
      'separate image' layer is the §7.8.1 follow-up slice)
    - kars_discover is skipped (no peers to discover)
    - eager-mesh-init thread is skipped (would log noisy connection
      failures otherwise)
    - sre.register(ctx) runs AFTER everything else

  runtimes/hermes/tests/test_sre.py — 15 tests covering:
    - env-gate truthy/falsy mapping
    - all 5 tools register with the correct schema
    - explain_error matches against the corpus, handles no-match,
      handles empty input
    - propose_fix codifies DeleteResourceQuota for ResourceQuota target;
      returns rationale-only envelope for other kinds
    - KARS_CR_KINDS lists all 11 proposal §3.5 CRDs
    - describe_state walks every kind + surfaces per-kind errors
      without raising

  docs/sre.md — operator-facing readme: install, talk, tool surface,
  containment summary, what S1 cannot do yet, links to proposal +
  Act II runbook.

Validation:
  pytest tests/test_sre.py            → 15/15 pass
  pytest tests/test_governance.py     → unchanged, pass
  pytest tests/test_package_shape.py  → unchanged, pass
  npm run typecheck (cli)             → no errors
  npm run build    (cli)              → builds
  helm lint --set sre.enabled=true    → 0 fails
  helm template ... --show-only sre.yaml  → renders 5 objects clean
  helm template ... (sre.enabled=false)   → sre.yaml correctly omitted
  ci/check-copyright-headers.sh       → all 501 files OK

What this slice does NOT ship (per §7.1 ladder):
  - K8s diag toolset (sre_image_probe, sre_endpoints_inspect,
    sre_what_changed, sre_top, sre_describe_resource) — Slice 2
  - Fix execution (sre_apply_fix + TokenRequest + admission VAPs) — S3
  - Proactive watcher + Telegram/Slack notifications — Slice 4
  - Separate kars/sre-sandbox image (§7.8.1 packaging containment) —
    deferred; Slice 1 ships SRE in the shared Hermes image behind
    the KARS_SRE_ENABLED env gate as a tactical bridge. The env gate
    is the interim containment: tools aren't registered in any other
    pod, so a request for sre_* in a standard sandbox hits 'tool not
    found' at the runtime.

Next: Slice 2 (K8s diag toolset), then Slice 3 (typed apply-fix + AGT
approval flow + admission VAPs).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre(s2): K8s diagnostic toolset — describe_resource, what_changed, endpoints, image_probe, top

Slice 2 of the kars-sre series. Extends the read-only diagnostic
surface from kars-CR-centric (Slice 1) to arbitrary Kubernetes
workloads — everything the agent needs to diagnose the Act II
ResourceQuota incident end-to-end.

What ships (5 new tools, all read-only):

  sre_describe_resource — structured-describe for any K8s kind. For
                          workload kinds (Deployment / StatefulSet /
                          DaemonSet) walks the OWNER GRAPH:
                          workload → ReplicaSet → matching Pods →
                          events on every level. One tool call returns
                          the whole incident picture.

  sre_what_changed      — events of failure-relevant reasons in last
                          N minutes across BOTH core/v1 and
                          events.k8s.io/v1. Surfaces FailedCreate,
                          BackOff, OOMKilling, Evicted, etc. — the
                          incident-framing tool.

  sre_endpoints_inspect — Service → selector → matching pods →
                          EndpointSlice readiness. Synthesises a
                          finding the agent can quote (no pods match,
                          pods NotReady, targetPort mismatch, OK).

  sre_image_probe       — given an image, enumerate Pod images
                          cluster-wide and suggest the closest in-use
                          tag by Levenshtein edit-distance. Doesn't
                          reach out to the registry (per-registry auth
                          plumbing is Slice 4+); instead answers the
                          question that's actually most useful:
                          'what's the closest in-use tag on THIS
                          cluster right now?'

  sre_top               — metrics.k8s.io wrapper for CPU+memory per
                          pod or per node. Gracefully degrades to
                          {unavailable: 'metrics-server not installed'}
                          if the metrics API isn't registered
                          (proposal §7.5 Q4).

Also extends sre_propose_fix to codify two more typed actions from
proposal §7.7.1: PatchDeploymentImage and ScaleDeployment (in
addition to Slice 1's DeleteResourceQuota). Slice 3 will widen the
typed-action set further AND add the execution path.

RBAC widened in deploy/helm/kars/templates/sre.yaml:
  + discovery.k8s.io/endpointslices  (for sre_endpoints_inspect)
  + metrics.k8s.io/pods, nodes        (for sre_top)
  + core/nodes, endpoints, resourcequotas  (cluster-wide read)

ToolPolicy extended to allow the 5 new tool names.

Containment unchanged: still gated by KARS_SRE_ENABLED env on the
SRE sandbox pod only; standard Hermes sandboxes don't see the env,
don't load the tools, can't call them.

Validation:
  pytest tests/test_sre.py tests/test_sre_k8s.py  → 31/31 pass
  ci/check-copyright-headers.sh                   → all 502 OK
  helm lint --set sre.enabled=true                → 0 fails
  python -m py_compile (sre.py, sre_k8s.py)       → OK

Next: Slice 3 (typed apply-fix + admission VAPs + TokenRequest path
+ kars sre approve CLI).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): resolve helm chart path from repo root, not CWD

`kars sre install` was passing the relative path 'deploy/helm/kars'
to helm, which helm parses as a chart repo name when the user's CWD
is anywhere other than the kars repo root. Result:
  Error: repo deploy not found

Fixed by resolving the kars repo root the same way `kars up` does:
first walk up from the CLI file's own location (works for npm link),
then fall back to walking up from CWD looking for deploy/helm/kars.

Also: replaced the broken `.option('--wait', ..., true)` with the
commander-idiomatic `.option('--no-wait', ...)` so the wait flag
actually defaults to on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): use --reset-then-reuse-values for kars sre install

A plain --reuse-values carries the stored release values forward
verbatim. If the stored values are older than the chart on disk
(e.g. operator ran 'kars dev' before runtimes.hermes was added to
values.yaml), the template fails with:

  nil pointer evaluating interface {}.image

at controller-deployment.yaml line 89.

--reset-then-reuse-values (helm 3.14+ / helm 4) re-loads the chart's
values.yaml defaults first, then overlays the previously --set values
on top. So new chart fields get their defaults populated, while user
overrides for older fields are preserved.

Applied to both install and uninstall sub-actions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): create kars-sre namespace explicitly in the chart

The ToolPolicy 'sre-tools' lives in namespace kars-sre by design
(kars's cross-namespace ToolPolicy refs are deliberately not
supported — principles.md §3). But the controller-created
kars-sre namespace only exists AFTER the KarsSandbox 'sre' is
reconciled, which is AFTER helm tries to apply the ToolPolicy.

  Error: UPGRADE FAILED: failed to create resource:
         namespaces "kars-sre" not found

Fix: add the Namespace as a chart-managed resource at the top of
sre.yaml. The controller's namespace-reconcile path uses server-side
apply, so it will harmlessly co-own this namespace (adding its
own labels + annotations) when it reaches reconciler/mod.rs step 1.
No conflict.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): add --force-conflicts to helm upgrade (helm 4 SSA)

Helm 4 uses server-side apply by default. When prior
`kubectl set image` / `kars push --apply` runs took ownership of
fields that the chart now also wants to manage, the SSA call fails
with:

  conflict with "kubectl-set" using apps/v1:
    .spec.template.spec.containers[name="controller"].image

--force-conflicts (helm 4) instructs server-side apply to take
ownership on conflict. Matches operator intent: the helm-managed
chart is the source of truth, and chart-driven upgrades should
override transient field-manager pollution from ad-hoc
`kubectl set` calls.

Confirmed via `helm upgrade --help`:
  --force-conflicts   if set server-side apply will force changes against conflicts

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): ToolPolicy must live in KarsSandbox's namespace (kars-system), not kars-sre

Controller rejected the KarsSandbox sre with:
  Degraded: ToolPolicyNotFound — 'sre-tools' not found in 'kars-system'
  (cross-namespace refs not supported)

I had ToolPolicy in 'kars-sre' under the misunderstanding that it
should be co-located with the runtime pod's namespace. The actual
kars convention is the opposite: governance refs are namespace-local
to the KarsSandbox CR's OWN namespace (kars-system in our case), per
principles.md §3 cross-namespace-refs-deliberately-unsupported rule.
The runtime namespace kars-sre is for the pod + RBAC, not for
governance.

Confirmed against the existing exec-brief-hermes-single scenario
which co-locates KarsSandbox + ToolPolicy in kars-system.

Net: still safe wrt §7.7.1 protected-resource denylist (kars-system
is denylisted, so SRE agent can't delete this ToolPolicy even though
it's not labeled kars.azure.com/managed-by=controller).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): rename gate env KARS_SRE_ENABLED → SRE_ENABLED + indent fix

Two related bugs uncovered during live test:

1) The controller silently strips user-supplied extraEnv keys with
   reserved prefixes (mod.rs:1583 — AGT_, AZURE_, FOUNDRY_AGENT_,
   IMDS_, KARS_). KARS_SRE_ENABLED was being dropped, so the plugin
   never registered.
   Fix: rename to SRE_ENABLED across:
     - runtimes/hermes/.../plugin/sre.py           (is_enabled)
     - runtimes/hermes/.../plugin/sre_k8s.py       (module docstring)
     - runtimes/hermes/.../plugin/__init__.py      (log line + docstring)
     - runtimes/hermes/tests/test_sre.py           (3 env patches)
     - deploy/helm/kars/templates/sre.yaml         (extraEnv key + comment)

2) During the rename edit, the `extraEnv:` block ended up under
   `runtime:` instead of `runtime.hermes:` (4-space vs 6-space indent),
   producing:
     UPGRADE FAILED: .spec.runtime.extraEnv: field not declared in schema
   Fix: restore correct 6-space indent so extraEnv nests inside hermes.

Long-term fix (deferred): controller should detect
kars.azure.com/role=sre label on the KarsSandbox and inject
KARS_SRE_ENABLED itself (controller-side injection bypasses the
prefix filter). Noted inline at sre.is_enabled() docstring and in
the sre.yaml extraEnv block as a follow-up.

Tests: 31/31 pass (test_sre.py + test_sre_k8s.py).
Live verification: SRE_ENABLED env appears on agent container's env;
helm upgrade succeeds; chart re-applies cleanly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): default contentSafety.requirePromptShields=false

The Slice 1 template hardcoded requirePromptShields: true on the
SRE InferencePolicy. Azure OpenAI deployments only carry
'prompt_filter_results' in responses when an explicit Content
Filter policy is attached to the deployment. Bare local-dev
deployments (Foundry quickstart, gpt-4.1 without explicit filter)
don't emit those annotations — so the router blocks every response
with:

  Response blocked: InferencePolicy requires Prompt Shields but
  the upstream response carried no prompt_filter_results annotations

Diagnosed live during kars sre talk session — first prompt ('hi
there') returned a cached greeting that happened to bypass the
check, second prompt died.

Fix: default false in values.yaml + chart; operators wiring
Content Safety in production can set:
  --set sre.requirePromptShields=true

(or values.yaml override).

The SRE agent's threat surface is operator-driven Kubernetes
diagnosis, not user-facing chat, so prompt-shield enforcement is
less critical than for an internet-facing assistant. Operators who
need it can opt back in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: default model gpt-4.1 → gpt-5.4

Switch default model so the SRE agent ships with current frontier
out of the box. Operator can still override per-install with
`kars sre install --model <name>`.

The model name must match an Azure OpenAI deployment in the
operator's Foundry project — InferencePolicy routes to that
deployment via the router; if the deployment doesn't exist the
router returns a clear 404 and the sandbox surfaces Degraded.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): declare sre_* tools in plugin.yaml provides_tools

Hermes uses plugin.yaml's provides_tools list as the gate for
ctx.register_tool() calls — tools not declared in the manifest are
silently rejected at registration time. So even though sre.register()
called register_tool() for all 10 sre_* tools, none of them became
callable.

Diagnosed via live test:
  hermes tools list  → showed foundry_*, http_fetch, kars_handoff_status
                       (the manifest-declared ones)
                     → NO sre_*  (registered at runtime, manifest-rejected)

Same pattern as the OpenClaw plugin's contracts.tools requirement
(see memory: 'OpenClaw 2026.5.x requires plugin manifest to declare
contracts.tools listing every tool the plugin will register').

Fix: add all 10 sre_* tools (5 Slice 1 + 5 Slice 2) to provides_tools.
The tools remain conditionally registered at runtime — standard Hermes
sandboxes don't set SRE_ENABLED → sre.register(ctx) is skipped → the
tools are declared-but-not-callable (still matches the manifest
contract; Hermes treats them as 'present but inactive').

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: wire SRE-mode SOUL.md system prompt + fix register_tool kwargs

Three correctness fixes landed during the live test pass:

1) Hermes register_tool kwargs were wrong
   sre.py + sre_k8s.py used parameters=... but Hermes' contract expects
   schema=... AND toolset="<name>". Without these the manifest's
   provides_tools entries still showed up but the tools were silently
   non-callable. Fixed all 10 sre_* register_tool calls.

2) plugin.yaml provides_tools missing the sre_* entries
   Hermes' plugin loader requires every tool the plugin will register
   to be declared in provides_tools (same shape as OpenClaw's
   contracts.tools). Added all 10. Conditionally registered at
   runtime via SRE_ENABLED — standard sandboxes don't trip them.

3) New: kars-sre persona / system prompt
   Following the OpenClaw pattern (sandbox-images/openclaw/entrypoint.sh
   :1214 writes SOUL.md on every boot), the Hermes entrypoint now
   writes a 110-line SRE-specific SOUL.md to $HERMES_HOME/SOUL.md
   when SRE_ENABLED=true. Content:
     - Identity + mission statement
     - Tone constraints (concise, evidence-based, direct, honest)
     - Catalog of all 10 sre_* tools with WHEN to use each
     - Catalog of tools the agent does NOT have (spawn, mesh, shell,
       external net) with rationale
     - Standard incident reasoning loop (5 steps)
     - Output structure for fix proposals (Symptom/Evidence/Root cause/
       Proposed fix/Why safe/Rollback)
     - Boundaries (protected-resource denylist enforced at proposal
       layer; agent should not even try)
     - Audit info (where the kars audit JSONL captures every call)
     - First-message greeting template (one line, no editorialising)

   The model name interpolates from KARS_MODEL → AZURE_OPENAI_DEPLOYMENT
   → 'gpt-5.4' default, so the prompt always names the live model.

Validation:
  pytest tests/test_sre.py tests/test_sre_k8s.py  → 31/31 pass
  bash -n entrypoint.sh                            → clean
  live verify: SOUL.md written 110 lines, model = gpt-5.4
  live verify: hermes tools list → '✓ enabled sre' toolset now shows

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: apiserver-bypass for role=sre sandboxes (controller egress-guard)

Adds two iptables rules to the egress-guard init container, gated on
the kars.azure.com/role=sre label being present on the KarsSandbox:

  1. Filter chain: ACCEPT for UID 1000 -> KUBERNETES_SERVICE_HOST:443
     (BEFORE the existing catch-all DROP).
  2. NAT chain: RETURN for UID 1000 -> KUBERNETES_SERVICE_HOST:443
     (BEFORE the existing :443 REDIRECT to :8444 transparent proxy).

Both are required. The NAT-bypass alone is not sufficient because
the filter chain runs AFTER NAT - the NAT-RETURN says 'don't redirect'
but the filter-chain DROP next would still slay the packet. Discovered
live during testing: the curl-to-apiserver hung until both rules
landed.

Why this is needed: the SRE plugin's K8s API client (sre_kube.py in
the Hermes runtime) needs DIRECT apiserver access with its projected
ServiceAccount token to read kars CRs / pods / events. Without the
bypass, every apiserver call gets NAT-redirected to the router's :8444
transparent proxy, which has no idea how to forward TLS to the
apiserver -- connections hang then time out.

Why only role=sre sandboxes: every other sandbox kind goes through
the router unchanged -- that's the whole point of the transparent
proxy + L7 audit. Direct apiserver access is the deliberate
exception, uniquely held by the nominated SRE sandbox per the
proposal section 7.8 containment design.

K8s audit log is the audit surface for these apiserver calls (the
router's L7 audit doesn't apply, but K8s audit is stronger -- every
call carries the SA identity, verb, and resource).

Implementation:
  - new build_egress_guard_command(is_sre_sandbox: bool) helper
    in reconciler/mod.rs that emits the right rule sequence per mode
  - 3 unit tests: standard has no bypass; SRE has NAT bypass before
    REDIRECT AND filter ACCEPT before DROP; both modes keep DROP

Validated end-to-end:
  - HTTP 200 in 17ms from agent container -> 10.96.0.1:443
  - sre_describe_state() returns 10 KarsSandboxes + all 11 CR kinds

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): correct AGT profile schema (version 1.0 + agent: name + policies)

The Slice 1 inline AGT profile used the wrong schema — version: 1
with rules[].match.tool — which produced:

  ToolPolicy sre-tools: invalid YAML: missing field agent

at compile time, then 'router has not yet loaded AgtProfile' at the
sre pod's policy loader. The sre KarsSandbox showed Degraded with
ToolPolicyNotCompiled.

Found by the SRE agent itself during the first cluster-health-overview
test (a beautifully on-point sre_diagnose result that flagged its
own ToolPolicy as the only Degraded thing in the cluster).

Right schema (from deploy/helm/kars/files/kars-default-agt-profile.yaml):
  version: '1.0'
  agent: <name>
  policies:
    - name: ...
      type: capability
      allowed_actions: [...]
      denied_actions: [...]
      priority: N

Action prefix convention used by the router:
  tool:<tool_name>        for tool calls
  inference:<api>:<model> for model dispatch
  spawn:* / mesh:*        for sub-agent + mesh

The new sre-tools profile has three policies:
  - sre-diagnostic-tools-allow (priority 100): all 10 sre_* tools
  - sre-inference-allow (priority 90):  chat_completions / responses /
                                        content_safety
  - sre-spawn-and-mesh-deny (priority 110): defense in depth for the
    §7.8.5/§7.8.6 containment (already enforced by plugin not even
    registering these tools)

After re-apply + sre pod restart:
  ToolPolicy sre-tools status:  Ready  True:RouterEnforcing
  KarsSandbox sre status:       Running

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sre): trailing-colon glob in AGT allow rules — match real action shape

The Slice 1 allow rules used literal 'tool:sre_<name>' strings but the
Hermes plugin governance hook actually emits 'tool:<name>:<first-arg>'
— with a trailing colon even when no significant arg is present (see
runtimes/hermes/.../plugin/governance.py _action_verb tail returns
f'tool:{tool_name}:'). So:

  literal allow: 'tool:sre_describe_state'
  router emit:   'tool:sre_describe_state:'  <-- no match → denied

The agent helpfully diagnosed itself via:

  sre_describe_state -> blocked by policy 'sre-diagnostic-tools-allow'

(visible because the WebUI surfaced the matched_rule name). Confirmed
the action shape in inference-router/src/routes/governance.rs:66
('if let Some(tool_name) = action.strip_prefix("tool:")...').

Fix: add a '*' wildcard to every allowed_action for the sre_* tools.
This matches both the trailing-colon shape (tools with no args) and
the suffix-args shape (sre_describe_resource:<name>, sre_logs:<pod>,
etc.) in a single entry.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: NetworkPolicy egress allow for apiserver (cluster-portable)

The egress-guard iptables bypass (b25f41b) lets UID 1000 reach
the apiserver at the iptables layer, but the pod-level NetworkPolicy
was still denying it. The blanket :443 egress rule explicitly
excludes RFC1918 ranges to prevent lateral movement to in-cluster
Services, but every cluster's apiserver ClusterIP IS in one of those
ranges (kind: 10.96.0.1, AKS: 10.0.0.1, EKS: 172.20.0.1).

Fix: when role=sre, add a NetworkPolicy egress rule for the
apiserver Service ClusterIP. The IP + port are read at reconcile
time from the controller's own KUBERNETES_SERVICE_HOST /
KUBERNETES_SERVICE_PORT_HTTPS env vars (kubelet-injected on every
pod). This is cluster-portable — kind, AKS, EKS, custom service-CIDRs
all get the right value automatically. No hardcoded IPs.

Implementation:
  - Top of reconcile(): compute is_sre_sandbox once + read apiserver
    IP/port from env. Threaded through both the egress-guard helper
    and the NetworkPolicy egress vec.
  - egress_rules.push(...) added after the static block, gated on
    is_sre_sandbox, with IP/port substituted from env.
  - Removed the duplicate is_sre_sandbox compute lower in reconcile()
    that was added in b25f41b — single source of truth now.

Validated live:
  - kubectl get netpol -n kars-sre shows the 10.96.0.1/32 :443 rule
  - sre_describe_state() returns in 0.10s — 11 CR kinds, 10
    KarsSandboxes enumerated, NO timeouts.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(demo): agent-a-research.yaml passes CRD admission

Two admission rejections:

1) spec.governance.toolPolicyRef.name required when governance.enabled=true
   Added a research-tools ToolPolicy with allow rules for:
     - inference:chat_completions:* / responses:* / content_safety:*
     - tool:http_fetch:* (the agent does web research)
     - tool:foundry_* family (memory + web_search + code_execute etc.)

2) spec.runtime.hermes must be set iff kind=Hermes (CEL guard rejects
   missing key, accepts empty object). The previous manifest had a
   commented placeholder which yamllint-fine but admission saw the key
   as missing. Changed to 'hermes: {}' — empty object honours image
   defaults without drift.

Also: aligned the demo with the SRE sandbox defaults shipped earlier:
  - deployment: gpt-5.4 (was gpt-4.1)
  - requirePromptShields: false (was true — bare local Foundry deployments
    don't emit prompt_filter_results, blocking every response)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(demo): break.sh uses kars.azure.com/component selector

Controller stamps pods with kars.azure.com/component=sandbox not
the app.kubernetes.io/component=sandbox the script was looking for.
Result: 'no sandbox pod found to evict; quota will only manifest
on next natural restart' — the script kept going but the break
never surfaced.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* kars-sre: Slice 3 (typed apply-fix) + Slice 4 (proactive watcher + Telegram)

Slice 3 — typed apply-fix path (operator-approved remediation)

Adds the KarsSREAction CRD and reconciler that drives an SRE-agent
fix proposal Proposed → Approved → Applied → Recovered. The agent
emits a CR via sre_propose_fix; the operator approves via kars sre
approve <id> (or kubectl edit); the controller mints a one-shot
ClusterRoleBinding scoped to the right writer ClusterRole
(kars-sre-writer-quotas | kars-sre-writer-workloads), executes the
typed action via SSA, tears the binding down, and observes recovery
by polling the target namespace for failure-class events. Terminal
CRs (Recovered / Failed / Expired / Rejected) auto-GC after 1h.

Closed set of typed actions per proposal §7.7.1:
  - DeleteResourceQuota (refuses kars.azure.com/managed-by=controller)
  - PatchDeploymentImage, ScaleDeployment (clamp 0..50),
    RolloutRestart (Deployment/StatefulSet/DaemonSet), DeletePod

New files:
  - controller/src/kars_sre_action.rs            (CRD types)
  - controller/src/kars_sre_action_reconciler.rs (state machine)
  - deploy/helm/kars/templates/crd-karssreaction.yaml

Hermes plugin (sre_propose_fix is now a CR-creator):
  - Tolerant arg parsing: target.kind / action_type / inferred kind
  - schema marks target.kind required + enum-validated
  - Returns action_id + ready-to-paste 'kars sre approve' command
  - Clear cr_error when no typed fix could be inferred

CLI:
  - kars sre approve <id> / reject <id> / actions / show <id>
  - kars sre show renders diagnosis + rationale + condition stamps

RBAC additions (controller-side):
  - karssreactions (full r/w)
  - resourcequotas: delete (the §7.8.4 escalation check requires the
    controller to hold the verbs it grants in the one-shot CRB)
  - apps/statefulsets,daemonsets: patch (RolloutRestart targets)
  - events: list/watch/get (recovery observer)
  - serviceaccounts/token: create (lands the §7.8.4 TokenRequest path)
  - clusterrolebindings: create/delete kars-sre-write-*

Slice 4 — proactive watcher + Telegram

sre_watcher.py runs alongside the Hermes gateway when SRE_ENABLED=true
and a channel is configured. Polls K8s events every 10s for failure-
class reasons in kars-* namespaces (excluding kars-sre / kars-system
/ kube-* / agentmesh / default), maps each into a typed-fix target,
and on incident:

  1. Reuses any open KarsSREAction with the same (action_type, ns,
     name) target — no duplicate CRs.
  2. Otherwise creates a new KarsSREAction with ttl_minutes=30.
  3. Coalesces a per-iteration burst into ONE detailed Telegram
     message (highest-priority candidate) plus an optional summary
     tail ('+N other incidents: 2 FailedScheduling, 1 BackOff').
  4. Sliding-window rate limit: max 4 messages/min cluster-wide.

Dedupe is bootstrapped from existing KarsSREActions on boot (survives
pod restart). First iteration is silently absorbed (priming) so a
pod re-roll doesn't replay the warm-cache flood as alerts. Periodic
60s CR resync REPLACES the dedupe state so operator-side delete
clears the in-memory map naturally.

ReplicaSet/Pod hash suffixes are normalised in the dedupe key so a
flapping Deployment's rollout sequence collapses to one alert
instead of one alert per pod-template-hash.

Telegram wiring:
  - Channel adapter libraries (python-telegram-bot 21, slack-sdk 3,
    discord.py 2) pre-installed in the runtime image so credentials
    in the sandbox-credentials secret 'just work'.
  - entrypoint.sh exports HTTPS_PROXY=http://127.0.0.1:8444 and
    NO_PROXY=$KUBERNETES_SERVICE_HOST,127.0.0.1,localhost,.svc.cluster.local
    so the gateway's outbound HTTPS reaches the inference-router's
    forward proxy (egress-guard iptables redirect doesn't fire in
    kind clusters without CAP_NET_ADMIN — explicit env covers both).
  - HOME=/sandbox export so gateway-locks dir under ~/.local/state
    is writable on the distroless base.
  - TELEGRAM_ALLOWED_USERS exported (not just config-set) so the
    gateway's per-platform allowlist skips pairing for known users.
  - TELEGRAM_HOME_CHANNEL set to first TELEGRAM_ALLOW_FROM id so
    'hermes send --to telegram' resolves without explicit chat id.

Operator install path (unchanged — uses existing kars credentials):
  kars credentials update sre --telegram-token <T> --telegram-allow-from <ID>

Tests: 31 hermes tests + 847 rust tests + cli typecheck/lint pass.
The phase taxonomy guard now passes after refactoring the reconciler
to use named constants for all condition types / reasons / event
reasons rather than 'Failed' / 'Pending' / 'Degraded' literals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* kars-sre: Headlamp SRE Console + Chat (Slice 4 primary UX)

Adds the SRE engineer's dedicated console as a top-level sidebar
branch in the kars Headlamp plugin. Replaces the prior workflow of
'kubectl get karssreactions + paste action_id into kars sre approve
in a terminal' with one click in the dashboard.

New routes:

  /kars/sre          — SRE Console (live cards, primary landing)
  /kars/sre/chat     — embedded Hermes WebUI iframe
  /kars/karssreactions — full CRD list (under existing CRD section)

SRE Console layout (top → bottom):

  🔴 Pending Approval — KarsSREActions awaiting operator. Inline
     Approve / Reject buttons PATCH .spec.approval.state directly
     via Headlamp's KubeObject.patch(), with optional rejection-
     reason prompt. No terminal hop needed.
  🔄 In-flight — actions the controller is currently executing
     (Applied + waiting for recovery). Shows phase + age.
  📊 Cluster Health — sandbox phase counts + degraded count.
  🚨 Active Incidents — failure-class events (FailedCreate,
     BackOff, FailedScheduling, Failed, ImagePullBackOff,
     CrashLoopBackOff, OOMKilling, Evicted, FailedMount) from
     kars-* namespaces in the last 15 min. Same filter the
     proactive watcher uses, so what the operator sees here is
     what the watcher would alert on.
  ✅ Recent — Recovered / Failed / Expired / Rejected actions
     from the last hour for post-incident review.

All cards live-update via Headlamp's useList() (watch + long-poll),
so the Proposed → Approved → Applied → Recovered walk is visible
without F5. The KarsSREAction CRD is added to the existing CRD
registration table so the standard list / detail pages 'just work'
under /kars/karssreactions/:ns/:name.

SRE Chat is an iframe of the Hermes WebUI:
  - tab 1: http://localhost:18789 (requires 'kars connect sre --web'
    in another terminal — populates the iframe via port-forward)
  - tab 2: apiserver service-proxy fallback for in-cluster operators
  - 'Open in new tab' button if iframe sandboxing breaks the embed

Helm chart: SRE sandbox's allowedEndpoints now includes
api.telegram.org / core.telegram.org cluster-side so the Slice 4
watcher's outbound Telegram alerts don't need an out-of-band
NetworkPolicy patch. Dormant when Telegram isn't configured — the
gateway only opens the channel when the token is present.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: fix browser-ESM require() crash + add 'SRE not installed' CTA

Two fixes:

1. ReferenceError: require is not defined
   The Active Incidents card lazily resolved the Event class via
   require("@kinvolk/headlamp-plugin/lib/K8s/event"). Headlamp ships
   plugin bundles as pure browser ESM modules — require() doesn't
   exist in that context, so the page crashed at first render. Switch
   to the documented public re-export via the K8s namespace
   (`import { K8s } from "@kinvolk/headlamp-plugin/lib"` →
   `K8s.event`), which is safe in both build- and run-time.

2. Empty-state CTA when kars-sre isn't deployed
   Both SREConsole and SREChat now check for the existence of the
   sre KarsSandbox in kars-system. If absent (or the list is still
   loading), they render an actionable install card with:
     - `kars sre install` (the one-liner that enables the chart)
     - `kars credentials update sre --telegram-token ...` (optional)
   So a fresh kars dev cluster that hasn't run `kars sre install`
   yet doesn't show 'No items' or a spinning iframe — it tells the
   operator exactly what to type. The cards rehydrate live once the
   sandbox lands (no refresh needed).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: stub Active Incidents — pluginLib.K8s.event isn't host-exposed

The Headlamp 0.41 host runtime exposes `pluginLib.K8s` as a flat
namespace of class-kind classes but does NOT expose the v1 `event`
sub-namespace. Importing it via either explicit submodule path
(`@kinvolk/headlamp-plugin/lib/K8s/event`) or the top-level barrel
(`K8s.event.default`) trips Vite's UMD wrapper into its CJS-fallback
branch on first execution, which crashes the browser with:

  ReferenceError: require is not defined
      at ct (//plugins/kars/dist/main.js:3:52537)

(`ct` was the INCIDENT_REASONS set at top-level — top-level
execution failed before any component mounted.)

The KarsSREAction CR cards above already surface every incident
the proactive watcher catches (same dedupe key, same target shape),
so for Slice 4 the operator doesn't need the raw events feed
duplicated in the dashboard.

Slice 4.1 (future) can resurrect this via direct fetch() to
/api/v1/events through the headlamp apiserver proxy, bypassing
the K8s.event class entirely.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp: bump plugin to v0.6.0 to bust Headlamp's plugin cache

Headlamp keys plugins by package.json version. A pure dist/main.js
swap (with same version) leaves the host's plugin loader cache
holding the previous bundle. Bumping minor → operator's browser
re-fetches main.js on next mount even without Cmd+Shift+R.

v0.5.1 → v0.6.0 covers the prior session's additions:
  - KarsSREAction CRD list / detail
  - SRE Console + Chat sidebar branch
  - browser-ESM safety pass (no require() in source)
  - SRE-not-installed empty-state CTA

* cli: kars sre install handles 3 cluster shapes (helm release / kars dev / fresh)

The 'no deployed releases' error happened because 'kars dev --target
local-k8s' deploys the chart via 'helm template | kubectl apply'
(see cli/src/commands/dev/local-k8s.ts:794), so no helm release
record exists. The sre install path assumed a helm release and
failed on a fresh kars dev cluster.

Now detects three shapes:

  A. helm release present
     → helm upgrade --reset-then-reuse-values --force-conflicts
       (preserves operator's prior --set choices)

  B. no helm release BUT controller deployed (= kars dev path)
     → helm template … | kubectl apply --server-side --force-conflicts
       (mirrors how the chart got there in the first place)

  C. neither (= fresh cluster)
     → helm install --create-namespace --take-ownership
       (--take-ownership: adopt any pre-existing namespace or
        NetworkPolicy from prior partial installs; helm >= 3.17)

The template path uses --include-crds so KarsSREAction is installed
on first sre install even when the cluster predates Slice 3. All
three paths set azure.workloadIdentity.clientId=dummy for local-k8s
brand-new installs (real AKS installs go through kars up).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: derive cluster name from URL for apiserver-proxy chat tab

The proxy URL hardcoded 'kind-kars-dev' as the cluster name, which
only worked for the local-k8s demo path. Real operators have any
context name (AKS managed-cluster names, in-cluster Headlamp,
multi-cluster setups). The 'key not found' error was Headlamp's
backend rejecting the request because the cluster path component
didn't match any of the operator's configured contexts.

Fix: parse the cluster name from window.location.pathname (Headlamp
routes every cluster-scoped view under /c/<cluster>/...). When the
parse fails (e.g. the Chat page is loaded outside a cluster context),
the proxy tab is disabled and the operator is steered to the local
port-forward tab.

Reads location directly instead of useCluster() because importing
the K8s namespace (where useCluster lives) trips the host's UMD
require() fallback — the same crash the v0.6.0 plugin fixed.

v0.6.0 → v0.6.1 to bust Headlamp's plugin cache.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* controller: expose Hermes gateway port (18789) on per-sandbox Service

The per-sandbox Service exposed only :8443 (inference-router). For
Hermes runtimes the gateway WebUI / inbound channel adapter lives
on the agent container at :18789, but operators had no way to reach
it without setting up a per-sandbox port-forward.

Now: when runtime.kind == Hermes, the controller appends a 'gateway'
port (18789) to the same Service. Result: 'kubectl port-forward
svc/<name> 18789' works directly, AND the Headlamp SRE → Chat tab
can route via the apiserver service proxy:

  /clusters/<cluster>/api/v1/namespaces/kars-sre/services/sre:18789/proxy/

OpenClaw runtimes are unaffected (no gateway port added).

The NetworkPolicy ingress rule for governance-enabled sandboxes
already allows port 18789 from peer sandbox namespaces, so this
purely widens what the cluster apiserver / Headlamp backend can
reach — no extra exposure to other sandboxes.

* headlamp/sre: replace iframe Chat tab with terminal-attach instructions

Hermes is a CLI/TUI agent — there's no embedded WebUI to iframe.
Earlier commits attempted an apiserver-proxy iframe pointing at
:18789 (Hermes admin port) — which only listens when the gateway
runs in channel mode, and even then doesn't serve a browser UI.

The SRE Chat page now shows three explicit operator paths in
copy-pasteable code blocks:

  1. kars sre talk  → kubectl exec REPL (live triage)
  2. kars credentials update sre --telegram-token …
                    → wire Telegram for proactive alerts
  3. kars sre status / actions / show <id>
                    → terminal-friendly snapshot

Plus a link back to /kars/sre (the Console) for the live approval
queue + cluster health cards. The 'iframe with connection refused'
error is gone; v0.6.1 → v0.6.2 to bust the host's plugin cache.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: replace internal Link with plain anchor + bump to 0.6.3

The internal <Link routeName="kars-sre-console"> in SREChat may
have been the source of the React error 310 — Headlamp's Link
implementation uses hooks internally and a routeName resolution
miss can fire conditional hook paths. Using a plain <a> anchor with
the canonical Headlamp URL avoids that branch entirely.

The bundle was also showing as stale (browser cached old dist) — v0.6.3
bumps the version to force a re-fetch on the host's plugin loader.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: embed hermes dashboard PTY chat in browser

Replaces the 'no embedded WebUI' instruction page with a real iframe
into the Hermes dashboard — an in-browser xterm.js PTY chat. The
operator can now talk to the SRE agent without leaving Headlamp.

How it works:

  1. sandbox image: pip-installs FastAPI + uvicorn + websockets +
     ptyprocess (the soft-optional deps hermes dashboard needs).
     Upgrades hermes-agent from 0.15.2 → 0.16.0 to pick up the
     dashboard_auth submodule that 0.15.2 was missing.

  2. entrypoint.sh: launches 'hermes dashboard --host 0.0.0.0
     --port 9119 --no-open --insecure --skip-build' alongside the
     gateway when SRE_ENABLED=true. HERMES_DASHBOARD_TUI=1 enables
     the embedded PTY tab. Opt-out via HERMES_DASHBOARD_ENABLED=false.

  3. controller: adds containerPort 9119 ('dashboard') to Hermes
     agent containers, and exposes it on the per-sandbox Service so
     the cluster apiserver proxy can reach it.

  4. Headlamp plugin: SREChat replaces the instruction page with an
     iframe pointing at
     /clusters/<cluster>/api/v1/namespaces/kars-sre/services/sre:9119/proxy/.
     Includes 'Open in new tab' fallback for cases where the sub-
     path proxy strips Hermes web bundle asset paths. v0.6.3 → v0.7.0
     to bust the host's plugin cache.

The '--insecure' flag is required when binding off-loopback inside
the pod — Hermes refuses non-127.0.0.1 binds without it. In our pod
the only reachers are the apiserver proxy + peer sandboxes (both
gated by RBAC + NetworkPolicy), so 'insecure' here doesn't mean
externally exposed.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: fix dashboard iframe — fetch HTML + rewrite asset URLs

Two changes work together to make the in-browser Hermes dashboard
load inside the Headlamp SRE Console iframe:

1. New runtimes/hermes/src/kars_runtime_hermes/dashboard_proxy.py
   Tiny FastAPI middleware wrapper around hermes_cli.web_server.app.
   Installs X-Forwarded-Prefix on every request from the
   HERMES_DASHBOARD_PREFIX env var. Hermes' dashboard reads that
   header to rewrite absolute asset URLs (/assets/...) for sub-path
   reverse proxies. K8s apiserver service proxy doesn't inject that
   header, so without this wrapper the SPA blank-loads in the iframe.

2. entrypoint.sh now boots that wrapper instead of 'hermes dashboard',
   with HERMES_DASHBOARD_PREFIX set to the K8s apiserver suffix:
     /api/v1/namespaces/<ns>/services/<svc>:9119/proxy

3. Headlamp SREChat fetches the dashboard HTML up front via the
   Headlamp proxy, rewrites asset paths to include /clusters/<cluster>
   (the Headlamp-added prefix that the in-pod wrapper can't know
   about), and injects via iframe srcDoc. Also injects <base href>
   so the SPA's relative fetch() calls resolve under the proxy.

v0.7.0 → v0.7.1 to bust the host's plugin cache.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* headlamp/sre: dashboard wrapper strips proxy prefix to dodge /api/* collision

The K8s apiserver-proxy URL prefix
/api/v1/namespaces/<ns>/services/<svc>:<port>/proxy starts with /api/v1
— which collides with Hermes' own /api/* route namespace. So when the
browser fetched a SPA asset like
/api/v1/namespaces/kars-sre/services/sre:9119/proxy/assets/index.js,
FastAPI matched it to its API router (401 Unauthorized) instead of
the static-file mount.

Fix: extend dashboard_proxy.py middleware to STRIP the prefix from
scope["path"] before FastAPI sees the request, while still injecting
X-Forwarded-Prefix so the SPA's index.html bootstrap rewrites asset
URLs with the absolute prefix. Result: browser fetches
.../proxy/assets/foo.js, middleware strips → FastAPI sees /assets/foo.js
→ static-file mount serves it → 200 OK.

Smoke test verified end-to-end:
  asset via prefix: HTTP 200
  index via prefix: HTTP 200

Headlamp SREChat still uses srcDoc + double-prefix rewrite because
Headlamp's apiserver proxy adds /clusters/<cluster> ON TOP of the
K8s suffix — the in-pod wrapper can't know <cluster>, so the
browser-side rewrite adds it.

v0.7.1 → v0.7.2 to bust the host's plugin cache.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: end-to-end embedded Hermes chat in Headlamp plugin

Three stacked bugs blocked the SRE Console's Chat tab from working
end-to-end. All fixed:

1. Headlamp's apiserver proxy demands Authorization: Bearer on every
   /clusters/<c>/api/v1/.../proxy/* call. Headlamp's SPA fetch wrapper
   attaches it; iframe asset loads bypass the wrapper and 403 as
   system:anonymous. Plugin v0.7.4 drops the apiserver-proxy approach
   entirely and iframes http://localhost:19119/ via a user-run
   port-forward. Cross-port = different origin so parent/child JS is
   isolated, but iframe document loads aren't same-origin-gated.

2. The dashboard_proxy wrapper bypasses Hermes' start_server() (to
   install X-Forwarded-Prefix middleware first), which is where Hermes
   sets app.state.bound_host/port. Without those, _build_gateway_ws_url
   returned None and the PTY-spawned hermes --tui child got no
   HERMES_TUI_GATEWAY_URL env var — accepting keystrokes but with
   nowhere to send them. _set_bind_state() mirrors what start_server
   does.

3. Azure Linux 3 ships Node 24; Hermes' ui-tui esbuild bundle was
   built against Node 22 and SIGSEGVs immediately on Node 24 (380MB
   core dumps). Dockerfile now pins Node 22.20.0 at /opt/node22/,
   entrypoint exports HERMES_NODE=/opt/node22/bin/node so Hermes'
   _node_bin() picks it up.

Plus:
- model.context_length: 200000 pinned so cold-start skips the slow
  /v1/models probe.
- GATEWAY_ALLOW_ALL_USERS=true on the SRE sandbox so the single-operator
  loopback deploy doesn't drop our own messages.
- entrypoint passes HOME/HERMES_HOME/HERMES_NODE through runuser's env
  reset via explicit env VAR=$VAR invocation.

Plugin bumped to 0.7.4. Verified end-to-end: chat opens, accepts
keystrokes, agent responds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: ACR-name typo + workload-aware SRE Cluster Health card

Two fixes that surfaced during demo dry-run:

1. ACR typo (introduced in 5c87e9de during the azureclaw\u2192kars rename)
   - 'kars.azurecr.io' was a search/replace artifact from 'azureclaw.azurecr.io';
     the actual ACR is 'karsjpdyyv.azurecr.io' (azd-suffixed). The canonical
     name we use in chart values + controller defaults is 'karsacr.azurecr.io'
     so operators have ONE name to re-publish to.
   - Symptom: when an existing sandbox spawned a sub-agent via kars_spawn, the
     controller minted the new Deployment with 'kars.azurecr.io/openclaw-sandbox:latest'.
     Kubelet did DNS on 'kars.azurecr.io' \u2192 NXDOMAIN \u2192 ImagePullBackOff loop.
   - Fixed in:
     - deploy/helm/kars/values.yaml (4 sites: controller, inference-router,
       sandbox, a2a-gateway)
     - cli/src/commands/dev/local-k8s.ts (inverted target/aliases shape:
       'target' is now the canonical name the controller expects, 'aliases'
       are the local build tags we look up + retag from)
     - tools/demo/scenarios/01-sandbox.yaml

2. Workload-aware Cluster Health (Headlamp plugin 0.7.5)
   - KarsSandbox CR's 'phase=Running' fires the moment the controller
     successfully reconciles the Deployment spec; it knows nothing about
     whether the pods inside actually pulled their image, passed readiness,
     or got OOM-killed. The old SREClusterHealthCard read phase only \u2192
     'all green' even when break.sh had killed every pod.
   - SREClusterHealthCard now cross-checks each sandbox against its
     underlying Deployment (kars-<name>/<name>) and surfaces three buckets:
       Healthy        \u2014 CR Running AND availableReplicas \u2265 desired
       Workload down  \u2014 CR Running BUT availableReplicas < desired (the
                        false-green case)
       CR-Degraded    \u2014 CR-level Degraded=True
   - Bonus: per-sandbox breakdown panel lists which ones are unhealthy
     and points the operator at 'kars-<name>' namespace for pod-level
     diagnosis. Matches the SRE agent's own diagnosis output.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(monitoring): include kars-ops dashboard in Grafana sidecar configmap

The grafana-dashboard-configmap.yaml only wrapped grafana-dashboard-kars-fleet.json
but not grafana-dashboard-kars-ops.json — even though both JSON files have lived
in deploy/monitoring/ since May 27. Result: the Headlamp plugin's SandboxMetricsCard
iframes a 'Dashboard not found' page (it targets uid=kars-ops).

Regenerated the configmap YAML from both .json files so the grafana-dashboard sidecar
picks up both on next kars dev run. No JSON content changed; just plumbing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* hermes: pre-warm AGT mesh registration in idle-gateway mode

`hermes gateway run --accept-hooks` in idle-daemon mode (no Telegram/Slack/
Discord channels configured) runs only the cron ticker — it never imports
the kars Hermes plugin, so the Phase A2.1 eager MeshClient init at
plugin load never fires. Result: a Hermes sandbox is invisible on
`kars_mesh_directory` listings until something else triggers a plugin
load (e.g. an interactive `hermes chat` invocation, which spins up a
short-lived process that registers + exits).

Adds a 5-line pre-warm in entrypoint.sh that runs `_get_or_init_client()`
in a short-lived background Python process at boot — register_self is
idempotent + restart-safe so re-runs are cheap. Guarded on:
  - SRE_ENABLED != true       (SRE agents are intentionally off-mesh)
  - KARS_MESH_PROVIDER == agt  (only run when the mesh is actually wired)

Verified on kind: research sandbox now logs '[kars-hermes] mesh pre-warm:
registered' within ~2s of pod boot, and shows up on the AGT registry's
live-agents endpoint before any chat invocation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* hermes: persistent mesh-keepalive (replaces short-lived pre-warm)

Followup to fcce016 — the short-lived pre-warm Python process registered
on the relay then EXITED, taking the MeshClient socket with it. Without
a live connection there's no relay heartbeat, so the AGT registry marks
the agent stale after ~90s and discovery tools hide it ('Stale/offline
filtered out').

Replaces the pre-warm with a long-lived 'kars-mesh-keepalive' process
that:

  1. Calls _get_or_init_client() to register + connect (same eager path
     the plugin would take if loaded by the gateway)
  2. Calls mesh_worker.start_worker() so the sandbox can REPLY to
     inbound mesh messages (not just appear in directory listings) —
     same auto-responder the controller wires into kars_spawn'd
     sub-agents via KARS_MESH_AUTO_RESPONDER=1
  3. Parks on threading.Event().wait() forever so the MeshClient stays
     alive and keeps heartbeating

Verified on kind: research's keepalive log shows registered + connected
+ worker started; dev-agent's mesh discover can now see research.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* hermes: enable AUTO_RESPONDER on the mesh keepalive process

Follow-up to 163e1de — the keepalive's mesh_worker.start_worker() was
draining inbound messages and silently dropping them because the worker
gates LLM replies behind KARS_MESH_AUTO_RESPONDER (mesh_worker.py:259).

Couldn't set the env var via the KarsSandbox CR's extraEnv because the
controller's reserved-prefix guard (reconciler/mod.rs:1820) strips any
user-supplied KARS_* env. Set it inline on the keepalive's exec env
instead — that's the only process that runs the worker, so a
process-local env var is sufficient.

After this fix: dev-agent → research mesh send now triggers an actual
Hermes-generated reply via the auto-responder.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* demo: bump dailyTokens cap to 2M for research + sre

The 500K cap (the InferencePolicy default when dailyTokens is unset)
exhausts trivially in a live demo — one 175K-context conversation
through a couple of turns already crosses it, after which the
inference-router throttles and the agent can't reply. The Headlamp
plugin's token-budget panel renders this as '100% used', looking like
a misconfiguration when it's actually intentional governance.

Sets explicit 2M for research (demo scenario) and sre (Helm template
default with a value-override path). Operators in production with
strict cost controls can override via:
  --set sre.dailyTokens=N
  edit the research scenario yaml inline

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* plugin: workload-aware Phase column on Overview + Sandboxes pages

Same false-Running problem as the SRE Cluster Health card (fixed in
5f1c2ee) affected the Overview's 'Ready' headline stat and the
Sandboxes list's Phase column. Both read KarsSandbox.status.phase,
which the controller sets to 'Running' the moment the Deployment
spec is reconciled — independent of whether the pods inside
actually pulled their image / passed readiness / etc.

Two visible bugs:
  - Overview's 'Ready' stat counted 'phase === "Ready"' but the
    controller never sets that — it uses 'Running'. So 'Ready'
    always showed 0 even with all sandboxes healthy.
  - Sandboxes Phase column showed 'Running' for a sandbox whose
    Deployment was at 0/1 available (ImagePullBackOff, OOMKilled,
    etc.) — directly contradicting reality.

Fixes both by pulling Deployments alongside KarsSandbox and
cross-checking availableReplicas >= spec.replicas before declaring a
sandbox 'Healthy'. Overview headline stats are now:
  Healthy        — CR Running AND workload available
  Workload down  — CR Running BUT workload unavailable
  CR-Degraded    — CR-level Degraded=True condition
Sandboxes list shows 'Workload down' (red StatusLabel) in the Phase
column when the underlying Deployment can't meet its replica count.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre-action: workload-aware recovery observer (no false Recovered)

The Slice 3 recovery observer declared an action 'Recovered' as soon
as there were no FailedCreate / BackOff / FailedScheduling events on
the target namespace in the last 30s. False positive on the canonical
DeleteResourceQuota path: deleting the quota silences new
FailedCreate events (no more ReplicaSet attempts), but the Deployment
can still sit at 0/1 because the ReplicaSet was scaled to 0 during
the failure cascade and no controller is going to scale it back up.

Result before this fix: action.phase=Recovered while the workload
was still down, directly contradicting what the operator sees in
Headlamp's plugin (the Sandboxes / Overview / Cluster Health cards
all show 'Workload down' for the same sandbox post-fix).

Tightens observe_recovery to require BOTH:
  (1) absence of recent failure events on the target namespace
      (existing gate), AND
  (2) every Deployment in the target namespace at
      availableReplicas >= spec.replicas
      (the gate the doc comment promised for Slice 4)

The Deployments gate runs first because it's the more authoritative
signal — if pods aren't available, recovery hasn't happened
regardless of what the event log shows.

Verified live on kind: created a test KarsSREAction targeting a
broken research deployment; the action stayed at phase=Applied
through 3 reconcile passes (workload still down), then flipped to
Recovered on the next pass after the deployment came back to 1/1.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* demo: 3 commented sandbox CRDs for the Act-I walkthrough

* controller: stop spamming LimitedSupport event on every McpServer reconcile

The McpServer reconciler emitted a Warning event with reason=LimitedSupport
on every successful reconcile (~15s cycle), repeating the same static
'singular spec.mcp binding today, plural lands in Slice 4' text. Result:
Headlamp's event view was permanently polluted with the same advisory
message for every McpServer CR, drowning out actually-actionable events.

The information belongs in CRD descriptions and design docs, not in the
per-incident K8s Event stream. Removed the call site; kept a breadcrumb
comment pointing future readers at the right places to publish the
roadmap (mcpserver.spec CRD description + crd-well-oiled-machine
blueprint).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: phase-changes-only watcher mode (Telegram pager, not event firehose)

Adds SRE_WATCHER_MODE=phase-changes-only which alerts ONLY on KarsSandbox
status.phase transitions (Running -> Failed -> Recovered) instead of the
default event stream. One Telegram message per real CR state change, no
pod-level event noise.

Default mode in the Helm chart is now phase-changes-only because that
matches what most operators actually want — a sandbox-level status pager.

Uses the same sre_kube.client() httpx singleton the event-mode watcher
uses (the distroless sandbox image has no kubectl). Verified live:
watcher primes with the current set of KarsSandboxes and only emits on
true transitions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre: overlay workload availability on synthetic phase

Both the proactive phase-changes-only Telegram watcher AND the
sre_diagnose chat tool only looked at KarsSandbox.status.phase, which
the controller doesn't flip when downstream pods break (e.g. evicted
pod can't re-admit due to a tight ResourceQuota, image-pull failure,
NodeAffinity unmet). The CR stayed Running while the Deployment was
0/1, so neither the pager nor the in-chat diagnose noticed.

Fix:
* sre_watcher._workload_state(): for each KarsSandbox, fetch the
  matching Deployment in kars-<name> and synthesize WorkloadDown(a/d)
  when available < desired. Transitions on that overlay fire one
  Telegram message per real state change — still no event-firehose
  noise.
* sre._impl_sre_diagnose: cross-checks Deployment availability for
  every KarsSandbox and adds WorkloadDown entries (with the affected
  ns + deploy name) to degraded_sandboxes. The LLM can now describe
  workload-level incidents accurately when the operator asks
  "what's wrong with my cluster?".

Verified live: research deployment was 0/1 (quota-violation, Act II
break.sh scenario). After healing the quota, the watcher fired one
Telegram alert: research: WorkloadDown(0/1) -> Running.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* sre-action: bump recovery window 5m→10m + late-recovery healer

Demo on 2026-06-11 hit a real-world false negative: SRE applied the
DeleteResourceQuota patch, observed for 5 min, marked the
KarsSREAction Failed — but research actually recovered ~1 min later.
The terminal Failed state then stuck even though the cluster was
fine, leaving the operator with a misleading state.

Two fixes:

1. RECOVERY_WINDOW_SECONDS 300 → 600. Real K8s recovery routinely
   exceeds 5 min on cold caches, RS back-offs, congested nodes.

2. Late-recovery healer (Failed → Recovered edge). For Failed CRs
   that DID reach Apply (i.e. have appliedAt set — pre-apply
   validation failures don't qualify), the terminal handler keeps
   running observe_recovery for LATE_RECOVERY_WINDOW_SECONDS = 30 min
   since appliedAt. If recovery is observed, flip phase back to
   Recovered with reason=LateRecovery. Polling cadence during this
   window is 60s (vs the standard 300s terminal requeue) so latency
   is bounded.

State-machine docs at the top of the file updated to reflect the new
Failed → Recovered edge. Existing tests (6) still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(security): audit for kars-sre demo-and-agent slice (Slices 0-4 + recovery healer)

Covers all 46 commits on this branch since main. Documents:
- T1: SRE writer SA escalation surface (mitigated: 7-layer a…
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request Jun 26, 2026
Address the remaining war-room recommendations:

- Above-the-fold "wow path" (#1): the hero code block now shows the full
  path to a chat (install -> `kars dev --release --target local-k8s` ->
  `kars connect dev-agent`), the first-run GIF moves directly under it (was
  buried in "Try it in five minutes"), and a "Full quickstart ->" link points
  to docs/quickstart.md. A visitor sees what it is, the three commands, and a
  GIF of it working without scrolling.
- Trim internal-detail overload (#4): compress the heaviest README paragraph
  (mesh crate pins / vendored-tgz mechanics) to the essential, verifiable
  security claim plus a link to architecture.md#the-mesh, which carries the
  full provenance. The CRD/runtime tables, inference-router spotlight, and
  honest "Known limitations" list are kept intact.
- Nit: drop two CONTRIBUTING references to the internal `plan.md` (public
  contributors can't see it) in favour of "open a follow-up tracking issue".

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request Jun 26, 2026
…s/internal (#468)

* docs: stop tracking docs/internal/ (was publicly exposed despite .gitignore)

The docs/internal/ tree (strategic plans, competitive analyses, security-audit
logs, announcement/blog drafts, POC lab-notes) was force-added to git in an
earlier change and would have shipped publicly at launch despite being listed
in .gitignore. Untrack the entire tree with `git rm -r --cached` — the files
remain on disk for maintainers but leave the public repository.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: rebuild information architecture (retire/relocate internal-bound pages)

Move maintainer-only and historical pages out of the public surface and rebuild
the navigation so every published page has exactly one home:

- Relocate to docs/internal/ (now untracked): PUBLISHING.md, the entra-agent-id
  POC lab-notes (00-poc-archive, 02-aci-token-flow, 03-original-findings,
  04-migration-guide), blueprint 07 (kars-sre proposal), and showcase/outline.
- Promote docs/sre.md to docs/runbooks/sre.md (now a shipped runbook, not a
  proposal) and add it to the nav.
- Rebuild docs/SUMMARY.md and docs/README.md: drop retired/moved pages, add the
  supply-chain posture page and the BYO runtime contract, nest the entra deep
  dive under Architecture so it renders.
- Trim the entra-agent-id index to the pages that remain public.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: per-page accuracy and consistency pass for launch

Ground every public claim in the code and remove overlap/contradictions:

- Counts: twelve CRDs (ten workload incl. KarsSREAction + two infrastructure)
  and eight runtimes, propagated everywhere; add a full KarsSREAction section to
  the CRD reference and lifecycle.
- Versions: v0.1.18 across README/installer/examples; package is @kars-runtime/cli
  (the only published npm package); org links microsoft/kars -> Azure/kars.
- Mesh provenance: correct the AGT SDK provenance and agentmesh 3.1.0 -> 4.0.0;
  describe the OpenClaw (TS SDK) vs Hermes/Python (Python mesh client) split.
- Status honesty: add build-from-source/roadmap banners to mesh-plugin,
  a2a-gateway (verifier-is-library), runtimes/CONTRACT (draft surface), and
  notation-ratify (no kars up --sign-images flag exists).
- getting-started: collapse to a single linear path; document the provider
  picker exactly once; make kind/local-k8s a first-class step; source build is
  secondary. README "Try it" reframed kind-first.
- Ecosystem: frame agentgateway and kubernetes-sigs/agent-sandbox as an aim to
  align (no discussions/integrations claimed).
- Fix broken relative links, de-orphan the supply-chain posture page, add a
  pricing disclaimer to the cost dashboard, and grammar ("an kars" -> "a kars").
- Update the sre.ts design-comment path to docs/runbooks/sre.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(site): commit custom mdbook theme and fix Mermaid rendering

The polished site styling lived in an untracked docs/site/theme/ directory, so a
clean CI/Pages build would fall back to stock mdbook and lose all of it. Commit
the theme and fix three rendering defects found in visual review:

- Track theme/index.hbs (kars top nav + logo), theme/css/custom.css (design
  tokens, CTA buttons, card tables/admonitions, framed code, diagram cards),
  and the brand favicons; wire them via book.toml (theme + additional-css).
- Mermaid contrast: subgraph titles rendered white-on-light on the dark site —
  add an explicit dark color to every light-fill classDef across the diagrams.
- Mermaid clipping: long cluster titles wrapped to a hidden second line — raise
  wrappingWidth and let cluster labels render at natural width via CSS.
- Mermaid init: guard getElementById theme-toggle hooks that were absent in the
  custom theme and threw a console error on every page.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(cli): don't require Rust toolchain for `kars dev --release`

`kars dev --release` pulls pre-built, signed images from ghcr.io/azure and
compiles nothing, but the preflight unconditionally required `cargo` on PATH —
breaking the documented "no Rust" quickstart promise. Require `cargo` only when
building locally from source (no `--release`). Extract the tool list into a pure,
exported `requiredToolsFor()` and add regression tests asserting `cargo` is
absent in `--release` mode and present otherwise.

Also remove the dead `--yes` flag from `kars upgrade` (no code read it; the
command is already non-interactive) so docs and code agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add quickstart, kars upgrade runbook, llms.txt, and top-OSS landing polish

Apply the launch-readiness scorecard derived from agentgateway / Cilium / Dapr /
cert-manager docs:

- Add docs/quickstart.md — a 3-command, ≤5-minute path to a running agent.
- Document `kars upgrade` (previously undocumented): a full CLI reference
  section plus docs/operations/upgrades.md runbook (verified against
  upgrade.ts), wired into SUMMARY + the operations index + getting-started.
- Add docs/llms.txt (machine-readable docs index for AI tooling) and its
  generator docs/site/gen-llms-txt.py, regenerated from SUMMARY.md.
- Landing page: architecture diagram above the fold, a "Why kars" comparison
  table, a Quickstart CTA, and a Feature-status link.
- Add "Last tested with kars v0.1.18" footers to the six blueprints and a
  "which loop do I want?" decision table to blueprint 01.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: accuracy pass from architect + doc-expert re-audit (zero vaporware)

Correct every claim flagged as an overclaim or stale against the code:

- Audit: "Merkle-chained" -> "SHA-256 hash-chained (AGT's AuditLogger)"; the
  Merkle module is library-only and explicitly not wired (README, SECURITY,
  agt-boundary already accurate).
- Cross-runtime mesh: scope the shipped/proven claim to OpenClaw + Hermes
  (tests/e2e/interop/hermes_openclaw_bidi.sh). The other Python adapters bundle
  the client but don't yet expose the mesh tools; LangGraph ships an A2A/HTTP
  helper, not a Signal session. Fixed in README + architecture.
- Hermes is at parity (Act 2 mesh is live): rewrite the stale "stub / NOT
  IMPLEMENTED" sections in hermes-plugin, CONTRACT, and runtimes/hermes/README.
- CRD counts: "eight peer workload CRDs" -> nine; lifecycle diagram 9 -> 12.
- AGT boundary: MeshProvider is agent-side only — the router has no Rust mesh
  impl (providers/mesh.rs); correct the provider table.
- Egress: the signed-OCI allowlist IS the enforced L7 source of truth
  (egress_allowlist_loader + forward_proxy), not advisory/roadmap.
- Token budgets: daily + monthly windows are enforced at the router, not
  reconciler-only.
- `kars upgrade --rollback` reverts the Helm revision, not image bits (workloads
  pin :latest) — document the accurate rollback semantics.
- Confidential (Kata + SEV-SNP) isolation is opt-in, not active by default.
- Remove nonexistent commands from docs (`claw attest` -> `kars attest`,
  `kars offload`); add missing `kars upgrade` / `sre` / `headlamp` to the CLI
  reference; OpenClaw language is TypeScript/Node, not Python; "separate router
  pod" -> sidecar container; runtime examples 8 -> 10; multi-runtime images are
  published (best-effort import), not "not yet published".
- Make the dev flow kind-first in getting-started + CONTRIBUTING; drop public
  references to gitignored docs/internal; refresh stale CLI command count.
- Fix broken links/anchors surfaced by the sweep (crd-reference deep links,
  roadmap anchors, demo-script/byo-contract/blocklists paths).

Verified: mdBook builds clean, 0 broken links/anchors, 0 orphans; 842 CLI tests
pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(readme): make the README kind-first, not a Docker ad

The "Try it in five minutes" hero and the modes table led with Docker
(`kars dev --release`) and treated the kind dev loop as a footnote, contradicting
the kind-first framing used everywhere else.

- "Try it" now leads with `kars dev --release --target local-k8s` (kind, the real
  production pod shape) as the recommended path; the single-container Docker
  target moves into a collapsible "just want the fastest smoke test?" aside.
- Rename "Two modes" → "Three ways to run it" with a three-column table
  (Local kind / Local Docker / Prod AKS), kind marked recommended, Docker framed
  as the fast inner loop rather than the default.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(cli): kind dev loop accepts any container runtime (Docker/Podman/nerdctl)

The local-k8s (kind) runner already drives docker, podman, or nerdctl via
`KIND_EXPERIMENTAL_PROVIDER` (`dev/local-k8s.ts` RUNTIME_PRIORITY), but the
preflight hard-required the `docker` binary for the local-k8s target — so a
Podman- or nerdctl-only user was blocked at preflight despite the runner
supporting them.

- Drop `docker` from the local-k8s must-each-be-present list; add an anyOf probe
  (`CONTAINER_RUNTIMES = docker | podman | nerdctl`) that passes if at least one
  is on PATH. The single-container `docker` target still requires the `docker`
  CLI (it shells out to it directly).
- Update the preflight tip and add regression tests (kind target must not pin
  docker; docker target must).
- Docs: make the runtime story accurate — the recommended kind loop accepts
  Docker/Podman/nerdctl; the fast single-container path uses the `docker` CLI.
  Reorder the getting-started prerequisites table kind-first.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: don't imply Docker-only where the kind path is runtime-agnostic

Sweep follow-up to the Podman/nerdctl fix. The single-container `docker` target
genuinely uses the `docker` CLI, so those references stay — but two spots tied to
the kind path / general first-run wrongly implied Docker specifically:

- blueprints/02 topology mermaid: "docker build && load" -> "build/pull + load
  images via the detected runtime" (kind drives docker/podman/nerdctl).
- getting-started troubleshooting: "Docker Desktop is not running / Start Docker"
  -> "the container runtime isn't running / start Docker Desktop, podman machine,
  or colima".

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add first-run demo (asciinema) to the README + quickstart hero

A real, recorded `kars dev --release --target local-k8s` first run on a local
kind cluster — provider picker (GitHub Copilot), bring-up of the controller +
encrypted mesh + sandbox, ending at a governed agent.

- docs/assets/kars-dev-firstrun.gif (1.9 MB, optimized) embedded in the README
  "Try it in five minutes" hero and docs/quickstart.md.
- docs/assets/kars-dev-firstrun.cast — the replayable asciinema cast, linked
  from the quickstart caption.

Recording hygiene: the cast is scrubbed of the recorder's internal Foundry
endpoint (-> my-foundry/my-project placeholder) and of all dead local tokens
(Headlamp SA JWT, OpenClaw gateway/WebUI tokens, did:mesh id, image digests).
Verified: no secrets remain; mdBook builds and the asset renders.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: CMO final-review fixes (consistency, conversion, credibility)

From the launch war-room review:

- Resolve a credibility contradiction: roadmap.md said aggregate token budgets
  are "not yet metered" while maturity.md (correctly) says daily+monthly are
  enforced. Reword the roadmap item to claim only what remains (per-hour windows
  + a configurable rejectOnExceed knob), matching budget.rs.
- Make docs/quickstart.md kind-first so the page matches its own first-run GIF
  and the README hero: step 2 is `kars dev --release --target local-k8s`,
  prerequisites are kind + kubectl + any container runtime, and the single
  container `docker` path is the explicit "even faster, less faithful" alt.
- README: surface the community on-ramp (good first issue / help wanted /
  Discussions) in the Contributing section instead of only doc links.
- README: replace the "print the current Azure subscription" sample prompt
  (odd in a no-Azure quickstart) with a neutral one; soften the slightly
  promissory "fit seamlessly" ecosystem wording to "can grow toward".

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: CMO final-review — above-the-fold wow path + trim README detail

Address the remaining war-room recommendations:

- Above-the-fold "wow path" (#1): the hero code block now shows the full
  path to a chat (install -> `kars dev --release --target local-k8s` ->
  `kars connect dev-agent`), the first-run GIF moves directly under it (was
  buried in "Try it in five minutes"), and a "Full quickstart ->" link points
  to docs/quickstart.md. A visitor sees what it is, the three commands, and a
  GIF of it working without scrolling.
- Trim internal-detail overload (#4): compress the heaviest README paragraph
  (mesh crate pins / vendored-tgz mechanics) to the essential, verifiable
  security claim plus a link to architecture.md#the-mesh, which carries the
  full provenance. The CRD/runtime tables, inference-router spotlight, and
  honest "Known limitations" list are kept intact.
- Nit: drop two CONTRIBUTING references to the internal `plan.md` (public
  contributors can't see it) in favour of "open a follow-up tracking issue".

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: complete docs/internal untracking (drop v0.1.19 audit docs too)

The launch branch stops tracking docs/internal/ (internal planning + security
audit content not meant for the public repo). Three audit docs were added on
main after this branch was cut (the v0.1.19 memory + dev fixes); remove them
too so docs/internal is fully untracked, consistent with 631b09e. The audit
content is preserved in git history.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file rust Pull requests that update rust code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant