Repository navigation
Bump jsonwebtoken from 9.3.1 to 10.3.0 - #1
Merged
Pal Lakatos-Toth (pallakatos) merged 1 commit intoMar 21, 2026
Merged
Conversation
Bumps [jsonwebtoken](https://github.com/Keats/jsonwebtoken) from 9.3.1 to 10.3.0. - [Changelog](https://github.com/Keats/jsonwebtoken/blob/master/CHANGELOG.md) - [Commits](Keats/jsonwebtoken@v9.3.1...v10.3.0) --- updated-dependencies: - dependency-name: jsonwebtoken dependency-version: 10.3.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com>
Collaborator
|
reviewed and approved - we can bump this too |
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
Apr 14, 2026
…wlist clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill). openclaw.ai enables `curl | bash` install vectors from inside sandboxes. Both were in the default Helm values and example CRD. Removed from: - deploy/helm/azureclaw/values.yaml (default egress allowlist) - examples/basic-agent/clawsandbox.yaml (example CRD) - PLAN.md (policy presets documentation) Defense is now three layers deep: 1. Egress proxy blocks clawhub.com/openclaw.ai (network) 2. Skills directory is root-owned, chmod 640/750 (filesystem) 3. Plugin code is root-owned, read-only for sandbox (code integrity) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Apr 14, 2026
* docs: add global AgentMesh handoff design document
Comprehensive design for agent live migration (local ↔ cloud):
- Identity succession protocol (Ed25519 signed, no key transfer)
- Reclamation protocol (co-signed reverse handoff)
- Sub-agent re-spawn with state injection
- Three-layer handoff endpoint auth (handoff token + no localhost bypass + mutual attestation)
- Security review: 11 threat findings with mitigations
- Handoff trigger security (confirmation token, time delay, AGT policy gate)
- UX design across webchat, TUI, and Telegram
- Demo script and implementation phases (H1-H4)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H1 — handoff module with three-layer auth
Router-side handoff infrastructure for agent live migration (local ↔ cloud):
## New module: handoff.rs (1368 lines)
- HandoffState, SubAgentSnapshot, HandoffMetadata, CredentialRef structs
- HandoffTokenStore: in-memory, TTL-based, one-at-a-time token management
- 32-byte random tokens, max 10min TTL, constant-time comparison
- Token hash logged for audit (never the token value)
- HandoffSession: phase tracking across the full handoff lifecycle
(idle → initialized → draining → snapshotting → transferring → restoring
→ verifying → decommissioning → complete | failed | aborted)
- DrainState: stops new work during handoff, tracks duration
- State serialization: JSON + gzip compression
- State encryption: AES-256-GCM with HKDF-SHA256 key derivation
Key derived from shared secret + salt using 'azureclaw-handoff-v1' info
- Verification: SHA-256 hash of plaintext for integrity checking
- 21 unit tests covering token store, serialization, encryption, sessions
## New endpoints (8 routes, three auth tiers)
1. POST /agt/handoff/init — admin token only, NO localhost bypass
2. POST /agt/handoff/snapshot — creates encrypted state blob
3. POST /agt/handoff/restore — decrypts, validates, restores state
4. POST /agt/handoff/verify — returns verification digest
5. POST /agt/handoff/drain — enters drain mode
6. POST /agt/handoff/decommission — agent goes dormant
7. POST /agt/handoff/abort — cancels in-progress handoff
8. GET /agt/handoff/status — read-only (localhost allowed)
## Security: three-layer authentication
- Layer 1: Handoff token (one-time, short-lived, CLI-only)
Token exists only in CLI process memory — never in pod env
- Layer 2: NO localhost bypass for mutation endpoints
Prevents prompt injection from exfiltrating state via localhost
- Layer 3: Mutual attestation via DH-encrypted state blob
(Phase H2 adds Ed25519 succession signature verification)
## All endpoints audit-logged with:
- Caller IP, timestamp, endpoint, success/failure
- Token hash (not value), state blob size, item counts
## Dependencies added:
aes-gcm 0.10, hkdf 0.12, sha2 0.10, rand 0.9, base64 0.22, flate2 1
## Test results:
- 77 unit tests pass (21 new handoff tests)
- 26 integration tests pass (updated for new AppState fields)
- 74 controller tests pass (unaffected)
- clippy clean (zero warnings)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H2 — registry mode, identity succession, and reclamation
Phase H2 of the agent handoff feature:
Registry topology (local vs global):
- Add RegistryMode enum to router config (AGT_REGISTRY_MODE env var)
- Handoff init returns 409 in local mode with clear guidance
- Global mode does startup health check on AGT_REGISTRY_URL
- Handoff status endpoint exposes registry_mode + handoff_available
Identity succession (A→B):
- SQL migration 008_succession.sql with succession_log table
- POST /v1/registry/succession endpoint with Ed25519 sig verification
- Canonical message format: succession:{pred}:{succ}:{timestamp}
- One-shot rule (unique index on active predecessor)
- Copies reputation A→B, marks predecessor dormant
Identity reclamation (B→A, co-signed):
- POST /v1/registry/reclamation with dual signature verification
- Original succession ref must match active event_hash
- Deactivates succession redirect, copies reputation back
- Sets original online, departing offline
Lookup follows succession redirects:
- lookup_agent checks succession_log for dormant predecessors
- Returns successor with succeeded_from + succession_hash metadata
- Max redirect depth = 1 (no chains)
Dormant presence status:
- New PresenceStatus::Dormant variant in registry
- Ghost cleanup skips dormant agents (preserves succession chains)
- Capability search excludes dormant agents
CLI --global-registry flag:
- azureclaw dev --global-registry <url> skips local registry stack
- Passes AGT_REGISTRY_MODE=global to router
- Health check on global registry at startup
- Status display shows "handoff enabled" for global mode
Tests: 177 Rust (77 unit + 26 integration + 74 controller) + 159 CLI
All passing, clippy clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H3 — CLI command and plugin tools
CLI command (cli/src/commands/handoff.ts):
- azureclaw handoff <name> --to cloud (forward: local → AKS)
- azureclaw handoff <name> --to local (reverse: AKS → local)
- azureclaw handoff <name> --status (check progress)
- azureclaw handoff <name> --abort (cancel in-progress)
- 7-step orchestration: verify → init → snapshot → drain → transfer → succession → summary
- Registry mode guard (blocks with clear guidance if not --global-registry)
- Shared secret derivation for snapshot encryption
- Full error recovery (abort on failure)
Plugin tools (cli/src/plugin.ts):
- azureclaw_handoff_status: check handoff availability and state
- azureclaw_handoff_request: LLM can request handoff, but returns CLI
command (tokens are CLI-only to prevent prompt injection attacks)
Security design:
- Handoff tokens live in CLI process memory only (never LLM-accessible)
- LLM tool returns the CLI command to run, not the token itself
- This prevents prompt injection from triggering unauthorized handoff
Registered in cli.ts under 'Agent mobility' section.
Tests: 177 Rust + 159 CLI, all passing, clippy clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security(handoff): close §9.9 security gaps — confirmation gate, injection protection, rate limits
§9.9.9 — Two-stage confirmation gate (HIGH priority):
- PendingHandoffStore with confirmation token (4-byte random hex)
- POST /agt/handoff/pending — creates pending request (rate limited: 1 per 5min)
- POST /agt/handoff/confirm — validates token, enforces 3s minimum delay
to prevent LLM self-confirm, consumes pending on success
- Plugin azureclaw_handoff_request now calls /pending (Stage 1)
- New plugin azureclaw_handoff_confirm tool (Stage 2)
- 6 unit tests for PendingHandoffStore (create, confirm, delay, rate limit,
cancel, wrong token)
§9.9.1 — State blob prompt injection protections:
- sanitize_chat_snapshot() strips messages matching 17 injection patterns
(system prompt override, handoff commands, instruction ignoring)
- User messages always preserved (legitimate user content)
- Non-UTF8 chat snapshots rejected entirely
- Trust scores capped at 750 on restore (cannot import max trust)
- 4 unit tests for chat sanitization
§9.9.4 — State blob size/DoS limits:
- 50MB blob size cap on both snapshot and restore
- MAX_WORKSPACE_FILES (100) and MAX_WORKSPACE_FILE_SIZE (10MB) constants
- PAYLOAD_TOO_LARGE (413) returned on violation
§9.9.3/§9.9.8 — Rate limits:
- Succession rate limit: 1 per AMID per 5 minutes (DB-backed)
- Reclamation rate limit: 1 per AMID per hour (DB-backed)
- check_succession_rate_limit() queries succession_log timestamps
§9.9.9 — AGT policy rule (belt-and-suspenders):
- handoff-tool-approval rule in azureclaw-default.yaml
- type: approval, priority: 75 (higher than tool-allow at 70)
- Requires operator approval for tool:azureclaw_handoff_request:*
and tool:azureclaw_handoff_confirm:*
Tests: 188 Rust (74 controller + 88 router + 26 integration) + 159 CLI
All passing, clippy clean, registry cargo check clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): implement global registry deployment — Ingress, OAuth, relay auth, CLI
Phase G1 implementation:
- G1a: AGIC Ingress manifest (deploy/agentmesh-ingress.yaml) with
NetworkPolicy (postgres locked to registry, registry/relay to AppGW),
Azure-managed TLS, WAF rate limiting, WebSocket support for relay
- G1b: Entra ID OAuth provider added to agentmesh-registry
(authorize, callback, token validation via Microsoft Graph).
Existing GitHub + Google providers untouched.
- G1c: Deployment manifest updated with OAuth secret references
(agentmesh-oauth-credentials), REGISTRY_URL for relay verification
- G1d: CLI 'azureclaw mesh auth' command — generates Ed25519 keypair,
runs browser-based OAuth flow, stores encrypted identity in
~/.azureclaw/mesh-identity.json (AES-256-GCM, machine-bound key).
Subcommands: auth, status, reset.
- G1e: CLI 'azureclaw up --global-registry <url>' skips local registry
deployment. '--expose-registry' deploys AGIC Ingress to make this
cluster's registry the global endpoint. Context persists registry mode.
- G1f: Relay registration verification — after Ed25519 signature check,
relay calls registry /v1/registry/lookup to confirm AMID is registered.
Unregistered/revoked agents rejected. Fails open on registry errors
(avoids cascading failures). Gated by REQUIRE_REGISTRATION=true.
Security: 4-layer auth chain (WAF → Ed25519 → registry check → OAuth).
PostgreSQL never exposed externally (NetworkPolicy enforced).
Private keys encrypted at rest (AES-256-GCM).
Tests: 188 Rust + 159 CLI passing, all clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test+docs(mesh): integration tests and security/architecture documentation
Tests:
- 28 new CLI mesh tests (mesh.test.ts): base58 encoding, Ed25519
keypair generation, AMID derivation, encrypt/decrypt roundtrip,
tamper detection, command structure verification
- 3 relay registry verifier tests (registry_verify.rs): disabled
verifier passthrough, env-based construction, enable logic
(compile-gated by pre-existing ed25519-dalek API mismatch in relay)
- All 188 Rust + 187 CLI tests passing
Documentation:
- architecture.md: new 'Global Registry Deployment' section — deployment
modes table, 4-layer auth chain diagram, NetworkPolicy enforcement,
identity management overview
- security.md: new 'Layer 9: Global Registry & Handoff Security' section
— relay auth layers table, handoff threat/mitigation matrix,
NetworkPolicy diagram, identity-at-rest encryption details
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(plugin): gate handoff mutation tools behind AGT_REGISTRY_MODE=global
In local registry mode, only azureclaw_handoff_status is registered.
The request and confirm tools are hidden from the LLM, preventing
unnecessary AGT governance prompts for tools that would 409 anyway.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): add promote/demote commands for registry global mode
azureclaw mesh promote — deploys AGIC Ingress + NetworkPolicies to expose
the cluster's AgentMesh registry and relay as public endpoints. Updates
deployment context to global mode.
azureclaw mesh demote — removes Ingress resources and reverts to
cluster-local registry. Disables cross-environment handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): add --allow-ip to promote for IP-based access control
mesh promote auto-detects your public IP (via ifconfig.me) and injects
the AGIC whitelist-source-range annotation into both Ingress resources.
Override with --allow-ip <cidr>. If detection fails, warns and leaves
the registry open.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): auto-detect AppGW IP and use sslip.io for zero-config DNS
mesh promote now queries the AGIC Application Gateway for its public IP
and generates sslip.io hostnames (e.g. registry.20-30-40-50.sslip.io).
No DNS setup needed for testing. TLS is disabled for sslip.io domains
(secured by IP allowlist instead). Use --domain for custom domains.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* refactor(mesh): switch promote/demote to LoadBalancer Services
Replace Ingress-based approach with direct LoadBalancer Service patching.
No ingress controller needed. promote patches registry + relay services
to LoadBalancer with loadBalancerSourceRanges for IP restriction, waits
for external IPs, builds sslip.io URLs. demote reverts to ClusterIP.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(plugin): LLM-driven handoff orchestration via E2E mesh
Replace the CLI-only handoff confirm flow with full LLM-driven
orchestration. After the user confirms the handoff code, the plugin
now executes the entire transfer autonomously:
1. Confirm → router creates handoff token (stays in plugin memory)
2. Snapshot → encrypted AES-256-GCM state blob
3. Drain → stop accepting new work
4. Spawn → create cloud target on AKS (or find existing)
5. Transfer → send state blob via E2E encrypted mesh (Signal Protocol)
6. Verify → target restores, sends verification digest back via mesh
7. Succession → registry identity chain update
8. Decommission → local agent enters dormant state
Key changes:
- handoff_confirm tool: full orchestration instead of returning CLI cmd
- onMessage handler: new handoff_transfer message type for target agent
to auto-restore state and send verification back
- _routerCallStrict: new helper that rejects on HTTP >= 400
- _readAdminToken: reads admin token from filesystem paths
- _routerCall: added extraHeaders parameter (backward compatible)
- agtReconnect: disconnect before connect to clear stale SDK state
Security model (§9.9.9): the LLM can REQUEST a handoff but never
EXECUTE one. The handoff token stays in plugin memory — the LLM
never sees it. All router calls use this token. Human confirmation
via the 2-stage code flow is the gate.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(mesh): promote checks health and reconnects stale port-forwards
When registry is already in global mode, 'azureclaw mesh promote' now:
- Checks registry HTTP health (/v1/health)
- Checks relay TCP connectivity
- If both healthy: reports status and exits
- If either dead: kills stale PIDs, clears held ports, restarts
fresh port-forward tunnels, verifies connectivity
Previously it just said 'already global' and exited, even when the
port-forwards had died (e.g. after IP change or sleep).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(plugin): use async import for fs in ESM context
_readAdminToken used require('node:fs') which is unavailable in ESM.
Changed to async function with await import('node:fs') and updated
both call sites to await the result.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): spawn AKS pod from dev mode for cloud handoff
In dev mode, the router's /sandbox/spawn endpoint was creating Docker
containers. For handoff (local→cloud), we need actual AKS pods.
Changes:
- Add HandoffMeta struct to SpawnRequest (mode + predecessor fields)
- When handoff.mode='restore' in dev mode, bypass Docker path and use
K8s CRD creation via kube-rs (kubeconfig mounted from host)
- Mount ~/.kube/config into dev container at /run/secrets/kubeconfig
so the router can reach the K8s API for handoff spawns
The controller already sets AGT_RELAY_URL and AGT_REGISTRY_URL on
spawned pods, and NetworkPolicy allows mesh egress — so the handoff
target automatically joins the global mesh.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): propagate trusted_peers and registry_mode to spawned pods
The handoff target was rejecting the source's KNOCK (trust score 0 <
threshold 500) because AGT_TRUSTED_PEERS wasn't propagated. Also,
AGT_REGISTRY_MODE wasn't set, so handoff tools were skipped.
Changes:
- CRD: add trusted_peers and registry_mode fields to GovernanceConfig
- Controller: propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to
the openclaw container env vars
- Spawn: write trusted_peers and registry_mode='global' into CRD
governance spec for handoff targets
- Spawn: use 'handoff'/'predecessor' labels instead of 'agent'/'parent'
for handoff-spawned CRDs (not sub-agents)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: add cloud handoff flow diagram (section 11)
Sequence diagram covering all 5 phases: two-stage confirm, snapshot/drain,
spawn on AKS, E2E mesh transfer, succession/decommission. Includes security
model diagram and current vs future (Entra OAuth) trust flow comparison.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): async orchestration with real-time progress tracking
Refactor handoff_confirm to return immediately and run orchestration
in the background via _runHandoffOrchestration(). The LLM polls
handoff_status every 3-5s and relays emoji step updates to the user
in real-time instead of blocking for 2-3 minutes.
Key changes:
- HandoffProgress interface tracks phase, steps[], status, error
- _hp() helper updates progress + logs at each step
- _runHandoffOrchestration() contains the full 7-step flow:
snapshot → drain → spawn → mesh-wait → transfer → verify →
succession → decommission
- handoff_status returns rich progress with active polling instruction
- Module-level _log set during register() for background access
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address security and reliability findings from handoff audit
sec-6: Fix message filter AND→OR — verification now rejects messages
unless BOTH from_amid AND from_agent match the expected target
sec-1: Propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to router
container (was only on openclaw container)
sec-2: Validate trusted_peers — reject values with control chars
sec-3: Validate registry_mode — only accept 'local'|'global'
rel-7: Wrap _runHandoffOrchestration in top-level try-catch
rel-6: Replace non-null assertions with explicit guard at completion
rel-3: Bump snapshot timeout 15s→60s, drain timeout 15s→30s
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: revise handoff flow diagrams with full review findings
Replace Section 11 with 7 comprehensive diagrams:
- 11.1: End-to-end sequence (both source + target sides, live progress)
- 11.2: Handoff state machine with known limitations noted
- 11.3: 7-layer security model (gate → isolation → auth → encryption →
injection hardening → identity → infrastructure)
- 11.4: Env var propagation showing both containers receive vars
- 11.5: Two orchestration paths (LLM vs CLI) and their differences
- 11.6: Trust flow (current unauthenticated vs future Entra OAuth)
- 11.7: Error recovery and planned improvements
Also add nohup.out to .gitignore (stale port-forward logs).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address remaining TS findings — orphan cleanup, guards, tests
rel-1: Clean up orphaned CRDs on abort when mesh transfer or discovery
fails after spawn (DELETE /sandbox/spawn/<name>)
rel-8: Concurrent handoff guard — reject confirm if handoff already running
sec-4: Respect $KUBECONFIG env var with fallback to ~/.kube/config
test-3: Fix 4 failing spawn error tests — tools handle unreachable router
gracefully (return status JSON), update assertions accordingly
(187/187 tests now pass)
dup-1: Document dual orchestration paths (CLI operator-mode vs plugin
LLM-mode) in handoff.ts header comment + architecture-diagrams.md
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address Rust findings — state machine, resume, tests
rel-2: Add try_transition() to enforce handoff phase ordering
(Idle→Init→Snapshot→Drain→Transfer→Restore→Verify→Decom→Complete)
rel-4: Add resume() + POST /agt/handoff/resume endpoint to cancel
drain state after abort (Aborted|Draining → Idle)
rel-5: Change body.unwrap() to expect() with message in spawn.rs
dead-1: Remove unused _used field from ActiveToken
test-1: Add 10 state machine transition tests (valid sequence,
invalid skip, abort, fail, resume, restart after complete)
test-2: Add 4 auth token tests (wrong value, no active, after
revoke, wrong pending confirmation code)
201 Rust tests pass (74 controller + 101 router + 26 integration)
Clippy clean with -D warnings
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* handoff: incremental progress polling, SDK reconnect fix, policy cleanup
- handoff_status tool: add since_step param for incremental polling,
returns only new_steps since last call so LLM relays one step at a time
- vendor SDK patch #9: AgentMeshClient.connect() no longer sets
connected=true when transport.connect() returns false, allowing retry
- Policy: remove handoff-tool-approval gate (two-step confirmation code
mechanism is sufficient; approval gate can return once native UI exists)
- config: add promoteMode to DeploymentContext for mesh promote tracking
All tests pass: 201 Rust (74+101+26), 187 CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(crd): add trustedPeers and registryMode to Helm CRD schema
The K8s API server was silently stripping these fields because the Helm
CRD template only defined enabled/toolPolicy/trustThreshold. spawn.rs
wrote the fields and reconciler.rs read them, but the schema validation
layer dropped them in between.
Root cause of handoff mesh registration failure: target pods never
received AGT_TRUSTED_PEERS or AGT_REGISTRY_MODE env vars because the
CRD never stored the values.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(router): read admin token from correct mount path
AppState::new() only checked /run/secrets/admin-token but the controller
mounts the secret at /etc/azureclaw/secrets/admin-token. This caused
'Server misconfiguration: no admin token' on target pods during handoff
verification. main.rs had the correct path but its token was only used
for the admin_auth_middleware, not the handoff middleware which reads
from state.admin_token.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): pre-build audit — snapshot strict, direction propagation, relay URLs
Three defensive fixes from comprehensive flow audit:
1. Snapshot endpoint now uses _routerCallStrict (was _routerCall) —
if snapshot fails, error surfaces immediately instead of continuing
with undefined blob data
2. Target-side handoff_transfer handler now reads direction from the
mesh message instead of hardcoding 'local_to_aks' — enables
reverse (aks_to_local) handoffs
3. Controller propagates AGT_RELAY_URL and AGT_REGISTRY_URL to the
openclaw container (was only on router container) — plugin no
longer relies on fallback to router proxy for relay connection
All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): enforce state machine — migrate all handlers to try_transition
All 5 handoff route handlers now use try_transition() instead of
set_phase(), returning 409 Conflict on invalid phase transitions.
Also fixed the transition rules: Decommissioning is now allowed from
Draining (source-side flow: Init→Snapshot→Drain→Decommission skips
Verify/Restore which happen on the target router, not the source).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): verification hash mismatch + synchronous progress
Two fixes:
1. Verification hash mismatch: verify endpoint was rebuilding a fresh
snapshot (new timestamp, nonce, hostname) instead of using the hash
from the restored data. Now restore stores the hash of the decrypted
compressed bytes, and verify reuses it. Falls back to build_snapshot
for source-side verify (where no restore happened).
2. No proactive progress: LLMs don't autonomously poll tools, so the
handoff_confirm tool now awaits _runHandoffOrchestration() and
returns all steps when complete, instead of firing-and-forgetting
and expecting the LLM to poll handoff_status.
All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(spawn): check Docker API HTTP status codes in docker_api
The docker_api helper only checked curl's exit code, not the HTTP
response status from Docker Engine. This caused silent failures —
e.g. container start returning HTTP 404 (network not found) was
swallowed and spawn reported success even though the container
never started.
Add -w flag to capture HTTP status code and return Err for 4xx/5xx
responses with the Docker error message extracted from the JSON body.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): propagate channel credentials to cloud target
During handoff spawn, collect channel/plugin credentials from the
source environment (TELEGRAM_BOT_TOKEN, SLACK_BOT_TOKEN, etc.) and
create a {name}-credentials K8s secret in the target namespace.
The controller already mounts this secret via envFrom (optional),
so the cloud agent inherits Telegram and other channels.
Also fix docker_api to check HTTP status codes — previously it only
checked curl's exit code, silently swallowing Docker Engine errors
like 'network not found' (HTTP 404).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): full state hydration — workspace, memory, conversations, Telegram
Source side (orchestration):
- Pack workspace tar from /sandbox/.openclaw/ (AGENTS.md, SOUL.md, etc.)
- Search Foundry Memory for recent context and include as chat_snapshot
- Include credential refs (channel/plugin names) in snapshot
Target router (restore):
- Extract workspace tar to /sandbox/ with path traversal protection
- Write chat_snapshot and metadata to /tmp/handoff/ for plugin
Target plugin (post-restore hydration):
- Create Foundry Conversation with replayed chat messages
- Store handoff event fact in Foundry Memory (update_memories)
- Write HANDOFF_CONTEXT.md to workspace (fallback context)
- Send 'handoff_ready' mesh message back to predecessor
The cloud agent now comes online with full context: workspace files,
conversation history in Foundry Conversations, semantic memory via
shared Memory Store, and proactively greets the user via Telegram.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: cross-container handoff — return state in response, not filesystem
The router and openclaw containers have separate filesystems in AKS.
Previously the router wrote workspace tar and chat snapshot to /tmp/
which the plugin couldn't read.
Changes:
- Router: return workspace_tar (base64) and chat_snapshot in restore
response JSON instead of writing to filesystem
- Plugin: extract workspace tar and parse chat snapshot from the
HTTP response body (runs in openclaw container where /sandbox/ lives)
- Remove unused extract_workspace_tar() from routes.rs
- Remove tar crate dependency (extraction now done by plugin via CLI tar)
- Add direction, initiated_at, restored_at to restore response
Future: workspace payloads >5MB will auto-transfer via Azure Blob
Storage (SAS URL in handoff state) — not yet implemented.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security: harden workspace tar extraction and chat snapshot parsing
Tar extraction:
- Pre-extract validation: list entries, reject any containing '..' or
starting with '/' (path traversal)
- Size guard: reject compressed payloads >5MB (decompression bomb)
- Unique temp dir per extraction (race condition prevention)
- --no-same-owner --no-overwrite-dir flags on extraction
- Temp dir cleaned up after extraction
- Removed '|| true' — errors now surface in logs
Chat snapshot:
- Schema validation: must be array, each entry must have string
role + content
- Cap at 100 messages, role capped at 20 chars, content at 10k chars
- Rejects non-conforming entries silently
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* perf: split sandbox Dockerfile into base + overlay for fast rebuilds
The sandbox image was 4.96 GB with ~4.0 GB of rarely-changing deps
(OpenClaw, Python wheels, Go tools, Node.js, CLI tools) rebuilt on
every code change. Now split into:
Dockerfile.base (~4.0 GB, rebuild weekly/on dep upgrade):
- Azure Linux 3 + system packages
- Node.js 22, Python 3 + 41 packages, Go CLI tools
- OpenClaw framework + extension symlinks + skills
- gh, ripgrep, 1password, himalaya
- User setup (sandbox:1000, router:1001)
Dockerfile (~50 MB, rebuild per commit in ~30s):
- FROM azureclaw-sandbox-base (all heavy deps pre-cached)
- CLI plugin builder reuses base image (has Node.js already)
- Router binary, plugin dist, vendored SDK overlay
- Entrypoint, proxy-bootstrap, skills, policies
All functionality preserved:
- UID separation, iptables egress guard, seccomp, read-only rootfs
- Channel plugins (Telegram/Slack/Discord/WhatsApp)
- Extension dep symlinks (grammy, carbon, bolt, etc.)
- Vendored SDK overlay, proxy-bootstrap, Control UI symlink
- ClawHub skills, npm CLI tools (clawhub, mcporter, oracle)
Build paths updated:
- azureclaw dev: auto-builds base if not cached, --build-base to force
- azureclaw push: --only sandbox-base to push base image
- Makefile: image-sandbox-base target added
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement reverse handoff (cloud → local)
CLI-driven reverse handoff orchestration:
- aksRouterExec: kubectl port-forward to AKS pod router
- wakeDormantDocker: detect and restart stopped containers
- readAksCrdSpec: inherit model/egress/isolation from CRD
- rehydrateCredentials: copy K8s secrets to Docker container
- Full 10-step reverse flow: connect → verify → init → snapshot →
drain → wake local → credentials → restore → succession →
decommission + delete CRD
Direction-aware source routing:
- sourceExec alias delegates to routerExec (Docker) or
aksRouterExec (AKS) based on direction
- Forward path unchanged at runtime (sourceExec === routerExec)
Plugin reverse handoff:
- handoff_request returns CLI command for aks_to_local direction
- Completion messages updated for both directions
- Decommission label direction-aware
Operator TUI:
- 'returning' handoff state for active aks_to_local handoff
- Table shows '<' icon and 'Returning' status
- ASCII-only table icons for reliable column alignment
- Column widths tightened
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(router): raise body limit on handoff routes to 50MB
Axum's default body limit is 2MB. Encrypted state snapshots easily
exceed this, causing HTTP 413 on /agt/handoff/snapshot and /restore.
Add DefaultBodyLimit::max(MAX_BLOB_SIZE_BYTES) layer to
handoff_protected_routes — matches the existing 50MB blob size
constant from §9.9.4.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): read AKS admin token from mounted secret
On AKS, the admin token is stored in K8s secret 'router-admin-token'
and mounted at /etc/azureclaw/secrets/admin-token — not as env var
or /tmp file. Updated getAksAdminToken() to:
1. Read from /etc/azureclaw/secrets/admin-token (router container)
2. Fallback: same path in openclaw container
3. Fallback: kubectl get secret (base64 decode)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): use POST for snapshot route (was GET → 405)
The /agt/handoff/snapshot route is POST-only but both forward and
reverse handoff paths were sending GET requests, causing HTTP 405.
Changed both to POST.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): reuse existing snapshot blob in reverse path
The reverse handoff was requesting a second snapshot at step 9, but
the state machine had already advanced to 'draining' after step 5.
The snapshot blob was already captured at step 3 — now the reverse
path uses snapshotResp.body.blob directly instead of re-fetching.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): pipe restore payload via stdin for large blobs
routerExec passes JSON as a curl -d argument, which hits shell
argument length limits for large encrypted snapshots. The reverse
handoff restore now uses 'docker exec -i ... curl -d @-' with the
payload piped via stdin, avoiding ARG_MAX issues.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(connect): handle Ctrl+C to disconnect port-forward
The kubectl port-forward child process with stdio:pipe did not
receive SIGINT from the terminal. Added explicit SIGINT/SIGTERM
handlers that terminate the child process and exit cleanly.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): init local handoff session + auth headers for restore
The local Docker router requires admin token + handoff token for
/agt/handoff/restore. The reverse path now:
1. Gets local admin token from Docker container
2. Inits a handoff session on the local router
3. Passes both auth headers to the restore curl call
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): correct step count + Telegram notification on handoff
- Fixed reverse handoff step counter: 13 steps (was 10, showing 11/10+)
- Added Telegram notification on handoff completion for both directions:
- local→cloud: 'I moved to the cloud'
- cloud→local: 'I am back on your local machine'
Best-effort — reads credentials from Docker container (reverse) or
env (forward), sends via Telegram Bot API. Failures are silent.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): clean up local handoff state after reverse restore
Two fixes for stale handoff sessions blocking subsequent handoffs:
1. CLI: After successful reverse restore, transition local router
through verify → decommission to reach a terminal state.
2. Router: Expand can_start() to allow re-init from Restoring,
Verifying, and Decommissioning phases. These indicate a previous
handoff that completed data transfer but wasn't properly finalized.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): extend verification timeout + retry mesh send
The verification timeout was 60s but AKS pods can take longer to
fully initialize their plugin message handlers after mesh registration.
The blob sent before the handler is ready gets silently dropped.
Fix:
- Extended timeout from 60s to 180s
- Re-send the handoff_transfer blob every 30s within the verification
loop, in case the target's message handler wasn't ready on first send
- Also fixes: router can_start() allows stale Restoring/Verifying states,
local handoff session cleaned up after reverse restore
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(relay): increase max_message_size to 1MB for handoff blobs
Handoff snapshots with real state (chat, audit, credentials) can be
80+ KB. After Signal Protocol encryption + base64 + JSON envelope,
they exceed the relay's 64KB default max_message_size. Bumped to 1MB.
Also includes verification timeout extension (60s→180s) with 30s
re-send retries, already committed in plugin.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: collect workspace/chat/credentials in CLI handoff snapshot
The CLI handoff command was sending an empty snapshot payload
(only shared_secret). The forward handoff via plugin.ts collected
workspace tar, Foundry memories, and credential refs — but the
CLI path (used for both forward and reverse) skipped this.
Changes:
- Collect workspace tar via kubectl exec (AKS) or docker exec (local)
- Collect Foundry Memory Store items as chat context
- Collect credential refs from container environment
- Fix snapshot response field name: size_bytes → snapshot_size_bytes
- Include items breakdown in snapshot response
- Fix step counter: move transfer step into forward branch only
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: stop stepper spinner and cleanup port-forward after handoff
stepper.step('Handoff summary...') started a spinner that was never
stopped with stepper.done(), keeping the event loop alive and requiring
Ctrl+C. Also aksPortForwardStop() was only called in the error path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: add /agt/handoff/succession router endpoint for Ed25519-signed succession
The registry's succession API requires the predecessor's Ed25519 signature
over a canonical message. The private key lives in the router's Governance
identity — inaccessible to the CLI.
New endpoint POST /agt/handoff/succession on the router:
- Takes {successor_amid, reason} from CLI
- Looks up predecessor (self) AMID from registry
- Looks up successor signing key from registry
- Signs canonical message 'succession:{pred}:{succ}:{timestamp}'
- Submits complete SuccessionRequest to registry
- Returns registry response
CLI + plugin updated to call /agt/handoff/succession instead of
/agt/registry/registry/succession (which lacked signing keys).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: sub-agent handoff — collect, snapshot, and re-spawn during handoff
Sub-agents spawned by a parent agent are now included in handoff state
transfer. The full pipeline:
Collection (source side):
- New GET /agt/handoff/sub-agents endpoint lists active sub-agents
and reconstructs SpawnRequest from CRD spec (K8s) or container
labels (Docker dev mode)
- CLI + plugin call this endpoint and inject sub_agent_snapshots
into the handoff snapshot payload
Re-spawn (target side):
- handoff_restore iterates sub_agent_snapshots after state hydration
- Calls create_sandbox() for each sub-agent with the stored config
- Returns per-sub-agent results (spawned/failed) in restore response
- Audit-logged as handoff:restore:sub-agent
Supporting changes:
- SpawnRequest, HandoffMeta, SubAgentSnapshot: added Clone derive
- Sub-agent results included in restore response as sub_agent_results
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: AMID remapping + sub-agent workspace collection during handoff
Two improvements to sub-agent handoff:
1. AMID remapping: When re-spawning sub-agents on the target, the old
parent AMID in trusted_peers is replaced with the new parent's AMID.
This ensures sub-agents trust the new parent for KNOCK handshakes.
The new parent AMID is looked up from the registry at restore time.
2. Workspace collection: The CLI now exec's into each sub-agent's
container (kubectl for AKS, docker for local) to collect workspace
tar before including it in the snapshot. Each sub-agent's workspace
is capped at 2MB. The plugin path is best-effort without workspace
(no container exec access from inside the sandbox).
Also adds sub_agents_respawned count to plugin restore metadata.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: collect sub-agent workspace via E2E mesh during handoff
Sub-agents now respond to handoff:workspace_request mesh messages by
tarring their /sandbox/.openclaw/workspace and sending it back via the
E2E encrypted relay. No shared volumes or cross-container exec needed.
Plugin message handler (sub-agent side):
- Receives handoff:workspace_request from parent
- Tars workspace (excludes extensions, node_modules, etc.)
- Sends base64-encoded tar back as handoff:workspace_response
- Falls back to empty response on error (parent doesn't hang)
Plugin handoff orchestration (parent side):
- After fetching sub-agent list from router, discovers each sub-agent
via registry search to get their AMID
- Sends handoff:workspace_request to each via mesh
- Polls agtInbox for handoff:workspace_response (up to 15s per agent)
- Enriches sub-agent snapshots with workspace tar before creating
the encrypted handoff snapshot
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: expand workspace tar to include cron/policies/agents + mesh 768KB cap
- Add WORKSPACE_TAR_CMD constant in handoff.ts for consistent tar commands
- Include .openclaw/cron, .openclaw/policies, .openclaw/agents in all 5 tar commands
- Refine extension exclusion: only exclude */dist and */node_modules (keep manifests)
- Cap mesh workspace response at 768KB (safe under relay's ~1MB limit)
- Add truncated flag in workspace_response messages
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: chunked mesh transfer for large sub-agent workspaces
Workspaces > 512KB are split into chunks sent via separate mesh messages,
then reassembled on the receiver side. This lifts the practical workspace
transfer limit from ~768KB to ~40MB (80 chunks × 512KB).
Sender (sub-agent):
- Small workspace (≤512KB): single handoff:workspace_response (fast path)
- Large workspace: N handoff:workspace_chunk messages + completion marker
- Max 80 chunks (leaves headroom in relay's 100-message offline queue)
Receiver (parent):
- Collects workspace_chunk messages into a Map keyed by chunk_index
- Reassembles in order when all chunks received or completion marker arrives
- 30s timeout with partial-chunk recovery (uses what was received)
- Backwards compatible: single-message responses still work unchanged
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: unified chunked mesh transport layer + file transfer tool
Implements a general-purpose auto-chunking transport layer for the E2E
encrypted mesh. Any payload exceeding 512KB is transparently split into
chunks with per-chunk SHA-256 integrity verification, then reassembled
on the receiver side before reaching application logic.
Transport layer (meshSend + meshHandleTransportMessage):
- meshSend(): auto-chunks large payloads into mesh:transfer_manifest +
N mesh:transfer_chunk messages. Small messages pass through directly.
- meshHandleTransportMessage(): intercepts transport messages in onMessage,
accumulates chunks, verifies SHA-256 hashes, reassembles, and delivers
the original message to the application layer.
- Per-chunk + manifest-level SHA-256 integrity verification
- 2-minute TTL with automatic stale transfer cleanup
- Max ~40MB per transfer (80 chunks × 512KB)
New tool — azureclaw_mesh_transfer_file:
- Agents can send files to each other via E2E encrypted mesh
- Files up to 30MB supported (auto-chunked transparently)
- Received files auto-saved to /sandbox/.openclaw/workspace/incoming/
- Path traversal protection (must be within /sandbox)
Consumers updated to use unified transport:
- mesh_send tool: auto-chunks large task messages
- Handoff blob transfer: auto-chunks encrypted snapshots
- Sub-agent workspace collection: auto-chunks workspace tars
- Handoff re-send loop: uses meshSend for retransmit
Limits raised:
- Router MAX_BLOB_SIZE_BYTES: 50MB → 200MB (sub-agent workspaces)
- Plugin workspace tar cap: 5MB → 50MB
- CLI workspace tar cap: 5MB → 50MB
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: configurable router URL + fix spawn test timeouts + add transport tests
- Make ROUTER and ROUTER_BASE configurable via AZURECLAW_ROUTER_URL env var
(defaults to http://127.0.0.1:8443 for backward compat)
- Fix 5 pre-existing spawn tool test timeouts caused by local dev router
on port 8443 — tests now point to unused port 19876 for immediate
ECONNREFUSED instead of 45s polling loop
- Fix test assertions to match actual error behavior (plain text errors,
not JSON, when router is unreachable)
- Add 11 new tests (198 total, up from 187):
- mesh_transfer_file: registration, schema, path traversal, abs path,
mesh-not-connected
- mesh_send: registration, params, error when disconnected
- handoff_status: registration, returns status JSON
- AZURECLAW_ROUTER_URL: spawn + spawn_status use configurable URL
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: address 4 direction-specific handoff gaps
1. Plugin reverse: retry with 60s backoff when discovering local target
(local agent may be waking from dormant state after CLI runs
wakeDormantDocker)
2. Direction validation: both plugin and router now validate that the
incoming handoff direction matches the environment (AZURECLAW_DEV_MODE).
Warn-only on mismatch — doesn't block to avoid false positives.
3. Forward credential rehydration: collect actual credential VALUES
from Docker (not just refs), create K8s secret on AKS target before
pod starts so envFrom can mount them. Closes the credential gap
where forward handoff lost Telegram/Slack/Brave tokens.
4. Symmetric cleanup: reverse handoff now scales deployment to 0
instead of deleting the CRD. This preserves the sandbox definition
for instant re-forward handoff while freeing all compute resources.
Falls back to CRD deletion if scale fails.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: graceful sub-agent interrupt during handoff
Add handoff:interrupt protocol so sub-agents can save in-progress work
before their workspace is collected during a handoff.
Plugin path (mesh-based):
- Parent sends handoff:interrupt to all sub-agents concurrently
- Sub-agents set handoffInterruptRequested flag
- processTaskWithTools checks the flag between LLM rounds
- On interrupt: saves task progress to .task-in-progress.json
(round, messages, last content, original task)
- Sends handoff:interrupt_ack back to parent
- Parent waits up to 10s for acks, then proceeds with workspace collection
CLI path (exec-based):
- CLI writes .handoff-interrupt sentinel file into each sub-agent container
- processTaskWithTools also checks for this file between rounds
- Same progress save behavior (.task-in-progress.json)
Both paths ensure the workspace tar includes the progress checkpoint,
so it survives the handoff and is available on the target for resumption.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: sub-agent workspace injection + task resumption after handoff
Complete the sub-agent handoff lifecycle — after re-spawn on the target,
sub-agents now receive their workspace and resume interrupted work.
Router changes:
- Restore response now includes sub_agent_workspaces array with each
sub-agent's workspace_tar, task_context, status, and checkpoint
Plugin — target side (post-restore):
- Waits up to 60s for each re-spawned sub-agent to register in mesh
- Sends workspace tar via meshSend (auto-chunked for large workspaces)
- Sends handoff:resume with task context and checkpoint info
Plugin — sub-agent side (new message handlers):
- handoff:workspace_inject: extracts received workspace tar into /sandbox/
with path traversal validation and size guard
- handoff:resume: reads .task-in-progress.json, sends resume_ack to parent
with status report ('Successfully restored in cloud. Resuming interrupted
work from round N: <task>'), then re-enters processTaskWithTools with a
contextual prompt that includes the original task, progress, and last output
The full sub-agent handoff lifecycle is now:
interrupt → save progress → collect workspace → transfer →
re-spawn → inject workspace → resume task → report to parent
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent handoff — Docker API encoding + name prefix + source cleanup
Three bugs prevented sub-agents from properly migrating during handoff:
1. collect_sub_agent_snapshots_docker passed raw JSON braces in the
Docker API URL — curl treated {} as glob patterns and the Docker
daemon couldn't parse the filter. Result: always returned 0
sub-agents, so the snapshot blob had no sub-agents to respawn.
Now uses URL-encoded filter via the docker_api() helper (matching
list_sandboxes_docker's pattern).
2. Same function used the Docker container name (azureclaw-{name})
as the agent name, causing respawn to create
azureclaw-azureclaw-{name} on the target. Now strips the prefix.
3. Source decommission only put the main agent dormant — sub-agent
containers kept running. Now destroys all source sub-agents via
the spawn API before decommissioning.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent handoff — wrong API URLs + missing report steps
Four fixes for sub-agent handoff orchestration:
1. Sub-agent list used GET /sandbox/spawn (wrong) — now GET /sandbox/list
2. Sub-agent delete used DELETE /sandbox/spawn/{name} (wrong) — now
DELETE /sandbox/{name} (matches actual router routes)
3. Response field was 'sub_agents' but endpoint returns 'sandboxes'
4. No sub-agent info in handoff progress report — added _hp() calls
for: discovery count, interrupt/checkpoint status, workspace
collection count, snapshot inclusion, cleanup status, and
sub_agents_transferred in the final result object
Also fixed orphan target cleanup URLs (2 places) that had the same
/sandbox/spawn/{name} → /sandbox/{name} issue.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: set trusted_peers for re-spawned sub-agents after handoff
The trusted_peers remapping in handoff_restore was dead code —
Docker snapshots had trusted_peers=None, so the if-let guard never
fired. Re-spawned sub-agents had no trusted parent AMID and rejected
handoff:workspace_inject + handoff:resume messages from the new
parent. This caused workspace injection to silently fail and task
resumption to never trigger.
Fix: always set trusted_peers to include the new parent's AMID,
regardless of whether the original snapshot had it set:
- If peers existed: remap old parent → new parent (existing logic)
- If peers existed but old parent absent: append new parent
- If peers was None: set to new parent entry (new case)
This ensures the sub-agent trusts the new parent on first KNOCK
and accepts workspace/resume messages immediately after spawn.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: include sub-agent status in Telegram greeting after handoff
Restructure the restore flow so sub-agent workspace injection and resume
happens BEFORE the Telegram greeting. After sending resume signals, wait
up to 8s for sub-agents to send resume_ack messages. The Telegram greeting
now includes a sub-agent status section showing each agent's name, state
(resumed/ready/starting/failed), and a task preview.
The handoff_ready mesh message back to the predecessor also now includes
sub_agents_restored count, sub_agents_resumed count, and per-agent details.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: use per-sandbox runtime for operator exec path, not global devMode
The operator's unified view fetches both Docker and AKS agents, but the
routerExec/fetchAgtQuick/fetchEgressDomains functions used the global
devMode flag to choose between docker-exec and kubectl-exec. This meant
Docker agents were queried via kubectl (which fails — no K8s namespace)
when the operator ran in non-dev mode, and vice versa.
Fix: use sb.runtime === 'docker' per-sandbox instead of the global devMode
flag in all 4 places that exec into agent containers:
- fetchSecurityState (routerExec + k8sCheck)
- fetchEgressDomains (routerCurl)
- fetchAgtQuick (docker/kubectl exec)
- seccomp profile inference
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent trust + workspace logging after handoff
Two fixes for broken parent↔sub-agent communication after handoff:
1. Stale trusted_peers: After handoff, re-spawned sub-agents get new
AMIDs but the parent's parentTrustedAmids set contains old AMIDs.
KNOCK handler rejects all incoming sub-agent messages (score=0 <
threshold=500). Fix: when the plugin discovers a sub-agent's new
AMID during workspace injection, register it in amidToName,
nameToAmid, parentTrustedAmids, and push baseline trust to router.
2. Silent from_value failure: serde_json::from_value for sub-agent
snapshots at POST /handoff/snapshot was wrapped in `if let Ok`
which silently swallowed deserialization errors, potentially losing
all sub-agent workspace data. Changed to match with tracing::warn
that logs the exact error + JSON preview for debugging.
Also adds roundtrip test for SubAgentSnapshot workspace_tar through
the full serialize→compress→encrypt→decrypt→decompress→deserialize
chain, and a test for the JS↔Rust base64 round-trip.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security: remove clawhub.com and openclaw.ai from default egress allowlist
clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill).
openclaw.ai enables `curl | bash` install vectors from inside sandboxes.
Both were in the default Helm values and example CRD. Removed from:
- deploy/helm/azureclaw/values.yaml (default egress allowlist)
- examples/basic-agent/clawsandbox.yaml (example CRD)
- PLAN.md (policy presets documentation)
Defense is now three layers deep:
1. Egress proxy blocks clawhub.com/openclaw.ai (network)
2. Skills directory is root-owned, chmod 640/750 (filesystem)
3. Plugin code is root-owned, read-only for sandbox (code integrity)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: use sub_agent_results as trust+resume loop driver, not sub_agent_workspaces
Root cause: the post-restore trust registration and resume signals were
gated on restoreResp.sub_agent_workspaces (workspace data), which could
be empty even when sub-agents were successfully spawned. This caused the
entire block to be skipped — no trust registration, no resume signals,
no Telegram sub-agent status.
Fix: use restoreResp.sub_agent_results (always populated when sub-agents
spawn) as the primary loop driver. Workspace data is looked up by name
from sub_agent_workspaces as a secondary source.
Also:
- Added workspace_inject_ack: sub-agent confirms extraction success/fail
with file_count + error back to parent before resume is sent
- Parent waits up to 15s for ack, logs result, passes workspace_delivered
flag in resume payload
- handoff_ready report now includes sub_agents_workspace_delivered count
- Telegram greeting shows 📦 icon per sub-agent when workspace arrived
- Increased mesh registration wait from 60s to 90s (AKS pods need boot)
- Router logs snapshot details when building sub_agent_workspaces
Tests added:
- Rust: sub_agent_workspaces builder filter (empty/non-empty workspace_tar)
- Rust: full encrypt→decrypt→restore round-trip with 2 sub-agents
- Rust: edge case — empty workspace + empty task_context filtered out
- TS: sub_agent_results drives loop even when sub_agent_workspaces empty
- TS: workspace_inject_ack protocol (success + failure paths)
- TS: handoff_ready includes workspace delivery status
- TS: only spawned sub-agents enter trust loop
- TS: missing sub_agent_results graceful fallback
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(sdk): reuse active Signal Protocol session instead of crashing
Vendor patch #10: SessionManager.initiateSession() threw 'Active session
already exists' when the crypto layer had a session established via
incoming KNOCK but client.activeSessions wasn't synced.
This broke mesh_transfer_file and any second send to the same peer.
Fix: return existing session info with reused=true flag instead of
throwing. establishSession detects reuse, syncs activeSessions, and
skips redundant KNOCK/activate.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: don't expose handoff confirmation code to LLM, add console.log diagnostics
Security fix: The handoff confirmation token was returned in the tool
result, allowing the LLM to self-confirm without user input. Now the
code is sent directly to Telegram (side-channel) and printed to console
for TUI users — the LLM never sees it.
Changes:
- Remove confirmation_token from azureclaw_handoff_request tool response
- Send code via Telegram sendMessage as a side-channel delivery
- Print code to console.log (visible in kubectl logs, not to LLM)
- Update tool descriptions to emphasize code comes from user input
- Bump CONFIRMATION_MIN_DELAY_SECS from 3s to 8s
- Add console.log diagnostics in post-restore IIFE for handoff debugging
- Add test: tool response must not contain confirmation_token
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: increase sub-agent tool-calling rounds from 10 to 25
10 rounds was too tight for data-heavy tasks — sub-agents hit the
cap and returned truncated results. 25 gives enough room for
multi-step research/collection while still preventing runaway loops.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: increase prekey retry window from 16s to 45s for sub-agent mesh send
Sub-agents take 20-30s+ after pod is Running to upload prekeys
(gateway start → plugin load → SDK init → relay connect → prekey
upload). The previous 8×2s=16s window wasn't enough, causing
'Cannot get prekeys' failures on task dispatch.
Now: 15 attempts × 3s = 45s max wait with clearer hint message.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: add file_transfer_ack for verified file delivery
The mesh_transfer_file tool had no delivery confirmation — it reported
'sent' but couldn't verify the file was actually written to disk on the
target agent. This caused silent failures where coordinator-notes.txt
appeared to transfer but never landed.
Now:
- Receiver sends file_transfer_ack with success/saved_to/error
- Sender waits up to 15s for ack
- Tool returns 'delivered' (with path) or 'sent_no_ack' (no confirmation)
- Write is verified with stat after writeFileSync
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: filter protocol messages from mesh_inbox, improve file transfer reliability
mesh_inbox now filters out internal protocol messages (handoff blobs,
acks, workspace inject/resume) so the LLM only sees actual sub-agent
replies. Shows filtered_protocol_messages count for visibility.
file_transfer now retries up to 3 times with ack verification — sends,
waits 15s for file_transfer_ack, retries with 3s backoff if no ack.
Returns 'delivered' with exact path or 'sent_no_ack' after all retries.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: auto-decode file_transfer content in mesh_inbox
file_transfer messages now show decoded text content (or a binary
placeholder) instead of raw base64 blobs. Uses null-byte detection
to distinguish text vs binary files. Text files are fully readable
in the inbox; binary files show filename and size.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: stale AMID cache poisoning breaks post-handoff mesh delivery
Root cause: trust+resume loop found OLD sub-agent AMIDs still in the
registry (Docker containers hadn't timed out yet) and cached them.
New AKS sub-agents registered with different AMIDs 26s later, but
the parent never discovered them — all messages went to dead relay
connections and were silently dropped.
Three-layer fix:
1. Stale AMID rejection: collect original_amid from handoff snapshots,
reject matching registry results, wait for NEW AMIDs to appear
2. Prekey readiness gate: verify E2E session is established before
sending workspace_inject (20 attempts × 3s = 60s max)
3. Workspace inject retry: 3 attempts with 20s ack wait each, catches
send errors and retries instead of fire-and-forget
Also filters protocol messages from mesh_inbox and auto-decodes
file_transfer base64 content so LLM sees readable text.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: write HANDOFF_FILES.md manifest after workspace inject
After extracting the workspace tar, writes a manifest listing restored
user-facing files to /sandbox/.openclaw/workspace/HANDOFF_FILES.md.
This makes injected files discoverable when the agent is asked about
its workspace contents post-handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: bump AGT rate limits for multi-agent handoff
Policy: 120 → 240 max_calls/60s for inference:* actions.
Router: 100 → 200 global req/s, 10 → 20 per-agent req/s.
Handoff with 3+ agents doing workspace inject + resume + relay
traffic was hitting the old limits.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: promote incoming/ files to workspace root after handoff inject
Copies files from incoming/ to the workspace root so the agent sees
them immediately when listing files, without needing to know about
the incoming/ directory convention.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: scale down sub-agent deployments during cloud→local decommission
After scaling the parent to 0, also scale down all sub-agent
deployments that were snapshotted. Prevents orphaned sub-agent
pods running on AKS after reverse handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: add bidirectional handoff architecture diagrams and changelog
- New section 12 in architecture-diagrams.md: agent lifecycle across
handoff, forward/reverse flows with sub-agents, stale AMID cache
poisoning problem & fix, workspace injection detail
- CHANGELOG: add handoff features, sub-agent support, rate limit bump
- README: add handoff to features list and CLI reference table
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced Apr 24, 2026
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Apr 27, 2026
…compile + helm CRD (S4) (#54) Phase 2 §8 entry 4. Ships the K8s primitive only — `InferencePolicy` is NOT a model-router (per §3 non-compete; model selection sits in Foundry). Sandbox-side budget / guardrail / safety policy CR, compiled to a JSON ConfigMap that the S7 router-side informer will load into the existing PolicyEnvelope. Per user direction 2026-04-27, runtime enforcement substrate stays on Phase 1: `inference-router::budget::TokenBudgetTracker` (env-fed) for tokens, Foundry Content Safety + `safety::report_content_flags_to_agt` → AGT BehaviorMonitor for safety. AGT-Rust 3.3.0 verified against `/Users/pallakatos/Private/Repos/agt/agent-governance-toolkit` — AGT-Python has BudgetTracker, AGT-Rust does not yet; the upstream port is an S7 decision and is explicitly out of scope here. Added: - controller/src/inference_policy.rs — CRD struct + spec sub-types (TokenBudget, ContentSafetyFloor, ModelPreference, ModelRef) + status reusing mcp_server::LocalObjectRef (4th semantic client). - controller/src/inference_policy_compile.rs — pure-fn compile_to_profile + version_hash, deterministic, key-canonical; output shape slots into PolicyEntry.payload, no parallel hot-reload. - controller/src/inference_policy_reconciler.rs — modeled on S3 a2a_agent_reconciler. Field manager azureclaw-controller/inferencepolicy (distinct per §10.4 #1), finalizer azureclaw.azure.com/inferencepolicy-cleanup. Conditions reuse status::conditions; closed-set error_class per §15.3. - 6 CEL admission rules in crd_validations.rs: monthlyTokens >= dailyTokens, monthlyTokens >= perRequestTokens, contentSafety severity ∈ {Safe,Low,Medium,High}, modelPreference primary/fallback non-empty provider+deployment, appliesTo.action ∈ {chat,responses,image,embeddings,*}. - deploy/helm/azureclaw/templates/crd-inferencepolicy.yaml — drift- checked by helm_inferencepolicy_crd_matches_rust_schema. - docs/security-audits/2026-04-27-phase2-inferencepolicy-reconciler.md — AGT boundary verification, STRIDE, out-of-scope list, two sign-offs. Tests: +20 (6 compile + 7 reconciler + 5 admission + 2 helm-drift). Controller suite 193 → 218. Workspace cargo test/fmt/clippy all green. §14.6: strengthens column 7 (Foundry / M365 integration) — primitive lands here; runtime consumers wired in S7. AGT crate pin unchanged: agentmesh = "3.3.0" from crates.io, no fork. `vendor/` directory untouched. Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Apr 28, 2026
) * phase2(s10.a1): introduce spec.runtime discriminated union (CRD + Helm only) S10.A1 step 1 of N — CRD schema spine for multi-runtime hosting. Replaces the legacy `spec.openclaw` field with a discriminated union `spec.runtime { kind, openclaw, openaiAgents, microsoftAgentFramework, byo }`. The `kind` discriminator selects which sibling struct is required; the others must be absent. Mutual exclusion enforced at admission via Helm CRD CEL `x-kubernetes-validations` (4 bidirectional rules `(self.kind == 'X') == has(self.x)`); controller-side defensive guard will land alongside the reconciler dispatch in a follow-up commit. Pre-release simplification: in-place v1alpha1 schema edit. No v1alpha2 cut, no conversion webhook (no installed base, per plan.md S10 + S13). One PR = one breaking change. What this commit delivers ------------------------- - `controller/src/crd.rs`: new `RuntimeSpec`, `RuntimeKind` enum, `OpenAIAgentsConfig`, `MicrosoftAgentFrameworkConfig`, `MafLanguage`, `AgentCodeRef { oci, git }`, `OciAgentCode`, `GitAgentCode`, `ByoRuntimeConfig` (with `contractVersion` REQUIRED — no silent default per rubber-duck #9). `ClawSandboxSpec.runtime` is required on the wire; `Default` retained for test ergonomics, returns `OpenClaw` with empty config. - `controller/src/crd.rs`: `ClawSandboxStatus.runtime_kind` Option field (`#[serde(skip_serializing_if = "Option::is_none")]` to avoid wiping a populated value via merge patch). - `controller/src/crd.rs`: 8 new tests — PascalCase wire-format guarantees for all 4 `RuntimeKind` variants, default-is-OpenClaw, per-variant round-trip, BYO contractVersion required-not-default, serializer omits absent variants, runtimeKind status absence. - `deploy/helm/azureclaw/templates/crd.yaml`: `spec.required` flips from `["openclaw", ...]` to `["runtime", ...]`. New `runtime` block with `kind` enum + 4 sibling structs + 4 CEL rules. Inner CEL on `agentCode` enforces `has(oci) != has(git)`. Status gains `runtimeKind`. `Runtime` printer column added. - `controller/src/reconciler/mod.rs`: minimal call-site fix — `spec.openclaw` → `spec.runtime.openclaw.clone()` to keep the build green. Full dispatch refactor (`RuntimeDeploymentPlan` per rubber-duck #2/#3) lands in step 2. What is NOT yet wired (intentional, follow-ups) ----------------------------------------------- - Reconciler dispatch per `runtime.kind` (single-seam plan struct). - `RuntimeReady` Condition machinery (folded into `build_running_status_patch` + `running_status_matches` per rubber-duck #1 to avoid status-merge churn). - OpenAI Agents / MAF deployment SKIP (must NOT silently use `ctx.sandbox_image` per rubber-duck #2; will stamp Degraded + AdapterMissing). - `validate_runtime_shape` controller-side guard. - Examples / fixtures / CLI templates / convert / from_kagent migration. - CHANGELOG.md, audit doc. Verification ------------ - `cargo test --package azureclaw-controller`: 284/284 pass (8 new RuntimeSpec tests included). - `cargo clippy --package azureclaw-controller --all-targets -- -D warnings`: clean. - `cargo fmt --all`: applied. - Helm CRD YAML: parses; verified via `yq` — 4 CEL rules on runtime block, byo.required = [image, contractVersion], runtimeKind status field present, Runtime printer column added. NOTE: This commit alone is NOT mergeable on its own. Without the fixture/CLI/example migrations + reconciler dispatch, every existing `spec.openclaw` manifest in-tree would fail admission. Branch `phase2-multi-runtime-crd` will accumulate the remaining steps before the PR opens. Refs: plan.md S10.A1; rubber-duck critique applied (status merge risk #1, OpenAI/MAF fall-through #2, single dispatch seam #3, CEL shape #6, contractVersion required #9, container name stays 'openclaw' #4). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: migrate spec.openclaw → spec.runtime.openclaw across emitters and fixtures Completes the in-place v1alpha1 schema migration for the multi-runtime CRD spine. CRD types + Helm schema + reconciler reader landed in d11d41d; this finishes the long tail of CRD-emitting / CRD-reading sites the user called out ("every aspect of the code — including the cloud offload"). Cloud offload (the user's explicit ask): - controller/src/mesh_peer/offload.rs: build offload ClawSandbox CRD with spec.runtime.{kind: OpenClaw, openclaw: {...}} shape; mutate spec.runtime.openclaw for OFFLOAD_* env injection (was spec.openclaw). - controller/src/reconciler/mod.rs:796: comment text aligned to new path. CLI emitters: - add.ts, up.ts: emit spec.runtime.{kind: OpenClaw, openclaw}. - convert.ts: ClawSandbox→upstream reads spec.runtime.openclaw and hard-fails with a clear error if runtime.kind != OpenClaw (no upstream Sandbox shape for non-OpenClaw runtimes); upstream→ClawSandbox emits the new shape. - migrate.ts: --image help text references spec.runtime.openclaw.image. - migrate/from_kagent.ts: emits spec.runtime.{kind: OpenClaw, openclaw}; warning message aligned. - handoff.ts: model inheritance reads spec.runtime.openclaw.config.agent.model. Fixtures + examples (8 yaml files): - examples/{basic,confidential,telegram}-agent/clawsandbox.yaml - examples/demo-clawshield/{fabrikam-legal,contoso-bank,northwind-trade}-agent.yaml - tests/compat/fixtures/null-provider-{prod-denied,devonly-ok}.yaml (note: pre-existing 'sandbox.isolation: strict' enum issue on the prod-denied fixture left unchanged — orthogonal to this migration; static scanner is the active enforcement, not CRD validation.) Tests updated to assert the new shape: - add.test.ts: 4 assertions - convert.test.ts: 6 assertion blocks + 1 multi-container test - from_kagent.test.ts: 4 assertions Verification: - cargo test --package azureclaw-controller: 284/284 pass - cargo clippy --package azureclaw-controller --all-targets -- -D warnings: clean - cli npm test: 435/435 pass + 2 skipped - cli npm run typecheck: clean - grep confirms no remaining spec.openclaw emission/read sites; only intentional docstring/comment references documenting the legacy shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: runtime-aware status surface + AdapterMissing dispatch guard Closes the S10.A1 spine of phase2-multi-runtime-crd. Builds on the prior two commits (d11d41d CRD spine; 3202d13 emitter migration) by wiring the runtime kind through the status surface and refusing to deploy a Pod for runtime kinds whose adapter has not yet shipped. Status surface - New `TYPE_RUNTIME_READY` Condition + `reason::ADAPTER_MISSING` in controller/src/status/conditions.rs. - `build_running_status_patch` / `running_status_matches` / `build_overlay_status_patch` / `overlay_status_matches` take `runtime_kind: &str` trailing arg; emit `status.runtimeKind` and append `RuntimeReady` to the conditions array (True/Reconciled on the running path, False/OverlayMode on overlay). Stamping inside the existing patch (rather than via a separate patch_status) avoids the merge-patch array overwrite that would erase the new Condition and re-introduce the resourceVersion-bump reconcile storm — see plan S10.A1 rubber-duck #1. - New `build_runtime_unsupported_status_patch` / `runtime_unsupported_status_matches` / `stamp_runtime_unsupported` helper trio mirrors the existing degraded_* trio. Stamps Degraded=True + Ready=False + RuntimeReady=False, all Reason=AdapterMissing. Reconciler dispatch - controller/src/reconciler/mod.rs:222-260 maps RuntimeKind to a static-str discriminator and explicitly skips namespace/SA/Deployment creation when the kind is not OpenClaw. Stamps AdapterMissing and returns Action::requeue(300s) BEFORE any K8s-resource builder is invoked — no silent fall-through to ctx.sandbox_image (plan S10.A1 rubber-duck #2). - Status-patch call sites at :1481-1498 thread the runtime_kind_str. Tests - 5 new tests for the AdapterMissing helper trio (stamp shape, status-missing/runtime-mismatch idempotency rejection, settled-status match, transition-time preservation across repeat patches). - 4 existing tests updated for new conditions array shape + runtimeKind field (Running: 2 conds, Overlay: 4 conds). - 289/289 controller tests pass (was 284); cargo clippy clean; CLI 435/435 still green. Docs - CHANGELOG.md: S10.A1 entry under Unreleased Phase 2 with breaking- change marker spanning the three commits. - docs/security-audits/2026-04-28-phase2-multi-runtime-crd.md: full audit doc with threat model (silent fallthrough, status churn, CEL-disabled, BYO contract bypass, convert hard-fail), existing- implementation survey, wire-format invariants, test matrix, and S10.A2-A5 deferral list. Deferred to S10.A2: RuntimeDeploymentPlan per-variant dispatch seam, per-variant image/entrypoint/env/agentCode resolution, BYO contract verifier, validate_runtime_shape defensive guard. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * S10.A1: scaffold Tier-2 runtime placeholders (SemanticKernel, LangGraph, Anthropic) Locks the CRD wire shape now for three additional declared-roadmap runtimes so adding their adapters in a later slice is not a breaking schema change. The CRD becomes a public roadmap signal: customers can pin `spec.runtime.kind` today and know the schema won't shift under them. Tiering: Tier 1 (Phase 2 adapters): OpenClaw, OpenAIAgents (S10.A3), MicrosoftAgentFramework (S10.A4) Tier 2 (placeholder, Phase 3+ adapters): SemanticKernel, LangGraph, Anthropic BYO (warn-only contract verifier in S10.A2) Schema changes - crd.rs: `RuntimeKind` gains 3 PascalCase variants. New config structs `SemanticKernelConfig` (language: python|dotnet|java), `LangGraphConfig` (language: python|typescript), `AnthropicConfig` (pythonVersion). All three carry the universal agentCode + entrypoint + extraEnv shape — same as OpenAIAgents/MAF. - helm crd.yaml: 3 new bidirectional CEL rules; 3 new schema property blocks; kind enum extended in both spec + status surfaces; nested AgentCodeRef exactly-one CEL on every variant that carries code. - reconciler: runtime_kind_str match extended; AdapterMissing message enumerates Tier-2 placeholders. Behavior - All three Tier-2 kinds short-circuit through the existing `stamp_runtime_unsupported` path: Degraded + Ready=False + RuntimeReady=False / AdapterMissing, requeue 300s. Zero new code paths; pure schema scaffold. Tests: 4 new round-trip tests (one per variant + a defaults check that SkLanguage and LangGraphLanguage default to python). 293/293 controller tests pass (was 289). Clippy clean. Docs: CHANGELOG + audit doc 2026-04-28-phase2-multi-runtime-crd.md updated to enumerate Tier-2 placeholders and reflect the 7-rule CEL matrix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Apr 28, 2026
… sub-slice) (#72) §10.4 #1 ("Server-Side Apply on every emitted object with stable field managers") landing in sub-slices. This is the first: central field-manager registry + replacement of every bare-string SSA site. - New `controller/src/field_managers.rs` — single source of truth. CLAWSANDBOX, PAIRING, MESH_PEER, MCP_SERVER, TOOL_POLICY, A2A_AGENT, INFERENCE_POLICY, CLAW_MEMORY, CLAW_EVAL constants + ROUTER_RECONCILER, PROVIDER_BRIDGE, MESH, RECONCILER. ALL_FIELD_MANAGERS registry + 4 invariant tests (uniqueness, namespaced-format, no bare-controller, legacy-string match). - 13 sites in `reconciler/mod.rs` and 3 in `pairing*.rs` previously used bare `"azureclaw-controller"` — now use namespaced constants. All sites had `.force()` so the field-ownership transition is transparent on existing clusters. - 3 sites in `mesh_peer/{offload,pair}.rs` previously used bare `"azureclaw-mesh-peer"` — now use the constant `MESH_PEER` whose value is the legacy string verbatim (zero migration). - 6 per-CRD `FIELD_MANAGER` constants in their respective reconciler files re-export from the central registry — same string, central source of truth. - `providers::field_managers` preserved as backwards-compat re-export. Controller tests: 324 → 328 (+4 invariant tests). Workspace clippy + fmt clean. Audit: docs/security-audits/2026-04-28-phase2-conditions-ssa-leader.md S7 sub-slices remaining: B (Conditions matrix), C (leader election + predicated informers), D (backoff + reconcile-DAG), E (workqueue metrics), F (VAP/MAP expansion). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 1, 2026
…e 3) (#58) * feat(ci): add governance workflows from AGT (#1) Adds security and supply chain governance workflows adapted from agent-governance-toolkit: - dependabot.yml: Automated dependency updates for Cargo, Docker, and GitHub Actions - dependency-review.yml: Block PRs introducing high-severity vulnerabilities - scorecard.yml: Weekly OpenSSF Scorecard analysis - secret-scanning.yml: TruffleHog secret detection on push and PR Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): add DID canonicalization for AGT identity parity (Phase 3) - Add did.ts: deriveCanonicalDid(), parseDid(), normalizeDid(), classifyAddress() - Extend MeshIdentity interface with computed 'did' field (not persisted) - Update buildFacade() to derive canonical DID from Ed25519 signing key - Add DID to connection logging for dual-identity observability - Add 20 tests covering derivation, parsing, normalization, classification Canonical format: did:agentmesh:<hex(sha256(pubkey))[:16]> Accepts AGT TS SDK (did:agentmesh:<id>:<fp>) and Python (did:mesh:*) formats. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 5, 2026
…hments, sibling trust race, final-deliverable rule Live multi-agent demo (analyst→viz→writer fan-out) surfaced four distinct breakages that each silently degraded the run while the parent agent still self-reported success. 1. Image generation 404 (router URL prefix regression) Foundry's account-scoped /openai/v1/images/generations endpoint does NOT accept the /api/projects/<project>/ URL prefix that chat-completions tolerates. Commit c13f302 unified everything through that prefix. In dev (raw Azure OpenAI account, no project path) it works; in AKS prod against a Foundry project endpoint the upstream returns a fast 404 and image_generation falls back to written descriptions. Fix: strip /api/projects/<name>/ from the upstream endpoint inside the images_generations handler before forwarding. Add strip_project_prefix helper + four unit tests (project-stripped, trailing-slash variant, AOAI passthrough, no-prefix passthrough). 2. foundry_code_execute drops container files The Responses API tool only walked output[].type=='message' for text. matplotlib PNGs / CSVs generated by code_interpreter live inside Foundry's per-run container and are referenced via container_file_citation annotations or { type: 'image' } entries in code_interpreter_call.outputs. They were never downloaded, so the demo's bar chart silently degraded to ASCII. Fix: collect every (container_id, file_id, filename) reference from both shapes, GET each via the new /openai/containers/... router route (added to foundry_standalone_routes), and write the bytes to /sandbox/.openclaw/workspace/. Append the local paths to the tool result so downstream tools (mesh_transfer_file, file_write) can ship them. Adds routerCallBinary helper for binary downloads through the router. 3. Sibling KNOCK race in parallel fan-out AGT_TRUSTED_PEERS is baked at spawn time and consumed once at sub-agent boot. When the parent spawns analyst → viz → writer in sequence, only writer (last) sees all siblings. analyst's parentTrustedAmids only contains parent — so when viz or writer later try to KNOCK analyst, the trust score is 0 + 0 = 0 and the KNOCK is rejected at threshold 500. The demo logs confirm only 1 of 3 sibling pairs ever opened a session. Fix: after every successful spawn, the parent broadcasts a peers_update message containing the new sibling's AMID to every already-running sibling. Each sub-agent now records the parent's AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and handles peers_update only from that AMID, extending its parentTrustedAmids set at runtime. 4. Sub-agent system prompt missing FINAL DELIVERABLE rule Sub-agents were told they could mesh_transfer_file artifacts to peers, but nothing forced them to mesh_transfer_file the FINAL artifact back to the parent before returning a summary. The writer's executive_brief.md sat in its local /sandbox forever while the parent reported success. Fix: append a hard rule to the sub-agent system prompt requiring mesh_transfer_file(to_agent='parent', ...) as the last action before any "task complete" reply, with one call per output file. Tests - inference-router: 643 lib tests pass (4 new strip_project_prefix tests). - inference-router: cargo clippy --all-targets clean. - runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only pre-existing warnings. - cli: 553 tests pass. Deployment - For #2: rebuild + push inference-router image. - For #1, #3, #4: rebuild + push sandbox image. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 11, 2026
Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.
Two coordinated fixes:
1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
(spawner seeds 'parent_name:parent_AMID' as the first entry), also
push PARENT_SANDBOX=<first_name> into the openclaw container env.
2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
lookup. The Symbol is set during runtime init from
AGT_TRUSTED_PEERS[0], so this works even on images built before fix
#1 lands. Skip in offload mode — 'parent' there is a protocol-level
routing token, not a mesh recipient name.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 11, 2026
…#245) * feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap Wire azureclaw runtime through createMeshTransport() factory so we can flip between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without code changes. Surface additions to IMeshTransport (both adapters now expose): - lookup(amid) — registry RPC for reputation/display name - submitReputation(...) — registry RPC for peer feedback - enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on) - onError(kind, from, detail) — diagnostic hook for decrypt + ws errors - onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal - onDisconnect(reason, code) — ws close / error fan-out mesh-plugin (vendored A adapter): - connection.ts delegates to the underlying SDK; lazy bind for hooks registered before connect() - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the contract so neither adapter can drop a method without CI failing mesh-plugin (AGT B adapter): - agt-transport.ts implements lookup/submitReputation as REST calls to the registry (AGT MeshClient is pure transport — registry RPCs intentionally not added to AGT upstream; they belong on a separate RegistryClient) - enableKnockEnforcement is a no-op (AGT MeshClient always enforces) - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods (added on local AGT branch azureclaw-meshclient-event-hooks, NOT pushed — AGT team owns the upstream PR) runtime (runtimes/openclaw): - Adds @azureclaw/mesh as a file: dependency - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)' when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value - Identity is generated once via vendored SDK regardless of provider, then raw Ed25519 keys are extracted via toData() and shared across both — same AMID either way - Banner now reports active provider (vendored vs agt) Docs: - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity, policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer + the wiring + migration path - Documents the 3 hooks added to local AGT branch and the 3 governance methods kept adapter-side Tests: - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat) - runtimes/openclaw: 118/118 pass - AGT (local branch): 387/387 pass with 8 new event-hook tests Open work for cleanup phase: once AGT publishes the version with our event hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var toggle. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): pre-build mesh-plugin for runtime CI + format reconciler PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): quote workflow step name containing colon YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): clippy fixes for Rust 1.95.0 CI runs Rust 1.95.0 which added new clippy lints: - doc_lazy_continuation: indent doc list items that span multiple lines. Added two-space indent to the trailing 'All three are populated...' paragraph so it is treated as a continuation of the preceding list item rather than its own malformed list item. - obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as if .. { a } else { b } per the lint suggestion. These were pre-existing on dev (CI only started failing once the runner picked up Rust 1.95.0); fixing here so PR #245 can land green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12 Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5 The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): mark gaps G1 and G2 fixed on local AGT branch Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass. * dev: add --mesh-provider <vendored|agt> selection with first-run prompt Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker mode before AKS rollout. Same flag, three integration points: CLI (cli/src/commands/dev.ts): - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball - First-run interactive prompt offers AGT only if the toolkit checkout is actually present locally — silently defaults to vendored otherwise (no pestering for users without AGT cloned). - --build branch: builds the right relay/registry images vendored → vendor/agentmesh-relay + agentmesh-registry (Rust) agt → agent-governance-python/agent-mesh/docker/Dockerfile with COMPONENT=relay / registry build-args - Sandbox image build: stages locally-packed AGT SDK tarball into .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL build-arg (auto-discovers if --agt-sdk-tarball not given). - Runtime branch: skips Postgres for AGT (in-memory registry), uses correct ports (AGT: 8083 relay, 8082 registry; vendored: 8765/8080) and health path (AGT: /healthz; vendored: /v1/health). - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the runtime transport-factory honors the user's choice. Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile): - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the sandbox npm-installs the local tarball instead of fetching the published @microsoft/agent-governance-sdk from npm. Lets us smoke-test the locally-patched AGT branch (G3/G4 fixes) end to end without round-tripping through npm publish. - .agt-sdk/ staging dir always exists (with .keep) so the COPY never fails when the user didn't stage a tarball. Defaults preserved: --mesh-provider=vendored, existing behavior is byte-identical for users who don't opt in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so tsc failed with TS2307 in the AGT build path. Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the build context, and strip its 'prepare' script (which would invoke tsc, not present in this stage; the pre-built dist/ is sufficient). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks). - connect() now passes autoRegister: true so the SDK uploads identity and prekeys instead of the adapter re-implementing that path with raw HTTP. - discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT. - lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}). - submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score clamped to [0,1]; the vendored /registry/feedback endpoint does not exist in AGT. - Replaced mapAgent with pickDisplayName helper: AGT puts display name in metadata.display_name (set by registerSelf), with the first capability as the fallback. Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and the bespoke fetchWithRetry helper — both are upstream concerns now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes the vendored registry wire shape. The vendored impl talks to /registry/* (the existing agentmesh-registry); the AGT impl talks to /v1/discover and /v1/agents/{did} on the upstream AGT registry. Both expose a single normalized RegistryEntry envelope, so callsites stay readable. getMeshRegistry(routerUrl) is the entry point. Provider selection follows AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base). Migrated all raw-HTTP registry callsites: - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName, resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid. - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS spawn discovery, AKS→local discovery. - core/agt-task-loop.ts (1 site): registry_capability_search tool. - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe, azureclaw_discover (mesh_discover) tool. - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe, heartbeat keepalive (no-op under AGT — relay does liveness via WS). The discover-on-router-unreachable test now asserts the new contract: empty list + count:0 instead of a 'Discovery failed' string. Registry hiccups must not break tool calls. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): wire AGT provider end-to-end (6 stackup bugs) End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced six bugs blocking the upstream AGT MeshClient swap. All fixed: 1. Final sandbox Docker stage didn't COPY mesh-plugin, so the file:../../mesh-plugin symlink dangled in node_modules. Plugin swap silently fell back to vendored with 'Cannot find package @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist} into /mesh-plugin/ in the final stage of the Dockerfile. 2. entrypoint.sh used cp -r when copying node_modules into the plugin extension dir, preserving the (now-broken-at-runtime-path) symlink. Switched to cp -rL so symlinks dereference into real files in the target tree. 3. mesh-plugin/src/index.ts imported createMeshTransport from ./transport-factory.js but never re-exported it. Runtime swap path couldn't find the factory. Added the missing re-export. 4. inference-router agt_registry_proxy unconditionally prepended '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents' became '/v1/v1/agents' at the upstream. Now: forward verbatim when path starts with 'v1/' or equals 'health', else prepend. Preserves vendored SDK behavior ('registry/register' → /v1/registry/register). 5. /agt/relay route only matched the bare path, but AGT MeshClient appends '/ws' to relayUrl. Added /agt/relay/ws route and made the upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt. 6. agt_registry_proxy route was declared get(...).post(...) only. AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey upload and DELETE for deregister — both 405'd at the router. Added .put() and .delete() to the route declaration. Bug #6 was invisible to vendored because the vendored SDK only ever uses GET/POST (registry/register, registry/prekeys, etc.). AGT's switch to REST verbs exposed the gap. Path allowlist also extended with 'v1/' prefix so AGT's REST paths (v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation. Verified end-to-end via azureclaw dev --mesh-provider=agt --build: - POST /v1/agents → 201 Created - PUT /v1/agents/{did}/prekeys → 200 OK - WebSocket /ws accepted, stable connection (no reconnect loop) - Plugin reports 'AGT mesh connected' + provider=agt Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): always route mesh registry through inference-router azureclaw_discover and other mesh registry callsites went via `process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`, intending to let out-of-sandbox sub-agents bypass the router. In practice, the sandbox launcher always sets AGT_REGISTRY_URL as the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082 in dev, the K8s service URL in prod). Since the runtime runs as UID 1000 and iptables egress-guard blocks UID 1000 from anything except localhost+DNS, the direct upstream URL ECONNREFUSEs and the catch-all silently returns []. Symptom: registered agents are invisible to azureclaw_discover even though they show up in `GET /v1/discover` when queried directly at the registry. Drop the env-var override — there's no in-sandbox runtime path where bypassing the router is correct. The router is the ONLY way out for UID 1000. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): break mesh_send infinite poll loop on dead sub-agent probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but routerCall actually resolves with the parsed JSON error body. When the sub-agent pod/container is gone the router returns 404 with { error: "Container '<name>' not found..." } and probeSubAgentAlive read status.phase = undefined → defaulted to "Unknown" → not in POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every 2s forever, blocking the LLM event loop ("LLM not responding" symptom). Also narrow the prekey transient retry test so permanent X3DH / signature-verification failures bubble up instead of being treated as "waiting for prekeys" and retried indefinitely. Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'. Before: registry log fills with GET /v1/discover?capability=echo-buddy every ~2s forever; LLM stops responding to new turns. After: mesh_send aborts with 'sub-agent sandbox not found' on first probe. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): suppress /v1/registry/* 404 leaks in AGT mode Three vendored-only registry paths were being called unconditionally in AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply cadence: 1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X. The AGT registry exposes GET /v1/discover?capability=X instead — display names live in the per-agent record, so the AGT path fans out to a second /v1/agents/{did} fetch per discover hit. Driven by the operator panel's /agt/reputation polling. 2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session. AGT has no per-session counter; per-agent reputation already submitted via MeshClient.submitReputation. No-op in AGT mode. 3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode. Also includes complementary debugging fixes from this session: - agt-transport: auto-call establishSessionWithPeer() before send() so AGT mode gets vendored-equivalent send-with-first-contact semantics. Without this, send() throws 'No encrypted session — call establishSession() first' and the retry loop spins forever. - cli operator fetchers: add 8–10s timeouts to kubectl get calls that were hanging when the cluster API was unreachable. cargo check: clean runtimes/openclaw: 118 vitest tests pass inference-router: 8 mesh tests pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode Fourth 404 leak revealed after deploying the previous fixes: the operator panel's ~30s /agt/reputation poll triggers governance::agt_reputation, which (after lookup_parent_amid succeeds) fetched the per-agent reputation score via the vendored-only GET /v1/registry/reputation/score?amid=X path. AGT registry has no such endpoint — the score is embedded as 'reputation_score: f64' in the per-agent record returned by /v1/agents/{did}. Provider-dispatch the URL; for AGT, wrap the agent record in a vendored-shaped payload (score / tier / raw) so downstream CLI fetchers and the operator panel stay schema-agnostic. cargo check: clean agt_governance_integration: 26/26 pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * runtime: hide Foundry tools in github-copilot mode (same as github-models) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest Phase B.1 of the AGT-on-AKS rollout (see session plan files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the existing `azureclaw dev --mesh-provider` flag so the same provider selection works for AKS pushes. When --mesh-provider=agt: * Builds relay+registry from the AGT upstream Dockerfile ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile) using COMPONENT=relay|registry build-args (matches dev.ts). * Tags as agentmesh-{relay,registry}-agt:latest so both vendored and AGT images can coexist on the same ACR and so the existing deploy/agentmesh-agt.yaml manifest picks them up unchanged. * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz) into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg. * Always passes MESH_PROVIDER build-arg to the sandbox image so the Dockerfile's conditional `npm install @microsoft/agent-governance-sdk` runs for AGT clusters. When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml, applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt, THEN rolls the controller (so the new pod reads AZURECLAW_MESH_PROVIDER=agt for new sandboxes). Auto-reverses when --apply --mesh-provider=vendored runs against a cluster currently on AGT (no Postgres deployment in the agentmesh ns). The image build loop also now supports absolute Dockerfile paths and absolute build contexts via a new `absoluteContext` field, needed because the AGT Dockerfile lives outside the azureclaw repo root. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks without rebuilding any images, assuming both image pairs were already seeded by 'azureclaw push'. Flow: 1. Detect current provider via 'kubectl get deploy/postgres -n agentmesh' (vendored has Postgres, AGT does not). 2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found 3. kubectl apply -f deploy/agentmesh-<target>.yaml 4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target> 5. kubectl rollout restart deploy/azureclaw-controller 6. With --restart-sandboxes: roll every azureclaw-managed Deployment so existing pods pick up the new AZURECLAW_MESH_PROVIDER value. Service names and ports are identical between the two manifests (agentmesh-relay:8765, agentmesh-registry:8080) so the controller's mesh_peer talks to either stack with no further config — the relay/ registry URLs already come from env vars (MESH_RELAY_URL / MESH_REGISTRY_URL). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(up): --mesh-provider=agt picks AGT manifest + flips helm value Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up' so first-time deploys can ship AGT instead of vendored. When --mesh-provider=agt: * helm install runs with --set mesh.provider=agt (controller env AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes). * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of deploy/agentmesh.yaml. * Skips the postgres ACR import and the agentmesh-db-credentials secret creation (both unused by AGT — its registry is in-memory). * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml vs .tmp-agentmesh.yaml) so concurrent provider switches don't collide. The deployAgentMesh signature gains a non-breaking 'meshProvider' option that defaults to 'vendored' (existing callers untouched). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): plumb --mesh-provider into local-k8s helm install Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(controller): AGT wire protocol adapter for mesh_peer Implement full AGT relay/registry wire support in the controller's mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt. Without this the controller's federation peer cannot connect to the AGT relay (different WS path, frame envelope, heartbeat, ack model) or the AGT registry (different HTTP shape, no signed body), and the leader fails-loops on AGT clusters — breaking the only cloud-offload control path. New module `mesh_peer/agt_wire.rs`: - `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error) with `#[serde(tag="type", rename_all="snake_case")]` matching `agentmesh/relay/app.py`. - `AgtRegisterAgentRequest` struct for `POST /v1/agents`. - 7 unit tests pinning the serialized shape. `mesh_peer/mod.rs`: - New `Provider` enum + `Provider::from_env()` selecting vendored (default) or AGT off `AZURECLAW_MESH_PROVIDER`. - `MeshPeerState.provider` carried through outbound + inbound paths. - `register_with_registry()` branches: vendored signs ts body; AGT posts `{did, public_key (base64url), capabilities, metadata}` with no signature; 409 treated as success for leader-failover idempotency. - `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>` (matches JS SDK `buildDid`), so every leader replica converges on the same DID without coordination. - Default `MESH_RELAY_URL` appends `/ws` for AGT. - `connect_and_listen()`: - AGT connect frame `{type:"connect", from:<did>, token?:<env>}` (token read from `AGENTMESH_RELAY_TOKEN` if set). - AGT has no `Connected` ack — mark `connected=true` immediately. - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored keeps `ping`). - `serialize_and_send_outbound()` / `send_to_peer()` now take `state` and branch outbound framing — AGT emits `message` frames `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex). - `handle_message()` dispatches to `handle_vendored_frame()` or `handle_agt_frame()`. AGT path: - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`. - Sends `Ack` reply (required — without it AGT redelivers on reconnect → duplicate offload processing). - Treats `Error` frames mentioning Authentication failed / Missing 'from' / session_replaced as fatal — drops connection for reconnect. `mesh_peer/offload.rs`: - All 8 `send_to_peer(...)` call sites updated to pass `&state` first. `main.rs`: - Remove the temporary AGT-skip guard around `mesh_peer::run`. The peer now starts unconditionally when enabled; provider is consumed inside `mesh_peer::run`. Build/test: - cargo build --release --package azureclaw-controller: OK - cargo test --package azureclaw-controller: 492 passed - cargo clippy --package azureclaw-controller --all-targets -D warnings: OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer - cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(deploy): AGT mesh probe path + Cilium pod-port NP allow deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz (see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness probes were 404'ing → CrashLoopBackOff/NotReady. operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane evaluates NetworkPolicy egress against the backend pod port (post-DNAT), not the Service port. AGT registry/relay listen on 8082/8083; the Service maps 8080->8082 and 8765->8083 so the Service-port allowlist (8080/8765) doesn't actually permit the post-DNAT flow. Add 8082/8083 alongside so both vendored (8080/8765 direct) and AGT paths work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh promote): AGT-compat health + WS upgrade paths azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): first-run picker for local vs remote mesh source azureclaw dev now asks new users where the mesh should live, just like the existing inference-provider picker: Where should the mesh live? ❯ Local (recommended; spin up relay + registry in Docker) Remote (auto port-forward to AKS cluster: <cluster-name>) Local (default) keeps the existing behaviour: docker-compose'd relay/registry/postgres on the user's laptop. Remote (advanced) federates with a previously-provisioned AKS mesh: - If ~/.azureclaw/context.json has a cached globalRegistryUrl from a prior 'azureclaw mesh promote', reuse it verbatim. - Otherwise default to http://localhost:18080 — the port-forward URL 'mesh promote --port-forward' uses — so the auto-promote fallback in the downstream global-registry block will spawn the tunnels on demand. - If there is no aksCluster in context at all, warn and fall back to local so the user isn't left with a broken sandbox. Skipped entirely when --global-registry was passed explicitly (the advanced flag overrides the prompt) or when the user is past their first run. Also fixed a latent AGT-compat bug in the same flow: the existing 'auto-promote' path probed only /v1/health, which 404s on AGT clusters. Replaced with a /health → /v1/health fallback (matches the same shape we used in checkRegistryHealth last commit). Verified: - npm run build / typecheck clean - 640 CLI tests pass (2 skipped, no regressions) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply back to their spawner, but on AGT the registry has no agent named or capability="parent" — the search returns 0 → no prekey bundle → send fails. The vendored runtime had this aliased only in the offload-mode task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the controller never set for AKS-spawned children. Two coordinated fixes: 1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set (spawner seeds 'parent_name:parent_AMID' as the first entry), also push PARENT_SANDBOX=<first_name> into the openclaw container env. 2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send and azureclaw_mesh_transfer_file, alias to_agent=='parent' → PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry lookup. The Symbol is set during runtime init from AGT_TRUSTED_PEERS[0], so this works even on images built before fix #1 lands. Skip in offload mode — 'parent' there is a protocol-level routing token, not a mesh recipient name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap mesh-plugin/src/agt-transport.ts.send() called this.client.establishSessionWithPeer(toAmid) before forwarding to client.send(). That method does not exist on AgentMeshClient — the real method is establishSession(toAmid, options) — so every parent → sub-agent send on AGT was failing with: establishSessionWithPeer is not a function It was also unnecessary: AgentMeshClient.send() already auto-bootstraps the X3DH handshake on first contact (see @agentmesh/sdk AgentMeshClient.send → cache miss → establishSession() fallthrough at dist/index.js:3321-3334). Calling establishSession() ourselves would also be wrong because it is not idempotent — it unconditionally writes activeSessions.set and starts a fresh X3DH. Fix: remove the pre-bootstrap entirely and let client.send() manage session lifecycle. The AgtSdkModule type loses the required establishSessionWithPeer member (now optional) since we no longer depend on it; test fakes remain valid as harmless extras. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call Previous commit aa7d28e wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * push: auto-detect mesh provider from live helm release When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * entrypoint: fail-open trust gate when running anonymous tier When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable The AGT registry has no autonomous presence model — `last_seen` is frozen at registration and the `update_last_seen()` store method is dead code with no HTTP handler calling it. Combined with the openclaw discover tool's 90s stale filter (agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes silently invisible 90s after spawn, breaking sibling-to-sibling peer discovery. Demo symptom: analyst/viz/writer all reported 'peer discovery did not return ...' even though mesh_send to those names succeeded with 'delivered_and_replied'. The relay was fine; only the registry's presence view was stale. Pair with the corresponding upstream registry change (AGT branch `azureclaw-meshclient-event-hooks`, commit adds POST /v1/agents/{did}/heartbeat -> store.update_last_seen). The new tick reuses the existing 30s relay-keepalive timer in connect(), so no extra timers and no extra event-loop pressure. Best-effort: 4xx/5xx are warned-once, network errors swallowed, loop survives a registry pod restart (next tick retries). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
Bump jsonwebtoken from 9.3.1 to 10.3.0
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
* docs: add global AgentMesh handoff design document
Comprehensive design for agent live migration (local ↔ cloud):
- Identity succession protocol (Ed25519 signed, no key transfer)
- Reclamation protocol (co-signed reverse handoff)
- Sub-agent re-spawn with state injection
- Three-layer handoff endpoint auth (handoff token + no localhost bypass + mutual attestation)
- Security review: 11 threat findings with mitigations
- Handoff trigger security (confirmation token, time delay, AGT policy gate)
- UX design across webchat, TUI, and Telegram
- Demo script and implementation phases (H1-H4)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H1 — handoff module with three-layer auth
Router-side handoff infrastructure for agent live migration (local ↔ cloud):
## New module: handoff.rs (1368 lines)
- HandoffState, SubAgentSnapshot, HandoffMetadata, CredentialRef structs
- HandoffTokenStore: in-memory, TTL-based, one-at-a-time token management
- 32-byte random tokens, max 10min TTL, constant-time comparison
- Token hash logged for audit (never the token value)
- HandoffSession: phase tracking across the full handoff lifecycle
(idle → initialized → draining → snapshotting → transferring → restoring
→ verifying → decommissioning → complete | failed | aborted)
- DrainState: stops new work during handoff, tracks duration
- State serialization: JSON + gzip compression
- State encryption: AES-256-GCM with HKDF-SHA256 key derivation
Key derived from shared secret + salt using 'azureclaw-handoff-v1' info
- Verification: SHA-256 hash of plaintext for integrity checking
- 21 unit tests covering token store, serialization, encryption, sessions
## New endpoints (8 routes, three auth tiers)
1. POST /agt/handoff/init — admin token only, NO localhost bypass
2. POST /agt/handoff/snapshot — creates encrypted state blob
3. POST /agt/handoff/restore — decrypts, validates, restores state
4. POST /agt/handoff/verify — returns verification digest
5. POST /agt/handoff/drain — enters drain mode
6. POST /agt/handoff/decommission — agent goes dormant
7. POST /agt/handoff/abort — cancels in-progress handoff
8. GET /agt/handoff/status — read-only (localhost allowed)
## Security: three-layer authentication
- Layer 1: Handoff token (one-time, short-lived, CLI-only)
Token exists only in CLI process memory — never in pod env
- Layer 2: NO localhost bypass for mutation endpoints
Prevents prompt injection from exfiltrating state via localhost
- Layer 3: Mutual attestation via DH-encrypted state blob
(Phase H2 adds Ed25519 succession signature verification)
## All endpoints audit-logged with:
- Caller IP, timestamp, endpoint, success/failure
- Token hash (not value), state blob size, item counts
## Dependencies added:
aes-gcm 0.10, hkdf 0.12, sha2 0.10, rand 0.9, base64 0.22, flate2 1
## Test results:
- 77 unit tests pass (21 new handoff tests)
- 26 integration tests pass (updated for new AppState fields)
- 74 controller tests pass (unaffected)
- clippy clean (zero warnings)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H2 — registry mode, identity succession, and reclamation
Phase H2 of the agent handoff feature:
Registry topology (local vs global):
- Add RegistryMode enum to router config (AGT_REGISTRY_MODE env var)
- Handoff init returns 409 in local mode with clear guidance
- Global mode does startup health check on AGT_REGISTRY_URL
- Handoff status endpoint exposes registry_mode + handoff_available
Identity succession (A→B):
- SQL migration 008_succession.sql with succession_log table
- POST /v1/registry/succession endpoint with Ed25519 sig verification
- Canonical message format: succession:{pred}:{succ}:{timestamp}
- One-shot rule (unique index on active predecessor)
- Copies reputation A→B, marks predecessor dormant
Identity reclamation (B→A, co-signed):
- POST /v1/registry/reclamation with dual signature verification
- Original succession ref must match active event_hash
- Deactivates succession redirect, copies reputation back
- Sets original online, departing offline
Lookup follows succession redirects:
- lookup_agent checks succession_log for dormant predecessors
- Returns successor with succeeded_from + succession_hash metadata
- Max redirect depth = 1 (no chains)
Dormant presence status:
- New PresenceStatus::Dormant variant in registry
- Ghost cleanup skips dormant agents (preserves succession chains)
- Capability search excludes dormant agents
CLI --global-registry flag:
- azureclaw dev --global-registry <url> skips local registry stack
- Passes AGT_REGISTRY_MODE=global to router
- Health check on global registry at startup
- Status display shows "handoff enabled" for global mode
Tests: 177 Rust (77 unit + 26 integration + 74 controller) + 159 CLI
All passing, clippy clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement Phase H3 — CLI command and plugin tools
CLI command (cli/src/commands/handoff.ts):
- azureclaw handoff <name> --to cloud (forward: local → AKS)
- azureclaw handoff <name> --to local (reverse: AKS → local)
- azureclaw handoff <name> --status (check progress)
- azureclaw handoff <name> --abort (cancel in-progress)
- 7-step orchestration: verify → init → snapshot → drain → transfer → succession → summary
- Registry mode guard (blocks with clear guidance if not --global-registry)
- Shared secret derivation for snapshot encryption
- Full error recovery (abort on failure)
Plugin tools (cli/src/plugin.ts):
- azureclaw_handoff_status: check handoff availability and state
- azureclaw_handoff_request: LLM can request handoff, but returns CLI
command (tokens are CLI-only to prevent prompt injection attacks)
Security design:
- Handoff tokens live in CLI process memory only (never LLM-accessible)
- LLM tool returns the CLI command to run, not the token itself
- This prevents prompt injection from triggering unauthorized handoff
Registered in cli.ts under 'Agent mobility' section.
Tests: 177 Rust + 159 CLI, all passing, clippy clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security(handoff): close §9.9 security gaps — confirmation gate, injection protection, rate limits
§9.9.9 — Two-stage confirmation gate (HIGH priority):
- PendingHandoffStore with confirmation token (4-byte random hex)
- POST /agt/handoff/pending — creates pending request (rate limited: 1 per 5min)
- POST /agt/handoff/confirm — validates token, enforces 3s minimum delay
to prevent LLM self-confirm, consumes pending on success
- Plugin azureclaw_handoff_request now calls /pending (Stage 1)
- New plugin azureclaw_handoff_confirm tool (Stage 2)
- 6 unit tests for PendingHandoffStore (create, confirm, delay, rate limit,
cancel, wrong token)
§9.9.1 — State blob prompt injection protections:
- sanitize_chat_snapshot() strips messages matching 17 injection patterns
(system prompt override, handoff commands, instruction ignoring)
- User messages always preserved (legitimate user content)
- Non-UTF8 chat snapshots rejected entirely
- Trust scores capped at 750 on restore (cannot import max trust)
- 4 unit tests for chat sanitization
§9.9.4 — State blob size/DoS limits:
- 50MB blob size cap on both snapshot and restore
- MAX_WORKSPACE_FILES (100) and MAX_WORKSPACE_FILE_SIZE (10MB) constants
- PAYLOAD_TOO_LARGE (413) returned on violation
§9.9.3/§9.9.8 — Rate limits:
- Succession rate limit: 1 per AMID per 5 minutes (DB-backed)
- Reclamation rate limit: 1 per AMID per hour (DB-backed)
- check_succession_rate_limit() queries succession_log timestamps
§9.9.9 — AGT policy rule (belt-and-suspenders):
- handoff-tool-approval rule in azureclaw-default.yaml
- type: approval, priority: 75 (higher than tool-allow at 70)
- Requires operator approval for tool:azureclaw_handoff_request:*
and tool:azureclaw_handoff_confirm:*
Tests: 188 Rust (74 controller + 88 router + 26 integration) + 159 CLI
All passing, clippy clean, registry cargo check clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): implement global registry deployment — Ingress, OAuth, relay auth, CLI
Phase G1 implementation:
- G1a: AGIC Ingress manifest (deploy/agentmesh-ingress.yaml) with
NetworkPolicy (postgres locked to registry, registry/relay to AppGW),
Azure-managed TLS, WAF rate limiting, WebSocket support for relay
- G1b: Entra ID OAuth provider added to agentmesh-registry
(authorize, callback, token validation via Microsoft Graph).
Existing GitHub + Google providers untouched.
- G1c: Deployment manifest updated with OAuth secret references
(agentmesh-oauth-credentials), REGISTRY_URL for relay verification
- G1d: CLI 'azureclaw mesh auth' command — generates Ed25519 keypair,
runs browser-based OAuth flow, stores encrypted identity in
~/.azureclaw/mesh-identity.json (AES-256-GCM, machine-bound key).
Subcommands: auth, status, reset.
- G1e: CLI 'azureclaw up --global-registry <url>' skips local registry
deployment. '--expose-registry' deploys AGIC Ingress to make this
cluster's registry the global endpoint. Context persists registry mode.
- G1f: Relay registration verification — after Ed25519 signature check,
relay calls registry /v1/registry/lookup to confirm AMID is registered.
Unregistered/revoked agents rejected. Fails open on registry errors
(avoids cascading failures). Gated by REQUIRE_REGISTRATION=true.
Security: 4-layer auth chain (WAF → Ed25519 → registry check → OAuth).
PostgreSQL never exposed externally (NetworkPolicy enforced).
Private keys encrypted at rest (AES-256-GCM).
Tests: 188 Rust + 159 CLI passing, all clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test+docs(mesh): integration tests and security/architecture documentation
Tests:
- 28 new CLI mesh tests (mesh.test.ts): base58 encoding, Ed25519
keypair generation, AMID derivation, encrypt/decrypt roundtrip,
tamper detection, command structure verification
- 3 relay registry verifier tests (registry_verify.rs): disabled
verifier passthrough, env-based construction, enable logic
(compile-gated by pre-existing ed25519-dalek API mismatch in relay)
- All 188 Rust + 187 CLI tests passing
Documentation:
- architecture.md: new 'Global Registry Deployment' section — deployment
modes table, 4-layer auth chain diagram, NetworkPolicy enforcement,
identity management overview
- security.md: new 'Layer 9: Global Registry & Handoff Security' section
— relay auth layers table, handoff threat/mitigation matrix,
NetworkPolicy diagram, identity-at-rest encryption details
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(plugin): gate handoff mutation tools behind AGT_REGISTRY_MODE=global
In local registry mode, only azureclaw_handoff_status is registered.
The request and confirm tools are hidden from the LLM, preventing
unnecessary AGT governance prompts for tools that would 409 anyway.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): add promote/demote commands for registry global mode
azureclaw mesh promote — deploys AGIC Ingress + NetworkPolicies to expose
the cluster's AgentMesh registry and relay as public endpoints. Updates
deployment context to global mode.
azureclaw mesh demote — removes Ingress resources and reverts to
cluster-local registry. Disables cross-environment handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): add --allow-ip to promote for IP-based access control
mesh promote auto-detects your public IP (via ifconfig.me) and injects
the AGIC whitelist-source-range annotation into both Ingress resources.
Override with --allow-ip <cidr>. If detection fails, warns and leaves
the registry open.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): auto-detect AppGW IP and use sslip.io for zero-config DNS
mesh promote now queries the AGIC Application Gateway for its public IP
and generates sslip.io hostnames (e.g. registry.20-30-40-50.sslip.io).
No DNS setup needed for testing. TLS is disabled for sslip.io domains
(secured by IP allowlist instead). Use --domain for custom domains.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* refactor(mesh): switch promote/demote to LoadBalancer Services
Replace Ingress-based approach with direct LoadBalancer Service patching.
No ingress controller needed. promote patches registry + relay services
to LoadBalancer with loadBalancerSourceRanges for IP restriction, waits
for external IPs, builds sslip.io URLs. demote reverts to ClusterIP.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(plugin): LLM-driven handoff orchestration via E2E mesh
Replace the CLI-only handoff confirm flow with full LLM-driven
orchestration. After the user confirms the handoff code, the plugin
now executes the entire transfer autonomously:
1. Confirm → router creates handoff token (stays in plugin memory)
2. Snapshot → encrypted AES-256-GCM state blob
3. Drain → stop accepting new work
4. Spawn → create cloud target on AKS (or find existing)
5. Transfer → send state blob via E2E encrypted mesh (Signal Protocol)
6. Verify → target restores, sends verification digest back via mesh
7. Succession → registry identity chain update
8. Decommission → local agent enters dormant state
Key changes:
- handoff_confirm tool: full orchestration instead of returning CLI cmd
- onMessage handler: new handoff_transfer message type for target agent
to auto-restore state and send verification back
- _routerCallStrict: new helper that rejects on HTTP >= 400
- _readAdminToken: reads admin token from filesystem paths
- _routerCall: added extraHeaders parameter (backward compatible)
- agtReconnect: disconnect before connect to clear stale SDK state
Security model (§9.9.9): the LLM can REQUEST a handoff but never
EXECUTE one. The handoff token stays in plugin memory — the LLM
never sees it. All router calls use this token. Human confirmation
via the 2-stage code flow is the gate.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(mesh): promote checks health and reconnects stale port-forwards
When registry is already in global mode, 'azureclaw mesh promote' now:
- Checks registry HTTP health (/v1/health)
- Checks relay TCP connectivity
- If both healthy: reports status and exits
- If either dead: kills stale PIDs, clears held ports, restarts
fresh port-forward tunnels, verifies connectivity
Previously it just said 'already global' and exited, even when the
port-forwards had died (e.g. after IP change or sleep).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(plugin): use async import for fs in ESM context
_readAdminToken used require('node:fs') which is unavailable in ESM.
Changed to async function with await import('node:fs') and updated
both call sites to await the result.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): spawn AKS pod from dev mode for cloud handoff
In dev mode, the router's /sandbox/spawn endpoint was creating Docker
containers. For handoff (local→cloud), we need actual AKS pods.
Changes:
- Add HandoffMeta struct to SpawnRequest (mode + predecessor fields)
- When handoff.mode='restore' in dev mode, bypass Docker path and use
K8s CRD creation via kube-rs (kubeconfig mounted from host)
- Mount ~/.kube/config into dev container at /run/secrets/kubeconfig
so the router can reach the K8s API for handoff spawns
The controller already sets AGT_RELAY_URL and AGT_REGISTRY_URL on
spawned pods, and NetworkPolicy allows mesh egress — so the handoff
target automatically joins the global mesh.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): propagate trusted_peers and registry_mode to spawned pods
The handoff target was rejecting the source's KNOCK (trust score 0 <
threshold 500) because AGT_TRUSTED_PEERS wasn't propagated. Also,
AGT_REGISTRY_MODE wasn't set, so handoff tools were skipped.
Changes:
- CRD: add trusted_peers and registry_mode fields to GovernanceConfig
- Controller: propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to
the openclaw container env vars
- Spawn: write trusted_peers and registry_mode='global' into CRD
governance spec for handoff targets
- Spawn: use 'handoff'/'predecessor' labels instead of 'agent'/'parent'
for handoff-spawned CRDs (not sub-agents)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: add cloud handoff flow diagram (section 11)
Sequence diagram covering all 5 phases: two-stage confirm, snapshot/drain,
spawn on AKS, E2E mesh transfer, succession/decommission. Includes security
model diagram and current vs future (Entra OAuth) trust flow comparison.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): async orchestration with real-time progress tracking
Refactor handoff_confirm to return immediately and run orchestration
in the background via _runHandoffOrchestration(). The LLM polls
handoff_status every 3-5s and relays emoji step updates to the user
in real-time instead of blocking for 2-3 minutes.
Key changes:
- HandoffProgress interface tracks phase, steps[], status, error
- _hp() helper updates progress + logs at each step
- _runHandoffOrchestration() contains the full 7-step flow:
snapshot → drain → spawn → mesh-wait → transfer → verify →
succession → decommission
- handoff_status returns rich progress with active polling instruction
- Module-level _log set during register() for background access
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address security and reliability findings from handoff audit
sec-6: Fix message filter AND→OR — verification now rejects messages
unless BOTH from_amid AND from_agent match the expected target
sec-1: Propagate AGT_TRUSTED_PEERS and AGT_REGISTRY_MODE to router
container (was only on openclaw container)
sec-2: Validate trusted_peers — reject values with control chars
sec-3: Validate registry_mode — only accept 'local'|'global'
rel-7: Wrap _runHandoffOrchestration in top-level try-catch
rel-6: Replace non-null assertions with explicit guard at completion
rel-3: Bump snapshot timeout 15s→60s, drain timeout 15s→30s
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: revise handoff flow diagrams with full review findings
Replace Section 11 with 7 comprehensive diagrams:
- 11.1: End-to-end sequence (both source + target sides, live progress)
- 11.2: Handoff state machine with known limitations noted
- 11.3: 7-layer security model (gate → isolation → auth → encryption →
injection hardening → identity → infrastructure)
- 11.4: Env var propagation showing both containers receive vars
- 11.5: Two orchestration paths (LLM vs CLI) and their differences
- 11.6: Trust flow (current unauthenticated vs future Entra OAuth)
- 11.7: Error recovery and planned improvements
Also add nohup.out to .gitignore (stale port-forward logs).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address remaining TS findings — orphan cleanup, guards, tests
rel-1: Clean up orphaned CRDs on abort when mesh transfer or discovery
fails after spawn (DELETE /sandbox/spawn/<name>)
rel-8: Concurrent handoff guard — reject confirm if handoff already running
sec-4: Respect $KUBECONFIG env var with fallback to ~/.kube/config
test-3: Fix 4 failing spawn error tests — tools handle unreachable router
gracefully (return status JSON), update assertions accordingly
(187/187 tests now pass)
dup-1: Document dual orchestration paths (CLI operator-mode vs plugin
LLM-mode) in handoff.ts header comment + architecture-diagrams.md
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(review): address Rust findings — state machine, resume, tests
rel-2: Add try_transition() to enforce handoff phase ordering
(Idle→Init→Snapshot→Drain→Transfer→Restore→Verify→Decom→Complete)
rel-4: Add resume() + POST /agt/handoff/resume endpoint to cancel
drain state after abort (Aborted|Draining → Idle)
rel-5: Change body.unwrap() to expect() with message in spawn.rs
dead-1: Remove unused _used field from ActiveToken
test-1: Add 10 state machine transition tests (valid sequence,
invalid skip, abort, fail, resume, restart after complete)
test-2: Add 4 auth token tests (wrong value, no active, after
revoke, wrong pending confirmation code)
201 Rust tests pass (74 controller + 101 router + 26 integration)
Clippy clean with -D warnings
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* handoff: incremental progress polling, SDK reconnect fix, policy cleanup
- handoff_status tool: add since_step param for incremental polling,
returns only new_steps since last call so LLM relays one step at a time
- vendor SDK patch #9: AgentMeshClient.connect() no longer sets
connected=true when transport.connect() returns false, allowing retry
- Policy: remove handoff-tool-approval gate (two-step confirmation code
mechanism is sufficient; approval gate can return once native UI exists)
- config: add promoteMode to DeploymentContext for mesh promote tracking
All tests pass: 201 Rust (74+101+26), 187 CLI
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(crd): add trustedPeers and registryMode to Helm CRD schema
The K8s API server was silently stripping these fields because the Helm
CRD template only defined enabled/toolPolicy/trustThreshold. spawn.rs
wrote the fields and reconciler.rs read them, but the schema validation
layer dropped them in between.
Root cause of handoff mesh registration failure: target pods never
received AGT_TRUSTED_PEERS or AGT_REGISTRY_MODE env vars because the
CRD never stored the values.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(router): read admin token from correct mount path
AppState::new() only checked /run/secrets/admin-token but the controller
mounts the secret at /etc/azureclaw/secrets/admin-token. This caused
'Server misconfiguration: no admin token' on target pods during handoff
verification. main.rs had the correct path but its token was only used
for the admin_auth_middleware, not the handoff middleware which reads
from state.admin_token.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): pre-build audit — snapshot strict, direction propagation, relay URLs
Three defensive fixes from comprehensive flow audit:
1. Snapshot endpoint now uses _routerCallStrict (was _routerCall) —
if snapshot fails, error surfaces immediately instead of continuing
with undefined blob data
2. Target-side handoff_transfer handler now reads direction from the
mesh message instead of hardcoding 'local_to_aks' — enables
reverse (aks_to_local) handoffs
3. Controller propagates AGT_RELAY_URL and AGT_REGISTRY_URL to the
openclaw container (was only on router container) — plugin no
longer relies on fallback to router proxy for relay connection
All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): enforce state machine — migrate all handlers to try_transition
All 5 handoff route handlers now use try_transition() instead of
set_phase(), returning 409 Conflict on invalid phase transitions.
Also fixed the transition rules: Decommissioning is now allowed from
Draining (source-side flow: Init→Snapshot→Drain→Decommission skips
Verify/Restore which happen on the target router, not the source).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): verification hash mismatch + synchronous progress
Two fixes:
1. Verification hash mismatch: verify endpoint was rebuilding a fresh
snapshot (new timestamp, nonce, hostname) instead of using the hash
from the restored data. Now restore stores the hash of the decrypted
compressed bytes, and verify reuses it. Falls back to build_snapshot
for source-side verify (where no restore happened).
2. No proactive progress: LLMs don't autonomously poll tools, so the
handoff_confirm tool now awaits _runHandoffOrchestration() and
returns all steps when complete, instead of firing-and-forgetting
and expecting the LLM to poll handoff_status.
All tests pass: 201 Rust (74+101+26), 187 CLI, clippy clean
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(spawn): check Docker API HTTP status codes in docker_api
The docker_api helper only checked curl's exit code, not the HTTP
response status from Docker Engine. This caused silent failures —
e.g. container start returning HTTP 404 (network not found) was
swallowed and spawn reported success even though the container
never started.
Add -w flag to capture HTTP status code and return Err for 4xx/5xx
responses with the Docker error message extracted from the JSON body.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): propagate channel credentials to cloud target
During handoff spawn, collect channel/plugin credentials from the
source environment (TELEGRAM_BOT_TOKEN, SLACK_BOT_TOKEN, etc.) and
create a {name}-credentials K8s secret in the target namespace.
The controller already mounts this secret via envFrom (optional),
so the cloud agent inherits Telegram and other channels.
Also fix docker_api to check HTTP status codes — previously it only
checked curl's exit code, silently swallowing Docker Engine errors
like 'network not found' (HTTP 404).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): full state hydration — workspace, memory, conversations, Telegram
Source side (orchestration):
- Pack workspace tar from /sandbox/.openclaw/ (AGENTS.md, SOUL.md, etc.)
- Search Foundry Memory for recent context and include as chat_snapshot
- Include credential refs (channel/plugin names) in snapshot
Target router (restore):
- Extract workspace tar to /sandbox/ with path traversal protection
- Write chat_snapshot and metadata to /tmp/handoff/ for plugin
Target plugin (post-restore hydration):
- Create Foundry Conversation with replayed chat messages
- Store handoff event fact in Foundry Memory (update_memories)
- Write HANDOFF_CONTEXT.md to workspace (fallback context)
- Send 'handoff_ready' mesh message back to predecessor
The cloud agent now comes online with full context: workspace files,
conversation history in Foundry Conversations, semantic memory via
shared Memory Store, and proactively greets the user via Telegram.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: cross-container handoff — return state in response, not filesystem
The router and openclaw containers have separate filesystems in AKS.
Previously the router wrote workspace tar and chat snapshot to /tmp/
which the plugin couldn't read.
Changes:
- Router: return workspace_tar (base64) and chat_snapshot in restore
response JSON instead of writing to filesystem
- Plugin: extract workspace tar and parse chat snapshot from the
HTTP response body (runs in openclaw container where /sandbox/ lives)
- Remove unused extract_workspace_tar() from routes.rs
- Remove tar crate dependency (extraction now done by plugin via CLI tar)
- Add direction, initiated_at, restored_at to restore response
Future: workspace payloads >5MB will auto-transfer via Azure Blob
Storage (SAS URL in handoff state) — not yet implemented.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security: harden workspace tar extraction and chat snapshot parsing
Tar extraction:
- Pre-extract validation: list entries, reject any containing '..' or
starting with '/' (path traversal)
- Size guard: reject compressed payloads >5MB (decompression bomb)
- Unique temp dir per extraction (race condition prevention)
- --no-same-owner --no-overwrite-dir flags on extraction
- Temp dir cleaned up after extraction
- Removed '|| true' — errors now surface in logs
Chat snapshot:
- Schema validation: must be array, each entry must have string
role + content
- Cap at 100 messages, role capped at 20 chars, content at 10k chars
- Rejects non-conforming entries silently
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* perf: split sandbox Dockerfile into base + overlay for fast rebuilds
The sandbox image was 4.96 GB with ~4.0 GB of rarely-changing deps
(OpenClaw, Python wheels, Go tools, Node.js, CLI tools) rebuilt on
every code change. Now split into:
Dockerfile.base (~4.0 GB, rebuild weekly/on dep upgrade):
- Azure Linux 3 + system packages
- Node.js 22, Python 3 + 41 packages, Go CLI tools
- OpenClaw framework + extension symlinks + skills
- gh, ripgrep, 1password, himalaya
- User setup (sandbox:1000, router:1001)
Dockerfile (~50 MB, rebuild per commit in ~30s):
- FROM azureclaw-sandbox-base (all heavy deps pre-cached)
- CLI plugin builder reuses base image (has Node.js already)
- Router binary, plugin dist, vendored SDK overlay
- Entrypoint, proxy-bootstrap, skills, policies
All functionality preserved:
- UID separation, iptables egress guard, seccomp, read-only rootfs
- Channel plugins (Telegram/Slack/Discord/WhatsApp)
- Extension dep symlinks (grammy, carbon, bolt, etc.)
- Vendored SDK overlay, proxy-bootstrap, Control UI symlink
- ClawHub skills, npm CLI tools (clawhub, mcporter, oracle)
Build paths updated:
- azureclaw dev: auto-builds base if not cached, --build-base to force
- azureclaw push: --only sandbox-base to push base image
- Makefile: image-sandbox-base target added
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(handoff): implement reverse handoff (cloud → local)
CLI-driven reverse handoff orchestration:
- aksRouterExec: kubectl port-forward to AKS pod router
- wakeDormantDocker: detect and restart stopped containers
- readAksCrdSpec: inherit model/egress/isolation from CRD
- rehydrateCredentials: copy K8s secrets to Docker container
- Full 10-step reverse flow: connect → verify → init → snapshot →
drain → wake local → credentials → restore → succession →
decommission + delete CRD
Direction-aware source routing:
- sourceExec alias delegates to routerExec (Docker) or
aksRouterExec (AKS) based on direction
- Forward path unchanged at runtime (sourceExec === routerExec)
Plugin reverse handoff:
- handoff_request returns CLI command for aks_to_local direction
- Completion messages updated for both directions
- Decommission label direction-aware
Operator TUI:
- 'returning' handoff state for active aks_to_local handoff
- Table shows '<' icon and 'Returning' status
- ASCII-only table icons for reliable column alignment
- Column widths tightened
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(router): raise body limit on handoff routes to 50MB
Axum's default body limit is 2MB. Encrypted state snapshots easily
exceed this, causing HTTP 413 on /agt/handoff/snapshot and /restore.
Add DefaultBodyLimit::max(MAX_BLOB_SIZE_BYTES) layer to
handoff_protected_routes — matches the existing 50MB blob size
constant from §9.9.4.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): read AKS admin token from mounted secret
On AKS, the admin token is stored in K8s secret 'router-admin-token'
and mounted at /etc/azureclaw/secrets/admin-token — not as env var
or /tmp file. Updated getAksAdminToken() to:
1. Read from /etc/azureclaw/secrets/admin-token (router container)
2. Fallback: same path in openclaw container
3. Fallback: kubectl get secret (base64 decode)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): use POST for snapshot route (was GET → 405)
The /agt/handoff/snapshot route is POST-only but both forward and
reverse handoff paths were sending GET requests, causing HTTP 405.
Changed both to POST.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): reuse existing snapshot blob in reverse path
The reverse handoff was requesting a second snapshot at step 9, but
the state machine had already advanced to 'draining' after step 5.
The snapshot blob was already captured at step 3 — now the reverse
path uses snapshotResp.body.blob directly instead of re-fetching.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): pipe restore payload via stdin for large blobs
routerExec passes JSON as a curl -d argument, which hits shell
argument length limits for large encrypted snapshots. The reverse
handoff restore now uses 'docker exec -i ... curl -d @-' with the
payload piped via stdin, avoiding ARG_MAX issues.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(connect): handle Ctrl+C to disconnect port-forward
The kubectl port-forward child process with stdio:pipe did not
receive SIGINT from the terminal. Added explicit SIGINT/SIGTERM
handlers that terminate the child process and exit cleanly.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): init local handoff session + auth headers for restore
The local Docker router requires admin token + handoff token for
/agt/handoff/restore. The reverse path now:
1. Gets local admin token from Docker container
2. Inits a handoff session on the local router
3. Passes both auth headers to the restore curl call
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): correct step count + Telegram notification on handoff
- Fixed reverse handoff step counter: 13 steps (was 10, showing 11/10+)
- Added Telegram notification on handoff completion for both directions:
- local→cloud: 'I moved to the cloud'
- cloud→local: 'I am back on your local machine'
Best-effort — reads credentials from Docker container (reverse) or
env (forward), sends via Telegram Bot API. Failures are silent.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): clean up local handoff state after reverse restore
Two fixes for stale handoff sessions blocking subsequent handoffs:
1. CLI: After successful reverse restore, transition local router
through verify → decommission to reach a terminal state.
2. Router: Expand can_start() to allow re-init from Restoring,
Verifying, and Decommissioning phases. These indicate a previous
handoff that completed data transfer but wasn't properly finalized.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(handoff): extend verification timeout + retry mesh send
The verification timeout was 60s but AKS pods can take longer to
fully initialize their plugin message handlers after mesh registration.
The blob sent before the handler is ready gets silently dropped.
Fix:
- Extended timeout from 60s to 180s
- Re-send the handoff_transfer blob every 30s within the verification
loop, in case the target's message handler wasn't ready on first send
- Also fixes: router can_start() allows stale Restoring/Verifying states,
local handoff session cleaned up after reverse restore
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(relay): increase max_message_size to 1MB for handoff blobs
Handoff snapshots with real state (chat, audit, credentials) can be
80+ KB. After Signal Protocol encryption + base64 + JSON envelope,
they exceed the relay's 64KB default max_message_size. Bumped to 1MB.
Also includes verification timeout extension (60s→180s) with 30s
re-send retries, already committed in plugin.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: collect workspace/chat/credentials in CLI handoff snapshot
The CLI handoff command was sending an empty snapshot payload
(only shared_secret). The forward handoff via plugin.ts collected
workspace tar, Foundry memories, and credential refs — but the
CLI path (used for both forward and reverse) skipped this.
Changes:
- Collect workspace tar via kubectl exec (AKS) or docker exec (local)
- Collect Foundry Memory Store items as chat context
- Collect credential refs from container environment
- Fix snapshot response field name: size_bytes → snapshot_size_bytes
- Include items breakdown in snapshot response
- Fix step counter: move transfer step into forward branch only
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: stop stepper spinner and cleanup port-forward after handoff
stepper.step('Handoff summary...') started a spinner that was never
stopped with stepper.done(), keeping the event loop alive and requiring
Ctrl+C. Also aksPortForwardStop() was only called in the error path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: add /agt/handoff/succession router endpoint for Ed25519-signed succession
The registry's succession API requires the predecessor's Ed25519 signature
over a canonical message. The private key lives in the router's Governance
identity — inaccessible to the CLI.
New endpoint POST /agt/handoff/succession on the router:
- Takes {successor_amid, reason} from CLI
- Looks up predecessor (self) AMID from registry
- Looks up successor signing key from registry
- Signs canonical message 'succession:{pred}:{succ}:{timestamp}'
- Submits complete SuccessionRequest to registry
- Returns registry response
CLI + plugin updated to call /agt/handoff/succession instead of
/agt/registry/registry/succession (which lacked signing keys).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: sub-agent handoff — collect, snapshot, and re-spawn during handoff
Sub-agents spawned by a parent agent are now included in handoff state
transfer. The full pipeline:
Collection (source side):
- New GET /agt/handoff/sub-agents endpoint lists active sub-agents
and reconstructs SpawnRequest from CRD spec (K8s) or container
labels (Docker dev mode)
- CLI + plugin call this endpoint and inject sub_agent_snapshots
into the handoff snapshot payload
Re-spawn (target side):
- handoff_restore iterates sub_agent_snapshots after state hydration
- Calls create_sandbox() for each sub-agent with the stored config
- Returns per-sub-agent results (spawned/failed) in restore response
- Audit-logged as handoff:restore:sub-agent
Supporting changes:
- SpawnRequest, HandoffMeta, SubAgentSnapshot: added Clone derive
- Sub-agent results included in restore response as sub_agent_results
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: AMID remapping + sub-agent workspace collection during handoff
Two improvements to sub-agent handoff:
1. AMID remapping: When re-spawning sub-agents on the target, the old
parent AMID in trusted_peers is replaced with the new parent's AMID.
This ensures sub-agents trust the new parent for KNOCK handshakes.
The new parent AMID is looked up from the registry at restore time.
2. Workspace collection: The CLI now exec's into each sub-agent's
container (kubectl for AKS, docker for local) to collect workspace
tar before including it in the snapshot. Each sub-agent's workspace
is capped at 2MB. The plugin path is best-effort without workspace
(no container exec access from inside the sandbox).
Also adds sub_agents_respawned count to plugin restore metadata.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: collect sub-agent workspace via E2E mesh during handoff
Sub-agents now respond to handoff:workspace_request mesh messages by
tarring their /sandbox/.openclaw/workspace and sending it back via the
E2E encrypted relay. No shared volumes or cross-container exec needed.
Plugin message handler (sub-agent side):
- Receives handoff:workspace_request from parent
- Tars workspace (excludes extensions, node_modules, etc.)
- Sends base64-encoded tar back as handoff:workspace_response
- Falls back to empty response on error (parent doesn't hang)
Plugin handoff orchestration (parent side):
- After fetching sub-agent list from router, discovers each sub-agent
via registry search to get their AMID
- Sends handoff:workspace_request to each via mesh
- Polls agtInbox for handoff:workspace_response (up to 15s per agent)
- Enriches sub-agent snapshots with workspace tar before creating
the encrypted handoff snapshot
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: expand workspace tar to include cron/policies/agents + mesh 768KB cap
- Add WORKSPACE_TAR_CMD constant in handoff.ts for consistent tar commands
- Include .openclaw/cron, .openclaw/policies, .openclaw/agents in all 5 tar commands
- Refine extension exclusion: only exclude */dist and */node_modules (keep manifests)
- Cap mesh workspace response at 768KB (safe under relay's ~1MB limit)
- Add truncated flag in workspace_response messages
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: chunked mesh transfer for large sub-agent workspaces
Workspaces > 512KB are split into chunks sent via separate mesh messages,
then reassembled on the receiver side. This lifts the practical workspace
transfer limit from ~768KB to ~40MB (80 chunks × 512KB).
Sender (sub-agent):
- Small workspace (≤512KB): single handoff:workspace_response (fast path)
- Large workspace: N handoff:workspace_chunk messages + completion marker
- Max 80 chunks (leaves headroom in relay's 100-message offline queue)
Receiver (parent):
- Collects workspace_chunk messages into a Map keyed by chunk_index
- Reassembles in order when all chunks received or completion marker arrives
- 30s timeout with partial-chunk recovery (uses what was received)
- Backwards compatible: single-message responses still work unchanged
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: unified chunked mesh transport layer + file transfer tool
Implements a general-purpose auto-chunking transport layer for the E2E
encrypted mesh. Any payload exceeding 512KB is transparently split into
chunks with per-chunk SHA-256 integrity verification, then reassembled
on the receiver side before reaching application logic.
Transport layer (meshSend + meshHandleTransportMessage):
- meshSend(): auto-chunks large payloads into mesh:transfer_manifest +
N mesh:transfer_chunk messages. Small messages pass through directly.
- meshHandleTransportMessage(): intercepts transport messages in onMessage,
accumulates chunks, verifies SHA-256 hashes, reassembles, and delivers
the original message to the application layer.
- Per-chunk + manifest-level SHA-256 integrity verification
- 2-minute TTL with automatic stale transfer cleanup
- Max ~40MB per transfer (80 chunks × 512KB)
New tool — azureclaw_mesh_transfer_file:
- Agents can send files to each other via E2E encrypted mesh
- Files up to 30MB supported (auto-chunked transparently)
- Received files auto-saved to /sandbox/.openclaw/workspace/incoming/
- Path traversal protection (must be within /sandbox)
Consumers updated to use unified transport:
- mesh_send tool: auto-chunks large task messages
- Handoff blob transfer: auto-chunks encrypted snapshots
- Sub-agent workspace collection: auto-chunks workspace tars
- Handoff re-send loop: uses meshSend for retransmit
Limits raised:
- Router MAX_BLOB_SIZE_BYTES: 50MB → 200MB (sub-agent workspaces)
- Plugin workspace tar cap: 5MB → 50MB
- CLI workspace tar cap: 5MB → 50MB
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: configurable router URL + fix spawn test timeouts + add transport tests
- Make ROUTER and ROUTER_BASE configurable via AZURECLAW_ROUTER_URL env var
(defaults to http://127.0.0.1:8443 for backward compat)
- Fix 5 pre-existing spawn tool test timeouts caused by local dev router
on port 8443 — tests now point to unused port 19876 for immediate
ECONNREFUSED instead of 45s polling loop
- Fix test assertions to match actual error behavior (plain text errors,
not JSON, when router is unreachable)
- Add 11 new tests (198 total, up from 187):
- mesh_transfer_file: registration, schema, path traversal, abs path,
mesh-not-connected
- mesh_send: registration, params, error when disconnected
- handoff_status: registration, returns status JSON
- AZURECLAW_ROUTER_URL: spawn + spawn_status use configurable URL
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: address 4 direction-specific handoff gaps
1. Plugin reverse: retry with 60s backoff when discovering local target
(local agent may be waking from dormant state after CLI runs
wakeDormantDocker)
2. Direction validation: both plugin and router now validate that the
incoming handoff direction matches the environment (AZURECLAW_DEV_MODE).
Warn-only on mismatch — doesn't block to avoid false positives.
3. Forward credential rehydration: collect actual credential VALUES
from Docker (not just refs), create K8s secret on AKS target before
pod starts so envFrom can mount them. Closes the credential gap
where forward handoff lost Telegram/Slack/Brave tokens.
4. Symmetric cleanup: reverse handoff now scales deployment to 0
instead of deleting the CRD. This preserves the sandbox definition
for instant re-forward handoff while freeing all compute resources.
Falls back to CRD deletion if scale fails.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: graceful sub-agent interrupt during handoff
Add handoff:interrupt protocol so sub-agents can save in-progress work
before their workspace is collected during a handoff.
Plugin path (mesh-based):
- Parent sends handoff:interrupt to all sub-agents concurrently
- Sub-agents set handoffInterruptRequested flag
- processTaskWithTools checks the flag between LLM rounds
- On interrupt: saves task progress to .task-in-progress.json
(round, messages, last content, original task)
- Sends handoff:interrupt_ack back to parent
- Parent waits up to 10s for acks, then proceeds with workspace collection
CLI path (exec-based):
- CLI writes .handoff-interrupt sentinel file into each sub-agent container
- processTaskWithTools also checks for this file between rounds
- Same progress save behavior (.task-in-progress.json)
Both paths ensure the workspace tar includes the progress checkpoint,
so it survives the handoff and is available on the target for resumption.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: sub-agent workspace injection + task resumption after handoff
Complete the sub-agent handoff lifecycle — after re-spawn on the target,
sub-agents now receive their workspace and resume interrupted work.
Router changes:
- Restore response now includes sub_agent_workspaces array with each
sub-agent's workspace_tar, task_context, status, and checkpoint
Plugin — target side (post-restore):
- Waits up to 60s for each re-spawned sub-agent to register in mesh
- Sends workspace tar via meshSend (auto-chunked for large workspaces)
- Sends handoff:resume with task context and checkpoint info
Plugin — sub-agent side (new message handlers):
- handoff:workspace_inject: extracts received workspace tar into /sandbox/
with path traversal validation and size guard
- handoff:resume: reads .task-in-progress.json, sends resume_ack to parent
with status report ('Successfully restored in cloud. Resuming interrupted
work from round N: <task>'), then re-enters processTaskWithTools with a
contextual prompt that includes the original task, progress, and last output
The full sub-agent handoff lifecycle is now:
interrupt → save progress → collect workspace → transfer →
re-spawn → inject workspace → resume task → report to parent
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent handoff — Docker API encoding + name prefix + source cleanup
Three bugs prevented sub-agents from properly migrating during handoff:
1. collect_sub_agent_snapshots_docker passed raw JSON braces in the
Docker API URL — curl treated {} as glob patterns and the Docker
daemon couldn't parse the filter. Result: always returned 0
sub-agents, so the snapshot blob had no sub-agents to respawn.
Now uses URL-encoded filter via the docker_api() helper (matching
list_sandboxes_docker's pattern).
2. Same function used the Docker container name (azureclaw-{name})
as the agent name, causing respawn to create
azureclaw-azureclaw-{name} on the target. Now strips the prefix.
3. Source decommission only put the main agent dormant — sub-agent
containers kept running. Now destroys all source sub-agents via
the spawn API before decommissioning.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent handoff — wrong API URLs + missing report steps
Four fixes for sub-agent handoff orchestration:
1. Sub-agent list used GET /sandbox/spawn (wrong) — now GET /sandbox/list
2. Sub-agent delete used DELETE /sandbox/spawn/{name} (wrong) — now
DELETE /sandbox/{name} (matches actual router routes)
3. Response field was 'sub_agents' but endpoint returns 'sandboxes'
4. No sub-agent info in handoff progress report — added _hp() calls
for: discovery count, interrupt/checkpoint status, workspace
collection count, snapshot inclusion, cleanup status, and
sub_agents_transferred in the final result object
Also fixed orphan target cleanup URLs (2 places) that had the same
/sandbox/spawn/{name} → /sandbox/{name} issue.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: set trusted_peers for re-spawned sub-agents after handoff
The trusted_peers remapping in handoff_restore was dead code —
Docker snapshots had trusted_peers=None, so the if-let guard never
fired. Re-spawned sub-agents had no trusted parent AMID and rejected
handoff:workspace_inject + handoff:resume messages from the new
parent. This caused workspace injection to silently fail and task
resumption to never trigger.
Fix: always set trusted_peers to include the new parent's AMID,
regardless of whether the original snapshot had it set:
- If peers existed: remap old parent → new parent (existing logic)
- If peers existed but old parent absent: append new parent
- If peers was None: set to new parent entry (new case)
This ensures the sub-agent trusts the new parent on first KNOCK
and accepts workspace/resume messages immediately after spawn.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat: include sub-agent status in Telegram greeting after handoff
Restructure the restore flow so sub-agent workspace injection and resume
happens BEFORE the Telegram greeting. After sending resume signals, wait
up to 8s for sub-agents to send resume_ack messages. The Telegram greeting
now includes a sub-agent status section showing each agent's name, state
(resumed/ready/starting/failed), and a task preview.
The handoff_ready mesh message back to the predecessor also now includes
sub_agents_restored count, sub_agents_resumed count, and per-agent details.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: use per-sandbox runtime for operator exec path, not global devMode
The operator's unified view fetches both Docker and AKS agents, but the
routerExec/fetchAgtQuick/fetchEgressDomains functions used the global
devMode flag to choose between docker-exec and kubectl-exec. This meant
Docker agents were queried via kubectl (which fails — no K8s namespace)
when the operator ran in non-dev mode, and vice versa.
Fix: use sb.runtime === 'docker' per-sandbox instead of the global devMode
flag in all 4 places that exec into agent containers:
- fetchSecurityState (routerExec + k8sCheck)
- fetchEgressDomains (routerCurl)
- fetchAgtQuick (docker/kubectl exec)
- seccomp profile inference
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: sub-agent trust + workspace logging after handoff
Two fixes for broken parent↔sub-agent communication after handoff:
1. Stale trusted_peers: After handoff, re-spawned sub-agents get new
AMIDs but the parent's parentTrustedAmids set contains old AMIDs.
KNOCK handler rejects all incoming sub-agent messages (score=0 <
threshold=500). Fix: when the plugin discovers a sub-agent's new
AMID during workspace injection, register it in amidToName,
nameToAmid, parentTrustedAmids, and push baseline trust to router.
2. Silent from_value failure: serde_json::from_value for sub-agent
snapshots at POST /handoff/snapshot was wrapped in `if let Ok`
which silently swallowed deserialization errors, potentially losing
all sub-agent workspace data. Changed to match with tracing::warn
that logs the exact error + JSON preview for debugging.
Also adds roundtrip test for SubAgentSnapshot workspace_tar through
the full serialize→compress→encrypt→decrypt→decompress→deserialize
chain, and a test for the JS↔Rust base64 round-trip.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* security: remove clawhub.com and openclaw.ai from default egress allowlist
clawhub.com had a 12% malware rate (Atomic Stealer was #1 skill).
openclaw.ai enables `curl | bash` install vectors from inside sandboxes.
Both were in the default Helm values and example CRD. Removed from:
- deploy/helm/azureclaw/values.yaml (default egress allowlist)
- examples/basic-agent/clawsandbox.yaml (example CRD)
- PLAN.md (policy presets documentation)
Defense is now three layers deep:
1. Egress proxy blocks clawhub.com/openclaw.ai (network)
2. Skills directory is root-owned, chmod 640/750 (filesystem)
3. Plugin code is root-owned, read-only for sandbox (code integrity)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: use sub_agent_results as trust+resume loop driver, not sub_agent_workspaces
Root cause: the post-restore trust registration and resume signals were
gated on restoreResp.sub_agent_workspaces (workspace data), which could
be empty even when sub-agents were successfully spawned. This caused the
entire block to be skipped — no trust registration, no resume signals,
no Telegram sub-agent status.
Fix: use restoreResp.sub_agent_results (always populated when sub-agents
spawn) as the primary loop driver. Workspace data is looked up by name
from sub_agent_workspaces as a secondary source.
Also:
- Added workspace_inject_ack: sub-agent confirms extraction success/fail
with file_count + error back to parent before resume is sent
- Parent waits up to 15s for ack, logs result, passes workspace_delivered
flag in resume payload
- handoff_ready report now includes sub_agents_workspace_delivered count
- Telegram greeting shows 📦 icon per sub-agent when workspace arrived
- Increased mesh registration wait from 60s to 90s (AKS pods need boot)
- Router logs snapshot details when building sub_agent_workspaces
Tests added:
- Rust: sub_agent_workspaces builder filter (empty/non-empty workspace_tar)
- Rust: full encrypt→decrypt→restore round-trip with 2 sub-agents
- Rust: edge case — empty workspace + empty task_context filtered out
- TS: sub_agent_results drives loop even when sub_agent_workspaces empty
- TS: workspace_inject_ack protocol (success + failure paths)
- TS: handoff_ready includes workspace delivery status
- TS: only spawned sub-agents enter trust loop
- TS: missing sub_agent_results graceful fallback
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(sdk): reuse active Signal Protocol session instead of crashing
Vendor patch #10: SessionManager.initiateSession() threw 'Active session
already exists' when the crypto layer had a session established via
incoming KNOCK but client.activeSessions wasn't synced.
This broke mesh_transfer_file and any second send to the same peer.
Fix: return existing session info with reused=true flag instead of
throwing. establishSession detects reuse, syncs activeSessions, and
skips redundant KNOCK/activate.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: don't expose handoff confirmation code to LLM, add console.log diagnostics
Security fix: The handoff confirmation token was returned in the tool
result, allowing the LLM to self-confirm without user input. Now the
code is sent directly to Telegram (side-channel) and printed to console
for TUI users — the LLM never sees it.
Changes:
- Remove confirmation_token from azureclaw_handoff_request tool response
- Send code via Telegram sendMessage as a side-channel delivery
- Print code to console.log (visible in kubectl logs, not to LLM)
- Update tool descriptions to emphasize code comes from user input
- Bump CONFIRMATION_MIN_DELAY_SECS from 3s to 8s
- Add console.log diagnostics in post-restore IIFE for handoff debugging
- Add test: tool response must not contain confirmation_token
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: increase sub-agent tool-calling rounds from 10 to 25
10 rounds was too tight for data-heavy tasks — sub-agents hit the
cap and returned truncated results. 25 gives enough room for
multi-step research/collection while still preventing runaway loops.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: increase prekey retry window from 16s to 45s for sub-agent mesh send
Sub-agents take 20-30s+ after pod is Running to upload prekeys
(gateway start → plugin load → SDK init → relay connect → prekey
upload). The previous 8×2s=16s window wasn't enough, causing
'Cannot get prekeys' failures on task dispatch.
Now: 15 attempts × 3s = 45s max wait with clearer hint message.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: add file_transfer_ack for verified file delivery
The mesh_transfer_file tool had no delivery confirmation — it reported
'sent' but couldn't verify the file was actually written to disk on the
target agent. This caused silent failures where coordinator-notes.txt
appeared to transfer but never landed.
Now:
- Receiver sends file_transfer_ack with success/saved_to/error
- Sender waits up to 15s for ack
- Tool returns 'delivered' (with path) or 'sent_no_ack' (no confirmation)
- Write is verified with stat after writeFileSync
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: filter protocol messages from mesh_inbox, improve file transfer reliability
mesh_inbox now filters out internal protocol messages (handoff blobs,
acks, workspace inject/resume) so the LLM only sees actual sub-agent
replies. Shows filtered_protocol_messages count for visibility.
file_transfer now retries up to 3 times with ack verification — sends,
waits 15s for file_transfer_ack, retries with 3s backoff if no ack.
Returns 'delivered' with exact path or 'sent_no_ack' after all retries.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: auto-decode file_transfer content in mesh_inbox
file_transfer messages now show decoded text content (or a binary
placeholder) instead of raw base64 blobs. Uses null-byte detection
to distinguish text vs binary files. Text files are fully readable
in the inbox; binary files show filename and size.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: stale AMID cache poisoning breaks post-handoff mesh delivery
Root cause: trust+resume loop found OLD sub-agent AMIDs still in the
registry (Docker containers hadn't timed out yet) and cached them.
New AKS sub-agents registered with different AMIDs 26s later, but
the parent never discovered them — all messages went to dead relay
connections and were silently dropped.
Three-layer fix:
1. Stale AMID rejection: collect original_amid from handoff snapshots,
reject matching registry results, wait for NEW AMIDs to appear
2. Prekey readiness gate: verify E2E session is established before
sending workspace_inject (20 attempts × 3s = 60s max)
3. Workspace inject retry: 3 attempts with 20s ack wait each, catches
send errors and retries instead of fire-and-forget
Also filters protocol messages from mesh_inbox and auto-decodes
file_transfer base64 content so LLM sees readable text.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: write HANDOFF_FILES.md manifest after workspace inject
After extracting the workspace tar, writes a manifest listing restored
user-facing files to /sandbox/.openclaw/workspace/HANDOFF_FILES.md.
This makes injected files discoverable when the agent is asked about
its workspace contents post-handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: bump AGT rate limits for multi-agent handoff
Policy: 120 → 240 max_calls/60s for inference:* actions.
Router: 100 → 200 global req/s, 10 → 20 per-agent req/s.
Handoff with 3+ agents doing workspace inject + resume + relay
traffic was hitting the old limits.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: promote incoming/ files to workspace root after handoff inject
Copies files from incoming/ to the workspace root so the agent sees
them immediately when listing files, without needing to know about
the incoming/ directory convention.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: scale down sub-agent deployments during cloud→local decommission
After scaling the parent to 0, also scale down all sub-agent
deployments that were snapshotted. Prevents orphaned sub-agent
pods running on AKS after reverse handoff.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs: add bidirectional handoff architecture diagrams and changelog
- New section 12 in architecture-diagrams.md: agent lifecycle across
handoff, forward/reverse flows with sub-agents, stale AMID cache
poisoning problem & fix, workspace injection detail
- CHANGELOG: add handoff features, sub-agent support, rate limit bump
- README: add handoff to features list and CLI reference table
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 12, 2026
Per docs/implementation-plan.md §5.4. The compat suite is the zero-regression wire (principle §0.2 #1) between existing user-facing flows and the Phase-0→Phase-4 decomposition. Before we split any of the 12 hotspot files or swap to AgtMeshProvider, every protected flow (§5.1, eight total) needs a spec here; after the change, the same spec must still pass. - tests/compat/package.json — vitest-based standalone npm package - tests/compat/tsconfig.json — ES2022 strict typecheck - tests/compat/vitest.config.ts — fork pool for per-spec isolation - tests/compat/README.md — charter, protected flows, authoring rules - tests/compat/harness/types.ts — the 8 protected-flow catalogue - tests/compat/harness/blessed-mock.ts — headless blessed + blessed-contrib surfaces (screen/box/log/list/table/ grid/line/bar/sparkline) with a snapshot() oracle + typeKey() driver - tests/compat/specs/operator-tui.spec.ts — 11 harness-sanity assertions green + 8 it.todo staging Phase 1 render-and-drive tests Locally: cd tests/compat && npm ci && npm test — 11 passed / 8 todo (19). No production code touched. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 12, 2026
Per docs/implementation-plan.md §5.4. Behavioral conformance corpus — the net that catches 'endpoint returned 200 but skipped the crypto step' bugs. tests/conformance/ layout - package.json / tsconfig.json / vitest.config.ts — own workspace, fork-isolated pool so ratchet state doesn't leak across specs. - README.md — corpus table (9 corpora across Phases 0, 1, 3), provider-axis rule (plan §11.3), invariant discipline. - fixtures/README.md — vendored-source policy per principle §0.2 #10. - harness/index.ts — empty scaffold; helpers land per-PR. Phase 0 specs (all 'it.todo', 59 invariants total — vitest reports 59 todo / 0 pass / 0 fail, no false-green): - signal-x3dh.spec.ts (14) — key-exchange shape, symmetric ratchet invariants, base64 input hygiene (vendor patches #3, #4). - signal-knock.spec.ts (15) — KNOCK happy-path, trust threshold, relay disruption, wire-shape parity (vendor patches #5, #7, #8, #1/#2 timestamps, vendored<->AGT byte-identical KNOCK). - signal-negative.spec.ts (13) — ciphertext integrity, replay, session clobber (vendor patch #10), DoS surfaces. - sandbox-isolation.spec.ts (17) — seccomp, Landlock, egress-guard, router-as-only-network-path. Guarded by CONFORMANCE_E2E=1 (requires Kind; compat suite Kind harness wires it in Phase 1). Principle §0.2 #8 ('solid, not look-alike'): it.todo is a documented pending assertion, not a silently-passing no-op. Each spec's top comment cites the vendor patch or production bug it exists to prevent recurring. Principle §0.2 #10: every invariant that references an external protocol cites its upstream source (libsignal, RFC3339 'Z' suffix, Signal Double Ratchet spec) either inline or in the README. ci/no-stubs.sh already allow-lists tests/ subtrees; gate remains PASS. No new dependencies outside the pinned vitest / typescript / @types dev-deps already used by tests/compat/. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 12, 2026
… validator
Implements plan §1.3 (Outage semantics) as a pure, deterministic decision
function. No I/O, no AGT import surface, clock injected.
inference-router/src/providers/outage.rs (new):
- OutageMode enum: Strict | CachedRead | DegradedDev
* serde camelCase; FromStr accepts camel/kebab/snake; Display; Default = Strict
- OutageConfig { mode, cached_ttl }
* validate_for_env(is_dev_env) rejects DegradedDev in prod
* rejects cached_ttl = 0 on CachedRead
* enforces MAX_CACHED_TTL = 15 min
- CachedDecision<T> with is_expired(ttl, now) — backwards-clock = expired
- OutageAction<T>: Deny { mode } | ServeCached { verdict } | AllowWithWarning
- decide_outage(config, cached, now) — pure, test-friendly
- 19 unit tests covering all three modes, cache freshness, clock skew,
serde round-trip, env validation, TTL bounds
inference-router/src/providers/mod.rs:
- remove placeholder OutageMode stub; re-export real types
controller/src/providers/mod.rs:
- mirrored OutageMode (from_spec, is_dev_only, validate_for_env)
- OutageModeError::DegradedDevInProd
- 4 new unit tests
docs/security-audits/2026-04-24-phase1-outage-semantics.md:
- STRIDE, principle-mapping, re-audit triggers; both sign-offs present.
No call-site in the router consumes decide_outage yet — that lands with
the first AGT provider. Landing the pure semantics first locks the
decision rules before any provider wiring pressures them.
Verification:
- cargo test --all: 106+155+15+26+3 = 305 passed (was 286, +19)
- cargo clippy --all-targets --all-features -- -D warnings: clean
- six CI gates PASS on the branch tip
Plan refs: §0.2 #1/#2/#3/#4/#5/#8/#9/#10 | §1.3 | §1.4
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 12, 2026
…hments, sibling trust race, final-deliverable rule Live multi-agent demo (analyst→viz→writer fan-out) surfaced four distinct breakages that each silently degraded the run while the parent agent still self-reported success. 1. Image generation 404 (router URL prefix regression) Foundry's account-scoped /openai/v1/images/generations endpoint does NOT accept the /api/projects/<project>/ URL prefix that chat-completions tolerates. Commit 582eb39 unified everything through that prefix. In dev (raw Azure OpenAI account, no project path) it works; in AKS prod against a Foundry project endpoint the upstream returns a fast 404 and image_generation falls back to written descriptions. Fix: strip /api/projects/<name>/ from the upstream endpoint inside the images_generations handler before forwarding. Add strip_project_prefix helper + four unit tests (project-stripped, trailing-slash variant, AOAI passthrough, no-prefix passthrough). 2. foundry_code_execute drops container files The Responses API tool only walked output[].type=='message' for text. matplotlib PNGs / CSVs generated by code_interpreter live inside Foundry's per-run container and are referenced via container_file_citation annotations or { type: 'image' } entries in code_interpreter_call.outputs. They were never downloaded, so the demo's bar chart silently degraded to ASCII. Fix: collect every (container_id, file_id, filename) reference from both shapes, GET each via the new /openai/containers/... router route (added to foundry_standalone_routes), and write the bytes to /sandbox/.openclaw/workspace/. Append the local paths to the tool result so downstream tools (mesh_transfer_file, file_write) can ship them. Adds routerCallBinary helper for binary downloads through the router. 3. Sibling KNOCK race in parallel fan-out AGT_TRUSTED_PEERS is baked at spawn time and consumed once at sub-agent boot. When the parent spawns analyst → viz → writer in sequence, only writer (last) sees all siblings. analyst's parentTrustedAmids only contains parent — so when viz or writer later try to KNOCK analyst, the trust score is 0 + 0 = 0 and the KNOCK is rejected at threshold 500. The demo logs confirm only 1 of 3 sibling pairs ever opened a session. Fix: after every successful spawn, the parent broadcasts a peers_update message containing the new sibling's AMID to every already-running sibling. Each sub-agent now records the parent's AMID at boot (first AGT_TRUSTED_PEERS entry, by convention) and handles peers_update only from that AMID, extending its parentTrustedAmids set at runtime. 4. Sub-agent system prompt missing FINAL DELIVERABLE rule Sub-agents were told they could mesh_transfer_file artifacts to peers, but nothing forced them to mesh_transfer_file the FINAL artifact back to the parent before returning a summary. The writer's executive_brief.md sat in its local /sandbox forever while the parent reported success. Fix: append a hard rule to the sub-agent system prompt requiring mesh_transfer_file(to_agent='parent', ...) as the last action before any "task complete" reply, with one call per output file. Tests - inference-router: 643 lib tests pass (4 new strip_project_prefix tests). - inference-router: cargo clippy --all-targets clean. - runtimes/openclaw: 118 tests pass; tsc clean; oxlint shows only pre-existing warnings. - cli: 553 tests pass. Deployment - For #2: rebuild + push inference-router image. - For #1, #3, #4: rebuild + push sandbox image. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
…#245) * feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap Wire azureclaw runtime through createMeshTransport() factory so we can flip between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without code changes. Surface additions to IMeshTransport (both adapters now expose): - lookup(amid) — registry RPC for reputation/display name - submitReputation(...) — registry RPC for peer feedback - enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on) - onError(kind, from, detail) — diagnostic hook for decrypt + ws errors - onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal - onDisconnect(reason, code) — ws close / error fan-out mesh-plugin (vendored A adapter): - connection.ts delegates to the underlying SDK; lazy bind for hooks registered before connect() - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the contract so neither adapter can drop a method without CI failing mesh-plugin (AGT B adapter): - agt-transport.ts implements lookup/submitReputation as REST calls to the registry (AGT MeshClient is pure transport — registry RPCs intentionally not added to AGT upstream; they belong on a separate RegistryClient) - enableKnockEnforcement is a no-op (AGT MeshClient always enforces) - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods (added on local AGT branch azureclaw-meshclient-event-hooks, NOT pushed — AGT team owns the upstream PR) runtime (runtimes/openclaw): - Adds @azureclaw/mesh as a file: dependency - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)' when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value - Identity is generated once via vendored SDK regardless of provider, then raw Ed25519 keys are extracted via toData() and shared across both — same AMID either way - Banner now reports active provider (vendored vs agt) Docs: - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity, policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer + the wiring + migration path - Documents the 3 hooks added to local AGT branch and the 3 governance methods kept adapter-side Tests: - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat) - runtimes/openclaw: 118/118 pass - AGT (local branch): 387/387 pass with 8 new event-hook tests Open work for cleanup phase: once AGT publishes the version with our event hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var toggle. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): pre-build mesh-plugin for runtime CI + format reconciler PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): quote workflow step name containing colon YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): clippy fixes for Rust 1.95.0 CI runs Rust 1.95.0 which added new clippy lints: - doc_lazy_continuation: indent doc list items that span multiple lines. Added two-space indent to the trailing 'All three are populated...' paragraph so it is treated as a continuation of the preceding list item rather than its own malformed list item. - obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as if .. { a } else { b } per the lint suggestion. These were pre-existing on dev (CI only started failing once the runner picked up Rust 1.95.0); fixing here so PR #245 can land green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12 Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5 The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): mark gaps G1 and G2 fixed on local AGT branch Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass. * dev: add --mesh-provider <vendored|agt> selection with first-run prompt Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker mode before AKS rollout. Same flag, three integration points: CLI (cli/src/commands/dev.ts): - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball - First-run interactive prompt offers AGT only if the toolkit checkout is actually present locally — silently defaults to vendored otherwise (no pestering for users without AGT cloned). - --build branch: builds the right relay/registry images vendored → vendor/agentmesh-relay + agentmesh-registry (Rust) agt → agent-governance-python/agent-mesh/docker/Dockerfile with COMPONENT=relay / registry build-args - Sandbox image build: stages locally-packed AGT SDK tarball into .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL build-arg (auto-discovers if --agt-sdk-tarball not given). - Runtime branch: skips Postgres for AGT (in-memory registry), uses correct ports (AGT: 8083 relay, 8082 registry; vendored: 8765/8080) and health path (AGT: /healthz; vendored: /v1/health). - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the runtime transport-factory honors the user's choice. Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile): - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the sandbox npm-installs the local tarball instead of fetching the published @microsoft/agent-governance-sdk from npm. Lets us smoke-test the locally-patched AGT branch (G3/G4 fixes) end to end without round-tripping through npm publish. - .agt-sdk/ staging dir always exists (with .keep) so the COPY never fails when the user didn't stage a tarball. Defaults preserved: --mesh-provider=vendored, existing behavior is byte-identical for users who don't opt in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so tsc failed with TS2307 in the AGT build path. Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the build context, and strip its 'prepare' script (which would invoke tsc, not present in this stage; the pre-built dist/ is sufficient). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks). - connect() now passes autoRegister: true so the SDK uploads identity and prekeys instead of the adapter re-implementing that path with raw HTTP. - discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT. - lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}). - submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score clamped to [0,1]; the vendored /registry/feedback endpoint does not exist in AGT. - Replaced mapAgent with pickDisplayName helper: AGT puts display name in metadata.display_name (set by registerSelf), with the first capability as the fallback. Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and the bespoke fetchWithRetry helper — both are upstream concerns now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes the vendored registry wire shape. The vendored impl talks to /registry/* (the existing agentmesh-registry); the AGT impl talks to /v1/discover and /v1/agents/{did} on the upstream AGT registry. Both expose a single normalized RegistryEntry envelope, so callsites stay readable. getMeshRegistry(routerUrl) is the entry point. Provider selection follows AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base). Migrated all raw-HTTP registry callsites: - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName, resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid. - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS spawn discovery, AKS→local discovery. - core/agt-task-loop.ts (1 site): registry_capability_search tool. - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe, azureclaw_discover (mesh_discover) tool. - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe, heartbeat keepalive (no-op under AGT — relay does liveness via WS). The discover-on-router-unreachable test now asserts the new contract: empty list + count:0 instead of a 'Discovery failed' string. Registry hiccups must not break tool calls. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): wire AGT provider end-to-end (6 stackup bugs) End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced six bugs blocking the upstream AGT MeshClient swap. All fixed: 1. Final sandbox Docker stage didn't COPY mesh-plugin, so the file:../../mesh-plugin symlink dangled in node_modules. Plugin swap silently fell back to vendored with 'Cannot find package @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist} into /mesh-plugin/ in the final stage of the Dockerfile. 2. entrypoint.sh used cp -r when copying node_modules into the plugin extension dir, preserving the (now-broken-at-runtime-path) symlink. Switched to cp -rL so symlinks dereference into real files in the target tree. 3. mesh-plugin/src/index.ts imported createMeshTransport from ./transport-factory.js but never re-exported it. Runtime swap path couldn't find the factory. Added the missing re-export. 4. inference-router agt_registry_proxy unconditionally prepended '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents' became '/v1/v1/agents' at the upstream. Now: forward verbatim when path starts with 'v1/' or equals 'health', else prepend. Preserves vendored SDK behavior ('registry/register' → /v1/registry/register). 5. /agt/relay route only matched the bare path, but AGT MeshClient appends '/ws' to relayUrl. Added /agt/relay/ws route and made the upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt. 6. agt_registry_proxy route was declared get(...).post(...) only. AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey upload and DELETE for deregister — both 405'd at the router. Added .put() and .delete() to the route declaration. Bug #6 was invisible to vendored because the vendored SDK only ever uses GET/POST (registry/register, registry/prekeys, etc.). AGT's switch to REST verbs exposed the gap. Path allowlist also extended with 'v1/' prefix so AGT's REST paths (v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation. Verified end-to-end via azureclaw dev --mesh-provider=agt --build: - POST /v1/agents → 201 Created - PUT /v1/agents/{did}/prekeys → 200 OK - WebSocket /ws accepted, stable connection (no reconnect loop) - Plugin reports 'AGT mesh connected' + provider=agt Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): always route mesh registry through inference-router azureclaw_discover and other mesh registry callsites went via `process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`, intending to let out-of-sandbox sub-agents bypass the router. In practice, the sandbox launcher always sets AGT_REGISTRY_URL as the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082 in dev, the K8s service URL in prod). Since the runtime runs as UID 1000 and iptables egress-guard blocks UID 1000 from anything except localhost+DNS, the direct upstream URL ECONNREFUSEs and the catch-all silently returns []. Symptom: registered agents are invisible to azureclaw_discover even though they show up in `GET /v1/discover` when queried directly at the registry. Drop the env-var override — there's no in-sandbox runtime path where bypassing the router is correct. The router is the ONLY way out for UID 1000. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): break mesh_send infinite poll loop on dead sub-agent probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but routerCall actually resolves with the parsed JSON error body. When the sub-agent pod/container is gone the router returns 404 with { error: "Container '<name>' not found..." } and probeSubAgentAlive read status.phase = undefined → defaulted to "Unknown" → not in POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every 2s forever, blocking the LLM event loop ("LLM not responding" symptom). Also narrow the prekey transient retry test so permanent X3DH / signature-verification failures bubble up instead of being treated as "waiting for prekeys" and retried indefinitely. Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'. Before: registry log fills with GET /v1/discover?capability=echo-buddy every ~2s forever; LLM stops responding to new turns. After: mesh_send aborts with 'sub-agent sandbox not found' on first probe. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): suppress /v1/registry/* 404 leaks in AGT mode Three vendored-only registry paths were being called unconditionally in AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply cadence: 1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X. The AGT registry exposes GET /v1/discover?capability=X instead — display names live in the per-agent record, so the AGT path fans out to a second /v1/agents/{did} fetch per discover hit. Driven by the operator panel's /agt/reputation polling. 2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session. AGT has no per-session counter; per-agent reputation already submitted via MeshClient.submitReputation. No-op in AGT mode. 3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode. Also includes complementary debugging fixes from this session: - agt-transport: auto-call establishSessionWithPeer() before send() so AGT mode gets vendored-equivalent send-with-first-contact semantics. Without this, send() throws 'No encrypted session — call establishSession() first' and the retry loop spins forever. - cli operator fetchers: add 8–10s timeouts to kubectl get calls that were hanging when the cluster API was unreachable. cargo check: clean runtimes/openclaw: 118 vitest tests pass inference-router: 8 mesh tests pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode Fourth 404 leak revealed after deploying the previous fixes: the operator panel's ~30s /agt/reputation poll triggers governance::agt_reputation, which (after lookup_parent_amid succeeds) fetched the per-agent reputation score via the vendored-only GET /v1/registry/reputation/score?amid=X path. AGT registry has no such endpoint — the score is embedded as 'reputation_score: f64' in the per-agent record returned by /v1/agents/{did}. Provider-dispatch the URL; for AGT, wrap the agent record in a vendored-shaped payload (score / tier / raw) so downstream CLI fetchers and the operator panel stay schema-agnostic. cargo check: clean agt_governance_integration: 26/26 pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * runtime: hide Foundry tools in github-copilot mode (same as github-models) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest Phase B.1 of the AGT-on-AKS rollout (see session plan files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the existing `azureclaw dev --mesh-provider` flag so the same provider selection works for AKS pushes. When --mesh-provider=agt: * Builds relay+registry from the AGT upstream Dockerfile ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile) using COMPONENT=relay|registry build-args (matches dev.ts). * Tags as agentmesh-{relay,registry}-agt:latest so both vendored and AGT images can coexist on the same ACR and so the existing deploy/agentmesh-agt.yaml manifest picks them up unchanged. * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz) into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg. * Always passes MESH_PROVIDER build-arg to the sandbox image so the Dockerfile's conditional `npm install @microsoft/agent-governance-sdk` runs for AGT clusters. When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml, applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt, THEN rolls the controller (so the new pod reads AZURECLAW_MESH_PROVIDER=agt for new sandboxes). Auto-reverses when --apply --mesh-provider=vendored runs against a cluster currently on AGT (no Postgres deployment in the agentmesh ns). The image build loop also now supports absolute Dockerfile paths and absolute build contexts via a new `absoluteContext` field, needed because the AGT Dockerfile lives outside the azureclaw repo root. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks without rebuilding any images, assuming both image pairs were already seeded by 'azureclaw push'. Flow: 1. Detect current provider via 'kubectl get deploy/postgres -n agentmesh' (vendored has Postgres, AGT does not). 2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found 3. kubectl apply -f deploy/agentmesh-<target>.yaml 4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target> 5. kubectl rollout restart deploy/azureclaw-controller 6. With --restart-sandboxes: roll every azureclaw-managed Deployment so existing pods pick up the new AZURECLAW_MESH_PROVIDER value. Service names and ports are identical between the two manifests (agentmesh-relay:8765, agentmesh-registry:8080) so the controller's mesh_peer talks to either stack with no further config — the relay/ registry URLs already come from env vars (MESH_RELAY_URL / MESH_REGISTRY_URL). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(up): --mesh-provider=agt picks AGT manifest + flips helm value Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up' so first-time deploys can ship AGT instead of vendored. When --mesh-provider=agt: * helm install runs with --set mesh.provider=agt (controller env AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes). * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of deploy/agentmesh.yaml. * Skips the postgres ACR import and the agentmesh-db-credentials secret creation (both unused by AGT — its registry is in-memory). * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml vs .tmp-agentmesh.yaml) so concurrent provider switches don't collide. The deployAgentMesh signature gains a non-breaking 'meshProvider' option that defaults to 'vendored' (existing callers untouched). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): plumb --mesh-provider into local-k8s helm install Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(controller): AGT wire protocol adapter for mesh_peer Implement full AGT relay/registry wire support in the controller's mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt. Without this the controller's federation peer cannot connect to the AGT relay (different WS path, frame envelope, heartbeat, ack model) or the AGT registry (different HTTP shape, no signed body), and the leader fails-loops on AGT clusters — breaking the only cloud-offload control path. New module `mesh_peer/agt_wire.rs`: - `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error) with `#[serde(tag="type", rename_all="snake_case")]` matching `agentmesh/relay/app.py`. - `AgtRegisterAgentRequest` struct for `POST /v1/agents`. - 7 unit tests pinning the serialized shape. `mesh_peer/mod.rs`: - New `Provider` enum + `Provider::from_env()` selecting vendored (default) or AGT off `AZURECLAW_MESH_PROVIDER`. - `MeshPeerState.provider` carried through outbound + inbound paths. - `register_with_registry()` branches: vendored signs ts body; AGT posts `{did, public_key (base64url), capabilities, metadata}` with no signature; 409 treated as success for leader-failover idempotency. - `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>` (matches JS SDK `buildDid`), so every leader replica converges on the same DID without coordination. - Default `MESH_RELAY_URL` appends `/ws` for AGT. - `connect_and_listen()`: - AGT connect frame `{type:"connect", from:<did>, token?:<env>}` (token read from `AGENTMESH_RELAY_TOKEN` if set). - AGT has no `Connected` ack — mark `connected=true` immediately. - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored keeps `ping`). - `serialize_and_send_outbound()` / `send_to_peer()` now take `state` and branch outbound framing — AGT emits `message` frames `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex). - `handle_message()` dispatches to `handle_vendored_frame()` or `handle_agt_frame()`. AGT path: - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`. - Sends `Ack` reply (required — without it AGT redelivers on reconnect → duplicate offload processing). - Treats `Error` frames mentioning Authentication failed / Missing 'from' / session_replaced as fatal — drops connection for reconnect. `mesh_peer/offload.rs`: - All 8 `send_to_peer(...)` call sites updated to pass `&state` first. `main.rs`: - Remove the temporary AGT-skip guard around `mesh_peer::run`. The peer now starts unconditionally when enabled; provider is consumed inside `mesh_peer::run`. Build/test: - cargo build --release --package azureclaw-controller: OK - cargo test --package azureclaw-controller: 492 passed - cargo clippy --package azureclaw-controller --all-targets -D warnings: OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer - cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(deploy): AGT mesh probe path + Cilium pod-port NP allow deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz (see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness probes were 404'ing → CrashLoopBackOff/NotReady. operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane evaluates NetworkPolicy egress against the backend pod port (post-DNAT), not the Service port. AGT registry/relay listen on 8082/8083; the Service maps 8080->8082 and 8765->8083 so the Service-port allowlist (8080/8765) doesn't actually permit the post-DNAT flow. Add 8082/8083 alongside so both vendored (8080/8765 direct) and AGT paths work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh promote): AGT-compat health + WS upgrade paths azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): first-run picker for local vs remote mesh source azureclaw dev now asks new users where the mesh should live, just like the existing inference-provider picker: Where should the mesh live? ❯ Local (recommended; spin up relay + registry in Docker) Remote (auto port-forward to AKS cluster: <cluster-name>) Local (default) keeps the existing behaviour: docker-compose'd relay/registry/postgres on the user's laptop. Remote (advanced) federates with a previously-provisioned AKS mesh: - If ~/.azureclaw/context.json has a cached globalRegistryUrl from a prior 'azureclaw mesh promote', reuse it verbatim. - Otherwise default to http://localhost:18080 — the port-forward URL 'mesh promote --port-forward' uses — so the auto-promote fallback in the downstream global-registry block will spawn the tunnels on demand. - If there is no aksCluster in context at all, warn and fall back to local so the user isn't left with a broken sandbox. Skipped entirely when --global-registry was passed explicitly (the advanced flag overrides the prompt) or when the user is past their first run. Also fixed a latent AGT-compat bug in the same flow: the existing 'auto-promote' path probed only /v1/health, which 404s on AGT clusters. Replaced with a /health → /v1/health fallback (matches the same shape we used in checkRegistryHealth last commit). Verified: - npm run build / typecheck clean - 640 CLI tests pass (2 skipped, no regressions) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply back to their spawner, but on AGT the registry has no agent named or capability="parent" — the search returns 0 → no prekey bundle → send fails. The vendored runtime had this aliased only in the offload-mode task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the controller never set for AKS-spawned children. Two coordinated fixes: 1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set (spawner seeds 'parent_name:parent_AMID' as the first entry), also push PARENT_SANDBOX=<first_name> into the openclaw container env. 2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send and azureclaw_mesh_transfer_file, alias to_agent=='parent' → PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry lookup. The Symbol is set during runtime init from AGT_TRUSTED_PEERS[0], so this works even on images built before fix #1 lands. Skip in offload mode — 'parent' there is a protocol-level routing token, not a mesh recipient name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap mesh-plugin/src/agt-transport.ts.send() called this.client.establishSessionWithPeer(toAmid) before forwarding to client.send(). That method does not exist on AgentMeshClient — the real method is establishSession(toAmid, options) — so every parent → sub-agent send on AGT was failing with: establishSessionWithPeer is not a function It was also unnecessary: AgentMeshClient.send() already auto-bootstraps the X3DH handshake on first contact (see @agentmesh/sdk AgentMeshClient.send → cache miss → establishSession() fallthrough at dist/index.js:3321-3334). Calling establishSession() ourselves would also be wrong because it is not idempotent — it unconditionally writes activeSessions.set and starts a fresh X3DH. Fix: remove the pre-bootstrap entirely and let client.send() manage session lifecycle. The AgtSdkModule type loses the required establishSessionWithPeer member (now optional) since we no longer depend on it; test fakes remain valid as harmless extras. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call Previous commit 34662c7 wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * push: auto-detect mesh provider from live helm release When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * entrypoint: fail-open trust gate when running anonymous tier When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable The AGT registry has no autonomous presence model — `last_seen` is frozen at registration and the `update_last_seen()` store method is dead code with no HTTP handler calling it. Combined with the openclaw discover tool's 90s stale filter (agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes silently invisible 90s after spawn, breaking sibling-to-sibling peer discovery. Demo symptom: analyst/viz/writer all reported 'peer discovery did not return ...' even though mesh_send to those names succeeded with 'delivered_and_replied'. The relay was fine; only the registry's presence view was stale. Pair with the corresponding upstream registry change (AGT branch `azureclaw-meshclient-event-hooks`, commit adds POST /v1/agents/{did}/heartbeat -> store.update_last_seen). The new tick reuses the existing 30s relay-keepalive timer in connect(), so no extra timers and no extra event-loop pressure. Best-effort: 4xx/5xx are warned-once, network errors swallowed, loop survives a registry pod restart (next tick retries). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 13, 2026
…D) (#283) Adds mtime-poll watchers for both /etc/azureclaw/inference and /etc/azureclaw/memory mount directories, mirroring the existing governance::Governance::spawn_policy_watcher pattern. Closes Slice 3 DoD #4 (router reloads within 5s of kubectl edit) explicitly and Slice 2 DoD #1 implicitly (the echo loop now closes Compiled→Ready on every change, not just the first). - spawn_inference_policy_watcher (INFERENCE_POLICY_WATCH_INTERVAL, default 5s) - spawn_memory_binding_watcher (MEMORY_BINDING_WATCH_INTERVAL, default 5s) - load_and_install now clears the handle on NoBinding/NoPolicy and preserves it on Error — proper hot-reload semantics for when an operator removes spec.memoryRef / spec.inferenceRef. - Both watchers wired in main.rs after governance. - 779 router lib tests (+5). Co-authored-by: Patrik Lagi <palagi@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced May 13, 2026
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 13, 2026
Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the existing singular mcpServerRef, which is now deprecated and honored as a length-1 alias via effective_mcp_server_refs(). Controller-side changes: - New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET in status/conditions.rs::reason. - Reconciler mirror loop refactored to iterate effective_mcp_server_refs(). Singular-field use emits tracing::warn with McpSingularDeprecated. len > 1 short-circuits via degrade! macro until Slice 4d.2 wires per-server addressing — principles §3 honest 'not-yet-enforced' signal. - 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty + plural-wins-when-both-set. Controller suite: 555 passing. Admission CEL on deploy/helm/azureclaw/templates/crd.yaml: - Mutex: singular and plural cannot both be set. - maxItems: 8 (router-side scheme is sized for this). - Per-name uniqueness across mcpServerRefs. Out of scope for 4d.1 (queued for 4d.2): - Per-server jwks-{name}.json / tools-{name}.json file scheme. - Router-side McpServerRegistry + namespaced tool dispatch. - Stale-file sweep (DoD #6). - e2e fixture with ≥ 3 servers (DoD #1). Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 13, 2026
) (#292) * Slice 4d.1 — mcpServerRefs plural CRD field + admission CEL Closes Slice 4 DoD #2 (mcpServerRef singular deprecation). Adds GovernanceConfig.mcpServerRefs (Vec<LocalObjectRef>) alongside the existing singular mcpServerRef, which is now deprecated and honored as a length-1 alias via effective_mcp_server_refs(). Controller-side changes: - New constants MCP_SINGULAR_DEPRECATED + PLURAL_MCP_SERVERS_UNSUPPORTED_YET in status/conditions.rs::reason. - Reconciler mirror loop refactored to iterate effective_mcp_server_refs(). Singular-field use emits tracing::warn with McpSingularDeprecated. len > 1 short-circuits via degrade! macro until Slice 4d.2 wires per-server addressing — principles §3 honest 'not-yet-enforced' signal. - 6 new unit tests: shim precedence (3 cases) + camelCase + omit-when-empty + plural-wins-when-both-set. Controller suite: 555 passing. Admission CEL on deploy/helm/azureclaw/templates/crd.yaml: - Mutex: singular and plural cannot both be set. - maxItems: 8 (router-side scheme is sized for this). - Per-name uniqueness across mcpServerRefs. Out of scope for 4d.1 (queued for 4d.2): - Per-server jwks-{name}.json / tools-{name}.json file scheme. - Router-side McpServerRegistry + namespaced tool dispatch. - Stale-file sweep (DoD #6). - e2e fixture with ≥ 3 servers (DoD #1). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Slice 4d.2 — per-server McpServer mounts + router discovery (DoD #1 + #6) Closes Slice 4 DoD #1 (≥3 plural McpServers reachable e2e) at the mount-and-discovery layer, and DoD #6 (stale-file sweep) via reconciler-driven volume rebuild. Multi-JWKS OAuth + namespaced tool dispatch (DoD #3) follow in Slice 4d.3. Controller: - reconciler/mod.rs: replace len>1 short-circuit with full iteration over effective_mcp_server_refs(). Per-name volumes (mcp-jwks-<name>, mcp-signing-<name>) mounted at /etc/azureclaw/mcp/<name>/ and /etc/azureclaw/mcp-signing/<name>/. - First entry (idx 0) keeps legacy MCP_JWKS_PATH + MCP_SIGNING_KEY_DIR env vars for backwards compat with current single-JWKS OAuth path. - New MCP_JWKS_DIR=/etc/azureclaw/mcp env set once per pod. - governance_mounts.rs: new inject_container_env helper for idempotent env-var injection without a volume mount. - Removed obsolete PLURAL_MCP_SERVERS_UNSUPPORTED_YET reason constant. - CRD doc comment updated for 4d.2 mount layout + 4d.3 forward-ref. Router: - New mcp/registry.rs with scan() + discover_from_env(). At startup, reads MCP_JWKS_DIR, enumerates subdirs with parseable jwks.json, emits tracing::info!(servers=?, count=N) per discovery + warn! per skipped candidate. - Empty/missing dir handled gracefully (sandbox with zero mcpServerRefs is a valid steady state). - 7 registry unit tests + 4 inject_container_env unit tests. Stale-file sweep (DoD #6): reconcile rebuilds pod-spec from current refs; removed refs disappear via SSA. No explicit sweep code needed. Verified: 559 controller + 804 router tests pass, clippy -D warnings clean across workspace, cargo fmt --check clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 13, 2026
Operator-facing surface for the BlockedBuffer the forward proxy has been populating since S12.f. Closes Slice 5 DoD #1. Producer side (inference-router): - New BlockedBuffer::snapshot_since(since_unix) and top_hosts(since,n) methods. snapshot_since sorts newest-first by last_seen_unix; top_hosts aggregates by hostname across (sandbox, port) and secondary-sorts deterministically by host name for tied counts. - 7 unit tests cover filter cutoffs, dedup count carry-through, multi-sandbox/multi-port aggregation, n=0 early return, and the truncate-to-n contract. Wire surface (routes/internal.rs): - GET /internal/egress/blocked?since=<rfc3339|unix|-Nm> - GET /internal/egress/blocked/top?window=<duration>&n=<int> Both mounted on the admin-gated 'protected' router. JSON envelopes carry schema_version: 1 and RFC 3339 strings alongside raw Unix seconds. Hand-rolled duration + RFC 3339 parsers (no chrono dep) with 7 unit tests + 6 integration tests covering bare seconds, s/m/h/d suffixes, relative -Nm form, malformed input → 0, the n≤100 cap, and the default 5m window. CLI (azureclaw egress blocked <sandbox>): - New subcommand under 'egress' so --watch/--top/--since don't collide with the existing flat options surface. - Mirrors azureclaw inspect token-resolution (router-admin-token secret first, in-pod admin-token file fallback) and uses in-pod kubectl exec curl to avoid port-forward collisions. - --watch loops every 5s with VT clear; --top overrides --since (window is its own filter); --json emits raw response. - 13 vitest unit tests for buildPath, renderers, and unixToIso(0). Tests: 849 router lib + 8 egress_blocked integration (up from 2) + 13 CLI unit. cargo clippy -D warnings + cargo fmt + npm typecheck/build/test/lint all clean. Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced May 13, 2026
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 16, 2026
…atency overclaims Second pass of the OSS-readiness audit, focused on the architecture diagrams (the first pass missed factual errors here). architecture-diagrams.md: - Diagram #3: audit record is hash-chained, append-only (NOT signed today — cryptographic signing of the chain head is on the roadmap, see security.md). Tamper-detection vs tamper-proof is a real distinction. - Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list), prose says nine + mentions ClawPairing as the controller-internal 10th. - Diagram #7: CRD relationship arrows were wrong against the actual Rust structs. Corrected: - policyRef → spec.governance.toolPolicyRef - mcpRefs → spec.governance.mcpServerRefs - inferenceRef → spec.inferenceRef (top-level, was already correct) - memoryRef → spec.memoryRef (top-level, was already correct) - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy. - CE 'sandboxRef' → spec.targetSandboxRef. - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected to every sandbox by the controller (no ref). Noted in prose. - EgressApproval added (it was missing from the diagram entirely); links to ClawSandbox via spec.sandbox (string name, not ref object). security.md: - Headline #1 'agent does not see Azure credentials. Period.' — softened with a dev-mode footnote matching the README hero. In azureclaw dev, agent+router share a container with separate UIDs but a kernel-level container escape defeats the boundary; the hard guarantee is the AKS path. Anchor points at architecture.md#two-modes. - Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation latency on the router hot path.' Microsecond was an overclaim without a benchmark to back it. architecture.md: - Controller row in the components table: 'watches the eight peer CRDs' → 'nine peer CRDs (plus controller-internal ClawPairing)'. All 9 mermaid blocks pass a bracket-balance sanity check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 16, 2026
…#325) * docs: OSS-readiness pass — strip internal jargon + correct overclaims Doc + comment-only audit across user-facing docs ahead of OSS launch. No code paths touched. Factual corrections (the heaviest changes): - docs/security.md headline guarantee #4: drop 'signed by the router' claim. The audit log is hash-chained and detects modification but is not signed today; signing is on the v1.1 roadmap. - docs/security.md Layer 6 inference-safety table: * Content Safety: 'Always on, server-side' → 'Always on for Foundry-provider requests; Copilot/GitHub-Models providers do not return prompt_filter_results'. * Token budget: was 'not yet aggregated'; the router DOES aggregate per-tenant daily and monthly UTC counters with on-disk persistence (inference-router/src/budget.rs:200-260). * New 'Operator escape hatches' subsection documents the two env-var knobs (AZURECLAW_SUPPRESS_CONTENT_FLAGS, AZURECLAW_CONTENT_FLAG_MIN_SEVERITY) honestly. - README hero: 'agent never sees an Azure key' softened to call out that this is the AKS guarantee; dev mode co-locates agent + router in one container. - README/architecture: '31 commands' → '30+ commands'; remove unverified '18 Foundry API groups' count. - docs/architecture.md design goal #1 + #4: explicit dev-vs-prod scoping; 'same code path' → 'same data-path code with documented AZURECLAW_DEV_MODE branches'. - docs/architecture/a2a-gateway.md: port 8445 is config-locked but the mTLS listener itself is still being wired; operators should set A2A_GATEWAY_UPSTREAM_URL explicitly until the listener is GA. - mesh-plugin/src/agt-identity.ts comment: replace 'encrypted at rest with per-host KEK' claim with honest 'chmod 0600 is the real boundary' note (mirrors PR #324 identity-store fix). - .github/copilot-instructions.md: '__AGT_INITIALIZED env guard' → 'Symbol.for(agt-mesh-client)' (matches current code in runtimes/openclaw/src/index.ts:458-480). Internal-jargon strip in user-facing docs: - docs/api/lifecycle.md: remove 'Slice 4', 'Slice 4d.3/4d.4', 'Slice 0', 'Slice 1c', 'Slice 2a/2b/2c/2d.1', 'Slice 2d.2', 'Slice 3a' references — replaced with descriptive prose. - docs/api/conditions.md: drop 'Slice 1c invariant' and 'principles.md §3' references. - docs/architecture/agt-boundary.md, docs/security-mcp-top10.md, docs/cli-reference.md: drop 'Phase 5.2' shibboleth — keep the facts (vendored AgentMesh fork was retired upstream). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: deep-dive diagram audit — fix CRD field-name labels + signing/latency overclaims Second pass of the OSS-readiness audit, focused on the architecture diagrams (the first pass missed factual errors here). architecture-diagrams.md: - Diagram #3: audit record is hash-chained, append-only (NOT signed today — cryptographic signing of the chain head is on the roadmap, see security.md). Tamper-detection vs tamper-proof is a real distinction. - Diagram #6: 8 CRDs → 9 CRDs (EgressApproval was missing from the list), prose says nine + mentions ClawPairing as the controller-internal 10th. - Diagram #7: CRD relationship arrows were wrong against the actual Rust structs. Corrected: - policyRef → spec.governance.toolPolicyRef - mcpRefs → spec.governance.mcpServerRefs - inferenceRef → spec.inferenceRef (top-level, was already correct) - memoryRef → spec.memoryRef (top-level, was already correct) - A2A 'sandboxRef' arrow was fake — A2AAgent has no sandboxRef. Real link is A2A -> ToolPolicy via spec.policyRefs.toolPolicy. - CE 'sandboxRef' → spec.targetSandboxRef. - TG 'trustRef' was fake — TrustGraph is cluster-scoped and projected to every sandbox by the controller (no ref). Noted in prose. - EgressApproval added (it was missing from the diagram entirely); links to ClawSandbox via spec.sandbox (string name, not ref object). security.md: - Headline #1 'agent does not see Azure credentials. Period.' — softened with a dev-mode footnote matching the README hero. In azureclaw dev, agent+router share a container with separate UIDs but a kernel-level container escape defeats the boundary; the hard guarantee is the AKS path. Anchor points at architecture.md#two-modes. - Layer 7: 'Sub-µs evaluation latency' → 'sub-millisecond evaluation latency on the router hot path.' Microsecond was an overclaim without a benchmark to back it. architecture.md: - Controller row in the components table: 'watches the eight peer CRDs' → 'nine peer CRDs (plus controller-internal ClawPairing)'. All 9 mermaid blocks pass a bracket-balance sanity check. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 16, 2026
…diagrams (#327) Link each provider bullet in 'Pluggable inference backend' to the relevant architecture chapter and diagram: - GitHub Copilot → architecture.md#dev-mode + diagram #1 (dev mode pod) - Foundry / Azure OpenAI → architecture.md#prod-mode + diagrams #2 (prod mode pod) and #3 (data path) - GitHub Models → security.md (Content Safety caveat) + architecture.md data-path section (provider routing notes) Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 28, 2026
Wire the previously-orphaned sidecar_client module into the inference-router and harden it against token substitution attacks. CONTROLLER (reconciler/mod.rs): - Extends agent_id_active tuple to carry tenant_id alongside agent identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via ProvisioningOutcome::Ready.auth_spec. - Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode sandboxes. Without it, router-side tid pinning is disabled and warns at boot (insecure, dev-only). ROUTER (sidecar_client.rs, auth.rs, lib.rs): - pub mod sidecar_client; — fix the orphaned-module bug. - WorkloadIdentityAuth now consults SidecarClient first; sidecar mode is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves per-sandbox audit attribution in downstream Azure RBAC. - from_env() now Result<Option<Self>>. Partial config (only URL or only PINNED set) returns Err; WorkloadIdentityAuth::new() panics → AKS surfaces as CrashLoopBackoff. No more silent fallback to a different identity model. - HARD-fail validate_token_claims on EVERY check (rubber-duck #1): * tid mismatch / missing (cross-tenant guard) * appid/azp mismatch when present (audit attribution guard) — if both present, BOTH must match * aud mismatch against per-service expected set (cache-poisoning guard) — Foundry, Graph, OpenAI, Management, Search mapped * exp in past or within 60s skew (stale-token guard) * exp missing when tid-pinning enabled (unbounded-lifetime guard) - Returns a TTL cap = exp - now - 60s; caller takes min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching. - JWT decoder tightened: requires EXACTLY 3 segments (rejects 2-segment unsecured JWS and 4-segment JWE). - aud claim normalized: accepts both single-string (Entra default) and array (RFC 7519); rejects out-of-spec shapes (number, bool). - Soft WARN when both appid and azp are absent (some MSI tokens legitimately omit both). - New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning state for operator visibility. TESTS: - 39 sidecar_client unit + wiremock integration tests, covering all HARD-fail branches, the cache-skip on validation failure, the TTL cap math, the partial-config Err path, JWT decoder edge cases (2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud, out-of-spec aud). - Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client. Rubber-duck review caught: (1) appid/azp should be hard not soft, (2) partial env config must fail closed, (3) 3-segment requirement, (4) aud validation against requested resource, (5) exp validation + TTL cap. All five addressed in this commit. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 29, 2026
* docs(entra-agent-id): import POC findings as architecture reference
Captures the end-to-end token-acquisition flow validated on real AKS
in the Microsoft tenant during the POC phase.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): scaffold Entra Agent ID auth machinery
Adds controller-side primitives for operating kars sandboxes as
per-sandbox Entra Agent Identities. Foundation only — pod-spec
integration lands in a follow-up commit.
New modules:
- auth_config.rs: KarsAuthConfig CRD (cluster-scoped singleton).
Tenant + blueprint + controller MI anchors plus optional
serviceManagementReference for Microsoft-style tenants.
- auth_config_reconciler.rs: materialises kars-auth-sidecar-env
ConfigMap with AzureAd__ and DownstreamApis__ env vars. Idle when
the CRD is absent (cluster stays in anonymous tier).
- agent_identity.rs: Graph client for per-sandbox agent identity SP
lifecycle. Implements IMDS to MI to blueprint to Graph chain
proven during POC.
- sidecar_injection.rs: pure-function pod-spec helpers (sidecar
container shape, pinned-identity env vars, egress-guard iptables).
CRD additions in crd.rs:
- KarsSandbox.spec.meshAuth.mode (auto/agent-id/anonymous)
- KarsSandbox.status.agentIdentity
786 controller tests pass, 14 new tests added.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli): auto-provision Entra Agent ID trust during kars up
Makes Entra Agent ID setup invisible to end users: when `kars up`
runs and the cluster does not yet have a `KarsAuthConfig/default`
resource, the new `mesh/agent_id_setup.ts` module idempotently
provisions the tenant trust anchor end-to-end:
1. Blueprint app via Graph (POST /v1.0/applications/ with
@odata.type=#Microsoft.Graph.AgentIdentityBlueprint),
including the optional serviceManagementReference required by
Microsoft-style enterprise tenants.
2. Blueprint service principal (visible in the Entra Agents portal).
3. Controller managed identity in the customer's subscription.
4. MI-as-FIC on the blueprint (issuer=login.microsoftonline.com),
the anti-loop-safe credential path proven by the POC.
5. KarsAuthConfig CR written to the cluster.
The step is non-fatal: if the user lacks `Agent ID Developer`,
`kars up` continues and the cluster runs in anonymous tier until
the role is granted and `kars mesh setup-trust` is rerun. This
matches the three-tier fallback model documented in
docs/architecture/entra-agent-id/.
New `--service-tree <guid>` flag on `kars up` (and KARS_SERVICE_TREE
env var) propagates the ServiceTree GUID to the blueprint creation.
No hardcoding — tenants that do not require it leave it empty.
Tests: 6 new unit tests for agent_id_setup (idempotence detection,
dry-run, env-var threading, error propagation). 775 CLI tests pass
(2 pre-existing skipped). typecheck + lint clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs+preflight: user-facing Entra Agent ID guidance
Adds the user-facing documentation layer that the previous controller
and CLI commits were missing. Every doc that mentioned the old
`api://agentmesh` flow now points to the new per-sandbox Entra Agent
ID model.
New docs:
- docs/agent-identity.md — day-1 and day-2 user guide. Prerequisites,
walk-through of what `kars up` does for auth, sub-agent semantics,
troubleshooting, teardown.
Updated docs:
- README.md — top-line mention of per-sandbox Entra Agent ID in the
`kars up` summary, linking to the new guide.
- docs/getting-started.md — Step 2.1 calls out the `Agent ID
Developer` role prerequisite; Step 2.2 expands the bring-up list to
include the preflight role check and the Entra trust provisioning.
- docs/permissions.md — rewrites the "Tenant-level (Entra ID)
considerations" section to describe the new model. Replaces
`api://agentmesh` failure rows with Entra-Agent-ID-specific ones
(CredentialInvalidLifetimeAsPerAppPolicy,
InvalidFederatedIdentityCredentialValue, missing sidecar).
- docs/cli-reference.md — documents the new `--service-tree` flag
+ Microsoft-corp example.
- docs/SUMMARY.md — wires agent-identity.md into the mdbook index.
New preflight check:
- cli/src/preflight.ts — calls `checkAgentIdRole` from agent_id_setup.
Warns (not blocks) when the signed-in user lacks the `Agent ID
Developer` directory role. The existing `api://agentmesh` warning
line is removed.
- cli/src/commands/mesh/agent_id_setup.ts — new exports:
- `checkAgentIdRole` returns hasRole/inconclusive/message via
Graph `/me/transitiveMemberOf`. Matches by role template id
(stable) AND display name (forward-compatible).
- `detectExistingBlueprint` returns whether the configured
blueprint already exists in Graph.
- `AgentIdSetupOptions.blueprintName` added so multi-cluster
deployments can share a tenant-wide blueprint.
Tests: 781 CLI tests pass (+6 new for checkAgentIdRole / blueprint
detect). Typecheck + lint clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(cli): UX polish — from-scratch resets context, surface CA block
Two papercuts surfaced when the user ran `kars up --from-scratch` on
a tenant that already had a previous deployment:
1. `--from-scratch` cleared the resume state but NOT the cached
deployment context, so the "fresh" run silently reused the prior
region/RG/Foundry endpoint instead of re-prompting. Fixed in
`cli/src/commands/up/preflight.ts` — when `fromScratch=true`, skip
the `loadContext()` prefill entirely and print a clear
"ignoring any cached deployment context" line. Adds the
`fromScratch?` field to `UpOptionsForPreflight`.
2. The Entra Agent ID preflight check correctly soft-failed on the
well-known AADSTS530084 Conditional Access token-binding block
(common in Microsoft-corporate tenants), but the warning was
generic and gave the user no actionable next step. Now detects
AADSTS530084 (and the related AADSTS65001/65002 missing-consent
codes) specifically and surfaces the exact `az login --scope
https://graph.microsoft.com//.default` workaround inline.
Docs: adds an `#az-cli-ca-block` anchor section to
`docs/agent-identity.md` so the inline preflight message can link
directly to the troubleshooting paragraph.
Tests: 2 new test cases in agent_id_setup.test.ts pinning the
AADSTS530084 + AADSTS65001 detection paths. Total CLI tests:
783 pass (+2 vs prior commit).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(controller): correct system namespace kars-system (was azureclaw-system)
Caught while auditing residual `azureclaw-` references after the
Azure/kars rebrand. The auth-config reconciler materialised the
sidecar env ConfigMap into `azureclaw-system`, but the actual
controller + helm chart deploy everything into `kars-system`. With
the wrong namespace the ConfigMap was unreachable from sandbox pods
even when the rest of the wiring was correct.
Two-line fix:
- controller/src/auth_config_reconciler.rs:61 — namespace constant.
- controller/src/auth_config.rs:43 — doc comment.
The rest of the controller already uses `kars-system` consistently
(pairing_reconciler, trust_graph_reconciler, signer_policy,
egress_approval_reconciler). This brings the new modules in line.
786 controller tests still pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli): kars mesh setup-trust --mode agent-id (standalone retry)
The Entra Agent ID auto-provisioning that `kars up` runs is now
also reachable as a standalone command — necessary for retrying
after a transient failure (e.g. AADSTS530084 Conditional Access
block on Microsoft Graph) without re-running every other `kars up`
phase.
cli/src/commands/mesh/setup-trust.ts grows a `--mode <agent-id|legacy>`
flag, defaulting to `agent-id`. Agent-id mode forwards to the same
`ensureAgentIdTrust` helper used by `kars up`. Legacy mode preserves
the original api://agentmesh app-registration flow for installations
that haven't migrated yet (slated for removal once all consumers
have switched).
Also surfaces:
- --service-tree <guid> for Microsoft-style tenants
- --cluster-name / --resource-group / --region for controller MI
override
- Clear AADSTS530084 + "Agent ID Developer missing" remediation
messages on failure.
Bonus: cli/src/commands/up.ts agentmesh-image import block from a
prior in-flight edit — imports agentmesh-relay-agt + agentmesh-
registry-agt from the public source ACR in --build mode so the
agentmesh deploy step does not block on ImagePullBackOff. (Future
improvement: build them locally from .agt-sdk when the AGT SDK
tarball is present.)
Tests + typecheck clean: 783 CLI tests pass, build green.
* chore(helm): install KarsAuthConfig CRD
Adds the helm template for the new cluster-scoped singleton CRD
introduced in c9ce68f. Without this, `helm upgrade` does not install
the CRD and `kars mesh setup-trust --mode agent-id` fails with
"the server doesn't have a resource type karsauthconfig" when it
tries to `kubectl apply` the CR.
The schema uses `x-kubernetes-preserve-unknown-fields: true` on
spec/status with required-field validation for the top-level
tenant / agentId / controller blocks. The canonical type definitions
remain in controller/src/auth_config.rs; the controller validates
spec shape on reconcile. A future PR will add karsauthconfig to the
existing helm-vs-Rust drift test in controller/src/helm_drift.rs.
helm lint clean. YAML parses.
* feat(cli+bicep): auto-fallback to Bicep when az CLI Graph is CA-blocked
End-to-end UX for Microsoft-corporate-style tenants where the Azure
CLI cannot acquire a Microsoft Graph token (AADSTS530084) — the same
block we hit repeatedly during the POC phase.
What changes:
- deploy/bicep/agent-id-trust.bicep — sub-scope Bicep template that
provisions everything the imperative path does: blueprint app + SP
(tagged EntraAgentId), controller MI, MI-as-FIC on the blueprint
using login.microsoftonline.com (universally allow-listed). Goes
through ARM + the Microsoft.Graph extension, which has its own auth
path and is not subject to az CLI CA token-binding policy.
- deploy/bicep/modules/controller-mi.bicep — RG-scope module for the
controller MI (Bicep needs RG scope for UAMI, parent runs at sub
scope to create the RG).
- deploy/bicep/bicepconfig.json — enables the Microsoft.Graph
extension (preview).
- cli/src/commands/mesh/agent_id_setup_bicep.ts — driver that runs
`az deployment sub create` against the template, parses outputs,
and writes the KarsAuthConfig CR.
- cli/src/commands/mesh/agent_id_setup.ts — adds
ensureAgentIdTrustAutoFallback() which tries the fast Graph REST
path first and transparently switches to the Bicep path on
AADSTS530084. Other Graph errors propagate unchanged (don't mask
real permission failures with a Bicep retry).
- cli/src/commands/up.ts — uses the auto-fallback wrapper so
`kars up` "just works" in any tenant.
- cli/src/commands/mesh/setup-trust.ts — same auto-fallback for
`--mode agent-id`, and a new `--mode bicep` for users who want to
skip the CLI attempt entirely.
- Also makes blueprintName configurable so multi-cluster deployments
in the same tenant share one blueprint by default.
Tests: 783 CLI pass, typecheck clean, helm-lint + bicep-build clean.
Operational note (not code): the AGT registry/relay images the user
manually built+pushed during this session were from a feature branch
(copilot/secure-mcp-agent-governance) that adds proof-of-possession
to /v1/agents — incompatible with SDK 3.7.0 which doesn't sign yet.
Rebuilding from tag v3.7.0 + restarting the deployments resolves the
422 flood. Documented for future kars releases to build from a known
AGT tag rather than HEAD.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(bicep): clean error reporting + linter-clean issuer URL
Three small fixes shaken out by the user's first end-to-end run of
`kars mesh setup-trust --mode bicep` in the Microsoft corporate
tenant:
1. deploy/bicep/agent-id-trust.bicep — the MI-as-FIC issuer was
hardcoded to `login.microsoftonline.com`. Bicep linter emits
`no-hardcoded-env-urls` as a warning on stderr, which the CLI
driver was mistaking for a deployment failure. Switched to
`environment().authentication.loginEndpoint` so the template is
cloud-portable AND the linter is silent.
2. cli/src/commands/mesh/agent_id_setup_bicep.ts — error capture
reworked to use execa's merged `all` stream so the real ARM
deployment error surfaces instead of being masked by stderr-only
linter warnings. Lines starting with `WARNING:` are stripped from
the error summary.
3. Same module — the final `kubectl apply KarsAuthConfig` is now
correctly recognised as a separate step from the Bicep deployment.
When the CRD isn't installed (e.g. older controller helm release),
the user sees:
"Bicep deployment succeeded, but KarsAuthConfig CRD is not
installed. All Entra resources are already created — just
install the CRD and re-run."
plus the exact `helm upgrade` command. The function returns the
Bicep result so callers know the Entra side is complete.
Also briefly probed the Microsoft.Graph Bicep extension's beta
channel (1.0.0) to see if it supports the `@odata.type` discriminator
for typed AgentIdentityBlueprint. It does not (`BCP037 The property
"@odata.type" is not allowed on objects of type
"Microsoft.Graph/applications"`). The Bicep path therefore creates a
regular Application tagged `EntraAgentId` — functional for runtime
but not visible under the Entra Agents portal page. Documented as a
limitation; users who need portal visibility for the typed form must
use Graph Explorer or PowerShell (both have different first-party
auth paths that aren't CA-blocked).
Bicep build clean, typecheck clean, 783 tests pass.
* fix(cli+docs): switch Graph calls from /v1.0 to /beta
User reported that the typed POST creating an
`#Microsoft.Graph.AgentIdentityBlueprint` works in Graph Explorer
against the /beta endpoint — and that's what makes the resulting
app visible in the Entra Agents portal. The /v1.0 endpoint accepts
the same body but does NOT route through the typed-resource
discriminator on the server, so the app ends up listed only under
App registrations.
Updates `cli/src/commands/mesh/agent_id_setup.ts` to use
/beta/applications, /beta/servicePrincipals, /beta/users, /beta/me,
and /beta/applications/{id}/federatedIdentityCredentials in every
Graph REST call. The `@odata.type` body field stays
`#Microsoft.Graph.AgentIdentityBlueprint` — only the URL prefix
changes.
docs/permissions.md updated to match (the manual escape-hatch
snippet showing `az rest --method POST --url
https://graph.microsoft.com/v1.0/applications/...` now reads `/beta/`).
Note: this only helps when the imperative Graph REST path runs
successfully. In tenants where the az CLI is Conditional-Access-
blocked from Graph (AADSTS530084), the CLI auto-falls back to the
Bicep ARM path. The Bicep Microsoft.Graph extension does NOT
support the @odata.type discriminator at v1.0 or beta (1.0.0), so
the Bicep-created app is functional but stays in the tag-based
detection mode (App registrations only, not Agents portal). This
is a current limitation of the extension, documented in
agent_id_setup_bicep.ts.
Tests: 783 CLI pass, typecheck clean.
* feat(cli): device-code re-login + exact-match Graph body for typed blueprint
Two coordinated changes to make the imperative Graph REST path
succeed in tenants with Conditional Access token-binding policy
(AADSTS530084) on the az CLI's first-party app:
1. Match Graph Explorer's exact working body shape:
- @odata.type value WITHOUT the leading `#` ("Microsoft.Graph.
AgentIdentityBlueprint", not "#Microsoft.Graph.…"). Both forms
are accepted by /v1.0; only the unprefixed form survives the
/beta route reliably per user testing.
- sponsors@odata.bind / owners@odata.bind URLs use /v1.0/users/.
Graph rejects /beta/users/ as an @odata.bind target.
- POST URL stays /beta/applications/ — the typed
AgentIdentityBlueprint resource only surfaces in the Entra
Agents portal page when created via the /beta route.
2. Auto device-code re-login on AADSTS530084:
- When `az rest` returns AADSTS530084, the helper does a one-shot
`az login --use-device-code --scope https://graph.microsoft.com//.default`
which goes through a different OAuth flow that often bypasses
the token-binding CA policy applied to the default interactive
flow.
- User is prompted in the terminal: "visit https://microsoft.com/
devicelogin, paste this code". After login, the original Graph
call is retried once. If it still fails, the
ensureAgentIdTrustAutoFallback wrapper proceeds to the Bicep
ARM path (untyped but functional).
Updates existing test fixture for checkAgentIdRole so the
device-code retry path is mocked. 783 CLI tests pass.
Together these should let Microsoft-corporate-tenant users get the
typed AgentIdentityBlueprint that surfaces in the Entra Agents
portal page, without needing to fall back to Bicep (which can only
produce the tag-based form).
* docs(agent-identity): document AADSTS530033 + Graph Explorer workaround
User hit the next layer of CA policy after device-code re-login:
`AADSTS530033` — "device must be Intune-managed" — applies to
both interactive and device-code flows of the Microsoft Azure CLI
first-party app. Bicep ARM path keeps working (different auth
surface), but it produces an untyped blueprint that only shows
under App registrations, not the Entra Agents portal page.
Updates the existing AADSTS troubleshooting section with:
- Table of auto-handled error codes + what the kars CLI does
- Explicit "even Bicep cannot produce the typed form" path:
Graph Explorer PATCH to upgrade the untyped blueprint app to
the typed AgentIdentityBlueprint discriminator in place
- Note that runtime is unaffected by typed-vs-untyped — only
portal categorisation differs
- Long-term Intune-enrolment recommendation
* fix(cli): kars mesh setup-trust short-circuits when already provisioned
Two bugs surfaced when the user re-ran setup-trust after the
initial successful Bicep + Graph Explorer flow:
1. `karsAuthConfigExists()` checked for "karsauthconfig/default" in
kubectl's `-o name` output, but newer clusters return the
fully-qualified form `karsauthconfig.kars.azure.com/default`.
The bug caused the existence check to always return false, so
the wrapper always tried the Graph REST path even when the trust
was already in place. Fix: match the invariant `/default`
suffix.
2. `kars mesh setup-trust --mode agent-id|bicep` did not consult
`karsAuthConfigExists()` at all, so every re-run triggered the
full provisioning attempt — which, in CA-blocked tenants, kicks
off a device-code login prompt the user has no reason to deal
with when nothing needs to be done. Fix: short-circuit at the
top of both modes when the CR already exists, with a clear
message and the `kubectl delete + retry` escape hatch for users
who genuinely want to re-provision.
Tests: 784 CLI pass (+1 new for the FQ-name kubectl output).
* fix(bicep-fallback): print full Graph Explorer PATCH for portal visibility
The Bicep `Microsoft.Graph` extension cannot set `@odata.type`, so the
Bicep-created blueprint is a plain `Application` that the kars runtime
uses fine but that the Entra portal's Agents page does not show. The
existing docs explained the workaround as a one-field PATCH that
sets `@odata.type` only — that does upgrade the type but the entry
*still* stays hidden in the portal because the Agents-page filter
also requires `sponsors` and `owners` to be set.
Changes:
- agent_id_setup_bicep.ts: at the end of every Bicep success path
(and also on the CRD-missing soft-failure branch), print a fully
populated Graph Explorer PATCH body the user can paste verbatim.
The body includes `@odata.type` + `sponsors@odata.bind` +
`owners@odata.bind` so the resulting typed blueprint actually
shows up under Entra portal → Identity → Agents.
We try `az ad signed-in-user show --query id -o tsv` to auto-fill
the user OID. In the very tenants where this matters (Microsoft
corp Macs without Intune enrollment) that command also fails with
AADSTS530084 — handled silently with a `<YOUR_USER_OID>` placeholder
and a one-line hint on where to find the OID in the Entra portal.
- docs/agent-identity.md: replace the misleading one-field PATCH
example with the full body, including notes about: no `#` prefix on
`@odata.type` in the request body, `/v1.0/users/` required in
`@odata.bind` even when the parent URL is `/beta/`, and a pointer
to the CLI's auto-generated copy-paste body.
Tests: 786 CLI pass (incl. 15 agent_id_setup tests).
* docs+cli: full delete+recreate runbook for typed blueprint (SP+FIC included)
A user hit the case where the in-place @odata.type PATCH was rejected
and they recreated the blueprint via Graph Explorer's POST /applications.
The recreated app then lacked an SP (so it couldn't receive RBAC role
assignments and stayed invisible in the Agents portal) and lacked the
MI-as-FIC (so the controller couldn't mint child identities).
Bicep creates all three resources (app + SP + FIC) but the Bicep
Microsoft.Graph extension cannot set the @odata.type discriminator
needed for a typed agentIdentityBlueprint — so Graph-Explorer recovery
remains the only path in CA-blocked tenants, and it must do all three
steps explicitly.
- agent_id_setup_bicep.ts: extend printPortalVisibilityHint() to print
the full 5-step recovery runbook (POST app, POST SP, POST FIC, kubectl
patch CR, optional DELETE old app) underneath the simpler in-place
PATCH path that's still tried first.
- docs/agent-identity.md: replace the one-line "delete + recreate"
hand-wave with the full POST sequence with all body shapes, plus the
kubectl jsonpath one-liner to extract tenantId + MI principalId from
the cluster.
Tests: 786 CLI pass.
* fix(controller): own egress-guard script in one place + fix iptables rule order
The original `agent_id_egress_rules()` shipped on `feat/entra-agent-id`
contained a security bug: every rule used `-A OUTPUT` (append). The
pre-existing baseline egress-guard script (currently emitted as a
`concat!` literal in `reconciler/mod.rs`) starts with
`-A OUTPUT --uid-owner 1000 -o lo -j ACCEPT`, so any later appended
rule blocking UID 1000 → 127.0.0.1:8080 would NEVER fire — the
loopback-allow would match first. This would have silently broken the
"agent cannot impersonate the router and mint downstream tokens"
boundary as soon as sidecar injection is wired up.
Rubber-duck critique caught this (finding #5). Fix:
- Refactor `agent_id_egress_rules()` to emit two correctly-positioned
rules. The sidecar-block uses `-I OUTPUT 1` (insert at chain head)
so it runs BEFORE the baseline loopback-allow. The router-IMDS
block can stay `-A` because no prior `--uid-owner 1001` rule exists.
- Drop the redundant UID 1000 → IMDS rule (catch-all `DROP UID 1000`
already covers it).
- Drop the explicit UID 1002 → IMDS ACCEPT (OUTPUT chain default is
ACCEPT and no UID 1002 restriction exists).
- New `build_egress_guard_command(agent_id_mode: bool)` composes the
full shell script: agent-id rules first (so `-I` semantics + script
text order both keep the security boundary), then the seven
baseline iptables lines, then a mode-specific echo. `&&`-chained so
any iptables failure aborts init-container startup (partial policy
is worse than no policy).
- Reconciler/mod.rs replaces the `concat!` literal with a call to the
new helper (currently always passing `false`; agent-id pass that
flips to `true` lands in a follow-up commit on the same branch).
Tests:
- New `egress_rules_use_insert_before_baseline_loopback_allow` — pins
the `-I OUTPUT 1` semantics. Direct regression for the security bug.
- New `egress_guard_command_legacy_mode_matches_existing_behaviour` —
byte-for-byte pin on the seven historical iptables lines so the
refactor is provably no-op for non-agent-id sandboxes.
- New `egress_guard_command_agent_id_mode_prepends_security_rules` —
asserts the sidecar REJECT appears BEFORE the loopback ACCEPT in
the script text (defence-in-depth even if a reader misreads -I).
- New `egress_guard_command_is_shell_safe_chained` — every step
starts with `iptables ` or `echo `.
789/789 controller tests pass.
* feat(controller): per-sandbox agent identity provisioning + sidecar injection
The end-to-end controller-side of agent-id mode. New module
`agent_id_provisioning` orchestrates the full flow that the rubber-duck
critique identified as the missing glue: resolve mesh-auth mode →
provision (or recover) the per-sandbox Entra Agent Identity via Graph
→ patch sandbox status → materialise the per-namespace sidecar env
ConfigMap → return a Ready outcome that the sandbox reconciler uses
to inject the sidecar container, pin the router env, and flip the
egress-guard into agent-id mode.
Architecture: model (D) from the critique — the controller provisions
the per-sandbox identity BEFORE pod creation, status-pins the appId,
and the router uses that pinned appId in every sidecar request. No
per-sandbox Secret; no sandbox-side workload identity binding.
## What this commit adds
### New module: `controller/src/agent_id_provisioning.rs`
- `ProvisionerCache` — process-wide cache of `AgentIdentityClient`s
keyed by blueprint client ID. Shares token caches and connection
pools across concurrent sandbox reconciles.
- `resolve_mesh_auth_mode` — pure-function 3-way resolution (Auto →
AgentId or Anonymous based on KarsAuthConfig readiness; explicit
AgentId without ready config surfaces a distinct `AuthConfigNotReady`
reason so operators can distinguish "tenant not set up" from
"auto-fallback to anonymous").
- `load_auth_config` — singleton fetch with explicit `Ok(None)` for
the 404 (anonymous-tier fallback) vs. `Err` for transient failures.
- `ensure_agent_identity_for_sandbox` — idempotent orchestration with
the three-step recovery flow the critique specified:
1. If `status.agentIdentity` recorded, GET via Graph; reuse on
200, reprovision on 404, requeue on 5xx.
2. If status empty, `list_cluster_agent_identities` filtered by
the `kars-sandbox-uid:<uid>` tag. Catches the
"Graph create succeeded but controller crashed before status
patch" crash window.
3. Otherwise create new, patch status, return.
- `materialise_sidecar_configmap` — copies the rendered sidecar env
into the sandbox namespace (`envFrom` cannot cross namespaces — the
critique caught this). Owned by the KarsSandbox via
`ownerReferences` so K8s garbage-collects on sandbox deletion.
### Modified: `controller/src/reconciler/mod.rs`
- `Context` gains `cluster_uid` (read from `kube-system` ns metadata
at startup — canonical "this cluster" identifier in K8s; falls back
to `KARS_CLUSTER_UID` env or generated string with a warning) and
`agent_id_cache: Arc<ProvisionerCache>`.
- `reconcile` calls `ensure_agent_identity_for_sandbox` BEFORE pod-spec
assembly. Match on outcome: `Skipped` → legacy path; `Ready` →
capture identity for downstream injection; `Failed` → patch
Degraded status with `AgentIdentityProvisioningFailed` reason and
requeue. No silent fallback to legacy on AgentId failure — explicit
user intent must not be downgraded.
- Pod-spec assembly:
- Egress-guard `command` flips to `build_egress_guard_command(true)`
when `agent_id_active.is_some()` — adds the security-critical
`-I OUTPUT 1` REJECT rule for UID 1000 → sidecar.
- Router `env` gains `PINNED_AGENT_IDENTITY_APP_ID` (the
per-sandbox appId) + `AUTH_SIDECAR_URL` (loopback :8080).
- Sidecar container is appended to `containers` after the
`runtimeClassName` block. Image pinned to the GA Microsoft
distroless build (overrideable via `KARS_SIDECAR_IMAGE`).
### Modified: `controller/src/auth_config_reconciler.rs`
- After successful ConfigMap apply, `patch_ready_status` patches
`phase=PHASE_READY` + `SidecarConfigMaterialized=True` condition.
This is the real readiness signal that the sandbox reconciler's
`resolve_mesh_auth_mode` gates on (per critique #7). Best-effort:
status-patch failure is logged but doesn't fail the reconcile.
- `build_condition_blueprint_ready` no longer `#[allow(dead_code)]` —
it's now consumed by `patch_ready_status`. Type renamed from
`BlueprintReady` (Graph-call check) to `SidecarConfigMaterialized`
(controller-side materialisation check) since the BlueprintReady
signal will come from a separate cluster-health probe.
## Test status
- 795/795 controller tests pass (+6 new for the new module).
- phase_taxonomy_guard pre-existing integration test passes (the
status patch uses `PHASE_READY` constant, not a string literal).
- Sidecar-injection tests still all pass (the iptables refactor from
the previous commit holds).
## Not yet in this commit (deferred to follow-up commits/PRs)
- Inference-router `sidecar_client.rs` that actually consumes the
pinned env vars (todays AgentIdentity unused on the router side).
- CLI changes: `kars up` VMSS-MI assignment; `kars mesh setup-trust
verify` end-to-end audit; Foundry RBAC assignment print.
- `agent_identity_reaper` for cleaning up orphan SPs.
- Sponsor user object IDs in KarsAuthConfig spec (currently empty
array passed; works in tenants where sponsors aren't required).
* fix(e2e): make Entra Agent ID chain work end-to-end on real AKS
Concludes a long live-debug session against kars-aks where we walked
the full controller → sidecar → router auth chain and fixed every
blocker until the sidecar successfully mints tokens from Entra and
relays them to the router. The remaining gate at end-of-chain is the
Microsoft corp tenant's Conditional Access policy on agent identities
(AADSTS53003 with a `capolids` claims challenge) — a tenant
configuration issue, not a code issue. The architecture is proven
correct.
- **Drop the GET-by-id verify on recorded identity.** Graph's
`GET /servicePrincipals/{id}` has multi-second eventual-consistency
after creation; treating a 404 as "stale, reprovision" creates a
runaway-creation loop (live cluster produced 70+ duplicate SPs per
minute before this fix). The reaper (separate PR) handles
out-of-band deletes via tag-scoped listing.
- **Idempotency guard on `patch_sandbox_status`.** Skip the SSA patch
when the recorded status already matches — same `lastTransitionTime`
drift pattern that bit the auth-config reconciler.
- **Drop ownerReference on the per-namespace sidecar CM.** KarsSandbox
lives in `kars-system`; the per-sandbox CM lives in `kars-<name>`.
K8s rejects cross-namespace owner refs with `OwnerRefInvalidNamespace`
and GC's the CM seconds after creation (debugged via the kubelet
events stream). Cleanup happens via namespace deletion instead.
- **Add apiVersion+kind to status patch body.** SSA patches without
the top-level type meta return `BadRequest: invalid object type:
/, Kind=`.
- **IMDS-first, WI fallback.** Workload-Identity-derived tokens cannot
be used as FIC assertions (Entra anti-loop AADSTS700231). The
controller MUST acquire its MI token via IMDS — which requires the
MI assigned to the AKS node-pool VMSS (`kars up` does this).
- **Permissive Graph list parser.** The agent identity list response
uses `agentAppId`, not `appId`. The strict-schema deserialiser
failed silently on every list call, so the tag-based recovery path
never found anything and the controller created a fresh SP every
reconcile. Fallback parser handles both field names.
- **Patch `phase=Ready`** so `resolve_mesh_auth_mode` has a real
signal to gate on (rubber-duck critique #7). No-op when the
observed status already matches the desired — avoids the SSA
`lastTransitionTime` reconcile loop.
- **`AUTH_SIDECAR_URL=http://localhost:8080`** not `http://127.0.0.1:8080`.
The Microsoft Entra SDK sidecar's HostFiltering middleware only
allows `Host: localhost`; calls with `Host: 127.0.0.1` are rejected
with `400 Bad Request - Invalid Hostname`. /etc/hosts in every pod
maps `localhost` → `127.0.0.1` so the loopback semantics are
identical.
- **`httpHeaders: Host=localhost:8080`** on the readiness/liveness
probes. The kubelet sends the pod IP as the default Host header,
which the sidecar rejects with the same 400. Without this override,
every sidecar stays ready=false even when fully functional.
- **emptyDir at /app/keys.** The sidecar's ASP.NET Data Protection
writes encryption keys to `/app/keys`; the distroless image has
/app owned by root and UID 1002 cannot mkdir there. Per-pod
emptyDir gives a writable scratch.
- Append the keys volume to `pod_spec.volumes` whenever the sidecar
is injected. Without this the container immediately crashes on
startup with `Access to the path '/app/keys' is denied`.
The Microsoft Entra SDK sidecar's response shape varies across
builds. We've observed three forms in the wild:
1. `{"AuthorizationHeader": "Bearer xxx", "ExpiresIn": 3600}` —
documented contract.
2. `"Bearer xxx"` — JSON-quoted string only.
3. `Bearer xxx` — plain text, no JSON.
The strict parser broke on shape 2/3 with `parse auth-sidecar JSON
response`, silently disabling agent-id auth for the entire pod. New
`parse_sidecar_body()` handles all three lenient + 5 dedicated tests.
Without this rule the controller's IMDS call times out (CILIUM
default-deny drops 169.254.169.254). IMDS is per-node link-local —
not externally routable, no exfiltration risk.
- KarsAuthConfig.spec.agentId.sponsorUserObjectIds — required in
tenants where Graph rejects agent identity creation without a
sponsor (live error: `No sponsor specified`).
- KarsSandbox.status.agentIdentity — appId/objectId/displayName/
createdAt. Without this in the CRD schema, SSA patches fail with
`field not declared in schema`.
- Grant the controller `get/list/watch/create/update/patch/delete`
on `karsauthconfigs` (and `/status`). Without this the auth-config
reconciler stays dormant ("KarsAuthConfig CRD not installed").
- 795/795 controller tests pass (+ 0 net change; the existing
sidecar_injection tests covered the URL constants we updated).
- 890/890 inference-router lib tests pass (+5 new
`parse_sidecar_body*` regression tests).
- agent_identity_reaper module (cleans Graph orphans by tag — needed
long-term to handle out-of-band deletes; for now duplicates require
a manual cleanup script).
- CLI flow: `kars up` should run `az identity federated-credential
create` for the controller SA so the WI fallback works in clusters
without VMSS-attached MI.
- Per-sandbox Foundry RBAC automation — verify already prints the
exact `az role assignment` commands; landing this would auto-grant.
```
✅ KarsAuthConfig provisioning (Bicep + Graph Explorer fallback)
✅ Controller MI assigned to AKS VMSSes
✅ Controller IMDS → MI → blueprint Graph token chain working
✅ Per-sandbox agent identity creation via Graph (with sponsors)
✅ Sandbox status.agentIdentity correctly recorded
✅ Per-namespace sidecar env ConfigMap materialized
✅ Sandbox pod contains 3 containers (openclaw + router + auth-sidecar)
✅ auth-sidecar healthy + listening on localhost:8080
✅ Router detects sidecar mode and routes auth through it (fail-closed)
✅ Sidecar mints Entra tokens for the pinned agent identity
✅ Sidecar relays token (or Entra error) back to router
⚠️ Foundry call returns 401 AADSTS53003 — Microsoft corp tenant's
Conditional Access policy blocks agent identities. This is the
same family of CA wall the user has been hitting all session for
their human account. Tenant configuration issue; the chain itself
is provably correct.
```
* fix(sidecar): bind on all interfaces so kubelet probes can reach /healthz
The Microsoft Entra SDK sidecar was bound to `127.0.0.1:8080` only,
which meant the kubelet's readiness/liveness probes (which connect to
the pod IP, not loopback) got `connection refused` at the TCP layer
before any HTTP exchange could happen. Overriding the probe's `Host`
header is useless when the connect itself fails.
Fix: bind on all interfaces (`http://+:8080`). Three defence-in-depth
layers keep cross-pod traffic out of the sidecar:
1. Sidecar's HostFiltering middleware accepts only `Host: localhost`;
anything else gets 400 Bad Request.
2. The sandbox NetworkPolicy has no ingress rule allowing port 8080,
so cross-pod packets are dropped before reaching the sidecar.
3. Egress-guard iptables rule REJECTs UID 1000 → 127.0.0.1:8080,
blocking in-pod privilege escalation from the agent container.
Verified live on kars-aks (sandbox kars-testrun):
- auth-sidecar: ready=true, restarts=0
- kubelet probe now succeeds via pod-IP connect + Host-header
override
- cross-pod connections are still rejected (no NP ingress rule)
This was the last container-level blocker keeping the chain from
running end-to-end. With this fix the sidecar mints real Entra tokens
and the router relays them to Foundry (where the only remaining gate
is the per-agent-identity Azure AI User role assignment).
* refactor(controller): pivot to shared-sidecar architecture (Phase 0)
This commit prepares the branch for the shared-sidecar redesign:
deletes the per-pod sidecar injection plumbing while preserving the
controller-side provisioning machinery, the Graph client, and all
hard-won bug fixes.
## What's removed
- `controller/src/sidecar_injection.rs` — per-pod container spec
helpers (build_sidecar_container, build_router_pinned_identity_env,
build_router_sidecar_url_env, egress-guard agent-id-mode rules).
In the new design the auth-sidecar runs ONCE per cluster as a
Helm-managed Deployment in `kars-system`; per-pod injection is no
longer needed.
- `cli/src/commands/mesh/vmss_mi_assign.ts` — VMSS managed-identity
assignment driver. Not needed when the sidecar runs in kars-system
with its own Workload Identity binding.
- `materialise_sidecar_configmap` per-namespace mirror in
`agent_id_provisioning.rs`. The shared sidecar consumes a single
`kars-system`-scoped ConfigMap managed by `auth_config_reconciler`.
- The agent-id-mode iptables rules in the egress-guard init container
(UID 1000 → REJECT 127.0.0.1:8080, UID 1001 → REJECT IMDS). Trust
boundary for cross-pod sidecar access moves to a NetworkPolicy on
the sidecar's namespace (added in a follow-up commit).
## What's preserved
- `controller/src/agent_identity.rs` — Graph client. Same provisioning
endpoints, same auth flow. With or without per-pod sidecar.
- `controller/src/agent_id_provisioning.rs` — provisioning loop with
idempotency, tag-based recovery from crashes, permissive Graph list
parser, status patch idempotency. Hard-won lessons preserved.
- `controller/src/auth_config.rs` + `auth_config_reconciler.rs` —
KarsAuthConfig CRD + its reconciler. Still cluster-level singleton
config; the sidecar just consumes a kars-system-scoped CM now.
- `inference-router/src/sidecar_client.rs` — URL-agnostic client.
The new design points it at the cluster Service DNS instead of
loopback; no code change needed.
- All Helm CRD templates, RBAC additions, IMDS NetworkPolicy egress
rule, the Bicep blueprint provisioning template, and the
Graph Explorer recovery runbook docs.
## What's reset to baseline behaviour
- `reconciler/mod.rs` egress-guard: back to the seven-line baseline
iptables script (the same one shipped on `main`). No agent-id-mode
branching; the security boundary for cross-pod sidecar access is
the sidecar-namespace NetworkPolicy, not pod-local iptables.
- Inference-router env injection: still pushes
`PINNED_AGENT_IDENTITY_APP_ID` + `AUTH_SIDECAR_URL` when an agent
identity is provisioned, but `AUTH_SIDECAR_URL` now points at
`http://entra-auth-sidecar.kars-system.svc:5000` instead of
loopback. Helm chart for the sidecar Deployment lands in the next
commit.
## Why this redesign
Microsoft's own Agent ID design-pattern docs and the Auth Sidecar
API support `?AgentIdentity=<appId>` so one sidecar instance can
mint tokens for any of the blueprint's children. We had been running
one sidecar per pod (N × 128MB), iptables-isolating cross-UID access
to localhost:8080. Switching to one shared sidecar (kars-system,
2 replicas, NetworkPolicy-gated) saves ~5x memory at 10 sandboxes,
eliminates the per-pod ASPNETCORE bind + probe gymnastics, and aligns
with the Microsoft.Identity.Web library's centralized MSAL cache.
Independent research review confirmed the per-pod model is one of
two documented patterns but is more resource-intensive; the
`?AgentIdentity=<appId>` shared-sidecar pattern is explicitly
documented as a valid alternative and the SDK's `AzureAd__ClientId`
takes the blueprint appId, making each child mintable on demand.
## Tests
785/785 controller tests pass. Inference-router tests untouched.
* feat(helm): shared entra-auth-sidecar Deployment (Phase 1)
One auth-sidecar Deployment per cluster (2 replicas, HA) serves all
sandboxes via the Microsoft AuthorizationHeaderUnauthenticated
?AgentIdentity=<appId> query parameter — replacing the per-pod
sidecar injection model from earlier iterations.
Templates added (all gated on `entraSidecar.enabled`, default false):
- auth-sidecar-serviceaccount.yaml - WI-annotated SA
- auth-sidecar-deployment.yaml - distroless sidecar, /healthz probes
- auth-sidecar-service.yaml - ClusterIP :5000
- auth-sidecar-networkpolicy.yaml - ingress only from sandbox routers
values.yaml gets the `entraSidecar:` config block populated by
`kars up` from the KarsAuthConfig CR.
Trust boundary rationale and credential-supply patterns (WI for OSS,
IMDS-MI for corp tenant) explained inline in each template.
Resource footprint: ~160MB cluster-total vs ~1.28GB at 10 sandboxes
under the per-pod model.
Tests: helm lint clean, helm template renders 4 resources when
enabled, 0 when disabled. Controller 785/785 unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): allow sandbox→kars-system:5000 egress for shared sidecar (Phase 2)
The shared entra-auth-sidecar lives in the kars-system namespace and
is reached by every sandbox's inference-router over TCP 5000. The
sandbox-policy NetworkPolicy now includes an explicit egress rule
selecting namespaces labeled
app.kubernetes.io/name=kars
app.kubernetes.io/component=system
on port 5000.
Also fixed the auth-sidecar's own ingress NetworkPolicy: the selector
now correctly targets pods labeled kars.azure.com/component=sandbox
(the pod label) rather than the non-existent
kars.azure.com/component=inference-router (which is a container, not
a pod, and so was never matchable).
Trust boundary is now two-sided:
- Sandbox-side: egress allows reaching only the kars-system namespace
on port 5000, where the sidecar Service is the only listener.
- Sidecar-side: ingress allows only pods in sandbox-labeled namespaces.
When entraSidecar.enabled=false at the Helm level, no auth-sidecar
exists and the egress rule is a harmless no-op (no destination to
reach).
Tests: controller 785/785, helm lint clean, helm template renders
correctly with enabled=true and enabled=false.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(router): tid+principal+aud+exp pinning on sidecar tokens (Phase 3)
Wire the previously-orphaned sidecar_client module into the
inference-router and harden it against token substitution attacks.
CONTROLLER (reconciler/mod.rs):
- Extends agent_id_active tuple to carry tenant_id alongside agent
identity. Sourced from KarsAuthConfig.spec.tenant.tenantId via
ProvisioningOutcome::Ready.auth_spec.
- Stamps EXPECTED_TENANT_ID env on the router for sidecar-mode
sandboxes. Without it, router-side tid pinning is disabled and
warns at boot (insecure, dev-only).
ROUTER (sidecar_client.rs, auth.rs, lib.rs):
- pub mod sidecar_client; — fix the orphaned-module bug.
- WorkloadIdentityAuth now consults SidecarClient first; sidecar mode
is the EXCLUSIVE auth path (no WI/IMDS/API-key fallback). Preserves
per-sandbox audit attribution in downstream Azure RBAC.
- from_env() now Result<Option<Self>>. Partial config (only URL or
only PINNED set) returns Err; WorkloadIdentityAuth::new() panics →
AKS surfaces as CrashLoopBackoff. No more silent fallback to a
different identity model.
- HARD-fail validate_token_claims on EVERY check (rubber-duck #1):
* tid mismatch / missing (cross-tenant guard)
* appid/azp mismatch when present (audit attribution guard)
— if both present, BOTH must match
* aud mismatch against per-service expected set (cache-poisoning
guard) — Foundry, Graph, OpenAI, Management, Search mapped
* exp in past or within 60s skew (stale-token guard)
* exp missing when tid-pinning enabled (unbounded-lifetime guard)
- Returns a TTL cap = exp - now - 60s; caller takes
min(sidecar-advertised, JWT-derived). Validation runs BEFORE caching.
- JWT decoder tightened: requires EXACTLY 3 segments (rejects
2-segment unsecured JWS and 4-segment JWE).
- aud claim normalized: accepts both single-string (Entra default)
and array (RFC 7519); rejects out-of-spec shapes (number, bool).
- Soft WARN when both appid and azp are absent (some MSI tokens
legitimately omit both).
- New env var const ENV_EXPECTED_TENANT_ID + boot-time log of pinning
state for operator visibility.
TESTS:
- 39 sidecar_client unit + wiremock integration tests, covering all
HARD-fail branches, the cache-skip on validation failure, the TTL
cap math, the partial-config Err path, JWT decoder edge cases
(2-seg, 4-seg, empty payload, bad base64, non-JSON, array aud,
out-of-spec aud).
- Router: 1014/1014. Controller: 785/785. Clippy clean on sidecar_client.
Rubber-duck review caught: (1) appid/azp should be hard not soft, (2)
partial env config must fail closed, (3) 3-segment requirement, (4)
aud validation against requested resource, (5) exp validation + TTL
cap. All five addressed in this commit.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(cli,bicep,controller): dual credential mode auto-detect (Phase 4)
Adds first-class support for two auth-sidecar credential supply
patterns and auto-detects which one works in the current Entra tenant.
PATTERN A (ManagedIdentityImds, default):
- Sidecar auths via SignedAssertionFromManagedIdentity against the
controller MI's IMDS endpoint. Corp-tenant safe; required when
the tenant's FIC issuer-allowlist policy blocks AKS OIDC
(Microsoft-corporate: InvalidFederatedIdentityCredentialValue).
- Bicep creates: blueprint + SP + controller MI + MI-as-FIC.
PATTERN B (WorkloadIdentity):
- Sidecar auths via SignedAssertionFilePath against the projected
K8s SA token. No per-cluster MI, no VMSS identity assignment.
- Bicep creates: blueprint + SP + SA-as-FIC pointing at the AKS
cluster's OIDC issuer URL.
CRD (controller/src/auth_config.rs):
- Added `controller.credentialMode` enum (ManagedIdentityImds default).
- Made `controller.managedIdentity{Client,Resource,Principal}Id`
Optional (required only in MI mode).
- Added is_valid_for_mode() validator.
CONTROLLER (auth_config_reconciler.rs):
- render_sidecar_env branches on credentialMode and emits the
correct AzureAd__ClientCredentials__0__SourceType. MI mode
emits ManagedIdentityClientId; WI mode emits
SignedAssertionFileDiskPath=/var/run/secrets/azure/tokens/azure-identity-token.
- Spec validation runs BEFORE ConfigMap materialisation. Invalid
specs (MI mode + empty clientId) surface as phase=Degraded
with InvalidCredentialMode condition — refuses to propagate
the misconfiguration to running sandboxes.
- New patch_degraded_status helper.
BICEP (deploy/bicep/agent-id-trust.bicep):
- New credentialMode parameter (allowed: ManagedIdentityImds |
WorkloadIdentity).
- Conditional resources: controller MI + MI-as-FIC only in
Pattern A; SA-as-FIC only in Pattern B.
- New aksOidcIssuerUrl parameter (required for Pattern B).
- New outputs: credentialMode, aksOidcIssuerUrl.
- az bicep build clean, no warnings.
CLI:
- ensureAgentIdTrust now accepts credentialMode (auto | WorkloadIdentity
| ManagedIdentityImds), aksClusterName, aksClusterResourceGroup.
- AUTO MODE: tries Pattern B first by discovering the AKS OIDC
issuer URL via az aks show, then creating an SA-as-FIC. On
InvalidFederatedIdentityCredentialValue (corp tenant signature),
falls back to creating the controller MI + MI-as-FIC.
- New TenantRejectedAksOidcIssuer sentinel for clean fallback
orchestration.
- New discoverAksOidcIssuerUrl helper — accepts explicit args or
walks kubeconfig + az aks list to guess.
- writeKarsAuthConfig strips MI fields when credentialMode is
WorkloadIdentity (matches the controller's Optional schema).
- setup-trust gains --credential-mode, --aks-cluster-name,
--aks-cluster-resource-group, --aks-oidc-issuer-url flags.
- Bicep wrapper threads credentialMode + aksOidcIssuerUrl through
to ARM.
TESTS:
- Controller: 789/789 (4 new — WI mode render, MI empty-field
rendering, is_valid_for_mode in MI and WI modes).
- Router: 1014/1014 (unchanged).
- CLI: 786/786 (2 new — credentialMode propagation through
dry-run with default + explicit).
- Bicep: az bicep build clean.
- Helm lint + template clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller,bicep,docs): security alignment (Phase 5)
Closes the rubber-duck research findings from the Phase 3 critique
and Microsoft's Entra Agent ID design-patterns audit.
CUSTOM SECURITY ATTRIBUTES (rubber-duck #3, MEDIUM):
- KarsSandbox.spec.meshAuth.customSecurityAttributes:
BTreeMap<set, BTreeMap<attr, Value>>. Operator declares which
attributes (from a tenant-declared set) the controller should
PATCH onto each per-sandbox agent identity.
- AgentIdentityClient::patch_custom_security_attributes Graph
client method. Constructs the documented CustomSecurityAttribute-
Value envelope with the required @odata.type per attribute,
inferred from the JSON value shape.
- odata_type_for_value helper: maps serde_json::Value →
'#String' | '#Int32' | '#Boolean' | '#Collection($T)' and
rejects floats, nulls, nested objects, mixed-type arrays, and
empty arrays with clear error messages BEFORE the call goes out.
- Wired into ensure_agent_identity_for_sandbox: PATCH runs on
every reconcile (idempotent on Graph). Failures surface as
ProvisioningOutcome::Failed → sandbox phase=Degraded, preventing
silent missing-attribute drift.
SCALE-OUT INVARIANT (rubber-duck #4):
- Documented in agent_id_provisioning.rs module doc: the agent
identity is keyed on KarsSandbox.metadata.uid (and cluster UID),
with NO per-pod / per-replica / per-ordinal dimension. All
replicas of one KarsSandbox share ONE agent identity.
- New test tag_layout_excludes_per_pod_attributes pins the tag
layout — any future PR that adds a 'kars-pod-' / 'kars-replica-'
/ 'kars-ordinal-' / 'kars-hostname-' / 'kars-podname-' tag prefix
breaks the test.
- Visibility change: AgentIdentityClient::tags_for is now
pub(crate) so the cross-module invariant test can call it.
BOOTSTRAP SCRIPTS (deploy/bicep/standalone/):
- custom-security-attributes.sh: declares the recommended
AgentGovernance set with 4 attributes — AgentClassification
(Standard|Restricted|Confidential), DataSensitivity
(Public|Internal|Confidential), ProductOwner, ManagedBy.
Idempotent via az rest against the Graph beta endpoint.
(The Microsoft.Graph Bicep extension does not yet ship typed
attributeSets / customSecurityAttributeDefinitions resources,
so this is a shell script that operators run once per tenant.)
- conditional-access-baseline.sh: applies Microsoft's
policy-autonomous-agents template — blocks sign-ins where the
risk level meets the configured threshold (default: high),
targeted via the ManagedBy=kars-controller attribute filter.
Defaults to report-only state for safe rollout. Idempotent via
upsert.
FOUNDRY RBAC (research R5):
- foundry-rbac.bicep grants Azure AI User on a Foundry resource
to the BLUEPRINT SP (not per-agent). All derived agent
identities inherit access — eliminates per-agent role-assignment
churn. Supports RG-scoped (default) and resource-scoped
assignment via foundryResourceName parameter. az bicep build
clean, no warnings.
DOCS:
- docs/architecture/entra-agent-id/05-security-alignment.md:
6.8KB operator-facing runbook covering the bootstrap order,
KarsSandbox YAML example, failure modes table, scale-out
invariant rationale, Foundry RBAC inheritance.
TESTS:
- Controller: 802/802 (+13 new — 12 odata_type_for_value variants
covering supported + rejected shapes, 1 scale-out invariant).
- Bicep: foundry-rbac.bicep builds clean. JSON compiled.
- Bash: both .sh scripts pass 'bash -n' syntax check.
- CLI: 786/786 (unchanged).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(phase-7): live deploy bug fixes + Foundry RBAC fallback (Phase 7)
End-to-end live validation of the Phase 0-5 shared-sidecar architecture
against real Microsoft corp tenant (kars-aks). Three discoveries
addressed in this commit:
1) ASP.NET Core HostFiltering rejects Service-DNS callers (router fix)
- Microsoft Entra SDK auth-sidecar's HostFiltering middleware rejects
non-localhost Host headers regardless of AllowedHosts=* env.
- inference-router/src/sidecar_client.rs: always send
Host: localhost:5000 on sidecar requests.
- deploy/helm/kars/templates/auth-sidecar-deployment.yaml: set
AllowedHosts=* env for completeness (sidecar config source doesn't
honour it for dynamic-binding path).
- NetworkPolicy remains the real ingress boundary; bypassing host
filter does not weaken security.
2) Azure RBAC inheritance correction (sandbox_bringup.ts)
- Phase 5 originally assumed Foundry RBAC would inherit blueprint ->
derived agent identity SPs.
- Microsoft docs (concept-agent-id-design-patterns) clarify that
inheritable permissions = Microsoft Graph delegated permissions only.
Azure RBAC is per-principal.
- Confirmed empirically live: granting on blueprint SP alone did NOT
unblock chat; direct grant on each agent identity SP did.
- cli/src/commands/up/sandbox_bringup.ts: inline Bicep now grants
Cognitive Services OpenAI User on the blueprint SP as a
break-glass / fallback (kept for the case where the controller
can't grant per-agent). Inheritable-permissions language removed.
- Operator-facing error message documents the per-agent grant path
when ARM deployment fails for permission reasons.
- Phase 5b (future PR): controller must assign Azure RBAC per agent.
3) Multi-agent exec-brief demo: chain proven multi-agent
- Demo applies CRDs, brings up parent execbrief + 3 sub-agents.
- Each sub-agent gets its own typed Entra agent identity SP.
- Verified live (with all 4 SPs granted Cognitive Services OpenAI
User AND Azure AI User on the Foundry account via az rest):
* All 4 sandboxes booted in fail-closed sidecar mode
* 65 successful Foundry 200s across all 4 sandboxes in 10 min
* ZERO PermissionDenied responses after roles in place
* Mesh routing: parent dispatched to analyst, replied
* E2E encrypted relay: file transfer (analyst.json) ACKed
* AGT trust scoring: +0.8 reputation submitted, accepted=true
* NetworkPolicy: 0 egress denials, 0 ingress drops
- Demo's verify.json shows 3/9 checks pass — passing checks are the
kars-runtime mechanisms (sub-agents active, egress 0 denials,
telegram skipped). 6 failing checks ALL trace to one root cause:
Foundry project's Bing Grounding connection is not configured
(foundry_web_search requires bing project_connection_id). This is
orthogonal to our work; the auth chain is proven.
TESTS:
- inference-router sidecar_client: 39/39 pass.
- CLI: 786 pass + 2 skipped (no regressions).
- Bicep: existing modules build clean.
- Helm: lint clean, template renders correctly.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* docs(entra-agent-id): top-level README + migration guide (Phase 8)
Consolidates the entra-agent-id architecture documentation into a
navigable index and adds a migration guide for operators upgrading
from earlier per-pod-sidecar branch heads.
- docs/architecture/entra-agent-id/README.md (new top-level index):
* TL;DR architecture diagram
* Pattern A/B selection table
* Phase ledger linking each commit to its scope
* Key files map
* Live validation snapshot (verified on kars-aks 2026-05-28)
- docs/architecture/entra-agent-id/00-poc-archive.md:
* Previous POC README, renamed to archive. The POC scaffolding
drove early design but is no longer the canonical reference.
- docs/architecture/entra-agent-id/04-migration-guide.md (new):
* Step-by-step upgrade path from per-pod-sidecar branch heads
* Helm upgrade command with the entraSidecar values
* Per-agent role grant runbook (az rest workaround for the
CA-blocked az role assignment create path)
* Rollback procedure
* Known caveats (HostFiltering, MSAL cache per replica,
Bing Grounding orthogonal issue)
The numbered ordering reflects deployment lifecycle:
00 — historical POC (archived)
01 — runtime token flow
02 — alternative ACI flow (reference)
03 — original POC findings
04 — migration guide (this commit)
05 — security alignment (Phase 5)
Closes Phase 8 of the feat/entra-agent-id PR.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): per-agent ARM RBAC + agent identity cleanup (Phase 5b)
Eliminates the manual az role assignment runbook step from the Phase 5
migration guide by having the controller assign and revoke Azure RBAC
roles per agent identity automatically. Closes the orphan SP risk by
adding agent-identity deprovision to the KarsSandbox deletion
finalizer.
CRD (auth_config.rs):
- New `KarsAuthConfig.spec.foundryRbac: Vec<FoundryRbacAssignment>`.
- Each assignment carries an ARM `scope` and a list of built-in role
definition GUIDs to PUT against every per-sandbox agent identity SP
at provisioning time.
- Empty list (default) preserves the manual-grant workflow for
backward compatibility with existing kars deployments.
ARM REST extensions (agent_identity.rs):
- `arm_token()` — MI token for management.azure.com audience.
- `assign_role_to_agent_identity()` — PUT roleAssignment with a
deterministic UUIDv4 derived from (scope, principal, role), so
repeated PUTs are idempotent on Azure's side. Treats 200/201 as
success and 409 RoleAssignmentExists as success.
- `delete_role_assignments_for_principal()` — GET assignments
filtered by `principalId` (NOT combined with `atScope()` — Azure
REST returns 400 UnsupportedQuery), narrows to the requested scope
client-side, then DELETEs each. Used by the deletion finalizer.
- New helpers: `deterministic_assignment_guid()` (SHA-256 → UUIDv4)
+ `extract_subscription_id()` (scope parser).
- 9 new unit tests cover GUID stability, case-insensitivity, UUID
format, scope parsing happy-path + rejected forms.
Provisioning (agent_id_provisioning.rs):
- After Graph create (step 3c): iterate `spec.foundry_rbac` and assign
every (scope, role) tuple to the new identity. WARN-and-continue on
failure so the sandbox still boots; ARM RBAC converges on retry.
- Early-return path (recorded identity): RE-ASSERT the same role
assignments so existing sandboxes pick up retroactive RBAC config
the first time the operator adds `foundryRbac` to KarsAuthConfig.
Idempotent via the deterministic GUID.
- Logs `Phase 5b reconcile: re-asserting ARM role assignments` with
`foundry_rbac_entries` count for visibility.
Cleanup (agent_id_provisioning.rs + reconciler/mod.rs):
- New `cleanup_agent_identity_for_sandbox()` orchestrates the two
Azure-side teardown steps on sandbox delete:
1. DELETE role assignments held by the agent identity SP at every
configured `foundryRbac` scope.
2. Graph DELETE the agent identity SP itself.
- Hooked into the KarsSandbox deletion finalizer in
reconciler/mod.rs. Best-effort: failures logged as WARN but do not
block finalizer removal, so the K8s CR doesn't get stuck Terminating
when Azure is degraded. The existing orphan reaper backstops what
this path misses.
VERIFIED LIVE on kars-aks:
- All 5 existing sandboxes' agent identities (execbrief + 3 sub-agents
+ kars-testrun) now have exactly 2 role assignments each on the
Foundry account (Cognitive Services OpenAI User + Azure AI User),
re-asserted by the new code on the early-return path.
- Deleting `kars-testrun` sandbox: controller logged
"deleted role assignments for principal" → 2 deleted
"agent identity SP deprovisioned"
Verified Foundry now reports 0 role assignments for the deleted
appId — full cleanup successful.
Tests: 811/811 controller (+9 new for the GUID + scope helpers).
Build clean, helm template + lint clean.
One-time operator prerequisite: the controller's identity (the MI
whose IMDS token the controller uses for ARM calls — typically the
AKS kubelet/agentpool MI) must hold
`Microsoft.Authorization/roleAssignments/write` at every scope listed
in `foundryRbac`. The kars-recommended way is to grant Role Based
Access Control Administrator on the Foundry RG. This is intentionally
a one-time operator action documented in the migration guide rather
than auto-granted, since it crosses cluster/customer trust boundaries.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(controller): scaffold MeshAuthBackend CRD field for Phase 6 design
Adds KarsAuthConfig.spec.meshAuthBackend (enum: Anonymous default,
EntraAgentIdentity opt-in) and optional meshAuthAudience override for
the next milestone — verified Entra-signed AGT mesh peer authentication.
Defaults preserve full backward compatibility (every existing cluster
keeps registering anonymously with no behaviour change). Operators on
clusters that have completed sidecar-based entrypoint mint + relay JWKS
verification (the next-PR work) flip the field to EntraAgentIdentity.
Includes a design doc (docs/architecture/entra-agent-id/06-mesh-trust-design.md)
capturing the three independent pieces that must land for end-to-end
enforcement (CRD scaffold here; sandbox entrypoint via sidecar; AGT
relay JWKS verification — last piece is upstream-coordination work).
Unit tests pin: default is Anonymous (backward compat), both variants
deserialize, unknown variants are rejected (no silent fallback).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(router,sandbox): /v1/mesh-token endpoint + Entra-signed mesh peer path
Phase 6.b — sandbox-side path for verified AGT mesh trust. When
KarsAuthConfig.spec.meshAuthBackend=EntraAgentIdentity, the controller
sets MESH_AUTH_BACKEND on the router, the router exposes
GET /v1/mesh-token, and entrypoint.sh acquires a verified-tier agent
identity token via the shared auth-sidecar instead of doing the legacy
direct Workload-Identity → Entra exchange.
Router changes:
- New inference-router/src/routes/mesh_token.rs route gated on
MESH_AUTH_BACKEND; returns 404 when disabled, 503 if sidecar not
configured, 502 on sidecar call failure, 200 with {access_token,
token_type, expires_in} on success
- Extended sidecar_client::resource_to_service_name to map
api://agentmesh* → AgentMesh service key
- Optional MESH_AUTH_AUDIENCE env override (defaults to
api://agentmesh/.default)
- 7 new mesh_token tests + agent-mesh resource-mapping test, all
serialised on a process-wide ENV_LOCK mutex to avoid env races
Controller changes:
- reconciler/mod.rs injects MESH_AUTH_BACKEND + MESH_AUTH_AUDIENCE env
on the sandbox's router container when the CRD field is set
- auth_config_reconciler.rs auto-emits the DownstreamApis__AgentMesh__*
cluster onto the shared sidecar when meshAuthBackend=EntraAgentIdentity
(skipped if operator has supplied an explicit AgentMesh entry)
- 4 new reconciler tests pin: anonymous default emits nothing, entra
variant emits expected env, custom audience honoured, operator-
supplied entry wins
Sandbox changes:
- entrypoint.sh: new branch BEFORE the legacy WI-exchange one. When
MESH_AUTH_BACKEND=EntraAgentIdentity, curl localhost:8443/v1/mesh-token,
export AGT_OAUTH_TOKEN on success, force AGT_TRUST_THRESHOLD=0 on
failure (anonymous-tier fail-open contract matches existing logic)
- Backward compat: when MESH_AUTH_BACKEND is unset (default), the
entrypoint follows the exact same path as before — no behaviour
change for existing clusters
Test results:
- 819 controller tests pass (was 815, +4)
- 930 router tests pass (was 923, +7)
- entrypoint.sh sh…
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 31, 2026
…cal-k8s, docker (#368) * docs(security-validation): cross-platform validation report — AKS, local-k8s, docker Comprehensive security & runtime validation across all three deploy modes, mapped to the 9-layer model in docs/security.md. Per-platform live evidence collected: - CRD presence + InferencePolicyCompiled/ToolPolicyCompiled/EgressAllowlistCompiled digests - Pod securityContext (UID, readOnlyRootFilesystem, capabilities, seccomp) - Workload Identity / Entra Agent ID federated token plumbing (AKS only) - entra-auth-sidecar token issuance + per-sandbox pinned_agent_id - Output authenticity (real Foundry calls, real URLs verified via HTTP 200) - AGT mesh KNOCK + E2E channel establishment per agent - Native AGT governance modules (PolicyEngine, AuditLogger, etc.) - NetworkPolicy enforcement + egress-guard caps - 9-check verify run on every platform Verified all three platforms 9/9 PASS with current main checks.py. Findings (no fixes applied — tracked for separate PRs): #1 HIGH (docker macOS) — UID 1000 reads Foundry API key (Docker Desktop UID virtualization) #2 HIGH (local-k8s) — API key + GitHub token in plaintext pod env (should be secretKeyRef) #3 MEDIUM (docker) — NET_ADMIN on persistent parent container (vs init-only in K8s mode) #4 MEDIUM (AKS, local-k8s) — blocklist-refresh CronJob failing every 6h (VAP collision) #5 MEDIUM (all) — audit JSONL not persisted (RO root FS, no emptyDir mount) #6 LOW (AKS) — AllowlistVerified=False on execbrief (inline endpoints, no cosign attestation) #7 LOW (AKS) — TrustGraph router-side enforcement not active (documented roadmap) All findings include a remediation plan in §10. None block the AKS verified-tier security story documented in docs/security.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> * docs(security-validation): add env-var inventory addendum for AKS containers Adds detailed per-container env-var analysis answering the question: 'do AKS containers have more env variables than they should?' Per-container inventories with full categorization: - openclaw container: 42 vars in 8 categories - inference-router container: 54 vars (router-only paths/toggles) Three additional findings on env-var hygiene: #8 LOW — OPENCLAW_GATEWAY_TOKEN exposed via env (should be file mount) #9 LOW — enableServiceLinks=true leaks internal cluster IPs (16 env vars) #10 LOW — possibly-redundant Foundry/Mesh-auth env vars on openclaw Headline confirmations: ✅ NO AZURE_OPENAI_API_KEY on either AKS container ✅ NO COPILOT_GITHUB_TOKEN on either AKS container ✅ Federated identity token mounted RO, never as env value ✅ Auth mode 'shared entra-auth-sidecar fail-closed, no WI/IMDS/API-key fallback' Comparison table across all 3 platforms included. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> --------- Signed-off-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
Jun 4, 2026
… working
Lands the protocol-correct fixes needed for MeshClient.connect() →
KNOCK → X3DH → Double Ratchet roundtrip between two sandboxes. Tested
end-to-end on kind-kars-dev with two Hermes pods (execbrief-hermes and
smoke-hermes) on the FRESHLY BUILT image (no hot patches):
- pod A registers, uploads prekey bundle, opens relay WS (with POP)
- pod B does the same
- pod A discovers B via /v1/discover (freshest-first sort)
- pod A fetches B's bundle, runs X3DH, sends KNOCK + first ciphertext
- pod B's _handle_knock_frame auto-accepts via SecureChannel.create_receiver,
decrypts plaintext 'hello from execbrief-hermes'
- pod B replies via send_by_did → encrypted message frame
- pod A decrypts 'pong from smoke-hermes'
## Critical protocol fixes
1. **Relay WS connect-frame POP** (relay_transport.py)
- Was: {type:'connect', from:did, ts:...}
- Now: full proof-of-possession (std-base64 pub_key + iso ts + sig
over ts), per AGT relay/app.py::_verify_connect_pop
- Without this, the relay rejects every connection with
'connect frame missing did/public_key/timestamp/signature'
2. **Registry auth header** (registry_client.py)
- Was: three separate X-Agent-DID/Timestamp/Signature headers,
signature over method+path+ts
- Now: single 'Authorization: Ed25519-Timestamp <did> <ts> <b64url-sig>',
signature over timestamp string only
- Matches AGT registry/app.py::verify_ed25519_timestamp_auth
3. **X3DH bootstrap missing** (client.py)
- Now connect() builds X3DHKeyManager + generates signed_pre_key
+ 10 OTKs + uploads bundle via PUT /v1/agents/{did}/prekeys
- Without this, peers couldn't fetch our bundle, X3DH initiation
would fail at the responder side
4. **KNOCK responder implemented** (client.py::_handle_knock_frame)
- Was: log-only stub ('responder path not implemented')
- Now: parses ChannelEstablishment, calls SecureChannel.create_receiver,
caches the channel, decrypts the bundled first ciphertext,
eagerly tops up the OTK pool for the next session
5. **Send fuses KNOCK + first message** (client.py::send_by_did)
- First call to a new peer DID sends {type:'knock', establishment, ciphertext}
- Subsequent calls send {type:'message', ciphertext}
- Matches the TS SDK wire convention (one RTT, not two)
6. **AAD directionality fix** (client.py)
- Initiator: f'{self_did}|{peer_did}'
- Responder: f'{from_did}|{self_did}' (reconstructs the same bytes)
7. **EncryptedMessage wire format** (client.py)
- Was: JSON of em.__dict__ (would fail at decoder)
- Now: EncryptedMessage.serialize() / .deserialize() (binary + b64url)
8. **PeerBundle flat shape** (registry_client.py + client.py)
- Was: nested dicts mirroring my best-guess wire format
- Now: matches agentmesh.encryption.x3dh.PreKeyBundle's flat dataclass
9. **register_self handles 409 gracefully** (registry_client.py)
- Was: raised MeshRegistryError, blocking every restart
- Now: logs and continues — the subsequent prekey PUT (with
Ed25519-Timestamp auth) proves we own the same key
10. **discover() sorts freshest-first** (registry_client.py)
- Avoids hitting stale ghost-DIDs when a sandbox restarts with
a new identity before the prior registration ages out
## Tests
- 9 kars-agt-mesh unit tests pass
- 83 Hermes unit tests pass
- Live bidirectional roundtrip verified on freshly-built image
(build hash c1dcdfc11475... loaded into kind-kars-dev)
## Security audit updated
docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
- Residual risk #1 (no KNOCK responder) removed — now implemented.
- Added residual risk #4 (stale registry entries — non-security).
- Added live bidirectional test description.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
Jun 8, 2026
User report:
> operator says "✓ Spawned" then nothing visible
> kubectl get karssandbox -A confirms the CR was never created
Two compounding silent-failure bugs:
1. kars add was log-then-exit-0 on caught errors.
The outer catch at cli/src/commands/add.ts line 601 (was: 531)
handled every exception by calling spinner.fail() + console.error()
and then RETURNING — letting Node exit 0 naturally. So
`kubectl apply -f -` failing (CRD missing, wrong context, schema
rejection on the bundle, etc.) surfaced as a clean exit code to
any caller. Operator's `execa("kars", args, { stdio: "pipe" })`
only logs `✗ Spawn fail` when execa REJECTS, so silent exit-0
masked every kars-add failure mode behind a green checkmark.
Fix: add `process.exit(1)` after the error logs. Preserves all
the existing error-message branching (controller-not-installed
hint, generic error text) — just stops lying about exit status.
2. Operator's spawn dialog was throwing away the real error text.
Previously logged only `(e.stderr || e.message)?.substring(0, 200)`
— execa's `.message` is usually `Command failed with exit code 1:
kars add ...`, NOT the underlying kars-add stderr. So even after
fix #1, the operator log would show "✗ Spawn fail: Command failed
with exit code 1: kars add testhermes --runtime hermes ..." with
no actual root cause.
Fix: prefer e.stderr (now populated thanks to fix #1) over
e.message, strip ANSI colour codes that kars add emits via chalk,
filter empty lines, keep the last 4 (which is where spinner.fail
+ error hints live), join with " | ", cap at 400 chars. Activity
log now shows e.g.:
✗ Spawn fail: Failed to create sandbox | Error: kubectl error:
KarsSandbox.kars.azure.com "testhermes" is invalid: spec.hermes:
Invalid value: ... | Connect: kars connect testhermes
Also: on SUCCESS, echo the last 3 lines of stdout (the
"Namespace / Model / Status / Connect" hints kars add prints) so
the operator sees useful follow-up info inline.
Verified:
npm run build + typecheck ⇒ clean
vitest run ⇒ 798 passed
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Jun 9, 2026
* fix(hermes): A1 docker smoke fixes — version pin + plugin opt-in
Two real bugs surfaced when running the first `docker build` +
end-to-end smoke test of the Hermes sandbox image:
1. **Hermes version pin wrong**
`ARG HERMES_VERSION=0.5.1` doesn't exist on PyPI. The 0.5.x
assumption came from misreading the Hermes README's Homebrew
formula tag (`5.1.14`); the actual `hermes-agent` PyPI package
uses 0.x.y numbering at 0.15.2 latest. Bumped to 0.15.2.
Hermes 0.15.2's plugin contract (PluginContext.register_tool,
register_hook, plugin.yaml with provides_tools/provides_hooks,
discovery via `$HERMES_HOME/plugins/`) matches what the A1
plugin code was already built for — verified by importing
hermes_cli.plugins and running discover_plugins() against our
materialized plugin tree.
2. **ripgrep not in Azure Linux 3**
`tdnf install -y` exits non-zero if ANY package is missing, and
Azure Linux 3 doesn't ship ripgrep. Hermes' built-in file_search
tool prefers ripgrep but falls back to grep, so dropping it is
safe. Image now builds in ~30s.
3. **kars plugin discovered but not loaded**
Hermes treats `standalone` plugins as opt-in via
`plugins.enabled` in config.yaml. The entrypoint was placing the
kars plugin into `$HERMES_HOME/plugins/kars/` (correct user
discovery path), but never adding `kars` to the enabled
allow-list — so it was discovered and silently skipped with
`error='not enabled in config'`.
The entrypoint now emits a `plugins.enabled: [kars]` block at
the top of every generated config.yaml. The awk-merge that
replaces prior `mcp_servers:` blocks was extended to also
replace prior `plugins:` blocks so re-runs are idempotent.
Verified end-to-end:
- `docker build` succeeds
- `discover_plugins()` loads kars plugin, registers 10 tools +
2 hooks (pre_tool_call + post_tool_call)
- Entrypoint generates correct config.yaml with both blocks
- `$HERMES_HOME/plugins/kars/` materialized from
`/opt/kars-hermes-stage/plugins/kars/` on every boot
- 83/83 python unit tests still pass inside the image
- Mock smoke run: `python3 -m hermes_cli.plugins discover` shows
kars: enabled=True, 17 total plugin tools across all enabled
plugins (10 from kars + 7 web/foundry from bundled providers)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(hermes): A1.2 CRD schema + gitignore for cross-compile artifacts
Two follow-ups from the kind-cluster end-to-end smoke test:
1. **Helm CRD schema missing Hermes enum** — controller's `crd.rs`
added `RuntimeKind::Hermes` in a7882b8 but the matching Helm
CRD YAML wasn't updated. Result: the API server rejected every
KarsSandbox with `runtime.kind: Hermes` BEFORE the controller
ever saw it. Verified by `kubectl apply --dry-run=server`
failing with "unknown enum value 'Hermes'".
Added:
- `Hermes` to the `runtime.kind` enum at line 85
- x-kubernetes-validations rule:
`(self.kind == 'Hermes') == has(self.hermes)`
- `runtime.hermes` properties block mirroring `pydanticAi`
shape (version, agentCode oci/git, entrypoint, extraEnv)
After the fix, `kubectl apply -f /tmp/hermes-sandbox.yaml`
succeeds, controller picks up the CR, and a 2-container pod
(`agent` + `inference-router`) reaches `2/2 Running` with the
kars plugin loaded (10 tools + 2 hooks registered).
2. **`.cargo-docker/` not gitignored** — when cross-compiling for
linux/arm64 via `docker run -v $PWD:/work … cargo build` (the
pattern used for kind-on-M-series), `CARGO_HOME=/work/.cargo-docker`
keeps container-arch crate cache out of the host's `~/.cargo`.
That directory was leaking into `git status`. Added rules:
- `.cargo-docker/` — explicit
- `/bin/` was already covered by `**/[Bb]in/*` (verified)
Verified end-to-end on kind cluster `kars-dev`:
$ kubectl get karssandbox,pods -n kars-smoke-hermes
NAME PHASE RUNTIME INFERENCEPOLICY ISOLATION
smoke-hermes Hermes smoke-inference standard
NAME READY STATUS RESTARTS
smoke-hermes-697c6bd557-q5xfr 2/2 Running 0
Plugin discovery inside the pod:
kars plugin: enabled=True, source=user
hooks : {'pre_tool_call': 1, 'post_tool_call': 1}
tools : http_fetch, kars_discover, kars_mesh_{send,inbox,
await,transfer_file}, kars_spawn{,_status,_destroy,
_list}
Router /healthz from the agent container: 200 ok
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(hermes): A1 e2e smoke — six bugs surfaced by kind cluster run
End-to-end Hermes smoke on kind cluster exposed and fixed six real
bugs blocking the runtime from being functional:
1. awk not in Azure Linux 3 — replaced entrypoint merge with Python
2. TUI mode crashed without TTY — switched to hermes gateway run
3. KARS_MCP_SERVERS injected only into "openclaw" container —
generalized to use agent_container_name based on runtime kind
4. Entrypoint scanned wrong path for MCP servers — aligned to the
KARS_MCP_SERVERS env + loopback router pattern
5. hermes config set used key=value (wrong) — fixed to two positional args
6. Router rustls CryptoProvider not pre-installed — added explicit
aws_lc_rs::default_provider().install_default() in main()
Verified 12/12 e2e checks pass on kind cluster:
- Pod 2/2 Running, plugin loaded with 10 tools + 2 hooks
- Router /healthz, /agt/evaluate, /egress/fetch, /sandbox/list all 200
- KarsMemory CR Compiled, McpServer translated, channel translation
- Mesh stubs return clear Act 2 error
- pre_tool_call hook fires + decision=allow
All 834 controller + 932 router Rust tests pass.
cargo clippy clean, cargo fmt applied.
Security audit:
docs/internal/security-audits/2026-06-04-hermes-act1-e2e-smoke-fixes.md
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(controller): always allow operator policy-echo ingress (NP)
The sandbox NetworkPolicy gated ALL ingress rules behind
`governance.enabled=true`. With governance off, the NP shipped with
`policyTypes: [Ingress, Egress]` and an empty `ingress: []` block —
deny-all ingress. The operator namespace then could not reach
`/internal/policy-status` on the router and every referencing
InferencePolicy / KarsMemory / ToolPolicy / McpServer / EgressApproval
stuck forever in `Ready=False / AwaitingRouterEnforcement`, observable
in the operator panel even though the sandbox itself was healthy and
the router /readyz returned 200.
Split into two ingress classes:
- **Operator policy-echo ingress** (router :8443 admin surface from
ns labeled `app.kubernetes.io/name=kars,component=system`) — emitted
UNCONDITIONALLY. Three orthogonal gates still protect it: bearer
token, constant-time compare, optional IP pinning.
- **Peer-sandbox mesh + gateway ingress** (8443 / 18789 / 18791 from
ns labeled `kars.azure.com/role=sandbox`) — kept gated on
governance.enabled (no peers when governance is off).
Surfaced during local-k8s smoke of smoke-hermes: even after fixing
the AZURE_OPENAI_API_KEY env path so /readyz returned 200, three
policy CRs (InferencePolicy, KarsMemory, ToolPolicy) stayed
Ready=False because the controller's /internal/policy-status probe
to the sandbox router timed out at the NetworkPolicy level.
After this fix, with governance off, the controller's HTTP probe
gets a 401 (admin-token gate doing its job) instead of a connection
timeout, and the policy reconcilers update status using the round
trip rather than reporting "router unreachable".
Verified end-to-end on kind cluster `kars-dev`:
$ kubectl get inferencepolicy smoke-inference -n kars-system -o jsonpath='{.status.conditions}' | jq
- Ready=True RouterEnforcing: all 1 referencing sandbox router(s) confirmed inference-policy digest
- Progressing=False Reconciled: router echo confirmed
$ kubectl get karsmemory smoke-mem -n kars-system -o jsonpath='{.status.conditions}' | jq
- Ready=True RouterEnforcing: all 1 referencing sandbox router(s) confirmed claw-memory binding digest
834 controller tests still pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(e2e): add exec-brief-hermes-single scenario (Hermes Act 1)
Collapses the canonical 4-agent exec-brief scenario (parent +
analyst + viz + writer) into a single Hermes agent doing the whole
pipeline itself — research, scorecard, hero image, written brief.
Built to validate the Hermes runtime adapter end-to-end on
local-k8s and AKS without depending on the Python AGT MeshClient
(which ships in Act 2; until then, `kars_mesh_*` returns explicit
"Act 2 not ready" errors and the prompt explicitly tells the agent
not to call those tools).
Scenario layout (mirrors exec-brief/):
- manifests/00-namespace.yaml ........ kars-execbrief-hermes ns
- manifests/01-inferencepolicy.yaml .. azure-openai gpt-5.4
- manifests/02-toolpolicy.yaml ....... allow-all AGT profile
- manifests/03-clawmemory.yaml ....... memory-execbrief-hermes store
- manifests/04-mcpserver.yaml ........ DeepWiki MCP (same as canonical)
- manifests/05-clawsandbox.yaml ...... runtime.kind: Hermes
- config.sh .......................... SCENARIO_SUB_SANDBOXES=()
- prompt.txt ......................... single-agent pipeline
- README.md .......................... what it exercises + skips
Verified on kind cluster `kars-dev`:
$ kubectl apply -f tools/e2e-harness/scenarios/exec-brief-hermes-single/manifests/
→ 6 resources created
$ kubectl get karssandbox execbrief-hermes -n kars-system
PHASE=healthy RUNTIME=Hermes
$ kubectl get pods -n kars-execbrief-hermes
execbrief-hermes-... 2/2 Running
All 5 CRs reach RouterEnforcing / Ready=True:
● execbrief-hermes-inference InferencePolicy router echo confirmed
● execbrief-hermes-toolpolicy ToolPolicy agt-profile digest confirmed
● execbrief-hermes-memory KarsMemory binding=bound
● execbrief-hermes-deepwiki McpServer healthy
● execbrief-hermes KarsSandbox healthy
In-pod verification:
- kars plugin: enabled=True source=user, 10 tools + 2 hooks
- foundry_memory store_name = memory-execbrief-hermes (matches CR)
- config.yaml mcp_servers.execbrief-hermes-deepwiki present
- KARS_MCP_SERVERS=execbrief-hermes-deepwiki in agent env
- Router /readyz: 200 ok
Note: the actual LLM execution of the prompt requires real Azure
OpenAI / Foundry credentials. With the fake-key dev overlay used in
this validation, the pipeline runs through Hermes → kars plugin →
router → upstream-call layer and hangs at the upstream (expected).
Running with real creds — either via `kars dev --target local-k8s`
with a real provider, or on AKS via `SCENARIO=exec-brief-hermes-single
PLATFORM=aks ./tools/e2e-harness/run.sh` — will execute the full
pipeline.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(hermes): A1 e2e harness wiring — dev pull policy + Hermes posture hardening
End-to-end run of the new `exec-brief-hermes-single` scenario on
local-k8s surfaced four more bugs that all gate the prompt from
actually reaching the model:
1. **`pull_policy=Always` for `:latest` images** in dev mode forced a
doomed registry pull (karsacr.azurecr.io/…) instead of using the
kind-cached image. The controller now picks `IfNotPresent` when
`KARS_DEV_PROFILE=true` is set on its own env. Production AKS
stays on `Always` for `:latest`.
2. **Hermes' `tirith` auto-download** from GitHub releases blocked
every cold start while the kars egress-guard slow-walked the
fetch. Entrypoint now sets `TIRITH_ENABLED=false` by default;
Hermes falls back to its built-in pattern-matching shell
checker. Operators can re-enable by pre-baking the binary at
`/usr/local/bin/tirith` and setting `TIRITH_ENABLED=true`.
3. **`HERMES_DISABLE_LAZY_INSTALLS=1`** suppresses Hermes' `pip
install` of discord.py / google-* / brotlicffi on first use of
bundled platform plugins. Saves 30–120s on every cold start;
operators wanting the extras re-bake into the image.
4. **`HERMES_SKIP_NODE_BOOTSTRAP=1`** suppresses Hermes' shell-based
Node.js 22 LTS auto-installer (scripts/install.sh). We pre-install
`nodejs` + `nodejs-npm` from the Azure Linux 3 base repo
(currently v20.14 — Hermes' dep_ensure accepts any modern node).
Browser tools that need a Chromium download still need to be
pre-baked separately.
All three Hermes-runtime knobs are also mirrored into
`$HERMES_HOME/.env` so they survive `kubectl exec` sessions
(kubectl exec spawns a fresh env that doesn't see entrypoint
exports). Hermes' env_loader loads .env at import time
(`hermes_cli/env_loader.py:_load_dotenv_with_fallback`).
After all four fixes verified end-to-end:
- smoke-hermes sandbox: phase=Running, 2/2 Ready
- Router /readyz: 200 ok (controller forwards real Foundry API
key from `kars-dev-creds` Secret via secretKeyRef)
- Router /v1/chat/completions: 200 with real gpt-5.4 reply ("OK"
in 1.1s, latency_checkpoint shows engine_ttft_ms=108)
- InferencePolicy / KarsMemory / ToolPolicy / McpServer all
Ready=True / RouterEnforcing
- Plugin loaded with 10 tools + 2 hooks + foundry_memory native
- Platform MCP block present in config.yaml when
FOUNDRY_PROJECT_ENDPOINT is bound
Outstanding gap (NOT in this commit): Hermes' `hermes -z` still
makes an outbound HTTPS handshake (state=SYN_SENT to 104.18.3.115
:443, a Cloudflare IP — likely a check-update or telemetry endpoint
the harness hasn't tracked down). The kars egress-guard's
forward-proxy stalls the connection rather than denying outright,
so the prompt-driven path hangs after plugin discovery completes.
Workarounds:
(a) `KARS_EGRESS_LEARN=true` to log unallowed hosts, then
explicitly allowlist in EgressAllowlist;
(b) find Hermes' env to disable check-update / telemetry — Act 1.x;
(c) drive Hermes via Telegram channel instead of `hermes -z`.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(hermes): A1 e2e — Hermes runs the full exec-brief pipeline on real Foundry
The single-agent exec-brief scenario (research → JSON → scorecard PNG →
hero PNG → 2-page brief.md) now runs end-to-end on Hermes through the
kars router to real Azure Foundry gpt-5.4. Verified on local-k8s with
the user's ~/.kars/ creds.
Four fixes were needed (each surfaced sequentially as the agent loop
progressed further):
1. **`OPENAI_API_KEY` env routes Hermes to openrouter** (and openrouter.ai
is blocked by the egress-guard). Switched the entrypoint's `.env`
mirror to `AZURE_FOUNDRY_API_KEY` + `AZURE_FOUNDRY_BASE_URL` so
resolve_provider() picks the `azure-foundry` provider (which has
no built-in Cloudflare callback).
2. **`agent_init.py` hardcodes `_codex_reasoning_replay_enabled = True`**
→ Hermes echoes `{"type": "reasoning", "encrypted_content": "..."}`
back to /v1/responses on every continuation, which Azure Foundry's
strict schema validator rejects with `invalid_payload`. OpenAI's
own Responses API accepts these. Hermes only learns to disable
replay when the upstream returns `invalid_encrypted_content` (a
different error code that Foundry doesn't emit).
Router fix: `build_upstream_url()` in proxy.rs now strips
`input[]` items of `type=reasoning` and the
`include=["reasoning.encrypted_content"]` field from any /v1/responses
request bound for Azure Foundry (NOT GitHub Models / Copilot —
their schemas accept the original shape).
3. **/v1/responses handler used `forward()` (non-streaming)** but Hermes
always opens these with `responses.create(stream=True)` and expects
an SSE `text/event-stream` response. The buffered JSON blob made
Hermes' SDK raise "Connection error" after ~15s and retry 6× before
giving up with `max_retries_exhausted`. Switched the handler to
`forward_stream()` so the SSE byte stream flows through unchanged.
4. **`forward_stream()` injected `stream_options.include_usage`** which
the OpenAI Responses API rejects (`unknown_parameter`). Skip the
injection for /v1/responses (Foundry already emits usage in the
terminating SSE event); was already skipped for Anthropic
/v1/messages — same exclusion now covers both shapes.
Plus the entrypoint now persists `model.{default,provider,base_url}` in
config.yaml on every boot (not just plugins+mcp_servers), so a fresh
pod doesn't need a one-time `hermes config set model` post-boot dance.
End-to-end run delivered:
/sandbox/incoming/brief.md 6,136 B (2 pages, real Markdown,
12 footnoted https citations,
references hero+scorecard PNGs
inline, all 4 control-domain
terms present)
/sandbox/incoming/analyst.json 5,025 B (foundry_web_search × 3 →
trends / control_categories /
runtimes / metrics)
/sandbox/incoming/hero.png 30,094 B (1024×1024, foundry_image_generation
gpt-image-1, "Defense in Depth"
isometric data-center cutaway)
/sandbox/incoming/scorecard.png 12,201 B (1024×640, foundry_code_execute
matplotlib grouped bar chart,
4 runtimes × 4 control columns)
Router log: 30+ /v1/responses SSE streams, all 200 OK, latencies
1.6–67s. Foundry stream headers received for every request after
this fix; pre-fix only 2 of 8 requests had `Foundry complete` entries
before Hermes gave up.
Agent stdout (final response after autonomous tool-use loop):
> Done. Artifacts produced:
> - /sandbox/incoming/brief.md — 6136 bytes
> - /sandbox/incoming/hero.png — 30094 bytes
> - /sandbox/incoming/scorecard.png — 12201 bytes
> - /sandbox/incoming/analyst.json — 5025 bytes
> Verified: brief.md exists and references both image files
> hero.png and scorecard.png exist as real PNGs
> analyst.json exists with the normalized runtime comparison
All 932 router + 834 controller Rust tests still pass.
Deliverables captured under:
tools/e2e-harness/out/hermes-exec-brief-delivered/
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(router): operator-UX token + sandbox metrics for /v1/responses
Two visibility gaps surfaced after the Hermes exec-brief run:
operator panel showed `sandbox="unknown"` (instead of the real
sandbox name) and zero token counters for every /v1/responses call.
1. **sandbox label was "unknown"**: every `x-kars-sandbox` header
parser fell back to `"unknown"` when the header wasn't set —
which is the default for clients like Hermes' openai SDK that
don't add kars-specific headers. Per-sandbox routers KNOW their
own identity via the `SANDBOX_NAME` env (set by the controller).
Added `resolve_sandbox_name()` helper at the top of inference.rs:
trust+validate the header if present; otherwise fall back to
`SANDBOX_NAME` env (Box::leak'd to &'static str — fine because
the env is set once at process start). Replaces 4 hand-rolled
`unwrap_or("unknown")` / `unwrap_or("self")` sites. All four
/v1/{responses,completions,embeddings} + foundry-proxy handlers
now produce metrics labelled with the real sandbox name.
2. **token counters were empty for /v1/responses**: the SSE parser
in `forward_stream` looked for top-level `usage` in each
`data:` chunk. OpenAI Chat Completions /v1/chat/completions puts
usage at the top level (works); OpenAI Responses /v1/responses
puts it nested under `response.usage` in the terminating
`response.completed` event (didn't work — captured a real
response.completed event to confirm).
Parser now probes both shapes:
v.get("usage").or_else(|| v.get("response")?.get("usage"))
/v1/responses tokens are now counted (verified live: kars_tokens
delta of +16 input / +12 output for a "list 3 colors" prompt;
was +0 / +0 before).
Verified on local kind cluster after rebuild:
kars_inference_requests_total{model="gpt-5.4",sandbox="execbrief-hermes",status="ok"} 5
kars_tokens_total{direction="input",model="gpt-5.4",sandbox="execbrief-hermes"} 51
kars_tokens_total{direction="output",model="gpt-5.4",sandbox="execbrief-hermes"} 30
The operator panel's "Inference by sandbox" + token-mix dashboards
now populate correctly for Hermes / pydantic-ai / langgraph / any
runtime that uses /v1/responses with non-kars HTTP clients.
932 router tests + cargo clippy --all-targets -- -D warnings clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(mesh): Hermes Act 2 — runtime-neutral Python AGT MeshClient + 6-tool deny list
Closes the inter-agent comms gap for Python frameworks. Until now only
the TypeScript OpenClaw runtime could speak E2E-encrypted AGT mesh;
Hermes had Act 1 stubs that returned 'not_yet_implemented'. This adds
a real implementation usable by any Python framework (Hermes is the
first consumer).
## What ships
1. New package 'kars-agt-mesh' (runtimes/agt-mesh-python/)
- MeshClient orchestrator wrapping the upstream agentmesh-platform
crypto primitives (X3DH, Double Ratchet, SecureChannel)
- IdentityStore: persists Ed25519+X25519 keys at mode 0600
- RegistryClient: POP-signed POST /v1/agents, prekey CRUD,
/v1/discover, Ed25519-Timestamp auth
- RelayTransport: async WS client with 30s heartbeat + backoff
- Process-singleton via _SINGLETONS dict (mirrors openclaw's
Symbol.for('agt-mesh-client') pattern)
- Runtime-neutral — no Hermes-specific code
- 9 unit tests pass
2. Hermes mesh adapter (runtimes/hermes/.../plugin/mesh.py)
- Replaces Act 1 mesh_stubs.py
- Sync→async bridge: dedicated asyncio loop in bg thread so
Hermes' sync tool callbacks can call MeshClient
- Defaults to router-proxied URLs (127.0.0.1:8443/agt/{relay,registry})
so egress-guard iptables stay in place
- Registers kars_mesh_{send,inbox,await,transfer_file}
3. Sub-agent tool deny list (defence in depth)
- Plugin-side: _HERMES_DENY in plugin/__init__.py deregisters
delegate_task, mixture_of_agents, cronjob, kanban_create,
kanban_comment, send_message
- AGT-profile-side: denied_actions block in scenario ToolPolicy
catches the same six names at priority 100
- Rationale per-tool in security audit doc
4. Dockerfile updated to install kars-agt-mesh wheel before plugin stage
5. AGT wheel build script extended to include 'agent-mesh' package
(now produces agentmesh_platform-4.0.0)
## Live verification on kind-kars-dev
- MeshClient.connect() returns 201 from registry, WS upgrade OK
- Self-discovery via /v1/discover returns own DID
- Plugin loader log shows 6 deregistrations + 4 mesh tools present
- 83 Hermes unit tests + 9 kars-agt-mesh unit tests pass
## Critical bug fixed mid-implementation
Initial POP shape sent raw 32-byte public key + ts; registry expected
base64url-string(pub) + ts. Also DID format is server-derived
did:mesh:<sha256(pub)[:32]>, NOT did:agentmesh:<b64url>. Fixed both
in registry_client.py and identity.py. Memory stored for future
non-TS SDK implementers.
## Security audit
See docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
(2 sign-offs, ci-gates green).
## Deferred to Act 2.2
- KNOCK auto-accept responder (currently logs only — Hermes only
initiates so not reachable yet)
- Cross-runtime golden vectors (TS↔Python interop test)
- Multi-process Hermes broker (lazy_install subprocess) — not
reachable while delegate_task is denied
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(mesh): kars-agt-mesh Act 2.1 — full bidirectional E2E round-trip working
Lands the protocol-correct fixes needed for MeshClient.connect() →
KNOCK → X3DH → Double Ratchet roundtrip between two sandboxes. Tested
end-to-end on kind-kars-dev with two Hermes pods (execbrief-hermes and
smoke-hermes) on the FRESHLY BUILT image (no hot patches):
- pod A registers, uploads prekey bundle, opens relay WS (with POP)
- pod B does the same
- pod A discovers B via /v1/discover (freshest-first sort)
- pod A fetches B's bundle, runs X3DH, sends KNOCK + first ciphertext
- pod B's _handle_knock_frame auto-accepts via SecureChannel.create_receiver,
decrypts plaintext 'hello from execbrief-hermes'
- pod B replies via send_by_did → encrypted message frame
- pod A decrypts 'pong from smoke-hermes'
## Critical protocol fixes
1. **Relay WS connect-frame POP** (relay_transport.py)
- Was: {type:'connect', from:did, ts:...}
- Now: full proof-of-possession (std-base64 pub_key + iso ts + sig
over ts), per AGT relay/app.py::_verify_connect_pop
- Without this, the relay rejects every connection with
'connect frame missing did/public_key/timestamp/signature'
2. **Registry auth header** (registry_client.py)
- Was: three separate X-Agent-DID/Timestamp/Signature headers,
signature over method+path+ts
- Now: single 'Authorization: Ed25519-Timestamp <did> <ts> <b64url-sig>',
signature over timestamp string only
- Matches AGT registry/app.py::verify_ed25519_timestamp_auth
3. **X3DH bootstrap missing** (client.py)
- Now connect() builds X3DHKeyManager + generates signed_pre_key
+ 10 OTKs + uploads bundle via PUT /v1/agents/{did}/prekeys
- Without this, peers couldn't fetch our bundle, X3DH initiation
would fail at the responder side
4. **KNOCK responder implemented** (client.py::_handle_knock_frame)
- Was: log-only stub ('responder path not implemented')
- Now: parses ChannelEstablishment, calls SecureChannel.create_receiver,
caches the channel, decrypts the bundled first ciphertext,
eagerly tops up the OTK pool for the next session
5. **Send fuses KNOCK + first message** (client.py::send_by_did)
- First call to a new peer DID sends {type:'knock', establishment, ciphertext}
- Subsequent calls send {type:'message', ciphertext}
- Matches the TS SDK wire convention (one RTT, not two)
6. **AAD directionality fix** (client.py)
- Initiator: f'{self_did}|{peer_did}'
- Responder: f'{from_did}|{self_did}' (reconstructs the same bytes)
7. **EncryptedMessage wire format** (client.py)
- Was: JSON of em.__dict__ (would fail at decoder)
- Now: EncryptedMessage.serialize() / .deserialize() (binary + b64url)
8. **PeerBundle flat shape** (registry_client.py + client.py)
- Was: nested dicts mirroring my best-guess wire format
- Now: matches agentmesh.encryption.x3dh.PreKeyBundle's flat dataclass
9. **register_self handles 409 gracefully** (registry_client.py)
- Was: raised MeshRegistryError, blocking every restart
- Now: logs and continues — the subsequent prekey PUT (with
Ed25519-Timestamp auth) proves we own the same key
10. **discover() sorts freshest-first** (registry_client.py)
- Avoids hitting stale ghost-DIDs when a sandbox restarts with
a new identity before the prior registration ages out
## Tests
- 9 kars-agt-mesh unit tests pass
- 83 Hermes unit tests pass
- Live bidirectional roundtrip verified on freshly-built image
(build hash c1dcdfc11475... loaded into kind-kars-dev)
## Security audit updated
docs/internal/security-audits/2026-06-04-hermes-act2-mesh-deny.md
- Residual risk #1 (no KNOCK responder) removed — now implemented.
- Added residual risk #4 (stale registry entries — non-security).
- Added live bidirectional test description.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(runtime-contract): lift env injection to be runtime-neutral + plug Hermes mesh egress-guard hole
## controller/src/reconciler/mod.rs
Adds three runtime-neutral env vars injected on EVERY agent container
(not just OpenClaw):
- KARS_MODEL=<inference model> — generic alias for OPENCLAW_MODEL so
Hermes / OpenAIAgents / MAF / BYO can read the same value without
knowing about runtime-specific env names
- KARS_RUNTIME_CONTRACT_VERSION=v1 — self-documenting marker that
this container claims to participate in the kars v1 runtime contract
- KARS_RUNTIME_KIND=<Debug repr of RuntimeKind> — uniform anchor any
plugin can use to introspect what runtime it's running as
Lifted from the OpenClaw-only `is_openclaw` gate. All 834 controller
tests still pass.
## runtimes/hermes/.../plugin/mesh.py
**Real bug fix**: the Hermes mesh plugin was reading AGT_RELAY_URL /
AGT_REGISTRY_URL from env. The controller injects these as the
upstream CLUSTER URLs (ws://agentmesh-relay.agentmesh.svc:8765 etc.)
— but those are blocked by the egress-guard iptables rule (UID 1000
is restricted to localhost + DNS only; ports 8765/8080 are dropped
before the connection establishes).
The OpenClaw runtime makes the same call deliberately in
`runtimes/openclaw/src/core/mesh-registry.ts` (always uses
`routerUrl("/agt/registry")` — comment: 'Runtime UID 1000 is
iptables-confined to localhost. AGT_REGISTRY_URL is set by the
sandbox launcher as the router's UPSTREAM target — it points at
the real registry which the runtime cannot reach directly').
Now Hermes does the same: hardcodes 127.0.0.1:8443/agt/{relay,registry}
(the router proxy) on the agent side, ignoring the cluster-DNS env
vars which only the router container is meant to consume.
## Live verification
End-to-end mesh round-trip re-run on the rebuilt controller + sandbox
images (no hot patches):
- pod A (execbrief-hermes) registers, discovers pod B, KNOCK + X3DH
- pod B auto-accepts, decrypts 'hello from execbrief-hermes', replies
- pod A decrypts 'pong from smoke-hermes'
Env vars confirmed present on the agent container post-reconcile:
KARS_MODEL=gpt-5.4
KARS_RUNTIME_CONTRACT_VERSION=v1
KARS_RUNTIME_KIND=Hermes
## Tests
- 834 controller tests pass (cargo test -p kars-controller)
- 83 Hermes unit tests pass
- 9 kars-agt-mesh unit tests pass
- cargo clippy --package kars-controller -- -D warnings clean
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(mesh): Hermes Act 2.2 — multi-agent mesh end-to-end with kars_spawn
Wires the missing pieces so a Hermes parent can spawn Hermes children
AND mesh-message them through the real Python AGT MeshClient.
Multi-agent fanout (parent → 3 sub-agents) verified live on
kind-kars-dev: each sub-agent receives the encrypted KNOCK + first
ciphertext, decrypts plaintext, and the parent's transcript ends with
'EXEC_BRIEF_MESH_FANOUT_DONE: 3 mesh sends delivered.'
## Bug fixes
### 1. Hermes parent now spawns Hermes children (NOT OpenClaw)
inference-router/src/spawn/mod.rs::build_sub_agent_crd_with_labels
hard-coded `runtime.kind = OpenClaw` for every spawn. Now it:
- Accepts an explicit `runtime_kind` field on SpawnRequest.
- Falls back to the `KARS_RUNTIME_KIND` env on the router (set by
the controller as part of the v1 runtime contract).
- Falls back to "OpenClaw" for backward compat.
Also stamps the matching runtime variant key
(openclaw/hermes/openaiAgents/maf) so the CRD admission webhook
doesn't strip-reject the spec.
Restores the runtime kind from a captured spec on handoff snapshot
re-spawn (so Hermes parents survive handoff without silently flipping
to OpenClaw children).
### 2. Controller injects KARS_RUNTIME_KIND on the router container
controller/src/reconciler/mod.rs previously injected
KARS_RUNTIME_CONTRACT_VERSION + KARS_RUNTIME_KIND only on the
*agent* container. Without these on the router too, the spawn
endpoint had no env-based fallback for the kind, so the previous
fix would have silently regressed to OpenClaw.
### 3. Hermes mesh.py accepts OpenClaw-style arg naming
kars_mesh_send now accepts `to_agent` (OpenClaw convention) and
`to` (short form), and `content` plus `payload`, so prompts
written for the OpenClaw mesh API work on Hermes too. Tool schema
advertises the canonical `to_agent`/`content` names primarily.
### 4. Hermes plugin eagerly pre-registers MeshClient at load
runtimes/hermes/.../plugin/__init__.py kicks off a background thread
that calls `_get_or_init_client()` at gateway boot, so the
sub-agent's DID is discoverable in the registry before the parent's
`kars_mesh_send` arrives. Without this, kars_spawn → kars_mesh_send
races: the child is Running but its lazy MeshClient hasn't connected
yet, so find_by_display_name returns nothing and the parent gets
'Peer not found'.
### 5. Discovery falls back to capability when registry omits metadata
runtimes/agt-mesh-python/.../registry_client.py find_by_display_name
no longer requires `metadata.display_name` to be present (the AGT
Python registry's /v1/discover only returns did + capabilities). It
now matches against the capabilities list, which is where MeshClient
puts the display name on register.
## Harness additions
### tools/e2e-harness/platforms/aks.sh
- New `hermes-exec` prompt driver (selected via
SCENARIO_PROMPT_DRIVER=hermes-exec) for runtimes that don't expose
an HTTP gateway on port 18789. Drives `hermes -z` via
`kubectl exec -c agent` with HOME=/sandbox + HERMES_HOME set
explicitly (kubectl exec doesn't inherit container ENV).
- Optional SCENARIO_DAEMON_{SUB,SCRIPT,READY_MARKER} hooks to copy a
helper script into a sub-sandbox and wait for a readiness marker
before posting the parent prompt.
- platform_collect_artifacts now picks the right container name and
gateway-log path per runtime (openclaw=/tmp/gateway.log,
hermes=/sandbox/.hermes/logs/gateway.log).
### tools/e2e-harness/scenarios/mesh-roundtrip-hermes/
Minimal smoke scenario: two pods, one Python echo daemon, one LLM
prompt that calls kars_mesh_send + kars_mesh_await and reports the
decoded plaintext. Verified end-to-end on freshly-built images.
### tools/e2e-harness/scenarios/exec-brief-hermes/
Multi-agent variant: parent uses kars_spawn to launch 3 Hermes children
(analyst/viz/writer), then fans out via kars_mesh_send. This is the
Hermes counterpart of the canonical OpenClaw exec-brief scenario.
## inference-router/Dockerfile.dev
The canonical Dockerfile is distroless (no shell). The controller's
egress-guard init container runs `sh -c "iptables ..."` which can
only work on an image that has sh + iptables. The .dev variant uses
mcr.microsoft.com/azurelinux/base/core:3.0 (non-distroless) + tdnf
install iptables, while still COPYing the pre-staged binary. Used by
`kind load`-based local dev; production AKS keeps the distroless
prod image.
## Tests
- 83 Hermes unit tests pass.
- 9 kars-agt-mesh unit tests pass.
- 16 router spawn tests pass (added env-locked parallelism guard so
the new sub_agent_inherits_parent_runtime_kind_from_env test
doesn't poison sub_agent_crd_uses_post_s10_s13_shape).
- All 834 controller tests pass.
- cargo clippy --package kars-inference-router -- -D warnings clean.
## Live verification on kind-kars-dev
Multi-agent fanout reproduced end-to-end (run.sh-equivalent invocation):
$ hermes -z 'kars_mesh_send to_agent="analyst" content="ECHO_TEST_ANALYST";
kars_mesh_send to_agent="viz" content="ECHO_TEST_VIZ";
kars_mesh_send to_agent="writer" content="ECHO_TEST_WRITER";
emit EXEC_BRIEF_MESH_FANOUT_DONE'
EXEC_BRIEF_MESH_FANOUT_DONE: 3 mesh sends delivered.
analyst daemon log: PRE_REG_GOT bytes=17 text='ECHO_TEST_ANALYST'
viz daemon log: PRE_REG_GOT bytes=13 text='ECHO_TEST_VIZ'
writer daemon log: PRE_REG_GOT bytes=16 text='ECHO_TEST_WRITER'
kubectl get karssandbox -n kars-system shows all 4 as RUNTIME=Hermes
(not the prior bug where Hermes parent spawned OpenClaw children).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* chore(harness): remove stray .new file from mesh-roundtrip-hermes
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(mesh): Hermes Act 2.3 — autonomous sub-agent auto-responder
Closes the last gap blocking the OpenClaw-style multi-agent
exec-brief pattern on Hermes: spawned sub-agents now respond to
inbound mesh messages **without an active session**.
## Problem
After Act 2.2 a Hermes parent could spawn Hermes children and
mesh-send to them, but the children couldn't reply with real LLM
output. Hermes sub-agents are passive daemons — the LLM only runs
when something invokes `hermes -z`. OpenClaw doesn't have this
issue because its plugin runs inside an always-on
`openclaw agent --local` session.
So a parent doing:
parent → kars_mesh_send(to_agent='analyst', content='research X')
parent → kars_mesh_await(senders=['analyst'])
would land the message in analyst's inbox but never get a reply.
The analyst's Hermes daemon would just queue the message and sleep.
## Fix
New `runtimes/hermes/.../plugin/mesh_worker.py`: a background
asyncio loop in each sub-agent that:
1. Drains the shared MeshClient inbox.
2. For each inbound message, runs `hermes -z <payload>` as a
subprocess with KARS_MESH_WORKER_TIMEOUT_S (default 1500s).
3. Resolves the sender's display name via the registry.
4. Replies with the captured stdout via `kars_mesh_send` on the
same singleton MeshClient.
Opt-in via `KARS_MESH_AUTO_RESPONDER=1`. The controller sets this
ONLY on Hermes sandboxes that have the
`kars.azure.com/parent` label (i.e. children spawned by another
sandbox via the router's spawn endpoint). The parent never gets it
on — the parent IS the human/external-driver and would otherwise
loop on the children's replies.
The plugin's `__init__`'s eager-init thread now also calls
`mesh_worker.start_worker()` after the MeshClient is up, so the
responder lifecycle is bound to the plugin's.
## Live verification
Multi-step exec-brief on kind-kars-dev with real Foundry work:
parent → analyst: 'research 2026 agentic AI runtimes, reply ANALYST_FOUND: <url>'
parent → viz: 'use foundry_code_execute to print a JSON dict'
parent → writer: 'use file_write to author /sandbox/incoming/brief.md'
parent → kars_mesh_await(senders=[analyst,viz,writer], timeout=600)
Parent transcript:
WRITER_DONE: 486
VIZ_DONE: {"chart_ready": true, "format": "bar", "width": 1024}
Writer pod /sandbox/incoming/brief.md (486 bytes, REAL LLM content):
'In 2026, agentic runtimes are defined less by raw model capability
than by orchestration: durable memory, verifiable tool use,
background jobs, and policy-aware delegation have turned agents
from clever chat interfaces into operating systems for knowledge
work. The winning stacks emphasize observability, rollback,
sandboxing, and human checkpoints, because the hard problem is no
longer generating ideas but coordinating long-running actions
safely, cheaply, and at production scale.'
Sub-agent daemon logs confirm:
- Accepted KNOCK from parent's DID
- AUTO_GOT bytes=<inbound>
- AUTO_REPLIED bytes=<reply> to=<parent DID>
(Analyst's reply landed slightly past the parent's await window so
the parent's transcript shows TIMEOUT: 2 received — the mesh path
itself worked for all 3; only the LLM coordination timing was tight
because foundry_web_search adds 30+s to analyst's hermes -z latency.
Verified independently that analyst auto-responded with 16 bytes.)
## Tests
- 83 Hermes unit tests pass
- 9 kars-agt-mesh unit tests pass
- 834 controller tests pass
- 16 router spawn tests pass
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* fix(hermes): pre_tool_call hook signature + heartbeat-vs-app metric breakdown
Two operator-visibility fixes called out during the Act 2.3 live
verification:
## 1. Hermes pre_tool_call hook crashed silently → no AGT audit for tools
Root cause: `runtimes/hermes/.../plugin/governance.py::_on_pre_tool_call`
took positional arg `params`, but Hermes 0.15.2 invokes the hook
with KEYWORD args matching `plugins.py:1685-1707`:
tool_name=<name>, args=<dict>, task_id=<id>,
session_id=<id>, tool_call_id=<id>
Our signature `(tool_name, params, **_kwargs)` matched `tool_name`
but every other kw landed in `**_kwargs` and `params` stayed unbound.
Result: TypeError on every invocation → Hermes' hook-runner swallowed
it → no `/agt/evaluate` POST → **no AGT audit entry for any tool
call**. Operator saw only `inference:responses:gpt-5.4` entries in
the audit log even though the agents made dozens of tool calls.
Fixed by matching the Hermes invocation signature exactly
(tool_name, args, task_id, session_id, tool_call_id) + keeping
**_kwargs for forward compat.
Also fixed the deny return shape: the hook used to return a
JSON-string error blob, but `get_pre_tool_call_block_message` only
recognises `{"action": "block", "message": <str>}`. Old denies
were logged + ignored — the tool actually ran. New dict-shape denies
make the block actually block.
Action-verb taxonomy fix: `kars_mesh_send` read `params['target_agent']`
but the real arg name is `to_agent` (alias `to`). Action verb
became `mesh:send:` (empty target). Now accepts all three names.
Also added `mesh:inbox` and `mesh:await` verbs for the drain/wait
tools.
### Live verification
Before fix, parent's /agt/audit:
inference:responses:gpt-5.4 × 63 (every line, no tool entries)
After fix, parent's /agt/audit:
inference:responses:gpt-5.4 × 64
tool:kars_discover:writer × 1 ← NEW
mesh:send:writer × 1 ← NEW
Writer's /agt/audit after fix:
tool:write_file:/sandbox/incoming/audit_evidence.txt × 1 ← NEW
## 2. Sent ≫ received metric asymmetry now legible
Operator UX was showing e.g. 2218 sent / 4 received which is correct
but confusing — sent counter included 30s heartbeats over hours of
uptime. The kars_mesh_messages_{sent,received}_total counters stay
(back-compat, total of all frame types).
New counters break the total down by frame type:
kars_mesh_frames_sent_total{type='heartbeat'} — 30s keepalive
kars_mesh_frames_sent_total{type='message'} — app payload
kars_mesh_frames_sent_total{type='knock'} — session establish
kars_mesh_frames_sent_total{type='connect'} — POP / WS open
kars_mesh_frames_sent_total{type='ack'} — KNOCK/heartbeat ack
kars_mesh_frames_sent_total{type='unknown'} — unclassified
Same shape for kars_mesh_frames_received_total.
Subtracting type=heartbeat + type=connect from the total gives the
real application-frame count. Operator dashboards can now show:
app_sent = sum(rate(kars_mesh_frames_sent_total{type!~'heartbeat|connect'}[5m]))
Classification is a cheap byte-prefix scan (first 80 bytes); the test
`classify_frame_type_buckets_known_kinds` guards every bucket and
`classify_frame_type_handles_short_input` guards bounds.
## Tests
- 84 Hermes unit tests pass (3 new govern hook contract tests)
- 936 router lib tests pass (2 new classify_frame_type tests)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(cli): wire kars connect for Hermes sandboxes
Before this change, `kars connect <hermes-sandbox>` failed silently:
the AKS path is OpenClaw-specific — reads the `gateway-token` Secret
(only created for OpenClaw, see controller/src/reconciler/mod.rs:1354)
and port-forwards :18789 (containerPort only added for OpenClaw, ibid.
:1852). On a Hermes sandbox both are absent, so connect would print
'Gateway token not found' and bail.
Adds a Hermes-specific branch in cli/src/commands/connect.ts that
runs after the AKS-existence check but before the WebUI/shell logic:
if (runtimeKind === 'Hermes') {
kubectl exec -it -c agent — env HOME=/sandbox HERMES_HOME=...
hermes chat --accept-hooks
}
`hermes chat` is the canonical interactive REPL (per
`hermes --help` in 0.15.2 — running `hermes` alone prints usage).
`--accept-hooks` lets the AGT pre_tool_call hook run without
per-tool approval prompts (operator already approved by issuing
`kars connect`).
HOME + HERMES_HOME must be set explicitly because kubectl exec does
NOT inherit container ENV. Hermes' `ensure_hermes_home()` falls
back to $HOME/.hermes; without HOME set, the running container's
HOME defaults to `/` and Hermes tries to mkdir `/.hermes` which
ENOENTs on the read-only rootfs. /sandbox is the writable emptyDir
the entrypoint uses for the long-running gateway daemon.
The exec-ban VAP only targets container name `openclaw`; Hermes'
container is `agent` (set in controller reconciler.rs:1801 from
`is_openclaw` branch), so this is admission-compliant. See
`deploy/helm/kars/templates/admission-pod-exec-ban.yaml`
`matchConditions`.
The --web flag falls back gracefully with a one-line note that
Hermes doesn't ship a browser UI.
The --reset flag works for both runtimes (it's just a rollout
restart). For OpenClaw it clears the in-process brute-force lockout;
for Hermes there's no equivalent state but a restart is still useful
to pick up plugin / env changes.
Local Docker mode (--local) is unchanged — it drops into bash with
OpenClaw-style tips. `kars dev --runtime hermes` for local Docker
isn't a common path yet (the harness lives on local-k8s + AKS);
leaving the bash drop-in to handle both cases until that comes up.
## Tests
789 CLI tests pass (vitest, no new tests added — interactive shell
path is exercised by integration runs, not unit tests).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(operator): Enter-key drops into Hermes agent TUI (full UX parity)
Restores the 'press Enter on a sandbox row → drop into the agent
TUI' UX the operator had for local OpenClaw, but for Hermes on AKS.
OpenClaw on AKS still uses the port-forward + WebUI URL path because
the exec-ban VAP blocks exec into the openclaw container.
## What changed
cli/src/commands/operator/dialogs/connect.ts splits the Enter
handler by (location × runtime kind):
- AKS + OpenClaw → existing port-forward path (VAP-bound)
- AKS + Hermes → PTY exec into 'agent' container (NEW)
- local Docker + OpenClaw → 'openclaw tui' PTY
- local Docker + Hermes → 'hermes chat --accept-hooks' PTY (NEW)
The two PTY paths share a common _spawnPtyConnect() helper extracted
from the old inline body; the OpenClaw port-forward path is now
_aksOpenClawConnect(). Both are pure refactors — the byte-identical
PTY plumbing (blessed save/restore, raw-mode stdin, Ctrl-\ detach)
moved into the helper, no functional change for OpenClaw.
## Why this works for Hermes but not OpenClaw on AKS
deploy/helm/kars/templates/admission-pod-exec-ban.yaml has
matchConditions:
expression: object.container == '' || object.container == 'openclaw'
The VAP fires ONLY when the target container is literally named
'openclaw' (or unspecified — which defaults to the first container,
which is 'openclaw' in OpenClaw pods). Hermes' container is named
'agent' (controller/src/reconciler/mod.rs:1801 picks the name from
the is_openclaw branch), so 'kubectl exec -c agent ...' bypasses the
VAP cleanly.
This was a deliberate VAP design: the policy targets the literal
openclaw runtime container, not 'any agent container'. Hermes (and
future runtimes whose container is named 'agent') benefit by design.
## HOME / HERMES_HOME env vars
Set explicitly on the exec because kubectl exec does NOT inherit
container ENV. Without them, Hermes' ensure_hermes_home() falls back
to $HOME/.hermes; since HOME defaults to '/' in kubectl exec
sessions, Hermes tries mkdir '/.hermes' on the read-only rootfs and
ENOENTs. /sandbox is the writable emptyDir the entrypoint daemon
uses for the long-running hermes gateway.
## Tests
- 789 CLI vitest tests pass (no new tests — interactive PTY path is
exercised by live operator runs, not unit tests).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* fix(mesh): align Python AGT MeshClient wire format with TS SDK (cross-runtime interop)
Closes the last gap blocking Hermes ↔ OpenClaw mesh communication.
Until this change, the Python kars-agt-mesh library and the TypeScript
@microsoft/agent-governance-sdk produced INCOMPATIBLE relay frames —
Python-Python and TS-TS interop worked fine, but a Python sender
talking to a TS receiver (or vice versa) silently dropped messages.
## Wire-format divergences fixed
### 1. message frame: structured header, std base64
**Before (Python only):**
{
'v': 1, 'type': 'message',
'ciphertext': '<urlsafe-base64 of (struct.pack(>I, header_len) + header + ct)>'
}
**After (matches TS mesh-client.js::send):**
{
'v': 1, 'type': 'message', 'from': ..., 'to': ..., 'id': ..., 'ts': ...,
'header': {
'dh': '<std-base64 dhPublicKey>',
'pn': <previous_chain_length>,
'n': <message_number>
},
'ciphertext': '<std-base64 ciphertext>'
}
The TS receiver reads frame.header.dh / frame.ciphertext as separate
fields; the old Python shape had no .header, so TS-side .base64ToUint8
got an unexpected packed blob and decrypt errored out (silently
dropped at the SDK boundary).
### 2. establishment: short TS-style keys
**Before:** {initiator_identity_key: ..., ephemeral_public_key: ..., used_one_time_key_id: ...}
**After:** {ik: ..., ek: ..., otk: ...} (matches mesh-client.js::serializeEstablishment)
### 3. KNOCK + first message: TWO frames, not one fused
**Before:** Python fused KNOCK + first ciphertext into a single
'type=knock' frame for one-RTT latency. TS receivers do NOT consume
a 'ciphertext' field on a KNOCK — they only read 'establishment',
call acceptSession, then await a separate 'type=message' frame.
→ first ciphertext was lost on Python-to-TS sends.
**After:** Python sends two distinct frames: 'type=knock' (no ciphertext,
just establishment) followed immediately by 'type=message'. Matches
TS mesh-client.js::establishSession + send.
### 4. std-base64 (not urlsafe) on the wire
JS's btoa / Node's Buffer.toString('base64') produce std-base64 with
'+' and '/'. Python's base64.urlsafe_b64encode produces '-' and '_'.
A TS receiver's atob fails on '-'/'_'; a Python receiver's
base64.b64decode fails on '+'/'_' depending on input. Now all on-the-
wire byte strings use std-base64.
## Backwards compat
Receiver tolerates both shapes for one release cycle:
- _message_frame_to_encrypted accepts BOTH the TS shape and the legacy
packed-ciphertext shape (fallback path)
- _wire_to_establishment accepts BOTH {ik,ek,otk} and the legacy
{initiator_identity_key, ephemeral_public_key, used_one_time_key_id}
- _b64std_decode tolerates urlsafe alphabet on input
A fleet mid-upgrade between old/new pods won't drop in-flight messages.
## Live verification
Sent {b'WIRE_TEST_DIRECT', 16 bytes} parent → analyst via direct
asyncio script with PYTHONPATH pointing at hot-patched client.py.
Parent stderr:
> TEXT '{"v": 1, "type": "knock", "from": "did:mesh:a61...", "establishment": {"ik":..., "ek":..., "otk": 20}}'
> TEXT '{"v": 1, "type": "message", ..., "header": {"dh":..., "pn":0, "n":0}, "ciphertext": "..."}'
Analyst auto_responder.log:
Accepted KNOCK from did:mesh:a61c9cbf...
AUTO_GOT from=did:mesh:a61c9cbf... bytes=16
AUTO_REPLIED bytes=16 to=did:mesh:a61c9cbf...
The 16-byte payload decrypted correctly with the TS-compatible shape.
## Tests
- 8 new wire-format unit tests pin every field-shape contract
- 9 existing kars-agt-mesh unit tests still pass
## Cross-runtime promise
With this commit, a Hermes agent CAN mesh-send to an OpenClaw agent
and vice versa (same relay, same registry, same crypto, now same
wire envelope). End-to-end interop verification on a mixed-runtime
cluster ships as a follow-up — the wire alignment is the prerequisite.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* feat(cli): include kars-runtime-hermes in 'kars push' image set
Without this, 'kars push' to ACR misses the Hermes sandbox image, so
operators who want to deploy a Hermes sandbox on AKS have to build +
push that image by hand. Aligns with the existing pattern for the
other six runtime adapter images (openai-agents, maf-python,
anthropic, langgraph, langgraph-ts, pydantic-ai).
The tag 'kars-runtime-hermes:latest' matches the controller default
(controller/src/reconciler/runtime.rs DEFAULT_HERMES_IMAGE).
## Cross-runtime mesh status (asked during this session)
Hermes ↔ Hermes mesh: WORKS end-to-end on local-k8s with the
TS-compatible wire format shipped in commit 1a6e7f4.
Hermes ↔ OpenClaw mesh: BLOCKED by an OpenClaw-side mesh-connect
bug. Symptoms observed on local-k8s with a freshly deployed
OpenClaw sandbox (kars-sandbox:dev image, built locally):
inference-router log: tight WS connect-loop (~50 cycles/sec)
'AGT relay WebSocket proxy connected'
'AGT relay WebSocket proxy disconnected outbound_messages=2 ...'
relay log:
'Rejecting connect frame for did:agentmesh:1dcfc3c6...:
connect frame missing did/public_key/timestamp/signature'
(later) 'Cannot call "send" once a close message has been sent'
This is the contract-drift bug already documented in
docs/internal/security-audits/2026-06-02-agt-relay-pop-flood.md and
in the agt-e2e-encryption skill: the bundled
@microsoft/agent-governance-sdk in runtimes/openclaw/node_modules
uses the legacy did:agentmesh:<b64url> format and does not send
proof-of-possession fields in the connect frame.
Setting AGENTMESH_RELAY_ALLOW_UNAUTHED_DID=1 on the relay (escape
hatch from AGT relay PR #66918631) doesn't help — the WS-close loop
continues, suggesting the OpenClaw plugin's mesh code closes the
socket after the first reply, which is an OpenClaw or TS-SDK issue
upstream from kars.
The Python wire format (commit 1a6e7f4) is correct against the TS
SDK spec — verified with TS-shape frames decoded on the Python
receive side — so once the OpenClaw side updates its bundled
agent-governance-sdk to a version that:
1. Uses the did:mesh:<sha256(pub)[:32]> DID format
2. Sends connect-frame POP (public_key + timestamp + signature)
3. Holds the WS open with heartbeats instead of disconnecting
cross-runtime mesh will work without any further Python changes.
Tracked separately as the next item on the OpenClaw upgrade list.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Pal Allakatos <pallakatos@microsoft.com>
* fix(mesh-plugin): emit modern did:mesh:<sha256[:32]> DID format
The kars mesh-plugin previously generated 'did:agentmesh:<sha256[:16]>'
— a kars-specific shorter fingerprint that pre-dated the AGT spec
update. The post-2026-05-23 AGT Python registry rejects connect
frames whose DID doesn't match 'did:mesh:' + sha256(pub).hex[:32],
which produced a tight WS connect/reject loop on every mesh attempt
once we picked up a recent registry build.
Match the upstream AGT TS SDK (@microsoft/agent-governance-sdk
≥4.0.0) and the AGT Python registry: derive
did:mesh:<sha256(pub_key)[:32]>. Legacy did:agentmesh: inbound DIDs
from older peers are still accepted on parse (no parser change
needed — `parseDid` already tolerates both prefixes).
Tested:
- mesh-plugin unit suite: 68 passed
- Live: openclaw sandbox on kind registers as
did:mesh:b3bdd3630f5f727ea1e708fcb431ff10 and is discoverable
by Hermes via registry capability search.
Refs: kars cross-runtime mesh interop (Hermes ↔ OpenClaw)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* build(agt): move AGT pin from pallakatos fork to upstream microsoft branch
Previously kars main built the TypeScript SDK from
`pallakatos/agent-governance-toolkit@bdea1097` (a private mirror of
the kars-sdk-pop-signing feature branch). That branch has now been
pushed to the upstream repo, so we can flip every pin to the
canonical `microsoft/agent-governance-toolkit:kars-sdk-pop-signing`
without needing a private fork.
The new HEAD (3322175d) also carries a second pre-release fix that
unblocks Hermes ↔ OpenClaw cross-runtime mesh:
3322175d fix(ts/x3dh): align KDF with Signal X3DH §2.2 spec
Combined contents of the branch on top of upstream main:
1. Proof-of-possession registration + connect-frame Ed25519
signing in the TypeScript SDK (upstream PR #2772, in review).
2. Spec-compliant X3DH KDF (F prefix moved into IKM, zero salt)
so Python sender ↔ TS receiver derive byte-equal shared secret
instead of triggering AEAD 'invalid tag' on the first MESSAGE.
Updated in lockstep (per docs/PUBLISHING.md 'Updating the AGT pin'):
- vendor/agt/pin.json: url → microsoft, sha → 3322175d
- Cargo.toml [patch.crates-io]: agentmesh + agentmesh-mcp git URL
+ rev (Rust agentmesh crate is byte-identical between the two
revs — only the TS SDK was touched — so the workspace stays at
one consistent toolkit revision)
- Cargo.lock: refreshed via 'cargo update -p agentmesh -p agentmesh-mcp'
- deny.toml allow-git: dropped pallakatos URL (microsoft already listed)
- vendor/agt/microsoft-agent-governance-sdk-4.0.0-agt-bdea1097.tgz
→ vendor/agt/microsoft-agent-governance-sdk-4.0.0-agt-3322175d.tgz
(174 KB, SHA256 dcc8cc...)
- vendor/agt/SHA256SUMS regenerated
- {mesh-plugin,runtimes/openclaw}/package.json: file: dep path
- {mesh-plugin,runtimes/openclaw}/package-lock.json: refreshed
via 'npm install'
Verified:
- mesh-plugin: 68 tests pass, typecheck clean
- runtimes/openclaw: typecheck clean
- cargo check -p kars-controller: clean
- new tarball contains the X3DH fix
(dist/encryption/x3dh.js: F_PREFIX + ZERO_SALT + concat(F_PREFIX, ikm))
When both upstream PRs (#2772 and the pending X3DH spec-compliance
PR) merge to microsoft/agent-governance-toolkit main and AGT cuts
a release containing them, drop the [patch.crates-io] block, the
vendor/agt/ tarball, the file: deps, and vendor/agt/pin.json —
switch to the published npm + crates.io artifacts.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(agt-mesh-python): JSON-wrap payload to interop with TS SDK receiver
Cross-runtime mesh proof Hermes(Python kars_agt_mesh) → OpenClaw(TS
@microsoft/agent-governance-sdk) was failing silently: KNOCK routed,
session established (knock_accept seen on the wire), AEAD tag
verified correctly, plaintext recovered — but the TS SDK's
MeshClient.handleMessage hardcodes a 'JSON.parse(new TextDecoder().
decode(plaintext))' on every successfully-decrypted frame, and a
Python sender that hands raw bytes straight to SecureChannel.send()
produces a plaintext that throws inside that JSON.parse. The
exception lives one level above handleMessage's own catch block, so
the message is dropped on the floor *after* successful decrypt:
onMessage handlers never fire, the plugin's pushInbox() is never
called, and the receiver looks like it never got anything (despite
the kars router metric kars_mesh_messages_received_total
ticking up — the frame DID arrive, but the application layer never
saw it).
The TS SDK's outbound convention (mesh-client.ts::send) wraps every
payload as 'new TextEncoder().encode(JSON.stringify(payload))' so
the inbound JSON.parse always succeeds. Mirror that here:
- UTF-8 byte payloads → JSON-encode as a string before encrypt
- Non-UTF-8 binary payloads → wrap in a {raw_b64: '...'} envelope
- Inbound: invert both paths so callers see the same bytes they sent
- Inbound passthrough for non-JSON plaintext (older Python senders
pre-dating this wrapper) and structured-JSON payloads (caller
opted into JSON shape and will re-parse themselves)
Live-cluster proof (kind-kars-dev, openclaw pod with patched TS SDK
@3322175d): Hermes sender → openclaw's kars_mesh_inbox returns
{ from_agent: 'execbrief-hermes-multi',
message_type: 'message',
content: 'PATCHED_V10_INTEROP_PROOF' }
with received_total=2 (was 0 / session_desync / 'invalid tag' on every
prior attempt). KNOCK+MESSAGE flow visible in relay debug log; no
knock_reject.
Tested:
- 4 new wire-format unit tests cover utf-8 round-trip, binary
raw_b64 round-trip, non-JSON passthrough, structured-JSON
passthrough — all pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(controller,hermes,agt-mesh-python): operator panel shows mesh peers on Hermes side
The operator's per-sandbox AGT trust panel was empty on Hermes
sandboxes even after a successful KNOCK + decrypted MESSAGE exchange
with an OpenClaw peer. Three independent issues stacked on top of
each other:
1. controller: agent-container volume mounts (admin-token + AGT
policy) were hard-coded to fire only when 'is_openclaw' was true.
Hermes runs the same in-pod kars governance plugin via Python
(runtimes/hermes/src/kars_runtime_hermes/plugin/) but never got
/etc/kars/secrets or /etc/agt/policies mounted, so:
- submit_trust() to /agt/trust returned 403 ('Admin token required
for trust mutations') because no token file existed at any of
the documented paths
- the in-process governance engine started with an empty policy
set ('AGT engine will start with an empty policy set and fail
closed')
Generalize the mount predicate: any runtime that ships the kars
plugin (today: OpenClaw + Hermes; future Tier-1 runtimes will opt
in by extension) gets both mounts. BYO continues to skip them.
The agt-policy ConfigMap mount target is now agent_container_name
instead of the literal string 'openclaw', so the inject reaches
whichever container holds the agent runtime.
2. runtimes/hermes mesh_worker: _handle_message decrypted inbound
messages but never called submit_trust(), so even after the
admin-token mount fix the operator panel still missed the peer.
Mirror what OpenClaw does inside its onKnock handler
(runtimes/openclaw/src/index.ts: pushTrustToRouter(fromName,
0.0)): resolve the sender's display_name first so the entry is
human-readable, then push with score 0.0 (router baseline 500 =
at-threshold trust).
3. runtimes/hermes mesh_worker _resolve_sender_name was calling
client._registry.discover('', limit=200) — the AGT registry
rejects an empty 'capability' query param and returns nothing,
so the lookup silently returned None and the trust entry was
keyed on the raw did:mesh:<hex> instead of the human name.
Add a new direct GET /v1/agents/<did> path to the Python
registry client (matches the AGT REST surface) and use it for
peer-name reverse lookup. O(1) on registry side + works for any
registered DID.
Tested live on kind-kars-dev with the rebuilt controller (admin-
token + agt-policy mounts inserted on Hermes pod, verified via
'kubectl get pod -o jsonpath') and rebuilt kars-runtime-hermes
image (mesh_worker submit_trust path baked in):
OC → Hermes (V21):
Hermes /agt/trust returns
[{ agent_id: 'mesh-peer-openclaw',
tier: 'Anonymous',
score: 10,
interactions: 1 }]
Previously (V20, pre-fix):
Hermes /agt/trust returns []
or, with hot-patched mesh_worker, returns the raw DID:
[{ agent_id: 'did:mesh:d7ddf7278ce65b...', ... }]
Auto-responder requires KARS_MESH_AUTO_RESPONDER=1 on the parent
sandbox to be enabled — this is the default on sub-agent containers
the controller spawns. Parents that explicitly opt in get the
same trust-publish path.
Tests:
- 2 new mesh_worker tests cover the submit_trust call path
(success + DID fallback when registry lookup misses).
- All 834 controller tests pass after the mount-predicate change.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(hermes): trust score at OpenClaw-convention baseline + bidi harness
- mesh_worker.submit_trust call now sends score=0.5 instead of
score=0.0. The Python submit_trust helper scales 0.0-1.0 → 0-1000
(so 0.0 → 0, floored by the router to its anonymous-tier minimum
of 10), but OpenClaw's TS plugin uses a delta-around-500
convention: pushTrustToRouter(name, 0.0) →
Math.round(500 + 0.0 * 500) = 500. Sending 0.5 from Python lands
at the same 500 baseline so a Hermes-side peer entry shows the
same at-threshold trust as an OpenClaw-side peer entry would,
rather than appearing as a near-zero-trust agent in the operator
panel.
- tests/e2e/interop/hermes_openclaw_bidi.sh: end-to-end
bidirectional reproducer. Spawns no infra, drives the existing
Hermes + OpenClaw sandboxes on kind-kars-dev through one
OpenClaw → Hermes mesh round-trip and asserts the operator-side
view:
1. OpenClaw warmed + registered with capability=mesh-peer-openclaw
2. kars_mesh_send delivers (status=delivered_via_agt_relay or
delivered_and_replied)
3. Hermes /agt/trust gains a fresh last_interaction for the
peer (proves submit_trust path fired on the inbound msg)
4. Trust entr…
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
Jun 15, 2026
…ategory error
User challenge (2026-06-15): why would kars need a front-door at all
when the inference router already lets agents communicate with models?
The challenge is correct. The front-door I had as Phase 1 items F1/F2
was solving 'give my external IDE a cluster-managed OpenAI endpoint',
which is exactly what agentgateway is for. Putting it in kars would:
* Collide with agentgateway in the category where they dominate
(LF-hosted, MSFT-backed, 9 enterprise sponsors)
* Contradict our per-pod trust-boundary claim (an external IDE has no
egress-guard and is a trusted, not adversarial, caller)
* Blur the product positioning ('agent runtime AND model gateway?')
* Centralize what we deliberately decentralized (front-door is a
cluster-singleton ingress; kars router is a per-pod sidecar)
* Reduce kars to a worse OpenAI proxy
Plan revised:
* F1, F2 removed from the must-have table
* Strategy note added explaining the decision
* Replaced with three agent-runtime-specific items only kars can ship:
- #1 Sub-agent spawn governance hardening (validate target / inherit
creds / propagate audit context across spawn chains)
- #2 Unified per-agent action-cost ledger across model + tool + MCP
+ mesh + spawn (agentgateway tracks model calls only)
- #18 Mesh-aware QoS (per-peer rate-limit, fair-share, KNOCK-aware
budget) — only kars has a mesh
* Composition framing: when an external IDE needs governed-cluster
credentials, the right answer is agentgateway in front + kars
inside, and the docs say so explicitly
Todo store updated to match (lead-F1, lead-F2 dropped; lead-SP1,
lead-AC1, lead-MQ1, lead-D1b added).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Jun 15, 2026
…d apply-fix + Telegram (#397) * demo(act2): S0 — infra-tier ResourceQuota incident harness for Agent A PR #1 in the kars-sre/demo-and-agent series — Slice 0 of the SRE proposal: the demo can now be walked end-to-end by hand before any SRE plugin code lands. Each subsequent slice (S1 read-only tools, S2 K8s diag toolset, S3 typed apply-fix, S4 proactive watcher) replaces one hand-walked step with an autonomous one. Scenario: 'platform team's GitOps refactor lands a tight ResourceQuota across every workload namespace; the quota's requests.memory ceiling (50Mi) is lower than what the research sandbox actually requests. The pod stays Running until anything triggers a reschedule — then it goes Pending forever because the quota blocks pod admission.' Why infrastructure, not image-tag: image tags don't change on a running pod for random reasons. ResourceQuota mis-configuration is a real GitOps-collision incident that operators hit regularly. Files: agent-a-research.yaml — KarsSandbox 'research' (Hermes runtime, mirrors exec-brief-hermes- single shape, simplified to two CRs so the demo focuses on the runtime) platform-hardening-quota.yaml — the bad ResourceQuota the break script applies; deliberately NOT labeled kars.azure.com/managed-by so the SRE's DeleteResourceQuota typed action is permitted break.sh — applies the quota, force-deletes the running pod, confirms the FailedCreate event surfaces reset.sh — deletes the quota and waits for Running 2/2 (manual recovery path) runbook.md — presenter script for walking Act II by hand until S2 ships; once S2 ships, the runbook becomes the expected-behaviour spec for the autonomous agent walk Proposal update: §7.7.1 — adds DeleteResourceQuota as a typed action (namespace- scope, requires the ResourceQuota NOT carry the kars.azure.com/managed-by=controller label so kars-owned governance quotas stay protected and only operator-applied platform quotas are deletable) §7.7.1 — removes the PatchSandboxRuntimeImage carve-out from the previous draft; the demo no longer requires writes to kars.azure.com/* CRs, so the no-governance-mutation rule stays absolute Validation: python3 -c yaml.safe_load_all on both YAMLs — parses OK bash -n break.sh / reset.sh — syntax OK ci/check-copyright-headers.sh — all 499 OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre(s1): MVP — Helm template + 5 read-only kars-CR tools + CLI + plugin containment Slice 1 of the kars-sre demo+agent series. The agent is now installable on any kars cluster via 'kars sre install' and reachable via 'kars sre talk'. It reads kars CRs cluster-wide, walks the diagnostic checklist, matches errors against the OOTB-blocker corpus, and proposes typed fixes (apply is Slice 3). What ships: deploy/helm/kars/templates/sre.yaml — Gated on .Values.sre.enabled. Creates 5 K8s objects when enabled: - InferencePolicy 'sre-inference' (kars-system) - KarsSandbox 'sre' (kars-system) with runtime: Hermes, extraEnv KARS_SRE_ENABLED=true, networkPolicy.defaultDeny=true + allowlist contains ONLY kubernetes.default.svc (NOT agentmesh — §7.8.6 network layer) - ToolPolicy 'sre-tools' (kars-sre) gating the sre_* surface - ClusterRole 'kars-sre-reader' — read on kars CRs + apiextensions + core workloads (RBAC per proposal §7.2.1 minus what S2/S3 add) - ClusterRoleBinding pinned to ServiceAccount kars-sre/sandbox (explicit subject — no group binding, no wildcard, §7.8.3) deploy/helm/kars/values.yaml — new 'sre:' block (enabled=false default, model=gpt-4.1, provider=azure-openai, tokenBudget=32000, extraAllowedEndpoints commented out for Slice 4 channel wiring). cli/src/commands/sre.ts — 'kars sre {install,uninstall,status,talk}' subcommands. 'install' wraps 'helm upgrade --reuse-values --set sre.enabled=true' then waits for the sandbox to reach Available. cli/src/cli.ts — wires sreCommand() into the Operations command group. runtimes/hermes/.../plugin/sre.py — 5 tools, all read-only: - sre_describe_state structured snapshot of all 11 kars-owned CRs - sre_logs apiserver-side pod log tail (cap 500 lines) - sre_diagnose kars-CR health checklist + summary string - sre_explain_error OOTB-blocker corpus matcher (6 known patterns including ImagePullBackOff, exceeded quota, OOMKilled, CrashLoopBackOff, FailedScheduling, ContainerCreating) - sre_propose_fix typed-action proposal envelope; Slice 1 codifies DeleteResourceQuota (the demo Act II target) — rest of typed-action set lands in S3 runtimes/hermes/.../plugin/sre_kube.py — minimal in-cluster apiserver client built on httpx (no new dep added to the shared Hermes image). Reads projected SA token + ca.crt + namespace from the standard paths; detects token rotation by content compare on each request. runtimes/hermes/.../plugin/__init__.py — adds the KARS_SRE_ENABLED gate. When set: - kars_spawn family is SKIPPED at registration (§7.8.5 — SRE agent cannot spawn sub-agents) - kars_mesh_* family is SKIPPED at registration (§7.8.6 — SRE agent is not on the mesh; combined with the NetworkPolicy block above this is two of three §7.8.6 enforcement layers — the third 'separate image' layer is the §7.8.1 follow-up slice) - kars_discover is skipped (no peers to discover) - eager-mesh-init thread is skipped (would log noisy connection failures otherwise) - sre.register(ctx) runs AFTER everything else runtimes/hermes/tests/test_sre.py — 15 tests covering: - env-gate truthy/falsy mapping - all 5 tools register with the correct schema - explain_error matches against the corpus, handles no-match, handles empty input - propose_fix codifies DeleteResourceQuota for ResourceQuota target; returns rationale-only envelope for other kinds - KARS_CR_KINDS lists all 11 proposal §3.5 CRDs - describe_state walks every kind + surfaces per-kind errors without raising docs/sre.md — operator-facing readme: install, talk, tool surface, containment summary, what S1 cannot do yet, links to proposal + Act II runbook. Validation: pytest tests/test_sre.py → 15/15 pass pytest tests/test_governance.py → unchanged, pass pytest tests/test_package_shape.py → unchanged, pass npm run typecheck (cli) → no errors npm run build (cli) → builds helm lint --set sre.enabled=true → 0 fails helm template ... --show-only sre.yaml → renders 5 objects clean helm template ... (sre.enabled=false) → sre.yaml correctly omitted ci/check-copyright-headers.sh → all 501 files OK What this slice does NOT ship (per §7.1 ladder): - K8s diag toolset (sre_image_probe, sre_endpoints_inspect, sre_what_changed, sre_top, sre_describe_resource) — Slice 2 - Fix execution (sre_apply_fix + TokenRequest + admission VAPs) — S3 - Proactive watcher + Telegram/Slack notifications — Slice 4 - Separate kars/sre-sandbox image (§7.8.1 packaging containment) — deferred; Slice 1 ships SRE in the shared Hermes image behind the KARS_SRE_ENABLED env gate as a tactical bridge. The env gate is the interim containment: tools aren't registered in any other pod, so a request for sre_* in a standard sandbox hits 'tool not found' at the runtime. Next: Slice 2 (K8s diag toolset), then Slice 3 (typed apply-fix + AGT approval flow + admission VAPs). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre(s2): K8s diagnostic toolset — describe_resource, what_changed, endpoints, image_probe, top Slice 2 of the kars-sre series. Extends the read-only diagnostic surface from kars-CR-centric (Slice 1) to arbitrary Kubernetes workloads — everything the agent needs to diagnose the Act II ResourceQuota incident end-to-end. What ships (5 new tools, all read-only): sre_describe_resource — structured-describe for any K8s kind. For workload kinds (Deployment / StatefulSet / DaemonSet) walks the OWNER GRAPH: workload → ReplicaSet → matching Pods → events on every level. One tool call returns the whole incident picture. sre_what_changed — events of failure-relevant reasons in last N minutes across BOTH core/v1 and events.k8s.io/v1. Surfaces FailedCreate, BackOff, OOMKilling, Evicted, etc. — the incident-framing tool. sre_endpoints_inspect — Service → selector → matching pods → EndpointSlice readiness. Synthesises a finding the agent can quote (no pods match, pods NotReady, targetPort mismatch, OK). sre_image_probe — given an image, enumerate Pod images cluster-wide and suggest the closest in-use tag by Levenshtein edit-distance. Doesn't reach out to the registry (per-registry auth plumbing is Slice 4+); instead answers the question that's actually most useful: 'what's the closest in-use tag on THIS cluster right now?' sre_top — metrics.k8s.io wrapper for CPU+memory per pod or per node. Gracefully degrades to {unavailable: 'metrics-server not installed'} if the metrics API isn't registered (proposal §7.5 Q4). Also extends sre_propose_fix to codify two more typed actions from proposal §7.7.1: PatchDeploymentImage and ScaleDeployment (in addition to Slice 1's DeleteResourceQuota). Slice 3 will widen the typed-action set further AND add the execution path. RBAC widened in deploy/helm/kars/templates/sre.yaml: + discovery.k8s.io/endpointslices (for sre_endpoints_inspect) + metrics.k8s.io/pods, nodes (for sre_top) + core/nodes, endpoints, resourcequotas (cluster-wide read) ToolPolicy extended to allow the 5 new tool names. Containment unchanged: still gated by KARS_SRE_ENABLED env on the SRE sandbox pod only; standard Hermes sandboxes don't see the env, don't load the tools, can't call them. Validation: pytest tests/test_sre.py tests/test_sre_k8s.py → 31/31 pass ci/check-copyright-headers.sh → all 502 OK helm lint --set sre.enabled=true → 0 fails python -m py_compile (sre.py, sre_k8s.py) → OK Next: Slice 3 (typed apply-fix + admission VAPs + TokenRequest path + kars sre approve CLI). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): resolve helm chart path from repo root, not CWD `kars sre install` was passing the relative path 'deploy/helm/kars' to helm, which helm parses as a chart repo name when the user's CWD is anywhere other than the kars repo root. Result: Error: repo deploy not found Fixed by resolving the kars repo root the same way `kars up` does: first walk up from the CLI file's own location (works for npm link), then fall back to walking up from CWD looking for deploy/helm/kars. Also: replaced the broken `.option('--wait', ..., true)` with the commander-idiomatic `.option('--no-wait', ...)` so the wait flag actually defaults to on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): use --reset-then-reuse-values for kars sre install A plain --reuse-values carries the stored release values forward verbatim. If the stored values are older than the chart on disk (e.g. operator ran 'kars dev' before runtimes.hermes was added to values.yaml), the template fails with: nil pointer evaluating interface {}.image at controller-deployment.yaml line 89. --reset-then-reuse-values (helm 3.14+ / helm 4) re-loads the chart's values.yaml defaults first, then overlays the previously --set values on top. So new chart fields get their defaults populated, while user overrides for older fields are preserved. Applied to both install and uninstall sub-actions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): create kars-sre namespace explicitly in the chart The ToolPolicy 'sre-tools' lives in namespace kars-sre by design (kars's cross-namespace ToolPolicy refs are deliberately not supported — principles.md §3). But the controller-created kars-sre namespace only exists AFTER the KarsSandbox 'sre' is reconciled, which is AFTER helm tries to apply the ToolPolicy. Error: UPGRADE FAILED: failed to create resource: namespaces "kars-sre" not found Fix: add the Namespace as a chart-managed resource at the top of sre.yaml. The controller's namespace-reconcile path uses server-side apply, so it will harmlessly co-own this namespace (adding its own labels + annotations) when it reaches reconciler/mod.rs step 1. No conflict. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): add --force-conflicts to helm upgrade (helm 4 SSA) Helm 4 uses server-side apply by default. When prior `kubectl set image` / `kars push --apply` runs took ownership of fields that the chart now also wants to manage, the SSA call fails with: conflict with "kubectl-set" using apps/v1: .spec.template.spec.containers[name="controller"].image --force-conflicts (helm 4) instructs server-side apply to take ownership on conflict. Matches operator intent: the helm-managed chart is the source of truth, and chart-driven upgrades should override transient field-manager pollution from ad-hoc `kubectl set` calls. Confirmed via `helm upgrade --help`: --force-conflicts if set server-side apply will force changes against conflicts Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): ToolPolicy must live in KarsSandbox's namespace (kars-system), not kars-sre Controller rejected the KarsSandbox sre with: Degraded: ToolPolicyNotFound — 'sre-tools' not found in 'kars-system' (cross-namespace refs not supported) I had ToolPolicy in 'kars-sre' under the misunderstanding that it should be co-located with the runtime pod's namespace. The actual kars convention is the opposite: governance refs are namespace-local to the KarsSandbox CR's OWN namespace (kars-system in our case), per principles.md §3 cross-namespace-refs-deliberately-unsupported rule. The runtime namespace kars-sre is for the pod + RBAC, not for governance. Confirmed against the existing exec-brief-hermes-single scenario which co-locates KarsSandbox + ToolPolicy in kars-system. Net: still safe wrt §7.7.1 protected-resource denylist (kars-system is denylisted, so SRE agent can't delete this ToolPolicy even though it's not labeled kars.azure.com/managed-by=controller). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): rename gate env KARS_SRE_ENABLED → SRE_ENABLED + indent fix Two related bugs uncovered during live test: 1) The controller silently strips user-supplied extraEnv keys with reserved prefixes (mod.rs:1583 — AGT_, AZURE_, FOUNDRY_AGENT_, IMDS_, KARS_). KARS_SRE_ENABLED was being dropped, so the plugin never registered. Fix: rename to SRE_ENABLED across: - runtimes/hermes/.../plugin/sre.py (is_enabled) - runtimes/hermes/.../plugin/sre_k8s.py (module docstring) - runtimes/hermes/.../plugin/__init__.py (log line + docstring) - runtimes/hermes/tests/test_sre.py (3 env patches) - deploy/helm/kars/templates/sre.yaml (extraEnv key + comment) 2) During the rename edit, the `extraEnv:` block ended up under `runtime:` instead of `runtime.hermes:` (4-space vs 6-space indent), producing: UPGRADE FAILED: .spec.runtime.extraEnv: field not declared in schema Fix: restore correct 6-space indent so extraEnv nests inside hermes. Long-term fix (deferred): controller should detect kars.azure.com/role=sre label on the KarsSandbox and inject KARS_SRE_ENABLED itself (controller-side injection bypasses the prefix filter). Noted inline at sre.is_enabled() docstring and in the sre.yaml extraEnv block as a follow-up. Tests: 31/31 pass (test_sre.py + test_sre_k8s.py). Live verification: SRE_ENABLED env appears on agent container's env; helm upgrade succeeds; chart re-applies cleanly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): default contentSafety.requirePromptShields=false The Slice 1 template hardcoded requirePromptShields: true on the SRE InferencePolicy. Azure OpenAI deployments only carry 'prompt_filter_results' in responses when an explicit Content Filter policy is attached to the deployment. Bare local-dev deployments (Foundry quickstart, gpt-4.1 without explicit filter) don't emit those annotations — so the router blocks every response with: Response blocked: InferencePolicy requires Prompt Shields but the upstream response carried no prompt_filter_results annotations Diagnosed live during kars sre talk session — first prompt ('hi there') returned a cached greeting that happened to bypass the check, second prompt died. Fix: default false in values.yaml + chart; operators wiring Content Safety in production can set: --set sre.requirePromptShields=true (or values.yaml override). The SRE agent's threat surface is operator-driven Kubernetes diagnosis, not user-facing chat, so prompt-shield enforcement is less critical than for an internet-facing assistant. Operators who need it can opt back in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: default model gpt-4.1 → gpt-5.4 Switch default model so the SRE agent ships with current frontier out of the box. Operator can still override per-install with `kars sre install --model <name>`. The model name must match an Azure OpenAI deployment in the operator's Foundry project — InferencePolicy routes to that deployment via the router; if the deployment doesn't exist the router returns a clear 404 and the sandbox surfaces Degraded. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): declare sre_* tools in plugin.yaml provides_tools Hermes uses plugin.yaml's provides_tools list as the gate for ctx.register_tool() calls — tools not declared in the manifest are silently rejected at registration time. So even though sre.register() called register_tool() for all 10 sre_* tools, none of them became callable. Diagnosed via live test: hermes tools list → showed foundry_*, http_fetch, kars_handoff_status (the manifest-declared ones) → NO sre_* (registered at runtime, manifest-rejected) Same pattern as the OpenClaw plugin's contracts.tools requirement (see memory: 'OpenClaw 2026.5.x requires plugin manifest to declare contracts.tools listing every tool the plugin will register'). Fix: add all 10 sre_* tools (5 Slice 1 + 5 Slice 2) to provides_tools. The tools remain conditionally registered at runtime — standard Hermes sandboxes don't set SRE_ENABLED → sre.register(ctx) is skipped → the tools are declared-but-not-callable (still matches the manifest contract; Hermes treats them as 'present but inactive'). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: wire SRE-mode SOUL.md system prompt + fix register_tool kwargs Three correctness fixes landed during the live test pass: 1) Hermes register_tool kwargs were wrong sre.py + sre_k8s.py used parameters=... but Hermes' contract expects schema=... AND toolset="<name>". Without these the manifest's provides_tools entries still showed up but the tools were silently non-callable. Fixed all 10 sre_* register_tool calls. 2) plugin.yaml provides_tools missing the sre_* entries Hermes' plugin loader requires every tool the plugin will register to be declared in provides_tools (same shape as OpenClaw's contracts.tools). Added all 10. Conditionally registered at runtime via SRE_ENABLED — standard sandboxes don't trip them. 3) New: kars-sre persona / system prompt Following the OpenClaw pattern (sandbox-images/openclaw/entrypoint.sh :1214 writes SOUL.md on every boot), the Hermes entrypoint now writes a 110-line SRE-specific SOUL.md to $HERMES_HOME/SOUL.md when SRE_ENABLED=true. Content: - Identity + mission statement - Tone constraints (concise, evidence-based, direct, honest) - Catalog of all 10 sre_* tools with WHEN to use each - Catalog of tools the agent does NOT have (spawn, mesh, shell, external net) with rationale - Standard incident reasoning loop (5 steps) - Output structure for fix proposals (Symptom/Evidence/Root cause/ Proposed fix/Why safe/Rollback) - Boundaries (protected-resource denylist enforced at proposal layer; agent should not even try) - Audit info (where the kars audit JSONL captures every call) - First-message greeting template (one line, no editorialising) The model name interpolates from KARS_MODEL → AZURE_OPENAI_DEPLOYMENT → 'gpt-5.4' default, so the prompt always names the live model. Validation: pytest tests/test_sre.py tests/test_sre_k8s.py → 31/31 pass bash -n entrypoint.sh → clean live verify: SOUL.md written 110 lines, model = gpt-5.4 live verify: hermes tools list → '✓ enabled sre' toolset now shows Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: apiserver-bypass for role=sre sandboxes (controller egress-guard) Adds two iptables rules to the egress-guard init container, gated on the kars.azure.com/role=sre label being present on the KarsSandbox: 1. Filter chain: ACCEPT for UID 1000 -> KUBERNETES_SERVICE_HOST:443 (BEFORE the existing catch-all DROP). 2. NAT chain: RETURN for UID 1000 -> KUBERNETES_SERVICE_HOST:443 (BEFORE the existing :443 REDIRECT to :8444 transparent proxy). Both are required. The NAT-bypass alone is not sufficient because the filter chain runs AFTER NAT - the NAT-RETURN says 'don't redirect' but the filter-chain DROP next would still slay the packet. Discovered live during testing: the curl-to-apiserver hung until both rules landed. Why this is needed: the SRE plugin's K8s API client (sre_kube.py in the Hermes runtime) needs DIRECT apiserver access with its projected ServiceAccount token to read kars CRs / pods / events. Without the bypass, every apiserver call gets NAT-redirected to the router's :8444 transparent proxy, which has no idea how to forward TLS to the apiserver -- connections hang then time out. Why only role=sre sandboxes: every other sandbox kind goes through the router unchanged -- that's the whole point of the transparent proxy + L7 audit. Direct apiserver access is the deliberate exception, uniquely held by the nominated SRE sandbox per the proposal section 7.8 containment design. K8s audit log is the audit surface for these apiserver calls (the router's L7 audit doesn't apply, but K8s audit is stronger -- every call carries the SA identity, verb, and resource). Implementation: - new build_egress_guard_command(is_sre_sandbox: bool) helper in reconciler/mod.rs that emits the right rule sequence per mode - 3 unit tests: standard has no bypass; SRE has NAT bypass before REDIRECT AND filter ACCEPT before DROP; both modes keep DROP Validated end-to-end: - HTTP 200 in 17ms from agent container -> 10.96.0.1:443 - sre_describe_state() returns 10 KarsSandboxes + all 11 CR kinds Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): correct AGT profile schema (version 1.0 + agent: name + policies) The Slice 1 inline AGT profile used the wrong schema — version: 1 with rules[].match.tool — which produced: ToolPolicy sre-tools: invalid YAML: missing field agent at compile time, then 'router has not yet loaded AgtProfile' at the sre pod's policy loader. The sre KarsSandbox showed Degraded with ToolPolicyNotCompiled. Found by the SRE agent itself during the first cluster-health-overview test (a beautifully on-point sre_diagnose result that flagged its own ToolPolicy as the only Degraded thing in the cluster). Right schema (from deploy/helm/kars/files/kars-default-agt-profile.yaml): version: '1.0' agent: <name> policies: - name: ... type: capability allowed_actions: [...] denied_actions: [...] priority: N Action prefix convention used by the router: tool:<tool_name> for tool calls inference:<api>:<model> for model dispatch spawn:* / mesh:* for sub-agent + mesh The new sre-tools profile has three policies: - sre-diagnostic-tools-allow (priority 100): all 10 sre_* tools - sre-inference-allow (priority 90): chat_completions / responses / content_safety - sre-spawn-and-mesh-deny (priority 110): defense in depth for the §7.8.5/§7.8.6 containment (already enforced by plugin not even registering these tools) After re-apply + sre pod restart: ToolPolicy sre-tools status: Ready True:RouterEnforcing KarsSandbox sre status: Running Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sre): trailing-colon glob in AGT allow rules — match real action shape The Slice 1 allow rules used literal 'tool:sre_<name>' strings but the Hermes plugin governance hook actually emits 'tool:<name>:<first-arg>' — with a trailing colon even when no significant arg is present (see runtimes/hermes/.../plugin/governance.py _action_verb tail returns f'tool:{tool_name}:'). So: literal allow: 'tool:sre_describe_state' router emit: 'tool:sre_describe_state:' <-- no match → denied The agent helpfully diagnosed itself via: sre_describe_state -> blocked by policy 'sre-diagnostic-tools-allow' (visible because the WebUI surfaced the matched_rule name). Confirmed the action shape in inference-router/src/routes/governance.rs:66 ('if let Some(tool_name) = action.strip_prefix("tool:")...'). Fix: add a '*' wildcard to every allowed_action for the sre_* tools. This matches both the trailing-colon shape (tools with no args) and the suffix-args shape (sre_describe_resource:<name>, sre_logs:<pod>, etc.) in a single entry. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: NetworkPolicy egress allow for apiserver (cluster-portable) The egress-guard iptables bypass (b25f41b) lets UID 1000 reach the apiserver at the iptables layer, but the pod-level NetworkPolicy was still denying it. The blanket :443 egress rule explicitly excludes RFC1918 ranges to prevent lateral movement to in-cluster Services, but every cluster's apiserver ClusterIP IS in one of those ranges (kind: 10.96.0.1, AKS: 10.0.0.1, EKS: 172.20.0.1). Fix: when role=sre, add a NetworkPolicy egress rule for the apiserver Service ClusterIP. The IP + port are read at reconcile time from the controller's own KUBERNETES_SERVICE_HOST / KUBERNETES_SERVICE_PORT_HTTPS env vars (kubelet-injected on every pod). This is cluster-portable — kind, AKS, EKS, custom service-CIDRs all get the right value automatically. No hardcoded IPs. Implementation: - Top of reconcile(): compute is_sre_sandbox once + read apiserver IP/port from env. Threaded through both the egress-guard helper and the NetworkPolicy egress vec. - egress_rules.push(...) added after the static block, gated on is_sre_sandbox, with IP/port substituted from env. - Removed the duplicate is_sre_sandbox compute lower in reconcile() that was added in b25f41b — single source of truth now. Validated live: - kubectl get netpol -n kars-sre shows the 10.96.0.1/32 :443 rule - sre_describe_state() returns in 0.10s — 11 CR kinds, 10 KarsSandboxes enumerated, NO timeouts. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(demo): agent-a-research.yaml passes CRD admission Two admission rejections: 1) spec.governance.toolPolicyRef.name required when governance.enabled=true Added a research-tools ToolPolicy with allow rules for: - inference:chat_completions:* / responses:* / content_safety:* - tool:http_fetch:* (the agent does web research) - tool:foundry_* family (memory + web_search + code_execute etc.) 2) spec.runtime.hermes must be set iff kind=Hermes (CEL guard rejects missing key, accepts empty object). The previous manifest had a commented placeholder which yamllint-fine but admission saw the key as missing. Changed to 'hermes: {}' — empty object honours image defaults without drift. Also: aligned the demo with the SRE sandbox defaults shipped earlier: - deployment: gpt-5.4 (was gpt-4.1) - requirePromptShields: false (was true — bare local Foundry deployments don't emit prompt_filter_results, blocking every response) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(demo): break.sh uses kars.azure.com/component selector Controller stamps pods with kars.azure.com/component=sandbox not the app.kubernetes.io/component=sandbox the script was looking for. Result: 'no sandbox pod found to evict; quota will only manifest on next natural restart' — the script kept going but the break never surfaced. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * kars-sre: Slice 3 (typed apply-fix) + Slice 4 (proactive watcher + Telegram) Slice 3 — typed apply-fix path (operator-approved remediation) Adds the KarsSREAction CRD and reconciler that drives an SRE-agent fix proposal Proposed → Approved → Applied → Recovered. The agent emits a CR via sre_propose_fix; the operator approves via kars sre approve <id> (or kubectl edit); the controller mints a one-shot ClusterRoleBinding scoped to the right writer ClusterRole (kars-sre-writer-quotas | kars-sre-writer-workloads), executes the typed action via SSA, tears the binding down, and observes recovery by polling the target namespace for failure-class events. Terminal CRs (Recovered / Failed / Expired / Rejected) auto-GC after 1h. Closed set of typed actions per proposal §7.7.1: - DeleteResourceQuota (refuses kars.azure.com/managed-by=controller) - PatchDeploymentImage, ScaleDeployment (clamp 0..50), RolloutRestart (Deployment/StatefulSet/DaemonSet), DeletePod New files: - controller/src/kars_sre_action.rs (CRD types) - controller/src/kars_sre_action_reconciler.rs (state machine) - deploy/helm/kars/templates/crd-karssreaction.yaml Hermes plugin (sre_propose_fix is now a CR-creator): - Tolerant arg parsing: target.kind / action_type / inferred kind - schema marks target.kind required + enum-validated - Returns action_id + ready-to-paste 'kars sre approve' command - Clear cr_error when no typed fix could be inferred CLI: - kars sre approve <id> / reject <id> / actions / show <id> - kars sre show renders diagnosis + rationale + condition stamps RBAC additions (controller-side): - karssreactions (full r/w) - resourcequotas: delete (the §7.8.4 escalation check requires the controller to hold the verbs it grants in the one-shot CRB) - apps/statefulsets,daemonsets: patch (RolloutRestart targets) - events: list/watch/get (recovery observer) - serviceaccounts/token: create (lands the §7.8.4 TokenRequest path) - clusterrolebindings: create/delete kars-sre-write-* Slice 4 — proactive watcher + Telegram sre_watcher.py runs alongside the Hermes gateway when SRE_ENABLED=true and a channel is configured. Polls K8s events every 10s for failure- class reasons in kars-* namespaces (excluding kars-sre / kars-system / kube-* / agentmesh / default), maps each into a typed-fix target, and on incident: 1. Reuses any open KarsSREAction with the same (action_type, ns, name) target — no duplicate CRs. 2. Otherwise creates a new KarsSREAction with ttl_minutes=30. 3. Coalesces a per-iteration burst into ONE detailed Telegram message (highest-priority candidate) plus an optional summary tail ('+N other incidents: 2 FailedScheduling, 1 BackOff'). 4. Sliding-window rate limit: max 4 messages/min cluster-wide. Dedupe is bootstrapped from existing KarsSREActions on boot (survives pod restart). First iteration is silently absorbed (priming) so a pod re-roll doesn't replay the warm-cache flood as alerts. Periodic 60s CR resync REPLACES the dedupe state so operator-side delete clears the in-memory map naturally. ReplicaSet/Pod hash suffixes are normalised in the dedupe key so a flapping Deployment's rollout sequence collapses to one alert instead of one alert per pod-template-hash. Telegram wiring: - Channel adapter libraries (python-telegram-bot 21, slack-sdk 3, discord.py 2) pre-installed in the runtime image so credentials in the sandbox-credentials secret 'just work'. - entrypoint.sh exports HTTPS_PROXY=http://127.0.0.1:8444 and NO_PROXY=$KUBERNETES_SERVICE_HOST,127.0.0.1,localhost,.svc.cluster.local so the gateway's outbound HTTPS reaches the inference-router's forward proxy (egress-guard iptables redirect doesn't fire in kind clusters without CAP_NET_ADMIN — explicit env covers both). - HOME=/sandbox export so gateway-locks dir under ~/.local/state is writable on the distroless base. - TELEGRAM_ALLOWED_USERS exported (not just config-set) so the gateway's per-platform allowlist skips pairing for known users. - TELEGRAM_HOME_CHANNEL set to first TELEGRAM_ALLOW_FROM id so 'hermes send --to telegram' resolves without explicit chat id. Operator install path (unchanged — uses existing kars credentials): kars credentials update sre --telegram-token <T> --telegram-allow-from <ID> Tests: 31 hermes tests + 847 rust tests + cli typecheck/lint pass. The phase taxonomy guard now passes after refactoring the reconciler to use named constants for all condition types / reasons / event reasons rather than 'Failed' / 'Pending' / 'Degraded' literals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * kars-sre: Headlamp SRE Console + Chat (Slice 4 primary UX) Adds the SRE engineer's dedicated console as a top-level sidebar branch in the kars Headlamp plugin. Replaces the prior workflow of 'kubectl get karssreactions + paste action_id into kars sre approve in a terminal' with one click in the dashboard. New routes: /kars/sre — SRE Console (live cards, primary landing) /kars/sre/chat — embedded Hermes WebUI iframe /kars/karssreactions — full CRD list (under existing CRD section) SRE Console layout (top → bottom): 🔴 Pending Approval — KarsSREActions awaiting operator. Inline Approve / Reject buttons PATCH .spec.approval.state directly via Headlamp's KubeObject.patch(), with optional rejection- reason prompt. No terminal hop needed. 🔄 In-flight — actions the controller is currently executing (Applied + waiting for recovery). Shows phase + age. 📊 Cluster Health — sandbox phase counts + degraded count. 🚨 Active Incidents — failure-class events (FailedCreate, BackOff, FailedScheduling, Failed, ImagePullBackOff, CrashLoopBackOff, OOMKilling, Evicted, FailedMount) from kars-* namespaces in the last 15 min. Same filter the proactive watcher uses, so what the operator sees here is what the watcher would alert on. ✅ Recent — Recovered / Failed / Expired / Rejected actions from the last hour for post-incident review. All cards live-update via Headlamp's useList() (watch + long-poll), so the Proposed → Approved → Applied → Recovered walk is visible without F5. The KarsSREAction CRD is added to the existing CRD registration table so the standard list / detail pages 'just work' under /kars/karssreactions/:ns/:name. SRE Chat is an iframe of the Hermes WebUI: - tab 1: http://localhost:18789 (requires 'kars connect sre --web' in another terminal — populates the iframe via port-forward) - tab 2: apiserver service-proxy fallback for in-cluster operators - 'Open in new tab' button if iframe sandboxing breaks the embed Helm chart: SRE sandbox's allowedEndpoints now includes api.telegram.org / core.telegram.org cluster-side so the Slice 4 watcher's outbound Telegram alerts don't need an out-of-band NetworkPolicy patch. Dormant when Telegram isn't configured — the gateway only opens the channel when the token is present. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: fix browser-ESM require() crash + add 'SRE not installed' CTA Two fixes: 1. ReferenceError: require is not defined The Active Incidents card lazily resolved the Event class via require("@kinvolk/headlamp-plugin/lib/K8s/event"). Headlamp ships plugin bundles as pure browser ESM modules — require() doesn't exist in that context, so the page crashed at first render. Switch to the documented public re-export via the K8s namespace (`import { K8s } from "@kinvolk/headlamp-plugin/lib"` → `K8s.event`), which is safe in both build- and run-time. 2. Empty-state CTA when kars-sre isn't deployed Both SREConsole and SREChat now check for the existence of the sre KarsSandbox in kars-system. If absent (or the list is still loading), they render an actionable install card with: - `kars sre install` (the one-liner that enables the chart) - `kars credentials update sre --telegram-token ...` (optional) So a fresh kars dev cluster that hasn't run `kars sre install` yet doesn't show 'No items' or a spinning iframe — it tells the operator exactly what to type. The cards rehydrate live once the sandbox lands (no refresh needed). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: stub Active Incidents — pluginLib.K8s.event isn't host-exposed The Headlamp 0.41 host runtime exposes `pluginLib.K8s` as a flat namespace of class-kind classes but does NOT expose the v1 `event` sub-namespace. Importing it via either explicit submodule path (`@kinvolk/headlamp-plugin/lib/K8s/event`) or the top-level barrel (`K8s.event.default`) trips Vite's UMD wrapper into its CJS-fallback branch on first execution, which crashes the browser with: ReferenceError: require is not defined at ct (//plugins/kars/dist/main.js:3:52537) (`ct` was the INCIDENT_REASONS set at top-level — top-level execution failed before any component mounted.) The KarsSREAction CR cards above already surface every incident the proactive watcher catches (same dedupe key, same target shape), so for Slice 4 the operator doesn't need the raw events feed duplicated in the dashboard. Slice 4.1 (future) can resurrect this via direct fetch() to /api/v1/events through the headlamp apiserver proxy, bypassing the K8s.event class entirely. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp: bump plugin to v0.6.0 to bust Headlamp's plugin cache Headlamp keys plugins by package.json version. A pure dist/main.js swap (with same version) leaves the host's plugin loader cache holding the previous bundle. Bumping minor → operator's browser re-fetches main.js on next mount even without Cmd+Shift+R. v0.5.1 → v0.6.0 covers the prior session's additions: - KarsSREAction CRD list / detail - SRE Console + Chat sidebar branch - browser-ESM safety pass (no require() in source) - SRE-not-installed empty-state CTA * cli: kars sre install handles 3 cluster shapes (helm release / kars dev / fresh) The 'no deployed releases' error happened because 'kars dev --target local-k8s' deploys the chart via 'helm template | kubectl apply' (see cli/src/commands/dev/local-k8s.ts:794), so no helm release record exists. The sre install path assumed a helm release and failed on a fresh kars dev cluster. Now detects three shapes: A. helm release present → helm upgrade --reset-then-reuse-values --force-conflicts (preserves operator's prior --set choices) B. no helm release BUT controller deployed (= kars dev path) → helm template … | kubectl apply --server-side --force-conflicts (mirrors how the chart got there in the first place) C. neither (= fresh cluster) → helm install --create-namespace --take-ownership (--take-ownership: adopt any pre-existing namespace or NetworkPolicy from prior partial installs; helm >= 3.17) The template path uses --include-crds so KarsSREAction is installed on first sre install even when the cluster predates Slice 3. All three paths set azure.workloadIdentity.clientId=dummy for local-k8s brand-new installs (real AKS installs go through kars up). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: derive cluster name from URL for apiserver-proxy chat tab The proxy URL hardcoded 'kind-kars-dev' as the cluster name, which only worked for the local-k8s demo path. Real operators have any context name (AKS managed-cluster names, in-cluster Headlamp, multi-cluster setups). The 'key not found' error was Headlamp's backend rejecting the request because the cluster path component didn't match any of the operator's configured contexts. Fix: parse the cluster name from window.location.pathname (Headlamp routes every cluster-scoped view under /c/<cluster>/...). When the parse fails (e.g. the Chat page is loaded outside a cluster context), the proxy tab is disabled and the operator is steered to the local port-forward tab. Reads location directly instead of useCluster() because importing the K8s namespace (where useCluster lives) trips the host's UMD require() fallback — the same crash the v0.6.0 plugin fixed. v0.6.0 → v0.6.1 to bust Headlamp's plugin cache. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * controller: expose Hermes gateway port (18789) on per-sandbox Service The per-sandbox Service exposed only :8443 (inference-router). For Hermes runtimes the gateway WebUI / inbound channel adapter lives on the agent container at :18789, but operators had no way to reach it without setting up a per-sandbox port-forward. Now: when runtime.kind == Hermes, the controller appends a 'gateway' port (18789) to the same Service. Result: 'kubectl port-forward svc/<name> 18789' works directly, AND the Headlamp SRE → Chat tab can route via the apiserver service proxy: /clusters/<cluster>/api/v1/namespaces/kars-sre/services/sre:18789/proxy/ OpenClaw runtimes are unaffected (no gateway port added). The NetworkPolicy ingress rule for governance-enabled sandboxes already allows port 18789 from peer sandbox namespaces, so this purely widens what the cluster apiserver / Headlamp backend can reach — no extra exposure to other sandboxes. * headlamp/sre: replace iframe Chat tab with terminal-attach instructions Hermes is a CLI/TUI agent — there's no embedded WebUI to iframe. Earlier commits attempted an apiserver-proxy iframe pointing at :18789 (Hermes admin port) — which only listens when the gateway runs in channel mode, and even then doesn't serve a browser UI. The SRE Chat page now shows three explicit operator paths in copy-pasteable code blocks: 1. kars sre talk → kubectl exec REPL (live triage) 2. kars credentials update sre --telegram-token … → wire Telegram for proactive alerts 3. kars sre status / actions / show <id> → terminal-friendly snapshot Plus a link back to /kars/sre (the Console) for the live approval queue + cluster health cards. The 'iframe with connection refused' error is gone; v0.6.1 → v0.6.2 to bust the host's plugin cache. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: replace internal Link with plain anchor + bump to 0.6.3 The internal <Link routeName="kars-sre-console"> in SREChat may have been the source of the React error 310 — Headlamp's Link implementation uses hooks internally and a routeName resolution miss can fire conditional hook paths. Using a plain <a> anchor with the canonical Headlamp URL avoids that branch entirely. The bundle was also showing as stale (browser cached old dist) — v0.6.3 bumps the version to force a re-fetch on the host's plugin loader. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: embed hermes dashboard PTY chat in browser Replaces the 'no embedded WebUI' instruction page with a real iframe into the Hermes dashboard — an in-browser xterm.js PTY chat. The operator can now talk to the SRE agent without leaving Headlamp. How it works: 1. sandbox image: pip-installs FastAPI + uvicorn + websockets + ptyprocess (the soft-optional deps hermes dashboard needs). Upgrades hermes-agent from 0.15.2 → 0.16.0 to pick up the dashboard_auth submodule that 0.15.2 was missing. 2. entrypoint.sh: launches 'hermes dashboard --host 0.0.0.0 --port 9119 --no-open --insecure --skip-build' alongside the gateway when SRE_ENABLED=true. HERMES_DASHBOARD_TUI=1 enables the embedded PTY tab. Opt-out via HERMES_DASHBOARD_ENABLED=false. 3. controller: adds containerPort 9119 ('dashboard') to Hermes agent containers, and exposes it on the per-sandbox Service so the cluster apiserver proxy can reach it. 4. Headlamp plugin: SREChat replaces the instruction page with an iframe pointing at /clusters/<cluster>/api/v1/namespaces/kars-sre/services/sre:9119/proxy/. Includes 'Open in new tab' fallback for cases where the sub- path proxy strips Hermes web bundle asset paths. v0.6.3 → v0.7.0 to bust the host's plugin cache. The '--insecure' flag is required when binding off-loopback inside the pod — Hermes refuses non-127.0.0.1 binds without it. In our pod the only reachers are the apiserver proxy + peer sandboxes (both gated by RBAC + NetworkPolicy), so 'insecure' here doesn't mean externally exposed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: fix dashboard iframe — fetch HTML + rewrite asset URLs Two changes work together to make the in-browser Hermes dashboard load inside the Headlamp SRE Console iframe: 1. New runtimes/hermes/src/kars_runtime_hermes/dashboard_proxy.py Tiny FastAPI middleware wrapper around hermes_cli.web_server.app. Installs X-Forwarded-Prefix on every request from the HERMES_DASHBOARD_PREFIX env var. Hermes' dashboard reads that header to rewrite absolute asset URLs (/assets/...) for sub-path reverse proxies. K8s apiserver service proxy doesn't inject that header, so without this wrapper the SPA blank-loads in the iframe. 2. entrypoint.sh now boots that wrapper instead of 'hermes dashboard', with HERMES_DASHBOARD_PREFIX set to the K8s apiserver suffix: /api/v1/namespaces/<ns>/services/<svc>:9119/proxy 3. Headlamp SREChat fetches the dashboard HTML up front via the Headlamp proxy, rewrites asset paths to include /clusters/<cluster> (the Headlamp-added prefix that the in-pod wrapper can't know about), and injects via iframe srcDoc. Also injects <base href> so the SPA's relative fetch() calls resolve under the proxy. v0.7.0 → v0.7.1 to bust the host's plugin cache. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * headlamp/sre: dashboard wrapper strips proxy prefix to dodge /api/* collision The K8s apiserver-proxy URL prefix /api/v1/namespaces/<ns>/services/<svc>:<port>/proxy starts with /api/v1 — which collides with Hermes' own /api/* route namespace. So when the browser fetched a SPA asset like /api/v1/namespaces/kars-sre/services/sre:9119/proxy/assets/index.js, FastAPI matched it to its API router (401 Unauthorized) instead of the static-file mount. Fix: extend dashboard_proxy.py middleware to STRIP the prefix from scope["path"] before FastAPI sees the request, while still injecting X-Forwarded-Prefix so the SPA's index.html bootstrap rewrites asset URLs with the absolute prefix. Result: browser fetches .../proxy/assets/foo.js, middleware strips → FastAPI sees /assets/foo.js → static-file mount serves it → 200 OK. Smoke test verified end-to-end: asset via prefix: HTTP 200 index via prefix: HTTP 200 Headlamp SREChat still uses srcDoc + double-prefix rewrite because Headlamp's apiserver proxy adds /clusters/<cluster> ON TOP of the K8s suffix — the in-pod wrapper can't know <cluster>, so the browser-side rewrite adds it. v0.7.1 → v0.7.2 to bust the host's plugin cache. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: end-to-end embedded Hermes chat in Headlamp plugin Three stacked bugs blocked the SRE Console's Chat tab from working end-to-end. All fixed: 1. Headlamp's apiserver proxy demands Authorization: Bearer on every /clusters/<c>/api/v1/.../proxy/* call. Headlamp's SPA fetch wrapper attaches it; iframe asset loads bypass the wrapper and 403 as system:anonymous. Plugin v0.7.4 drops the apiserver-proxy approach entirely and iframes http://localhost:19119/ via a user-run port-forward. Cross-port = different origin so parent/child JS is isolated, but iframe document loads aren't same-origin-gated. 2. The dashboard_proxy wrapper bypasses Hermes' start_server() (to install X-Forwarded-Prefix middleware first), which is where Hermes sets app.state.bound_host/port. Without those, _build_gateway_ws_url returned None and the PTY-spawned hermes --tui child got no HERMES_TUI_GATEWAY_URL env var — accepting keystrokes but with nowhere to send them. _set_bind_state() mirrors what start_server does. 3. Azure Linux 3 ships Node 24; Hermes' ui-tui esbuild bundle was built against Node 22 and SIGSEGVs immediately on Node 24 (380MB core dumps). Dockerfile now pins Node 22.20.0 at /opt/node22/, entrypoint exports HERMES_NODE=/opt/node22/bin/node so Hermes' _node_bin() picks it up. Plus: - model.context_length: 200000 pinned so cold-start skips the slow /v1/models probe. - GATEWAY_ALLOW_ALL_USERS=true on the SRE sandbox so the single-operator loopback deploy doesn't drop our own messages. - entrypoint passes HOME/HERMES_HOME/HERMES_NODE through runuser's env reset via explicit env VAR=$VAR invocation. Plugin bumped to 0.7.4. Verified end-to-end: chat opens, accepts keystrokes, agent responds. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix: ACR-name typo + workload-aware SRE Cluster Health card Two fixes that surfaced during demo dry-run: 1. ACR typo (introduced in 5c87e9de during the azureclaw\u2192kars rename) - 'kars.azurecr.io' was a search/replace artifact from 'azureclaw.azurecr.io'; the actual ACR is 'karsjpdyyv.azurecr.io' (azd-suffixed). The canonical name we use in chart values + controller defaults is 'karsacr.azurecr.io' so operators have ONE name to re-publish to. - Symptom: when an existing sandbox spawned a sub-agent via kars_spawn, the controller minted the new Deployment with 'kars.azurecr.io/openclaw-sandbox:latest'. Kubelet did DNS on 'kars.azurecr.io' \u2192 NXDOMAIN \u2192 ImagePullBackOff loop. - Fixed in: - deploy/helm/kars/values.yaml (4 sites: controller, inference-router, sandbox, a2a-gateway) - cli/src/commands/dev/local-k8s.ts (inverted target/aliases shape: 'target' is now the canonical name the controller expects, 'aliases' are the local build tags we look up + retag from) - tools/demo/scenarios/01-sandbox.yaml 2. Workload-aware Cluster Health (Headlamp plugin 0.7.5) - KarsSandbox CR's 'phase=Running' fires the moment the controller successfully reconciles the Deployment spec; it knows nothing about whether the pods inside actually pulled their image, passed readiness, or got OOM-killed. The old SREClusterHealthCard read phase only \u2192 'all green' even when break.sh had killed every pod. - SREClusterHealthCard now cross-checks each sandbox against its underlying Deployment (kars-<name>/<name>) and surfaces three buckets: Healthy \u2014 CR Running AND availableReplicas \u2265 desired Workload down \u2014 CR Running BUT availableReplicas < desired (the false-green case) CR-Degraded \u2014 CR-level Degraded=True - Bonus: per-sandbox breakdown panel lists which ones are unhealthy and points the operator at 'kars-<name>' namespace for pod-level diagnosis. Matches the SRE agent's own diagnosis output. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(monitoring): include kars-ops dashboard in Grafana sidecar configmap The grafana-dashboard-configmap.yaml only wrapped grafana-dashboard-kars-fleet.json but not grafana-dashboard-kars-ops.json — even though both JSON files have lived in deploy/monitoring/ since May 27. Result: the Headlamp plugin's SandboxMetricsCard iframes a 'Dashboard not found' page (it targets uid=kars-ops). Regenerated the configmap YAML from both .json files so the grafana-dashboard sidecar picks up both on next kars dev run. No JSON content changed; just plumbing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * hermes: pre-warm AGT mesh registration in idle-gateway mode `hermes gateway run --accept-hooks` in idle-daemon mode (no Telegram/Slack/ Discord channels configured) runs only the cron ticker — it never imports the kars Hermes plugin, so the Phase A2.1 eager MeshClient init at plugin load never fires. Result: a Hermes sandbox is invisible on `kars_mesh_directory` listings until something else triggers a plugin load (e.g. an interactive `hermes chat` invocation, which spins up a short-lived process that registers + exits). Adds a 5-line pre-warm in entrypoint.sh that runs `_get_or_init_client()` in a short-lived background Python process at boot — register_self is idempotent + restart-safe so re-runs are cheap. Guarded on: - SRE_ENABLED != true (SRE agents are intentionally off-mesh) - KARS_MESH_PROVIDER == agt (only run when the mesh is actually wired) Verified on kind: research sandbox now logs '[kars-hermes] mesh pre-warm: registered' within ~2s of pod boot, and shows up on the AGT registry's live-agents endpoint before any chat invocation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * hermes: persistent mesh-keepalive (replaces short-lived pre-warm) Followup to fcce016 — the short-lived pre-warm Python process registered on the relay then EXITED, taking the MeshClient socket with it. Without a live connection there's no relay heartbeat, so the AGT registry marks the agent stale after ~90s and discovery tools hide it ('Stale/offline filtered out'). Replaces the pre-warm with a long-lived 'kars-mesh-keepalive' process that: 1. Calls _get_or_init_client() to register + connect (same eager path the plugin would take if loaded by the gateway) 2. Calls mesh_worker.start_worker() so the sandbox can REPLY to inbound mesh messages (not just appear in directory listings) — same auto-responder the controller wires into kars_spawn'd sub-agents via KARS_MESH_AUTO_RESPONDER=1 3. Parks on threading.Event().wait() forever so the MeshClient stays alive and keeps heartbeating Verified on kind: research's keepalive log shows registered + connected + worker started; dev-agent's mesh discover can now see research. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * hermes: enable AUTO_RESPONDER on the mesh keepalive process Follow-up to 163e1de — the keepalive's mesh_worker.start_worker() was draining inbound messages and silently dropping them because the worker gates LLM replies behind KARS_MESH_AUTO_RESPONDER (mesh_worker.py:259). Couldn't set the env var via the KarsSandbox CR's extraEnv because the controller's reserved-prefix guard (reconciler/mod.rs:1820) strips any user-supplied KARS_* env. Set it inline on the keepalive's exec env instead — that's the only process that runs the worker, so a process-local env var is sufficient. After this fix: dev-agent → research mesh send now triggers an actual Hermes-generated reply via the auto-responder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * demo: bump dailyTokens cap to 2M for research + sre The 500K cap (the InferencePolicy default when dailyTokens is unset) exhausts trivially in a live demo — one 175K-context conversation through a couple of turns already crosses it, after which the inference-router throttles and the agent can't reply. The Headlamp plugin's token-budget panel renders this as '100% used', looking like a misconfiguration when it's actually intentional governance. Sets explicit 2M for research (demo scenario) and sre (Helm template default with a value-override path). Operators in production with strict cost controls can override via: --set sre.dailyTokens=N edit the research scenario yaml inline Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * plugin: workload-aware Phase column on Overview + Sandboxes pages Same false-Running problem as the SRE Cluster Health card (fixed in 5f1c2ee) affected the Overview's 'Ready' headline stat and the Sandboxes list's Phase column. Both read KarsSandbox.status.phase, which the controller sets to 'Running' the moment the Deployment spec is reconciled — independent of whether the pods inside actually pulled their image / passed readiness / etc. Two visible bugs: - Overview's 'Ready' stat counted 'phase === "Ready"' but the controller never sets that — it uses 'Running'. So 'Ready' always showed 0 even with all sandboxes healthy. - Sandboxes Phase column showed 'Running' for a sandbox whose Deployment was at 0/1 available (ImagePullBackOff, OOMKilled, etc.) — directly contradicting reality. Fixes both by pulling Deployments alongside KarsSandbox and cross-checking availableReplicas >= spec.replicas before declaring a sandbox 'Healthy'. Overview headline stats are now: Healthy — CR Running AND workload available Workload down — CR Running BUT workload unavailable CR-Degraded — CR-level Degraded=True condition Sandboxes list shows 'Workload down' (red StatusLabel) in the Phase column when the underlying Deployment can't meet its replica count. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre-action: workload-aware recovery observer (no false Recovered) The Slice 3 recovery observer declared an action 'Recovered' as soon as there were no FailedCreate / BackOff / FailedScheduling events on the target namespace in the last 30s. False positive on the canonical DeleteResourceQuota path: deleting the quota silences new FailedCreate events (no more ReplicaSet attempts), but the Deployment can still sit at 0/1 because the ReplicaSet was scaled to 0 during the failure cascade and no controller is going to scale it back up. Result before this fix: action.phase=Recovered while the workload was still down, directly contradicting what the operator sees in Headlamp's plugin (the Sandboxes / Overview / Cluster Health cards all show 'Workload down' for the same sandbox post-fix). Tightens observe_recovery to require BOTH: (1) absence of recent failure events on the target namespace (existing gate), AND (2) every Deployment in the target namespace at availableReplicas >= spec.replicas (the gate the doc comment promised for Slice 4) The Deployments gate runs first because it's the more authoritative signal — if pods aren't available, recovery hasn't happened regardless of what the event log shows. Verified live on kind: created a test KarsSREAction targeting a broken research deployment; the action stayed at phase=Applied through 3 reconcile passes (workload still down), then flipped to Recovered on the next pass after the deployment came back to 1/1. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * demo: 3 commented sandbox CRDs for the Act-I walkthrough * controller: stop spamming LimitedSupport event on every McpServer reconcile The McpServer reconciler emitted a Warning event with reason=LimitedSupport on every successful reconcile (~15s cycle), repeating the same static 'singular spec.mcp binding today, plural lands in Slice 4' text. Result: Headlamp's event view was permanently polluted with the same advisory message for every McpServer CR, drowning out actually-actionable events. The information belongs in CRD descriptions and design docs, not in the per-incident K8s Event stream. Removed the call site; kept a breadcrumb comment pointing future readers at the right places to publish the roadmap (mcpserver.spec CRD description + crd-well-oiled-machine blueprint). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: phase-changes-only watcher mode (Telegram pager, not event firehose) Adds SRE_WATCHER_MODE=phase-changes-only which alerts ONLY on KarsSandbox status.phase transitions (Running -> Failed -> Recovered) instead of the default event stream. One Telegram message per real CR state change, no pod-level event noise. Default mode in the Helm chart is now phase-changes-only because that matches what most operators actually want — a sandbox-level status pager. Uses the same sre_kube.client() httpx singleton the event-mode watcher uses (the distroless sandbox image has no kubectl). Verified live: watcher primes with the current set of KarsSandboxes and only emits on true transitions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre: overlay workload availability on synthetic phase Both the proactive phase-changes-only Telegram watcher AND the sre_diagnose chat tool only looked at KarsSandbox.status.phase, which the controller doesn't flip when downstream pods break (e.g. evicted pod can't re-admit due to a tight ResourceQuota, image-pull failure, NodeAffinity unmet). The CR stayed Running while the Deployment was 0/1, so neither the pager nor the in-chat diagnose noticed. Fix: * sre_watcher._workload_state(): for each KarsSandbox, fetch the matching Deployment in kars-<name> and synthesize WorkloadDown(a/d) when available < desired. Transitions on that overlay fire one Telegram message per real state change — still no event-firehose noise. * sre._impl_sre_diagnose: cross-checks Deployment availability for every KarsSandbox and adds WorkloadDown entries (with the affected ns + deploy name) to degraded_sandboxes. The LLM can now describe workload-level incidents accurately when the operator asks "what's wrong with my cluster?". Verified live: research deployment was 0/1 (quota-violation, Act II break.sh scenario). After healing the quota, the watcher fired one Telegram alert: research: WorkloadDown(0/1) -> Running. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * sre-action: bump recovery window 5m→10m + late-recovery healer Demo on 2026-06-11 hit a real-world false negative: SRE applied the DeleteResourceQuota patch, observed for 5 min, marked the KarsSREAction Failed — but research actually recovered ~1 min later. The terminal Failed state then stuck even though the cluster was fine, leaving the operator with a misleading state. Two fixes: 1. RECOVERY_WINDOW_SECONDS 300 → 600. Real K8s recovery routinely exceeds 5 min on cold caches, RS back-offs, congested nodes. 2. Late-recovery healer (Failed → Recovered edge). For Failed CRs that DID reach Apply (i.e. have appliedAt set — pre-apply validation failures don't qualify), the terminal handler keeps running observe_recovery for LATE_RECOVERY_WINDOW_SECONDS = 30 min since appliedAt. If recovery is observed, flip phase back to Recovered with reason=LateRecovery. Polling cadence during this window is 60s (vs the standard 300s terminal requeue) so latency is bounded. State-machine docs at the top of the file updated to reflect the new Failed → Recovered edge. Existing tests (6) still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(security): audit for kars-sre demo-and-agent slice (Slices 0-4 + recovery healer) Covers all 46 commits on this branch since main. Documents: - T1: SRE writer SA escalation surface (mitigated: 7-layer a…
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
Jun 26, 2026
Address the remaining war-room recommendations: - Above-the-fold "wow path" (#1): the hero code block now shows the full path to a chat (install -> `kars dev --release --target local-k8s` -> `kars connect dev-agent`), the first-run GIF moves directly under it (was buried in "Try it in five minutes"), and a "Full quickstart ->" link points to docs/quickstart.md. A visitor sees what it is, the three commands, and a GIF of it working without scrolling. - Trim internal-detail overload (#4): compress the heaviest README paragraph (mesh crate pins / vendored-tgz mechanics) to the essential, verifiable security claim plus a link to architecture.md#the-mesh, which carries the full provenance. The CRD/runtime tables, inference-router spotlight, and honest "Known limitations" list are kept intact. - Nit: drop two CONTRIBUTING references to the internal `plan.md` (public contributors can't see it) in favour of "open a follow-up tracking issue". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
Jun 26, 2026
…s/internal (#468) * docs: stop tracking docs/internal/ (was publicly exposed despite .gitignore) The docs/internal/ tree (strategic plans, competitive analyses, security-audit logs, announcement/blog drafts, POC lab-notes) was force-added to git in an earlier change and would have shipped publicly at launch despite being listed in .gitignore. Untrack the entire tree with `git rm -r --cached` — the files remain on disk for maintainers but leave the public repository. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: rebuild information architecture (retire/relocate internal-bound pages) Move maintainer-only and historical pages out of the public surface and rebuild the navigation so every published page has exactly one home: - Relocate to docs/internal/ (now untracked): PUBLISHING.md, the entra-agent-id POC lab-notes (00-poc-archive, 02-aci-token-flow, 03-original-findings, 04-migration-guide), blueprint 07 (kars-sre proposal), and showcase/outline. - Promote docs/sre.md to docs/runbooks/sre.md (now a shipped runbook, not a proposal) and add it to the nav. - Rebuild docs/SUMMARY.md and docs/README.md: drop retired/moved pages, add the supply-chain posture page and the BYO runtime contract, nest the entra deep dive under Architecture so it renders. - Trim the entra-agent-id index to the pages that remain public. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: per-page accuracy and consistency pass for launch Ground every public claim in the code and remove overlap/contradictions: - Counts: twelve CRDs (ten workload incl. KarsSREAction + two infrastructure) and eight runtimes, propagated everywhere; add a full KarsSREAction section to the CRD reference and lifecycle. - Versions: v0.1.18 across README/installer/examples; package is @kars-runtime/cli (the only published npm package); org links microsoft/kars -> Azure/kars. - Mesh provenance: correct the AGT SDK provenance and agentmesh 3.1.0 -> 4.0.0; describe the OpenClaw (TS SDK) vs Hermes/Python (Python mesh client) split. - Status honesty: add build-from-source/roadmap banners to mesh-plugin, a2a-gateway (verifier-is-library), runtimes/CONTRACT (draft surface), and notation-ratify (no kars up --sign-images flag exists). - getting-started: collapse to a single linear path; document the provider picker exactly once; make kind/local-k8s a first-class step; source build is secondary. README "Try it" reframed kind-first. - Ecosystem: frame agentgateway and kubernetes-sigs/agent-sandbox as an aim to align (no discussions/integrations claimed). - Fix broken relative links, de-orphan the supply-chain posture page, add a pricing disclaimer to the cost dashboard, and grammar ("an kars" -> "a kars"). - Update the sre.ts design-comment path to docs/runbooks/sre.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(site): commit custom mdbook theme and fix Mermaid rendering The polished site styling lived in an untracked docs/site/theme/ directory, so a clean CI/Pages build would fall back to stock mdbook and lose all of it. Commit the theme and fix three rendering defects found in visual review: - Track theme/index.hbs (kars top nav + logo), theme/css/custom.css (design tokens, CTA buttons, card tables/admonitions, framed code, diagram cards), and the brand favicons; wire them via book.toml (theme + additional-css). - Mermaid contrast: subgraph titles rendered white-on-light on the dark site — add an explicit dark color to every light-fill classDef across the diagrams. - Mermaid clipping: long cluster titles wrapped to a hidden second line — raise wrappingWidth and let cluster labels render at natural width via CSS. - Mermaid init: guard getElementById theme-toggle hooks that were absent in the custom theme and threw a console error on every page. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(cli): don't require Rust toolchain for `kars dev --release` `kars dev --release` pulls pre-built, signed images from ghcr.io/azure and compiles nothing, but the preflight unconditionally required `cargo` on PATH — breaking the documented "no Rust" quickstart promise. Require `cargo` only when building locally from source (no `--release`). Extract the tool list into a pure, exported `requiredToolsFor()` and add regression tests asserting `cargo` is absent in `--release` mode and present otherwise. Also remove the dead `--yes` flag from `kars upgrade` (no code read it; the command is already non-interactive) so docs and code agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: add quickstart, kars upgrade runbook, llms.txt, and top-OSS landing polish Apply the launch-readiness scorecard derived from agentgateway / Cilium / Dapr / cert-manager docs: - Add docs/quickstart.md — a 3-command, ≤5-minute path to a running agent. - Document `kars upgrade` (previously undocumented): a full CLI reference section plus docs/operations/upgrades.md runbook (verified against upgrade.ts), wired into SUMMARY + the operations index + getting-started. - Add docs/llms.txt (machine-readable docs index for AI tooling) and its generator docs/site/gen-llms-txt.py, regenerated from SUMMARY.md. - Landing page: architecture diagram above the fold, a "Why kars" comparison table, a Quickstart CTA, and a Feature-status link. - Add "Last tested with kars v0.1.18" footers to the six blueprints and a "which loop do I want?" decision table to blueprint 01. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: accuracy pass from architect + doc-expert re-audit (zero vaporware) Correct every claim flagged as an overclaim or stale against the code: - Audit: "Merkle-chained" -> "SHA-256 hash-chained (AGT's AuditLogger)"; the Merkle module is library-only and explicitly not wired (README, SECURITY, agt-boundary already accurate). - Cross-runtime mesh: scope the shipped/proven claim to OpenClaw + Hermes (tests/e2e/interop/hermes_openclaw_bidi.sh). The other Python adapters bundle the client but don't yet expose the mesh tools; LangGraph ships an A2A/HTTP helper, not a Signal session. Fixed in README + architecture. - Hermes is at parity (Act 2 mesh is live): rewrite the stale "stub / NOT IMPLEMENTED" sections in hermes-plugin, CONTRACT, and runtimes/hermes/README. - CRD counts: "eight peer workload CRDs" -> nine; lifecycle diagram 9 -> 12. - AGT boundary: MeshProvider is agent-side only — the router has no Rust mesh impl (providers/mesh.rs); correct the provider table. - Egress: the signed-OCI allowlist IS the enforced L7 source of truth (egress_allowlist_loader + forward_proxy), not advisory/roadmap. - Token budgets: daily + monthly windows are enforced at the router, not reconciler-only. - `kars upgrade --rollback` reverts the Helm revision, not image bits (workloads pin :latest) — document the accurate rollback semantics. - Confidential (Kata + SEV-SNP) isolation is opt-in, not active by default. - Remove nonexistent commands from docs (`claw attest` -> `kars attest`, `kars offload`); add missing `kars upgrade` / `sre` / `headlamp` to the CLI reference; OpenClaw language is TypeScript/Node, not Python; "separate router pod" -> sidecar container; runtime examples 8 -> 10; multi-runtime images are published (best-effort import), not "not yet published". - Make the dev flow kind-first in getting-started + CONTRIBUTING; drop public references to gitignored docs/internal; refresh stale CLI command count. - Fix broken links/anchors surfaced by the sweep (crd-reference deep links, roadmap anchors, demo-script/byo-contract/blocklists paths). Verified: mdBook builds clean, 0 broken links/anchors, 0 orphans; 842 CLI tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(readme): make the README kind-first, not a Docker ad The "Try it in five minutes" hero and the modes table led with Docker (`kars dev --release`) and treated the kind dev loop as a footnote, contradicting the kind-first framing used everywhere else. - "Try it" now leads with `kars dev --release --target local-k8s` (kind, the real production pod shape) as the recommended path; the single-container Docker target moves into a collapsible "just want the fastest smoke test?" aside. - Rename "Two modes" → "Three ways to run it" with a three-column table (Local kind / Local Docker / Prod AKS), kind marked recommended, Docker framed as the fast inner loop rather than the default. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(cli): kind dev loop accepts any container runtime (Docker/Podman/nerdctl) The local-k8s (kind) runner already drives docker, podman, or nerdctl via `KIND_EXPERIMENTAL_PROVIDER` (`dev/local-k8s.ts` RUNTIME_PRIORITY), but the preflight hard-required the `docker` binary for the local-k8s target — so a Podman- or nerdctl-only user was blocked at preflight despite the runner supporting them. - Drop `docker` from the local-k8s must-each-be-present list; add an anyOf probe (`CONTAINER_RUNTIMES = docker | podman | nerdctl`) that passes if at least one is on PATH. The single-container `docker` target still requires the `docker` CLI (it shells out to it directly). - Update the preflight tip and add regression tests (kind target must not pin docker; docker target must). - Docs: make the runtime story accurate — the recommended kind loop accepts Docker/Podman/nerdctl; the fast single-container path uses the `docker` CLI. Reorder the getting-started prerequisites table kind-first. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: don't imply Docker-only where the kind path is runtime-agnostic Sweep follow-up to the Podman/nerdctl fix. The single-container `docker` target genuinely uses the `docker` CLI, so those references stay — but two spots tied to the kind path / general first-run wrongly implied Docker specifically: - blueprints/02 topology mermaid: "docker build && load" -> "build/pull + load images via the detected runtime" (kind drives docker/podman/nerdctl). - getting-started troubleshooting: "Docker Desktop is not running / Start Docker" -> "the container runtime isn't running / start Docker Desktop, podman machine, or colima". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: add first-run demo (asciinema) to the README + quickstart hero A real, recorded `kars dev --release --target local-k8s` first run on a local kind cluster — provider picker (GitHub Copilot), bring-up of the controller + encrypted mesh + sandbox, ending at a governed agent. - docs/assets/kars-dev-firstrun.gif (1.9 MB, optimized) embedded in the README "Try it in five minutes" hero and docs/quickstart.md. - docs/assets/kars-dev-firstrun.cast — the replayable asciinema cast, linked from the quickstart caption. Recording hygiene: the cast is scrubbed of the recorder's internal Foundry endpoint (-> my-foundry/my-project placeholder) and of all dead local tokens (Headlamp SA JWT, OpenClaw gateway/WebUI tokens, did:mesh id, image digests). Verified: no secrets remain; mdBook builds and the asset renders. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: CMO final-review fixes (consistency, conversion, credibility) From the launch war-room review: - Resolve a credibility contradiction: roadmap.md said aggregate token budgets are "not yet metered" while maturity.md (correctly) says daily+monthly are enforced. Reword the roadmap item to claim only what remains (per-hour windows + a configurable rejectOnExceed knob), matching budget.rs. - Make docs/quickstart.md kind-first so the page matches its own first-run GIF and the README hero: step 2 is `kars dev --release --target local-k8s`, prerequisites are kind + kubectl + any container runtime, and the single container `docker` path is the explicit "even faster, less faithful" alt. - README: surface the community on-ramp (good first issue / help wanted / Discussions) in the Contributing section instead of only doc links. - README: replace the "print the current Azure subscription" sample prompt (odd in a no-Azure quickstart) with a neutral one; soften the slightly promissory "fit seamlessly" ecosystem wording to "can grow toward". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: CMO final-review — above-the-fold wow path + trim README detail Address the remaining war-room recommendations: - Above-the-fold "wow path" (#1): the hero code block now shows the full path to a chat (install -> `kars dev --release --target local-k8s` -> `kars connect dev-agent`), the first-run GIF moves directly under it (was buried in "Try it in five minutes"), and a "Full quickstart ->" link points to docs/quickstart.md. A visitor sees what it is, the three commands, and a GIF of it working without scrolling. - Trim internal-detail overload (#4): compress the heaviest README paragraph (mesh crate pins / vendored-tgz mechanics) to the essential, verifiable security claim plus a link to architecture.md#the-mesh, which carries the full provenance. The CRD/runtime tables, inference-router spotlight, and honest "Known limitations" list are kept intact. - Nit: drop two CONTRIBUTING references to the internal `plan.md` (public contributors can't see it) in favour of "open a follow-up tracking issue". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: complete docs/internal untracking (drop v0.1.19 audit docs too) The launch branch stops tracking docs/internal/ (internal planning + security audit content not meant for the public repo). Three audit docs were added on main after this branch was cut (the v0.1.19 memory + dev fixes); remove them too so docs/internal is fully untracked, consistent with 631b09e. The audit content is preserved in git history. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps jsonwebtoken from 9.3.1 to 10.3.0.
Changelog
Sourced from jsonwebtoken's changelog.
Commits
abbc307Fix type confusione99740dfix: bump minimal version requirements (#481)50d15e0Use try_sign to avoid panics (#479)245858fBump some dep122c2edBump action number in CI72e0c7fExpose cryptography backends via CryptoProvider (#452)53a3fc2Do not fail for clippy3226cfcPrepare for releasedfe58f9Remove unnecessary Clone bounds from decode functions (#458)9b3e19cFix function names in README (#457)Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting
@dependabot rebase.Dependabot commands and options
You can trigger Dependabot actions by commenting on this PR:
@dependabot rebasewill rebase this PR@dependabot recreatewill recreate this PR, overwriting any edits that have been made to it@dependabot show <dependency name> ignore conditionswill show all of the ignore conditions of the specified dependency@dependabot ignore this major versionwill close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this minor versionwill close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this dependencywill close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)You can disable automated security fix PRs for this repo from the Security Alerts page.