Repository navigation
Security hardening phase 1 + Foundry integration improvements - #15
Merged
Merged
Conversation
Phase 1 — Secrets: - Remove hardcoded PostgreSQL password from agentmesh.yaml, use K8s Secret ref - Read API key from file at runtime instead of hardcoding in env export - Add chmod 0400 on /tmp/azure-openai-key fallback - Remove /tmp/azure-openai-key fallback from router auth.rs Phase 2 — Authentication: - Auto-generate ADMIN_TOKEN from /dev/urandom when not configured - Add admin token auth to sidecar GET endpoints (/trust, /audit, /audit/verify) - Validate x-azureclaw-sandbox header against K8s name regex Phase 3 — Fail-closed: - Router: fail-closed after 3 consecutive sidecar failures (grace window for cold start) - Plugin: same fail-closed pattern in evaluateAGTPolicy() - Spawn endpoint: fail-closed after grace window Phase 4 — Command injection: - Replace execSync(curl) with Node http.request() for http_fetch - Gate fallback exec_command through AGT policy before execution - Validate sandbox name format at controller reconcile entry point Phase 5 — Cryptographic: - Replace Math.random() with crypto.randomUUID() in 3 locations - Fix timing attack in bytesEqual() — use XOR accumulation Phase 6 — Policy/Trust: - Allowlist registry proxy paths (prevent path traversal) - Add threading.Lock around trust score get+update (race condition fix) All 321 tests pass (47 router + 80 controller + 148 TS + 46 Python). Clippy clean with -D warnings. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase 5 — Crypto: - Add AES-GCM nonce tracking (collision detection via Set) in encrypted audit - X25519 key length validation (32-byte check) + all-zero DH output rejection - X3DH signature failure logging for audit trail Phase 6 — Policy: - Strengthen recon-deny regex patterns (context-aware boundaries vs \b) Phase 7 — Network: - Truncate domains >100 chars in log output (prevent log flooding) - Guard IMDS fallback: require AZURE_TENANT_ID before attempting IMDS token Phase 8 — Remaining: - Log sanitization: strip ANSI escapes + collapse newlines in user input - SANDBOX_NAME format validation (K8s name regex) - Rate limit sidecar GET endpoints (/status, /trust, /audit) - Helm values: document image tag pinning for production All 321 tests pass. Clippy clean with -D warnings. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…via env The AGT sidecar was starting before the admin token file was written, and the token file (chmod 400, owned sandbox:sandbox) was unreadable by the sidecar process (UID 1001/router). This left ADMIN_TOKEN empty, causing _check_admin_token() to skip auth on /trust and /audit endpoints. Fixes: - Generate ROUTER_ADMIN_TOKEN before starting the sidecar - Pass ADMIN_TOKEN env var directly to the sidecar process - server.py now prefers env var over file paths (avoids permission issues) - Token file chmod 440 with router added to sandbox group (belt-and-suspenders) - Log warning when no admin token is configured Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Root cause: spawn polling called GET /sandbox/{name} but the route is
GET /sandbox/{name}/status — every poll returned 405, the catch block
swallowed it, and the loop ran all 24×5s=120s iterations.
Fixes:
- Correct polling path to /sandbox/{name}/status
- Reduce poll interval: 5s→1s, cap at 45s (was 120s)
- Replace blind 3s sleep with registry discovery poll (1s×10 max)
- Log every 5th poll instead of every poll (reduce noise)
- Add 'registry/' prefix to registry path allowlist (fixes
azureclaw_discover 400 'Invalid registry path' error)
- Replace entrypoint sleeps with health-gate loops (saves ~4s)
- Router: sleep 1 → poll /healthz every 200ms
- Gateway: sleep 1 → sleep 0.5
- Auto-approve: sleep 2 → poll gateway healthz every 500ms
Expected spawn timeline: ~10-15s (container boot + poll detect + registry)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
After security hardening added admin token auth on sidecar GET /trust,
/audit, /audit/verify, the router's SidecarProxy was forwarding
requests without the token. This caused:
1. Operator view: 'no audit entries' — GET /agt/audit → sidecar
returned 403 because router didn't include admin token
2. Mesh trust lookup: incoming mesh messages had trust score 0
because the plugin's direct sidecar call got 403
Proper fix (not a hack):
- SidecarProxy reads ADMIN_TOKEN env var at init, includes it as
Authorization header on all forwarded requests (same shared secret)
- Plugin mesh trust lookup routes through /agt/trust/{id} on the
router (port 8443) instead of directly hitting sidecar (port 8081)
- Works in both dev (same container) and AKS (same pod, shared secret
via K8s Secret volume mount)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Switch foundry_image_generation tool from the Responses API (which
required an agent_reference + tool orchestration) to the standard
Azure OpenAI Images API:
POST /openai/deployments/{model}/images/generations
This matches the working Python SDK call:
client.images.generate(model='gpt-image-1', prompt='...', n=1, size='1024x1024')
Changes:
- Router: add /openai/deployments/{deployment}/images/generations route
with AGT policy gate and proxy to Azure OpenAI endpoint
- Plugin: call standard images API instead of Responses API with tools
- Remove unnecessary orchestrator model parameter
- Verified: direct API call returns 200 with b64_json image data
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Switch proxy to unified /openai/v1/ URL format (model in body) for both dev (API key) and AKS (Entra) modes — all models use same endpoint - Auth always uses Authorization: Bearer (works with both API keys and Entra tokens on the unified endpoint) - Add /v1/responses route for Responses-only models (e.g. gpt-5.4-pro) with AGT policy gate and budget tracking - Auto-fallback in chat_completions: if Azure returns 400 'unsupported', transparently retry via Responses API with format translation (messages→input, max_completion_tokens→max_output_tokens) - Convert Responses API output back to chat/completions format so OpenClaw works transparently without changes - Fix image generation path (was double-deploying deployment in URL) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agents spawned by the same parent can now communicate via mesh without KNOCK rejection. Trust is established securely: - Parent's router passes AGT_TRUSTED_PEERS env var at spawn time containing parent-verified AMIDs (name:AMID pairs) - Sub-agent plugin reads env var at startup, pre-seeds trust store via admin-token-protected /agt/trust endpoint (local sidecar) - KNOCK handler gives +500 affinity bonus to parent-verified AMIDs (stored in parentTrustedAmids set, separate from general lookups) - Parent's spawn tool builds peer list from its own AMID + all existing siblings' AMIDs (from amidToName map) Security properties: - Trust source is the parent's router (sets env var), not self-reported - Env vars are set at container creation time, not modifiable by sandbox - parentTrustedAmids is a separate set from amidToName to prevent trust escalation via arbitrary registry lookups - Admin token required for trust mutations (unchanged) - Works identically on Docker (dev) and AKS (production) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…attempts Models like gpt-5.4-pro don't support chat/completions — every request was doing a wasted round-trip (400 in ~0.8s) before falling back to Responses API. Now the first fallback caches the model name in an in-memory HashSet, and all subsequent requests go directly to /openai/v1/responses. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Handle tool_calls → function_call items in input conversion - Handle tool role → function_call_output items - Handle string content → typed content blocks (input_text/output_text) - Convert system role → developer for Responses API - Add type:'message' wrapper for regular messages - Handle function_call → tool_calls in response conversion - Set finish_reason to 'tool_calls' when function calls present - Add debug logging for upstream requests (URL, body_len, resp_len) - Add stream status logging - Add 4 unit tests for conversion functions Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The spawn tool was unconditionally force-removing existing containers before creating new ones. When the LLM calls spawn twice for the same sub-agent (retry or re-invocation), the running container gets killed mid-operation, causing: - 'Connection reset without closing handshake' in relay logs - New AMID on reconnect (new identity = broken sessions) - 'Failed to send message: closed connection' for stale AMIDs Now checks if container is already running and reuses it instead. Only removes stopped/dead containers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
On AKS, if the LLM calls spawn twice for the same sub-agent, the K8s API returns 409 Conflict. Previously this was a hard error — now it returns success with 'already running', consistent with the Docker path fix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Three optimizations that reduce initial mesh communication lag: 1. AMID cache-hit in mesh_send: skip registry search if AMID was pre-cached during spawn. Falls back to re-discovery if cached AMID is stale (agent not found / connection closed). Saves: 2-12s on first message to known agents. 2. Parallel spawn readiness: poll container status AND registry search simultaneously instead of sequentially. Saves: 5-10s during spawn (registry search overlaps boot). 3. SDK parallel registry calls: fetch prekeys and agent info via Promise.all() instead of sequential awaits. Saves: 1-2s per session establishment. Total: ~8-24s faster initial mesh communication. All E2E encryption unchanged — Signal Protocol handshake intact. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Azure Responses API includes 'error': null in every successful response.
The error check used resp.get('error').is_some() which returned true for
null values, causing the function to return the raw Responses API format
instead of converting to chat/completions format. The OpenAI SDK then
failed to parse it: 'Expected property name at position 1'.
Fixed by checking for non-null error values only.
Added test covering the null error case.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reasoning models (gpt-5.4-pro) take 30-50s via Responses API. The
previous buffered approach held the HTTP connection with no data,
causing the OpenClaw client to timeout ('surface_error reason=timeout').
Now sends SSE comment keepalives (': keepalive') every 5 seconds while
the Responses API call is in progress. SSE comments are ignored by
parsers but keep the HTTP connection alive. The actual data is sent
as a single SSE frame when the response arrives.
Uses tokio::select! with a spawned task and mpsc channel to stream
the keepalive comments and final data to the client.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…+ image_gen The native delegateToNativeAgent path failed because: 1. OpenClaw gateway requires device pairing (blocks headless exec) 2. OpenClaw exec allowlist double-gates commands already checked by AGT sidecar Changes: - Use processTaskWithTools directly for AGT mesh task_request messages - Add mesh_send tool to sub-agent toolset (agent-to-agent E2E encrypted comms) - Add foundry_image_generation tool to sub-agent toolset - Fix system prompt to match actual available tools (remove phantom tools) - Save generated images to temp file instead of discarding base64 data Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The sed -i on .bashrc failed with exit code 2 when the file didn't exist at runtime (emptyDir mount or permission issues). With set -e in the script, this killed the entrypoint and caused CrashLoopBackOff for all sandbox pods. Fix: touch + || true guard before sed operation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The plugin was in plugins.allow but not in plugins.entries, so OpenClaw
permitted it but never activated it. In dev mode the gateway auto-discovers
extensions from the directory, but AKS requires explicit entries with
{"enabled": true}. Initialize PLUGINS_ENTRIES with the azureclaw entry.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Two bugs in up.ts caused the Foundry project endpoint (services.ai.azure.com) to be lost when running from cached context: 1. Cache prefill (line 255-258) set options.foundryEndpoint to the OpenAI endpoint (from cachedCtx.foundryEndpoint), overwriting the project URL 2. saveContext stored options.foundryEndpoint (already overwritten) as foundryProjectEndpoint instead of the original foundryEndpoint variable Fix: cache prefill now restores OpenAI endpoint to options.openaiEndpoint and project endpoint to options.foundryEndpoint (separate variables). saveContext uses the local foundryEndpoint variable which holds the original project URL. Also fixed cached context file with correct project endpoint URL. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add ui.assistant.name/avatar in openclaw.json config template (proper config-driven approach, no upstream file patching) - Add Identity section to MEMORY.md with strong AzureClaw persona - Update sub-agent system prompt with AzureClaw branding - Remove phantom tools from MEMORY.md (foundry_conversations, foundry_evaluations, foundry_deployments, foundry_agents) - Add azureclaw_discover to tools list - Derive clean SANDBOX_NAME from pod hostname (strip hash suffixes) - AZURECLAW_DISPLAY_NAME env var for custom branding (default: AzureClaw) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add ~/.azureclaw/secrets.json (mode 600) to store all secrets: - Azure OpenAI API key (migrated from legacy credentials file) - Channel tokens (Telegram, Slack, Discord) - Plugin API keys (Brave, Tavily, Exa, Firecrawl, Perplexity, OpenAI) New CLI commands: - azureclaw credentials set <key> <value> — save a secret - azureclaw credentials list — show stored secrets (masked) - azureclaw credentials remove <key> — remove a secret Resolution priority: CLI flag > secrets.json > host env var dev.ts now auto-loads channel tokens from secrets.json, so 'azureclaw dev' works without --telegram-token flags once secrets are configured via 'azureclaw credentials set'. 10 new tests for secrets CRUD, migration, and resolution. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Channel tokens and plugin API keys for 'azureclaw add' now resolve via: CLI flag > secrets.json > host env var. This means 'azureclaw add my-agent --channels telegram' works without --telegram-token if the token is already stored via 'azureclaw credentials set telegram-token <token>'. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw credentials (no subcommand) is now a guided interactive wizard with masked input: - Menu: Azure OpenAI, Telegram, Slack, Discord, Search APIs, Other - Each prompts with masked password input (•••• on screen) - Channel tokens support dot-suffixed variants for multiple bots (e.g. telegram-token.cloud, telegram-token.dev) - Shows existing stored secrets at the top - Looping menu — configure multiple categories in one session azureclaw credentials set <key> [value]: - Value argument is now optional - If omitted, prompts interactively with masked input - Validates dot-suffixed keys against base key Operator TUI spawn dialog: - Pre-fills channel token from secrets.json on startup - Shows token field for ALL channels (not just Telegram) - ←→ cycles through stored token variants when multiple exist - Auto-fills default token when switching channels - Passes correct --<channel>-token flag for any channel type listSecretVariants() finds base + dot-suffixed keys for picker UIs. 2 new tests for variant listing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agent tool gaps fixed:
- discover: query AGT registry for other agents (names, trust, status)
- mesh_inbox: check for incoming messages from sibling agents
Sub-agents now have 10 tools: exec_command, http_fetch,
foundry_{web_search,code_execute,file_search,memory,image_generation},
mesh_send, mesh_inbox, discover
Credentials wizard (azureclaw credentials):
- Interactive guided flow with masked password input
- Category menu: Azure OpenAI, Telegram, Slack, Discord, Search, Other
- Dot-suffixed multi-token support (telegram-token.cloud, .dev)
- Shows existing secrets, loops for multi-category setup
Operator TUI:
- Pre-fills channel tokens from secrets.json
- Token field for all channels (not just Telegram)
- ←→ cycles stored token variants
- Auto-fills default token on channel switch
Config:
- listSecretVariants() for multi-token picker UIs
- 2 new tests for variant listing
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reverts the processTaskWithTools-only approach from e157a0f. Sub-agents now use delegateToNativeAgent first (full OpenClaw toolset: exec, file editing, git, browser, foundry_*, azureclaw_*, etc.) with processTaskWithTools as fallback if native agent fails. The original switch was made because openclaw agent required device pairing — but the gateway already has dangerouslyDisableDeviceAuth:true, and the actual failure was a missing --session-id flag. The native agent now works correctly on both Docker and AKS. processTaskWithTools still serves as a reliable fallback with 10 tools (exec_command, http_fetch, foundry_*, mesh_send/inbox, discover). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- task:execute AGT policy check gates entire task before delegateToNativeAgent - Container sandbox (seccomp, netpol, cgroups) is enforcement boundary - No mixing with OpenClaw tools.deny — AGT is the single policy source - Plugin tools inside native agent still check AGT per-call - Falls back to processTaskWithTools if native agent fails Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
max_tokens is rejected by newer models (GPT-5.x). max_completion_tokens works universally across all models (GPT-4.1+, GPT-5.x). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Fix TypeScript type for credential prompt map (TS7053) - Define cmd variable from tool args in exec_command handler (TS2304) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When no base key exists (e.g. telegram-token) but variants do (telegram-token.dev, telegram-token.cloud), pick the first variant. Fixes credential resolution for 'azureclaw dev --channels telegram'. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Supports dot-suffix on channel names in --channels flag: azureclaw dev --channels telegram.cloud azureclaw dev --channels telegram.dev,slack Resolves telegram-token.cloud from secrets.json when variant is specified, falls back to base key otherwise. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- _routerCall now accepts configurable timeout (default 15s) - Image generation uses 90s timeout (was 15s, causing timeouts) - Upgrade Go builder to 1.25 (patches 38 Go stdlib CVEs) - Run npm audit fix after OpenClaw install (patches tar, minimatch, glob) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Set tools.exec.security=full in openclaw.json — AGT is the sole policy authority. Removes flaky post-startup 'openclaw approvals set' hack that could hang or fail silently. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Uses OpenClaw's registerImageGenerationProvider SDK to register 'azure-foundry' provider. Routes through inference router (managed identity on AKS, API key on dev). OpenClaw's native image_generate tool renders images inline in the chat UI automatically. Config: imageGenerationModel = azure-foundry/gpt-image-1 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add /v1/images/generations route to inference router (OpenAI-compat) - Route extracts model from body, delegates to deployment-based handler - Configure OpenAI provider in openclaw.json pointing to inference router - Set imageGenerationModel = openai/gpt-image-1 in agent defaults - OpenClaw's built-in image_generate now renders images inline in chat - Works on both AKS (managed identity) and dev (API key) - Remove custom registerImageGenerationProvider (not needed) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
enabled auto-merge (squash)
April 1, 2026 12:51
…port - Operator spawn form shows variant labels (cloud/dev) instead of masked values - Added Allow From field with ←→ cycling and auto-correlation to token variant - Added --telegram-allow-from flag to 'add' command (parity with 'dev') - credentials list shows parent label for dot-suffixed variant keys - telegram-allow-from now supports per-variant suffixes (.cloud, .dev) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The sandbox Docker build (36 stages) exhausts ubuntu-latest (2 vCPU, 7GB). Switch to ubuntu-latest-xl (4 vCPU, 16GB) to prevent runner OOM/disconnect. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
pushed a commit
that referenced
this pull request
May 10, 2026
…ocally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 11, 2026
…#245) * feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap Wire azureclaw runtime through createMeshTransport() factory so we can flip between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without code changes. Surface additions to IMeshTransport (both adapters now expose): - lookup(amid) — registry RPC for reputation/display name - submitReputation(...) — registry RPC for peer feedback - enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on) - onError(kind, from, detail) — diagnostic hook for decrypt + ws errors - onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal - onDisconnect(reason, code) — ws close / error fan-out mesh-plugin (vendored A adapter): - connection.ts delegates to the underlying SDK; lazy bind for hooks registered before connect() - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the contract so neither adapter can drop a method without CI failing mesh-plugin (AGT B adapter): - agt-transport.ts implements lookup/submitReputation as REST calls to the registry (AGT MeshClient is pure transport — registry RPCs intentionally not added to AGT upstream; they belong on a separate RegistryClient) - enableKnockEnforcement is a no-op (AGT MeshClient always enforces) - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods (added on local AGT branch azureclaw-meshclient-event-hooks, NOT pushed — AGT team owns the upstream PR) runtime (runtimes/openclaw): - Adds @azureclaw/mesh as a file: dependency - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)' when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value - Identity is generated once via vendored SDK regardless of provider, then raw Ed25519 keys are extracted via toData() and shared across both — same AMID either way - Banner now reports active provider (vendored vs agt) Docs: - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity, policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer + the wiring + migration path - Documents the 3 hooks added to local AGT branch and the 3 governance methods kept adapter-side Tests: - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat) - runtimes/openclaw: 118/118 pass - AGT (local branch): 387/387 pass with 8 new event-hook tests Open work for cleanup phase: once AGT publishes the version with our event hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var toggle. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): pre-build mesh-plugin for runtime CI + format reconciler PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): quote workflow step name containing colon YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): clippy fixes for Rust 1.95.0 CI runs Rust 1.95.0 which added new clippy lints: - doc_lazy_continuation: indent doc list items that span multiple lines. Added two-space indent to the trailing 'All three are populated...' paragraph so it is treated as a continuation of the preceding list item rather than its own malformed list item. - obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as if .. { a } else { b } per the lint suggestion. These were pre-existing on dev (CI only started failing once the runner picked up Rust 1.95.0); fixing here so PR #245 can land green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12 Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5 The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): mark gaps G1 and G2 fixed on local AGT branch Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass. * dev: add --mesh-provider <vendored|agt> selection with first-run prompt Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker mode before AKS rollout. Same flag, three integration points: CLI (cli/src/commands/dev.ts): - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball - First-run interactive prompt offers AGT only if the toolkit checkout is actually present locally — silently defaults to vendored otherwise (no pestering for users without AGT cloned). - --build branch: builds the right relay/registry images vendored → vendor/agentmesh-relay + agentmesh-registry (Rust) agt → agent-governance-python/agent-mesh/docker/Dockerfile with COMPONENT=relay / registry build-args - Sandbox image build: stages locally-packed AGT SDK tarball into .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL build-arg (auto-discovers if --agt-sdk-tarball not given). - Runtime branch: skips Postgres for AGT (in-memory registry), uses correct ports (AGT: 8083 relay, 8082 registry; vendored: 8765/8080) and health path (AGT: /healthz; vendored: /v1/health). - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the runtime transport-factory honors the user's choice. Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile): - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the sandbox npm-installs the local tarball instead of fetching the published @microsoft/agent-governance-sdk from npm. Lets us smoke-test the locally-patched AGT branch (G3/G4 fixes) end to end without round-tripping through npm publish. - .agt-sdk/ staging dir always exists (with .keep) so the COPY never fails when the user didn't stage a tarball. Defaults preserved: --mesh-provider=vendored, existing behavior is byte-identical for users who don't opt in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so tsc failed with TS2307 in the AGT build path. Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the build context, and strip its 'prepare' script (which would invoke tsc, not present in this stage; the pre-built dist/ is sufficient). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks). - connect() now passes autoRegister: true so the SDK uploads identity and prekeys instead of the adapter re-implementing that path with raw HTTP. - discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT. - lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}). - submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score clamped to [0,1]; the vendored /registry/feedback endpoint does not exist in AGT. - Replaced mapAgent with pickDisplayName helper: AGT puts display name in metadata.display_name (set by registerSelf), with the first capability as the fallback. Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and the bespoke fetchWithRetry helper — both are upstream concerns now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes the vendored registry wire shape. The vendored impl talks to /registry/* (the existing agentmesh-registry); the AGT impl talks to /v1/discover and /v1/agents/{did} on the upstream AGT registry. Both expose a single normalized RegistryEntry envelope, so callsites stay readable. getMeshRegistry(routerUrl) is the entry point. Provider selection follows AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base). Migrated all raw-HTTP registry callsites: - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName, resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid. - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS spawn discovery, AKS→local discovery. - core/agt-task-loop.ts (1 site): registry_capability_search tool. - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe, azureclaw_discover (mesh_discover) tool. - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe, heartbeat keepalive (no-op under AGT — relay does liveness via WS). The discover-on-router-unreachable test now asserts the new contract: empty list + count:0 instead of a 'Discovery failed' string. Registry hiccups must not break tool calls. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): wire AGT provider end-to-end (6 stackup bugs) End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced six bugs blocking the upstream AGT MeshClient swap. All fixed: 1. Final sandbox Docker stage didn't COPY mesh-plugin, so the file:../../mesh-plugin symlink dangled in node_modules. Plugin swap silently fell back to vendored with 'Cannot find package @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist} into /mesh-plugin/ in the final stage of the Dockerfile. 2. entrypoint.sh used cp -r when copying node_modules into the plugin extension dir, preserving the (now-broken-at-runtime-path) symlink. Switched to cp -rL so symlinks dereference into real files in the target tree. 3. mesh-plugin/src/index.ts imported createMeshTransport from ./transport-factory.js but never re-exported it. Runtime swap path couldn't find the factory. Added the missing re-export. 4. inference-router agt_registry_proxy unconditionally prepended '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents' became '/v1/v1/agents' at the upstream. Now: forward verbatim when path starts with 'v1/' or equals 'health', else prepend. Preserves vendored SDK behavior ('registry/register' → /v1/registry/register). 5. /agt/relay route only matched the bare path, but AGT MeshClient appends '/ws' to relayUrl. Added /agt/relay/ws route and made the upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt. 6. agt_registry_proxy route was declared get(...).post(...) only. AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey upload and DELETE for deregister — both 405'd at the router. Added .put() and .delete() to the route declaration. Bug #6 was invisible to vendored because the vendored SDK only ever uses GET/POST (registry/register, registry/prekeys, etc.). AGT's switch to REST verbs exposed the gap. Path allowlist also extended with 'v1/' prefix so AGT's REST paths (v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation. Verified end-to-end via azureclaw dev --mesh-provider=agt --build: - POST /v1/agents → 201 Created - PUT /v1/agents/{did}/prekeys → 200 OK - WebSocket /ws accepted, stable connection (no reconnect loop) - Plugin reports 'AGT mesh connected' + provider=agt Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): always route mesh registry through inference-router azureclaw_discover and other mesh registry callsites went via `process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`, intending to let out-of-sandbox sub-agents bypass the router. In practice, the sandbox launcher always sets AGT_REGISTRY_URL as the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082 in dev, the K8s service URL in prod). Since the runtime runs as UID 1000 and iptables egress-guard blocks UID 1000 from anything except localhost+DNS, the direct upstream URL ECONNREFUSEs and the catch-all silently returns []. Symptom: registered agents are invisible to azureclaw_discover even though they show up in `GET /v1/discover` when queried directly at the registry. Drop the env-var override — there's no in-sandbox runtime path where bypassing the router is correct. The router is the ONLY way out for UID 1000. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): break mesh_send infinite poll loop on dead sub-agent probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but routerCall actually resolves with the parsed JSON error body. When the sub-agent pod/container is gone the router returns 404 with { error: "Container '<name>' not found..." } and probeSubAgentAlive read status.phase = undefined → defaulted to "Unknown" → not in POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every 2s forever, blocking the LLM event loop ("LLM not responding" symptom). Also narrow the prekey transient retry test so permanent X3DH / signature-verification failures bubble up instead of being treated as "waiting for prekeys" and retried indefinitely. Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'. Before: registry log fills with GET /v1/discover?capability=echo-buddy every ~2s forever; LLM stops responding to new turns. After: mesh_send aborts with 'sub-agent sandbox not found' on first probe. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): suppress /v1/registry/* 404 leaks in AGT mode Three vendored-only registry paths were being called unconditionally in AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply cadence: 1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X. The AGT registry exposes GET /v1/discover?capability=X instead — display names live in the per-agent record, so the AGT path fans out to a second /v1/agents/{did} fetch per discover hit. Driven by the operator panel's /agt/reputation polling. 2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session. AGT has no per-session counter; per-agent reputation already submitted via MeshClient.submitReputation. No-op in AGT mode. 3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode. Also includes complementary debugging fixes from this session: - agt-transport: auto-call establishSessionWithPeer() before send() so AGT mode gets vendored-equivalent send-with-first-contact semantics. Without this, send() throws 'No encrypted session — call establishSession() first' and the retry loop spins forever. - cli operator fetchers: add 8–10s timeouts to kubectl get calls that were hanging when the cluster API was unreachable. cargo check: clean runtimes/openclaw: 118 vitest tests pass inference-router: 8 mesh tests pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode Fourth 404 leak revealed after deploying the previous fixes: the operator panel's ~30s /agt/reputation poll triggers governance::agt_reputation, which (after lookup_parent_amid succeeds) fetched the per-agent reputation score via the vendored-only GET /v1/registry/reputation/score?amid=X path. AGT registry has no such endpoint — the score is embedded as 'reputation_score: f64' in the per-agent record returned by /v1/agents/{did}. Provider-dispatch the URL; for AGT, wrap the agent record in a vendored-shaped payload (score / tier / raw) so downstream CLI fetchers and the operator panel stay schema-agnostic. cargo check: clean agt_governance_integration: 26/26 pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * runtime: hide Foundry tools in github-copilot mode (same as github-models) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest Phase B.1 of the AGT-on-AKS rollout (see session plan files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the existing `azureclaw dev --mesh-provider` flag so the same provider selection works for AKS pushes. When --mesh-provider=agt: * Builds relay+registry from the AGT upstream Dockerfile ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile) using COMPONENT=relay|registry build-args (matches dev.ts). * Tags as agentmesh-{relay,registry}-agt:latest so both vendored and AGT images can coexist on the same ACR and so the existing deploy/agentmesh-agt.yaml manifest picks them up unchanged. * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz) into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg. * Always passes MESH_PROVIDER build-arg to the sandbox image so the Dockerfile's conditional `npm install @microsoft/agent-governance-sdk` runs for AGT clusters. When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml, applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt, THEN rolls the controller (so the new pod reads AZURECLAW_MESH_PROVIDER=agt for new sandboxes). Auto-reverses when --apply --mesh-provider=vendored runs against a cluster currently on AGT (no Postgres deployment in the agentmesh ns). The image build loop also now supports absolute Dockerfile paths and absolute build contexts via a new `absoluteContext` field, needed because the AGT Dockerfile lives outside the azureclaw repo root. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks without rebuilding any images, assuming both image pairs were already seeded by 'azureclaw push'. Flow: 1. Detect current provider via 'kubectl get deploy/postgres -n agentmesh' (vendored has Postgres, AGT does not). 2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found 3. kubectl apply -f deploy/agentmesh-<target>.yaml 4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target> 5. kubectl rollout restart deploy/azureclaw-controller 6. With --restart-sandboxes: roll every azureclaw-managed Deployment so existing pods pick up the new AZURECLAW_MESH_PROVIDER value. Service names and ports are identical between the two manifests (agentmesh-relay:8765, agentmesh-registry:8080) so the controller's mesh_peer talks to either stack with no further config — the relay/ registry URLs already come from env vars (MESH_RELAY_URL / MESH_REGISTRY_URL). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(up): --mesh-provider=agt picks AGT manifest + flips helm value Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up' so first-time deploys can ship AGT instead of vendored. When --mesh-provider=agt: * helm install runs with --set mesh.provider=agt (controller env AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes). * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of deploy/agentmesh.yaml. * Skips the postgres ACR import and the agentmesh-db-credentials secret creation (both unused by AGT — its registry is in-memory). * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml vs .tmp-agentmesh.yaml) so concurrent provider switches don't collide. The deployAgentMesh signature gains a non-breaking 'meshProvider' option that defaults to 'vendored' (existing callers untouched). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): plumb --mesh-provider into local-k8s helm install Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(controller): AGT wire protocol adapter for mesh_peer Implement full AGT relay/registry wire support in the controller's mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt. Without this the controller's federation peer cannot connect to the AGT relay (different WS path, frame envelope, heartbeat, ack model) or the AGT registry (different HTTP shape, no signed body), and the leader fails-loops on AGT clusters — breaking the only cloud-offload control path. New module `mesh_peer/agt_wire.rs`: - `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error) with `#[serde(tag="type", rename_all="snake_case")]` matching `agentmesh/relay/app.py`. - `AgtRegisterAgentRequest` struct for `POST /v1/agents`. - 7 unit tests pinning the serialized shape. `mesh_peer/mod.rs`: - New `Provider` enum + `Provider::from_env()` selecting vendored (default) or AGT off `AZURECLAW_MESH_PROVIDER`. - `MeshPeerState.provider` carried through outbound + inbound paths. - `register_with_registry()` branches: vendored signs ts body; AGT posts `{did, public_key (base64url), capabilities, metadata}` with no signature; 409 treated as success for leader-failover idempotency. - `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>` (matches JS SDK `buildDid`), so every leader replica converges on the same DID without coordination. - Default `MESH_RELAY_URL` appends `/ws` for AGT. - `connect_and_listen()`: - AGT connect frame `{type:"connect", from:<did>, token?:<env>}` (token read from `AGENTMESH_RELAY_TOKEN` if set). - AGT has no `Connected` ack — mark `connected=true` immediately. - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored keeps `ping`). - `serialize_and_send_outbound()` / `send_to_peer()` now take `state` and branch outbound framing — AGT emits `message` frames `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex). - `handle_message()` dispatches to `handle_vendored_frame()` or `handle_agt_frame()`. AGT path: - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`. - Sends `Ack` reply (required — without it AGT redelivers on reconnect → duplicate offload processing). - Treats `Error` frames mentioning Authentication failed / Missing 'from' / session_replaced as fatal — drops connection for reconnect. `mesh_peer/offload.rs`: - All 8 `send_to_peer(...)` call sites updated to pass `&state` first. `main.rs`: - Remove the temporary AGT-skip guard around `mesh_peer::run`. The peer now starts unconditionally when enabled; provider is consumed inside `mesh_peer::run`. Build/test: - cargo build --release --package azureclaw-controller: OK - cargo test --package azureclaw-controller: 492 passed - cargo clippy --package azureclaw-controller --all-targets -D warnings: OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer - cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(deploy): AGT mesh probe path + Cilium pod-port NP allow deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz (see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness probes were 404'ing → CrashLoopBackOff/NotReady. operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane evaluates NetworkPolicy egress against the backend pod port (post-DNAT), not the Service port. AGT registry/relay listen on 8082/8083; the Service maps 8080->8082 and 8765->8083 so the Service-port allowlist (8080/8765) doesn't actually permit the post-DNAT flow. Add 8082/8083 alongside so both vendored (8080/8765 direct) and AGT paths work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh promote): AGT-compat health + WS upgrade paths azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): first-run picker for local vs remote mesh source azureclaw dev now asks new users where the mesh should live, just like the existing inference-provider picker: Where should the mesh live? ❯ Local (recommended; spin up relay + registry in Docker) Remote (auto port-forward to AKS cluster: <cluster-name>) Local (default) keeps the existing behaviour: docker-compose'd relay/registry/postgres on the user's laptop. Remote (advanced) federates with a previously-provisioned AKS mesh: - If ~/.azureclaw/context.json has a cached globalRegistryUrl from a prior 'azureclaw mesh promote', reuse it verbatim. - Otherwise default to http://localhost:18080 — the port-forward URL 'mesh promote --port-forward' uses — so the auto-promote fallback in the downstream global-registry block will spawn the tunnels on demand. - If there is no aksCluster in context at all, warn and fall back to local so the user isn't left with a broken sandbox. Skipped entirely when --global-registry was passed explicitly (the advanced flag overrides the prompt) or when the user is past their first run. Also fixed a latent AGT-compat bug in the same flow: the existing 'auto-promote' path probed only /v1/health, which 404s on AGT clusters. Replaced with a /health → /v1/health fallback (matches the same shape we used in checkRegistryHealth last commit). Verified: - npm run build / typecheck clean - 640 CLI tests pass (2 skipped, no regressions) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply back to their spawner, but on AGT the registry has no agent named or capability="parent" — the search returns 0 → no prekey bundle → send fails. The vendored runtime had this aliased only in the offload-mode task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the controller never set for AKS-spawned children. Two coordinated fixes: 1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set (spawner seeds 'parent_name:parent_AMID' as the first entry), also push PARENT_SANDBOX=<first_name> into the openclaw container env. 2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send and azureclaw_mesh_transfer_file, alias to_agent=='parent' → PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry lookup. The Symbol is set during runtime init from AGT_TRUSTED_PEERS[0], so this works even on images built before fix #1 lands. Skip in offload mode — 'parent' there is a protocol-level routing token, not a mesh recipient name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap mesh-plugin/src/agt-transport.ts.send() called this.client.establishSessionWithPeer(toAmid) before forwarding to client.send(). That method does not exist on AgentMeshClient — the real method is establishSession(toAmid, options) — so every parent → sub-agent send on AGT was failing with: establishSessionWithPeer is not a function It was also unnecessary: AgentMeshClient.send() already auto-bootstraps the X3DH handshake on first contact (see @agentmesh/sdk AgentMeshClient.send → cache miss → establishSession() fallthrough at dist/index.js:3321-3334). Calling establishSession() ourselves would also be wrong because it is not idempotent — it unconditionally writes activeSessions.set and starts a fresh X3DH. Fix: remove the pre-bootstrap entirely and let client.send() manage session lifecycle. The AgtSdkModule type loses the required establishSessionWithPeer member (now optional) since we no longer depend on it; test fakes remain valid as harmless extras. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call Previous commit aa7d28e wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * push: auto-detect mesh provider from live helm release When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * entrypoint: fail-open trust gate when running anonymous tier When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable The AGT registry has no autonomous presence model — `last_seen` is frozen at registration and the `update_last_seen()` store method is dead code with no HTTP handler calling it. Combined with the openclaw discover tool's 90s stale filter (agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes silently invisible 90s after spawn, breaking sibling-to-sibling peer discovery. Demo symptom: analyst/viz/writer all reported 'peer discovery did not return ...' even though mesh_send to those names succeeded with 'delivered_and_replied'. The relay was fine; only the registry's presence view was stale. Pair with the corresponding upstream registry change (AGT branch `azureclaw-meshclient-event-hooks`, commit adds POST /v1/agents/{did}/heartbeat -> store.update_last_seen). The new tick reuses the existing 30s relay-keepalive timer in connect(), so no extra timers and no extra event-loop pressure. Best-effort: 4xx/5xx are warned-once, network errors swallowed, loop survives a registry pod restart (next tick retries). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
Security hardening phase 1 + Foundry integration improvements
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
…#245) * feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap Wire azureclaw runtime through createMeshTransport() factory so we can flip between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without code changes. Surface additions to IMeshTransport (both adapters now expose): - lookup(amid) — registry RPC for reputation/display name - submitReputation(...) — registry RPC for peer feedback - enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on) - onError(kind, from, detail) — diagnostic hook for decrypt + ws errors - onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal - onDisconnect(reason, code) — ws close / error fan-out mesh-plugin (vendored A adapter): - connection.ts delegates to the underlying SDK; lazy bind for hooks registered before connect() - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the contract so neither adapter can drop a method without CI failing mesh-plugin (AGT B adapter): - agt-transport.ts implements lookup/submitReputation as REST calls to the registry (AGT MeshClient is pure transport — registry RPCs intentionally not added to AGT upstream; they belong on a separate RegistryClient) - enableKnockEnforcement is a no-op (AGT MeshClient always enforces) - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods (added on local AGT branch azureclaw-meshclient-event-hooks, NOT pushed — AGT team owns the upstream PR) runtime (runtimes/openclaw): - Adds @azureclaw/mesh as a file: dependency - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)' when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value - Identity is generated once via vendored SDK regardless of provider, then raw Ed25519 keys are extracted via toData() and shared across both — same AMID either way - Banner now reports active provider (vendored vs agt) Docs: - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity, policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer + the wiring + migration path - Documents the 3 hooks added to local AGT branch and the 3 governance methods kept adapter-side Tests: - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat) - runtimes/openclaw: 118/118 pass - AGT (local branch): 387/387 pass with 8 new event-hook tests Open work for cleanup phase: once AGT publishes the version with our event hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var toggle. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): pre-build mesh-plugin for runtime CI + format reconciler PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): quote workflow step name containing colon YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): clippy fixes for Rust 1.95.0 CI runs Rust 1.95.0 which added new clippy lints: - doc_lazy_continuation: indent doc list items that span multiple lines. Added two-space indent to the trailing 'All three are populated...' paragraph so it is treated as a continuation of the preceding list item rather than its own malformed list item. - obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as if .. { a } else { b } per the lint suggestion. These were pre-existing on dev (CI only started failing once the runner picked up Rust 1.95.0); fixing here so PR #245 can land green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12 Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5 The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): mark gaps G1 and G2 fixed on local AGT branch Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass. * dev: add --mesh-provider <vendored|agt> selection with first-run prompt Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker mode before AKS rollout. Same flag, three integration points: CLI (cli/src/commands/dev.ts): - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball - First-run interactive prompt offers AGT only if the toolkit checkout is actually present locally — silently defaults to vendored otherwise (no pestering for users without AGT cloned). - --build branch: builds the right relay/registry images vendored → vendor/agentmesh-relay + agentmesh-registry (Rust) agt → agent-governance-python/agent-mesh/docker/Dockerfile with COMPONENT=relay / registry build-args - Sandbox image build: stages locally-packed AGT SDK tarball into .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL build-arg (auto-discovers if --agt-sdk-tarball not given). - Runtime branch: skips Postgres for AGT (in-memory registry), uses correct ports (AGT: 8083 relay, 8082 registry; vendored: 8765/8080) and health path (AGT: /healthz; vendored: /v1/health). - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the runtime transport-factory honors the user's choice. Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile): - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the sandbox npm-installs the local tarball instead of fetching the published @microsoft/agent-governance-sdk from npm. Lets us smoke-test the locally-patched AGT branch (G3/G4 fixes) end to end without round-tripping through npm publish. - .agt-sdk/ staging dir always exists (with .keep) so the COPY never fails when the user didn't stage a tarball. Defaults preserved: --mesh-provider=vendored, existing behavior is byte-identical for users who don't opt in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so tsc failed with TS2307 in the AGT build path. Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the build context, and strip its 'prepare' script (which would invoke tsc, not present in this stage; the pre-built dist/ is sufficient). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks). - connect() now passes autoRegister: true so the SDK uploads identity and prekeys instead of the adapter re-implementing that path with raw HTTP. - discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT. - lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}). - submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score clamped to [0,1]; the vendored /registry/feedback endpoint does not exist in AGT. - Replaced mapAgent with pickDisplayName helper: AGT puts display name in metadata.display_name (set by registerSelf), with the first capability as the fallback. Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and the bespoke fetchWithRetry helper — both are upstream concerns now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes the vendored registry wire shape. The vendored impl talks to /registry/* (the existing agentmesh-registry); the AGT impl talks to /v1/discover and /v1/agents/{did} on the upstream AGT registry. Both expose a single normalized RegistryEntry envelope, so callsites stay readable. getMeshRegistry(routerUrl) is the entry point. Provider selection follows AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base). Migrated all raw-HTTP registry callsites: - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName, resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid. - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS spawn discovery, AKS→local discovery. - core/agt-task-loop.ts (1 site): registry_capability_search tool. - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe, azureclaw_discover (mesh_discover) tool. - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe, heartbeat keepalive (no-op under AGT — relay does liveness via WS). The discover-on-router-unreachable test now asserts the new contract: empty list + count:0 instead of a 'Discovery failed' string. Registry hiccups must not break tool calls. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): wire AGT provider end-to-end (6 stackup bugs) End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced six bugs blocking the upstream AGT MeshClient swap. All fixed: 1. Final sandbox Docker stage didn't COPY mesh-plugin, so the file:../../mesh-plugin symlink dangled in node_modules. Plugin swap silently fell back to vendored with 'Cannot find package @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist} into /mesh-plugin/ in the final stage of the Dockerfile. 2. entrypoint.sh used cp -r when copying node_modules into the plugin extension dir, preserving the (now-broken-at-runtime-path) symlink. Switched to cp -rL so symlinks dereference into real files in the target tree. 3. mesh-plugin/src/index.ts imported createMeshTransport from ./transport-factory.js but never re-exported it. Runtime swap path couldn't find the factory. Added the missing re-export. 4. inference-router agt_registry_proxy unconditionally prepended '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents' became '/v1/v1/agents' at the upstream. Now: forward verbatim when path starts with 'v1/' or equals 'health', else prepend. Preserves vendored SDK behavior ('registry/register' → /v1/registry/register). 5. /agt/relay route only matched the bare path, but AGT MeshClient appends '/ws' to relayUrl. Added /agt/relay/ws route and made the upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt. 6. agt_registry_proxy route was declared get(...).post(...) only. AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey upload and DELETE for deregister — both 405'd at the router. Added .put() and .delete() to the route declaration. Bug #6 was invisible to vendored because the vendored SDK only ever uses GET/POST (registry/register, registry/prekeys, etc.). AGT's switch to REST verbs exposed the gap. Path allowlist also extended with 'v1/' prefix so AGT's REST paths (v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation. Verified end-to-end via azureclaw dev --mesh-provider=agt --build: - POST /v1/agents → 201 Created - PUT /v1/agents/{did}/prekeys → 200 OK - WebSocket /ws accepted, stable connection (no reconnect loop) - Plugin reports 'AGT mesh connected' + provider=agt Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): always route mesh registry through inference-router azureclaw_discover and other mesh registry callsites went via `process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`, intending to let out-of-sandbox sub-agents bypass the router. In practice, the sandbox launcher always sets AGT_REGISTRY_URL as the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082 in dev, the K8s service URL in prod). Since the runtime runs as UID 1000 and iptables egress-guard blocks UID 1000 from anything except localhost+DNS, the direct upstream URL ECONNREFUSEs and the catch-all silently returns []. Symptom: registered agents are invisible to azureclaw_discover even though they show up in `GET /v1/discover` when queried directly at the registry. Drop the env-var override — there's no in-sandbox runtime path where bypassing the router is correct. The router is the ONLY way out for UID 1000. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): break mesh_send infinite poll loop on dead sub-agent probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but routerCall actually resolves with the parsed JSON error body. When the sub-agent pod/container is gone the router returns 404 with { error: "Container '<name>' not found..." } and probeSubAgentAlive read status.phase = undefined → defaulted to "Unknown" → not in POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every 2s forever, blocking the LLM event loop ("LLM not responding" symptom). Also narrow the prekey transient retry test so permanent X3DH / signature-verification failures bubble up instead of being treated as "waiting for prekeys" and retried indefinitely. Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'. Before: registry log fills with GET /v1/discover?capability=echo-buddy every ~2s forever; LLM stops responding to new turns. After: mesh_send aborts with 'sub-agent sandbox not found' on first probe. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): suppress /v1/registry/* 404 leaks in AGT mode Three vendored-only registry paths were being called unconditionally in AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply cadence: 1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X. The AGT registry exposes GET /v1/discover?capability=X instead — display names live in the per-agent record, so the AGT path fans out to a second /v1/agents/{did} fetch per discover hit. Driven by the operator panel's /agt/reputation polling. 2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session. AGT has no per-session counter; per-agent reputation already submitted via MeshClient.submitReputation. No-op in AGT mode. 3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode. Also includes complementary debugging fixes from this session: - agt-transport: auto-call establishSessionWithPeer() before send() so AGT mode gets vendored-equivalent send-with-first-contact semantics. Without this, send() throws 'No encrypted session — call establishSession() first' and the retry loop spins forever. - cli operator fetchers: add 8–10s timeouts to kubectl get calls that were hanging when the cluster API was unreachable. cargo check: clean runtimes/openclaw: 118 vitest tests pass inference-router: 8 mesh tests pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode Fourth 404 leak revealed after deploying the previous fixes: the operator panel's ~30s /agt/reputation poll triggers governance::agt_reputation, which (after lookup_parent_amid succeeds) fetched the per-agent reputation score via the vendored-only GET /v1/registry/reputation/score?amid=X path. AGT registry has no such endpoint — the score is embedded as 'reputation_score: f64' in the per-agent record returned by /v1/agents/{did}. Provider-dispatch the URL; for AGT, wrap the agent record in a vendored-shaped payload (score / tier / raw) so downstream CLI fetchers and the operator panel stay schema-agnostic. cargo check: clean agt_governance_integration: 26/26 pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * runtime: hide Foundry tools in github-copilot mode (same as github-models) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest Phase B.1 of the AGT-on-AKS rollout (see session plan files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the existing `azureclaw dev --mesh-provider` flag so the same provider selection works for AKS pushes. When --mesh-provider=agt: * Builds relay+registry from the AGT upstream Dockerfile ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile) using COMPONENT=relay|registry build-args (matches dev.ts). * Tags as agentmesh-{relay,registry}-agt:latest so both vendored and AGT images can coexist on the same ACR and so the existing deploy/agentmesh-agt.yaml manifest picks them up unchanged. * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz) into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg. * Always passes MESH_PROVIDER build-arg to the sandbox image so the Dockerfile's conditional `npm install @microsoft/agent-governance-sdk` runs for AGT clusters. When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml, applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt, THEN rolls the controller (so the new pod reads AZURECLAW_MESH_PROVIDER=agt for new sandboxes). Auto-reverses when --apply --mesh-provider=vendored runs against a cluster currently on AGT (no Postgres deployment in the agentmesh ns). The image build loop also now supports absolute Dockerfile paths and absolute build contexts via a new `absoluteContext` field, needed because the AGT Dockerfile lives outside the azureclaw repo root. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks without rebuilding any images, assuming both image pairs were already seeded by 'azureclaw push'. Flow: 1. Detect current provider via 'kubectl get deploy/postgres -n agentmesh' (vendored has Postgres, AGT does not). 2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found 3. kubectl apply -f deploy/agentmesh-<target>.yaml 4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target> 5. kubectl rollout restart deploy/azureclaw-controller 6. With --restart-sandboxes: roll every azureclaw-managed Deployment so existing pods pick up the new AZURECLAW_MESH_PROVIDER value. Service names and ports are identical between the two manifests (agentmesh-relay:8765, agentmesh-registry:8080) so the controller's mesh_peer talks to either stack with no further config — the relay/ registry URLs already come from env vars (MESH_RELAY_URL / MESH_REGISTRY_URL). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(up): --mesh-provider=agt picks AGT manifest + flips helm value Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up' so first-time deploys can ship AGT instead of vendored. When --mesh-provider=agt: * helm install runs with --set mesh.provider=agt (controller env AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes). * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of deploy/agentmesh.yaml. * Skips the postgres ACR import and the agentmesh-db-credentials secret creation (both unused by AGT — its registry is in-memory). * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml vs .tmp-agentmesh.yaml) so concurrent provider switches don't collide. The deployAgentMesh signature gains a non-breaking 'meshProvider' option that defaults to 'vendored' (existing callers untouched). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): plumb --mesh-provider into local-k8s helm install Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(controller): AGT wire protocol adapter for mesh_peer Implement full AGT relay/registry wire support in the controller's mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt. Without this the controller's federation peer cannot connect to the AGT relay (different WS path, frame envelope, heartbeat, ack model) or the AGT registry (different HTTP shape, no signed body), and the leader fails-loops on AGT clusters — breaking the only cloud-offload control path. New module `mesh_peer/agt_wire.rs`: - `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error) with `#[serde(tag="type", rename_all="snake_case")]` matching `agentmesh/relay/app.py`. - `AgtRegisterAgentRequest` struct for `POST /v1/agents`. - 7 unit tests pinning the serialized shape. `mesh_peer/mod.rs`: - New `Provider` enum + `Provider::from_env()` selecting vendored (default) or AGT off `AZURECLAW_MESH_PROVIDER`. - `MeshPeerState.provider` carried through outbound + inbound paths. - `register_with_registry()` branches: vendored signs ts body; AGT posts `{did, public_key (base64url), capabilities, metadata}` with no signature; 409 treated as success for leader-failover idempotency. - `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>` (matches JS SDK `buildDid`), so every leader replica converges on the same DID without coordination. - Default `MESH_RELAY_URL` appends `/ws` for AGT. - `connect_and_listen()`: - AGT connect frame `{type:"connect", from:<did>, token?:<env>}` (token read from `AGENTMESH_RELAY_TOKEN` if set). - AGT has no `Connected` ack — mark `connected=true` immediately. - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored keeps `ping`). - `serialize_and_send_outbound()` / `send_to_peer()` now take `state` and branch outbound framing — AGT emits `message` frames `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex). - `handle_message()` dispatches to `handle_vendored_frame()` or `handle_agt_frame()`. AGT path: - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`. - Sends `Ack` reply (required — without it AGT redelivers on reconnect → duplicate offload processing). - Treats `Error` frames mentioning Authentication failed / Missing 'from' / session_replaced as fatal — drops connection for reconnect. `mesh_peer/offload.rs`: - All 8 `send_to_peer(...)` call sites updated to pass `&state` first. `main.rs`: - Remove the temporary AGT-skip guard around `mesh_peer::run`. The peer now starts unconditionally when enabled; provider is consumed inside `mesh_peer::run`. Build/test: - cargo build --release --package azureclaw-controller: OK - cargo test --package azureclaw-controller: 492 passed - cargo clippy --package azureclaw-controller --all-targets -D warnings: OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer - cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(deploy): AGT mesh probe path + Cilium pod-port NP allow deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz (see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness probes were 404'ing → CrashLoopBackOff/NotReady. operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane evaluates NetworkPolicy egress against the backend pod port (post-DNAT), not the Service port. AGT registry/relay listen on 8082/8083; the Service maps 8080->8082 and 8765->8083 so the Service-port allowlist (8080/8765) doesn't actually permit the post-DNAT flow. Add 8082/8083 alongside so both vendored (8080/8765 direct) and AGT paths work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh promote): AGT-compat health + WS upgrade paths azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): first-run picker for local vs remote mesh source azureclaw dev now asks new users where the mesh should live, just like the existing inference-provider picker: Where should the mesh live? ❯ Local (recommended; spin up relay + registry in Docker) Remote (auto port-forward to AKS cluster: <cluster-name>) Local (default) keeps the existing behaviour: docker-compose'd relay/registry/postgres on the user's laptop. Remote (advanced) federates with a previously-provisioned AKS mesh: - If ~/.azureclaw/context.json has a cached globalRegistryUrl from a prior 'azureclaw mesh promote', reuse it verbatim. - Otherwise default to http://localhost:18080 — the port-forward URL 'mesh promote --port-forward' uses — so the auto-promote fallback in the downstream global-registry block will spawn the tunnels on demand. - If there is no aksCluster in context at all, warn and fall back to local so the user isn't left with a broken sandbox. Skipped entirely when --global-registry was passed explicitly (the advanced flag overrides the prompt) or when the user is past their first run. Also fixed a latent AGT-compat bug in the same flow: the existing 'auto-promote' path probed only /v1/health, which 404s on AGT clusters. Replaced with a /health → /v1/health fallback (matches the same shape we used in checkRegistryHealth last commit). Verified: - npm run build / typecheck clean - 640 CLI tests pass (2 skipped, no regressions) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply back to their spawner, but on AGT the registry has no agent named or capability="parent" — the search returns 0 → no prekey bundle → send fails. The vendored runtime had this aliased only in the offload-mode task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the controller never set for AKS-spawned children. Two coordinated fixes: 1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set (spawner seeds 'parent_name:parent_AMID' as the first entry), also push PARENT_SANDBOX=<first_name> into the openclaw container env. 2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send and azureclaw_mesh_transfer_file, alias to_agent=='parent' → PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry lookup. The Symbol is set during runtime init from AGT_TRUSTED_PEERS[0], so this works even on images built before fix #1 lands. Skip in offload mode — 'parent' there is a protocol-level routing token, not a mesh recipient name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap mesh-plugin/src/agt-transport.ts.send() called this.client.establishSessionWithPeer(toAmid) before forwarding to client.send(). That method does not exist on AgentMeshClient — the real method is establishSession(toAmid, options) — so every parent → sub-agent send on AGT was failing with: establishSessionWithPeer is not a function It was also unnecessary: AgentMeshClient.send() already auto-bootstraps the X3DH handshake on first contact (see @agentmesh/sdk AgentMeshClient.send → cache miss → establishSession() fallthrough at dist/index.js:3321-3334). Calling establishSession() ourselves would also be wrong because it is not idempotent — it unconditionally writes activeSessions.set and starts a fresh X3DH. Fix: remove the pre-bootstrap entirely and let client.send() manage session lifecycle. The AgtSdkModule type loses the required establishSessionWithPeer member (now optional) since we no longer depend on it; test fakes remain valid as harmless extras. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call Previous commit 34662c7 wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * push: auto-detect mesh provider from live helm release When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * entrypoint: fail-open trust gate when running anonymous tier When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable The AGT registry has no autonomous presence model — `last_seen` is frozen at registration and the `update_last_seen()` store method is dead code with no HTTP handler calling it. Combined with the openclaw discover tool's 90s stale filter (agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes silently invisible 90s after spawn, breaking sibling-to-sibling peer discovery. Demo symptom: analyst/viz/writer all reported 'peer discovery did not return ...' even though mesh_send to those names succeeded with 'delivered_and_replied'. The relay was fine; only the registry's presence view was stale. Pair with the corresponding upstream registry change (AGT branch `azureclaw-meshclient-event-hooks`, commit adds POST /v1/agents/{did}/heartbeat -> store.update_last_seen). The new tick reuses the existing 30s relay-keepalive timer in connect(), so no extra timers and no extra event-loop pressure. Best-effort: 4xx/5xx are warned-once, network errors swallowed, loop survives a registry pod restart (next tick retries). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Security hardening, Foundry API compatibility fixes, and infrastructure improvements for alpha release.
Security
Foundry / AOAI Compatibility
max_completion_tokensfor GPT-5.x models (replaces rejectedmax_tokens)/v1/images/generationsroute stripsresponse_format(Azure rejects it)Infrastructure
telegram-token.dev/.cloud)--channels telegram.cloud34 commits, 160 tests passing, CLI + router compile clean