Skip to content

Security hardening phase 1 + Foundry integration improvements - #15

Merged
Pal Lakatos-Toth (pallakatos) merged 38 commits into
mainfrom
security/hardening-phase1
Apr 1, 2026
Merged

Pal Lakatos-Toth (pallakatos) merged 38 commits into
mainfrom
security/hardening-phase1

Conversation

@pallakatos

Copy link
Copy Markdown
Collaborator

Summary

Security hardening, Foundry API compatibility fixes, and infrastructure improvements for alpha release.

Security

  • Crypto, policy, logging, and validation hardening (phase 1 + phase 2)
  • AGT as sole policy authority for sub-agent tasks
  • OpenClaw exec approvals disabled via config (AGT handles governance)
  • Go 1.25 builder + npm audit fix to reduce dependency CVEs
  • Sidecar admin token race fix

Foundry / AOAI Compatibility

  • max_completion_tokens for GPT-5.x models (replaces rejected max_tokens)
  • Responses API proxy with auto chat↔responses format translation
  • Image generation: /v1/images/generations route strips response_format (Azure rejects it)
  • Image gen timeout extended to 90s
  • Cache Responses-only models to skip redundant attempts

Infrastructure

  • Credential variant resolution (telegram-token.dev / .cloud)
  • Channel variant syntax: --channels telegram.cloud
  • Unified secrets store for all credentials
  • Spawn latency optimization (~2min → ~15s)
  • Mesh handshake latency optimization
  • Parent-mediated trust delegation for mesh siblings

34 commits, 160 tests passing, CLI + router compile clean

Pal Lakatos-Toth and others added 30 commits March 31, 2026 17:54
Phase 1 — Secrets:
- Remove hardcoded PostgreSQL password from agentmesh.yaml, use K8s Secret ref
- Read API key from file at runtime instead of hardcoding in env export
- Add chmod 0400 on /tmp/azure-openai-key fallback
- Remove /tmp/azure-openai-key fallback from router auth.rs

Phase 2 — Authentication:
- Auto-generate ADMIN_TOKEN from /dev/urandom when not configured
- Add admin token auth to sidecar GET endpoints (/trust, /audit, /audit/verify)
- Validate x-azureclaw-sandbox header against K8s name regex

Phase 3 — Fail-closed:
- Router: fail-closed after 3 consecutive sidecar failures (grace window for cold start)
- Plugin: same fail-closed pattern in evaluateAGTPolicy()
- Spawn endpoint: fail-closed after grace window

Phase 4 — Command injection:
- Replace execSync(curl) with Node http.request() for http_fetch
- Gate fallback exec_command through AGT policy before execution
- Validate sandbox name format at controller reconcile entry point

Phase 5 — Cryptographic:
- Replace Math.random() with crypto.randomUUID() in 3 locations
- Fix timing attack in bytesEqual() — use XOR accumulation

Phase 6 — Policy/Trust:
- Allowlist registry proxy paths (prevent path traversal)
- Add threading.Lock around trust score get+update (race condition fix)

All 321 tests pass (47 router + 80 controller + 148 TS + 46 Python).
Clippy clean with -D warnings.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase 5 — Crypto:
- Add AES-GCM nonce tracking (collision detection via Set) in encrypted audit
- X25519 key length validation (32-byte check) + all-zero DH output rejection
- X3DH signature failure logging for audit trail

Phase 6 — Policy:
- Strengthen recon-deny regex patterns (context-aware boundaries vs \b)

Phase 7 — Network:
- Truncate domains >100 chars in log output (prevent log flooding)
- Guard IMDS fallback: require AZURE_TENANT_ID before attempting IMDS token

Phase 8 — Remaining:
- Log sanitization: strip ANSI escapes + collapse newlines in user input
- SANDBOX_NAME format validation (K8s name regex)
- Rate limit sidecar GET endpoints (/status, /trust, /audit)
- Helm values: document image tag pinning for production

All 321 tests pass. Clippy clean with -D warnings.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…via env

The AGT sidecar was starting before the admin token file was written,
and the token file (chmod 400, owned sandbox:sandbox) was unreadable by
the sidecar process (UID 1001/router). This left ADMIN_TOKEN empty,
causing _check_admin_token() to skip auth on /trust and /audit endpoints.

Fixes:
- Generate ROUTER_ADMIN_TOKEN before starting the sidecar
- Pass ADMIN_TOKEN env var directly to the sidecar process
- server.py now prefers env var over file paths (avoids permission issues)
- Token file chmod 440 with router added to sandbox group (belt-and-suspenders)
- Log warning when no admin token is configured

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Root cause: spawn polling called GET /sandbox/{name} but the route is
GET /sandbox/{name}/status — every poll returned 405, the catch block
swallowed it, and the loop ran all 24×5s=120s iterations.

Fixes:
- Correct polling path to /sandbox/{name}/status
- Reduce poll interval: 5s→1s, cap at 45s (was 120s)
- Replace blind 3s sleep with registry discovery poll (1s×10 max)
- Log every 5th poll instead of every poll (reduce noise)
- Add 'registry/' prefix to registry path allowlist (fixes
  azureclaw_discover 400 'Invalid registry path' error)
- Replace entrypoint sleeps with health-gate loops (saves ~4s)
  - Router: sleep 1 → poll /healthz every 200ms
  - Gateway: sleep 1 → sleep 0.5
  - Auto-approve: sleep 2 → poll gateway healthz every 500ms

Expected spawn timeline: ~10-15s (container boot + poll detect + registry)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
After security hardening added admin token auth on sidecar GET /trust,
/audit, /audit/verify, the router's SidecarProxy was forwarding
requests without the token. This caused:

1. Operator view: 'no audit entries' — GET /agt/audit → sidecar
   returned 403 because router didn't include admin token
2. Mesh trust lookup: incoming mesh messages had trust score 0
   because the plugin's direct sidecar call got 403

Proper fix (not a hack):
- SidecarProxy reads ADMIN_TOKEN env var at init, includes it as
  Authorization header on all forwarded requests (same shared secret)
- Plugin mesh trust lookup routes through /agt/trust/{id} on the
  router (port 8443) instead of directly hitting sidecar (port 8081)
- Works in both dev (same container) and AKS (same pod, shared secret
  via K8s Secret volume mount)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Switch foundry_image_generation tool from the Responses API (which
required an agent_reference + tool orchestration) to the standard
Azure OpenAI Images API:

  POST /openai/deployments/{model}/images/generations

This matches the working Python SDK call:
  client.images.generate(model='gpt-image-1', prompt='...', n=1, size='1024x1024')

Changes:
- Router: add /openai/deployments/{deployment}/images/generations route
  with AGT policy gate and proxy to Azure OpenAI endpoint
- Plugin: call standard images API instead of Responses API with tools
- Remove unnecessary orchestrator model parameter
- Verified: direct API call returns 200 with b64_json image data

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Switch proxy to unified /openai/v1/ URL format (model in body) for both
  dev (API key) and AKS (Entra) modes — all models use same endpoint
- Auth always uses Authorization: Bearer (works with both API keys and
  Entra tokens on the unified endpoint)
- Add /v1/responses route for Responses-only models (e.g. gpt-5.4-pro)
  with AGT policy gate and budget tracking
- Auto-fallback in chat_completions: if Azure returns 400 'unsupported',
  transparently retry via Responses API with format translation
  (messages→input, max_completion_tokens→max_output_tokens)
- Convert Responses API output back to chat/completions format so
  OpenClaw works transparently without changes
- Fix image generation path (was double-deploying deployment in URL)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agents spawned by the same parent can now communicate via mesh
without KNOCK rejection. Trust is established securely:

- Parent's router passes AGT_TRUSTED_PEERS env var at spawn time
  containing parent-verified AMIDs (name:AMID pairs)
- Sub-agent plugin reads env var at startup, pre-seeds trust store
  via admin-token-protected /agt/trust endpoint (local sidecar)
- KNOCK handler gives +500 affinity bonus to parent-verified AMIDs
  (stored in parentTrustedAmids set, separate from general lookups)
- Parent's spawn tool builds peer list from its own AMID + all
  existing siblings' AMIDs (from amidToName map)

Security properties:
- Trust source is the parent's router (sets env var), not self-reported
- Env vars are set at container creation time, not modifiable by sandbox
- parentTrustedAmids is a separate set from amidToName to prevent
  trust escalation via arbitrary registry lookups
- Admin token required for trust mutations (unchanged)
- Works identically on Docker (dev) and AKS (production)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…attempts

Models like gpt-5.4-pro don't support chat/completions — every request
was doing a wasted round-trip (400 in ~0.8s) before falling back to
Responses API. Now the first fallback caches the model name in an
in-memory HashSet, and all subsequent requests go directly to
/openai/v1/responses.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Handle tool_calls → function_call items in input conversion
- Handle tool role → function_call_output items
- Handle string content → typed content blocks (input_text/output_text)
- Convert system role → developer for Responses API
- Add type:'message' wrapper for regular messages
- Handle function_call → tool_calls in response conversion
- Set finish_reason to 'tool_calls' when function calls present
- Add debug logging for upstream requests (URL, body_len, resp_len)
- Add stream status logging
- Add 4 unit tests for conversion functions

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The spawn tool was unconditionally force-removing existing containers
before creating new ones. When the LLM calls spawn twice for the same
sub-agent (retry or re-invocation), the running container gets killed
mid-operation, causing:
- 'Connection reset without closing handshake' in relay logs
- New AMID on reconnect (new identity = broken sessions)
- 'Failed to send message: closed connection' for stale AMIDs

Now checks if container is already running and reuses it instead.
Only removes stopped/dead containers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
On AKS, if the LLM calls spawn twice for the same sub-agent, the K8s
API returns 409 Conflict. Previously this was a hard error — now it
returns success with 'already running', consistent with the Docker
path fix.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Three optimizations that reduce initial mesh communication lag:

1. AMID cache-hit in mesh_send: skip registry search if AMID was
   pre-cached during spawn. Falls back to re-discovery if cached
   AMID is stale (agent not found / connection closed).
   Saves: 2-12s on first message to known agents.

2. Parallel spawn readiness: poll container status AND registry
   search simultaneously instead of sequentially.
   Saves: 5-10s during spawn (registry search overlaps boot).

3. SDK parallel registry calls: fetch prekeys and agent info via
   Promise.all() instead of sequential awaits.
   Saves: 1-2s per session establishment.

Total: ~8-24s faster initial mesh communication.
All E2E encryption unchanged — Signal Protocol handshake intact.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Azure Responses API includes 'error': null in every successful response.
The error check used resp.get('error').is_some() which returned true for
null values, causing the function to return the raw Responses API format
instead of converting to chat/completions format. The OpenAI SDK then
failed to parse it: 'Expected property name at position 1'.

Fixed by checking for non-null error values only.
Added test covering the null error case.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reasoning models (gpt-5.4-pro) take 30-50s via Responses API. The
previous buffered approach held the HTTP connection with no data,
causing the OpenClaw client to timeout ('surface_error reason=timeout').

Now sends SSE comment keepalives (': keepalive') every 5 seconds while
the Responses API call is in progress. SSE comments are ignored by
parsers but keep the HTTP connection alive. The actual data is sent
as a single SSE frame when the response arrives.

Uses tokio::select! with a spawned task and mpsc channel to stream
the keepalive comments and final data to the client.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…+ image_gen

The native delegateToNativeAgent path failed because:
1. OpenClaw gateway requires device pairing (blocks headless exec)
2. OpenClaw exec allowlist double-gates commands already checked by AGT sidecar

Changes:
- Use processTaskWithTools directly for AGT mesh task_request messages
- Add mesh_send tool to sub-agent toolset (agent-to-agent E2E encrypted comms)
- Add foundry_image_generation tool to sub-agent toolset
- Fix system prompt to match actual available tools (remove phantom tools)
- Save generated images to temp file instead of discarding base64 data

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The sed -i on .bashrc failed with exit code 2 when the file didn't exist
at runtime (emptyDir mount or permission issues). With set -e in the
script, this killed the entrypoint and caused CrashLoopBackOff for all
sandbox pods.

Fix: touch + || true guard before sed operation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The plugin was in plugins.allow but not in plugins.entries, so OpenClaw
permitted it but never activated it. In dev mode the gateway auto-discovers
extensions from the directory, but AKS requires explicit entries with
{"enabled": true}. Initialize PLUGINS_ENTRIES with the azureclaw entry.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Two bugs in up.ts caused the Foundry project endpoint (services.ai.azure.com)
to be lost when running from cached context:

1. Cache prefill (line 255-258) set options.foundryEndpoint to the OpenAI
   endpoint (from cachedCtx.foundryEndpoint), overwriting the project URL
2. saveContext stored options.foundryEndpoint (already overwritten) as
   foundryProjectEndpoint instead of the original foundryEndpoint variable

Fix: cache prefill now restores OpenAI endpoint to options.openaiEndpoint
and project endpoint to options.foundryEndpoint (separate variables).
saveContext uses the local foundryEndpoint variable which holds the
original project URL.

Also fixed cached context file with correct project endpoint URL.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add ui.assistant.name/avatar in openclaw.json config template
  (proper config-driven approach, no upstream file patching)
- Add Identity section to MEMORY.md with strong AzureClaw persona
- Update sub-agent system prompt with AzureClaw branding
- Remove phantom tools from MEMORY.md (foundry_conversations,
  foundry_evaluations, foundry_deployments, foundry_agents)
- Add azureclaw_discover to tools list
- Derive clean SANDBOX_NAME from pod hostname (strip hash suffixes)
- AZURECLAW_DISPLAY_NAME env var for custom branding (default: AzureClaw)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add ~/.azureclaw/secrets.json (mode 600) to store all secrets:
- Azure OpenAI API key (migrated from legacy credentials file)
- Channel tokens (Telegram, Slack, Discord)
- Plugin API keys (Brave, Tavily, Exa, Firecrawl, Perplexity, OpenAI)

New CLI commands:
- azureclaw credentials set <key> <value>  — save a secret
- azureclaw credentials list               — show stored secrets (masked)
- azureclaw credentials remove <key>       — remove a secret

Resolution priority: CLI flag > secrets.json > host env var

dev.ts now auto-loads channel tokens from secrets.json,
so 'azureclaw dev' works without --telegram-token flags
once secrets are configured via 'azureclaw credentials set'.

10 new tests for secrets CRUD, migration, and resolution.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Channel tokens and plugin API keys for 'azureclaw add' now
resolve via: CLI flag > secrets.json > host env var.

This means 'azureclaw add my-agent --channels telegram' works
without --telegram-token if the token is already stored via
'azureclaw credentials set telegram-token <token>'.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw credentials (no subcommand) is now a guided interactive
wizard with masked input:
- Menu: Azure OpenAI, Telegram, Slack, Discord, Search APIs, Other
- Each prompts with masked password input (•••• on screen)
- Channel tokens support dot-suffixed variants for multiple bots
  (e.g. telegram-token.cloud, telegram-token.dev)
- Shows existing stored secrets at the top
- Looping menu — configure multiple categories in one session

azureclaw credentials set <key> [value]:
- Value argument is now optional
- If omitted, prompts interactively with masked input
- Validates dot-suffixed keys against base key

Operator TUI spawn dialog:
- Pre-fills channel token from secrets.json on startup
- Shows token field for ALL channels (not just Telegram)
- ←→ cycles through stored token variants when multiple exist
- Auto-fills default token when switching channels
- Passes correct --<channel>-token flag for any channel type

listSecretVariants() finds base + dot-suffixed keys for picker UIs.
2 new tests for variant listing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agent tool gaps fixed:
- discover: query AGT registry for other agents (names, trust, status)
- mesh_inbox: check for incoming messages from sibling agents
Sub-agents now have 10 tools: exec_command, http_fetch,
foundry_{web_search,code_execute,file_search,memory,image_generation},
mesh_send, mesh_inbox, discover

Credentials wizard (azureclaw credentials):
- Interactive guided flow with masked password input
- Category menu: Azure OpenAI, Telegram, Slack, Discord, Search, Other
- Dot-suffixed multi-token support (telegram-token.cloud, .dev)
- Shows existing secrets, loops for multi-category setup

Operator TUI:
- Pre-fills channel tokens from secrets.json
- Token field for all channels (not just Telegram)
- ←→ cycles stored token variants
- Auto-fills default token on channel switch

Config:
- listSecretVariants() for multi-token picker UIs
- 2 new tests for variant listing

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reverts the processTaskWithTools-only approach from e157a0f.
Sub-agents now use delegateToNativeAgent first (full OpenClaw toolset:
exec, file editing, git, browser, foundry_*, azureclaw_*, etc.)
with processTaskWithTools as fallback if native agent fails.

The original switch was made because openclaw agent required device
pairing — but the gateway already has dangerouslyDisableDeviceAuth:true,
and the actual failure was a missing --session-id flag. The native
agent now works correctly on both Docker and AKS.

processTaskWithTools still serves as a reliable fallback with 10 tools
(exec_command, http_fetch, foundry_*, mesh_send/inbox, discover).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- task:execute AGT policy check gates entire task before delegateToNativeAgent
- Container sandbox (seccomp, netpol, cgroups) is enforcement boundary
- No mixing with OpenClaw tools.deny — AGT is the single policy source
- Plugin tools inside native agent still check AGT per-call
- Falls back to processTaskWithTools if native agent fails

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
max_tokens is rejected by newer models (GPT-5.x). max_completion_tokens
works universally across all models (GPT-4.1+, GPT-5.x).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Fix TypeScript type for credential prompt map (TS7053)
- Define cmd variable from tool args in exec_command handler (TS2304)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When no base key exists (e.g. telegram-token) but variants do
(telegram-token.dev, telegram-token.cloud), pick the first variant.
Fixes credential resolution for 'azureclaw dev --channels telegram'.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Supports dot-suffix on channel names in --channels flag:
  azureclaw dev --channels telegram.cloud
  azureclaw dev --channels telegram.dev,slack

Resolves telegram-token.cloud from secrets.json when variant
is specified, falls back to base key otherwise.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth and others added 4 commits April 1, 2026 12:03
- _routerCall now accepts configurable timeout (default 15s)
- Image generation uses 90s timeout (was 15s, causing timeouts)
- Upgrade Go builder to 1.25 (patches 38 Go stdlib CVEs)
- Run npm audit fix after OpenClaw install (patches tar, minimatch, glob)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Set tools.exec.security=full in openclaw.json — AGT is the sole
policy authority. Removes flaky post-startup 'openclaw approvals set'
hack that could hang or fail silently.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Uses OpenClaw's registerImageGenerationProvider SDK to register
'azure-foundry' provider. Routes through inference router (managed
identity on AKS, API key on dev). OpenClaw's native image_generate
tool renders images inline in the chat UI automatically.

Config: imageGenerationModel = azure-foundry/gpt-image-1

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add /v1/images/generations route to inference router (OpenAI-compat)
- Route extracts model from body, delegates to deployment-based handler
- Configure OpenAI provider in openclaw.json pointing to inference router
- Set imageGenerationModel = openai/gpt-image-1 in agent defaults
- OpenClaw's built-in image_generate now renders images inline in chat
- Works on both AKS (managed identity) and dev (API key)
- Remove custom registerImageGenerationProvider (not needed)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread sidecar-images/agt-governance/server.py Fixed
Comment thread sidecar-images/agt-governance/server.py Fixed
Comment thread cli/src/commands/credentials.ts Fixed
Comment thread cli/src/commands/credentials.ts Fixed
Comment thread cli/src/commands/credentials.ts Fixed
Comment thread cli/src/config.test.ts Fixed
Pal Lakatos-Toth and others added 2 commits April 1, 2026 14:23
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…port

- Operator spawn form shows variant labels (cloud/dev) instead of masked values
- Added Allow From field with ←→ cycling and auto-correlation to token variant
- Added --telegram-allow-from flag to 'add' command (parity with 'dev')
- credentials list shows parent label for dot-suffixed variant keys
- telegram-allow-from now supports per-variant suffixes (.cloud, .dev)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The sandbox Docker build (36 stages) exhausts ubuntu-latest (2 vCPU, 7GB).
Switch to ubuntu-latest-xl (4 vCPU, 16GB) to prevent runner OOM/disconnect.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@pallakatos
Pal Lakatos-Toth (pallakatos) merged commit a1473c6 into main Apr 1, 2026
11 of 12 checks passed
@pallakatos
Pal Lakatos-Toth (pallakatos) deleted the security/hardening-phase1 branch April 27, 2026 12:59
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 10, 2026
…ocally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 11, 2026
…#245)

* feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap

Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): pre-build mesh-plugin for runtime CI + format reconciler

PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): quote workflow step name containing colon

YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): clippy fixes for Rust 1.95.0

CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12

Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5

The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): mark gaps G1 and G2 fixed on local AGT branch

Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.

* dev: add --mesh-provider <vendored|agt> selection with first-run prompt

Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites

Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): wire AGT provider end-to-end (6 stackup bugs)

End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): always route mesh registry through inference-router

azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): break mesh_send infinite poll loop on dead sub-agent

probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): suppress /v1/registry/* 404 leaks in AGT mode

Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode

Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s

The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* runtime: hide Foundry tools in github-copilot mode (same as github-models)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch

Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(up): --mesh-provider=agt picks AGT manifest + flips helm value

Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): plumb --mesh-provider into local-k8s helm install

Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): AGT wire protocol adapter for mesh_peer

Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer

- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(deploy): AGT mesh probe path + Cilium pod-port NP allow

deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh promote): AGT-compat health + WS upgrade paths

azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): first-run picker for local vs remote mesh source

azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container

On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh

Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap

mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call

Previous commit aa7d28e wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* push: auto-detect mesh provider from live helm release

When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* entrypoint: fail-open trust gate when running anonymous tier

When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop

Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable

The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log

CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
Security hardening phase 1 + Foundry integration improvements
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
…#245)

* feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap

Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): pre-build mesh-plugin for runtime CI + format reconciler

PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): quote workflow step name containing colon

YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): clippy fixes for Rust 1.95.0

CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12

Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5

The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): mark gaps G1 and G2 fixed on local AGT branch

Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.

* dev: add --mesh-provider <vendored|agt> selection with first-run prompt

Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites

Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): wire AGT provider end-to-end (6 stackup bugs)

End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): always route mesh registry through inference-router

azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): break mesh_send infinite poll loop on dead sub-agent

probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): suppress /v1/registry/* 404 leaks in AGT mode

Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode

Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s

The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* runtime: hide Foundry tools in github-copilot mode (same as github-models)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch

Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(up): --mesh-provider=agt picks AGT manifest + flips helm value

Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): plumb --mesh-provider into local-k8s helm install

Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): AGT wire protocol adapter for mesh_peer

Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer

- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(deploy): AGT mesh probe path + Cilium pod-port NP allow

deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh promote): AGT-compat health + WS upgrade paths

azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): first-run picker for local vs remote mesh source

azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container

On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh

Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap

mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call

Previous commit 34662c7 wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* push: auto-detect mesh provider from live helm release

When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* entrypoint: fail-open trust gate when running anonymous tier

When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop

Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable

The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log

CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant