Skip to content

feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap - #245

Merged
Pal Lakatos-Toth (pallakatos) merged 38 commits into
devfrom
agt-runtime-swap-phase2
May 11, 2026
Merged

Pal Lakatos-Toth (pallakatos) merged 38 commits into
devfrom
agt-runtime-swap-phase2

Conversation

@pallakatos

Copy link
Copy Markdown
Collaborator

Wire the runtime through createMeshTransport() so we can flip between the vendored @agentmesh/sdk (default) and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without touching any caller code.

Builds on PR #244 (the factory scaffold) by closing the API gap between the two SDKs and refactoring the runtime to use the factory.

What changed

IMeshTransport now requires (both adapters implement)

  • lookup(amid) — registry RPC
  • submitReputation(toAmid, sessionId, score, tags) — registry RPC
  • enableKnockEnforcement() — vendored toggle / no-op on AGT
  • onError(kind, fromAmid, detail) — diagnostic
  • onE2EVerified(peerAmid, isFirstPeer) — diagnostic
  • onDisconnect(reason, code) — diagnostic

Vendored adapter (mesh-plugin/src/connection.ts)

  • Delegates to underlying AgentMeshClient
  • Lazy-binds hooks registered before connect()

AGT adapter (mesh-plugin/src/agt-transport.ts)

  • lookup / submitReputation → REST against the registry directly. AGT's MeshClient is pure transport; registry RPCs belong on a separate client (out of scope for AGT upstream).
  • enableKnockEnforcement → no-op (AGT MeshClient always enforces)
  • Event hooks → delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods

AGT upstream changes (LOCAL branch only — NOT pushed)

Branch azureclaw-meshclient-event-hooks on /Users/pallakatos/Private/Repos/agt/agent-governance-toolkit adds the 3 event hooks + 8 tests. The AGT team owns the upstream PR — we hold this local until they're ready.

Until that ships in a published @microsoft/agent-governance-sdk release, our adapter's optional-chain calls (client.onError?.(...)) make the hooks no-ops on the published 3.5.0. Provider stays functional, just without the diagnostic hooks.

Runtime (runtimes/openclaw/src/index.ts)

  • Adds @azureclaw/mesh as file:../../mesh-plugin
  • Replaces new sdk.AgentMeshClient(...) with await createMeshTransport({...}) when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value (typo, empty, unset)
  • Identity is generated once via vendored SDK regardless of provider, raw Ed25519 keys extracted via toData(), then shared with both → same AMID either way
  • Banner now reports active provider

Tests

Suite Result
mesh-plugin (incl. new 16-test compat suite) 97/97 ✅
runtimes/openclaw 118/118 ✅
AGT local branch (8 new event-hook tests) 387/387 ✅

Docs

docs/agt-vs-vendored-sdk.md — exhaustive side-by-side analysis covering every functional surface (identity, policy, trust, audit, mesh-client, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer) + wiring + 4-phase migration plan.

Migration path

  • Phase 3 (next): flip default to agt in sandbox image; soak in dev with cross-provider parent↔child interop
  • Phase 4 (cleanup): once AGT upstream merges + publishes the event hooks, drop vendor/agentmesh-sdk/ entirely (5 protocol patches no longer needed) and remove the env-var toggle

Reviewer notes

  • The factory respects an unknown / typo / empty env var by falling back to vendored — pods can never accidentally boot on AGT
  • Compat test fails CI if either adapter regresses on any of the 6 methods
  • AGT side stays optional (optionalDependencies) so pods built without AGT installed still boot on default

cc @Azure/azureclaw-maintainers

Pal Lakatos-Toth and others added 2 commits May 10, 2026 01:36
Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread .github/workflows/ci.yml Fixed
Pal Lakatos-Toth and others added 14 commits May 10, 2026 07:17
YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ocally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.
Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…solves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…gistry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread mesh-plugin/src/agt-transport.ts Fixed
Pal Lakatos-Toth and others added 12 commits May 11, 2026 11:28
Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…dels)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread cli/src/commands/dev.ts Dismissed
Pal Lakatos-Toth and others added 4 commits May 11, 2026 14:23
On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previous commit aa7d28e wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread mesh-plugin/src/agt-transport.ts Fixed
Pal Lakatos-Toth and others added 5 commits May 11, 2026 15:44
When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Comment thread mesh-plugin/src/agt-transport.ts Fixed
CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Imran Siddique (imran-siddique) pushed a commit to microsoft/agent-governance-toolkit that referenced this pull request May 11, 2026
…ify, heartbeat (#2090)

* feat(mesh-client): add onError, onDisconnect, onE2EVerified event hooks

Adds three observer-registration methods to MeshClient to align with the
AzureClaw vendored AgentMesh SDK surface so consumers can swap providers
behind a single transport interface:

  onError(handler)        — fires on ws errors and decrypt failures
  onDisconnect(handler)   — fires on ws.close with reason+code
  onE2EVerified(handler)  — fires on first successful decrypt per peer

Pure additions; no behaviour change for existing flows. Decrypt-failure
and missing-session paths now drop the message AND notify error
observers instead of silently dropping (still safe — unhandled events
are swallowed).

Tests: 8 new in mesh-client-event-hooks.test.ts; full suite 387/387 green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-client): close gaps G1 (KNOCK auto-bootstrap) and G2 (auto-reconnect)

Two protocol-level gaps were identified during the AzureClaw vendored-SDK
audit (see Azure/kars#245 docs/agt-vs-vendored-sdk.md). Both block
moving fully off the vendored agentmesh-sdk fork onto upstream AGT.

## G1: receiver-side X3DH auto-bootstrap from KNOCK

Before: the responder side of an encrypted session was created only by
calling acceptSession(peerId, establishment) explicitly. But the
ChannelEstablishment was never carried on the wire — neither in
knock_accept nor in the first message frame — so the receiver had no way
to obtain it. Result: any fresh encrypted session failed with
"No encrypted session" the first time a ciphertext arrived.

Mirrors vendored agentmesh-sdk patch #4b. Behavior:

- Sender (establishSession): builds the SecureChannel first, then sends
  KNOCK with the ChannelEstablishment embedded as
  { ik: base64, ek: base64, otk?: number }.
- Receiver (handleKnock): if the knock contains `establishment` AND no
  prior session exists, deserialize and call acceptSession() automatically
  before responding with knock_accept. On bad establishment data, the
  knock is rejected and onError("knock", ...) fires.
- Backwards-compatible: legacy peers that don't embed establishment
  continue to work; caller must invoke acceptSession() manually as before.

## G2: auto-reconnect on transport drop

Before: MeshClient.reconnect() existed but was never called automatically.
Network blips, relay restarts, AKS node OOM-evicts left the client
disconnected forever — agents went mesh-deaf.

Mirrors vendored agentmesh-sdk patch #9. Behavior:

- New options: autoReconnect (default true), maxReconnectAttempts
  (default Number.POSITIVE_INFINITY), reconnectBaseDelayMs (default
  1000), reconnectMaxDelayMs (default 60000).
- On non-1000 ws.onclose: schedule reconnect with exponential backoff
  capped at reconnectMaxDelayMs. Light jitter (±20%) avoids
  thundering-herd reconnects across many sandboxes.
- onDisconnect handlers fire BEFORE the reconnect is scheduled so
  observers can see the drop.
- After maxReconnectAttempts, fires onError("ws", ..., "auto-reconnect
  gave up after N attempts") and stops.
- disconnect() cancels any pending reconnect timer.
- A connect() failure inside the reconnect path schedules another retry
  via the existing onclose path (and recursively from the catch handler
  in scheduleReconnect for the case where ws.onopen never fired).

## Tests

11 new tests across two files:
- tests/mesh-client-knock-bootstrap.test.ts: 5 tests covering sender-side
  embedding, receiver auto-bootstrap, legacy fallback, malformed
  establishment, and end-to-end happy-path encrypted send.
- tests/mesh-client-auto-reconnect.test.ts: 6 tests covering server-close
  reconnect, client-close no-reconnect, opt-out, give-up after max
  attempts, disconnect cancels pending timer, no duplicate scheduling.

Full AGT TS suite: 398/398 pass (was 387 before this commit). Build clean.

Both gaps were identified in the AzureClaw audit doc:
docs/agt-vs-vendored-sdk.md (Patch-by-patch audit section).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): close vendored-parity gaps G3, G4, G5

G3 — session-desync teardown (vendored agentmesh-sdk patch #13 equiv):
when Double Ratchet decryption fails for an existing session, the local
ratchet state is irrecoverable. Previously we only fired
onError('decrypt'); now we tear down the session (delete from
this.sessions; clear knockAccepted) and fire a distinct
onError('session_desync') so callers can re-run establishSession() and
resume communication.

G4 — pre-KNOCK encrypted-message buffer (vendored agentmesh-sdk
patch #16 equiv): under transport reordering the relay can deliver an
encrypted 'message' frame BEFORE the matching 'knock' arrives. Without
buffering, that first frame was dropped silently. Now MeshClient buffers
encrypted frames per-peer (default cap 5, TTL 3000ms), replays them
when the corresponding knock is accepted, and drops them when rejected
or on disconnect. Disabled by setting preKnockBufferSize: 0.

G5 — eager ghost-connection close on rebind (vendored agentmesh-relay
patch #2 equiv): when an agent reconnects with the same DID, the prior
'ghost' WebSocket is now closed eagerly with code 1000 'session_replaced'
instead of waiting for the 90s heartbeat-eviction timer. The 'finally'
cleanup now compares socket identity to avoid removing the freshly
rebound connection when the old socket's handler unwinds.

Tests: +5 (G4 pre-knock buffer behaviour), +2 (G3 type + lifecycle
safety), +1 (G5 ghost rebind). 405 TS tests pass; 18 relay tests pass.

NOTE: held local-only on branch azureclaw-meshclient-event-hooks; will
be coordinated upstream with the AGT team.

* feat(mesh): add RegistryClient and auto-register on MeshClient.connect()

Closes the structural gap where MeshClientOptions declared registryUrl
but no code path ever used it. Without this, agents can connect to the
relay and exchange ciphertext but are invisible to peers — they never
appear in /v1/discover and no one can fetch their X3DH pre-key bundle.

This commit adds the missing registry-side glue, in upstream-quality
shape, so AGT MeshClient becomes a drop-in replacement for vendored
forks that wired their own registry calls (e.g. AzureClaw vendor/agentmesh-sdk).

What's new:

- src/encryption/registry-client.ts: typed HTTP client wrapping the
  AgentMesh registry's REST surface (POST /v1/agents,
  PUT/GET /v1/agents/{did}/prekeys, GET /v1/discover, GET/DELETE
  /v1/agents/{did}). base64url <-> Uint8Array marshalling, optional
  Ed25519-Timestamp Authorization signer, configurable retry/timeout.

- src/encryption/mesh-client.ts:
  * New options: registryClient, registryClientOptions, capabilities,
    registrationMetadata, oneTimePrekeyCount, autoRegister.
  * connect() now calls registerSelf() at the end of a successful WS
    handshake when autoRegister is true (default) and a registry is
    configured. Idempotent — 409 (already registered) is treated as
    success; reconnect doesn't re-register.
  * registerSelf(): generates signed-prekey + N one-time prekeys via
    keyManager, then POSTs /v1/agents (capabilities = [displayName,
    ...options.capabilities]) and PUTs the prekey bundle.
  * discover(capability): wraps registry.discover.
  * establishSessionWithPeer(peerId): fetches peer prekeys from the
    registry and calls existing establishSession.
  * getRegistry(): exposes the underlying client for advanced cases.

- tests/registry-client.test.ts: 11 new tests covering wire format,
  base64url round-trip, idempotent register (409), 5xx retry, and
  Authorization header. Also covers MeshClient auto-register on connect
  end-to-end with a fake fetch + fake WebSocket.

- tests/mesh-client-*.test.ts: existing fixtures updated to pass
  autoRegister: false (no registry available in those tests). Behavior
  unchanged for the production path.

All 71 tests pass (60 previously + 11 new).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(registry): add POST /v1/agents/{did}/heartbeat to bump last_seen

The registry's last_seen field is set on registration but never
updated by any HTTP handler — update_last_seen() exists in
store.py but is dead code in production. As a result, every
registered agent looks 'offline' (online=false on /presence) 90s
after spawn, and any client-side discover stale-filter rejects
all live agents.

Add an unauthenticated POST /v1/agents/{did}/heartbeat that calls
store.update_last_seen(). Idempotent. Returns 404 for unknown DIDs
so callers can detect a registry restart and re-register.

Mirrors the agent-side ping cadence (30s, well within the 90s
presence threshold).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): include Ed25519 identity_key_ed in prekey bundle so verifyBundle() works

PreKeyBundle.identityKeyEd was previously aliased to the X25519 identity_key
because the registry didn't store the Ed25519 signing key. As a result,
verifyBundle() either passed only because the stub test never exercised it,
or quietly accepted any signed pre-key — a soundness gap in the X3DH wrapper.

This commit threads identityKeyEd through the full path:

- agent-mesh/registry/store.py (AgentRecord): add identity_key_ed field
  (X25519 long-term key already present as identity_key; Ed25519 distinct).
- agent-governance-typescript/src/encryption/x3dh.ts: expose
  X3DHKeyManager.identityKeyEd getter.
- agent-governance-typescript/src/encryption/registry-client.ts:
  uploadPrekeys() now accepts identityKeyEd (32 bytes, validated) and
  serializes it as identity_key_ed; fetchPrekeys() deserializes it back into
  PreKeyBundle.identityKeyEd. Older bundles without the field fall back to
  the X25519 key — verifyBundle() will then correctly reject (forensic
  visibility, no silent bypass).
- agent-governance-typescript/src/encryption/mesh-client.ts: registerSelf()
  passes keyManager.identityKeyEd through to the registry.
- tests/registry-client.test.ts + tests/test_registry.py: assert the new
  field is round-tripped on PUT/GET.

Wire compatibility: PUT now requires identity_key_ed; GET emits it when set.
Existing AGT clients that don't upload it will get verifyBundle() rejection
on the receiving side — peers must upgrade together. (No silent regression
on the wire — the field is new, not a breaking change to existing fields.)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test(mesh): set autoRegister:false in upstream knock + malformed-frame tests

These two upstream tests construct MeshClient without autoRegister:false,
which now triggers a real fetch to the registry on connect() (added by
our registry-client commit). The tests have no fake registry, so they
fail with 'fetch failed'. autoRegister:false skips the auto-registration
path and matches the pattern used in our 6 sibling mesh-client-*.test.ts
files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@pallakatos
Pal Lakatos-Toth (pallakatos) merged commit 27ed755 into dev May 11, 2026
21 checks passed
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
…#245)

* feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap

Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.

Surface additions to IMeshTransport (both adapters now expose):
  - lookup(amid)                  — registry RPC for reputation/display name
  - submitReputation(...)         — registry RPC for peer feedback
  - enableKnockEnforcement()      — vendored toggle (no-op on AGT, always-on)
  - onError(kind, from, detail)   — diagnostic hook for decrypt + ws errors
  - onE2EVerified(peer, isFirst)  — first-decrypt-per-peer signal
  - onDisconnect(reason, code)    — ws close / error fan-out

mesh-plugin (vendored A adapter):
  - connection.ts delegates to the underlying SDK; lazy bind for hooks
    registered before connect()
  - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
    contract so neither adapter can drop a method without CI failing

mesh-plugin (AGT B adapter):
  - agt-transport.ts implements lookup/submitReputation as REST calls to the
    registry (AGT MeshClient is pure transport — registry RPCs intentionally
    not added to AGT upstream; they belong on a separate RegistryClient)
  - enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
  - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
    methods (added on local AGT branch azureclaw-meshclient-event-hooks,
    NOT pushed — AGT team owns the upstream PR)

runtime (runtimes/openclaw):
  - Adds @azureclaw/mesh as a file: dependency
  - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
    when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
  - Identity is generated once via vendored SDK regardless of provider, then
    raw Ed25519 keys are extracted via toData() and shared across both — same
    AMID either way
  - Banner now reports active provider (vendored vs agt)

Docs:
  - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
    policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
    plaintext peers, file transfer + the wiring + migration path
  - Documents the 3 hooks added to local AGT branch and the 3 governance
    methods kept adapter-side

Tests:
  - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
  - runtimes/openclaw: 118/118 pass
  - AGT (local branch): 387/387 pass with 8 new event-hook tests

Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): pre-build mesh-plugin for runtime CI + format reconciler

PR #245 CI failures:
1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' —
   the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the
   CI workflow only ran 'npm install' inside runtimes/openclaw, which does
   not build the file: dep's dist/. Add an explicit pre-build step that
   installs vendored agentmesh-sdk + mesh-plugin and runs its build before
   the runtime install.
2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift
   inherited from PR #244. Run cargo fmt --all.

Also added a 'prepare' script to mesh-plugin/package.json so any future
file: consumer auto-builds on install (defensive — the explicit CI step
above is still the primary fix).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): quote workflow step name containing colon

YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because
'file:' was interpreted as a mapping key. Quote the string.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): clippy fixes for Rust 1.95.0

CI runs Rust 1.95.0 which added new clippy lints:

- doc_lazy_continuation: indent doc list items that span multiple lines.
  Added two-space indent to the trailing 'All three are populated...'
  paragraph so it is treated as a continuation of the preceding list
  item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
  if .. { a } else { b } per the lint suggestion.

These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12

Audit findings (docs/agt-vs-vendored-sdk.md):
- Verified each of the 9 vendored SDK patches against AGT MeshClient
- Verified all 4 vendored relay + 4 vendored registry patches
- Identified 5 protocol-level gaps that block Phase 3:
  * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg)
  * G2: no auto-reconnect loop (manual reconnect() only)
  * G3: registry RPCs not in MeshClient (compensated in adapter)
  * G4: fast-fail handshake edge (defensive)
  * G5: connect frame incompatibility with vendored relay (BLOCKING)
- Documented which gaps require AGT-upstream changes vs adapter fixes
- Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5

Adapter-side fixes (mesh-plugin/src/agt-transport.ts):
- Patch #7 port: submitReputation now logs status + body on non-2xx
  and logs network errors (vendored swallowed both silently)
- Patch #12 port: registry fetches now use bounded retry with
  exponential backoff (250ms, 750ms, 2000ms) — applied to lookup,
  submitReputation, and discovery search

Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing
unreachable-registry tests now also exercise retry path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5

The previous audit framed gaps as 'AGT vs vendored relay' which is the
wrong question — when we move fully upstream, AGT will use its own Python
relay and registry, not ours. So wire-format compat with the vendored
relay (the old G5) is irrelevant by design.

Re-audit against the full AGT upstream stack (TS SDK + Python relay +
Python registry):

- Confirmed AGT registry already does Ed25519-over-raw-timestamp
  signature verification (registry/app.py:54-98) — same approach we
  patched into the vendored registry. No port needed.
- Confirmed AGT relay has /health, heartbeat, and 90s offline threshold.
- Confirmed AGT registry tracks last_seen with 90s online window.
- AGT relay overwrites duplicate connections without explicit close —
  slower than our 4001 SessionReplaced but functionally similar.
- AGT uses shared-secret token auth on relay (no per-frame sig) — different
  security model than our vendored relay; flagged for review but not a
  functional regression.

Real gaps that block moving upstream remain only 2:
- G1: receiver-side X3DH bootstrap (acceptSession() exists but
  ChannelEstablishment is never serialized onto the wire)
- G2: no auto-reconnect loop in MeshClient (manual reconnect() only)

Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks
on the local AGT branch are a prerequisite for cleanly implementing G2.

Migration strategy updated to reflect that A↔B cross-provider message
interop is not a goal (different relays by design); the swap unit is the
sandbox, not the message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): mark gaps G1 and G2 fixed on local AGT branch

Audit doc updated to reflect that both protocol gaps identified during
the vendored-vs-AGT audit are now closed on the local AGT branch
`azureclaw-meshclient-event-hooks` (commit `d75ea37b`):

- G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on
  the wire; `handleKnock()` auto-calls `acceptSession()` on receipt.
  Backwards-compatible with legacy peers.
- G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter)
  on non-1000 close; `autoReconnect: true` by default; opt-out via
  options.

AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across
`mesh-client-knock-bootstrap.test.ts` and
`mesh-client-auto-reconnect.test.ts`).

The AGT branch is held locally — NOT pushed — pending coordination
with the AGT team for an upstream PR. From AzureClaw's perspective,
the upstream-AGT scenario is now feature-complete: every vendored
patch has either been merged upstream, has an equivalent in AGT, lives
in our adapter, or is fixed on the local AGT branch.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally

Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited
patches: SDK #10 (idempotent initiateSession), #11 (wsFactory +
plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different
KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow),
#18 (simpler closeSession-based recovery).

Documents three additional real gaps now fixed locally on the AGT
branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2):

- G3 (vendored SDK #13): MeshClient now tears down the session on
  decrypt failure and fires onError('session_desync', ...) so the
  caller can re-run establishSession() to recover. Without G3, a
  single ratchet drift permanently jams the channel.

- G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer
  (default cap 5, TTL 3000ms) when no session exists yet, drains on
  knock_accept, drops on knock_reject. Without G4, relay frame
  reorder silently loses the first message of every fresh handshake.

- G5 (vendored relay #2): AGT relay now closes the previous WebSocket
  with code 1000 'session_replaced' before overwriting the
  _connections entry on rebind. Without G5, the old socket lingers
  for up to 90 seconds and messages route to a dead connection. The
  finally cleanup now compares socket identity to avoid removing the
  fresh connection on the old handler's unwind.

Also updates the chunked file-transfer reliability note: G3 + G4 are
both required for robust mesh_file_transfer because chunked transfers
amplify silent-drop and ratchet-drift bugs into stuck transfers with
no error surface.

Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different-
but-equivalent, 5 real gaps all fixed locally on AGT branch
(NOT pushed; awaiting upstream PR coordination with the AGT team).
AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.

* dev: add --mesh-provider <vendored|agt> selection with first-run prompt

Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:

CLI (cli/src/commands/dev.ts):
  - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
  - First-run interactive prompt offers AGT only if the toolkit
    checkout is actually present locally — silently defaults to
    vendored otherwise (no pestering for users without AGT cloned).
  - --build branch: builds the right relay/registry images
      vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
      agt      → agent-governance-python/agent-mesh/docker/Dockerfile
                 with COMPONENT=relay / registry build-args
  - Sandbox image build: stages locally-packed AGT SDK tarball into
    .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
    build-arg (auto-discovers if --agt-sdk-tarball not given).
  - Runtime branch: skips Postgres for AGT (in-memory registry),
    uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
    8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
  - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
    runtime transport-factory honors the user's choice.

Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
  - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
    sandbox npm-installs the local tarball instead of fetching the
    published @microsoft/agent-governance-sdk from npm. Lets us
    smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
    end without round-tripping through npm publish.
  - .agt-sdk/ staging dir always exists (with .keep) so the COPY
    never fails when the user didn't stage a tarball.

Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves

The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.

Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API

Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).

- connect() now passes autoRegister: true so the SDK uploads identity and
  prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
  (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
  clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
  in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
  metadata.display_name (set by registerSelf), with the first capability as
  the fallback.

Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites

Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.

getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).

Migrated all raw-HTTP registry callsites:
  - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
    resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
  - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
    spawn discovery, AKS→local discovery.
  - core/agt-task-loop.ts (1 site): registry_capability_search tool.
  - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
    azureclaw_discover (mesh_discover) tool.
  - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
    heartbeat keepalive (no-op under AGT — relay does liveness via WS).

The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): wire AGT provider end-to-end (6 stackup bugs)

End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:

1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
   file:../../mesh-plugin symlink dangled in node_modules. Plugin
   swap silently fell back to vendored with 'Cannot find package
   @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
   into /mesh-plugin/ in the final stage of the Dockerfile.

2. entrypoint.sh used cp -r when copying node_modules into the
   plugin extension dir, preserving the (now-broken-at-runtime-path)
   symlink. Switched to cp -rL so symlinks dereference into real
   files in the target tree.

3. mesh-plugin/src/index.ts imported createMeshTransport from
   ./transport-factory.js but never re-exported it. Runtime swap
   path couldn't find the factory. Added the missing re-export.

4. inference-router agt_registry_proxy unconditionally prepended
   '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
   became '/v1/v1/agents' at the upstream. Now: forward verbatim
   when path starts with 'v1/' or equals 'health', else prepend.
   Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).

5. /agt/relay route only matched the bare path, but AGT MeshClient
   appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
   upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.

6. agt_registry_proxy route was declared get(...).post(...) only.
   AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
   upload and DELETE for deregister — both 405'd at the router.
   Added .put() and .delete() to the route declaration.

   Bug #6 was invisible to vendored because the vendored SDK only
   ever uses GET/POST (registry/register, registry/prekeys, etc.).
   AGT's switch to REST verbs exposed the gap.

Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.

Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
  - POST /v1/agents → 201 Created
  - PUT /v1/agents/{did}/prekeys → 200 OK
  - WebSocket /ws accepted, stable connection (no reconnect loop)
  - Plugin reports 'AGT mesh connected' + provider=agt

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): always route mesh registry through inference-router

azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.

In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.

Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): break mesh_send infinite poll loop on dead sub-agent

probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).

Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.

Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): suppress /v1/registry/* 404 leaks in AGT mode

Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:

1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
   The AGT registry exposes GET /v1/discover?capability=X instead — display
   names live in the per-agent record, so the AGT path fans out to a
   second /v1/agents/{did} fetch per discover hit. Driven by the operator
   panel's /agt/reputation polling.

2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
   AGT has no per-session counter; per-agent reputation already submitted
   via MeshClient.submitReputation. No-op in AGT mode.

3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
   on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
   for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.

Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
  mode gets vendored-equivalent send-with-first-contact semantics. Without
  this, send() throws 'No encrypted session — call establishSession() first'
  and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
  were hanging when the cluster API was unreachable.

cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode

Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.

Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.

cargo check: clean
agt_governance_integration: 26/26 pass

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s

The AGT Python relay (agentmesh/relay/app.py) marks any connection
stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then
routes subsequent messages for that DID to its OFFLINE STORE instead
of live delivery. Stored frames are only replayed on (re)connect via
_deliver_pending — so a long-lived parent that never reconnects loses
every reply that arrives more than 90s after it last connected.

The AGT MeshClient exposes sendHeartbeat() but never auto-schedules
it. Vendored mode worked despite the same gap because the vendored
Rust relay has no time-based stale check (only checks broken
channels). For AGT mode we run our own 30s ticker (matches relay's
HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear
it down in disconnect(). The ticker is .unref()'d so it doesn't keep
the Node event loop alive on its own.

Reproduces deterministically when a sub-agent's reply lands >90s
after the parent's connect timestamp:

  parent connect t=0
  parent sends t=t1 (<90s)            -> messages_routed += 1
  child sends reply t=t2 (>90s)       -> stored offline, never delivered
  relay /health: messages_delivered=0 (forever)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* runtime: hide Foundry tools in github-copilot mode (same as github-models)

The Foundry tool catalog only makes sense when there is a real Azure
Foundry project bound to the sandbox. Both GH-token providers
(github-models, github-copilot) talk to GitHub-hosted models directly
and have no Foundry project — exposing the 6 foundry_* tools just
burns context with verbose JSON-schema and tempts the model to call
endpoints the router will 404.

Three call-sites were checking the provider:

1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`,
   now matches either GH-token provider. The DuckDuckGo-backed
   web_search + memory fallbacks are appended in both modes.

2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives
   the prompt's tool-block descriptions and the slim 'Mode note' so the
   sub-agent sees the same tool catalog the LLM was given. Mode-note
   string adjusted to identify which provider is active.

3. runtimes/openclaw/src/index.ts — parent-side foundry tool
   registration in github-copilot mode. Was registering the full
   Foundry catalog with no upstream to call.

Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved
per request) in github-copilot mode, and the 6 dead-end foundry_*
tools no longer appear as options.

Tests: runtimes/openclaw 118/118 pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest

Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.

When --mesh-provider=agt:
  * Builds relay+registry from the AGT upstream Dockerfile
    ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
    using COMPONENT=relay|registry build-args (matches dev.ts).
  * Tags as agentmesh-{relay,registry}-agt:latest so both vendored
    and AGT images can coexist on the same ACR and so the existing
    deploy/agentmesh-agt.yaml manifest picks them up unchanged.
  * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
    in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
    into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
  * Always passes MESH_PROVIDER build-arg to the sandbox image so the
    Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
    runs for AGT clusters.

When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).

Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).

The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch

Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.

Flow:
  1. Detect current provider via 'kubectl get deploy/postgres -n
     agentmesh' (vendored has Postgres, AGT does not).
  2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
  3. kubectl apply  -f deploy/agentmesh-<target>.yaml
  4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
  5. kubectl rollout restart deploy/azureclaw-controller
  6. With --restart-sandboxes: roll every azureclaw-managed Deployment
     so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.

Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(up): --mesh-provider=agt picks AGT manifest + flips helm value

Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.

When --mesh-provider=agt:
  * helm install runs with --set mesh.provider=agt (controller env
    AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
  * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
    deploy/agentmesh.yaml.
  * Skips the postgres ACR import and the agentmesh-db-credentials
    secret creation (both unused by AGT — its registry is in-memory).
  * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
    vs .tmp-agentmesh.yaml) so concurrent provider switches don't
    collide.

The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): plumb --mesh-provider into local-k8s helm install

Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s'
now forwards through runLocalK8s() → helmInstall() as
'--set mesh.provider=<value>', so the controller deployed into the
kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and
spawns sandboxes against the chosen mesh stack.

NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all
(the plan notes this as a Phase 3 pre-req blocked on AGT upstream
patches G1/G2/G5). This commit only handles the helm-value plumbing;
adding actual relay/registry deploy to local-k8s will land once the
AGT fixes are upstream so we can prove end-to-end mesh roundtrip on
local kind.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(controller): AGT wire protocol adapter for mesh_peer

Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.

Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.

New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
  with `#[serde(tag="type", rename_all="snake_case")]` matching
  `agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.

`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
  (default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
  AGT posts `{did, public_key (base64url), capabilities, metadata}`
  with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
  (matches JS SDK `buildDid`), so every leader replica converges on
  the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
  - AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
    (token read from `AGENTMESH_RELAY_TOKEN` if set).
  - AGT has no `Connected` ack — mark `connected=true` immediately.
  - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
    keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
  and branch outbound framing — AGT emits `message` frames
  `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
  `handle_agt_frame()`. AGT path:
  - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
  - Sends `Ack` reply (required — without it AGT redelivers on
    reconnect → duplicate offload processing).
  - Treats `Error` frames mentioning Authentication failed /
    Missing 'from' / session_replaced as fatal — drops connection
    for reconnect.

`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.

`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
  peer now starts unconditionally when enabled; provider is consumed
  inside `mesh_peer::run`.

Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer

- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs,
  inference-router/governance.rs).
- mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer`
  to FakeClient interface + mock — pre-existing test gap exposed
  by the post-606f5b0 send path that calls it before send().

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(deploy): AGT mesh probe path + Cilium pod-port NP allow

deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.

operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh promote): AGT-compat health + WS upgrade paths

azureclaw mesh promote ran post-promote health checks against
vendored-only paths and would 404 on AGT clusters:

- Registry probe hit /v1/health. AGT only exposes /health (vendored
  exposes both). Probe /health first, fall back to /v1/health for
  vendored compatibility with older deployments that may have only
  served the /v1/ alias.
- Relay WebSocket upgrade was attempted on /. AGT only serves WS on
  /ws (vendored uses /). Try /ws first, fall back to /.
- 'Test: curl' hint pointed at /v1/health — also updated to /health
  so the suggested command works on both providers.

Verified live against AGT cluster:
  Registry healthy (agentmesh-registry)
  Relay healthy (WebSocket upgrade on localhost:19991/ws)

640 CLI tests still pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(dev): first-run picker for local vs remote mesh source

azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:

  Where should the mesh live?
  ❯ Local   (recommended; spin up relay + registry in Docker)
    Remote  (auto port-forward to AKS cluster: <cluster-name>)

Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.

Remote (advanced) federates with a previously-provisioned AKS mesh:
  - If ~/.azureclaw/context.json has a cached globalRegistryUrl
    from a prior 'azureclaw mesh promote', reuse it verbatim.
  - Otherwise default to http://localhost:18080 — the port-forward
    URL 'mesh promote --port-forward' uses — so the auto-promote
    fallback in the downstream global-registry block will spawn
    the tunnels on demand.
  - If there is no aksCluster in context at all, warn and fall
    back to local so the user isn't left with a broken sandbox.

Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.

Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).

Verified:
  - npm run build / typecheck clean
  - 640 CLI tests pass (2 skipped, no regressions)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container

On AKS the inference-router runs as a separate sidecar with its own
env array, unlike local docker where it shares the openclaw container's
env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to
decide whether to upgrade the relay WS on `/` (vendored) or `/ws`
(AGT), and likewise for the registry discover endpoint. The controller
was only injecting the var into the openclaw container, so on AGT
clusters the router defaulted to vendored and got 403 Forbidden in a
tight reconnect loop against the AGT FastAPI relay.

Push the same normalized provider value into router_agt_env (which is
extended into router_env) so both containers agree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh

Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.

Two coordinated fixes:

1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
   (spawner seeds 'parent_name:parent_AMID' as the first entry), also
   push PARENT_SANDBOX=<first_name> into the openclaw container env.

2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
   and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
   PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
   lookup. The Symbol is set during runtime init from
   AGT_TRUSTED_PEERS[0], so this works even on images built before fix
   #1 lands. Skip in offload mode — 'parent' there is a protocol-level
   routing token, not a mesh recipient name.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap

mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:

    establishSessionWithPeer is not a function

It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.

Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call

Previous commit 34662c7 wrongly removed the establishSessionWithPeer()
pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the
upstream @agentmesh/sdk API surface. The mesh-plugin actually loads
@microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned
to ^3.5.0), which:

  • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 —
    a high-level helper that fetches the prekey bundle and runs
    X3DH+KNOCK, idempotent on cache-hit
  • does NOT auto-bootstrap in send(): the path at L341 explicitly
    throws 'No encrypted session with <peer>. Call establishSession()
    first.' when no SecureChannel exists yet

Symptom of the bad fix: parent → sub-agent mesh_send failed with
'No encrypted session with <amid>. Call establishSession() first.'
on every first contact post-rollout.

Restoring the pre-bootstrap with the correct rationale documented and
the SDK source citations. AgtSdkModule type keeps the method optional
for forward-compat with SDKs that auto-bootstrap; the runtime call
uses non-null assertion since AGT SDK 3.5.0 ships the method.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* push: auto-detect mesh provider from live helm release

When running 'azureclaw push --only sandbox --apply' without an explicit
--mesh-provider flag, the CLI silently defaulted to 'vendored'. On a
cluster already flipped to AGT (mesh.provider=agt), this caused the
sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/
— npm would install the public @microsoft/agent-governance-sdk@3.5.0
which lacks establishSessionWithPeer/discover/registerSelf helpers.
Result: parent throws 'this.client.establishSessionWithPeer is not a
function' on every mesh send.

Auto-detect by reading 'mesh.provider' from the live helm release and
respect it when --mesh-provider was not passed on the command line.
Explicit flag still wins.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* entrypoint: fail-open trust gate when running anonymous tier

When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when
the Entra token exchange exhausts its retries, every sandbox registers
as anonymous tier with registry reputation score 0. The KNOCK trust
gate compares (registry_score * 1000 + affinity_bonus) against
AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity:

  - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus
  - effectiveScore = 0 < 500 → KNOCK rejected
  - whole mesh appears 'blocked' even though discovery + X3DH succeed

Trust scoring is meaningless without OAuth identity. When we know we're
in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation
in onKnock still runs, and the SDK's X3DH still proves cryptographic
identity end-to-end — we just stop using a meaningless score as a gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop

Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough +
multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts
to add a `web_search` branch (DuckDuckGo for slim-mode) and a
`memory` branch, but in doing so deleted the
`} else if (fnName === "foundry_web_search" || foundry_code_execute
|| foundry_file_search) {` else-if opener and forgot to put it back
after the memory branch closes.

The result: the entire foundry_web_search / foundry_code_execute /
foundry_file_search dispatch block (lines 333-548) got silently
nested INSIDE the memory branch — only reachable when
`fnName === "memory"`, in which case none of its inner
`fnName === "foundry_*"` checks match. Dead code.

Symptom from this morning's demo: sub-agents calling
foundry_web_search fell through every else-if and hit the final
`echo 'no command'` exec fallback, returning the literal string
"no command" — which the model then dutifully reported as
"Foundry web search returned no command" in a loop.

Parent agent was unaffected because the parent's foundry tools go
through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427),
not the sub-agent dispatcher. That's why foundry_web_search "always
worked" for the user — the parent path is a totally different code
path.

Fix: add back the missing else-if opener between the memory branch
close and the existing foundry_* body. tsc clean. The dispatcher
chain is now:
  file_write → http_fetch → web_search → memory → foundry_web_search
  → foundry_download_file → foundry_memory → foundry_image_generation
  → mesh_send → mesh_transfer_file → discover → mesh_inbox
  → mesh_await → exec_command fallback

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable

The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.

Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.

Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).

The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening

Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime
emits strict-conformant tool schemas (additionalProperties:false, all-required,
nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when
slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via
a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL.

Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search,
foundry_code_execute, foundry_memory, foundry_file_search, mesh_send.

Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file,
mesh_inbox, mesh_await, discover, foundry_image_generation,
foundry_download_file, web_search, memory.

Skipped (free-form schema): http_fetch (variable headers object).

Plumbing:
- runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set,
  STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate.
- runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule
  in sub-agent prompt, parse-error hint pointing to
  foundry_code_execute → json.dump → mesh_transfer_file, boot observability log.
- runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument
  resilience (matches new prompt guidance).
- controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into
  openclaw container env when enabled on controller.
- deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default).
- deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env
  injection block.

CodeQL hardening (pre-existing alerts on this branch):
- mesh-plugin/src/agt-transport.ts: log error class instead of full message
  to avoid clear-text-logging-of-sensitive-information.
- cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch
  to satisfy js/file-access-to-http.

Verified live on demoagtmesh + analyst/viz/writer with file-first prompt
fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4
with zero tool-call parse failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log

CodeQL js/clear-text-logging was still flagging the truncated toAmid
prefix as taint from process.env. Log only a fixed string + error
class; full error preserved on throw so caller's /prekey/i matcher
still works.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
@pallakatos
Pal Lakatos-Toth (pallakatos) deleted the agt-runtime-swap-phase2 branch June 1, 2026 14:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants