Repository navigation
feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap - #245
Merged
Merged
Conversation
Wire azureclaw runtime through createMeshTransport() factory so we can flip
between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk
via AZURECLAW_MESH_PROVIDER without code changes.
Surface additions to IMeshTransport (both adapters now expose):
- lookup(amid) — registry RPC for reputation/display name
- submitReputation(...) — registry RPC for peer feedback
- enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on)
- onError(kind, from, detail) — diagnostic hook for decrypt + ws errors
- onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal
- onDisconnect(reason, code) — ws close / error fan-out
mesh-plugin (vendored A adapter):
- connection.ts delegates to the underlying SDK; lazy bind for hooks
registered before connect()
- 16-test compatibility suite (transport-phase2-compat.test.ts) pins the
contract so neither adapter can drop a method without CI failing
mesh-plugin (AGT B adapter):
- agt-transport.ts implements lookup/submitReputation as REST calls to the
registry (AGT MeshClient is pure transport — registry RPCs intentionally
not added to AGT upstream; they belong on a separate RegistryClient)
- enableKnockEnforcement is a no-op (AGT MeshClient always enforces)
- Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified}
methods (added on local AGT branch azureclaw-meshclient-event-hooks,
NOT pushed — AGT team owns the upstream PR)
runtime (runtimes/openclaw):
- Adds @azureclaw/mesh as a file: dependency
- Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)'
when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value
- Identity is generated once via vendored SDK regardless of provider, then
raw Ed25519 keys are extracted via toData() and shared across both — same
AMID either way
- Banner now reports active provider (vendored vs agt)
Docs:
- docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity,
policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK,
plaintext peers, file transfer + the wiring + migration path
- Documents the 3 hooks added to local AGT branch and the 3 governance
methods kept adapter-side
Tests:
- mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat)
- runtimes/openclaw: 118/118 pass
- AGT (local branch): 387/387 pass with 8 new event-hook tests
Open work for cleanup phase: once AGT publishes the version with our event
hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var
toggle.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI runs Rust 1.95.0 which added new clippy lints:
- doc_lazy_continuation: indent doc list items that span multiple lines.
Added two-space indent to the trailing 'All three are populated...'
paragraph so it is treated as a continuation of the preceding list
item rather than its own malformed list item.
- obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as
if .. { a } else { b } per the lint suggestion.
These were pre-existing on dev (CI only started failing once the runner
picked up Rust 1.95.0); fixing here so PR #245 can land green.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ocally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass.
Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker
mode before AKS rollout. Same flag, three integration points:
CLI (cli/src/commands/dev.ts):
- New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball
- First-run interactive prompt offers AGT only if the toolkit
checkout is actually present locally — silently defaults to
vendored otherwise (no pestering for users without AGT cloned).
- --build branch: builds the right relay/registry images
vendored → vendor/agentmesh-relay + agentmesh-registry (Rust)
agt → agent-governance-python/agent-mesh/docker/Dockerfile
with COMPONENT=relay / registry build-args
- Sandbox image build: stages locally-packed AGT SDK tarball into
.agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL
build-arg (auto-discovers if --agt-sdk-tarball not given).
- Runtime branch: skips Postgres for AGT (in-memory registry),
uses correct ports (AGT: 8083 relay, 8082 registry; vendored:
8765/8080) and health path (AGT: /healthz; vendored: /v1/health).
- Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the
runtime transport-factory honors the user's choice.
Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile):
- New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the
sandbox npm-installs the local tarball instead of fetching the
published @microsoft/agent-governance-sdk from npm. Lets us
smoke-test the locally-patched AGT branch (G3/G4 fixes) end to
end without round-tripping through npm publish.
- .agt-sdk/ staging dir always exists (with .keep) so the COPY
never fails when the user didn't stage a tarball.
Defaults preserved: --mesh-provider=vendored, existing behavior is
byte-identical for users who don't opt in.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…solves
The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT
provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so
tsc failed with TS2307 in the AGT build path.
Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the
build context, and strip its 'prepare' script (which would invoke tsc,
not present in this stage; the pre-built dist/ is sufficient).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…gistry API
Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream
AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks).
- connect() now passes autoRegister: true so the SDK uploads identity and
prekeys instead of the adapter re-implementing that path with raw HTTP.
- discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover
(not /registry/search), so the previous raw-HTTP path was 404-ing under AGT.
- lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}).
- submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score
clamped to [0,1]; the vendored /registry/feedback endpoint does not exist
in AGT.
- Replaced mapAgent with pickDisplayName helper: AGT puts display name in
metadata.display_name (set by registerSelf), with the first capability as
the fallback.
Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and
the bespoke fetchWithRetry helper — both are upstream concerns now.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes
the vendored registry wire shape. The vendored impl talks to /registry/* (the
existing agentmesh-registry); the AGT impl talks to /v1/discover and
/v1/agents/{did} on the upstream AGT registry. Both expose a single normalized
RegistryEntry envelope, so callsites stay readable.
getMeshRegistry(routerUrl) is the entry point. Provider selection follows
AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with
AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base).
Migrated all raw-HTTP registry callsites:
- core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName,
resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid.
- core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS
spawn discovery, AKS→local discovery.
- core/agt-task-loop.ts (1 site): registry_capability_search tool.
- core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe,
azureclaw_discover (mesh_discover) tool.
- index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe,
heartbeat keepalive (no-op under AGT — relay does liveness via WS).
The discover-on-router-unreachable test now asserts the new contract: empty
list + count:0 instead of a 'Discovery failed' string. Registry hiccups must
not break tool calls.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced
six bugs blocking the upstream AGT MeshClient swap. All fixed:
1. Final sandbox Docker stage didn't COPY mesh-plugin, so the
file:../../mesh-plugin symlink dangled in node_modules. Plugin
swap silently fell back to vendored with 'Cannot find package
@azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist}
into /mesh-plugin/ in the final stage of the Dockerfile.
2. entrypoint.sh used cp -r when copying node_modules into the
plugin extension dir, preserving the (now-broken-at-runtime-path)
symlink. Switched to cp -rL so symlinks dereference into real
files in the target tree.
3. mesh-plugin/src/index.ts imported createMeshTransport from
./transport-factory.js but never re-exported it. Runtime swap
path couldn't find the factory. Added the missing re-export.
4. inference-router agt_registry_proxy unconditionally prepended
'/v1/' to every path, so AGT SDK's already-qualified 'v1/agents'
became '/v1/v1/agents' at the upstream. Now: forward verbatim
when path starts with 'v1/' or equals 'health', else prepend.
Preserves vendored SDK behavior ('registry/register' → /v1/registry/register).
5. /agt/relay route only matched the bare path, but AGT MeshClient
appends '/ws' to relayUrl. Added /agt/relay/ws route and made the
upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt.
6. agt_registry_proxy route was declared get(...).post(...) only.
AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey
upload and DELETE for deregister — both 405'd at the router.
Added .put() and .delete() to the route declaration.
Bug #6 was invisible to vendored because the vendored SDK only
ever uses GET/POST (registry/register, registry/prekeys, etc.).
AGT's switch to REST verbs exposed the gap.
Path allowlist also extended with 'v1/' prefix so AGT's REST paths
(v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation.
Verified end-to-end via azureclaw dev --mesh-provider=agt --build:
- POST /v1/agents → 201 Created
- PUT /v1/agents/{did}/prekeys → 200 OK
- WebSocket /ws accepted, stable connection (no reconnect loop)
- Plugin reports 'AGT mesh connected' + provider=agt
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw_discover and other mesh registry callsites went via
`process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`,
intending to let out-of-sandbox sub-agents bypass the router.
In practice, the sandbox launcher always sets AGT_REGISTRY_URL as
the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082
in dev, the K8s service URL in prod). Since the runtime runs as
UID 1000 and iptables egress-guard blocks UID 1000 from anything
except localhost+DNS, the direct upstream URL ECONNREFUSEs and the
catch-all silently returns []. Symptom: registered agents are
invisible to azureclaw_discover even though they show up in
`GET /v1/discover` when queried directly at the registry.
Drop the env-var override — there's no in-sandbox runtime path
where bypassing the router is correct. The router is the ONLY
way out for UID 1000.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but
routerCall actually resolves with the parsed JSON error body. When the
sub-agent pod/container is gone the router returns 404 with
{ error: "Container '<name>' not found..." } and probeSubAgentAlive
read status.phase = undefined → defaulted to "Unknown" → not in
POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every
2s forever, blocking the LLM event loop ("LLM not responding" symptom).
Also narrow the prekey transient retry test so permanent X3DH /
signature-verification failures bubble up instead of being treated as
"waiting for prekeys" and retried indefinitely.
Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'.
Before: registry log fills with GET /v1/discover?capability=echo-buddy
every ~2s forever; LLM stops responding to new turns.
After: mesh_send aborts with 'sub-agent sandbox not found' on first probe.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Three vendored-only registry paths were being called unconditionally in
AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply
cadence:
1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X.
The AGT registry exposes GET /v1/discover?capability=X instead — display
names live in the per-agent record, so the AGT path fans out to a
second /v1/agents/{did} fetch per discover hit. Driven by the operator
panel's /agt/reputation polling.
2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session.
AGT has no per-session counter; per-agent reputation already submitted
via MeshClient.submitReputation. No-op in AGT mode.
3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke
on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter
for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode.
Also includes complementary debugging fixes from this session:
- agt-transport: auto-call establishSessionWithPeer() before send() so AGT
mode gets vendored-equivalent send-with-first-contact semantics. Without
this, send() throws 'No encrypted session — call establishSession() first'
and the retry loop spins forever.
- cli operator fetchers: add 8–10s timeouts to kubectl get calls that
were hanging when the cluster API was unreachable.
cargo check: clean
runtimes/openclaw: 118 vitest tests pass
inference-router: 8 mesh tests pass
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fourth 404 leak revealed after deploying the previous fixes: the
operator panel's ~30s /agt/reputation poll triggers
governance::agt_reputation, which (after lookup_parent_amid succeeds)
fetched the per-agent reputation score via the vendored-only
GET /v1/registry/reputation/score?amid=X path. AGT registry has no
such endpoint — the score is embedded as 'reputation_score: f64' in
the per-agent record returned by /v1/agents/{did}.
Provider-dispatch the URL; for AGT, wrap the agent record in a
vendored-shaped payload (score / tier / raw) so downstream CLI
fetchers and the operator panel stay schema-agnostic.
cargo check: clean
agt_governance_integration: 26/26 pass
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…dels) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ifest
Phase B.1 of the AGT-on-AKS rollout (see session plan
files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the
existing `azureclaw dev --mesh-provider` flag so the same provider
selection works for AKS pushes.
When --mesh-provider=agt:
* Builds relay+registry from the AGT upstream Dockerfile
($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile)
using COMPONENT=relay|registry build-args (matches dev.ts).
* Tags as agentmesh-{relay,registry}-agt:latest so both vendored
and AGT images can coexist on the same ACR and so the existing
deploy/agentmesh-agt.yaml manifest picks them up unchanged.
* Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered
in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz)
into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg.
* Always passes MESH_PROVIDER build-arg to the sandbox image so the
Dockerfile's conditional `npm install @microsoft/agent-governance-sdk`
runs for AGT clusters.
When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml,
applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt,
THEN rolls the controller (so the new pod reads
AZURECLAW_MESH_PROVIDER=agt for new sandboxes).
Auto-reverses when --apply --mesh-provider=vendored runs against a
cluster currently on AGT (no Postgres deployment in the agentmesh ns).
The image build loop also now supports absolute Dockerfile paths and
absolute build contexts via a new `absoluteContext` field, needed
because the AGT Dockerfile lives outside the azureclaw repo root.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks
without rebuilding any images, assuming both image pairs were already
seeded by 'azureclaw push'.
Flow:
1. Detect current provider via 'kubectl get deploy/postgres -n
agentmesh' (vendored has Postgres, AGT does not).
2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found
3. kubectl apply -f deploy/agentmesh-<target>.yaml
4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target>
5. kubectl rollout restart deploy/azureclaw-controller
6. With --restart-sandboxes: roll every azureclaw-managed Deployment
so existing pods pick up the new AZURECLAW_MESH_PROVIDER value.
Service names and ports are identical between the two manifests
(agentmesh-relay:8765, agentmesh-registry:8080) so the controller's
mesh_peer talks to either stack with no further config — the relay/
registry URLs already come from env vars (MESH_RELAY_URL /
MESH_REGISTRY_URL).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up'
so first-time deploys can ship AGT instead of vendored.
When --mesh-provider=agt:
* helm install runs with --set mesh.provider=agt (controller env
AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes).
* deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of
deploy/agentmesh.yaml.
* Skips the postgres ACR import and the agentmesh-db-credentials
secret creation (both unused by AGT — its registry is in-memory).
* Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml
vs .tmp-agentmesh.yaml) so concurrent provider switches don't
collide.
The deployAgentMesh signature gains a non-breaking 'meshProvider'
option that defaults to 'vendored' (existing callers untouched).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Implement full AGT relay/registry wire support in the controller's
mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt.
Without this the controller's federation peer cannot connect to the
AGT relay (different WS path, frame envelope, heartbeat, ack model)
or the AGT registry (different HTTP shape, no signed body), and the
leader fails-loops on AGT clusters — breaking the only cloud-offload
control path.
New module `mesh_peer/agt_wire.rs`:
- `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error)
with `#[serde(tag="type", rename_all="snake_case")]` matching
`agentmesh/relay/app.py`.
- `AgtRegisterAgentRequest` struct for `POST /v1/agents`.
- 7 unit tests pinning the serialized shape.
`mesh_peer/mod.rs`:
- New `Provider` enum + `Provider::from_env()` selecting vendored
(default) or AGT off `AZURECLAW_MESH_PROVIDER`.
- `MeshPeerState.provider` carried through outbound + inbound paths.
- `register_with_registry()` branches: vendored signs ts body;
AGT posts `{did, public_key (base64url), capabilities, metadata}`
with no signature; 409 treated as success for leader-failover idempotency.
- `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>`
(matches JS SDK `buildDid`), so every leader replica converges on
the same DID without coordination.
- Default `MESH_RELAY_URL` appends `/ws` for AGT.
- `connect_and_listen()`:
- AGT connect frame `{type:"connect", from:<did>, token?:<env>}`
(token read from `AGENTMESH_RELAY_TOKEN` if set).
- AGT has no `Connected` ack — mark `connected=true` immediately.
- Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored
keeps `ping`).
- `serialize_and_send_outbound()` / `send_to_peer()` now take `state`
and branch outbound framing — AGT emits `message` frames
`{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex).
- `handle_message()` dispatches to `handle_vendored_frame()` or
`handle_agt_frame()`. AGT path:
- Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`.
- Sends `Ack` reply (required — without it AGT redelivers on
reconnect → duplicate offload processing).
- Treats `Error` frames mentioning Authentication failed /
Missing 'from' / session_replaced as fatal — drops connection
for reconnect.
`mesh_peer/offload.rs`:
- All 8 `send_to_peer(...)` call sites updated to pass `&state` first.
`main.rs`:
- Remove the temporary AGT-skip guard around `mesh_peer::run`. The
peer now starts unconditionally when enabled; provider is consumed
inside `mesh_peer::run`.
Build/test:
- cargo build --release --package azureclaw-controller: OK
- cargo test --package azureclaw-controller: 492 passed
- cargo clippy --package azureclaw-controller --all-targets -D warnings: OK
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz
(see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness
probes were 404'ing → CrashLoopBackOff/NotReady.
operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane
evaluates NetworkPolicy egress against the backend pod port
(post-DNAT), not the Service port. AGT registry/relay listen on
8082/8083; the Service maps 8080->8082 and 8765->8083 so the
Service-port allowlist (8080/8765) doesn't actually permit the
post-DNAT flow. Add 8082/8083 alongside so both vendored
(8080/8765 direct) and AGT paths work.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
azureclaw dev now asks new users where the mesh should live, just
like the existing inference-provider picker:
Where should the mesh live?
❯ Local (recommended; spin up relay + registry in Docker)
Remote (auto port-forward to AKS cluster: <cluster-name>)
Local (default) keeps the existing behaviour: docker-compose'd
relay/registry/postgres on the user's laptop.
Remote (advanced) federates with a previously-provisioned AKS mesh:
- If ~/.azureclaw/context.json has a cached globalRegistryUrl
from a prior 'azureclaw mesh promote', reuse it verbatim.
- Otherwise default to http://localhost:18080 — the port-forward
URL 'mesh promote --port-forward' uses — so the auto-promote
fallback in the downstream global-registry block will spawn
the tunnels on demand.
- If there is no aksCluster in context at all, warn and fall
back to local so the user isn't left with a broken sandbox.
Skipped entirely when --global-registry was passed explicitly (the
advanced flag overrides the prompt) or when the user is past their
first run.
Also fixed a latent AGT-compat bug in the same flow: the existing
'auto-promote' path probed only /v1/health, which 404s on AGT
clusters. Replaced with a /health → /v1/health fallback (matches
the same shape we used in checkRegistryHealth last commit).
Verified:
- npm run build / typecheck clean
- 640 CLI tests pass (2 skipped, no regressions)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply
back to their spawner, but on AGT the registry has no agent named or
capability="parent" — the search returns 0 → no prekey bundle → send
fails. The vendored runtime had this aliased only in the offload-mode
task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the
controller never set for AKS-spawned children.
Two coordinated fixes:
1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set
(spawner seeds 'parent_name:parent_AMID' as the first entry), also
push PARENT_SANDBOX=<first_name> into the openclaw container env.
2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send
and azureclaw_mesh_transfer_file, alias to_agent=='parent' →
PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry
lookup. The Symbol is set during runtime init from
AGT_TRUSTED_PEERS[0], so this works even on images built before fix
#1 lands. Skip in offload mode — 'parent' there is a protocol-level
routing token, not a mesh recipient name.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
mesh-plugin/src/agt-transport.ts.send() called
this.client.establishSessionWithPeer(toAmid) before forwarding to
client.send(). That method does not exist on AgentMeshClient — the
real method is establishSession(toAmid, options) — so every parent →
sub-agent send on AGT was failing with:
establishSessionWithPeer is not a function
It was also unnecessary: AgentMeshClient.send() already auto-bootstraps
the X3DH handshake on first contact (see @agentmesh/sdk
AgentMeshClient.send → cache miss → establishSession() fallthrough at
dist/index.js:3321-3334). Calling establishSession() ourselves would
also be wrong because it is not idempotent — it unconditionally writes
activeSessions.set and starts a fresh X3DH.
Fix: remove the pre-bootstrap entirely and let client.send() manage
session lifecycle. The AgtSdkModule type loses the required
establishSessionWithPeer member (now optional) since we no longer
depend on it; test fakes remain valid as harmless extras.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previous commit aa7d28e wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Commit 9f48f87 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The AGT registry has no autonomous presence model — `last_seen`
is frozen at registration and the `update_last_seen()` store
method is dead code with no HTTP handler calling it. Combined
with the openclaw discover tool's 90s stale filter
(agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes
silently invisible 90s after spawn, breaking sibling-to-sibling
peer discovery.
Demo symptom: analyst/viz/writer all reported 'peer discovery
did not return ...' even though mesh_send to those names
succeeded with 'delivered_and_replied'. The relay was fine; only
the registry's presence view was stale.
Pair with the corresponding upstream registry change (AGT branch
`azureclaw-meshclient-event-hooks`, commit adds
POST /v1/agents/{did}/heartbeat -> store.update_last_seen).
The new tick reuses the existing 30s relay-keepalive timer in
connect(), so no extra timers and no extra event-loop pressure.
Best-effort: 4xx/5xx are warned-once, network errors swallowed,
loop survives a registry pod restart (next tick retries).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Imran Siddique (imran-siddique)
pushed a commit
to microsoft/agent-governance-toolkit
that referenced
this pull request
May 11, 2026
…ify, heartbeat (#2090) * feat(mesh-client): add onError, onDisconnect, onE2EVerified event hooks Adds three observer-registration methods to MeshClient to align with the AzureClaw vendored AgentMesh SDK surface so consumers can swap providers behind a single transport interface: onError(handler) — fires on ws errors and decrypt failures onDisconnect(handler) — fires on ws.close with reason+code onE2EVerified(handler) — fires on first successful decrypt per peer Pure additions; no behaviour change for existing flows. Decrypt-failure and missing-session paths now drop the message AND notify error observers instead of silently dropping (still safe — unhandled events are swallowed). Tests: 8 new in mesh-client-event-hooks.test.ts; full suite 387/387 green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-client): close gaps G1 (KNOCK auto-bootstrap) and G2 (auto-reconnect) Two protocol-level gaps were identified during the AzureClaw vendored-SDK audit (see Azure/kars#245 docs/agt-vs-vendored-sdk.md). Both block moving fully off the vendored agentmesh-sdk fork onto upstream AGT. ## G1: receiver-side X3DH auto-bootstrap from KNOCK Before: the responder side of an encrypted session was created only by calling acceptSession(peerId, establishment) explicitly. But the ChannelEstablishment was never carried on the wire — neither in knock_accept nor in the first message frame — so the receiver had no way to obtain it. Result: any fresh encrypted session failed with "No encrypted session" the first time a ciphertext arrived. Mirrors vendored agentmesh-sdk patch #4b. Behavior: - Sender (establishSession): builds the SecureChannel first, then sends KNOCK with the ChannelEstablishment embedded as { ik: base64, ek: base64, otk?: number }. - Receiver (handleKnock): if the knock contains `establishment` AND no prior session exists, deserialize and call acceptSession() automatically before responding with knock_accept. On bad establishment data, the knock is rejected and onError("knock", ...) fires. - Backwards-compatible: legacy peers that don't embed establishment continue to work; caller must invoke acceptSession() manually as before. ## G2: auto-reconnect on transport drop Before: MeshClient.reconnect() existed but was never called automatically. Network blips, relay restarts, AKS node OOM-evicts left the client disconnected forever — agents went mesh-deaf. Mirrors vendored agentmesh-sdk patch #9. Behavior: - New options: autoReconnect (default true), maxReconnectAttempts (default Number.POSITIVE_INFINITY), reconnectBaseDelayMs (default 1000), reconnectMaxDelayMs (default 60000). - On non-1000 ws.onclose: schedule reconnect with exponential backoff capped at reconnectMaxDelayMs. Light jitter (±20%) avoids thundering-herd reconnects across many sandboxes. - onDisconnect handlers fire BEFORE the reconnect is scheduled so observers can see the drop. - After maxReconnectAttempts, fires onError("ws", ..., "auto-reconnect gave up after N attempts") and stops. - disconnect() cancels any pending reconnect timer. - A connect() failure inside the reconnect path schedules another retry via the existing onclose path (and recursively from the catch handler in scheduleReconnect for the case where ws.onopen never fired). ## Tests 11 new tests across two files: - tests/mesh-client-knock-bootstrap.test.ts: 5 tests covering sender-side embedding, receiver auto-bootstrap, legacy fallback, malformed establishment, and end-to-end happy-path encrypted send. - tests/mesh-client-auto-reconnect.test.ts: 6 tests covering server-close reconnect, client-close no-reconnect, opt-out, give-up after max attempts, disconnect cancels pending timer, no duplicate scheduling. Full AGT TS suite: 398/398 pass (was 387 before this commit). Build clean. Both gaps were identified in the AzureClaw audit doc: docs/agt-vs-vendored-sdk.md (Patch-by-patch audit section). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): close vendored-parity gaps G3, G4, G5 G3 — session-desync teardown (vendored agentmesh-sdk patch #13 equiv): when Double Ratchet decryption fails for an existing session, the local ratchet state is irrecoverable. Previously we only fired onError('decrypt'); now we tear down the session (delete from this.sessions; clear knockAccepted) and fire a distinct onError('session_desync') so callers can re-run establishSession() and resume communication. G4 — pre-KNOCK encrypted-message buffer (vendored agentmesh-sdk patch #16 equiv): under transport reordering the relay can deliver an encrypted 'message' frame BEFORE the matching 'knock' arrives. Without buffering, that first frame was dropped silently. Now MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms), replays them when the corresponding knock is accepted, and drops them when rejected or on disconnect. Disabled by setting preKnockBufferSize: 0. G5 — eager ghost-connection close on rebind (vendored agentmesh-relay patch #2 equiv): when an agent reconnects with the same DID, the prior 'ghost' WebSocket is now closed eagerly with code 1000 'session_replaced' instead of waiting for the 90s heartbeat-eviction timer. The 'finally' cleanup now compares socket identity to avoid removing the freshly rebound connection when the old socket's handler unwinds. Tests: +5 (G4 pre-knock buffer behaviour), +2 (G3 type + lifecycle safety), +1 (G5 ghost rebind). 405 TS tests pass; 18 relay tests pass. NOTE: held local-only on branch azureclaw-meshclient-event-hooks; will be coordinated upstream with the AGT team. * feat(mesh): add RegistryClient and auto-register on MeshClient.connect() Closes the structural gap where MeshClientOptions declared registryUrl but no code path ever used it. Without this, agents can connect to the relay and exchange ciphertext but are invisible to peers — they never appear in /v1/discover and no one can fetch their X3DH pre-key bundle. This commit adds the missing registry-side glue, in upstream-quality shape, so AGT MeshClient becomes a drop-in replacement for vendored forks that wired their own registry calls (e.g. AzureClaw vendor/agentmesh-sdk). What's new: - src/encryption/registry-client.ts: typed HTTP client wrapping the AgentMesh registry's REST surface (POST /v1/agents, PUT/GET /v1/agents/{did}/prekeys, GET /v1/discover, GET/DELETE /v1/agents/{did}). base64url <-> Uint8Array marshalling, optional Ed25519-Timestamp Authorization signer, configurable retry/timeout. - src/encryption/mesh-client.ts: * New options: registryClient, registryClientOptions, capabilities, registrationMetadata, oneTimePrekeyCount, autoRegister. * connect() now calls registerSelf() at the end of a successful WS handshake when autoRegister is true (default) and a registry is configured. Idempotent — 409 (already registered) is treated as success; reconnect doesn't re-register. * registerSelf(): generates signed-prekey + N one-time prekeys via keyManager, then POSTs /v1/agents (capabilities = [displayName, ...options.capabilities]) and PUTs the prekey bundle. * discover(capability): wraps registry.discover. * establishSessionWithPeer(peerId): fetches peer prekeys from the registry and calls existing establishSession. * getRegistry(): exposes the underlying client for advanced cases. - tests/registry-client.test.ts: 11 new tests covering wire format, base64url round-trip, idempotent register (409), 5xx retry, and Authorization header. Also covers MeshClient auto-register on connect end-to-end with a fake fetch + fake WebSocket. - tests/mesh-client-*.test.ts: existing fixtures updated to pass autoRegister: false (no registry available in those tests). Behavior unchanged for the production path. All 71 tests pass (60 previously + 11 new). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(registry): add POST /v1/agents/{did}/heartbeat to bump last_seen The registry's last_seen field is set on registration but never updated by any HTTP handler — update_last_seen() exists in store.py but is dead code in production. As a result, every registered agent looks 'offline' (online=false on /presence) 90s after spawn, and any client-side discover stale-filter rejects all live agents. Add an unauthenticated POST /v1/agents/{did}/heartbeat that calls store.update_last_seen(). Idempotent. Returns 404 for unknown DIDs so callers can detect a registry restart and re-register. Mirrors the agent-side ping cadence (30s, well within the 90s presence threshold). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): include Ed25519 identity_key_ed in prekey bundle so verifyBundle() works PreKeyBundle.identityKeyEd was previously aliased to the X25519 identity_key because the registry didn't store the Ed25519 signing key. As a result, verifyBundle() either passed only because the stub test never exercised it, or quietly accepted any signed pre-key — a soundness gap in the X3DH wrapper. This commit threads identityKeyEd through the full path: - agent-mesh/registry/store.py (AgentRecord): add identity_key_ed field (X25519 long-term key already present as identity_key; Ed25519 distinct). - agent-governance-typescript/src/encryption/x3dh.ts: expose X3DHKeyManager.identityKeyEd getter. - agent-governance-typescript/src/encryption/registry-client.ts: uploadPrekeys() now accepts identityKeyEd (32 bytes, validated) and serializes it as identity_key_ed; fetchPrekeys() deserializes it back into PreKeyBundle.identityKeyEd. Older bundles without the field fall back to the X25519 key — verifyBundle() will then correctly reject (forensic visibility, no silent bypass). - agent-governance-typescript/src/encryption/mesh-client.ts: registerSelf() passes keyManager.identityKeyEd through to the registry. - tests/registry-client.test.ts + tests/test_registry.py: assert the new field is round-tripped on PUT/GET. Wire compatibility: PUT now requires identity_key_ed; GET emits it when set. Existing AGT clients that don't upload it will get verifyBundle() rejection on the receiving side — peers must upgrade together. (No silent regression on the wire — the field is new, not a breaking change to existing fields.) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test(mesh): set autoRegister:false in upstream knock + malformed-frame tests These two upstream tests construct MeshClient without autoRegister:false, which now triggers a real fetch to the registry on connect() (added by our registry-client commit). The tests have no fake registry, so they fail with 'fetch failed'. autoRegister:false skips the auto-registration path and matches the pattern used in our 6 sibling mesh-client-*.test.ts files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos)
added a commit
that referenced
this pull request
May 12, 2026
…#245) * feat(mesh): Phase 2 — provider-agnostic IMeshTransport + runtime swap Wire azureclaw runtime through createMeshTransport() factory so we can flip between the vendored @agentmesh/sdk and Microsoft's @microsoft/agent-governance-sdk via AZURECLAW_MESH_PROVIDER without code changes. Surface additions to IMeshTransport (both adapters now expose): - lookup(amid) — registry RPC for reputation/display name - submitReputation(...) — registry RPC for peer feedback - enableKnockEnforcement() — vendored toggle (no-op on AGT, always-on) - onError(kind, from, detail) — diagnostic hook for decrypt + ws errors - onE2EVerified(peer, isFirst) — first-decrypt-per-peer signal - onDisconnect(reason, code) — ws close / error fan-out mesh-plugin (vendored A adapter): - connection.ts delegates to the underlying SDK; lazy bind for hooks registered before connect() - 16-test compatibility suite (transport-phase2-compat.test.ts) pins the contract so neither adapter can drop a method without CI failing mesh-plugin (AGT B adapter): - agt-transport.ts implements lookup/submitReputation as REST calls to the registry (AGT MeshClient is pure transport — registry RPCs intentionally not added to AGT upstream; they belong on a separate RegistryClient) - enableKnockEnforcement is a no-op (AGT MeshClient always enforces) - Event hooks delegate to AGT MeshClient's new on{Error,Disconnect,E2EVerified} methods (added on local AGT branch azureclaw-meshclient-event-hooks, NOT pushed — AGT team owns the upstream PR) runtime (runtimes/openclaw): - Adds @azureclaw/mesh as a file: dependency - Replaces 'new sdk.AgentMeshClient(...)' with 'await createMeshTransport(...)' when AZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value - Identity is generated once via vendored SDK regardless of provider, then raw Ed25519 keys are extracted via toData() and shared across both — same AMID either way - Banner now reports active provider (vendored vs agt) Docs: - docs/agt-vs-vendored-sdk.md — full side-by-side analysis covering identity, policy, trust, audit, transport, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer + the wiring + migration path - Documents the 3 hooks added to local AGT branch and the 3 governance methods kept adapter-side Tests: - mesh-plugin: 97/97 pass (81 pre-Phase 2 + 16 new compat) - runtimes/openclaw: 118/118 pass - AGT (local branch): 387/387 pass with 8 new event-hook tests Open work for cleanup phase: once AGT publishes the version with our event hooks merged, drop vendor/agentmesh-sdk/ entirely and remove the env-var toggle. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): pre-build mesh-plugin for runtime CI + format reconciler PR #245 CI failures: 1. Runtime job failed with TS2307 'Cannot find module @azureclaw/mesh' — the runtime depends on mesh-plugin via 'file:../../mesh-plugin' but the CI workflow only ran 'npm install' inside runtimes/openclaw, which does not build the file: dep's dist/. Add an explicit pre-build step that installs vendored agentmesh-sdk + mesh-plugin and runs its build before the runtime install. 2. Rust fmt check failed on controller/src/reconciler/mod.rs — drift inherited from PR #244. Run cargo fmt --all. Also added a 'prepare' script to mesh-plugin/package.json so any future file: consumer auto-builds on install (defensive — the explicit CI step above is still the primary fix). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): quote workflow step name containing colon YAML parser rejected 'Build mesh-plugin (file: dep of runtime)' because 'file:' was interpreted as a mapping key. Quote the string. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): clippy fixes for Rust 1.95.0 CI runs Rust 1.95.0 which added new clippy lints: - doc_lazy_continuation: indent doc list items that span multiple lines. Added two-space indent to the trailing 'All three are populated...' paragraph so it is treated as a continuation of the preceding list item rather than its own malformed list item. - obfuscated_if_else: rewrite is_empty().then_some(a).unwrap_or(b) as if .. { a } else { b } per the lint suggestion. These were pre-existing on dev (CI only started failing once the runner picked up Rust 1.95.0); fixing here so PR #245 can land green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): full patch-by-patch audit + adapter-side fixes for #7/#12 Audit findings (docs/agt-vs-vendored-sdk.md): - Verified each of the 9 vendored SDK patches against AGT MeshClient - Verified all 4 vendored relay + 4 vendored registry patches - Identified 5 protocol-level gaps that block Phase 3: * G1: receiver-side X3DH bootstrap (no auto-create on first encrypted msg) * G2: no auto-reconnect loop (manual reconnect() only) * G3: registry RPCs not in MeshClient (compensated in adapter) * G4: fast-fail handshake edge (defensive) * G5: connect frame incompatibility with vendored relay (BLOCKING) - Documented which gaps require AGT-upstream changes vs adapter fixes - Updated migration strategy: Phase 3 BLOCKED until AGT lands G1, G2, G5 Adapter-side fixes (mesh-plugin/src/agt-transport.ts): - Patch #7 port: submitReputation now logs status + body on non-2xx and logs network errors (vendored swallowed both silently) - Patch #12 port: registry fetches now use bounded retry with exponential backoff (250ms, 750ms, 2000ms) — applied to lookup, submitReputation, and discovery search Tests: 97/97 mesh-plugin tests pass (no new tests needed — existing unreachable-registry tests now also exercise retry path). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): reframe audit for upstream-AGT scenario, drop invalid Gap G5 The previous audit framed gaps as 'AGT vs vendored relay' which is the wrong question — when we move fully upstream, AGT will use its own Python relay and registry, not ours. So wire-format compat with the vendored relay (the old G5) is irrelevant by design. Re-audit against the full AGT upstream stack (TS SDK + Python relay + Python registry): - Confirmed AGT registry already does Ed25519-over-raw-timestamp signature verification (registry/app.py:54-98) — same approach we patched into the vendored registry. No port needed. - Confirmed AGT relay has /health, heartbeat, and 90s offline threshold. - Confirmed AGT registry tracks last_seen with 90s online window. - AGT relay overwrites duplicate connections without explicit close — slower than our 4001 SessionReplaced but functionally similar. - AGT uses shared-secret token auth on relay (no per-frame sig) — different security model than our vendored relay; flagged for review but not a functional regression. Real gaps that block moving upstream remain only 2: - G1: receiver-side X3DH bootstrap (acceptSession() exists but ChannelEstablishment is never serialized onto the wire) - G2: no auto-reconnect loop in MeshClient (manual reconnect() only) Both are well-scoped fixes to AGT's mesh-client.ts. The 3 event hooks on the local AGT branch are a prerequisite for cleanly implementing G2. Migration strategy updated to reflect that A↔B cross-provider message interop is not a goal (different relays by design); the swap unit is the sandbox, not the message. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): mark gaps G1 and G2 fixed on local AGT branch Audit doc updated to reflect that both protocol gaps identified during the vendored-vs-AGT audit are now closed on the local AGT branch `azureclaw-meshclient-event-hooks` (commit `d75ea37b`): - G1 KNOCK auto-bootstrap: `establishSession()` embeds X3DH params on the wire; `handleKnock()` auto-calls `acceptSession()` on receipt. Backwards-compatible with legacy peers. - G2 auto-reconnect loop: exponential backoff (1s → 60s, ±20% jitter) on non-1000 close; `autoReconnect: true` by default; opt-out via options. AGT TS test suite: 398/398 pass (was 387 before; 11 new tests across `mesh-client-knock-bootstrap.test.ts` and `mesh-client-auto-reconnect.test.ts`). The AGT branch is held locally — NOT pushed — pending coordination with the AGT team for an upstream PR. From AzureClaw's perspective, the upstream-AGT scenario is now feature-complete: every vendored patch has either been merged upstream, has an equivalent in AGT, lives in our adapter, or is fixed on the local AGT branch. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(agt): complete patch-by-patch audit with gaps G3, G4, G5 fixed locally Extends the AGT-vs-vendored-SDK audit to cover the previously-unaudited patches: SDK #10 (idempotent initiateSession), #11 (wsFactory + plaintextPeers), #13, #14 (vendored-dist-only bug), #15 (different KNOCK-once model), #16, #17 (Buffer.from-based, no spread overflow), #18 (simpler closeSession-based recovery). Documents three additional real gaps now fixed locally on the AGT branch (azureclaw-meshclient-event-hooks, commit 3a96a0f2): - G3 (vendored SDK #13): MeshClient now tears down the session on decrypt failure and fires onError('session_desync', ...) so the caller can re-run establishSession() to recover. Without G3, a single ratchet drift permanently jams the channel. - G4 (vendored SDK #16): MeshClient buffers encrypted frames per-peer (default cap 5, TTL 3000ms) when no session exists yet, drains on knock_accept, drops on knock_reject. Without G4, relay frame reorder silently loses the first message of every fresh handshake. - G5 (vendored relay #2): AGT relay now closes the previous WebSocket with code 1000 'session_replaced' before overwriting the _connections entry on rebind. Without G5, the old socket lingers for up to 90 seconds and messages route to a dead connection. The finally cleanup now compares socket identity to avoid removing the fresh connection on the old handler's unwind. Also updates the chunked file-transfer reliability note: G3 + G4 are both required for robust mesh_file_transfer because chunked transfers amplify silent-drop and ratchet-drift bugs into stuck transfers with no error surface. Summary table now: 12 already-in-AGT, 3 adapter-side, 7 different- but-equivalent, 5 real gaps all fixed locally on AGT branch (NOT pushed; awaiting upstream PR coordination with the AGT team). AGT TS test suite: 405/405 pass; AGT Python relay test suite: 18/18 pass. * dev: add --mesh-provider <vendored|agt> selection with first-run prompt Phase 3 prep: enables E2E testing the AGT runtime swap locally in Docker mode before AKS rollout. Same flag, three integration points: CLI (cli/src/commands/dev.ts): - New flags: --mesh-provider, --agt-repo, --agt-sdk-tarball - First-run interactive prompt offers AGT only if the toolkit checkout is actually present locally — silently defaults to vendored otherwise (no pestering for users without AGT cloned). - --build branch: builds the right relay/registry images vendored → vendor/agentmesh-relay + agentmesh-registry (Rust) agt → agent-governance-python/agent-mesh/docker/Dockerfile with COMPONENT=relay / registry build-args - Sandbox image build: stages locally-packed AGT SDK tarball into .agt-sdk/ build-context dir and forwards it via AGT_SDK_TARBALL build-arg (auto-discovers if --agt-sdk-tarball not given). - Runtime branch: skips Postgres for AGT (in-memory registry), uses correct ports (AGT: 8083 relay, 8082 registry; vendored: 8765/8080) and health path (AGT: /healthz; vendored: /v1/health). - Sandbox env: AZURECLAW_MESH_PROVIDER passed through so the runtime transport-factory honors the user's choice. Sandbox Dockerfile (sandbox-images/openclaw/Dockerfile): - New AGT_SDK_TARBALL build-arg. When set + MESH_PROVIDER=agt, the sandbox npm-installs the local tarball instead of fetching the published @microsoft/agent-governance-sdk from npm. Lets us smoke-test the locally-patched AGT branch (G3/G4 fixes) end to end without round-tripping through npm publish. - .agt-sdk/ staging dir always exists (with .keep) so the COPY never fails when the user didn't stage a tarball. Defaults preserved: --mesh-provider=vendored, existing behavior is byte-identical for users who don't opt in. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sandbox): copy mesh-plugin into cli-builder so @azureclaw/mesh resolves The runtime now imports @azureclaw/mesh (file:../../mesh-plugin) for AGT provider swap. The cli-builder Docker stage didn't copy mesh-plugin, so tsc failed with TS2307 in the AGT build path. Fix: copy mesh-plugin/{package.json,package-lock.json,dist/} into the build context, and strip its 'prepare' script (which would invoke tsc, not present in this stage; the pre-built dist/ is sufficient). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh-plugin): collapse agt-transport onto upstream MeshClient registry API Use the new MeshClient.registerSelf/discover/getRegistry surface from upstream AGT (microsoft/agent-governance-toolkit branch azureclaw-meshclient-event-hooks). - connect() now passes autoRegister: true so the SDK uploads identity and prekeys instead of the adapter re-implementing that path with raw HTTP. - discover() → meshClient.discover(capability); the AGT endpoint is /v1/discover (not /registry/search), so the previous raw-HTTP path was 404-ing under AGT. - lookup() → meshClient.getRegistry().getAgent() (correct /v1/agents/{did}). - submitReputation() ports to AGT POST /v1/agents/{did}/reputation with score clamped to [0,1]; the vendored /registry/feedback endpoint does not exist in AGT. - Replaced mapAgent with pickDisplayName helper: AGT puts display name in metadata.display_name (set by registerSelf), with the first capability as the fallback. Removes the manual generateSignedPreKey()/generateOneTimePreKeys() dance and the bespoke fetchWithRetry helper — both are upstream concerns now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(runtime): mesh-registry abstraction + migrate raw-HTTP callsites Introduce IMeshRegistry provider abstraction so the runtime no longer hardcodes the vendored registry wire shape. The vendored impl talks to /registry/* (the existing agentmesh-registry); the AGT impl talks to /v1/discover and /v1/agents/{did} on the upstream AGT registry. Both expose a single normalized RegistryEntry envelope, so callsites stay readable. getMeshRegistry(routerUrl) is the entry point. Provider selection follows AZURECLAW_MESH_PROVIDER (vendored|agt). Sub-agents can override with AGT_REGISTRY_URL for a direct endpoint. Cached per (provider, base). Migrated all raw-HTTP registry callsites: - core/amid-cache.ts (5 sites): resolveAmidByName, resolveAmidToName, resolveSigningKey, registryLookupDisplayName, registrySearchFreshestAmid. - core/agt-handoff.ts (3 sites): sub-agent interrupt lookup, local→AKS spawn discovery, AKS→local discovery. - core/agt-task-loop.ts (1 site): registry_capability_search tool. - core/agt-tools/agt.ts (2 sites): azureclaw_status mesh_registered probe, azureclaw_discover (mesh_discover) tool. - index.ts (3 sites): REQUIRE_VERIFIED_TIER lookup, post-spawn AMID probe, heartbeat keepalive (no-op under AGT — relay does liveness via WS). The discover-on-router-unreachable test now asserts the new contract: empty list + count:0 instead of a 'Discovery failed' string. Registry hiccups must not break tool calls. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): wire AGT provider end-to-end (6 stackup bugs) End-to-end Docker test of azureclaw dev --mesh-provider=agt surfaced six bugs blocking the upstream AGT MeshClient swap. All fixed: 1. Final sandbox Docker stage didn't COPY mesh-plugin, so the file:../../mesh-plugin symlink dangled in node_modules. Plugin swap silently fell back to vendored with 'Cannot find package @azureclaw/mesh'. Fixed by staging mesh-plugin/{package.json,dist} into /mesh-plugin/ in the final stage of the Dockerfile. 2. entrypoint.sh used cp -r when copying node_modules into the plugin extension dir, preserving the (now-broken-at-runtime-path) symlink. Switched to cp -rL so symlinks dereference into real files in the target tree. 3. mesh-plugin/src/index.ts imported createMeshTransport from ./transport-factory.js but never re-exported it. Runtime swap path couldn't find the factory. Added the missing re-export. 4. inference-router agt_registry_proxy unconditionally prepended '/v1/' to every path, so AGT SDK's already-qualified 'v1/agents' became '/v1/v1/agents' at the upstream. Now: forward verbatim when path starts with 'v1/' or equals 'health', else prepend. Preserves vendored SDK behavior ('registry/register' → /v1/registry/register). 5. /agt/relay route only matched the bare path, but AGT MeshClient appends '/ws' to relayUrl. Added /agt/relay/ws route and made the upstream WS URL auto-append /ws when AZURECLAW_MESH_PROVIDER=agt. 6. agt_registry_proxy route was declared get(...).post(...) only. AGT RegistryClient uses PUT /v1/agents/{did}/prekeys for prekey upload and DELETE for deregister — both 405'd at the router. Added .put() and .delete() to the route declaration. Bug #6 was invisible to vendored because the vendored SDK only ever uses GET/POST (registry/register, registry/prekeys, etc.). AGT's switch to REST verbs exposed the gap. Path allowlist also extended with 'v1/' prefix so AGT's REST paths (v1/agents, v1/agents/{did}/prekeys, v1/discover) pass validation. Verified end-to-end via azureclaw dev --mesh-provider=agt --build: - POST /v1/agents → 201 Created - PUT /v1/agents/{did}/prekeys → 200 OK - WebSocket /ws accepted, stable connection (no reconnect loop) - Plugin reports 'AGT mesh connected' + provider=agt Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): always route mesh registry through inference-router azureclaw_discover and other mesh registry callsites went via `process.env.AGT_REGISTRY_URL || routerUrl("/agt/registry")`, intending to let out-of-sandbox sub-agents bypass the router. In practice, the sandbox launcher always sets AGT_REGISTRY_URL as the ROUTER'S upstream target (e.g., http://azureclaw-agt-registry:8082 in dev, the K8s service URL in prod). Since the runtime runs as UID 1000 and iptables egress-guard blocks UID 1000 from anything except localhost+DNS, the direct upstream URL ECONNREFUSEs and the catch-all silently returns []. Symptom: registered agents are invisible to azureclaw_discover even though they show up in `GET /v1/discover` when queried directly at the registry. Drop the env-var override — there's no in-sandbox runtime path where bypassing the router is correct. The router is the ONLY way out for UID 1000. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): break mesh_send infinite poll loop on dead sub-agent probeSubAgentAlive() relied on routerCall throwing on HTTP 4xx, but routerCall actually resolves with the parsed JSON error body. When the sub-agent pod/container is gone the router returns 404 with { error: "Container '<name>' not found..." } and probeSubAgentAlive read status.phase = undefined → defaulted to "Unknown" → not in POD_DEAD_PHASES → mesh_send retry loop kept polling /v1/discover every 2s forever, blocking the LLM event loop ("LLM not responding" symptom). Also narrow the prekey transient retry test so permanent X3DH / signature-verification failures bubble up instead of being treated as "waiting for prekeys" and retried indefinitely. Repro: spawn echo-buddy, destroy it, send mesh_send to_agent='echo-buddy'. Before: registry log fills with GET /v1/discover?capability=echo-buddy every ~2s forever; LLM stops responding to new turns. After: mesh_send aborts with 'sub-agent sandbox not found' on first probe. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): suppress /v1/registry/* 404 leaks in AGT mode Three vendored-only registry paths were being called unconditionally in AGT mode, producing 404 spam in the registry logs at ~30s/per-mesh-reply cadence: 1. lookup_parent_amid (router): hardcoded GET /v1/registry/search?capability=X. The AGT registry exposes GET /v1/discover?capability=X instead — display names live in the per-agent record, so the AGT path fans out to a second /v1/agents/{did} fetch per discover hit. Driven by the operator panel's /agt/reputation polling. 2. recordMeshSession (runtime): POST /agt/registry/registry/reputation/session. AGT has no per-session counter; per-agent reputation already submitted via MeshClient.submitReputation. No-op in AGT mode. 3. registerRevokeShutdownHook (runtime): POST /agt/registry/registry/revoke on SIGTERM. AGT uses WS-disconnect + receiver-side 90s last_seen filter for pruning; no /v1/registry/revoke endpoint exists. Skip in AGT mode. Also includes complementary debugging fixes from this session: - agt-transport: auto-call establishSessionWithPeer() before send() so AGT mode gets vendored-equivalent send-with-first-contact semantics. Without this, send() throws 'No encrypted session — call establishSession() first' and the retry loop spins forever. - cli operator fetchers: add 8–10s timeouts to kubectl get calls that were hanging when the cluster API was unreachable. cargo check: clean runtimes/openclaw: 118 vitest tests pass inference-router: 8 mesh tests pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): use /v1/agents/{did} for reputation lookup in AGT mode Fourth 404 leak revealed after deploying the previous fixes: the operator panel's ~30s /agt/reputation poll triggers governance::agt_reputation, which (after lookup_parent_amid succeeds) fetched the per-agent reputation score via the vendored-only GET /v1/registry/reputation/score?amid=X path. AGT registry has no such endpoint — the score is embedded as 'reputation_score: f64' in the per-agent record returned by /v1/agents/{did}. Provider-dispatch the URL; for AGT, wrap the agent record in a vendored-shaped payload (score / tier / raw) so downstream CLI fetchers and the operator panel stay schema-agnostic. cargo check: clean agt_governance_integration: 26/26 pass Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh): auto-tick AGT MeshClient sendHeartbeat every 30s The AGT Python relay (agentmesh/relay/app.py) marks any connection stale after OFFLINE_THRESHOLD = 90s without a 'heartbeat' frame, then routes subsequent messages for that DID to its OFFLINE STORE instead of live delivery. Stored frames are only replayed on (re)connect via _deliver_pending — so a long-lived parent that never reconnects loses every reply that arrives more than 90s after it last connected. The AGT MeshClient exposes sendHeartbeat() but never auto-schedules it. Vendored mode worked despite the same gap because the vendored Rust relay has no time-based stale check (only checks broken channels). For AGT mode we run our own 30s ticker (matches relay's HEARTBEAT_INTERVAL constant) inside AgtTransport.connect() and tear it down in disconnect(). The ticker is .unref()'d so it doesn't keep the Node event loop alive on its own. Reproduces deterministically when a sub-agent's reply lands >90s after the parent's connect timestamp: parent connect t=0 parent sends t=t1 (<90s) -> messages_routed += 1 child sends reply t=t2 (>90s) -> stored offline, never delivered relay /health: messages_delivered=0 (forever) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * runtime: hide Foundry tools in github-copilot mode (same as github-models) The Foundry tool catalog only makes sense when there is a real Azure Foundry project bound to the sandbox. Both GH-token providers (github-models, github-copilot) talk to GitHub-hosted models directly and have no Foundry project — exposing the 6 foundry_* tools just burns context with verbose JSON-schema and tempts the model to call endpoints the router will 404. Three call-sites were checking the provider: 1. agt-task-tools.ts:getTaskTools() — was `provider === "github-models"`, now matches either GH-token provider. The DuckDuckGo-backed web_search + memory fallbacks are appended in both modes. 2. agt-task-loop.ts:slim — was `provider === "github-models"`. Drives the prompt's tool-block descriptions and the slim 'Mode note' so the sub-agent sees the same tool catalog the LLM was given. Mode-note string adjusted to identify which provider is active. 3. runtimes/openclaw/src/index.ts — parent-side foundry tool registration in github-copilot mode. Was registering the full Foundry catalog with no upstream to call. Sub-agent tools-array shrinks 11,859 → 9,478 chars (~595 tokens saved per request) in github-copilot mode, and the 6 dead-end foundry_* tools no longer appear as options. Tests: runtimes/openclaw 118/118 pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(push): --mesh-provider=agt builds AGT relay/registry + swaps manifest Phase B.1 of the AGT-on-AKS rollout (see session plan files/agt-aks-end-to-end-plan.md). `azureclaw push` now mirrors the existing `azureclaw dev --mesh-provider` flag so the same provider selection works for AKS pushes. When --mesh-provider=agt: * Builds relay+registry from the AGT upstream Dockerfile ($AZURECLAW_AGT_REPO/agent-governance-python/agent-mesh/docker/Dockerfile) using COMPONENT=relay|registry build-args (matches dev.ts). * Tags as agentmesh-{relay,registry}-agt:latest so both vendored and AGT images can coexist on the same ACR and so the existing deploy/agentmesh-agt.yaml manifest picks them up unchanged. * Stages the AGT SDK tarball (--agt-sdk-tarball or auto-discovered in $agtRepo/agent-governance-typescript/microsoft-agent-governance-sdk-*.tgz) into .agt-sdk/ and passes AGT_SDK_TARBALL build-arg. * Always passes MESH_PROVIDER build-arg to the sandbox image so the Dockerfile's conditional `npm install @microsoft/agent-governance-sdk` runs for AGT clusters. When --apply --mesh-provider=agt: deletes deploy/agentmesh.yaml, applies deploy/agentmesh-agt.yaml, helm-upgrades with mesh.provider=agt, THEN rolls the controller (so the new pod reads AZURECLAW_MESH_PROVIDER=agt for new sandboxes). Auto-reverses when --apply --mesh-provider=vendored runs against a cluster currently on AGT (no Postgres deployment in the agentmesh ns). The image build loop also now supports absolute Dockerfile paths and absolute build contexts via a new `absoluteContext` field, needed because the AGT Dockerfile lives outside the azureclaw repo root. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(mesh): add 'azureclaw mesh provider <vendored|agt>' live switch Phase B.2 of AGT-on-AKS. Lets a deployed cluster flip mesh stacks without rebuilding any images, assuming both image pairs were already seeded by 'azureclaw push'. Flow: 1. Detect current provider via 'kubectl get deploy/postgres -n agentmesh' (vendored has Postgres, AGT does not). 2. kubectl delete -f deploy/agentmesh-<current>.yaml --ignore-not-found 3. kubectl apply -f deploy/agentmesh-<target>.yaml 4. helm upgrade azureclaw --reuse-values --set mesh.provider=<target> 5. kubectl rollout restart deploy/azureclaw-controller 6. With --restart-sandboxes: roll every azureclaw-managed Deployment so existing pods pick up the new AZURECLAW_MESH_PROVIDER value. Service names and ports are identical between the two manifests (agentmesh-relay:8765, agentmesh-registry:8080) so the controller's mesh_peer talks to either stack with no further config — the relay/ registry URLs already come from env vars (MESH_RELAY_URL / MESH_REGISTRY_URL). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(up): --mesh-provider=agt picks AGT manifest + flips helm value Phase B.3 of AGT-on-AKS. Adds -m/--mesh-provider to 'azureclaw up' so first-time deploys can ship AGT instead of vendored. When --mesh-provider=agt: * helm install runs with --set mesh.provider=agt (controller env AZURECLAW_MESH_PROVIDER=agt propagates to sandboxes). * deployAgentMesh() applies deploy/agentmesh-agt.yaml instead of deploy/agentmesh.yaml. * Skips the postgres ACR import and the agentmesh-db-credentials secret creation (both unused by AGT — its registry is in-memory). * Uses a per-provider temp manifest filename (.tmp-agentmesh-agt.yaml vs .tmp-agentmesh.yaml) so concurrent provider switches don't collide. The deployAgentMesh signature gains a non-breaking 'meshProvider' option that defaults to 'vendored' (existing callers untouched). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): plumb --mesh-provider into local-k8s helm install Phase D piece: --mesh-provider on 'azureclaw dev --target local-k8s' now forwards through runLocalK8s() → helmInstall() as '--set mesh.provider=<value>', so the controller deployed into the kind cluster carries the matching AZURECLAW_MESH_PROVIDER env var and spawns sandboxes against the chosen mesh stack. NOTE: local-k8s does not yet deploy agentmesh-relay/registry at all (the plan notes this as a Phase 3 pre-req blocked on AGT upstream patches G1/G2/G5). This commit only handles the helm-value plumbing; adding actual relay/registry deploy to local-k8s will land once the AGT fixes are upstream so we can prove end-to-end mesh roundtrip on local kind. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(controller): AGT wire protocol adapter for mesh_peer Implement full AGT relay/registry wire support in the controller's mesh_peer so cloud-offload works when AZURECLAW_MESH_PROVIDER=agt. Without this the controller's federation peer cannot connect to the AGT relay (different WS path, frame envelope, heartbeat, ack model) or the AGT registry (different HTTP shape, no signed body), and the leader fails-loops on AGT clusters — breaking the only cloud-offload control path. New module `mesh_peer/agt_wire.rs`: - `AgtFrame` enum (Connect/Message/Ack/Heartbeat/Disconnect/Error) with `#[serde(tag="type", rename_all="snake_case")]` matching `agentmesh/relay/app.py`. - `AgtRegisterAgentRequest` struct for `POST /v1/agents`. - 7 unit tests pinning the serialized shape. `mesh_peer/mod.rs`: - New `Provider` enum + `Provider::from_env()` selecting vendored (default) or AGT off `AZURECLAW_MESH_PROVIDER`. - `MeshPeerState.provider` carried through outbound + inbound paths. - `register_with_registry()` branches: vendored signs ts body; AGT posts `{did, public_key (base64url), capabilities, metadata}` with no signature; 409 treated as success for leader-failover idempotency. - `agt_did_for_identity()` derives `did:agentmesh:<base64url(pk)>` (matches JS SDK `buildDid`), so every leader replica converges on the same DID without coordination. - Default `MESH_RELAY_URL` appends `/ws` for AGT. - `connect_and_listen()`: - AGT connect frame `{type:"connect", from:<did>, token?:<env>}` (token read from `AGENTMESH_RELAY_TOKEN` if set). - AGT has no `Connected` ack — mark `connected=true` immediately. - Keepalive: AGT sends `{type:"heartbeat"}` every 30s (vendored keeps `ping`). - `serialize_and_send_outbound()` / `send_to_peer()` now take `state` and branch outbound framing — AGT emits `message` frames `{type, to, from, id, payload}` with `new_msg_id()` (16-byte hex). - `handle_message()` dispatches to `handle_vendored_frame()` or `handle_agt_frame()`. AGT path: - Parses `AgtFrame`, dispatches `Message` to `handle_peer_message()`. - Sends `Ack` reply (required — without it AGT redelivers on reconnect → duplicate offload processing). - Treats `Error` frames mentioning Authentication failed / Missing 'from' / session_replaced as fatal — drops connection for reconnect. `mesh_peer/offload.rs`: - All 8 `send_to_peer(...)` call sites updated to pass `&state` first. `main.rs`: - Remove the temporary AGT-skip guard around `mesh_peer::run`. The peer now starts unconditionally when enabled; provider is consumed inside `mesh_peer::run`. Build/test: - cargo build --release --package azureclaw-controller: OK - cargo test --package azureclaw-controller: 492 passed - cargo clippy --package azureclaw-controller --all-targets -D warnings: OK Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(ci): rustfmt + mesh-plugin fake-client establishSessionWithPeer - cargo fmt --all (controller/agt_wire.rs, mesh_peer/mod.rs, inference-router/governance.rs). - mesh-plugin agt-transport.test.ts: add `establishSessionWithPeer` to FakeClient interface + mock — pre-existing test gap exposed by the post-606f5b0 send path that calls it before send(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(deploy): AGT mesh probe path + Cilium pod-port NP allow deploy/agentmesh-agt.yaml: AGT FastAPI exposes /health, not /healthz (see agent-mesh/.../{registry,relay}/app.py). Liveness/readiness probes were 404'ing → CrashLoopBackOff/NotReady. operator-default-deny-networkpolicy.yaml: AKS Cilium dataplane evaluates NetworkPolicy egress against the backend pod port (post-DNAT), not the Service port. AGT registry/relay listen on 8082/8083; the Service maps 8080->8082 and 8765->8083 so the Service-port allowlist (8080/8765) doesn't actually permit the post-DNAT flow. Add 8082/8083 alongside so both vendored (8080/8765 direct) and AGT paths work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh promote): AGT-compat health + WS upgrade paths azureclaw mesh promote ran post-promote health checks against vendored-only paths and would 404 on AGT clusters: - Registry probe hit /v1/health. AGT only exposes /health (vendored exposes both). Probe /health first, fall back to /v1/health for vendored compatibility with older deployments that may have only served the /v1/ alias. - Relay WebSocket upgrade was attempted on /. AGT only serves WS on /ws (vendored uses /). Try /ws first, fall back to /. - 'Test: curl' hint pointed at /v1/health — also updated to /health so the suggested command works on both providers. Verified live against AGT cluster: Registry healthy (agentmesh-registry) Relay healthy (WebSocket upgrade on localhost:19991/ws) 640 CLI tests still pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(dev): first-run picker for local vs remote mesh source azureclaw dev now asks new users where the mesh should live, just like the existing inference-provider picker: Where should the mesh live? ❯ Local (recommended; spin up relay + registry in Docker) Remote (auto port-forward to AKS cluster: <cluster-name>) Local (default) keeps the existing behaviour: docker-compose'd relay/registry/postgres on the user's laptop. Remote (advanced) federates with a previously-provisioned AKS mesh: - If ~/.azureclaw/context.json has a cached globalRegistryUrl from a prior 'azureclaw mesh promote', reuse it verbatim. - Otherwise default to http://localhost:18080 — the port-forward URL 'mesh promote --port-forward' uses — so the auto-promote fallback in the downstream global-registry block will spawn the tunnels on demand. - If there is no aksCluster in context at all, warn and fall back to local so the user isn't left with a broken sandbox. Skipped entirely when --global-registry was passed explicitly (the advanced flag overrides the prompt) or when the user is past their first run. Also fixed a latent AGT-compat bug in the same flow: the existing 'auto-promote' path probed only /v1/health, which 404s on AGT clusters. Replaced with a /health → /v1/health fallback (matches the same shape we used in checkRegistryHealth last commit). Verified: - npm run build / typecheck clean - 640 CLI tests pass (2 skipped, no regressions) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(controller): propagate AZURECLAW_MESH_PROVIDER to router container On AKS the inference-router runs as a separate sidecar with its own env array, unlike local docker where it shares the openclaw container's env. The router's mesh code paths read AZURECLAW_MESH_PROVIDER to decide whether to upgrade the relay WS on `/` (vendored) or `/ws` (AGT), and likewise for the registry discover endpoint. The controller was only injecting the var into the openclaw container, so on AGT clusters the router defaulted to vendored and got 403 Forbidden in a tight reconnect loop against the AGT FastAPI relay. Push the same normalized provider value into router_agt_env (which is extended into router_env) so both containers agree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt): resolve 'parent' alias for spawned sub-agents on AGT mesh Sub-agent LLMs routinely call mesh_send(to_agent="parent") to reply back to their spawner, but on AGT the registry has no agent named or capability="parent" — the search returns 0 → no prekey bundle → send fails. The vendored runtime had this aliased only in the offload-mode task loop (agt-task-loop.ts), gated on $PARENT_SANDBOX, which the controller never set for AKS-spawned children. Two coordinated fixes: 1. controller/src/reconciler/mod.rs: when AGT_TRUSTED_PEERS is set (spawner seeds 'parent_name:parent_AMID' as the first entry), also push PARENT_SANDBOX=<first_name> into the openclaw container env. 2. runtimes/openclaw/src/core/agt-tools/agt.ts: in azureclaw_mesh_send and azureclaw_mesh_transfer_file, alias to_agent=='parent' → PARENT_SANDBOX || Symbol.for('agt-parent-name') before the registry lookup. The Symbol is set during runtime init from AGT_TRUSTED_PEERS[0], so this works even on images built before fix #1 lands. Skip in offload mode — 'parent' there is a protocol-level routing token, not a mesh recipient name. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop bogus establishSessionWithPeer() pre-bootstrap mesh-plugin/src/agt-transport.ts.send() called this.client.establishSessionWithPeer(toAmid) before forwarding to client.send(). That method does not exist on AgentMeshClient — the real method is establishSession(toAmid, options) — so every parent → sub-agent send on AGT was failing with: establishSessionWithPeer is not a function It was also unnecessary: AgentMeshClient.send() already auto-bootstraps the X3DH handshake on first contact (see @agentmesh/sdk AgentMeshClient.send → cache miss → establishSession() fallthrough at dist/index.js:3321-3334). Calling establishSession() ourselves would also be wrong because it is not idempotent — it unconditionally writes activeSessions.set and starts a fresh X3DH. Fix: remove the pre-bootstrap entirely and let client.send() manage session lifecycle. The AgtSdkModule type loses the required establishSessionWithPeer member (now optional) since we no longer depend on it; test fakes remain valid as harmless extras. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert 'drop establishSessionWithPeer pre-bootstrap' — was correct call Previous commit 34662c7 wrongly removed the establishSessionWithPeer() pre-bootstrap in mesh-plugin/agt-transport.ts based on a misread of the upstream @agentmesh/sdk API surface. The mesh-plugin actually loads @microsoft/agent-governance-sdk (see loadAgtSdk(), package.json pinned to ^3.5.0), which: • exposes establishSessionWithPeer(peerId) at mesh-client.js L230 — a high-level helper that fetches the prekey bundle and runs X3DH+KNOCK, idempotent on cache-hit • does NOT auto-bootstrap in send(): the path at L341 explicitly throws 'No encrypted session with <peer>. Call establishSession() first.' when no SecureChannel exists yet Symptom of the bad fix: parent → sub-agent mesh_send failed with 'No encrypted session with <amid>. Call establishSession() first.' on every first contact post-rollout. Restoring the pre-bootstrap with the correct rationale documented and the SDK source citations. AgtSdkModule type keeps the method optional for forward-compat with SDKs that auto-bootstrap; the runtime call uses non-null assertion since AGT SDK 3.5.0 ships the method. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * push: auto-detect mesh provider from live helm release When running 'azureclaw push --only sandbox --apply' without an explicit --mesh-provider flag, the CLI silently defaulted to 'vendored'. On a cluster already flipped to AGT (mesh.provider=agt), this caused the sandbox build to skip staging the local AGT SDK tarball into .agt-sdk/ — npm would install the public @microsoft/agent-governance-sdk@3.5.0 which lacks establishSessionWithPeer/discover/registerSelf helpers. Result: parent throws 'this.client.establishSessionWithPeer is not a function' on every mesh send. Auto-detect by reading 'mesh.provider' from the live helm release and respect it when --mesh-provider was not passed on the command line. Explicit flag still wins. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * entrypoint: fail-open trust gate when running anonymous tier When AGT_SKIP_ENTRA=1 (operator intentionally disabled OAuth) or when the Entra token exchange exhausts its retries, every sandbox registers as anonymous tier with registry reputation score 0. The KNOCK trust gate compares (registry_score * 1000 + affinity_bonus) against AGT_TRUST_THRESHOLD, which defaults to 500. Without OAuth identity: - sibling-to-sibling KNOCKs get no parent-trust or spawner bonus - effectiveScore = 0 < 500 → KNOCK rejected - whole mesh appears 'blocked' even though discovery + X3DH succeed Trust scoring is meaningless without OAuth identity. When we know we're in anonymous-tier mode, force AGT_TRUST_THRESHOLD=0. Policy evaluation in onKnock still runs, and the SDK's X3DH still proves cryptographic identity end-to-end — we just stop using a meaningless score as a gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(runtime): restore foundry_* dispatcher branch in sub-agent task loop Commit 073e759 ("GitHub Copilot provider + Anthropic passthrough + multi-agent peer roster", 2026-05-08) refactored agt-task-loop.ts to add a `web_search` branch (DuckDuckGo for slim-mode) and a `memory` branch, but in doing so deleted the `} else if (fnName === "foundry_web_search" || foundry_code_execute || foundry_file_search) {` else-if opener and forgot to put it back after the memory branch closes. The result: the entire foundry_web_search / foundry_code_execute / foundry_file_search dispatch block (lines 333-548) got silently nested INSIDE the memory branch — only reachable when `fnName === "memory"`, in which case none of its inner `fnName === "foundry_*"` checks match. Dead code. Symptom from this morning's demo: sub-agents calling foundry_web_search fell through every else-if and hit the final `echo 'no command'` exec fallback, returning the literal string "no command" — which the model then dutifully reported as "Foundry web search returned no command" in a loop. Parent agent was unaffected because the parent's foundry tools go through openclaw's plugin `registerTool` (agt-tools/foundry.ts:427), not the sub-agent dispatcher. That's why foundry_web_search "always worked" for the user — the parent path is a totally different code path. Fix: add back the missing else-if opener between the memory branch close and the existing foundry_* body. tsc clean. The dispatcher chain is now: file_write → http_fetch → web_search → memory → foundry_web_search → foundry_download_file → foundry_memory → foundry_image_generation → mesh_send → mesh_transfer_file → discover → mesh_inbox → mesh_await → exec_command fallback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(agt-mesh): ping registry /heartbeat every 30s to stay discoverable The AGT registry has no autonomous presence model — `last_seen` is frozen at registration and the `update_last_seen()` store method is dead code with no HTTP handler calling it. Combined with the openclaw discover tool's 90s stale filter (agt-tools/agt.ts STALE_AFTER_MS), every alive sub-agent goes silently invisible 90s after spawn, breaking sibling-to-sibling peer discovery. Demo symptom: analyst/viz/writer all reported 'peer discovery did not return ...' even though mesh_send to those names succeeded with 'delivered_and_replied'. The relay was fine; only the registry's presence view was stale. Pair with the corresponding upstream registry change (AGT branch `azureclaw-meshclient-event-hooks`, commit adds POST /v1/agents/{did}/heartbeat -> store.update_last_seen). The new tick reuses the existing 30s relay-keepalive timer in connect(), so no extra timers and no extra event-loop pressure. Best-effort: 4xx/5xx are warned-once, network errors swallowed, loop survives a registry pod restart (next tick retries). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * feat(strict-tools): opt-in OpenAI strict-mode + file-first transport hardening Adds AZURECLAW_STRICT_TOOLS gate, defaulted OFF. When enabled the runtime emits strict-conformant tool schemas (additionalProperties:false, all-required, nullable optionals) for 15 of 16 task-loop tools. Skipped automatically when slim-mode is active or the active model is non-OpenAI (Claude/Gemini/etc.) via a regex allowlist on AZURECLAW_MODEL || OPENCLAW_MODEL || OPENAI_MODEL. Strict-eligible (zero refactor): exec_command, file_write, foundry_web_search, foundry_code_execute, foundry_memory, foundry_file_search, mesh_send. Strict via STRICT_SCHEMA_OVERRIDES (nullable refactor): mesh_transfer_file, mesh_inbox, mesh_await, discover, foundry_image_generation, foundry_download_file, web_search, memory. Skipped (free-form schema): http_fetch (variable headers object). Plumbing: - runtimes/openclaw/src/core/agt-task-tools.ts: STRICT_ELIGIBLE set, STRICT_SCHEMA_OVERRIDES map, applyStrict() helper, model-allowlist gate. - runtimes/openclaw/src/core/agt-task-loop.ts: file-first transport hard-rule in sub-agent prompt, parse-error hint pointing to foundry_code_execute → json.dump → mesh_transfer_file, boot observability log. - runtimes/openclaw/src/core/agt-tools/agt.ts: tool-call argument resilience (matches new prompt guidance). - controller/src/reconciler/mod.rs: propagate AZURECLAW_STRICT_TOOLS into openclaw container env when enabled on controller. - deploy/helm/azureclaw/values.yaml: strictTools.enabled: false (default). - deploy/helm/azureclaw/templates/controller-deployment.yaml: conditional env injection block. CodeQL hardening (pre-existing alerts on this branch): - mesh-plugin/src/agt-transport.ts: log error class instead of full message to avoid clear-text-logging-of-sensitive-information. - cli/src/commands/dev.ts: validate --global-registry URL scheme before fetch to satisfy js/file-access-to-http. Verified live on demoagtmesh + analyst/viz/writer with file-first prompt fix alone (no strict): writer pushed 191KB request bodies through gpt-5.4 with zero tool-call parse failures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(mesh-plugin): drop toAmid from establishSessionWithPeer error log CodeQL js/clear-text-logging was still flagging the truncated toAmid prefix as taint from process.env. Log only a fixed string + error class; full error preserved on throw so caller's /prekey/i matcher still works. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Pal Lakatos-Toth <palakatosth@microsoft.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Wire the runtime through
createMeshTransport()so we can flip between the vendored@agentmesh/sdk(default) and Microsoft's@microsoft/agent-governance-sdkviaAZURECLAW_MESH_PROVIDERwithout touching any caller code.Builds on PR #244 (the factory scaffold) by closing the API gap between the two SDKs and refactoring the runtime to use the factory.
What changed
IMeshTransport now requires (both adapters implement)
lookup(amid)— registry RPCsubmitReputation(toAmid, sessionId, score, tags)— registry RPCenableKnockEnforcement()— vendored toggle / no-op on AGTonError(kind, fromAmid, detail)— diagnosticonE2EVerified(peerAmid, isFirstPeer)— diagnosticonDisconnect(reason, code)— diagnosticVendored adapter (mesh-plugin/src/connection.ts)
AgentMeshClientconnect()AGT adapter (mesh-plugin/src/agt-transport.ts)
lookup/submitReputation→ REST against the registry directly. AGT'sMeshClientis pure transport; registry RPCs belong on a separate client (out of scope for AGT upstream).enableKnockEnforcement→ no-op (AGT MeshClient always enforces)on{Error,Disconnect,E2EVerified}methodsAGT upstream changes (LOCAL branch only — NOT pushed)
Branch
azureclaw-meshclient-event-hookson/Users/pallakatos/Private/Repos/agt/agent-governance-toolkitadds the 3 event hooks + 8 tests. The AGT team owns the upstream PR — we hold this local until they're ready.Until that ships in a published
@microsoft/agent-governance-sdkrelease, our adapter's optional-chain calls (client.onError?.(...)) make the hooks no-ops on the published 3.5.0. Provider stays functional, just without the diagnostic hooks.Runtime (runtimes/openclaw/src/index.ts)
@azureclaw/meshasfile:../../mesh-pluginnew sdk.AgentMeshClient(...)withawait createMeshTransport({...})whenAZURECLAW_MESH_PROVIDER=agt; falls back to vendored on any other value (typo, empty, unset)toData(), then shared with both → same AMID either wayTests
Docs
docs/agt-vs-vendored-sdk.md— exhaustive side-by-side analysis covering every functional surface (identity, policy, trust, audit, mesh-client, registry, relay, X3DH, ratchet, KNOCK, plaintext peers, file transfer) + wiring + 4-phase migration plan.Migration path
agtin sandbox image; soak in dev with cross-provider parent↔child interopvendor/agentmesh-sdk/entirely (5 protocol patches no longer needed) and remove the env-var toggleReviewer notes
optionalDependencies) so pods built without AGT installed still boot on defaultcc @Azure/azureclaw-maintainers