Skip to content

feat(mesh): azureclaw mesh setup-trust — provision api://agentmesh app reg - #218

Merged
Pal Lakatos-Toth (pallakatos) merged 2 commits into
devfrom
mesh-setup-trust
May 5, 2026
Merged

Pal Lakatos-Toth (pallakatos) merged 2 commits into
devfrom
mesh-setup-trust

Conversation

@pallakatos

Copy link
Copy Markdown
Collaborator

Why

Audit item #3 from the pre-launch weakness list — until this PR, AzureClaw documented the three az calls needed to provision the tenant-wide api://agentmesh Entra app registration but did not help an operator run them. Without that registration every sandbox falls back to the AGT anonymous tier (trust score 0).

What

One new idempotent CLI subcommand:

azureclaw mesh setup-trust [--display-name <name>] [--dry-run]
  • Verifies the az session, prints tenant + subscription + user before mutating anything
  • Queries Graph for an existing api://agentmesh app reg — if found, just prints IDs and exits
  • Handles the half-finished case (app reg exists, SP missing) by creating only the SP
  • Otherwise creates app reg (AzureADMyOrg audience) + SP
  • Prints tenant ID, client ID, and the revert command

Why this resolves the issue without code in the cluster

The sandbox entrypoint already has the verified-tier code path — it requests a token for api://agentmesh/.default via Workload Identity → federated client assertion. That call is what currently returns AADSTS500011. Once the app reg + SP exist in the tenant, the same code path succeeds and AGT_OAUTH_TOKEN gets exported to AGT. Zero controller change, zero Helm upgrade, zero pod-image rebuild.

Error handling

Three real failure modes are surfaced with specific guidance, not just the raw az error:

  • AADSTS530084 (conditional access on Graph) — suggests az login --scope https://graph.microsoft.com//.default
  • Authorization_RequestDenied / Insufficient privileges — Application Administrator requirement
  • generic Graph 403 — Directory.Read.All hint

Tests

3 new structural tests in cli/src/commands/mesh.test.ts: 30 / 30 mesh tests pass.

Manual verification

azureclaw mesh setup-trust --dry-run against a corporate tenant with conditional access correctly identifies the tenant, prints the user, and surfaces the AADSTS530084 message with the right re-auth hint.

Out of scope

  • Auto-running this from azureclaw up preflight (Option B from the audit table — natural follow-up).
  • Bicep deployment-script alternative (Option C — only useful for tenants where CLI access is restricted).

Pal Lakatos-Toth and others added 2 commits May 5, 2026 11:48
…app reg

Closes the gap from the pre-launch audit: until now, AzureClaw
documented the manual `az ad app create` steps a tenant admin had to
run for sandboxes to register as the AGT *verified* tier, but did
not actually help the operator run them. Without that registration,
every sandbox falls back to the *anonymous* tier (trust score 0) and
peer KNOCKs are gated against that floor.

This adds a single idempotent CLI command:

    azureclaw mesh setup-trust [--display-name <name>] [--dry-run]

Behaviour:

  1. Confirms `az` is signed in and prints tenant + subscription +
     user, so the operator sees exactly which tenant they're about
     to mutate before anything is created.
  2. Queries Microsoft Graph (`az ad app list --identifier-uri
     api://agentmesh`). If the app reg already exists, prints its
     IDs and exits cleanly — re-running this command on a tenant
     that's already set up is safe.
  3. If the app reg exists but the service principal does not (a
     half-finished previous run), creates only the SP.
  4. Otherwise creates both: app reg with identifier URI
     api://agentmesh, sign-in audience AzureADMyOrg (single-tenant),
     then the matching SP.
  5. Prints the tenant ID + client ID + a one-line revert command
     (`az ad app delete --id <id>`).

Error handling specifically calls out the three real-world failure
modes:

  - AADSTS530084 (conditional access blocking the Graph token in
    this `az` session): suggests `az logout && az login --scope
    https://graph.microsoft.com//.default`.
  - Authorization_RequestDenied / Insufficient privileges: explains
    the Application Administrator requirement.
  - Generic Graph 403: hints at Directory.Read.All.

The verified-tier upgrade then happens automatically — no controller
restart, no Helm upgrade. The existing entrypoint Workload-Identity
→ Entra token-exchange code path simply succeeds where it currently
sees AADSTS500011.

Updates `docs/permissions.md` to lead with the new CLI command and
keep the manual `az` snippet as the underlying-mechanics
explanation.

Tests: 3 new structural tests in cli/src/commands/mesh.test.ts —
30 mesh tests pass total (was 27).

Verified manually: `azureclaw mesh setup-trust --dry-run` against
a Microsoft tenant with conditional access correctly identifies the
tenant, surfaces the AADSTS530084 message verbatim, and prints the
specific re-auth guidance.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…le, example

Three additions to the existing mesh section in docs/cli-reference.md:
  - subcommand table row with one-line description
  - options table (--display-name, --dry-run) plus a note about the
    Application Administrator requirement and idempotency
  - example block entry between 'mesh auth' and 'mesh status'

The 27-top-level-command count at the top of the file is unchanged
because setup-trust is a subcommand of mesh, not a new top-level
command.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) pushed a commit that referenced this pull request May 5, 2026
Once PR #218 lands, the trust-tier staleness statement 'the controller
and CLI do not perform it for you' becomes wrong. Update the three
spots that talked about manual provisioning to point at the new CLI
helper instead, while keeping docs/permissions.md as the canonical
home for the underlying 'az ad app create' invocation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@pallakatos
Pal Lakatos-Toth (pallakatos) merged commit e51aa1c into dev May 5, 2026
20 checks passed
@pallakatos
Pal Lakatos-Toth (pallakatos) deleted the mesh-setup-trust branch May 5, 2026 09:53
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 5, 2026
… tiers + vendored-fork workflow (#216)

* docs: launch readiness — known limitations, Az OpenAI prereq, trust tiers, vendored-fork workflow

Pre-launch documentation pass covering four soft spots flagged in the
weakness audit. No code changes; pure docs.

README.md
- New "Known limitations" section between "Project status" and
  "Contributing & support". Calls out: anonymous-tier mesh default
  pending api://agentmesh provisioning, multi-runtime images not yet
  published to a public registry, Semantic Kernel + MAF .NET CRD-wired
  but adapter-incomplete, attestation router-and-audit only, no managed
  service equivalent.
- Add a one-line callout under "Try it in five minutes" linking to the
  new Az OpenAI prereq snippet so first-time users without a deployment
  can self-serve.

docs/getting-started.md
- New "Don't have an Azure OpenAI deployment yet?" subsection under
  "Prerequisites" with three az commands (account create, deployment
  create, read endpoint+key) plus link to the official quickstart and
  pointer to azureclaw up for the Foundry-managed alternative.

docs/security.md
- New "Trust tiers and the api://agentmesh prerequisite" subsection
  under Layer 8. Tier table (Anonymous 0 vs Verified 600), explanation
  of why fail-open is the default, three resolution paths (lower
  threshold / provision app reg / set AGT_SKIP_ENTRA=1), the three log
  lines an operator may see at sandbox start, and the framing that none
  of them are errors.

CONTRIBUTING.md
- New "Working with the vendored AgentMesh forks" section between the
  credentials secret convention and Pull Requests. Two ground rules
  (upstream PR first, no quiet rebases past the audit gate), the dist/
  overlay flow for the SDK, the Rust 1.94+ toolchain note for relay/
  registry, and the copyright-header exclusion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: prereq snippet uses Foundry (AIServices), not legacy OpenAI

The rest of the codebase (router, controller, CLI) integrates with
Azure AI Foundry — not standalone Azure OpenAI accounts. The prereq
snippet I added in the previous commit used --kind OpenAI, which is
inconsistent with how azureclaw up provisions things and with the
18 Foundry API groups the router actually proxies.

- docs/getting-started.md: rename subsection to "Don't have an Azure
  AI Foundry deployment yet?", switch --kind to AIServices, explain
  why (Content Safety, Memory Store, agents, the rest of the AI
  Services surface), point to the Foundry quickstart instead of the
  Azure OpenAI quickstart.
- docs/getting-started.md prereq table: clarify "Azure AI Foundry (or
  Azure OpenAI)" to match — local mode still accepts a bare AOAI
  endpoint, but Foundry is the recommended shape.
- README.md: update the cross-link anchor to match the renamed
  heading (#dont-have-an-azure-ai-foundry-deployment-yet).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: reference azureclaw mesh setup-trust as the canonical resolution

Once PR #218 lands, the trust-tier staleness statement 'the controller
and CLI do not perform it for you' becomes wrong. Update the three
spots that talked about manual provisioning to point at the new CLI
helper instead, while keeping docs/permissions.md as the canonical
home for the underlying 'az ad app create' invocation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
…app reg (#218)

* feat(mesh): `azureclaw mesh setup-trust` — provision api://agentmesh app reg

Closes the gap from the pre-launch audit: until now, AzureClaw
documented the manual `az ad app create` steps a tenant admin had to
run for sandboxes to register as the AGT *verified* tier, but did
not actually help the operator run them. Without that registration,
every sandbox falls back to the *anonymous* tier (trust score 0) and
peer KNOCKs are gated against that floor.

This adds a single idempotent CLI command:

    azureclaw mesh setup-trust [--display-name <name>] [--dry-run]

Behaviour:

  1. Confirms `az` is signed in and prints tenant + subscription +
     user, so the operator sees exactly which tenant they're about
     to mutate before anything is created.
  2. Queries Microsoft Graph (`az ad app list --identifier-uri
     api://agentmesh`). If the app reg already exists, prints its
     IDs and exits cleanly — re-running this command on a tenant
     that's already set up is safe.
  3. If the app reg exists but the service principal does not (a
     half-finished previous run), creates only the SP.
  4. Otherwise creates both: app reg with identifier URI
     api://agentmesh, sign-in audience AzureADMyOrg (single-tenant),
     then the matching SP.
  5. Prints the tenant ID + client ID + a one-line revert command
     (`az ad app delete --id <id>`).

Error handling specifically calls out the three real-world failure
modes:

  - AADSTS530084 (conditional access blocking the Graph token in
    this `az` session): suggests `az logout && az login --scope
    https://graph.microsoft.com//.default`.
  - Authorization_RequestDenied / Insufficient privileges: explains
    the Application Administrator requirement.
  - Generic Graph 403: hints at Directory.Read.All.

The verified-tier upgrade then happens automatically — no controller
restart, no Helm upgrade. The existing entrypoint Workload-Identity
→ Entra token-exchange code path simply succeeds where it currently
sees AADSTS500011.

Updates `docs/permissions.md` to lead with the new CLI command and
keep the manual `az` snippet as the underlying-mechanics
explanation.

Tests: 3 new structural tests in cli/src/commands/mesh.test.ts —
30 mesh tests pass total (was 27).

Verified manually: `azureclaw mesh setup-trust --dry-run` against
a Microsoft tenant with conditional access correctly identifies the
tenant, surfaces the AADSTS530084 message verbatim, and prints the
specific re-auth guidance.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs(cli-reference): add mesh setup-trust subcommand row, options table, example

Three additions to the existing mesh section in docs/cli-reference.md:
  - subcommand table row with one-line description
  - options table (--display-name, --dry-run) plus a note about the
    Application Administrator requirement and idempotency
  - example block entry between 'mesh auth' and 'mesh status'

The 27-top-level-command count at the top of the file is unchanged
because setup-trust is a subcommand of mesh, not a new top-level
command.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
… tiers + vendored-fork workflow (#216)

* docs: launch readiness — known limitations, Az OpenAI prereq, trust tiers, vendored-fork workflow

Pre-launch documentation pass covering four soft spots flagged in the
weakness audit. No code changes; pure docs.

README.md
- New "Known limitations" section between "Project status" and
  "Contributing & support". Calls out: anonymous-tier mesh default
  pending api://agentmesh provisioning, multi-runtime images not yet
  published to a public registry, Semantic Kernel + MAF .NET CRD-wired
  but adapter-incomplete, attestation router-and-audit only, no managed
  service equivalent.
- Add a one-line callout under "Try it in five minutes" linking to the
  new Az OpenAI prereq snippet so first-time users without a deployment
  can self-serve.

docs/getting-started.md
- New "Don't have an Azure OpenAI deployment yet?" subsection under
  "Prerequisites" with three az commands (account create, deployment
  create, read endpoint+key) plus link to the official quickstart and
  pointer to azureclaw up for the Foundry-managed alternative.

docs/security.md
- New "Trust tiers and the api://agentmesh prerequisite" subsection
  under Layer 8. Tier table (Anonymous 0 vs Verified 600), explanation
  of why fail-open is the default, three resolution paths (lower
  threshold / provision app reg / set AGT_SKIP_ENTRA=1), the three log
  lines an operator may see at sandbox start, and the framing that none
  of them are errors.

CONTRIBUTING.md
- New "Working with the vendored AgentMesh forks" section between the
  credentials secret convention and Pull Requests. Two ground rules
  (upstream PR first, no quiet rebases past the audit gate), the dist/
  overlay flow for the SDK, the Rust 1.94+ toolchain note for relay/
  registry, and the copyright-header exclusion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: prereq snippet uses Foundry (AIServices), not legacy OpenAI

The rest of the codebase (router, controller, CLI) integrates with
Azure AI Foundry — not standalone Azure OpenAI accounts. The prereq
snippet I added in the previous commit used --kind OpenAI, which is
inconsistent with how azureclaw up provisions things and with the
18 Foundry API groups the router actually proxies.

- docs/getting-started.md: rename subsection to "Don't have an Azure
  AI Foundry deployment yet?", switch --kind to AIServices, explain
  why (Content Safety, Memory Store, agents, the rest of the AI
  Services surface), point to the Foundry quickstart instead of the
  Azure OpenAI quickstart.
- docs/getting-started.md prereq table: clarify "Azure AI Foundry (or
  Azure OpenAI)" to match — local mode still accepts a bare AOAI
  endpoint, but Foundry is the recommended shape.
- README.md: update the cross-link anchor to match the renamed
  heading (#dont-have-an-azure-ai-foundry-deployment-yet).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: reference azureclaw mesh setup-trust as the canonical resolution

Once PR #218 lands, the trust-tier staleness statement 'the controller
and CLI do not perform it for you' becomes wrong. Update the three
spots that talked about manual provisioning to point at the new CLI
helper instead, while keeping docs/permissions.md as the canonical
home for the underlying 'az ad app create' invocation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant