Skip to content

build(deps): Bump github/codeql-action from 3 to 4 - #9

Merged
Imran Siddique (imran-siddique) merged 1 commit into
mainfrom
dependabot/github_actions/github/codeql-action-4
Mar 4, 2026
Merged

Imran Siddique (imran-siddique) merged 1 commit into
mainfrom
dependabot/github_actions/github/codeql-action-4

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Mar 4, 2026

Copy link
Copy Markdown
Contributor

Bumps github/codeql-action from 3 to 4.

Release notes

Sourced from github/codeql-action's releases.

v3.32.5

  • Repositories owned by an organization can now set up the github-codeql-disable-overlay custom repository property to disable improved incremental analysis for CodeQL. First, create a custom repository property with the name github-codeql-disable-overlay and the type "True/false" in the organization's settings. Then in the repository's settings, set this property to true to disable improved incremental analysis. For more information, see Managing custom properties for repositories in your organization. This feature is not yet available on GitHub Enterprise Server. #3507
  • Added an experimental change so that when improved incremental analysis fails on a runner — potentially due to insufficient disk space — the failure is recorded in the Actions cache so that subsequent runs will automatically skip improved incremental analysis until something changes (e.g. a larger runner is provisioned or a new CodeQL version is released). We expect to roll this change out to everyone in March. #3487
  • The minimum memory check for improved incremental analysis is now skipped for CodeQL 2.24.3 and later, which has reduced peak RAM usage. #3515
  • Reduced log levels for best-effort private package registry connection check failures to reduce noise from workflow annotations. #3516
  • Added an experimental change which lowers the minimum disk space requirement for improved incremental analysis, enabling it to run on standard GitHub Actions runners. We expect to roll this change out to everyone in March. #3498
  • Added an experimental change which allows the start-proxy action to resolve the CodeQL CLI version from feature flags instead of using the linked CLI bundle version. We expect to roll this change out to everyone in March. #3512
  • The previously experimental changes from versions 4.32.3, 4.32.4, 3.32.3 and 3.32.4 are now enabled by default. #3503, #3504

v3.32.4

  • Update default CodeQL bundle version to 2.24.2. #3493
  • Added an experimental change which improves how certificates are generated for the authentication proxy that is used by the CodeQL Action in Default Setup when private package registries are configured. This is expected to generate more widely compatible certificates and should have no impact on analyses which are working correctly already. We expect to roll this change out to everyone in February. #3473
  • When the CodeQL Action is run with debugging enabled in Default Setup and private package registries are configured, the "Setup proxy for registries" step will output additional diagnostic information that can be used for troubleshooting. #3486
  • Added a setting which allows the CodeQL Action to enable network debugging for Java programs. This will help GitHub staff support customers with troubleshooting issues in GitHub-managed CodeQL workflows, such as Default Setup. This setting can only be enabled by GitHub staff. #3485
  • Added a setting which enables GitHub-managed workflows, such as Default Setup, to use a nightly CodeQL CLI release instead of the latest, stable release that is used by default. This will help GitHub staff support customers whose analyses for a given repository or organization require early access to a change in an upcoming CodeQL CLI release. This setting can only be enabled by GitHub staff. #3484

v3.32.3

  • Added experimental support for testing connections to private package registries. This feature is not currently enabled for any analysis. In the future, it may be enabled by default for Default Setup. #3466

v3.32.2

  • Update default CodeQL bundle version to 2.24.1. #3460

v3.32.1

  • A warning is now shown in Default Setup workflow logs if a private package registry is configured using a GitHub Personal Access Token (PAT), but no username is configured. #3422
  • Fixed a bug which caused the CodeQL Action to fail when repository properties cannot successfully be retrieved. #3421

v3.32.0

  • Update default CodeQL bundle version to 2.24.0. #3425

v3.31.11

  • When running a Default Setup workflow with Actions debugging enabled, the CodeQL Action will now use more unique names when uploading logs from the Dependabot authentication proxy as workflow artifacts. This ensures that the artifact names do not clash between multiple jobs in a build matrix. #3409
  • Improved error handling throughout the CodeQL Action. #3415
  • Added experimental support for automatically excluding generated files from the analysis. This feature is not currently enabled for any analysis. In the future, it may be enabled by default for some GitHub-managed analyses. #3318
  • The changelog extracts that are included with releases of the CodeQL Action are now shorter to avoid duplicated information from appearing in Dependabot PRs. #3403

v3.31.10

CodeQL Action Changelog

See the releases page for the relevant changes to the CodeQL CLI and language packs.

3.31.10 - 12 Jan 2026

  • Update default CodeQL bundle version to 2.23.9. #3393

See the full CHANGELOG.md for more information.

v3.31.9

CodeQL Action Changelog

See the releases page for the relevant changes to the CodeQL CLI and language packs.

... (truncated)

Changelog

Sourced from github/codeql-action's changelog.

4.32.3 - 13 Feb 2026

  • Added experimental support for testing connections to private package registries. This feature is not currently enabled for any analysis. In the future, it may be enabled by default for Default Setup. #3466

4.32.2 - 05 Feb 2026

  • Update default CodeQL bundle version to 2.24.1. #3460

4.32.1 - 02 Feb 2026

  • A warning is now shown in Default Setup workflow logs if a private package registry is configured using a GitHub Personal Access Token (PAT), but no username is configured. #3422
  • Fixed a bug which caused the CodeQL Action to fail when repository properties cannot successfully be retrieved. #3421

4.32.0 - 26 Jan 2026

  • Update default CodeQL bundle version to 2.24.0. #3425

4.31.11 - 23 Jan 2026

  • When running a Default Setup workflow with Actions debugging enabled, the CodeQL Action will now use more unique names when uploading logs from the Dependabot authentication proxy as workflow artifacts. This ensures that the artifact names do not clash between multiple jobs in a build matrix. #3409
  • Improved error handling throughout the CodeQL Action. #3415
  • Added experimental support for automatically excluding generated files from the analysis. This feature is not currently enabled for any analysis. In the future, it may be enabled by default for some GitHub-managed analyses. #3318
  • The changelog extracts that are included with releases of the CodeQL Action are now shorter to avoid duplicated information from appearing in Dependabot PRs. #3403

4.31.10 - 12 Jan 2026

  • Update default CodeQL bundle version to 2.23.9. #3393

4.31.9 - 16 Dec 2025

No user facing changes.

4.31.8 - 11 Dec 2025

  • Update default CodeQL bundle version to 2.23.8. #3354

4.31.7 - 05 Dec 2025

  • Update default CodeQL bundle version to 2.23.7. #3343

4.31.6 - 01 Dec 2025

No user facing changes.

4.31.5 - 24 Nov 2025

  • Update default CodeQL bundle version to 2.23.6. #3321

4.31.4 - 18 Nov 2025

... (truncated)

Commits
  • ba1288c Merge branch 'main' into dependabot/npm_and_yarn/globals-17.3.0
  • 29765a3 Skip overlay memory check for CodeQL 2.24.3 and later
  • 068e80c Rebuild
  • 154969e Merge branch 'main' into dependabot/npm_and_yarn/npm-minor-e1092f1102
  • b0ed4de Merge pull request #3511 from github/henrymercer/merge-queue
  • 3c83f57 Merge pull request #3516 from github/mbg/start-proxy/reduce-connection-check-...
  • See full diff in compare view

Dependabot compatibility score

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Bumps [github/codeql-action](https://github.com/github/codeql-action) from 3 to 4.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@v3...v4)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added the dependencies Pull requests that update a dependency file label Mar 4, 2026
@imran-siddique
Imran Siddique (imran-siddique) merged commit 2a0fc50 into main Mar 4, 2026
17 of 19 checks passed
@dependabot
dependabot Bot deleted the dependabot/github_actions/github/codeql-action-4 branch March 4, 2026 22:34
Imran Siddique (imran-siddique) added a commit that referenced this pull request Apr 21, 2026
…eue, wsFactory (#1301)

High-level mesh client for the TS SDK, addressing three AzureClaw
compatibility requirements:

- plaintextPeers: bypass E2E encryption for legacy peers (Rust
  controller uses base64(JSON), not Signal). addPlaintextPeer/
  removePlaintextPeer/isPlaintextPeer API.
- wsFactory: custom WebSocket constructor hook for HTTPS_PROXY
  CONNECT tunneling (Node 22 global fetch/undici quirk).
- KNOCK pending queue: when a message arrives for a peer with an
  in-flight KNOCK, await resolution instead of rejecting. Fixes
  the race condition documented in vendored patch #5.

Also handles:
- Session reuse (returns existing session, no crash — patch #10)
- Buffer-based base64 (avoids stack overflow on >100KB — patch #9)
- Heartbeat sending

Clean-room: implements against Wire Protocol spec Sections 9, 10, 12.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Imran Siddique (imran-siddique) pushed a commit that referenced this pull request May 11, 2026
…ify, heartbeat (#2090)

* feat(mesh-client): add onError, onDisconnect, onE2EVerified event hooks

Adds three observer-registration methods to MeshClient to align with the
AzureClaw vendored AgentMesh SDK surface so consumers can swap providers
behind a single transport interface:

  onError(handler)        — fires on ws errors and decrypt failures
  onDisconnect(handler)   — fires on ws.close with reason+code
  onE2EVerified(handler)  — fires on first successful decrypt per peer

Pure additions; no behaviour change for existing flows. Decrypt-failure
and missing-session paths now drop the message AND notify error
observers instead of silently dropping (still safe — unhandled events
are swallowed).

Tests: 8 new in mesh-client-event-hooks.test.ts; full suite 387/387 green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-client): close gaps G1 (KNOCK auto-bootstrap) and G2 (auto-reconnect)

Two protocol-level gaps were identified during the AzureClaw vendored-SDK
audit (see Azure/kars#245 docs/agt-vs-vendored-sdk.md). Both block
moving fully off the vendored agentmesh-sdk fork onto upstream AGT.

## G1: receiver-side X3DH auto-bootstrap from KNOCK

Before: the responder side of an encrypted session was created only by
calling acceptSession(peerId, establishment) explicitly. But the
ChannelEstablishment was never carried on the wire — neither in
knock_accept nor in the first message frame — so the receiver had no way
to obtain it. Result: any fresh encrypted session failed with
"No encrypted session" the first time a ciphertext arrived.

Mirrors vendored agentmesh-sdk patch #4b. Behavior:

- Sender (establishSession): builds the SecureChannel first, then sends
  KNOCK with the ChannelEstablishment embedded as
  { ik: base64, ek: base64, otk?: number }.
- Receiver (handleKnock): if the knock contains `establishment` AND no
  prior session exists, deserialize and call acceptSession() automatically
  before responding with knock_accept. On bad establishment data, the
  knock is rejected and onError("knock", ...) fires.
- Backwards-compatible: legacy peers that don't embed establishment
  continue to work; caller must invoke acceptSession() manually as before.

## G2: auto-reconnect on transport drop

Before: MeshClient.reconnect() existed but was never called automatically.
Network blips, relay restarts, AKS node OOM-evicts left the client
disconnected forever — agents went mesh-deaf.

Mirrors vendored agentmesh-sdk patch #9. Behavior:

- New options: autoReconnect (default true), maxReconnectAttempts
  (default Number.POSITIVE_INFINITY), reconnectBaseDelayMs (default
  1000), reconnectMaxDelayMs (default 60000).
- On non-1000 ws.onclose: schedule reconnect with exponential backoff
  capped at reconnectMaxDelayMs. Light jitter (±20%) avoids
  thundering-herd reconnects across many sandboxes.
- onDisconnect handlers fire BEFORE the reconnect is scheduled so
  observers can see the drop.
- After maxReconnectAttempts, fires onError("ws", ..., "auto-reconnect
  gave up after N attempts") and stops.
- disconnect() cancels any pending reconnect timer.
- A connect() failure inside the reconnect path schedules another retry
  via the existing onclose path (and recursively from the catch handler
  in scheduleReconnect for the case where ws.onopen never fired).

## Tests

11 new tests across two files:
- tests/mesh-client-knock-bootstrap.test.ts: 5 tests covering sender-side
  embedding, receiver auto-bootstrap, legacy fallback, malformed
  establishment, and end-to-end happy-path encrypted send.
- tests/mesh-client-auto-reconnect.test.ts: 6 tests covering server-close
  reconnect, client-close no-reconnect, opt-out, give-up after max
  attempts, disconnect cancels pending timer, no duplicate scheduling.

Full AGT TS suite: 398/398 pass (was 387 before this commit). Build clean.

Both gaps were identified in the AzureClaw audit doc:
docs/agt-vs-vendored-sdk.md (Patch-by-patch audit section).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): close vendored-parity gaps G3, G4, G5

G3 — session-desync teardown (vendored agentmesh-sdk patch #13 equiv):
when Double Ratchet decryption fails for an existing session, the local
ratchet state is irrecoverable. Previously we only fired
onError('decrypt'); now we tear down the session (delete from
this.sessions; clear knockAccepted) and fire a distinct
onError('session_desync') so callers can re-run establishSession() and
resume communication.

G4 — pre-KNOCK encrypted-message buffer (vendored agentmesh-sdk
patch #16 equiv): under transport reordering the relay can deliver an
encrypted 'message' frame BEFORE the matching 'knock' arrives. Without
buffering, that first frame was dropped silently. Now MeshClient buffers
encrypted frames per-peer (default cap 5, TTL 3000ms), replays them
when the corresponding knock is accepted, and drops them when rejected
or on disconnect. Disabled by setting preKnockBufferSize: 0.

G5 — eager ghost-connection close on rebind (vendored agentmesh-relay
patch #2 equiv): when an agent reconnects with the same DID, the prior
'ghost' WebSocket is now closed eagerly with code 1000 'session_replaced'
instead of waiting for the 90s heartbeat-eviction timer. The 'finally'
cleanup now compares socket identity to avoid removing the freshly
rebound connection when the old socket's handler unwinds.

Tests: +5 (G4 pre-knock buffer behaviour), +2 (G3 type + lifecycle
safety), +1 (G5 ghost rebind). 405 TS tests pass; 18 relay tests pass.

NOTE: held local-only on branch azureclaw-meshclient-event-hooks; will
be coordinated upstream with the AGT team.

* feat(mesh): add RegistryClient and auto-register on MeshClient.connect()

Closes the structural gap where MeshClientOptions declared registryUrl
but no code path ever used it. Without this, agents can connect to the
relay and exchange ciphertext but are invisible to peers — they never
appear in /v1/discover and no one can fetch their X3DH pre-key bundle.

This commit adds the missing registry-side glue, in upstream-quality
shape, so AGT MeshClient becomes a drop-in replacement for vendored
forks that wired their own registry calls (e.g. AzureClaw vendor/agentmesh-sdk).

What's new:

- src/encryption/registry-client.ts: typed HTTP client wrapping the
  AgentMesh registry's REST surface (POST /v1/agents,
  PUT/GET /v1/agents/{did}/prekeys, GET /v1/discover, GET/DELETE
  /v1/agents/{did}). base64url <-> Uint8Array marshalling, optional
  Ed25519-Timestamp Authorization signer, configurable retry/timeout.

- src/encryption/mesh-client.ts:
  * New options: registryClient, registryClientOptions, capabilities,
    registrationMetadata, oneTimePrekeyCount, autoRegister.
  * connect() now calls registerSelf() at the end of a successful WS
    handshake when autoRegister is true (default) and a registry is
    configured. Idempotent — 409 (already registered) is treated as
    success; reconnect doesn't re-register.
  * registerSelf(): generates signed-prekey + N one-time prekeys via
    keyManager, then POSTs /v1/agents (capabilities = [displayName,
    ...options.capabilities]) and PUTs the prekey bundle.
  * discover(capability): wraps registry.discover.
  * establishSessionWithPeer(peerId): fetches peer prekeys from the
    registry and calls existing establishSession.
  * getRegistry(): exposes the underlying client for advanced cases.

- tests/registry-client.test.ts: 11 new tests covering wire format,
  base64url round-trip, idempotent register (409), 5xx retry, and
  Authorization header. Also covers MeshClient auto-register on connect
  end-to-end with a fake fetch + fake WebSocket.

- tests/mesh-client-*.test.ts: existing fixtures updated to pass
  autoRegister: false (no registry available in those tests). Behavior
  unchanged for the production path.

All 71 tests pass (60 previously + 11 new).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(registry): add POST /v1/agents/{did}/heartbeat to bump last_seen

The registry's last_seen field is set on registration but never
updated by any HTTP handler — update_last_seen() exists in
store.py but is dead code in production. As a result, every
registered agent looks 'offline' (online=false on /presence) 90s
after spawn, and any client-side discover stale-filter rejects
all live agents.

Add an unauthenticated POST /v1/agents/{did}/heartbeat that calls
store.update_last_seen(). Idempotent. Returns 404 for unknown DIDs
so callers can detect a registry restart and re-register.

Mirrors the agent-side ping cadence (30s, well within the 90s
presence threshold).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): include Ed25519 identity_key_ed in prekey bundle so verifyBundle() works

PreKeyBundle.identityKeyEd was previously aliased to the X25519 identity_key
because the registry didn't store the Ed25519 signing key. As a result,
verifyBundle() either passed only because the stub test never exercised it,
or quietly accepted any signed pre-key — a soundness gap in the X3DH wrapper.

This commit threads identityKeyEd through the full path:

- agent-mesh/registry/store.py (AgentRecord): add identity_key_ed field
  (X25519 long-term key already present as identity_key; Ed25519 distinct).
- agent-governance-typescript/src/encryption/x3dh.ts: expose
  X3DHKeyManager.identityKeyEd getter.
- agent-governance-typescript/src/encryption/registry-client.ts:
  uploadPrekeys() now accepts identityKeyEd (32 bytes, validated) and
  serializes it as identity_key_ed; fetchPrekeys() deserializes it back into
  PreKeyBundle.identityKeyEd. Older bundles without the field fall back to
  the X25519 key — verifyBundle() will then correctly reject (forensic
  visibility, no silent bypass).
- agent-governance-typescript/src/encryption/mesh-client.ts: registerSelf()
  passes keyManager.identityKeyEd through to the registry.
- tests/registry-client.test.ts + tests/test_registry.py: assert the new
  field is round-tripped on PUT/GET.

Wire compatibility: PUT now requires identity_key_ed; GET emits it when set.
Existing AGT clients that don't upload it will get verifyBundle() rejection
on the receiving side — peers must upgrade together. (No silent regression
on the wire — the field is new, not a breaking change to existing fields.)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test(mesh): set autoRegister:false in upstream knock + malformed-frame tests

These two upstream tests construct MeshClient without autoRegister:false,
which now triggers a real fetch to the registry on connect() (added by
our registry-client commit). The tests have no fake registry, so they
fail with 'fetch failed'. autoRegister:false skips the auto-registration
path and matches the pattern used in our 6 sibling mesh-client-*.test.ts
files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 29, 2026
…ently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 29, 2026
Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 29, 2026
…ently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 29, 2026
Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 30, 2026
…ently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 30, 2026
Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 30, 2026
…ently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Jack Batzner (jackbatzner) pushed a commit to jackbatzner/agent-governance-toolkit that referenced this pull request May 30, 2026
Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Imran Siddique (imran-siddique) pushed a commit that referenced this pull request May 30, 2026
…xecute API (#2644)

* fix(agent-os): close authorization bypasses in stateless kernel and execute API

Three same-class authorization fixes identified in security review:

1. stateless._check_policies: caller-supplied params['approved']=True no longer satisfies requires_approval gates. Approval must flow through the trusted IntentManager path; unplanned drift on restricted actions is now denied. The legacy flag is stripped from params before action execution.

2. server/app.py /api/v1/execute: caller-supplied agent_id is no longer trusted when authentication is bypassed. The legacy AGENT_OS_ALLOW_UNAUTHENTICATED_EXECUTE env var now raises ValueError at construction time. The replacement AGENT_OS_UNSAFE_ALLOW_UNAUTHENTICATED_EXECUTE is gated on AGENT_OS_ENV in {dev,development,local}; the server-side identity is fixed by AGENT_OS_UNSAFE_LOCAL_EXECUTE_AGENT_ID (default local-dev-agent); mismatched caller agent_id is rejected with 422 (unsafe) or 403 (authenticated).

3. mcp-kernel-server KernelExecuteTool._check_policies: same params.get('approved') bypass pattern as (1); now ignored with a warning log and the action is denied with guidance pointing to a trusted host approval workflow.

Tests added/updated for all three paths. Tangential sweep covered other auth surfaces (mcp_gateway approval callback, AGENT_OS_* env vars, REST endpoints) and found no further in-class bugs in agent-os core; module-level FastAPI surfaces in caas/iatp/observability are out of scope for this PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(mcp-scan): regression for env-poisoning RCE + cwd hijack -- currently FAILING

Red-team findings #1 + #2: mcp-scan CLI accepts arbitrary environment keys (LD_PRELOAD, PYTHONPATH, NODE_OPTIONS, ...) and untrusted cwd paths when launching subprocesses, enabling pre-exec code injection.

These regression tests assert the SECURE behavior (refusal). They FAIL on this commit because the helpers _blocked_command_env_keys and _validate_launch_cwd do not exist, proving the vuln surface is present.

Failure mode: 28 errors in TestLaunchEnvAndCwdGuards (AttributeError on missing helpers). Fix applied in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(mcp-scan): restore env-key blocklist and untrusted-cwd guard

Closes red-team findings #1 + #2. Restores _blocked_command_env_keys and _validate_launch_cwd helpers. Red->Green: 28 errors -> 129 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for approval-key bypasses + provider edge cases -- currently FAILING

Red-team findings #8 (confusable/nested approved keys bypass strip), #10 (non-strict-True provider return treated as allow), #11 (log injection via CR/LF in caller fields), #12 (provider BaseException leaks past approval check).

Failure mode: 15 failures across stateless + mcp_kernel_server.tools. Cyrillic 'approvеd', uppercased 'Approved', nested dict values, truthy-non-bool returns ('yes', 1, object), and SystemExit/KeyboardInterrupt all currently bypass the gate. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): harden approval-key strip, strict-bool, BaseException, log sanitization

Closes red-team #8, #10, #11, #12. NFKC + casefold approved-key match, recursive strip into nested dicts/lists, strict 'is True', except BaseException, _sanitize_log_field. Red->Green: 15 failed -> 141 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for empty-policies bypass + non-loopback execute -- currently FAILING

Red-team findings #3 (no policy match -> action allowed even when requires_approval declared elsewhere) and #5 (unsafe execute mode trusted from arbitrary remote peers).

Failure mode: test_execute_global_approval_blocks_empty_policy_list FAILS because StatelessKernel falls through to allow when no policy entry matches. test_execute_unsafe_escape_hatch_rejects_non_loopback_peer FAILS because _authenticate_execute_request does not inspect request.client. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): close empty-policies bypass and enforce loopback for unsafe execute

Closes #3 + #5. _globally_protected_actions enforced after per-policy loop; _is_loopback_client rejects non-127.x/::1 peers with 403. Red->Green: 2 failed -> 94 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(intent): regression for cross-agent intent reuse -- currently FAILING

Red-team finding #4: IntentManager.check_action does not verify that the caller's agent_id matches the intent's agent_id, so agent B can reuse agent A's stored intent record to perform privileged actions under A's policy context.

Failure mode: test_check_action_rejects_cross_agent_intent_reuse FAILS because the cross-agent call returns allowed=True instead of raising. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(intent): bind intent to declaring agent_id

Closes #4. Asserts intent.agent_id == caller agent_id in check_action. Red->Green: 1 failed -> 41 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(iatp): regression for weak/short trusted-override tokens -- currently FAILING

Red-team finding #9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(iatp): reject weak/short trusted-override tokens

Closes #9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(policies): regression for plaintext OPA over network -- currently FAILING

Red-team finding #7: OPABackend remote mode follows http:// URLs to non-loopback hosts without warning. An on-path attacker on the OPA route flips allow=true and the kernel approves any action.

Failure mode: test_plaintext_remote_non_loopback_denied and test_plaintext_opt_in_without_local_env_denied FAIL because _evaluate_remote performs the HTTP call without protocol gating. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(policies): require HTTPS for remote OPA unless explicitly opted in

Closes #7. _evaluate_remote rejects non-HTTPS unless loopback host OR (AGENT_OS_OPA_ALLOW_PLAINTEXT=1 + AGENT_OS_ENV in {local,dev,development}). Plaintext non-loopback returns error='plaintext_opa_blocked'. Red->Green: 2 failed -> 77 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(caas): regression for unauthenticated FastAPI surface gate -- currently FAILING

Red-team finding #6: caas.api.server only LOGS a warning when started outside local env; misconfigured deployment exposes every CaaS route silently.

Failure mode: 13 failures because _caas_unauth_gate_satisfied does not exist and startup hook does not raise. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(caas): require explicit env gate to start unauthenticated CaaS surface

Closes #6. Startup hook raises RuntimeError unless AGENT_OS_ENV in {local,dev,development} OR CAAS_UNSAFE_ALLOW_UNAUTH=1. Red->Green: 13 failed -> 13 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(agent-os): clear no-stubs/no-crypto/spell-check/safety-critical CI gates

- Reword TODO(security) doc comments to 'Future hardening (security)' in caas/api/server.py, iatp/main.py (x2 including proxy_task cross-ref), iatp/sidecar/__init__.py so the no-stubs CI gate accepts the docs without losing the design-followup intent.

- Replace inline 'import hmac; hmac.compare_digest' with 'import secrets; secrets.compare_digest' in iatp/main.py so the no-custom-crypto CI gate is happy (secrets.compare_digest is the stdlib re-export of hmac.compare_digest, same constant-time guarantee).

- Add 19 project-specific terms to .cspell-repo-terms.txt (ASGI, NFKC, casefold, confusables, multitenant, normalisation, sanitised, unicodedata, testclient, monkeypatched, baseexception, rsplit, hdrs, oncall, madmin, backendunavailable, changeme, shortone, approv) for the spell-check-changed-files job.

- Update tests/test_safety_critical.py::TestPolicyEdgeCases::test_empty_policies_list_allows to reflect the new fail-closed behavior from fix #3: an empty policies list must DENY requires_approval actions (file_write). Renamed to test_empty_policies_list_denies_protected_actions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(spell-check): allow cyrillic-e 'approv\u0435d' confusable used in unicode normalization tests

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

---------

Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
MohammadHaroonAbuomar pushed a commit to MohammadHaroonAbuomar/agt-acs that referenced this pull request Jun 1, 2026
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 3 to 4.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@v3...v4)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
MohammadHaroonAbuomar pushed a commit to MohammadHaroonAbuomar/agt-acs that referenced this pull request Jun 1, 2026
…eue, wsFactory (microsoft#1301)

High-level mesh client for the TS SDK, addressing three AzureClaw
compatibility requirements:

- plaintextPeers: bypass E2E encryption for legacy peers (Rust
  controller uses base64(JSON), not Signal). addPlaintextPeer/
  removePlaintextPeer/isPlaintextPeer API.
- wsFactory: custom WebSocket constructor hook for HTTPS_PROXY
  CONNECT tunneling (Node 22 global fetch/undici quirk).
- KNOCK pending queue: when a message arrives for a peer with an
  in-flight KNOCK, await resolution instead of rejecting. Fixes
  the race condition documented in vendored patch microsoft#5.

Also handles:
- Session reuse (returns existing session, no crash — patch microsoft#10)
- Buffer-based base64 (avoids stack overflow on >100KB — patch microsoft#9)
- Heartbeat sending

Clean-room: implements against Wire Protocol spec Sections 9, 10, 12.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
MohammadHaroonAbuomar pushed a commit to MohammadHaroonAbuomar/agt-acs that referenced this pull request Jun 1, 2026
…ify, heartbeat (microsoft#2090)

* feat(mesh-client): add onError, onDisconnect, onE2EVerified event hooks

Adds three observer-registration methods to MeshClient to align with the
AzureClaw vendored AgentMesh SDK surface so consumers can swap providers
behind a single transport interface:

  onError(handler)        — fires on ws errors and decrypt failures
  onDisconnect(handler)   — fires on ws.close with reason+code
  onE2EVerified(handler)  — fires on first successful decrypt per peer

Pure additions; no behaviour change for existing flows. Decrypt-failure
and missing-session paths now drop the message AND notify error
observers instead of silently dropping (still safe — unhandled events
are swallowed).

Tests: 8 new in mesh-client-event-hooks.test.ts; full suite 387/387 green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh-client): close gaps G1 (KNOCK auto-bootstrap) and G2 (auto-reconnect)

Two protocol-level gaps were identified during the AzureClaw vendored-SDK
audit (see Azure/kars#245 docs/agt-vs-vendored-sdk.md). Both block
moving fully off the vendored agentmesh-sdk fork onto upstream AGT.

## G1: receiver-side X3DH auto-bootstrap from KNOCK

Before: the responder side of an encrypted session was created only by
calling acceptSession(peerId, establishment) explicitly. But the
ChannelEstablishment was never carried on the wire — neither in
knock_accept nor in the first message frame — so the receiver had no way
to obtain it. Result: any fresh encrypted session failed with
"No encrypted session" the first time a ciphertext arrived.

Mirrors vendored agentmesh-sdk patch #4b. Behavior:

- Sender (establishSession): builds the SecureChannel first, then sends
  KNOCK with the ChannelEstablishment embedded as
  { ik: base64, ek: base64, otk?: number }.
- Receiver (handleKnock): if the knock contains `establishment` AND no
  prior session exists, deserialize and call acceptSession() automatically
  before responding with knock_accept. On bad establishment data, the
  knock is rejected and onError("knock", ...) fires.
- Backwards-compatible: legacy peers that don't embed establishment
  continue to work; caller must invoke acceptSession() manually as before.

## G2: auto-reconnect on transport drop

Before: MeshClient.reconnect() existed but was never called automatically.
Network blips, relay restarts, AKS node OOM-evicts left the client
disconnected forever — agents went mesh-deaf.

Mirrors vendored agentmesh-sdk patch microsoft#9. Behavior:

- New options: autoReconnect (default true), maxReconnectAttempts
  (default Number.POSITIVE_INFINITY), reconnectBaseDelayMs (default
  1000), reconnectMaxDelayMs (default 60000).
- On non-1000 ws.onclose: schedule reconnect with exponential backoff
  capped at reconnectMaxDelayMs. Light jitter (±20%) avoids
  thundering-herd reconnects across many sandboxes.
- onDisconnect handlers fire BEFORE the reconnect is scheduled so
  observers can see the drop.
- After maxReconnectAttempts, fires onError("ws", ..., "auto-reconnect
  gave up after N attempts") and stops.
- disconnect() cancels any pending reconnect timer.
- A connect() failure inside the reconnect path schedules another retry
  via the existing onclose path (and recursively from the catch handler
  in scheduleReconnect for the case where ws.onopen never fired).

## Tests

11 new tests across two files:
- tests/mesh-client-knock-bootstrap.test.ts: 5 tests covering sender-side
  embedding, receiver auto-bootstrap, legacy fallback, malformed
  establishment, and end-to-end happy-path encrypted send.
- tests/mesh-client-auto-reconnect.test.ts: 6 tests covering server-close
  reconnect, client-close no-reconnect, opt-out, give-up after max
  attempts, disconnect cancels pending timer, no duplicate scheduling.

Full AGT TS suite: 398/398 pass (was 387 before this commit). Build clean.

Both gaps were identified in the AzureClaw audit doc:
docs/agt-vs-vendored-sdk.md (Patch-by-patch audit section).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(mesh): close vendored-parity gaps G3, G4, G5

G3 — session-desync teardown (vendored agentmesh-sdk patch microsoft#13 equiv):
when Double Ratchet decryption fails for an existing session, the local
ratchet state is irrecoverable. Previously we only fired
onError('decrypt'); now we tear down the session (delete from
this.sessions; clear knockAccepted) and fire a distinct
onError('session_desync') so callers can re-run establishSession() and
resume communication.

G4 — pre-KNOCK encrypted-message buffer (vendored agentmesh-sdk
patch microsoft#16 equiv): under transport reordering the relay can deliver an
encrypted 'message' frame BEFORE the matching 'knock' arrives. Without
buffering, that first frame was dropped silently. Now MeshClient buffers
encrypted frames per-peer (default cap 5, TTL 3000ms), replays them
when the corresponding knock is accepted, and drops them when rejected
or on disconnect. Disabled by setting preKnockBufferSize: 0.

G5 — eager ghost-connection close on rebind (vendored agentmesh-relay
patch microsoft#2 equiv): when an agent reconnects with the same DID, the prior
'ghost' WebSocket is now closed eagerly with code 1000 'session_replaced'
instead of waiting for the 90s heartbeat-eviction timer. The 'finally'
cleanup now compares socket identity to avoid removing the freshly
rebound connection when the old socket's handler unwinds.

Tests: +5 (G4 pre-knock buffer behaviour), +2 (G3 type + lifecycle
safety), +1 (G5 ghost rebind). 405 TS tests pass; 18 relay tests pass.

NOTE: held local-only on branch azureclaw-meshclient-event-hooks; will
be coordinated upstream with the AGT team.

* feat(mesh): add RegistryClient and auto-register on MeshClient.connect()

Closes the structural gap where MeshClientOptions declared registryUrl
but no code path ever used it. Without this, agents can connect to the
relay and exchange ciphertext but are invisible to peers — they never
appear in /v1/discover and no one can fetch their X3DH pre-key bundle.

This commit adds the missing registry-side glue, in upstream-quality
shape, so AGT MeshClient becomes a drop-in replacement for vendored
forks that wired their own registry calls (e.g. AzureClaw vendor/agentmesh-sdk).

What's new:

- src/encryption/registry-client.ts: typed HTTP client wrapping the
  AgentMesh registry's REST surface (POST /v1/agents,
  PUT/GET /v1/agents/{did}/prekeys, GET /v1/discover, GET/DELETE
  /v1/agents/{did}). base64url <-> Uint8Array marshalling, optional
  Ed25519-Timestamp Authorization signer, configurable retry/timeout.

- src/encryption/mesh-client.ts:
  * New options: registryClient, registryClientOptions, capabilities,
    registrationMetadata, oneTimePrekeyCount, autoRegister.
  * connect() now calls registerSelf() at the end of a successful WS
    handshake when autoRegister is true (default) and a registry is
    configured. Idempotent — 409 (already registered) is treated as
    success; reconnect doesn't re-register.
  * registerSelf(): generates signed-prekey + N one-time prekeys via
    keyManager, then POSTs /v1/agents (capabilities = [displayName,
    ...options.capabilities]) and PUTs the prekey bundle.
  * discover(capability): wraps registry.discover.
  * establishSessionWithPeer(peerId): fetches peer prekeys from the
    registry and calls existing establishSession.
  * getRegistry(): exposes the underlying client for advanced cases.

- tests/registry-client.test.ts: 11 new tests covering wire format,
  base64url round-trip, idempotent register (409), 5xx retry, and
  Authorization header. Also covers MeshClient auto-register on connect
  end-to-end with a fake fetch + fake WebSocket.

- tests/mesh-client-*.test.ts: existing fixtures updated to pass
  autoRegister: false (no registry available in those tests). Behavior
  unchanged for the production path.

All 71 tests pass (60 previously + 11 new).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* feat(registry): add POST /v1/agents/{did}/heartbeat to bump last_seen

The registry's last_seen field is set on registration but never
updated by any HTTP handler — update_last_seen() exists in
store.py but is dead code in production. As a result, every
registered agent looks 'offline' (online=false on /presence) 90s
after spawn, and any client-side discover stale-filter rejects
all live agents.

Add an unauthenticated POST /v1/agents/{did}/heartbeat that calls
store.update_last_seen(). Idempotent. Returns 404 for unknown DIDs
so callers can detect a registry restart and re-register.

Mirrors the agent-side ping cadence (30s, well within the 90s
presence threshold).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(mesh): include Ed25519 identity_key_ed in prekey bundle so verifyBundle() works

PreKeyBundle.identityKeyEd was previously aliased to the X25519 identity_key
because the registry didn't store the Ed25519 signing key. As a result,
verifyBundle() either passed only because the stub test never exercised it,
or quietly accepted any signed pre-key — a soundness gap in the X3DH wrapper.

This commit threads identityKeyEd through the full path:

- agent-mesh/registry/store.py (AgentRecord): add identity_key_ed field
  (X25519 long-term key already present as identity_key; Ed25519 distinct).
- agent-governance-typescript/src/encryption/x3dh.ts: expose
  X3DHKeyManager.identityKeyEd getter.
- agent-governance-typescript/src/encryption/registry-client.ts:
  uploadPrekeys() now accepts identityKeyEd (32 bytes, validated) and
  serializes it as identity_key_ed; fetchPrekeys() deserializes it back into
  PreKeyBundle.identityKeyEd. Older bundles without the field fall back to
  the X25519 key — verifyBundle() will then correctly reject (forensic
  visibility, no silent bypass).
- agent-governance-typescript/src/encryption/mesh-client.ts: registerSelf()
  passes keyManager.identityKeyEd through to the registry.
- tests/registry-client.test.ts + tests/test_registry.py: assert the new
  field is round-tripped on PUT/GET.

Wire compatibility: PUT now requires identity_key_ed; GET emits it when set.
Existing AGT clients that don't upload it will get verifyBundle() rejection
on the receiving side — peers must upgrade together. (No silent regression
on the wire — the field is new, not a breaking change to existing fields.)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* test(mesh): set autoRegister:false in upstream knock + malformed-frame tests

These two upstream tests construct MeshClient without autoRegister:false,
which now triggers a real fetch to the registry on connect() (added by
our registry-client commit). The tests have no fake registry, so they
fail with 'fetch failed'. autoRegister:false skips the auto-registration
path and matches the pattern used in our 6 sibling mesh-client-*.test.ts
files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
MohammadHaroonAbuomar pushed a commit to MohammadHaroonAbuomar/agt-acs that referenced this pull request Jun 1, 2026
…xecute API (microsoft#2644)

* fix(agent-os): close authorization bypasses in stateless kernel and execute API

Three same-class authorization fixes identified in security review:

1. stateless._check_policies: caller-supplied params['approved']=True no longer satisfies requires_approval gates. Approval must flow through the trusted IntentManager path; unplanned drift on restricted actions is now denied. The legacy flag is stripped from params before action execution.

2. server/app.py /api/v1/execute: caller-supplied agent_id is no longer trusted when authentication is bypassed. The legacy AGENT_OS_ALLOW_UNAUTHENTICATED_EXECUTE env var now raises ValueError at construction time. The replacement AGENT_OS_UNSAFE_ALLOW_UNAUTHENTICATED_EXECUTE is gated on AGENT_OS_ENV in {dev,development,local}; the server-side identity is fixed by AGENT_OS_UNSAFE_LOCAL_EXECUTE_AGENT_ID (default local-dev-agent); mismatched caller agent_id is rejected with 422 (unsafe) or 403 (authenticated).

3. mcp-kernel-server KernelExecuteTool._check_policies: same params.get('approved') bypass pattern as (1); now ignored with a warning log and the action is denied with guidance pointing to a trusted host approval workflow.

Tests added/updated for all three paths. Tangential sweep covered other auth surfaces (mcp_gateway approval callback, AGENT_OS_* env vars, REST endpoints) and found no further in-class bugs in agent-os core; module-level FastAPI surfaces in caas/iatp/observability are out of scope for this PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(mcp-scan): regression for env-poisoning RCE + cwd hijack -- currently FAILING

Red-team findings microsoft#1 + microsoft#2: mcp-scan CLI accepts arbitrary environment keys (LD_PRELOAD, PYTHONPATH, NODE_OPTIONS, ...) and untrusted cwd paths when launching subprocesses, enabling pre-exec code injection.

These regression tests assert the SECURE behavior (refusal). They FAIL on this commit because the helpers _blocked_command_env_keys and _validate_launch_cwd do not exist, proving the vuln surface is present.

Failure mode: 28 errors in TestLaunchEnvAndCwdGuards (AttributeError on missing helpers). Fix applied in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(mcp-scan): restore env-key blocklist and untrusted-cwd guard

Closes red-team findings microsoft#1 + microsoft#2. Restores _blocked_command_env_keys and _validate_launch_cwd helpers. Red->Green: 28 errors -> 129 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for approval-key bypasses + provider edge cases -- currently FAILING

Red-team findings microsoft#8 (confusable/nested approved keys bypass strip), microsoft#10 (non-strict-True provider return treated as allow), microsoft#11 (log injection via CR/LF in caller fields), microsoft#12 (provider BaseException leaks past approval check).

Failure mode: 15 failures across stateless + mcp_kernel_server.tools. Cyrillic 'approvеd', uppercased 'Approved', nested dict values, truthy-non-bool returns ('yes', 1, object), and SystemExit/KeyboardInterrupt all currently bypass the gate. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): harden approval-key strip, strict-bool, BaseException, log sanitization

Closes red-team microsoft#8, microsoft#10, microsoft#11, microsoft#12. NFKC + casefold approved-key match, recursive strip into nested dicts/lists, strict 'is True', except BaseException, _sanitize_log_field. Red->Green: 15 failed -> 141 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for empty-policies bypass + non-loopback execute -- currently FAILING

Red-team findings microsoft#3 (no policy match -> action allowed even when requires_approval declared elsewhere) and microsoft#5 (unsafe execute mode trusted from arbitrary remote peers).

Failure mode: test_execute_global_approval_blocks_empty_policy_list FAILS because StatelessKernel falls through to allow when no policy entry matches. test_execute_unsafe_escape_hatch_rejects_non_loopback_peer FAILS because _authenticate_execute_request does not inspect request.client. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): close empty-policies bypass and enforce loopback for unsafe execute

Closes microsoft#3 + microsoft#5. _globally_protected_actions enforced after per-policy loop; _is_loopback_client rejects non-127.x/::1 peers with 403. Red->Green: 2 failed -> 94 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(intent): regression for cross-agent intent reuse -- currently FAILING

Red-team finding microsoft#4: IntentManager.check_action does not verify that the caller's agent_id matches the intent's agent_id, so agent B can reuse agent A's stored intent record to perform privileged actions under A's policy context.

Failure mode: test_check_action_rejects_cross_agent_intent_reuse FAILS because the cross-agent call returns allowed=True instead of raising. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(intent): bind intent to declaring agent_id

Closes microsoft#4. Asserts intent.agent_id == caller agent_id in check_action. Red->Green: 1 failed -> 41 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(iatp): regression for weak/short trusted-override tokens -- currently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(iatp): reject weak/short trusted-override tokens

Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(policies): regression for plaintext OPA over network -- currently FAILING

Red-team finding microsoft#7: OPABackend remote mode follows http:// URLs to non-loopback hosts without warning. An on-path attacker on the OPA route flips allow=true and the kernel approves any action.

Failure mode: test_plaintext_remote_non_loopback_denied and test_plaintext_opt_in_without_local_env_denied FAIL because _evaluate_remote performs the HTTP call without protocol gating. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(policies): require HTTPS for remote OPA unless explicitly opted in

Closes microsoft#7. _evaluate_remote rejects non-HTTPS unless loopback host OR (AGENT_OS_OPA_ALLOW_PLAINTEXT=1 + AGENT_OS_ENV in {local,dev,development}). Plaintext non-loopback returns error='plaintext_opa_blocked'. Red->Green: 2 failed -> 77 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(caas): regression for unauthenticated FastAPI surface gate -- currently FAILING

Red-team finding microsoft#6: caas.api.server only LOGS a warning when started outside local env; misconfigured deployment exposes every CaaS route silently.

Failure mode: 13 failures because _caas_unauth_gate_satisfied does not exist and startup hook does not raise. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(caas): require explicit env gate to start unauthenticated CaaS surface

Closes microsoft#6. Startup hook raises RuntimeError unless AGENT_OS_ENV in {local,dev,development} OR CAAS_UNSAFE_ALLOW_UNAUTH=1. Red->Green: 13 failed -> 13 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(agent-os): clear no-stubs/no-crypto/spell-check/safety-critical CI gates

- Reword TODO(security) doc comments to 'Future hardening (security)' in caas/api/server.py, iatp/main.py (x2 including proxy_task cross-ref), iatp/sidecar/__init__.py so the no-stubs CI gate accepts the docs without losing the design-followup intent.

- Replace inline 'import hmac; hmac.compare_digest' with 'import secrets; secrets.compare_digest' in iatp/main.py so the no-custom-crypto CI gate is happy (secrets.compare_digest is the stdlib re-export of hmac.compare_digest, same constant-time guarantee).

- Add 19 project-specific terms to .cspell-repo-terms.txt (ASGI, NFKC, casefold, confusables, multitenant, normalisation, sanitised, unicodedata, testclient, monkeypatched, baseexception, rsplit, hdrs, oncall, madmin, backendunavailable, changeme, shortone, approv) for the spell-check-changed-files job.

- Update tests/test_safety_critical.py::TestPolicyEdgeCases::test_empty_policies_list_allows to reflect the new fail-closed behavior from fix microsoft#3: an empty policies list must DENY requires_approval actions (file_write). Renamed to test_empty_policies_list_denies_protected_actions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(spell-check): allow cyrillic-e 'approv\u0435d' confusable used in unicode normalization tests

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

---------

Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Dhinesh Ponnarasan (DhineshPonnarasan) pushed a commit to DhineshPonnarasan/agent-governance-toolkit that referenced this pull request Jun 1, 2026
…xecute API (microsoft#2644)

* fix(agent-os): close authorization bypasses in stateless kernel and execute API

Three same-class authorization fixes identified in security review:

1. stateless._check_policies: caller-supplied params['approved']=True no longer satisfies requires_approval gates. Approval must flow through the trusted IntentManager path; unplanned drift on restricted actions is now denied. The legacy flag is stripped from params before action execution.

2. server/app.py /api/v1/execute: caller-supplied agent_id is no longer trusted when authentication is bypassed. The legacy AGENT_OS_ALLOW_UNAUTHENTICATED_EXECUTE env var now raises ValueError at construction time. The replacement AGENT_OS_UNSAFE_ALLOW_UNAUTHENTICATED_EXECUTE is gated on AGENT_OS_ENV in {dev,development,local}; the server-side identity is fixed by AGENT_OS_UNSAFE_LOCAL_EXECUTE_AGENT_ID (default local-dev-agent); mismatched caller agent_id is rejected with 422 (unsafe) or 403 (authenticated).

3. mcp-kernel-server KernelExecuteTool._check_policies: same params.get('approved') bypass pattern as (1); now ignored with a warning log and the action is denied with guidance pointing to a trusted host approval workflow.

Tests added/updated for all three paths. Tangential sweep covered other auth surfaces (mcp_gateway approval callback, AGENT_OS_* env vars, REST endpoints) and found no further in-class bugs in agent-os core; module-level FastAPI surfaces in caas/iatp/observability are out of scope for this PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(mcp-scan): regression for env-poisoning RCE + cwd hijack -- currently FAILING

Red-team findings #1 + microsoft#2: mcp-scan CLI accepts arbitrary environment keys (LD_PRELOAD, PYTHONPATH, NODE_OPTIONS, ...) and untrusted cwd paths when launching subprocesses, enabling pre-exec code injection.

These regression tests assert the SECURE behavior (refusal). They FAIL on this commit because the helpers _blocked_command_env_keys and _validate_launch_cwd do not exist, proving the vuln surface is present.

Failure mode: 28 errors in TestLaunchEnvAndCwdGuards (AttributeError on missing helpers). Fix applied in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(mcp-scan): restore env-key blocklist and untrusted-cwd guard

Closes red-team findings #1 + microsoft#2. Restores _blocked_command_env_keys and _validate_launch_cwd helpers. Red->Green: 28 errors -> 129 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for approval-key bypasses + provider edge cases -- currently FAILING

Red-team findings microsoft#8 (confusable/nested approved keys bypass strip), microsoft#10 (non-strict-True provider return treated as allow), microsoft#11 (log injection via CR/LF in caller fields), microsoft#12 (provider BaseException leaks past approval check).

Failure mode: 15 failures across stateless + mcp_kernel_server.tools. Cyrillic 'approvеd', uppercased 'Approved', nested dict values, truthy-non-bool returns ('yes', 1, object), and SystemExit/KeyboardInterrupt all currently bypass the gate. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): harden approval-key strip, strict-bool, BaseException, log sanitization

Closes red-team microsoft#8, microsoft#10, microsoft#11, microsoft#12. NFKC + casefold approved-key match, recursive strip into nested dicts/lists, strict 'is True', except BaseException, _sanitize_log_field. Red->Green: 15 failed -> 141 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(authz): regression for empty-policies bypass + non-loopback execute -- currently FAILING

Red-team findings microsoft#3 (no policy match -> action allowed even when requires_approval declared elsewhere) and microsoft#5 (unsafe execute mode trusted from arbitrary remote peers).

Failure mode: test_execute_global_approval_blocks_empty_policy_list FAILS because StatelessKernel falls through to allow when no policy entry matches. test_execute_unsafe_escape_hatch_rejects_non_loopback_peer FAILS because _authenticate_execute_request does not inspect request.client. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(authz): close empty-policies bypass and enforce loopback for unsafe execute

Closes microsoft#3 + microsoft#5. _globally_protected_actions enforced after per-policy loop; _is_loopback_client rejects non-127.x/::1 peers with 403. Red->Green: 2 failed -> 94 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(intent): regression for cross-agent intent reuse -- currently FAILING

Red-team finding microsoft#4: IntentManager.check_action does not verify that the caller's agent_id matches the intent's agent_id, so agent B can reuse agent A's stored intent record to perform privileged actions under A's policy context.

Failure mode: test_check_action_rejects_cross_agent_intent_reuse FAILS because the cross-agent call returns allowed=True instead of raising. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(intent): bind intent to declaring agent_id

Closes microsoft#4. Asserts intent.agent_id == caller agent_id in check_action. Red->Green: 1 failed -> 41 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(iatp): regression for weak/short trusted-override tokens -- currently FAILING

Red-team finding microsoft#9: AGENT_OS_IATP_TRUSTED_OVERRIDE_TOKEN accepts any non-empty string -- 'true', 'admin', 'password', 'x' -- so a misconfigured operator (or attacker who can set one env var) trivially enables the X-User-Override path.

Failure mode: 18 failures in test_blacklisted_weak_token_disables_gate (main+sidecar paths) and test_short_token_disables_gate. Each demonstrates a weak/short token still bypassing the override check. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(iatp): reject weak/short trusted-override tokens

Closes microsoft#9. _load_trusted_override_token enforces 16-char minimum and blacklists {true,yes,admin,password,...}. Sidecar delegates to iatp.main to prevent drift. Red->Green: 18 failed -> 30 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(policies): regression for plaintext OPA over network -- currently FAILING

Red-team finding microsoft#7: OPABackend remote mode follows http:// URLs to non-loopback hosts without warning. An on-path attacker on the OPA route flips allow=true and the kernel approves any action.

Failure mode: test_plaintext_remote_non_loopback_denied and test_plaintext_opt_in_without_local_env_denied FAIL because _evaluate_remote performs the HTTP call without protocol gating. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(policies): require HTTPS for remote OPA unless explicitly opted in

Closes microsoft#7. _evaluate_remote rejects non-HTTPS unless loopback host OR (AGENT_OS_OPA_ALLOW_PLAINTEXT=1 + AGENT_OS_ENV in {local,dev,development}). Plaintext non-loopback returns error='plaintext_opa_blocked'. Red->Green: 2 failed -> 77 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* test(caas): regression for unauthenticated FastAPI surface gate -- currently FAILING

Red-team finding microsoft#6: caas.api.server only LOGS a warning when started outside local env; misconfigured deployment exposes every CaaS route silently.

Failure mode: 13 failures because _caas_unauth_gate_satisfied does not exist and startup hook does not raise. Fix in next commit.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* fix(caas): require explicit env gate to start unauthenticated CaaS surface

Closes microsoft#6. Startup hook raises RuntimeError unless AGENT_OS_ENV in {local,dev,development} OR CAAS_UNSAFE_ALLOW_UNAUTH=1. Red->Green: 13 failed -> 13 passed.

Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(agent-os): clear no-stubs/no-crypto/spell-check/safety-critical CI gates

- Reword TODO(security) doc comments to 'Future hardening (security)' in caas/api/server.py, iatp/main.py (x2 including proxy_task cross-ref), iatp/sidecar/__init__.py so the no-stubs CI gate accepts the docs without losing the design-followup intent.

- Replace inline 'import hmac; hmac.compare_digest' with 'import secrets; secrets.compare_digest' in iatp/main.py so the no-custom-crypto CI gate is happy (secrets.compare_digest is the stdlib re-export of hmac.compare_digest, same constant-time guarantee).

- Add 19 project-specific terms to .cspell-repo-terms.txt (ASGI, NFKC, casefold, confusables, multitenant, normalisation, sanitised, unicodedata, testclient, monkeypatched, baseexception, rsplit, hdrs, oncall, madmin, backendunavailable, changeme, shortone, approv) for the spell-check-changed-files job.

- Update tests/test_safety_critical.py::TestPolicyEdgeCases::test_empty_policies_list_allows to reflect the new fail-closed behavior from fix microsoft#3: an empty policies list must DENY requires_approval actions (file_write). Renamed to test_empty_policies_list_denies_protected_actions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

* ci(spell-check): allow cyrillic-e 'approv\u0435d' confusable used in unicode normalization tests

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>

---------

Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jack Batzner <jackbatzner@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
dylanyunlon added a commit to dylanyunlon/agent-governance-toolkit that referenced this pull request Sep 17, 2026
…t#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>
dylanyunlon added a commit to dylanyunlon/agent-governance-toolkit that referenced this pull request Sep 17, 2026
…t#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>
dylanyunlon added a commit to dylanyunlon/agent-governance-toolkit that referenced this pull request Sep 17, 2026
…t#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>
dylanyunlon added a commit to dylanyunlon/agent-governance-toolkit that referenced this pull request Sep 17, 2026
…t#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>
dylanyunlon added a commit to dylanyunlon/agent-governance-toolkit that referenced this pull request Sep 18, 2026
…t#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>
MohammadHaroonAbuomar pushed a commit that referenced this pull request Sep 18, 2026
…ust (#3933) (#3955)

* fix: credential redactor boundary anchors across all SDKs (#3933)

McpCredentialRedactor (C#), AuditLogger (TypeScript), CredentialRedactor
(Rust), and CredentialRedactor (Python) all contained boundary anchors
that treated _ as a boundary-blocking character, so a valid GitHub,
OpenAI, AWS, or Google API secret annotated with _old, _deprecated, or
_rotated (or preceded by session_, env_, etc.) passed through completely
unredacted.

Root cause: _ included in lookaround exclusion sets for GitHub/OpenAI
patterns, and \b word boundaries (which treat _ as a word character in
all regex engines) for AWS/Google patterns. In Rust the same logic was
expressed procedurally via is_left_boundary_char/is_right_boundary_char.

Fix: all four SDKs now use (?<![A-Za-z0-9])...(?![A-Za-z0-9]) lookaround
anchors (alphanumeric-only), matching the existing SlackToken pattern
which was already correct.

Also fixes the content_scanner.py SSN pattern in agent-rag-governance to
accept space and dot separators and use the consistent lookaround anchor,
closing the detection-disagreement gap with credential_redactor.py (#3815).

Closes #3933
Ref #3815

Test coverage:
- C#: McpCredentialRedactorTests (right-edge, left-edge, both-edges,
  multi-credential, still-rejects-alphanumeric), McpResponseSanitizerTests
  (pipeline boundary), McpGatewayTests (gateway-level BLOCK/SANITIZE)
- TypeScript: policy-audit.test.ts (all four patterns through AuditLogger)
- Rust: redactor.rs inline tests + response.rs scanner pipeline tests
- Python: test_credential_redactor.py, test_mcp_response_scanner.py,
  test_mcp_pii_and_response_gateway.py, test_content_scanner.py
- CI: test_regression_credential_boundary.py (source-level guard across
  all four SDKs)

Security audit: docs/security/audits/2026-09-14-credential-boundary-redaction-cross-sdk.md

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address review feedback on credential boundary anchors (#3933)

Changes responding to review on PR #3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review #5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review #3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review #2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review #4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review #7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review #8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review #3, #6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review #9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-2 review feedback (#3933)

Round-2 review fixes:

- DCO: both commits now carry Signed-off-by (rebase --signoff)

- cspell: added AKIAIOSFODNN, AKIAXXXX, fooghp, lookarounds to
  .cspell-repo-terms.txt

- C# overlap with #3934: dropped McpCredentialRedactor.cs and
  McpCredentialRedactorTests.cs changes; C# fixes land in #3934

- Rust Google tail: is_right_boundary_char now accepts a trailing
  hyphen for GoogleApiKey via last_consumed check, mirroring
  (?:(?![A-Za-z0-9])|(?<=-)) in the regex SDKs

- Rust OpenAI left-edge widening: is_left_boundary_char no longer
  blocks '-' for OpenAiToken, aligning with Python. Added pinning
  test redacts_openai_token_preceded_by_hyphen_left_edge_widening

- TS OpenAI left-edge widening: added pinning test in
  policy-audit.test.ts

- E2E fixture: changed to access_token=AKIAIOSFODNN7EXAMPLE_old
  so MuteAgent api_key builtin matches it

- Regression guard: removed C# from regex file list (now in #3934),
  removed dead _WB_NEAR_TOKEN, added Rust-specific match-arm tests
  (test_rust_left_boundary_non_slack_rejects_only_alphanumeric,
  test_rust_right_boundary_non_slack_rejects_alphanumeric)

- Audit doc: rewritten with accurate per-SDK delta vs main table,
  OpenAI widening section now includes Rust

- CHANGELOG: accurately describes this PR's delta vs main, notes
  C# fixes land in #3934

- CONTRIBUTING.md + copilot-instructions.md: Rust uses match arms
  not regex, guard covers regex SDKs + Rust separately

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-3 review feedback (#3933)

Round-3 review fixes (MohammadHaroonAbuomar, 2026-09-17T14:46):

- C# test suite red: removed 3 boundary tests from McpGatewayTests.cs
  and 5 boundary tests from McpResponseSanitizerTests.cs that asserted
  glued-key redaction against unchanged main patterns. C# source changes
  and all C# boundary tests belong in #3934.

- C# AGENTS.md: reverted to main; anchor guidance belongs in #3934.

- TS Google superset: added pinning test in policy-audit.test.ts for
  the (?<=-)  tail branch (AIza + 34*A + '-X' -> [REDACTED]).

- cspell: added AKIAAAAAAAAAAAAAAAAA to .cspell-repo-terms.txt (the
  round-2 fix added AKIAIOSFODNN but missed the all-A test literal).

- Audit doc :19: changed 'landed' to 'will land' for #3934 (still
  open at time of writing).

- Audit doc :76: guard coverage now accurately states TypeScript
  (audit.ts) and Python (credential_redactor.py); Rust is checked
  via separate match-arm assertions, not regex scanning.

- Regression guard: replaced naive startswith('}') function-end
  detection with brace-depth counting. The old logic stopped at the
  first inner '}' (e.g. GoogleApiKey's block arm), so mutating the
  catch-all arm below it was undetected (verified by MohammadHaroonAbuomar
  with a multi-line arm containing '_').

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-4 review feedback (#3933)

Round-4 review fixes (MohammadHaroonAbuomar, 2026-09-18T03:36):

- CodeQL js/insufficient-password-hash false positive: renamed TS test
  fixtures from fakeGoogleApiKey/fakeOpenAiToken/fakeAwsAccessKey to
  googleKeyFixture/openAiTokenFixture/awsKeyFixture. The name heuristic
  flagged 'fakeGoogleApiKey' flowing into the audit hash chain.

- PR text still said 'all four SDKs' / listed C#: updated CONTRIBUTING.md
  :440, copilot-instructions.md :384, and regression guard docstring to say
  'TypeScript and Python (C# is in #3934)'. Removed '.NET' from the guard
  module docstring.

- no-stubs.sh XXX match: changed credential_redactor.py :57 comment from
  'AKIAXXXX_old' to 'AKIA<key>_old'.

- gitleaks generic-api-key: added .gitleaksignore entry for commit
  4013555:test_mcp_pii_and_response_gateway.py:343.

- C# stray blank lines: restored McpGatewayTests.cs and
  McpResponseSanitizerTests.cs to exact main content (zero diff).

- Regression guard multi-line arm bypass: rewrote the Rust boundary
  function guards using _extract_fn_body + _split_match_arms helpers
  that parse brace-delimited match arms holistically instead of
  per-line. Mutations that previously passed now fail:
  - multi-line left arm with '_ in block: detected
  - catch-all '_ => { false }' across lines: detected
  - Google right arm with added '_ check: detected

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: assert boundary guard functions exist before checking arms (#3933)

Round-5 review fix (MohammadHaroonAbuomar, 2026-09-18T04:38):

- Rust boundary guard passes vacuously if is_left_boundary_char or
  is_right_boundary_char is renamed or moved: the for-loop over arms
  iterates zero times and the test passes with no arms checked. Added
  `assert start_line > 0 and arms` after each _extract_fn_body +
  _split_match_arms call so the guard fails loudly when the target
  function is absent.

- PR title and body updated via GitHub API: removed all C# references
  (Problem table row, AST chain, test coverage bullets, dotnet checkbox).
  Title now reads 'TypeScript, Python and Rust'. C# is tracked in #3934.

Signed-off-by: dylanyunlon <dogechat@163.com>

---------

Signed-off-by: dylanyunlon <dogechat@163.com>
Yuvraj Singh (yuvrajsingh2428) pushed a commit to yuvrajsingh2428/agent-governance-toolkit that referenced this pull request Oct 1, 2026
…ust (microsoft#3933) (microsoft#3955)

* fix: credential redactor boundary anchors across all SDKs (microsoft#3933)

McpCredentialRedactor (C#), AuditLogger (TypeScript), CredentialRedactor
(Rust), and CredentialRedactor (Python) all contained boundary anchors
that treated _ as a boundary-blocking character, so a valid GitHub,
OpenAI, AWS, or Google API secret annotated with _old, _deprecated, or
_rotated (or preceded by session_, env_, etc.) passed through completely
unredacted.

Root cause: _ included in lookaround exclusion sets for GitHub/OpenAI
patterns, and \b word boundaries (which treat _ as a word character in
all regex engines) for AWS/Google patterns. In Rust the same logic was
expressed procedurally via is_left_boundary_char/is_right_boundary_char.

Fix: all four SDKs now use (?<![A-Za-z0-9])...(?![A-Za-z0-9]) lookaround
anchors (alphanumeric-only), matching the existing SlackToken pattern
which was already correct.

Also fixes the content_scanner.py SSN pattern in agent-rag-governance to
accept space and dot separators and use the consistent lookaround anchor,
closing the detection-disagreement gap with credential_redactor.py (microsoft#3815).

Closes microsoft#3933
Ref microsoft#3815

Test coverage:
- C#: McpCredentialRedactorTests (right-edge, left-edge, both-edges,
  multi-credential, still-rejects-alphanumeric), McpResponseSanitizerTests
  (pipeline boundary), McpGatewayTests (gateway-level BLOCK/SANITIZE)
- TypeScript: policy-audit.test.ts (all four patterns through AuditLogger)
- Rust: redactor.rs inline tests + response.rs scanner pipeline tests
- Python: test_credential_redactor.py, test_mcp_response_scanner.py,
  test_mcp_pii_and_response_gateway.py, test_content_scanner.py
- CI: test_regression_credential_boundary.py (source-level guard across
  all four SDKs)

Security audit: docs/security/audits/2026-09-14-credential-boundary-redaction-cross-sdk.md

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address review feedback on credential boundary anchors (microsoft#3933)

Changes responding to review on PR microsoft#3955:

- Google API key: C# and TS tail anchors use strict superset
  (?:(?![A-Za-z0-9])|(?<=-)) matching the Python SDK, so a key
  ending in '-' followed by alnum is not missed (review microsoft#5)

- Rust is_right_boundary_char: now rejects ASCII alphanumerics for
  all non-Slack kinds, aligning with (?![A-Za-z0-9]) in C#/TS/Python.
  Previously returned false for all non-Slack kinds (review microsoft#3)

- Rust test: updated to assert detection — the old test encoded the
  buggy behaviour where prefix_ghp_... was missed (review microsoft#2)

- Regression guard regex: fixed to match both lookbehind (?<!) and
  lookahead (?!) forms (review microsoft#4)

- Rego comment: corrected to state the actual reason (RE2 engine
  limitation, not detection-only) (review microsoft#7)

- E2E assertion: added _old glued form to fixture so
  assert_no_raw_secrets is exercised (review microsoft#8)

- Audit doc: rewritten with per-SDK table and explicit OpenAI
  left-edge widening acknowledgement (review microsoft#3, microsoft#6)

- Doc drift: fixed copilot-instructions numbering, Rust AGENTS.md
  Slack exception, reverted README.md date-only change (review microsoft#9)

OpenAI left anchor kept as (?<![A-Za-z0-9]): aligns with Python SDK,
ensures env-sk-... prefixed secrets are detected. Documented in audit.

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-2 review feedback (microsoft#3933)

Round-2 review fixes:

- DCO: both commits now carry Signed-off-by (rebase --signoff)

- cspell: added AKIAIOSFODNN, AKIAXXXX, fooghp, lookarounds to
  .cspell-repo-terms.txt

- C# overlap with microsoft#3934: dropped McpCredentialRedactor.cs and
  McpCredentialRedactorTests.cs changes; C# fixes land in microsoft#3934

- Rust Google tail: is_right_boundary_char now accepts a trailing
  hyphen for GoogleApiKey via last_consumed check, mirroring
  (?:(?![A-Za-z0-9])|(?<=-)) in the regex SDKs

- Rust OpenAI left-edge widening: is_left_boundary_char no longer
  blocks '-' for OpenAiToken, aligning with Python. Added pinning
  test redacts_openai_token_preceded_by_hyphen_left_edge_widening

- TS OpenAI left-edge widening: added pinning test in
  policy-audit.test.ts

- E2E fixture: changed to access_token=AKIAIOSFODNN7EXAMPLE_old
  so MuteAgent api_key builtin matches it

- Regression guard: removed C# from regex file list (now in microsoft#3934),
  removed dead _WB_NEAR_TOKEN, added Rust-specific match-arm tests
  (test_rust_left_boundary_non_slack_rejects_only_alphanumeric,
  test_rust_right_boundary_non_slack_rejects_alphanumeric)

- Audit doc: rewritten with accurate per-SDK delta vs main table,
  OpenAI widening section now includes Rust

- CHANGELOG: accurately describes this PR's delta vs main, notes
  C# fixes land in microsoft#3934

- CONTRIBUTING.md + copilot-instructions.md: Rust uses match arms
  not regex, guard covers regex SDKs + Rust separately

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-3 review feedback (microsoft#3933)

Round-3 review fixes (MohammadHaroonAbuomar, 2026-09-17T14:46):

- C# test suite red: removed 3 boundary tests from McpGatewayTests.cs
  and 5 boundary tests from McpResponseSanitizerTests.cs that asserted
  glued-key redaction against unchanged main patterns. C# source changes
  and all C# boundary tests belong in microsoft#3934.

- C# AGENTS.md: reverted to main; anchor guidance belongs in microsoft#3934.

- TS Google superset: added pinning test in policy-audit.test.ts for
  the (?<=-)  tail branch (AIza + 34*A + '-X' -> [REDACTED]).

- cspell: added AKIAAAAAAAAAAAAAAAAA to .cspell-repo-terms.txt (the
  round-2 fix added AKIAIOSFODNN but missed the all-A test literal).

- Audit doc :19: changed 'landed' to 'will land' for microsoft#3934 (still
  open at time of writing).

- Audit doc :76: guard coverage now accurately states TypeScript
  (audit.ts) and Python (credential_redactor.py); Rust is checked
  via separate match-arm assertions, not regex scanning.

- Regression guard: replaced naive startswith('}') function-end
  detection with brace-depth counting. The old logic stopped at the
  first inner '}' (e.g. GoogleApiKey's block arm), so mutating the
  catch-all arm below it was undetected (verified by MohammadHaroonAbuomar
  with a multi-line arm containing '_').

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: address round-4 review feedback (microsoft#3933)

Round-4 review fixes (MohammadHaroonAbuomar, 2026-09-18T03:36):

- CodeQL js/insufficient-password-hash false positive: renamed TS test
  fixtures from fakeGoogleApiKey/fakeOpenAiToken/fakeAwsAccessKey to
  googleKeyFixture/openAiTokenFixture/awsKeyFixture. The name heuristic
  flagged 'fakeGoogleApiKey' flowing into the audit hash chain.

- PR text still said 'all four SDKs' / listed C#: updated CONTRIBUTING.md
  :440, copilot-instructions.md :384, and regression guard docstring to say
  'TypeScript and Python (C# is in microsoft#3934)'. Removed '.NET' from the guard
  module docstring.

- no-stubs.sh XXX match: changed credential_redactor.py :57 comment from
  'AKIAXXXX_old' to 'AKIA<key>_old'.

- gitleaks generic-api-key: added .gitleaksignore entry for commit
  4013555:test_mcp_pii_and_response_gateway.py:343.

- C# stray blank lines: restored McpGatewayTests.cs and
  McpResponseSanitizerTests.cs to exact main content (zero diff).

- Regression guard multi-line arm bypass: rewrote the Rust boundary
  function guards using _extract_fn_body + _split_match_arms helpers
  that parse brace-delimited match arms holistically instead of
  per-line. Mutations that previously passed now fail:
  - multi-line left arm with '_ in block: detected
  - catch-all '_ => { false }' across lines: detected
  - Google right arm with added '_ check: detected

Signed-off-by: dylanyunlon <dogechat@163.com>

* fix: assert boundary guard functions exist before checking arms (microsoft#3933)

Round-5 review fix (MohammadHaroonAbuomar, 2026-09-18T04:38):

- Rust boundary guard passes vacuously if is_left_boundary_char or
  is_right_boundary_char is renamed or moved: the for-loop over arms
  iterates zero times and the test passes with no arms checked. Added
  `assert start_line > 0 and arms` after each _extract_fn_body +
  _split_match_arms call so the guard fails loudly when the target
  function is absent.

- PR title and body updated via GitHub API: removed all C# references
  (Problem table row, AST chain, test coverage bullets, dotnet checkbox).
  Title now reads 'TypeScript, Python and Rust'. C# is tracked in microsoft#3934.

Signed-off-by: dylanyunlon <dogechat@163.com>

---------

Signed-off-by: dylanyunlon <dogechat@163.com>
Signed-off-by: yuvrajsingh2428 <offcyuvi2428@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant