Skip to content

fix(agent-os): harden foundational skill-aware audit trail for issue #1609 - #2572

Merged
Imran Siddique (imran-siddique) merged 6 commits into
microsoft:mainfrom
DhineshPonnarasan:fix/1609-skill-audit-hardening
Jun 9, 2026
Merged

Imran Siddique (imran-siddique) merged 6 commits into
microsoft:mainfrom
DhineshPonnarasan:fix/1609-skill-audit-hardening

Conversation

@DhineshPonnarasan

@DhineshPonnarasan Dhinesh Ponnarasan (DhineshPonnarasan) commented May 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Hardens and normalizes the foundational skill-aware audit trail for issue #1609 without expanding scope.

Problem

The initial implementation had three review blockers:

  • Skill provenance could be inferred from generic mappings and user-controlled payloads.
  • Context hashing had non-deterministic fallback behavior, reducing forensic value.
  • Timestamp and adapter emission semantics were inconsistent across framework integrations.

This update also tightens intent lifecycle validation so verification is only valid during active execution.

Changes

File What changed
base.py Added trusted provenance abstraction, trusted-only extraction flow, Literal-typed trust marker, deterministic canonical hashing with safe-null fallback, and debug visibility in hash canonicalization failure path.
google_adk_adapter.py Removed duplicate skill-field computation in pre-tool path and added centralized post-tool POLICY_CHECK emission for lifecycle symmetry.
autogen_adapter.py Removed duplicate skill-field computation by reusing emitted payload fields; kept trusted-only provenance extraction and UTC audit timestamp behavior.
semantic_kernel_adapter.py Normalized skill provenance from trusted SK metadata surfaces and updated invocation audit timestamp to UTC ISO format.
langchain_adapter.py Removed payload-derived provenance inference and normalized tool invocation timestamp to UTC ISO format.
openai_agents_sdk.py Standardized trusted-source wiring for skill-audit enrichment across lifecycle events.
crewai_adapter.py Standardized centralized audit emission to pass only trusted context-derived provenance sources.
intent.py Verification guard now requires EXECUTING state only, enforcing strict declare -> approve -> execute -> verify lifecycle and preventing verification-finalization from APPROVED without execution.
test_skill_audit_helpers.py Added and refined coverage for trusted-only extraction, spoof resistance, deterministic nested hash behavior, safe-null hashing for non-serializable input, and clearer nested-ordering test naming.
test_google_adk_adapter.py Added assertions for trust field and spoof-resistance in ADK audit events.
test_crewai_hooks.py Added trust-field and spoofed metadata resistance assertions.
test_semantic_kernel_hooks.py Added provenance trust assertion for SK invocation audit records.
test_openai_agents_sdk_adapter.py Added trust-field assertions and spoofed tool-args provenance resistance checks.
test_langchain_middleware.py Updated provenance tests to trusted metadata surfaces and added adversarial spoof tests for tool args.
test_autogen_hooks.py Added trust-field assertion and spoofed FunctionCall argument resistance test.
test_intent.py Updated async helper compatibility and intent lifecycle coverage under the stricter verify guard.
test_intent_hardened.py Updated async helper compatibility behavior for current runtime.
test_llamafirewall_adapter.py Updated async helper compatibility behavior for current runtime.
CHANGELOG.md Clarified trusted provenance boundary, deterministic hash semantics, additive nullable fields, and safe-null behavior for non-canonical payloads.
ci.yml Added npm dependency install fallback for matrix paths without lockfiles.
.cspell-repo-terms.txt Added missing repo terms from CI spellcheck feedback.

Testing

Focused regression runs:

  • pytest tests/test_skill_audit_helpers.py tests/test_crewai_hooks.py tests/test_semantic_kernel_hooks.py tests/test_openai_agents_sdk_adapter.py tests/test_langchain_middleware.py tests/test_autogen_hooks.py -q
    Result: 224 passed

  • pytest tests/test_google_adk_adapter.py::TestAuditAndStats::test_audit_event_fields tests/test_google_adk_adapter.py::TestAuditAndStats::test_skill_metadata_extracted_into_audit_event tests/test_google_adk_adapter.py::TestAuditAndStats::test_spoofed_skill_metadata_in_tool_args_is_ignored -q
    Result: 3 passed

  • Combined targeted pass of the above suites
    Result: 227 passed

  • pytest tests/test_skill_audit_helpers.py tests/test_intent.py tests/test_intent_hardened.py tests/test_llamafirewall_adapter.py -q
    Result: 98 passed

Note: Full ADK async-marked suite in this local environment still shows pre-existing async plugin configuration warnings or failures not introduced by this PR.

Scope Guardrails

This PR intentionally does not add signature systems, sandbox enforcement, skill policy engines, marketplace trust infrastructure, or heavy forensic snapshot and diff systems.

This remains first-deliverable scope: foundational, framework-agnostic, backward-compatible skill-aware audit hardening.

Follow-up

Remaining naive timestamp calls in base integration utilities will be normalized in a follow-up PR for end-to-end UTC consistency.

Closes

Closes #1609

Imran Siddique (@imran-siddique) Nishar Miya (@miyannishar) 风 (Feng) (@fengtrace) could you please review when possible?

@github-actions github-actions Bot added documentation Improvements or additions to documentation tests size/XL Extra large PR (500+ lines) labels May 25, 2026
@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown
🤖 AI Agent: code-reviewer — View details

AI-generated review output. Treat it as untrusted analysis and verify before acting.

TL;DR: 0 blockers, 2 warnings. The changes enhance security and auditability but introduce minor areas for follow-up.

# Sev Issue Where
1 Warn Lack of cryptographic guarantees for context hashing in audit trails base.py
2 Warn Missing end-to-end UTC normalization for timestamps Multiple adapters and utilities

Action items: None, as no blockers were identified.

Warnings (fine as follow-up PRs):

  1. Ensure context hashing uses cryptographic guarantees if provenance integrity is critical.
  2. Normalize all timestamp calls to UTC across the codebase for consistency.

@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown
🤖 AI Agent: security-scanner — View details

AI-generated review output. Treat it as untrusted analysis and verify before acting.

No security issues found.

@github-actions

Copy link
Copy Markdown

🟡 Contributor Check: MEDIUM

Check Result
Profile MEDIUM
Credential NONE
Overall MEDIUM

Automated check by AGT Contributor Check.

@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown
🤖 AI Agent: test-generator — `agent-os/src/agent_os/integrations/base.py`

AI-generated review output. Treat it as untrusted analysis and verify before acting.

agent-os/src/agent_os/integrations/base.py

  • test_hash_context_null_input -- Validate safe-null behavior when None is passed to hash_context.
  • test_hash_context_non_serializable_input -- Ensure non-serializable input to hash_context returns None without exceptions.
  • test_trusted_skill_metadata_source_invalid_input -- Test that invalid inputs (non-string or empty values) are handled correctly in trusted_skill_metadata_source.
  • test_extract_skill_metadata_untrusted_sources -- Verify that extract_skill_metadata ignores untrusted sources.
  • test_extract_skill_metadata_default_origin -- Check that extract_skill_metadata correctly applies the default_origin fallback.

agent-os/src/agent_os/integrations/autogen_adapter.py

  • test_emit_skill_audit_event -- Verify that emit_skill_audit_event correctly populates all fields, including trusted_sources and context_hash.
  • test_on_send_trusted_metadata -- Ensure on_send correctly extracts and emits trusted metadata from message.

agent-os/src/agent_os/integrations/semantic_kernel_adapter.py

  • test_normalized_skill_provenance -- Validate that skill provenance is normalized from trusted SK metadata surfaces.
  • test_audit_timestamp_format -- Ensure audit timestamps are emitted in UTC ISO format.

agent-os/src/agent_os/integrations/langchain_adapter.py

  • test_provenance_inference_removal -- Confirm that payload-derived provenance inference is removed.
  • test_tool_invocation_timestamp -- Validate that tool invocation timestamps are normalized to UTC ISO format.

agent-os/src/agent_os/integrations/openai_agents_sdk.py

  • test_trusted_source_wiring -- Verify that trusted-source wiring is correctly implemented for skill-audit enrichment.
  • test_provenance_resistance -- Ensure spoofed tool-args provenance is ignored.

agent-os/src/agent_os/integrations/crewai_adapter.py

  • test_centralized_audit_emission -- Validate centralized audit emission with trusted context-derived provenance sources.

@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown
🤖 AI Agent: breaking-change-detector — API Compatibility

AI-generated review output. Treat it as untrusted analysis and verify before acting.

API Compatibility

Severity Change Impact
High intent.py: Verification guard now requires EXECUTING state only, enforcing stricter lifecycle validation. This change may break existing workflows that rely on verification-finalization from the APPROVED state without execution.
High base.py: Added SkillAuditMetadata and TrustedSkillMetadataSource classes with stricter trust boundaries for metadata extraction. Existing integrations relying on user-controlled data for skill metadata may break due to the new trust boundary enforcement.
Medium autogen_adapter.py: Modified audit trail emission to include only trusted metadata sources and normalized timestamps to UTC. This may affect integrations relying on untrusted or non-UTC metadata.

@github-actions github-actions Bot added the needs-review:MEDIUM Contributor check flagged MEDIUM risk label May 25, 2026
@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown
🤖 AI Agent: docs-sync-checker — Docs Sync

AI-generated review output. Treat it as untrusted analysis and verify before acting.

Docs Sync

  • SkillAuditMetadata in base.py -- missing docstring
  • TrustedSkillMetadataSource in base.py -- missing docstring
  • hash_context() in base.py -- missing docstring
  • trusted_skill_metadata_source() in base.py -- missing docstring
  • trusted_skill_metadata_from_mapping() in base.py -- missing docstring
  • trusted_sources() in base.py -- missing docstring
  • trusted_sources_from_attrs() in base.py -- missing docstring
  • extract_skill_metadata() in base.py -- missing docstring
  • README.md -- no updates found for new skill-aware audit trail features
  • CHANGELOG.md -- entry found for skill-aware audit trail changes

@github-actions

github-actions Bot commented May 25, 2026 •

Copy link
Copy Markdown

PR Review Summary

Check Status Details
🔍 Code Review ⚠️ Missing No current-run comment
🛡️ Security Scan ⚠️ Missing No current-run comment
🔄 Breaking Changes ⚠️ Missing No current-run comment
📝 Docs Sync ⚠️ Missing No current-run comment
🧪 Test Coverage ⚠️ Missing No current-run comment

Verdict: ⚠️ AI review incomplete; ready for human review

AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims.

@Ricky-G Ricky Gummadi (Ricky-G) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: extract the repeated trusted-source boilerplate into a BaseIntegration helper

Every adapter touched in this PR now repeats the same 9-line block, and it appears 8+ times across the diff:

trusted_skill_sources = tuple(
    source
    for source in (
        kernel.trusted_skill_metadata_source(
            skill_name=getattr(x, "skill_name", None),
            skill_origin=getattr(x, "skill_origin", None),
        ),
    )
    if source is not None
)

Consider extracting a small helper on BaseIntegration, e.g.:

@staticmethod
def trusted_sources_from_attrs(*objs) -> tuple[TrustedSkillMetadataSource, ...]:
    return tuple(
        s for s in (
            BaseIntegration.trusted_skill_metadata_source(
                skill_name=getattr(o, "skill_name", None),
                skill_origin=getattr(o, "skill_origin", None),
            ) for o in objs
        ) if s is not None
    )

Each adapter call site then collapses to a single line:

trusted_skill_sources = kernel.trusted_sources_from_attrs(context)
# or for adapters that read from multiple objects:
trusted_skill_sources = kernel.trusted_sources_from_attrs(request, getattr(request, "skill_metadata", None))

Why this is worth doing:

  • Removes ~70+ lines of duplicated boilerplate across autogen_adapter.py, crewai_adapter.py, google_adk_adapter.py, langchain_adapter.py, openai_agents_sdk.py, and semantic_kernel_adapter.py.
  • Centralizes the trust-extraction pattern so the next adapter (or the next framework hook added to an existing adapter) can't accidentally diverge — e.g. forget the if source is not None filter, or call getattr without a default.
  • Keeps the trust boundary enforced in exactly one place, which is the whole point of TrustedSkillMetadataSource.

Not a blocker for merging, but high value-to-risk ratio and a natural extension of the abstraction already introduced here.

@DhineshPonnarasan

Copy link
Copy Markdown
Contributor Author

Thanks Ricky Gummadi (@Ricky-G) great suggestion.
Implemented in this PR.

What I changed:

  • Added shared helper methods on BaseIntegration for trusted-source extraction/filtering.
  • Replaced duplicated trusted-source tuple blocks across the touched adapters with helper calls.
  • Added helper-focused test coverage in the skill-audit helper tests.
  • This removes the repeated boilerplate and keeps trusted metadata extraction centralized so future adapter hooks stay consistent.

I also included small spell-check dictionary updates for intentional project terms that were blocking CI.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the contribution. The audit trail hardening changes look reasonable and well-tested (good anti-spoofing checks for skill metadata).

However, .cspell-repo-terms.txt\ contains unresolved merge conflict markers (<<<<<<<, =======, >>>>>>>). Please rebase on current main and resolve the conflicts before this can be merged.

@DhineshPonnarasan

Copy link
Copy Markdown
Contributor Author

Thanks for the contribution. The audit trail hardening changes look reasonable and well-tested (good anti-spoofing checks for skill metadata).

However, .cspell-repo-terms.txt\ contains unresolved merge conflict markers (<<<<<<<, =======, >>>>>>>). Please rebase on current main and resolve the conflicts before this can be merged.

Imran Siddique (@imran-siddique) Thanks for catching this. I rebased the branch on current main, resolved the conflict markers in .cspell-repo-terms.txt, and force-pushed the updated branch. Could you please take another look when you have a chance?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good work on this overall. The trust boundary design is solid and the spoof resistance tests across all six adapters are exactly what this needed. A few things worth addressing before merge:

Double computation of skill fields

In autogen_adapter.py (lines ~49 to 69) and google_adk_adapter.py before_tool_callback, build_skill_audit_fields is called first to get skill_fields, and then emit_skill_audit_event is called immediately after. The problem is that emit_skill_audit_event internally calls build_skill_audit_fields again. So the fields are computed twice for the same invocation. You could either have emit_skill_audit_event accept pre-built fields, or restructure so the caller reads the returned payload for log appending instead of calling build_skill_audit_fields separately.

UTC timestamp normalization is incomplete

The PR description lists UTC timestamp normalization as one of the three blockers being fixed, and autogen_adapter.py was updated correctly. But semantic_kernel_adapter.py line 876 and langchain_adapter.py _record_tool_invocation line 598 still use datetime.now().isoformat() with no timezone info. These need the same fix.

Silent exception in hash_context

except Exception:
    return None

The fail-safe behavior makes sense, but there is zero visibility into why a context value returned None. In production, if hashes are None across all events it will be impossible to tell whether it is by design or because payloads are consistently non-serializable. A single logger.debug call in that branch would help a lot without any real cost.

Asymmetric audit emission in google_adk_adapter.py

before_tool_callback calls both emit_skill_audit_event and _record. after_tool_callback only calls _record. This means consumers listening on GovernanceEventType.POLICY_CHECK will not receive a post-tool event from the ADK adapter, while they do from crewai (after_tool_call) and openai_agents (on_tool_end). The inconsistency across adapters for the same lifecycle position will cause surprises downstream.

provenance_source_trust as a string literal

provenance_source_trust: str | None = "trusted" if skill_name else None

Tests assert == "trusted". This works for now but if this field ever needs more granular values the string approach will cause inconsistent comparisons across consumers. A small Enum or at minimum Literal["trusted"] would make this safer with very little added complexity.

Minor observations

test_context_hash_stable_for_nested_ordering tests dict key ordering stability within list elements, not list element order stability. The name implies broader coverage than it provides. Worth a rename or a clarifying comment so someone does not assume list reordering is also stable (it is not).

The TrustedSkillMetadataSource constructible-by-anyone point is already acknowledged in the docstrings and is fine as a first-deliverable tradeoff. Just worth a TODO comment so it is not forgotten when signature-backed trust gets added later.

Overall the design is in good shape. Main blockers before merge in my view are the timestamp gap in SK and LangChain adapters, the asymmetric ADK emit, and the double computation issue.

@DhineshPonnarasan

Copy link
Copy Markdown
Contributor Author

Good work on this overall. The trust boundary design is solid and the spoof resistance tests across all six adapters are exactly what this needed. A few things worth addressing before merge:

Double computation of skill fields

In autogen_adapter.py (lines ~49 to 69) and google_adk_adapter.py before_tool_callback, build_skill_audit_fields is called first to get skill_fields, and then emit_skill_audit_event is called immediately after. The problem is that emit_skill_audit_event internally calls build_skill_audit_fields again. So the fields are computed twice for the same invocation. You could either have emit_skill_audit_event accept pre-built fields, or restructure so the caller reads the returned payload for log appending instead of calling build_skill_audit_fields separately.

UTC timestamp normalization is incomplete

The PR description lists UTC timestamp normalization as one of the three blockers being fixed, and autogen_adapter.py was updated correctly. But semantic_kernel_adapter.py line 876 and langchain_adapter.py _record_tool_invocation line 598 still use datetime.now().isoformat() with no timezone info. These need the same fix.

Silent exception in hash_context

except Exception:
    return None

The fail-safe behavior makes sense, but there is zero visibility into why a context value returned None. In production, if hashes are None across all events it will be impossible to tell whether it is by design or because payloads are consistently non-serializable. A single logger.debug call in that branch would help a lot without any real cost.

Asymmetric audit emission in google_adk_adapter.py

before_tool_callback calls both emit_skill_audit_event and _record. after_tool_callback only calls _record. This means consumers listening on GovernanceEventType.POLICY_CHECK will not receive a post-tool event from the ADK adapter, while they do from crewai (after_tool_call) and openai_agents (on_tool_end). The inconsistency across adapters for the same lifecycle position will cause surprises downstream.

provenance_source_trust as a string literal

provenance_source_trust: str | None = "trusted" if skill_name else None

Tests assert == "trusted". This works for now but if this field ever needs more granular values the string approach will cause inconsistent comparisons across consumers. A small Enum or at minimum Literal["trusted"] would make this safer with very little added complexity.

Minor observations

test_context_hash_stable_for_nested_ordering tests dict key ordering stability within list elements, not list element order stability. The name implies broader coverage than it provides. Worth a rename or a clarifying comment so someone does not assume list reordering is also stable (it is not).

The TrustedSkillMetadataSource constructible-by-anyone point is already acknowledged in the docstrings and is fine as a first-deliverable tradeoff. Just worth a TODO comment so it is not forgotten when signature-backed trust gets added later.

Overall the design is in good shape. Main blockers before merge in my view are the timestamp gap in SK and LangChain adapters, the asymmetric ADK emit, and the double computation issue.

Dipika Ranabhat (@qubeena07) Thanks for the detailed review. I’ve addressed all the points you called out in the latest update (commit 3d0cfe3):

  • Removed double computation of skill audit fields in AutoGen and ADK before_tool_callback by reusing emitted payload fields.
  • Completed UTC timestamp normalization for the previously missed SK and LangChain locations.
  • Added debug logging in the hash_context exception path so None hashes are observable.
  • Added centralized post-tool POLICY_CHECK emission in ADK after_tool_callback for lifecycle symmetry.
  • Tightened provenance_source_trust typing to Literal["trusted"].
  • Renamed the nested hash-ordering test for clarity and added a TODO note on future signature-backed trust attestation.

Focused tests for touched areas pass locally. Please take another look when you have time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All 7 claimed changes look good. The double computation removal for AutoGen and ADK is clean, UTC normalization is correct for SK and LangChain, debug logging in the hash_context exception path is useful, the Literal["trusted"] tightening is right, and the test rename is clearer.

One question before I can sign off: line 681 in intent.py changes the verify_intent guard from checking both EXECUTING and APPROVED states to only EXECUTING. Previously an intent in APPROVED state could be verified, now it raises IntentStateError. This is a behavior change not mentioned in the PR description. Was this intentional? If so please add a note explaining why APPROVED should no longer be a valid state for verification. If not, this line should be reverted.

Also noticed base.py still has naive datetime.now() calls without UTC at lines 1303, 1345, 1432, 1453. Not a blocker for this PR but worth a follow-up to keep things consistent with the normalization done here.

Comment thread agent-governance-python/agent-os/src/agent_os/intent.py

@imran-siddique Imran Siddique (imran-siddique) left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(duplicate)

@github-actions github-actions Bot added dependencies Pull requests that update a dependency file agent-mesh agent-mesh package agent-hypervisor agent-hypervisor package agent-sre agent-sre package security Security-related issues labels Jun 1, 2026
@Ricky-G

Copy link
Copy Markdown
Contributor

Review status: my and Imran Siddique (@imran-siddique)'s requested changes are addressed

Confirmed against the current branch:

Reviewer Requested change Status
Ricky Gummadi (@Ricky-G) Extract the repeated trusted-source boilerplate into a BaseIntegration helper Done. trusted_sources_from_attrs() was added in base.py and the adapters now call kernel.trusted_sources_from_attrs(...) instead of repeating the inline block.
Imran Siddique (@imran-siddique) Resolve the merge conflict markers in .cspell-repo-terms.txt Done. No conflict markers remain and main has been merged into the branch.

Both of these are resolved, so I am approving and merging.

Remaining feedback from Dipika Ranabhat (@qubeena07) (tracked, not blocking)

Dipika Ranabhat (@qubeena07)'s points (documenting the verify_intent lifecycle narrowing, finishing UTC normalization in base.py, and the optional enum for provenance_source_trust) are not blockers. The author already added an explanatory comment and replied in the thread. To keep this PR moving, I have moved these to a follow-up: #2908.

Scope note: this is Python only

The skill-aware audit trail hardening here lands only in the Python agent-os adapters. The TypeScript, .NET, Go, and Rust SDKs have framework discovery but do not yet implement these skill-audit primitives (trusted skill metadata source, provenance_source_trust, UTC-normalized audit timestamps, deterministic context hashing). Cross-language parity is tracked in #2907.

Thanks Dhinesh Ponnarasan (@DhineshPonnarasan) for the thorough work and tests.

@Ricky-G Ricky Gummadi (Ricky-G) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. My helper-extraction request and Imran Siddique (@imran-siddique)'s cspell conflict-marker fix are both addressed on the current branch. Dipika Ranabhat (@qubeena07)'s remaining points are tracked in #2908 and cross-language parity in #2907. This is Python only and good to merge.

@Ricky-G
Ricky Gummadi (Ricky-G) dismissed stale reviews from Imran Siddique (imran-siddique), Imran Siddique (imran-siddique), and Imran Siddique (imran-siddique) June 9, 2026 10:34

Resolved: the .cspell-repo-terms.txt merge conflict markers have been fixed and main has been merged into the branch. Dismissing this stale changes-requested review so the PR can merge. Remaining non-blocking follow-ups are tracked in #2908.

@Ricky-G
Ricky Gummadi (Ricky-G) enabled auto-merge (squash) June 9, 2026 10:37
@Ricky-G Ricky Gummadi (Ricky-G) added the stale Inactive issue/PR label Jun 9, 2026
Signed-off-by: Dhinesh Ponnarasan <dhineshponnarasan@gmail.com>
Signed-off-by: Dhinesh Ponnarasan <dhineshponnarasan@gmail.com>
Signed-off-by: Dhinesh Ponnarasan <dhineshponnarasan@gmail.com>
Signed-off-by: Dhinesh Ponnarasan <dhineshponnarasan@gmail.com>
Signed-off-by: Dhinesh Ponnarasan <dhineshponnarasan@gmail.com>
auto-merge was automatically disabled June 9, 2026 12:44

Head branch was pushed to by a user without write access

@DhineshPonnarasan

Copy link
Copy Markdown
Contributor Author

Hi Team,
Quick update: after auto-merge was enabled, CI surfaced a number of follow-up issues, including DCO and quality-gate failures, along with downstream CI checks that were blocked by those failures.

I investigated the reported issues, addressed the missing DCO signoff and the TODO quality-gate violation, validated the changes locally, and force-pushed the fixes.

The workflow approvals and reviews appear to have been marked stale automatically after the branch update. Aside from these CI-related fixes, the implementation itself remains unchanged.

Thank you again for the reviews and feedback.

@imran-siddique
Imran Siddique (imran-siddique) merged commit 16e392c into microsoft:main Jun 9, 2026
10 of 11 checks passed
Imran Siddique (imran-siddique) added a commit that referenced this pull request Jun 9, 2026
…regressions from #2572

* fix(agent-os): fill missing bodies in four empty if blocks in crewai_adapter

Commit 16e392c added four if statements with no body, causing
IndentationError at import time and failing lint, test, and
docker-compose-test jobs on main.

- allowed_tools check: return False when tool not in allowlist
- after_tool_call pattern match: raise PolicyViolationError
- before_llm_call pre_execute check: return False with log
- after_llm_call pattern match: raise PolicyViolationError

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>

* fix(agent-os): remove duplicate import blocks in google_adk, langchain, sk adapters

Same commit (16e392c) left duplicate import blocks in three adapters,
causing F811 ruff errors masked by the SyntaxError in crewai_adapter.

- google_adk_adapter: remove redundant second 'from .base import' line
- langchain_adapter: merge Callable, remove duplicate datetime/typing/.base imports
- semantic_kernel_adapter: same as langchain_adapter

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>

* fix(nexus): update test_client to use real Ed25519 keypairs for signing

PR #2782 made private_key_bytes required when local_mode=True registers
an agent (registry now verifies the manifest signature). Tests were not
updated and failed with ValueError on every register() call.

Use generate_keypair() to create a real keypair per test client and
set the manifest verification_key to the matching public key.

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>

* fix(nexus): update test_registry to use real Ed25519 keypairs for signing

Registry verifies signatures on register/deregister. Tests using fake
keys ("ed25519:test_key_123") and fake sigs ("sig") caused
InvalidSignatureError. Replace create_test_manifest with
create_signed_manifest returning a real keypair, and recompute the
deregister signature as sign(private_key_bytes, agent_did.encode()).

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(nexus): wire escrow and DMZ signing in NexusClient; fix ProofOfOutcome tests

client.create_escrow now passes requester_signature via _sign_escrow().
sign_dmz_policy replaced nonexistent _generate_signature with an inline
sign(private_key_bytes, transfer_id.encode()) call.
TestProofOfOutcome tests generate a real keypair and pass requester_signature
to poo.create_escrow(), satisfying the ValueError guard added in 16e392c.

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Imran Siddique (imran-siddique) added a commit that referenced this pull request Jun 9, 2026
Promote all 4.0.x packages to 4.1.0 to reflect the changes since
the v4.0.0 release on 2026-06-01:
- skill-aware audit trail hardening (#2572)
- Ed25519 signature verification in Nexus registry/escrow (#2782)
- ring enforcement wiring in sandbox providers (#2868)
- dynamic policy conditions: time-based, cost-aware, quota-aware (#2870)
- policy regression testing framework (agt test) (#2869)
- crewai adapter if-body syntax fix (#2911)
- various security fixes, dependabot updates

Also promote agt-policies from 5.0.0a1 to 5.0.0 (alpha has been
live on PyPI since the June 9 dry-run confirmed the build was clean).

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>
Imran Siddique (imran-siddique) added a commit that referenced this pull request Jun 9, 2026
Promote all 4.0.x packages to 4.1.0 to reflect the changes since
the v4.0.0 release on 2026-06-01:
- skill-aware audit trail hardening (#2572)
- Ed25519 signature verification in Nexus registry/escrow (#2782)
- ring enforcement wiring in sandbox providers (#2868)
- dynamic policy conditions: time-based, cost-aware, quota-aware (#2870)
- policy regression testing framework (agt test) (#2869)
- crewai adapter if-body syntax fix (#2911)
- various security fixes, dependabot updates

Also promote agt-policies from 5.0.0a1 to 5.0.0 (alpha has been
live on PyPI since the June 9 dry-run confirmed the build was clean).

Signed-off-by: Imran Siddique <imran.siddique@opaque.co>
Ricky Gummadi (Ricky-G) added a commit that referenced this pull request Jun 13, 2026
…ization in base.py

- Add CHANGELOG entry under [Unreleased] > Changed documenting that
  verify_intent now requires EXECUTING state only (previously APPROVED
  was also accepted). The lifecycle is strictly declare -> approve ->
  execute -> verify. Closes #2908.
- Expand the guard comment in intent.py to explain why APPROVED is
  rejected: without an execution phase there are no records to compare
  against the declared plan.
- Replace all 5 naive datetime.now() calls in integrations/base.py with
  datetime.now(timezone.utc) for consistency with UTC normalization
  done in the adapters (PR #2572):
  - ExecutionContext.start_time default factory
  - event_base timestamp in _run_policy_checks
  - elapsed-time / timeout comparison
  - DRIFT_DETECTED event timestamp
  - CHECKPOINT_CREATED event timestamp

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>
Imran Siddique (imran-siddique) pushed a commit that referenced this pull request Jun 14, 2026
…ization in base.py (#2994)

* fix: document verify_intent lifecycle narrowing and finish UTC normalization in base.py

- Add CHANGELOG entry under [Unreleased] > Changed documenting that
  verify_intent now requires EXECUTING state only (previously APPROVED
  was also accepted). The lifecycle is strictly declare -> approve ->
  execute -> verify. Closes #2908.
- Expand the guard comment in intent.py to explain why APPROVED is
  rejected: without an execution phase there are no records to compare
  against the declared plan.
- Replace all 5 naive datetime.now() calls in integrations/base.py with
  datetime.now(timezone.utc) for consistency with UTC normalization
  done in the adapters (PR #2572):
  - ExecutionContext.start_time default factory
  - event_base timestamp in _run_policy_checks
  - elapsed-time / timeout comparison
  - DRIFT_DETECTED event timestamp
  - CHECKPOINT_CREATED event timestamp

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add fastembed to REGISTERED_PACKAGES in dep-confusion scan

fastembed is a real PyPI package (fast embedding generation from Qdrant)
referenced in agent-os/pyproject.toml line 61. The dep-confusion scanner
was flagging it as unregistered because it was missing from REGISTERED_PACKAGES.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add fastembed to cspell word list and clean up comment

- Add 'fastembed' and 'FastEmbed' to .cspell.json words list so the
  spell checker does not flag the package name string and comment
- Remove brand name from dep-confusion comment to avoid future
  spell-check churn on proper nouns

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: update tests to use timezone-aware datetimes for start_time

The UTC normalization of ExecutionContext.start_time (datetime.now() ->
datetime.now(timezone.utc)) broke tests that explicitly set start_time
to a naive datetime. Update all four affected test files to use
datetime.now(timezone.utc) so the subtraction in _run_policy_checks
does not raise TypeError.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: defensive UTC normalization for naive start_time and drop out-of-scope changes

Address review feedback from @imran-siddique:

1. Add defensive normalization at base.py timeout check: if ctx.start_time
   is naive (set by external callers before the default factory was made
   tz-aware), convert it to UTC via astimezone() before subtracting.
   astimezone() correctly interprets the naive value as local time and
   converts it to UTC; replace(tzinfo=...) would not adjust the value.
   This protects external callers that still pass naive datetimes without
   breaking the UTC-aware fast path.

2. Add test_blocked_when_timeout_exceeded_naive_start_time to explicitly
   cover the defensive normalization path with a naive start_time.

3. Revert .cspell.json and scripts/check_dependency_confusion.py to main
   (fastembed additions are already on main and out of scope for #2908).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add datetimes to cspell wordlist

The word 'datetimes' appears in the defensive normalization comment added
in base.py and is not in the default cspell dictionary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

---------

Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
jlaportebot (jlaportebot) pushed a commit to jlaportebot/agent-governance-toolkit that referenced this pull request Jun 17, 2026
…ization in base.py (microsoft#2994)

* fix: document verify_intent lifecycle narrowing and finish UTC normalization in base.py

- Add CHANGELOG entry under [Unreleased] > Changed documenting that
  verify_intent now requires EXECUTING state only (previously APPROVED
  was also accepted). The lifecycle is strictly declare -> approve ->
  execute -> verify. Closes microsoft#2908.
- Expand the guard comment in intent.py to explain why APPROVED is
  rejected: without an execution phase there are no records to compare
  against the declared plan.
- Replace all 5 naive datetime.now() calls in integrations/base.py with
  datetime.now(timezone.utc) for consistency with UTC normalization
  done in the adapters (PR microsoft#2572):
  - ExecutionContext.start_time default factory
  - event_base timestamp in _run_policy_checks
  - elapsed-time / timeout comparison
  - DRIFT_DETECTED event timestamp
  - CHECKPOINT_CREATED event timestamp

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add fastembed to REGISTERED_PACKAGES in dep-confusion scan

fastembed is a real PyPI package (fast embedding generation from Qdrant)
referenced in agent-os/pyproject.toml line 61. The dep-confusion scanner
was flagging it as unregistered because it was missing from REGISTERED_PACKAGES.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add fastembed to cspell word list and clean up comment

- Add 'fastembed' and 'FastEmbed' to .cspell.json words list so the
  spell checker does not flag the package name string and comment
- Remove brand name from dep-confusion comment to avoid future
  spell-check churn on proper nouns

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: update tests to use timezone-aware datetimes for start_time

The UTC normalization of ExecutionContext.start_time (datetime.now() ->
datetime.now(timezone.utc)) broke tests that explicitly set start_time
to a naive datetime. Update all four affected test files to use
datetime.now(timezone.utc) so the subtraction in _run_policy_checks
does not raise TypeError.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: defensive UTC normalization for naive start_time and drop out-of-scope changes

Address review feedback from @imran-siddique:

1. Add defensive normalization at base.py timeout check: if ctx.start_time
   is naive (set by external callers before the default factory was made
   tz-aware), convert it to UTC via astimezone() before subtracting.
   astimezone() correctly interprets the naive value as local time and
   converts it to UTC; replace(tzinfo=...) would not adjust the value.
   This protects external callers that still pass naive datetimes without
   breaking the UTC-aware fast path.

2. Add test_blocked_when_timeout_exceeded_naive_start_time to explicitly
   cover the defensive normalization path with a naive start_time.

3. Revert .cspell.json and scripts/check_dependency_confusion.py to main
   (fastembed additions are already on main and out of scope for microsoft#2908).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

* fix: add datetimes to cspell wordlist

The word 'datetimes' appears in the defensive normalization comment added
in base.py and is not in the default cspell dictionary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>

---------

Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: jlaportebot <jlaportebot@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation needs-review:MEDIUM Contributor check flagged MEDIUM risk scripts/ci/cd size/XL Extra large PR (500+ lines) stale Inactive issue/PR tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

RFC: Govern agent Skills (context injection) across ADK, CrewAI, Semantic Kernel, and OpenAI Agents SDK

4 participants