Skip to content

feat(agent-compliance): PromptDefense + GovernanceVerifier integration example - #1837

Merged
Imran Siddique (imran-siddique) merged 1 commit into
microsoft:mainfrom
lawcontinue:feat/prompt-defense-governance-example
May 9, 2026
Merged

Imran Siddique (imran-siddique) merged 1 commit into
microsoft:mainfrom
lawcontinue:feat/prompt-defense-governance-example

Conversation

@lawcontinue

Copy link
Copy Markdown
Contributor

Summary

Adds an integration example demonstrating how PromptDefenseEvaluator works with GovernanceVerifier for pre-deployment governance checks.

Follow-up to the co-authoring discussion in #854 with MinYi Xie (@ppcvote).

What it shows

  1. Weak vs strong prompt comparison — grade F (1/12) blocked vs grade A (11/12) allowed
  2. Batch evaluation — using evaluate_batch() across multiple agents
  3. Audit entry generation — to_audit_entry() output for MerkleAuditChain integration
  4. Deployment gating — is_blocking() decision based on configurable min grade

Integration points

  • PromptDefenseEvaluator → evaluate() / evaluate_batch() / evaluate_file()
  • PromptDefenseReport → is_blocking() for deployment decisions
  • to_audit_entry() → MerkleAuditChain compatible audit log entries
  • GovernanceVerifier → verify() + evidence_checks for full governance attestation

Testing

$ pip install agent-governance-toolkit
$ python prompt_defense_governance.py

--- Demo 1: Weak prompt ---
Grade: F | Score: 8 | Coverage: 1/12 | Decision: BLOCKED

--- Demo 2: Well-defended prompt ---
Grade: A | Score: 92 | Coverage: 11/12 | Decision: ALLOWED

--- Demo 3: Batch evaluation ---
  ❌ chatbot-agent: Grade F (1/12)
  ✅ data-query-agent: Grade A (11/12)

--- Demo 4: Audit entry (JSON) ---
{"event_type": "prompt.defense.evaluated", "agent_did": "agent:demo-002", ...}

Checklist

  • Code follows existing example conventions (governed_agent.py, quickstart.py)
  • MIT license header included
  • Works with graceful fallback when optional dependencies unavailable
  • No new dependencies required

…ation example

Demonstrates how PromptDefenseEvaluator integrates with GovernanceVerifier
for pre-deployment governance checks:

- Scan system prompts for missing defenses (12 attack vectors)
- Block deployment if prompt grade falls below threshold
- Generate audit entries for MerkleAuditChain
- Batch evaluation across multiple agents

Includes 4 progressive demos: weak vs strong prompt comparison,
batch evaluation, and audit entry generation.

Follow-up to microsoft#854 discussion with @ppcvote.
@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown

Welcome to the Agent Governance Toolkit! Thanks for your first pull request.
Please ensure tests pass, code follows style (ruff check), and you have signed the CLA.
See our Contributing Guide.

@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown
🤖 AI Agent: breaking-change-detector — API Compatibility

API Compatibility

No breaking changes detected.

@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown
🤖 AI Agent: docs-sync-checker — Docs Sync

Docs Sync

  • Documentation is in sync.

@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown
🤖 AI Agent: security-scanner — View details

No security issues found.

@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown
🤖 AI Agent: test-generator — `prompt_defense_governance.py`

prompt_defense_governance.py

  • test_evaluate_weak_prompt -- Verify PromptDefenseEvaluator.evaluate() correctly identifies and blocks a weak prompt.
  • test_evaluate_strong_prompt -- Verify PromptDefenseEvaluator.evaluate() correctly allows a strong prompt.
  • test_evaluate_batch -- Validate evaluate_batch() processes multiple agents and returns correct results.
  • test_to_evidence_check_error_handling -- Test to_evidence_check() behavior when PromptDefenseEvaluator is unavailable or fails.
  • test_governance_integration -- Ensure GovernanceVerifier.verify() integrates correctly with prompt defense results.

@github-actions github-actions Bot added the size/L Large PR (< 500 lines) label May 9, 2026
@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown
🤖 AI Agent: code-reviewer — Action Items:

TL;DR: 0 blockers, 1 warning. Example implementation is functional and well-structured, but a minor improvement is suggested.

# Sev Issue Where
1 Warn Missing test coverage for run_prompt_defense_governance function. prompt_defense_governance.py

Action Items:

  • None.

Warnings:

# Issue Where Follow-up
1 Add unit tests for run_prompt_defense_governance to ensure robustness. prompt_defense_governance.py Fine as follow-up PR.

@github-actions

github-actions Bot commented May 9, 2026

Copy link
Copy Markdown

PR Review Summary

Check Status Details
🔍 Code Review ⚠️ Warning See details
🛡️ Security Scan ✅ Passed No issues found
🔄 Breaking Changes ✅ Passed No issues found
📝 Docs Sync ✅ Passed No issues found
🧪 Test Coverage ✅ Completed Analysis complete

Verdict: ⚠️ Ready for human review

@ppcvote

Copy link
Copy Markdown
Contributor

lawcontinue — thanks for taking this through to a concrete integration example. Seeing PromptDefenseEvaluator plug directly into GovernanceVerifier.verify() + MerkleAuditChain is exactly what the v0.x usage docs were missing.

A few notes that might tighten the demo, drawn from the production gap data we baselined this April (gap-20260405.json — 1,646 prompts, 4 sources):

  1. Threshold tuning evidence. The current min_grade='B' decision boundary is reasonable but conservative. Empirically, 78.3% of production prompts in the corpus score F (<45/100). For the demo's "weak prompt" (1/12 coverage), a more representative comparison might be a "median real prompt" (~36/100) — it shows the gate triggers on something most deployments would actually have, rather than a strawman.

  2. v1.5 vectors land in May 2026. Three new vectors went into prompt-defense-audit v1.5 last week: skill-provenance (CurXecute / CVE-2025-54135 class), least-agency (Replit prod-DB-wipe class), implicit-memory-hygiene (cross-session memory poisoning). The integration example would benefit from including at least least-agency since it's the most agentic-toolkit-relevant — agents that can take destructive actions need this in evidence_checks.

  3. Mandarin / multilingual extension hook. The current 12-vector regex coverage is English + Traditional Chinese. For Microsoft customers in APAC (especially regulated TW finance / healthcare), there's a real gap — multilingual prompt injection success rates match English per arxiv:2512.23684, but the keyword filters are language-asymmetric. Happy to PR a MultilingualPromptDefenseEvaluator companion if the toolkit roadmap has space for it.

I'd be glad to follow up with a PR adding (1) a "median prompt" demo case using anonymized gap-corpus data and (2) the v1.5 vector pass — let me know if either fits the project direction.

— Min Yi (ppcvote / Ultra Lab)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean integration example. Good use of fallback imports for standalone usage.

@imran-siddique
Imran Siddique (imran-siddique) merged commit ddb581d into microsoft:main May 9, 2026
13 of 14 checks passed
MohammadHaroonAbuomar pushed a commit to MohammadHaroonAbuomar/agt-acs that referenced this pull request Jun 1, 2026
…ation example (microsoft#1837)

Demonstrates how PromptDefenseEvaluator integrates with GovernanceVerifier
for pre-deployment governance checks:

- Scan system prompts for missing defenses (12 attack vectors)
- Block deployment if prompt grade falls below threshold
- Generate audit entries for MerkleAuditChain
- Batch evaluation across multiple agents

Includes 4 progressive demos: weak vs strong prompt comparison,
batch evaluation, and audit entry generation.

Follow-up to microsoft#854 discussion with @ppcvote.

Co-authored-by: deepsearch <deepsearch@deepsearchdeMac-mini.local>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/L Large PR (< 500 lines)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants