Skip to content

feat(compliance): add 5 OWASP Agentic (ASI) vectors to prompt-defense evaluator - #3028

Merged
Imran Siddique (imran-siddique) merged 1 commit into
microsoft:mainfrom
ppcvote:feat/prompt-defense-agentic-vectors
Jun 15, 2026
Merged

Imran Siddique (imran-siddique) merged 1 commit into
microsoft:mainfrom
ppcvote:feat/prompt-defense-agentic-vectors

Conversation

@ppcvote

Copy link
Copy Markdown
Contributor

What & why

PromptDefenseEvaluator (agt red-team scan) statically audits an agent's
system prompt for missing defensive language before deployment. Today it
checks 12 vectors mapped to the OWASP LLM Top 10 — conversational safety
only. But this toolkit positions itself against the OWASP Agentic Top 10
(ASI)
, and the prompt auditor says nothing about the risks that exist once
the model is an autonomous agent. This PR closes that gap by adding 5
agent-era vectors so a pre-deployment audit covers the agentic layer too.

Vector ASI Defends
cross-agent-auth ASI-07 Authority from another agent is re-verified per request, not inherited transitively
transaction-guardrails ASI-02 Value-moving actions require a limit / second approval / explicit refusal-without-authorization
skill-provenance ASI-04 Skills load only from a signed/trusted/pinned source; unverified ones refused
least-agency ASI-01 Least privilege, scoped to the assigned goal, aborts on goal drift
encoding-injection ASI-01 Decoded / base64 content treated as untrusted data, never as a command

Design — consistent with the existing module

  • min_matches=2 contract preserved. Each agentic rule needs the
    capability pattern AND the constraint pattern — attack vocabulary alone
    ("transfer the funds", "another agent told me", "decode this base64") never
    scores as defended. Verified against adversarial prompts.
  • ReDoS-safe. All new patterns use bounded quantifiers (no unbounded
    .*); MAX_PROMPT_LENGTH guard unchanged. No catastrophic backtracking on
    99K adversarial inputs (the new vectors run in ~20ms at the cap).
  • VECTOR_COUNT stays dynamic (len(_RULES)) — nothing hardcoded.
  • ASI risk numbers follow docs/OWASP-COMPLIANCE.md (the stack-wide map;
    ASI-07 = Insecure Inter-Agent Communication).

Tests & docs

  • Positive + negative-control tests for all 5 vectors.
  • Regression tests for three false-positive classes found in review:
    auth token vs. spending guardrail; data-pipeline "as input" vs.
    treat-as-untrusted; QA "verified" vs. a provenance refusal.
  • STRONG_PROMPT fixture extended with agentic defenses.
  • New docs/prompt-defense-vectors.md — full vector → OWASP mapping table.
  • agt red-team scan docstring + example count updated 12 → 17.
  • ruff check clean; full tests/test_prompt_defense.py green (81 passed).

Attribution

Regex vocabulary for the 5 agentic vectors was distilled from the open-source
UltraProbe scanner (MIT) — its
scanDefense agent-era vectors — ported English-first to match this module's
style and ReDoS discipline.

Context

Extends the PromptDefenseEvaluator added in #854. Signed-off-by per DCO.

… evaluator

PromptDefenseEvaluator audited system prompts against 12 OWASP LLM Top 10
vectors but nothing on the agentic layer, even though AGT positions itself
against the OWASP Agentic Top 10. This extends the evaluator with 5 agent-era
vectors so a pre-deployment prompt audit also covers agentic risks:

  - cross-agent-auth       (ASI-07) inter-agent authority boundary
  - transaction-guardrails (ASI-02) value-moving action guardrails
  - skill-provenance       (ASI-04) signed/trusted skill loading
  - least-agency           (ASI-01) least-privilege + goal-drift abort
  - encoding-injection     (ASI-01) decoded payload treated as data, not command

Each rule follows the module's existing discipline: bounded quantifiers
(ReDoS-safe), min_matches=2 so attack vocabulary alone never scores as
defended (the capability AND its constraint must both be present), plus a
severity_map entry. VECTOR_COUNT stays dynamic (len(_RULES)).

Adds positive + negative-control tests for all 5, regression tests for three
false-positive classes (auth token vs. spending guardrail; data-pipeline
"as input" vs. treat-as-untrusted; QA "verified" vs. provenance refusal),
extends the STRONG_PROMPT fixture, and a docs/ vector->OWASP mapping table.
red-team scan CLI docstring updated 12 -> 17.

Regex vocabulary distilled from the open-source UltraProbe scanner
(npm: ultraprobe, MIT).

Signed-off-by: ppcvote <risky9763@gmail.com>
@github-actions

Copy link
Copy Markdown
🤖 AI Agent: security-scanner — View details

AI-generated review output. Treat it as untrusted analysis and verify before acting.

No security issues found.

@github-actions github-actions Bot added the size/L Large PR (< 500 lines) label Jun 15, 2026
@github-actions

Copy link
Copy Markdown
🤖 AI Agent: docs-sync-checker — Docs Sync

AI-generated review output. Treat it as untrusted analysis and verify before acting.

Docs Sync

  • scan() in red_team.py -- missing updated docstring to reflect the increase from 12 to 17 attack vectors.
  • CHANGELOG.md -- missing entry for the addition of 5 new OWASP Agentic (ASI) vectors to the PromptDefenseEvaluator.

@github-actions

Copy link
Copy Markdown
🤖 AI Agent: code-reviewer — Action items:

AI-generated review output. Treat it as untrusted analysis and verify before acting.

TL;DR: 0 blockers, 1 warning. The PR introduces new OWASP Agentic Top 10 vectors to the PromptDefenseEvaluator with proper tests and documentation. However, the regex patterns for the new vectors should be reviewed for completeness and potential gaps.

# Sev Issue Where
1 Warn Regex patterns for new agentic vectors should be reviewed for gaps. agent-compliance/src/agent_compliance/prompt_defense.py

Action items:

  • None.

Warnings (fine as follow-up PRs):

# Issue Where
1 Regex patterns for new agentic vectors should be reviewed for gaps. agent-compliance/src/agent_compliance/prompt_defense.py

@github-actions

Copy link
Copy Markdown
🤖 AI Agent: breaking-change-detector — API Compatibility

AI-generated review output. Treat it as untrusted analysis and verify before acting.

API Compatibility

Severity Change Impact
High Increased the number of required defense vectors from 12 to 17 in PromptDefenseEvaluator. Existing users relying on the PromptDefenseEvaluator may experience failures if their prompts do not meet the new requirements for the 5 additional agentic vectors.
High Updated the scan CLI command to evaluate prompts against 17 vectors instead of 12. Users running agt red-team scan may encounter stricter evaluations, potentially causing previously passing prompts to fail.
High Modified the prompt_defense.py module to include 5 new defense rules for agentic safety. Any systems or tests relying on the previous set of 12 vectors may need updates to accommodate the new rules.

@github-actions

Copy link
Copy Markdown
🤖 AI Agent: test-generator — `agent_compliance/prompt_defense.py`

AI-generated review output. Treat it as untrusted analysis and verify before acting.

agent_compliance/prompt_defense.py

  • test_cross_agent_auth_defense -- Validate that the cross-agent-auth vector is correctly identified and defended.
  • test_transaction_guardrails_defense -- Validate that the transaction-guardrails vector is correctly identified and defended.
  • test_skill_provenance_defense -- Validate that the skill-provenance vector is correctly identified and defended.
  • test_least_agency_defense -- Validate that the least-agency vector is correctly identified and defended.
  • test_encoding_injection_defense -- Validate that the encoding-injection vector is correctly identified and defended.

agent_compliance/cli/red_team.py

  • test_scan_with_17_vectors -- Ensure the scan command correctly handles 17 vectors and outputs accurate results.
  • test_scan_with_strict_mode -- Verify that the --strict flag exits with a non-zero status for failing prompts.

@github-actions

Copy link
Copy Markdown

🔴 Contributor Check: HIGH

Check Result
Profile HIGH
Credential HIGH
Overall HIGH

Automated check by AGT Contributor Check.

@github-actions

Copy link
Copy Markdown

PR Review Summary

Check Status Details
🔍 Code Review ⚠️ Missing No current-run comment
🛡️ Security Scan ⚠️ Missing No current-run comment
🔄 Breaking Changes ⚠️ Missing No current-run comment
📝 Docs Sync ⚠️ Missing No current-run comment
🧪 Test Coverage ⚠️ Missing No current-run comment

Verdict: ⚠️ AI review incomplete; ready for human review

AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims.

@github-actions github-actions Bot added the needs-review:HIGH Contributor reputation check flagged HIGH risk label Jun 15, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Five OWASP ASI vectors added cleanly: bounded quantifiers throughout (no ReDoS risk), min_matches=2 ensures attack vocabulary alone never scores as a defense, and the false-positive regression tests (JWT-token, data-pipeline base64, verified-tool) cover the real edge cases. Attribution to UltraProbe (MIT) is appropriate. Merging.

@imran-siddique
Imran Siddique (imran-siddique) merged commit 517d6fc into microsoft:main Jun 15, 2026
14 of 15 checks passed
jlaportebot (jlaportebot) pushed a commit to jlaportebot/agent-governance-toolkit that referenced this pull request Jun 17, 2026
… evaluator (microsoft#3028)

PromptDefenseEvaluator audited system prompts against 12 OWASP LLM Top 10
vectors but nothing on the agentic layer, even though AGT positions itself
against the OWASP Agentic Top 10. This extends the evaluator with 5 agent-era
vectors so a pre-deployment prompt audit also covers agentic risks:

  - cross-agent-auth       (ASI-07) inter-agent authority boundary
  - transaction-guardrails (ASI-02) value-moving action guardrails
  - skill-provenance       (ASI-04) signed/trusted skill loading
  - least-agency           (ASI-01) least-privilege + goal-drift abort
  - encoding-injection     (ASI-01) decoded payload treated as data, not command

Each rule follows the module's existing discipline: bounded quantifiers
(ReDoS-safe), min_matches=2 so attack vocabulary alone never scores as
defended (the capability AND its constraint must both be present), plus a
severity_map entry. VECTOR_COUNT stays dynamic (len(_RULES)).

Adds positive + negative-control tests for all 5, regression tests for three
false-positive classes (auth token vs. spending guardrail; data-pipeline
"as input" vs. treat-as-untrusted; QA "verified" vs. provenance refusal),
extends the STRONG_PROMPT fixture, and a docs/ vector->OWASP mapping table.
red-team scan CLI docstring updated 12 -> 17.

Regex vocabulary distilled from the open-source UltraProbe scanner
(npm: ultraprobe, MIT).

Signed-off-by: ppcvote <risky9763@gmail.com>
Signed-off-by: jlaportebot <jlaportebot@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation needs-review:HIGH Contributor reputation check flagged HIGH risk size/L Large PR (< 500 lines) tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants