Repository navigation
feat(compliance): add 5 OWASP Agentic (ASI) vectors to prompt-defense evaluator - #3028
Conversation
… evaluator PromptDefenseEvaluator audited system prompts against 12 OWASP LLM Top 10 vectors but nothing on the agentic layer, even though AGT positions itself against the OWASP Agentic Top 10. This extends the evaluator with 5 agent-era vectors so a pre-deployment prompt audit also covers agentic risks: - cross-agent-auth (ASI-07) inter-agent authority boundary - transaction-guardrails (ASI-02) value-moving action guardrails - skill-provenance (ASI-04) signed/trusted skill loading - least-agency (ASI-01) least-privilege + goal-drift abort - encoding-injection (ASI-01) decoded payload treated as data, not command Each rule follows the module's existing discipline: bounded quantifiers (ReDoS-safe), min_matches=2 so attack vocabulary alone never scores as defended (the capability AND its constraint must both be present), plus a severity_map entry. VECTOR_COUNT stays dynamic (len(_RULES)). Adds positive + negative-control tests for all 5, regression tests for three false-positive classes (auth token vs. spending guardrail; data-pipeline "as input" vs. treat-as-untrusted; QA "verified" vs. provenance refusal), extends the STRONG_PROMPT fixture, and a docs/ vector->OWASP mapping table. red-team scan CLI docstring updated 12 -> 17. Regex vocabulary distilled from the open-source UltraProbe scanner (npm: ultraprobe, MIT). Signed-off-by: ppcvote <risky9763@gmail.com>
🤖 AI Agent: security-scanner — View details
No security issues found. |
🤖 AI Agent: docs-sync-checker — Docs Sync
Docs Sync
|
🤖 AI Agent: code-reviewer — Action items:
TL;DR: 0 blockers, 1 warning. The PR introduces new OWASP Agentic Top 10 vectors to the
Action items:
Warnings (fine as follow-up PRs):
|
🤖 AI Agent: breaking-change-detector — API Compatibility
API Compatibility
|
🤖 AI Agent: test-generator — `agent_compliance/prompt_defense.py`
|
|
🔴 Contributor Check: HIGH
Automated check by AGT Contributor Check. |
PR Review Summary
Verdict: AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims. |
Imran Siddique (imran-siddique)
left a comment
There was a problem hiding this comment.
Reviewed. Five OWASP ASI vectors added cleanly: bounded quantifiers throughout (no ReDoS risk), min_matches=2 ensures attack vocabulary alone never scores as a defense, and the false-positive regression tests (JWT-token, data-pipeline base64, verified-tool) cover the real edge cases. Attribution to UltraProbe (MIT) is appropriate. Merging.
517d6fc
into
microsoft:main
… evaluator (microsoft#3028) PromptDefenseEvaluator audited system prompts against 12 OWASP LLM Top 10 vectors but nothing on the agentic layer, even though AGT positions itself against the OWASP Agentic Top 10. This extends the evaluator with 5 agent-era vectors so a pre-deployment prompt audit also covers agentic risks: - cross-agent-auth (ASI-07) inter-agent authority boundary - transaction-guardrails (ASI-02) value-moving action guardrails - skill-provenance (ASI-04) signed/trusted skill loading - least-agency (ASI-01) least-privilege + goal-drift abort - encoding-injection (ASI-01) decoded payload treated as data, not command Each rule follows the module's existing discipline: bounded quantifiers (ReDoS-safe), min_matches=2 so attack vocabulary alone never scores as defended (the capability AND its constraint must both be present), plus a severity_map entry. VECTOR_COUNT stays dynamic (len(_RULES)). Adds positive + negative-control tests for all 5, regression tests for three false-positive classes (auth token vs. spending guardrail; data-pipeline "as input" vs. treat-as-untrusted; QA "verified" vs. provenance refusal), extends the STRONG_PROMPT fixture, and a docs/ vector->OWASP mapping table. red-team scan CLI docstring updated 12 -> 17. Regex vocabulary distilled from the open-source UltraProbe scanner (npm: ultraprobe, MIT). Signed-off-by: ppcvote <risky9763@gmail.com> Signed-off-by: jlaportebot <jlaportebot@gmail.com>
What & why
PromptDefenseEvaluator(agt red-team scan) statically audits an agent'ssystem prompt for missing defensive language before deployment. Today it
checks 12 vectors mapped to the OWASP LLM Top 10 — conversational safety
only. But this toolkit positions itself against the OWASP Agentic Top 10
(ASI), and the prompt auditor says nothing about the risks that exist once
the model is an autonomous agent. This PR closes that gap by adding 5
agent-era vectors so a pre-deployment audit covers the agentic layer too.
cross-agent-authtransaction-guardrailsskill-provenanceleast-agencyencoding-injectionDesign — consistent with the existing module
min_matches=2contract preserved. Each agentic rule needs thecapability pattern AND the constraint pattern — attack vocabulary alone
("transfer the funds", "another agent told me", "decode this base64") never
scores as defended. Verified against adversarial prompts.
.*);MAX_PROMPT_LENGTHguard unchanged. No catastrophic backtracking on99K adversarial inputs (the new vectors run in ~20ms at the cap).
VECTOR_COUNTstays dynamic (len(_RULES)) — nothing hardcoded.docs/OWASP-COMPLIANCE.md(the stack-wide map;ASI-07 = Insecure Inter-Agent Communication).
Tests & docs
auth token vs. spending guardrail; data-pipeline "as input" vs.
treat-as-untrusted; QA "verified" vs. a provenance refusal.
STRONG_PROMPTfixture extended with agentic defenses.docs/prompt-defense-vectors.md— full vector → OWASP mapping table.agt red-team scandocstring + example count updated 12 → 17.ruff checkclean; fulltests/test_prompt_defense.pygreen (81 passed).Attribution
Regex vocabulary for the 5 agentic vectors was distilled from the open-source
UltraProbe scanner (MIT) — its
scanDefenseagent-era vectors — ported English-first to match this module'sstyle and ReDoS discipline.
Context
Extends the
PromptDefenseEvaluatoradded in #854. Signed-off-by per DCO.