An agent skill for opencode and Claude Code: a complete, signal-driven methodology for security-testing LLM applications, RAG pipelines, AI agents, and MCP tools — plus the defensive, governance, and general-security context around them, grounded in modern frameworks.
- 176 catalog techniques across 19 sections, organized by the five stages where untrusted content or model output crosses a trust boundary.
- 240 ready-to-use payloads using safe
GSL-*canary markers. - Modern (2024–2026) additions — new techniques, agentic/MCP coverage, and model/data-extraction attacks.
- Defense & governance — CaMeL, dual-LLM, plan-then-execute, information-flow control; OWASP Agentic/MCP hardening; guardrail, eval, and observability tooling.
- Framework grounding — OWASP LLM Top 10 (2026), OWASP Agentic Top 10 (2026), OWASP MCP Top 10, MITRE ATLAS, NIST AI RMF + Generative AI Profile, ISO/IEC 42001, EU AI Act, OWASP Top 10 (2025) / ASVS 5.0 / API Security Top 10, CWE Top 25.
- General application security — authN/authZ, injection, secrets, supply chain, cloud/K8s, logging/IR, privacy & compliance, secure SDLC.
- Severity + reporting — 13-field finding template and framework mapping.
Authorized testing only. Use these techniques and payloads solely on systems you own or have explicit written permission to test. The maintainers accept no liability for misuse.
ai-llm-security-testing/
SKILL.md # opencode skill — methodology, grounding, engagement workflow, index
references/ # shared reference content (see table below)
.claude/skills/ai-llm-security-testing/
SKILL.md # Claude Code skill (same content)
references/ # bundled copy for a self-contained Claude Code install
| File | Purpose |
|---|---|
foundations.md |
Framework grounding, control layers, deterministic vs probabilistic, trust-boundary model |
standards.md |
OWASP LLM/Agentic/MCP, MITRE ATLAS + mitigations, OWASP Top 10 2025, ASVS 5.0, API 2023, CWE Top 25 |
attack-surface.md |
Recon questions + risk-by-capability |
techniques.md |
Full 176-technique catalog (Test / Signal / pivot) |
payloads.md |
240 payloads + encoder/mutator patterns |
modern-techniques.md |
2024–2026 technique additions, confidence-tagged |
agentic-security.md |
OWASP ASI + MCP Top 10, MCP hardening, identity/delegation, HITL, browser/coding agents |
defenses.md |
Architecture-first mitigations + threat→defense map |
security-tooling.md |
Red-team, guardrail, observability tooling + benchmarks |
general-security.md |
General AppSec / privacy / compliance with AI relevance |
incidents.md |
Notable 2024–2026 AI incidents/CVEs with lessons |
red-team-methodology.md |
Scoping, execution, evidence capture, reporting |
reporting.md |
Severity scenarios, report template, finding-type → report-language table |
The repo root is the opencode skill folder. Drop it into your global skills directory:
git clone https://github.com/TheRealAlexV/ai-llm-security-testing.git \
~/.config/opencode/skills/ai-llm-security-testing
# or symlink an existing checkout
ln -s "$(pwd)/ai-llm-security-testing" ~/.config/opencode/skills/ai-llm-security-testingAlternatively, register the path in opencode.json:
{ "skills": { "paths": ["/abs/path/to/ai-llm-security-testing"] } }Restart opencode after installing — skills load at startup.
Install the bundled Claude Code skill into ~/.claude/skills/:
git clone https://github.com/TheRealAlexV/ai-llm-security-testing.git /tmp/ai-llm-security-testing
# copy the .claude skill folder into Claude Code's global skills dir
cp -r /tmp/ai-llm-security-testing/.claude/skills/ai-llm-security-testing \
~/.claude/skills/ai-llm-security-testing
# or symlink it
ln -s "$(pwd)/ai-llm-security-testing/.claude/skills/ai-llm-security-testing" \
~/.claude/skills/ai-llm-security-testingThe Claude Code skill is self-contained (SKILL.md + a bundled references/
copy), so the .claude/ folder can be copied anywhere without the rest of the
repo.
Invoke the skill with phrases like:
- "security-test this LLM app"
- "red-team the RAG pipeline"
- "pentest the AI agent / MCP tool / multi-agent system"
- "threat-model our AI coding assistant"
- "harden our agent's defenses"
- "map these findings to OWASP LLM / Agentic Top 10"
The skill guides the agent through: grounding in the applicable frameworks → mapping the attack surface → seeding canaries → running the applicable techniques signal-first (with negative controls) → stating the control verdict → writing severity-rated, standards-mapped findings.
Three anchors drive every test:
- Trust boundary — keep user and retrieved content as data, never instructions.
- Control point — enforce permissions outside the model, in deterministic code.
- Evidence — record impact, reproduction, and remediation for every finding.
Controls are labelled deterministic (guarantees) or probabilistic (likelihood reduction) — filters never count as a fix.