Pre-action content firewall — wire prompt-injection defense over untrusted inputs
Part of umbrella #10542 · Theme A (collaboration core) · security — highest-risk gap
Problem
A PromptInjectionDetector exists but only guards context files. It is not wired into the agent's untrusted-input paths — MCP tool output, web-fetched content, RAG-retrieved documents, file reads, command stdout. An agent that executes commands in a Docker sandbox and ingests external content is exactly the threat model for indirect prompt injection (a malicious web page / repo / doc telling the agent to exfiltrate secrets or run destructive commands). Today there's no firewall on that boundary.
What exists today
security/prompt_injection_detector.py — PromptInjectionDetector / InjectionRisk (detects "ignore previous instructions" etc.) — scoped to context files (tests/test_injection_detection.py).
services/tool_output_filter.py — formatting cleaner (truncate/dedup/ansi), not a security boundary.
a2a/trust_score.py — trust scoring primitive.
web_fetch/extractors.py — where web content enters.
- Memory note: "prompt-injection in tool output — ignore + flag once" is currently a manual discipline, not enforced.
What to build
- A content firewall that runs
PromptInjectionDetector (+ data/instruction separation) over every untrusted input before it reaches the model: MCP tool results, web fetch, RAG documents, file reads, command stdout.
- Policy on detection: strip / quarantine / require human approval based on
InjectionRisk, surfaced in the trajectory and to the human.
- Mark provenance of untrusted spans so the model treats them as data, not instructions (delimiting + system reminder).
- Wire at the choke points (parallel executor, MCP client, RAG response builder, web_fetch) — one shared filter, not per-site copies.
Acceptance criteria
Entry points
autobot-backend/security/prompt_injection_detector.py · autobot-backend/tools/parallel/executor.py · autobot-backend/skills/sync/mcp_client.py · autobot-backend/knowledge/search_components/response_builder.py · autobot-backend/web_fetch/extractors.py · autobot-backend/agent_loop/approval_workflow.py.
Pre-action content firewall — wire prompt-injection defense over untrusted inputs
Part of umbrella #10542 · Theme A (collaboration core) · security — highest-risk gap
Problem
A
PromptInjectionDetectorexists but only guards context files. It is not wired into the agent's untrusted-input paths — MCP tool output, web-fetched content, RAG-retrieved documents, file reads, command stdout. An agent that executes commands in a Docker sandbox and ingests external content is exactly the threat model for indirect prompt injection (a malicious web page / repo / doc telling the agent to exfiltrate secrets or run destructive commands). Today there's no firewall on that boundary.What exists today
security/prompt_injection_detector.py—PromptInjectionDetector/InjectionRisk(detects "ignore previous instructions" etc.) — scoped to context files (tests/test_injection_detection.py).services/tool_output_filter.py— formatting cleaner (truncate/dedup/ansi), not a security boundary.a2a/trust_score.py— trust scoring primitive.web_fetch/extractors.py— where web content enters.What to build
PromptInjectionDetector(+ data/instruction separation) over every untrusted input before it reaches the model: MCP tool results, web fetch, RAG documents, file reads, command stdout.InjectionRisk, surfaced in the trajectory and to the human.Acceptance criteria
Entry points
autobot-backend/security/prompt_injection_detector.py·autobot-backend/tools/parallel/executor.py·autobot-backend/skills/sync/mcp_client.py·autobot-backend/knowledge/search_components/response_builder.py·autobot-backend/web_fetch/extractors.py·autobot-backend/agent_loop/approval_workflow.py.