Problem
Pattern detection against injected instructions can never be complete — #16354 is the local proof (a zero-width character defeats the match). The layer that has to hold is the one that decides what the model is allowed to do once untrusted text is in its context. Traced end to end, no provenance signal reaches any approval decision.
The approval decisions and their actual inputs
| Gate |
Signature |
Inputs |
| Chat tool calls |
chat_workflow/graph.py:494 _tool_call_needs_approval(tool_call, ctx) |
tool name + ctx.requires_approval_before |
| Command risk |
services/command_approval_manager.py:247 needs_approval(agent_role, command_risk, permissions) |
role + risk of the command string |
| Shell execution |
security_layer.py:583 _should_force_approval(command, user_role) |
command string + role |
None takes provenance. Risk is derived from the text of the action, never from where the instruction that produced it came from.
Two specific consequences
1. The chat tool gate defaults open.
declared = getattr(ctx, "requires_approval_before", None) or []
if not declared:
return False
A session with no declared approval categories — an ordinary chat, since declarations come from an LLC work item — approves every planned tool call unconditionally. The risk-based gates above still cover shell command execution, but every other tool (file writes, HTTP, KB writes, deploys) passes on the declaration check alone.
2. The taint bit already exists and is discarded.
ctx carries used_knowledge (chat_workflow/graph.py:103, set at llm_handler.py:727) — the gate receives the very object that records that retrieved content entered the prompt, and never reads it.
ContentFirewall computes FirewallAction.QUARANTINE, risk and detected_patterns (security/content_firewall.py:94-121), but every caller reads only .blocked and .content (advanced_rag_optimizer.py:1055, web_fetch/extractors.py:192, chat_workflow/tool_handler.py:2979, tools/parallel/executor.py:479, skills/sync/mcp_client.py:160,221). A QUARANTINE verdict appends a warning string to the prompt and changes nothing downstream.
So the system detects "this content was suspicious", writes it into a verdict, and then throws that verdict away before the decision that matters.
Why it compounds with agent roles
AUTOMATION_AGENT and SYSTEM_AGENT carry auto_approve_moderate=True, allow_high=True; ADMIN_AGENT adds allow_dangerous=True (command_approval_manager.py:134-150). These roles retrieve KB facts like any other. An injected instruction that yields a MODERATE or HIGH command under those roles executes with no human in the loop.
Acceptance criteria
Related
Completes the chain with #16770 (ingestion sanitizing bypassed) and #16771 (retrieval not firewalled). #16354 is why the detector cannot be the last line. #13250 tracks the two-gate split this issue also touches.
Problem
Pattern detection against injected instructions can never be complete — #16354 is the local proof (a zero-width character defeats the match). The layer that has to hold is the one that decides what the model is allowed to do once untrusted text is in its context. Traced end to end, no provenance signal reaches any approval decision.
The approval decisions and their actual inputs
chat_workflow/graph.py:494_tool_call_needs_approval(tool_call, ctx)ctx.requires_approval_beforeservices/command_approval_manager.py:247needs_approval(agent_role, command_risk, permissions)security_layer.py:583_should_force_approval(command, user_role)None takes provenance. Risk is derived from the text of the action, never from where the instruction that produced it came from.
Two specific consequences
1. The chat tool gate defaults open.
A session with no declared approval categories — an ordinary chat, since declarations come from an LLC work item — approves every planned tool call unconditionally. The risk-based gates above still cover shell command execution, but every other tool (file writes, HTTP, KB writes, deploys) passes on the declaration check alone.
2. The taint bit already exists and is discarded.
ctxcarriesused_knowledge(chat_workflow/graph.py:103, set atllm_handler.py:727) — the gate receives the very object that records that retrieved content entered the prompt, and never reads it.ContentFirewallcomputesFirewallAction.QUARANTINE,riskanddetected_patterns(security/content_firewall.py:94-121), but every caller reads only.blockedand.content(advanced_rag_optimizer.py:1055,web_fetch/extractors.py:192,chat_workflow/tool_handler.py:2979,tools/parallel/executor.py:479,skills/sync/mcp_client.py:160,221). A QUARANTINE verdict appends a warning string to the prompt and changes nothing downstream.So the system detects "this content was suspicious", writes it into a verdict, and then throws that verdict away before the decision that matters.
Why it compounds with agent roles
AUTOMATION_AGENTandSYSTEM_AGENTcarryauto_approve_moderate=True, allow_high=True;ADMIN_AGENTaddsallow_dangerous=True(command_approval_manager.py:134-150). These roles retrieve KB facts like any other. An injected instruction that yields a MODERATE or HIGH command under those roles executes with no human in the loop.Acceptance criteria
_tool_call_needs_approvalconsults it: while tainted, tool calls above a configured risk floor require approval even when the run declared no categories.CommandApprovalManager.needs_approvalaccepts the marker and suppressesauto_approve_moderate/allow_highwhile tainted, for every agent role.Related
Completes the chain with #16770 (ingestion sanitizing bypassed) and #16771 (retrieval not firewalled). #16354 is why the detector cannot be the last line. #13250 tracks the two-gate split this issue also touches.