Skip to content

security(agents): no taint propagation — approval gates never learn that untrusted content entered the context #16776

Description

@mrveiss

Problem

Pattern detection against injected instructions can never be complete — #16354 is the local proof (a zero-width character defeats the match). The layer that has to hold is the one that decides what the model is allowed to do once untrusted text is in its context. Traced end to end, no provenance signal reaches any approval decision.

The approval decisions and their actual inputs

Gate Signature Inputs
Chat tool calls chat_workflow/graph.py:494 _tool_call_needs_approval(tool_call, ctx) tool name + ctx.requires_approval_before
Command risk services/command_approval_manager.py:247 needs_approval(agent_role, command_risk, permissions) role + risk of the command string
Shell execution security_layer.py:583 _should_force_approval(command, user_role) command string + role

None takes provenance. Risk is derived from the text of the action, never from where the instruction that produced it came from.

Two specific consequences

1. The chat tool gate defaults open.

declared = getattr(ctx, "requires_approval_before", None) or []
if not declared:
    return False

A session with no declared approval categories — an ordinary chat, since declarations come from an LLC work item — approves every planned tool call unconditionally. The risk-based gates above still cover shell command execution, but every other tool (file writes, HTTP, KB writes, deploys) passes on the declaration check alone.

2. The taint bit already exists and is discarded.

  • ctx carries used_knowledge (chat_workflow/graph.py:103, set at llm_handler.py:727) — the gate receives the very object that records that retrieved content entered the prompt, and never reads it.
  • ContentFirewall computes FirewallAction.QUARANTINE, risk and detected_patterns (security/content_firewall.py:94-121), but every caller reads only .blocked and .content (advanced_rag_optimizer.py:1055, web_fetch/extractors.py:192, chat_workflow/tool_handler.py:2979, tools/parallel/executor.py:479, skills/sync/mcp_client.py:160,221). A QUARANTINE verdict appends a warning string to the prompt and changes nothing downstream.

So the system detects "this content was suspicious", writes it into a verdict, and then throws that verdict away before the decision that matters.

Why it compounds with agent roles

AUTOMATION_AGENT and SYSTEM_AGENT carry auto_approve_moderate=True, allow_high=True; ADMIN_AGENT adds allow_dangerous=True (command_approval_manager.py:134-150). These roles retrieve KB facts like any other. An injected instruction that yields a MODERATE or HIGH command under those roles executes with no human in the loop.

Acceptance criteria

  • A session/run carries a taint marker set when content from an untrusted source (RAG, web, MCP, tool output, file) enters the prompt — sourced from the firewall verdict, not re-derived.
  • The marker records the highest risk seen and the source, not just a boolean.
  • _tool_call_needs_approval consults it: while tainted, tool calls above a configured risk floor require approval even when the run declared no categories.
  • CommandApprovalManager.needs_approval accepts the marker and suppresses auto_approve_moderate / allow_high while tainted, for every agent role.
  • A QUARANTINE verdict sets the marker — today it is inert.
  • Tests: a chat turn whose KB context carries an injection payload cannot auto-approve a tool call that the same turn would auto-approve with clean context; and the same for an automation-role command at MODERATE risk.
  • The marker is visible in the trajectory/audit record, so a reviewer can see that a decision was taken under tainted context.

Related

Completes the chain with #16770 (ingestion sanitizing bypassed) and #16771 (retrieval not firewalled). #16354 is why the detector cannot be the last line. #13250 tracks the two-gate split this issue also touches.

Activity

  1. mrveiss commented on Sep 16, 2026

    @mrveiss
    OwnerAuthor

    Implementation approach, from autobot-ai-21's analysis while building #16791 — recorded here so it survives the session.

    Do not attach the taint signal to RAGMetrics. It is advanced_search()'s second return value and looks like the natural carrier, but services/knowledge/service.py:240 (_search_filter_and_format) does results, _ = await self.rag_service.advanced_search(...) and throws it away. Carrying it to run level would mean threading a parameter through a function that deliberately discards it.

    Attach to citations instead. They already reach run level unconditionally: chat_workflow/llm_handler.py:727 sets session.metadata["last_citations"] = citations on every RAG turn, in the same place used_knowledge and query_intent are set — which is precisely where the approval gate should be reading. #16791 already adds firewall_safe_content to each citation dict in format_citations() and at the doc-search citation site in conversation_aware_retrieve(); firewall_action and firewall_risk are one attribute access away on the same SearchResult.

    The run-level marker is then an aggregation at that existing site: take the worst firewall_action present across the citations and write it beside used_knowledge. No signature changes to advanced_search() or get_optimized_context().

    One gap that must close with it: firewall_filter_doc_results() in rag_content_firewall.py attaches only firewall_safe_content on the dict path, not action or risk. A taint signal covering the fact path but not the doc path would let a tainted doc chunk arrive unmarked and show the gate a clean run — a half-guard that reads as coverage. Both functions in that module get the same treatment.

    Adds to the acceptance criteria above: the marker carries source and highest risk (not a boolean), and both retrieval sources feed it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions