Skip to content

security: content safety and prompt shields fail-open is a bypass risk #5

Description

Summary

Both Content Safety and Prompt Shields checks in inference-router/src/safety.rs are configured as fail-open — if the Azure AI Content Safety API is unreachable or returns an error, the request proceeds without safety checks.

Current Behavior (safety.rs)

// Content Safety check — fail-open
match check_content_safety(&text, &endpoint, &token).await {
    Ok(result) => { /* enforce if severity >= threshold */ }
    Err(e) => {
        warn!("Content Safety check failed, allowing request: {}", e);
        // REQUEST PROCEEDS WITHOUT SAFETY CHECK
    }
}

Same pattern for Prompt Shields.

Risk

An attacker who can cause the Content Safety endpoint to be unreachable (DNS manipulation, network partition, DoS on the endpoint) can bypass all safety checks. In a sandbox environment where the agent has shell access, this is a significant risk.

Proposal

  1. Add a configurable fail_mode per safety check: fail-open (current default) or fail-closed
  2. For production deployments, recommend fail-closed with a circuit breaker (after N failures, block all requests until the safety endpoint recovers)
  3. Add a governance policy option: content_safety_fail_mode: closed in ClawSandbox CRD

AGT's agent-sre package has circuit breaker patterns that could be applied here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions