Skip to content

About

A framework for safe, autonomous AI agent loops. A routing gate triggers LLM calls only when needed, a token-bucket limiter controls costs, a safety firewall validates actions, a policy gate routes sensitive decisions to humans, and every decision is logged for auditability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

closed-loop-autonomous-ops

A production-grade Rust framework and reference template for building autonomous AI agent loops that are safe to run unattended in mission-critical environments.

It implements a battle-tested closed-loop architecture:

  1. TriggerFilter fast-paths routine events to avoid unnecessary LLM calls (saving 80–90% of token costs).
  2. RateLimiter paces calls against your API budget with native async backpressure.
  3. SafetyFirewall clamps or rejects out-of-bounds proposals using hard mathematical/numeric invariants, completely independent of model confidence.
  4. PolicyGate separates routine execution from high-risk decisions, routing escalations to human sign-off ("Boss's Desk").
  5. AuditSink writes structured, queryable logs for every single action and routing decision.

Zero-Infrastructure Local Testing: Ships with an offline MockLlmReasoner and InMemoryAuditSink so the entire pipeline compiles, runs, and tests offline in milliseconds without API keys, cloud infrastructure, or external dependencies.


What Problem Does This Template Solve?

Deploying autonomous AI agents into production operations (SRE, procurement, FinOps, infrastructure management) typically leads to one of two dead-ends:

  1. The "Unbounded Risk" Trap: Agents are given tool access (APIs, shells, database mutations) with only prompt-level instructions to "be careful". When an LLM inevitably hallucinates, experiences prompt drift, or encounters malicious injection, it can trigger catastrophic actions—such as dropping production clusters, over-ordering millions of dollars in inventory, or spamming third-party APIs.
  2. The "Perpetual Bottleneck" Trap: Because teams fear autonomous hallucinations, every single operational decision requires human sign-off. This completely negates the speed and scalability benefits of autonomous agents.

What Happens Without This Template vs. With This Template

Problem Area Without This Template (Bare LLM Agent Loop) With This Template (closed-loop-autonomous-ops)
Safety & Hallucinations Dependent on prompt compliance and model self-confidence. Out-of-bounds actions slip through. SafetyFirewall enforces rigid quantitative invariants after the model. Wild proposals (e.g. 50x limits) are rejected outright; slight excesses are clamped.
LLM Operating Costs Every heartbeat, alert, and routine check invokes an LLM API call. Costs balloon exponentially. TriggerFilter evaluates event state locally via fast rules. 80–90% of healthy events fast-path immediately with $0 token cost.
API Rate Limits & 429s Ad-hoc retries and backoffs cause thundering herds, queue overflow, and crashed workers during traffic spikes. RateLimiter with async backpressure (.acquire()) naturally throttles queue consumers (Kafka/RabbitMQ) when budget is low.
Human-in-the-Loop Binary: either 100% manual or 100% autonomous. No granular policy differentiation. PolicyGate automatically auto-approves safe, low-cost actions while seamlessly escalating expensive or clamped actions to human review.
Compliance & Auditing Decisions are scattered across unstructured application logs, making post-incident forensics nearly impossible. Structured AuditRecord captures timestamp, entity ID, proposed reasoning, firewall verdicts, and final policy decisions in queryable JSON.
Testing & CI/CD Testing requires live API keys, mock HTTP servers, and incurs real token costs with non-deterministic test runs. 100% offline, deterministic testing using built-in MockLlmReasoner and HallucinatingReasoner.

Who Should Use This?

This template is designed for engineering teams building autonomous systems where agents have write access to real-world state:

  • Autonomous DevOps & SRE Teams:
    • Auto-remediating alerts (restarting crash-looping services, clearing caches, re-routing traffic) while mathematically ensuring the agent can never delete a primary database or scale a cluster beyond budget caps.
  • AI Supply Chain, Inventory & Procurement:
    • Agents that continuously monitor SKU levels and place reorders autonomously for routine deficits, but automatically require a manager's sign-off for expensive bulk orders.
  • FinOps & Cloud Cost Optimization:
    • Unattended agents that decommission idle infrastructure, downscale unused test environments, and purchase compute reservations within strict daily dollar budgets.
  • Autonomous IT & Security Operations (SecOps):
    • Incident response agents that isolate compromised hosts or rotate API keys within deterministic bounds, escalating privileged access modifications to human operators.
  • Enterprise Platform & Agent Developers:
    • Any engineering team that needs to ship an agentic loop into production that will satisfy strict enterprise security, risk management, and compliance reviews.

How This Template Paces Up Your Development (In Detail)

Building a closed-loop autonomous system from scratch usually takes several months of architectural trial and error. This template dramatically accelerates your time-to-production:

  1. Skips 3–6 Months of Core Systems Engineering:
    • You don't have to design and debug token-bucket rate limiters, backpressure queues, dual-phase validation gates (firewall vs. policy), or audit logging infrastructure. The architectural blueprint is already built, tested, and validated.
  2. Zero-API-Key Local Development & Rapid CI/CD:
    • Start coding your domain logic immediately. The included MockLlmReasoner simulates realistic model responses offline. Your entire test suite runs in under 100ms on local machines and CI without internet access, third-party credentials, or API bills.
  3. Pluggable, Decoupled Trait Boundaries:
    • Clean abstractions allow you to swap components in minutes:
      • Swap MockLlmReasoner for OpenAI, Anthropic, Groq, or local Ollama with ~15 lines of code.
      • Swap InMemoryAuditSink for PostgreSQL, ClickHouse, or DynamoDB without changing a single line of orchestrator code.
  4. Instant Security & Compliance Approvals:
    • Getting autonomous agents approved by InfoSec and Risk teams is notoriously difficult. By presenting a deterministic post-model firewall and an un-bypassable policy gate with full audit trails, you can clear governance reviews in days rather than months.
  5. Turnkey Async Backpressure for Streaming Architectures:
    • Avoid concurrency bugs and 429 cascade failures. The built-in rate limiter plugs directly into Tokio async loops, enabling queue consumers (Kafka, RabbitMQ, SQS) to naturally pause and resume based on token availability.
  6. Robust Type-Safe Rust Foundation:
    • Leverage Rust’s type system and compiler guarantees to eliminate null pointer bugs, race conditions, and unhandled edge cases across your agent pipeline.

Pipeline Architecture

Incoming Event / State
         │
         ▼
┌──────────────────┐
│  Trigger Filter  │ ──FastPathAck──▶ Done (No LLM call, $0 cost)
└──────────────────┘
         │ EscalateToAi
         ▼
┌──────────────────┐
│   Rate Limiter   │ ◀── Async .acquire() backpressures the caller/queue
└──────────────────┘
         │
         ▼
┌──────────────────┐
│    LLM Bridge    │ ◀── LlmReasoner trait (Mock, OpenAI, Anthropic, etc.)
└──────────────────┘
         │ ProposedAction
         ▼
┌──────────────────┐
│ Safety Firewall  │ ◀── Hard physical/numeric invariant checks (Clamps or Rejects)
└──────────────────┘
         │ FirewallVerdict (Approved | ClampedPendingReview | Rejected)
         ▼
┌──────────────────┐
│   Policy Gate    │ ◀── Business risk checks (AutoApproved | RequiresHumanApproval | Blocked)
└──────────────────┘
         │
         ▼
┌──────────────────┐
│ Governance/Audit │ ◀── AuditSink (Persists record with full rationale)
└──────────────────┘
         │
         ▼
Action Execution / Notification

How to Use This Template (Step-by-Step Integration Guide)

Integrating this template into your service requires just a few straightforward steps:

Step 1: Define Your Domain State

Represent your operational entity and its current health metrics as a standard Rust struct:

#[derive(Debug, Clone)]
pub struct ServiceHealthState {
    pub service_name: String,
    pub error_rate: f64,        // e.g. percentage 0.0 - 100.0
    pub p99_latency_ms: f64,
    pub current_replicas: u32,
    pub max_allowed_replicas: u32,
}

Step 2: Configure the Trigger Filter (Fast-Path Rules)

Define rules that identify when an event actually requires AI reasoning. If no rules match, the event fast-paths with zero latency and zero LLM cost:

use closed_loop_autonomous_ops::trigger_filter::TriggerFilter;

let trigger = TriggerFilter::new()
    .with_rule("high_error_rate", |s: &ServiceHealthState| s.error_rate > 5.0)
    .with_rule("latency_spike", |s: &ServiceHealthState| s.p99_latency_ms > 1000.0);

Step 3: Set Hard Firewall Limits & Policy Ceilings

Configure your quantitative boundaries. The firewall enforces physical safety, while the policy gate sets your auto-execution budget:

use closed_loop_autonomous_ops::safety_firewall::FirewallLimits;
use closed_loop_autonomous_ops::OrchestratorConfig;

let config = OrchestratorConfig {
    firewall_limits: FirewallLimits {
        max_quantity: 20.0,                   // e.g. max replicas an agent can add
        max_cost: 500.0,                      // absolute max spend permitted
        human_review_cost_threshold: 100.0,   // above this, flag for human confirmation
    },
    auto_approve_cost_ceiling: 50.0,          // auto-execute if below this dollar threshold
    tokens_per_minute: 60_000.0,              // LLM API rate limit (TPM)
    burst_capacity: 10_000.0,                 // allowable burst tokens
    tokens_per_call: 500.0,                   // estimated tokens consumed per call
};

Step 4: Wire Your LLM Reasoner

Use the built-in MockLlmReasoner during development/testing, or implement LlmReasoner for your chosen model provider (OpenAI, Claude, Groq, Ollama):

use closed_loop_autonomous_ops::llm_bridge::{
    LlmReasoner, ReasoningRequest, ProposedAction, LlmError
};
use async_trait::async_trait;

pub struct ProductionLlmReasoner {
    pub api_key: String,
    pub client: reqwest::Client,
}

#[async_trait]
impl LlmReasoner for ProductionLlmReasoner {
    async fn reason(&self, request: ReasoningRequest) -> Result<ProposedAction, LlmError> {
        // Send prompt and context JSON to your LLM API...
        // Parse the structured output into a ProposedAction:
        Ok(ProposedAction {
            action_type: "scale_service".to_string(),
            quantity: 3.0,
            estimated_cost: 45.0,
            reasoning: "Error rate is 8.2%; adding 3 replicas to relieve load.".to_string(),
        })
    }
}

Step 5: Wire an Audit Sink

Use InMemoryAuditSink for local development or write a database-backed sink (PostgreSQL, ClickHouse) for production:

use closed_loop_autonomous_ops::governance::{AuditSink, AuditRecord};
use async_trait::async_trait;

pub struct PostgresAuditSink {
    pub pool: sqlx::PgPool,
}

#[async_trait]
impl AuditSink for PostgresAuditSink {
    async fn record(&self, record: AuditRecord) {
        let decision_json = serde_json::to_value(&record.decision).unwrap();
        sqlx::query!(
            "INSERT INTO agent_audit_logs (id, timestamp, entity_id, decision) VALUES ($1, $2, $3, $4)",
            record.id,
            record.timestamp,
            record.entity_id,
            decision_json
        )
        .execute(&self.pool)
        .await
        .expect("Failed to write audit log");
    }

    async fn recent(&self, _limit: usize) -> Vec<AuditRecord> {
        // Query recent records from DB...
        vec![]
    }
}

Step 6: Initialize the Orchestrator & Process Events

Instantiate ClosedLoopOrchestrator and wrap your message stream (Kafka, RabbitMQ, SQS, or webhook handler):

use closed_loop_autonomous_ops::{ClosedLoopOrchestrator, EventOutcome};
use closed_loop_autonomous_ops::policy_gate::PolicyDecision;
use serde_json::json;
use std::sync::Arc;

let orchestrator = ClosedLoopOrchestrator::new(
    config,
    trigger,
    Arc::new(ProductionLlmReasoner { api_key: "sk-...".into(), client: reqwest::Client::new() }),
    Arc::new(PostgresAuditSink { pool: db_pool }),
);

// In your event consumer loop:
let outcome = orchestrator
    .handle_event(&state.service_name, &state, |s| {
        ReasoningRequest {
            context: json!({
                "error_rate": s.error_rate,
                "latency_p99": s.p99_latency_ms,
                "current_replicas": s.current_replicas,
            }),
            prompt: format!("Recommend scaling adjustment for {}", s.service_name),
        }
    })
    .await;

// Step 7: Act on the Outcome
match outcome {
    EventOutcome::FastPathAck => {
        // System healthy: acknowledge event, no action needed ($0 cost)
    }
    EventOutcome::Decided(PolicyDecision::AutoApproved(action)) => {
        // Safe and within budget: dispatch action to infrastructure API
        execute_remediation(&action).await;
    }
    EventOutcome::Decided(PolicyDecision::RequiresHumanApproval { action, reason }) => {
        // High risk or clamped: send ticket/alert to on-call dashboard or Slack
        notify_human_approver(&action, &reason).await;
    }
    EventOutcome::Decided(PolicyDecision::Blocked { reason }) => {
        // Dangerous or hallucinated proposal: alert engineering of firewall trip
        log_security_alert(&reason).await;
    }
    EventOutcome::ReasoningFailed(err) => {
        // LLM provider down or parse error: fallback to static heuristic
        fallback_static_remediation(&err).await;
    }
}

Quickstart

Run the included end-to-end autonomous replenishment simulation:

cargo run --example autonomous_replenishment

What this simulation demonstrates:

  • 3 healthy SKUs fast-path instantly without calling the LLM ($0 cost, 0ms latency).
  • 1 low-stock SKU is analyzed by the reasoner and auto-approved by the firewall and policy gate.
  • 1 low-stock but expensive SKU is analyzed and safely routed to RequiresHumanApproval.
  • An immutable audit trail is logged and printed at the end. See examples/autonomous_replenishment.rs.

Running Offline Tests

# Run all unit tests offline (17 tests)
cargo test

# Run the safety verification test (proves hallucinated 10,000,000-unit order gets blocked)
cargo test --test firewall_safety_test

See tests/firewall_safety_test.rs.


Production Deployment Checklist

  • Tune auto_approve_cost_ceiling: Keep this well below what a single bad week can absorb, independent of the firewall's max_cost.
  • Validate Firewall Limits: Define absolute upper bounds (max_quantity, max_cost) based on physical or business constraints. Note that proposals exceeding 50x limits are rejected outright as hallucinations rather than clamped.
  • Token Usage Tracking: Update tokens_per_call dynamically from response metadata returned by provider APIs for ultra-precise TPM budgeting.
  • Persistent Audit Sink: Replace InMemoryAuditSink with PostgreSQL, ClickHouse, or Snowflake before enabling unattended auto-execution.
  • Dead-Letter Queue for Blocked Events: Route PolicyDecision::Blocked occurrences to your security/observability platform (Datadog, Sentry, PagerDuty) to catch model drift or prompt injection attempts early.

License

MIT

About

A framework for safe, autonomous AI agent loops. A routing gate triggers LLM calls only when needed, a token-bucket limiter controls costs, a safety firewall validates actions, a policy gate routes sensitive decisions to humans, and every decision is logged for auditability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages