A production-grade Rust framework and reference template for building autonomous AI agent loops that are safe to run unattended in mission-critical environments.
It implements a battle-tested closed-loop architecture:
TriggerFilterfast-paths routine events to avoid unnecessary LLM calls (saving 80–90% of token costs).RateLimiterpaces calls against your API budget with native async backpressure.SafetyFirewallclamps or rejects out-of-bounds proposals using hard mathematical/numeric invariants, completely independent of model confidence.PolicyGateseparates routine execution from high-risk decisions, routing escalations to human sign-off ("Boss's Desk").AuditSinkwrites structured, queryable logs for every single action and routing decision.
Zero-Infrastructure Local Testing: Ships with an offline
MockLlmReasonerandInMemoryAuditSinkso the entire pipeline compiles, runs, and tests offline in milliseconds without API keys, cloud infrastructure, or external dependencies.
Deploying autonomous AI agents into production operations (SRE, procurement, FinOps, infrastructure management) typically leads to one of two dead-ends:
- The "Unbounded Risk" Trap: Agents are given tool access (APIs, shells, database mutations) with only prompt-level instructions to "be careful". When an LLM inevitably hallucinates, experiences prompt drift, or encounters malicious injection, it can trigger catastrophic actions—such as dropping production clusters, over-ordering millions of dollars in inventory, or spamming third-party APIs.
- The "Perpetual Bottleneck" Trap: Because teams fear autonomous hallucinations, every single operational decision requires human sign-off. This completely negates the speed and scalability benefits of autonomous agents.
| Problem Area | Without This Template (Bare LLM Agent Loop) | With This Template (closed-loop-autonomous-ops) |
|---|---|---|
| Safety & Hallucinations | Dependent on prompt compliance and model self-confidence. Out-of-bounds actions slip through. | SafetyFirewall enforces rigid quantitative invariants after the model. Wild proposals (e.g. 50x limits) are rejected outright; slight excesses are clamped. |
| LLM Operating Costs | Every heartbeat, alert, and routine check invokes an LLM API call. Costs balloon exponentially. | TriggerFilter evaluates event state locally via fast rules. 80–90% of healthy events fast-path immediately with $0 token cost. |
| API Rate Limits & 429s | Ad-hoc retries and backoffs cause thundering herds, queue overflow, and crashed workers during traffic spikes. | RateLimiter with async backpressure (.acquire()) naturally throttles queue consumers (Kafka/RabbitMQ) when budget is low. |
| Human-in-the-Loop | Binary: either 100% manual or 100% autonomous. No granular policy differentiation. | PolicyGate automatically auto-approves safe, low-cost actions while seamlessly escalating expensive or clamped actions to human review. |
| Compliance & Auditing | Decisions are scattered across unstructured application logs, making post-incident forensics nearly impossible. | Structured AuditRecord captures timestamp, entity ID, proposed reasoning, firewall verdicts, and final policy decisions in queryable JSON. |
| Testing & CI/CD | Testing requires live API keys, mock HTTP servers, and incurs real token costs with non-deterministic test runs. | 100% offline, deterministic testing using built-in MockLlmReasoner and HallucinatingReasoner. |
This template is designed for engineering teams building autonomous systems where agents have write access to real-world state:
- Autonomous DevOps & SRE Teams:
- Auto-remediating alerts (restarting crash-looping services, clearing caches, re-routing traffic) while mathematically ensuring the agent can never delete a primary database or scale a cluster beyond budget caps.
- AI Supply Chain, Inventory & Procurement:
- Agents that continuously monitor SKU levels and place reorders autonomously for routine deficits, but automatically require a manager's sign-off for expensive bulk orders.
- FinOps & Cloud Cost Optimization:
- Unattended agents that decommission idle infrastructure, downscale unused test environments, and purchase compute reservations within strict daily dollar budgets.
- Autonomous IT & Security Operations (SecOps):
- Incident response agents that isolate compromised hosts or rotate API keys within deterministic bounds, escalating privileged access modifications to human operators.
- Enterprise Platform & Agent Developers:
- Any engineering team that needs to ship an agentic loop into production that will satisfy strict enterprise security, risk management, and compliance reviews.
Building a closed-loop autonomous system from scratch usually takes several months of architectural trial and error. This template dramatically accelerates your time-to-production:
- Skips 3–6 Months of Core Systems Engineering:
- You don't have to design and debug token-bucket rate limiters, backpressure queues, dual-phase validation gates (firewall vs. policy), or audit logging infrastructure. The architectural blueprint is already built, tested, and validated.
- Zero-API-Key Local Development & Rapid CI/CD:
- Start coding your domain logic immediately. The included
MockLlmReasonersimulates realistic model responses offline. Your entire test suite runs in under 100ms on local machines and CI without internet access, third-party credentials, or API bills.
- Start coding your domain logic immediately. The included
- Pluggable, Decoupled Trait Boundaries:
- Clean abstractions allow you to swap components in minutes:
- Swap
MockLlmReasonerfor OpenAI, Anthropic, Groq, or local Ollama with ~15 lines of code. - Swap
InMemoryAuditSinkfor PostgreSQL, ClickHouse, or DynamoDB without changing a single line of orchestrator code.
- Swap
- Clean abstractions allow you to swap components in minutes:
- Instant Security & Compliance Approvals:
- Getting autonomous agents approved by InfoSec and Risk teams is notoriously difficult. By presenting a deterministic post-model firewall and an un-bypassable policy gate with full audit trails, you can clear governance reviews in days rather than months.
- Turnkey Async Backpressure for Streaming Architectures:
- Avoid concurrency bugs and 429 cascade failures. The built-in rate limiter plugs directly into Tokio async loops, enabling queue consumers (Kafka, RabbitMQ, SQS) to naturally pause and resume based on token availability.
- Robust Type-Safe Rust Foundation:
- Leverage Rust’s type system and compiler guarantees to eliminate null pointer bugs, race conditions, and unhandled edge cases across your agent pipeline.
Incoming Event / State
│
▼
┌──────────────────┐
│ Trigger Filter │ ──FastPathAck──▶ Done (No LLM call, $0 cost)
└──────────────────┘
│ EscalateToAi
▼
┌──────────────────┐
│ Rate Limiter │ ◀── Async .acquire() backpressures the caller/queue
└──────────────────┘
│
▼
┌──────────────────┐
│ LLM Bridge │ ◀── LlmReasoner trait (Mock, OpenAI, Anthropic, etc.)
└──────────────────┘
│ ProposedAction
▼
┌──────────────────┐
│ Safety Firewall │ ◀── Hard physical/numeric invariant checks (Clamps or Rejects)
└──────────────────┘
│ FirewallVerdict (Approved | ClampedPendingReview | Rejected)
▼
┌──────────────────┐
│ Policy Gate │ ◀── Business risk checks (AutoApproved | RequiresHumanApproval | Blocked)
└──────────────────┘
│
▼
┌──────────────────┐
│ Governance/Audit │ ◀── AuditSink (Persists record with full rationale)
└──────────────────┘
│
▼
Action Execution / Notification
Integrating this template into your service requires just a few straightforward steps:
Represent your operational entity and its current health metrics as a standard Rust struct:
#[derive(Debug, Clone)]
pub struct ServiceHealthState {
pub service_name: String,
pub error_rate: f64, // e.g. percentage 0.0 - 100.0
pub p99_latency_ms: f64,
pub current_replicas: u32,
pub max_allowed_replicas: u32,
}Define rules that identify when an event actually requires AI reasoning. If no rules match, the event fast-paths with zero latency and zero LLM cost:
use closed_loop_autonomous_ops::trigger_filter::TriggerFilter;
let trigger = TriggerFilter::new()
.with_rule("high_error_rate", |s: &ServiceHealthState| s.error_rate > 5.0)
.with_rule("latency_spike", |s: &ServiceHealthState| s.p99_latency_ms > 1000.0);Configure your quantitative boundaries. The firewall enforces physical safety, while the policy gate sets your auto-execution budget:
use closed_loop_autonomous_ops::safety_firewall::FirewallLimits;
use closed_loop_autonomous_ops::OrchestratorConfig;
let config = OrchestratorConfig {
firewall_limits: FirewallLimits {
max_quantity: 20.0, // e.g. max replicas an agent can add
max_cost: 500.0, // absolute max spend permitted
human_review_cost_threshold: 100.0, // above this, flag for human confirmation
},
auto_approve_cost_ceiling: 50.0, // auto-execute if below this dollar threshold
tokens_per_minute: 60_000.0, // LLM API rate limit (TPM)
burst_capacity: 10_000.0, // allowable burst tokens
tokens_per_call: 500.0, // estimated tokens consumed per call
};Use the built-in MockLlmReasoner during development/testing, or implement LlmReasoner for your chosen model provider (OpenAI, Claude, Groq, Ollama):
use closed_loop_autonomous_ops::llm_bridge::{
LlmReasoner, ReasoningRequest, ProposedAction, LlmError
};
use async_trait::async_trait;
pub struct ProductionLlmReasoner {
pub api_key: String,
pub client: reqwest::Client,
}
#[async_trait]
impl LlmReasoner for ProductionLlmReasoner {
async fn reason(&self, request: ReasoningRequest) -> Result<ProposedAction, LlmError> {
// Send prompt and context JSON to your LLM API...
// Parse the structured output into a ProposedAction:
Ok(ProposedAction {
action_type: "scale_service".to_string(),
quantity: 3.0,
estimated_cost: 45.0,
reasoning: "Error rate is 8.2%; adding 3 replicas to relieve load.".to_string(),
})
}
}Use InMemoryAuditSink for local development or write a database-backed sink (PostgreSQL, ClickHouse) for production:
use closed_loop_autonomous_ops::governance::{AuditSink, AuditRecord};
use async_trait::async_trait;
pub struct PostgresAuditSink {
pub pool: sqlx::PgPool,
}
#[async_trait]
impl AuditSink for PostgresAuditSink {
async fn record(&self, record: AuditRecord) {
let decision_json = serde_json::to_value(&record.decision).unwrap();
sqlx::query!(
"INSERT INTO agent_audit_logs (id, timestamp, entity_id, decision) VALUES ($1, $2, $3, $4)",
record.id,
record.timestamp,
record.entity_id,
decision_json
)
.execute(&self.pool)
.await
.expect("Failed to write audit log");
}
async fn recent(&self, _limit: usize) -> Vec<AuditRecord> {
// Query recent records from DB...
vec![]
}
}Instantiate ClosedLoopOrchestrator and wrap your message stream (Kafka, RabbitMQ, SQS, or webhook handler):
use closed_loop_autonomous_ops::{ClosedLoopOrchestrator, EventOutcome};
use closed_loop_autonomous_ops::policy_gate::PolicyDecision;
use serde_json::json;
use std::sync::Arc;
let orchestrator = ClosedLoopOrchestrator::new(
config,
trigger,
Arc::new(ProductionLlmReasoner { api_key: "sk-...".into(), client: reqwest::Client::new() }),
Arc::new(PostgresAuditSink { pool: db_pool }),
);
// In your event consumer loop:
let outcome = orchestrator
.handle_event(&state.service_name, &state, |s| {
ReasoningRequest {
context: json!({
"error_rate": s.error_rate,
"latency_p99": s.p99_latency_ms,
"current_replicas": s.current_replicas,
}),
prompt: format!("Recommend scaling adjustment for {}", s.service_name),
}
})
.await;
// Step 7: Act on the Outcome
match outcome {
EventOutcome::FastPathAck => {
// System healthy: acknowledge event, no action needed ($0 cost)
}
EventOutcome::Decided(PolicyDecision::AutoApproved(action)) => {
// Safe and within budget: dispatch action to infrastructure API
execute_remediation(&action).await;
}
EventOutcome::Decided(PolicyDecision::RequiresHumanApproval { action, reason }) => {
// High risk or clamped: send ticket/alert to on-call dashboard or Slack
notify_human_approver(&action, &reason).await;
}
EventOutcome::Decided(PolicyDecision::Blocked { reason }) => {
// Dangerous or hallucinated proposal: alert engineering of firewall trip
log_security_alert(&reason).await;
}
EventOutcome::ReasoningFailed(err) => {
// LLM provider down or parse error: fallback to static heuristic
fallback_static_remediation(&err).await;
}
}Run the included end-to-end autonomous replenishment simulation:
cargo run --example autonomous_replenishmentWhat this simulation demonstrates:
- 3 healthy SKUs fast-path instantly without calling the LLM ($0 cost, 0ms latency).
- 1 low-stock SKU is analyzed by the reasoner and auto-approved by the firewall and policy gate.
- 1 low-stock but expensive SKU is analyzed and safely routed to
RequiresHumanApproval. - An immutable audit trail is logged and printed at the end. See examples/autonomous_replenishment.rs.
# Run all unit tests offline (17 tests)
cargo test
# Run the safety verification test (proves hallucinated 10,000,000-unit order gets blocked)
cargo test --test firewall_safety_testSee tests/firewall_safety_test.rs.
- Tune
auto_approve_cost_ceiling: Keep this well below what a single bad week can absorb, independent of the firewall'smax_cost. - Validate Firewall Limits: Define absolute upper bounds (
max_quantity,max_cost) based on physical or business constraints. Note that proposals exceeding 50x limits are rejected outright as hallucinations rather than clamped. - Token Usage Tracking: Update
tokens_per_calldynamically from response metadata returned by provider APIs for ultra-precise TPM budgeting. - Persistent Audit Sink: Replace
InMemoryAuditSinkwith PostgreSQL, ClickHouse, or Snowflake before enabling unattended auto-execution. - Dead-Letter Queue for Blocked Events: Route
PolicyDecision::Blockedoccurrences to your security/observability platform (Datadog, Sentry, PagerDuty) to catch model drift or prompt injection attempts early.
MIT