Skip to content

Architecture & System Design

mrveiss edited this page Jul 22, 2026 · 2 revisions

Architecture & System Design

High-Level Overview

AutoBot is a microservices-based platform where a Vue.js frontend communicates with a FastAPI backend. The backend orchestrates AI inference, vector search, relational data, caching, fleet operations, and browser automation as separate concerns. Inference is routed through an LLM gateway so you can choose local models, a hosted provider, or a mix — your data stays on your infrastructure either way. A Service Lifecycle Manager (SLM) handles deploy, operate, and scale for the platform itself.

graph TB
    User["User (Browser)"]
    Frontend["Frontend (Vue.js)"]
    Backend["Backend (FastAPI)"]

    Redis["Redis (Cache/Queue)"]
    PostgreSQL["PostgreSQL (Data)"]
    ChromaDB["ChromaDB (Vectors)"]

    Gateway["LLM Gateway (routing/fallback)"]
    LocalLLM["Local Inference (Ollama)"]
    Provider["External Providers (claude_api, openrouter)"]

    SLM["Service Lifecycle Manager (deploy/operate/scale)"]

    Ansible["Ansible (Fleet Ops)"]
    Browser["Browser Automation (Chromium)"]

    User -->|HTTP/WebSocket| Frontend
    Frontend -->|REST API| Backend

    Backend -->|Read/Write| Redis
    Backend -->|Query| PostgreSQL
    Backend -->|Vector Search| ChromaDB
    Backend -->|Inference| Gateway
    Gateway -->|Local| LocalLLM
    Gateway -->|Optional provider| Provider

    Backend -->|Execute Playbooks| Ansible
    Backend -->|Control| Browser
    SLM -.->|Deploy / operate / scale| Backend
Loading

Components

Frontend (autobot-frontend)

  • Framework: Vue.js 3 with Composition API
  • Build: Vite
  • UI: Custom component library, xterm.js for terminal views
  • Served by: nginx (also handles TLS termination and reverse proxy)
  • Communicates with: Backend via REST and WebSocket

Backend (autobot-backend)

  • Framework: FastAPI (Python)
  • Async: Full async/await with asyncio
  • Auth: JWT tokens with multi-user accounts and role-based access control (RBAC) — the platform is multi-user; permissions are enforced per role on every route
  • Key modules:
    • api/ — REST route handlers
    • agents/ — AI agent loop and orchestration
    • agent_loop/ — Multi-turn conversation management
    • advanced_rag_optimizer.py — RAG retrieval and reranking

Database (autobot-database)

  • Engine: PostgreSQL
  • Stores: Users, sessions, fleet inventory, knowledge base metadata, workflow definitions, audit logs

Redis

  • Role: Task queue (Celery/background jobs) + caching layer
  • Used for: Chat session state, async job results, rate limiting

ChromaDB

  • Role: Vector database for RAG (Retrieval Augmented Generation)
  • Stores: Embeddings for knowledge base documents, code, runbooks
  • Used for: Semantic search across all indexed content

LLM Gateway

  • Role: Single entry point for all model inference; routes each request to the configured backend
  • Providers: Local inference (Ollama) and external providers (claude_api, openrouter) — you choose per deployment
  • Features: Provider routing and fallback, so a request can fail over to an alternate provider
  • Your data, your choice: run fully local, use a hosted provider, or mix — your data stays on your infrastructure and no provider is contacted unless you configure one

Local Inference (Ollama)

  • Role: Local AI inference backend behind the gateway
  • Runs: Open-source models (Llama, Mistral, etc.) entirely on your hardware
  • Use when: you want inference to stay on-box with no external calls

Service Lifecycle Manager (SLM)

  • Role: Deploys, operates, and scales the AutoBot platform itself
  • Handles: installation, upgrades/self-update, and multi-node scaling of the stack
  • Driven by: Ansible playbooks (e.g. the installer's deploy-slm-manager.yml)
  • Note: "SLM" always means Service Lifecycle Manager — it is not a language-model term

Ansible (autobot-infrastructure)

  • Role: Fleet operations engine
  • Executes: Playbooks for multi-server deployments, updates, configuration management
  • Driven by: Backend API commands triggered from chat or workflows

Browser Worker (autobot-browser-worker)

  • Role: Browser automation via Chromium with vision-in-the-loop
  • Used for: Automated browsing where the agent captures the page, reasons over what it sees, and drives the next action; also web scraping and screenshot-based troubleshooting

Data Flows

Chat Request Flow

User types message
    → Frontend (WebSocket)
    → Backend: agent_loop receives message
    → Backend: RAG query to ChromaDB (if knowledge base attached)
    → Backend: context assembled (history + retrieved docs)
    → LLM Gateway: inference request routed to the configured provider (local Ollama or an external provider)
    → Backend: streams response tokens back
    → Frontend: renders streaming response to user

Knowledge Base Indexing Flow

User uploads document
    → Backend: file received, stored in PostgreSQL (metadata)
    → Backend: document chunked and embedded
    → ChromaDB: embeddings stored
    → Available for RAG queries immediately

Fleet Operation Flow

User issues fleet command in chat
    → Backend: intent parsed by agent
    → Backend: Ansible playbook selected or generated
    → Ansible: executes against target nodes
    → Backend: captures stdout/stderr, updates job status
    → Frontend: shows real-time progress
    → PostgreSQL: audit log written

Tech Stack Summary

Layer Technology
Frontend Vue.js 3, Vite, nginx
Backend FastAPI, Python 3.11+, asyncio
Relational DB PostgreSQL 15+
Cache / Queue Redis 7+
Vector DB ChromaDB
AI Inference LLM gateway — local models via Ollama and/or external providers (claude_api, openrouter)
Fleet Ops Ansible
Browser Chromium (headless), vision-in-the-loop
Platform lifecycle Service Lifecycle Manager (SLM) — deploy / operate / scale
Containers Docker Compose
Infrastructure Ansible (deployment)

Security Model

  • All inter-service traffic stays on the internal Docker network
  • Only ports 80/443 are exposed externally
  • JWT-based authentication with configurable expiry
  • Multi-user accounts with role-based access control (RBAC) enforced on all backend routes
  • You choose where inference runs: local-only inference keeps everything on the host, or you can plug in an external provider through the LLM gateway. Your data lives on your infrastructure; nothing is sent to a provider you did not configure

Scalability

For larger deployments:

  • Backend can run multiple replicas behind a load balancer
  • Redis handles distributed task queuing
  • PostgreSQL supports read replicas for analytics workloads
  • ChromaDB can be replaced with a managed vector DB for very large corpora

See Deployment Guide for production topology examples.