Repository navigation
Architecture & System Design
mrveiss edited this page Jul 22, 2026
·
2 revisions
AutoBot is a microservices-based platform where a Vue.js frontend communicates with a FastAPI backend. The backend orchestrates AI inference, vector search, relational data, caching, fleet operations, and browser automation as separate concerns. Inference is routed through an LLM gateway so you can choose local models, a hosted provider, or a mix — your data stays on your infrastructure either way. A Service Lifecycle Manager (SLM) handles deploy, operate, and scale for the platform itself.
graph TB
User["User (Browser)"]
Frontend["Frontend (Vue.js)"]
Backend["Backend (FastAPI)"]
Redis["Redis (Cache/Queue)"]
PostgreSQL["PostgreSQL (Data)"]
ChromaDB["ChromaDB (Vectors)"]
Gateway["LLM Gateway (routing/fallback)"]
LocalLLM["Local Inference (Ollama)"]
Provider["External Providers (claude_api, openrouter)"]
SLM["Service Lifecycle Manager (deploy/operate/scale)"]
Ansible["Ansible (Fleet Ops)"]
Browser["Browser Automation (Chromium)"]
User -->|HTTP/WebSocket| Frontend
Frontend -->|REST API| Backend
Backend -->|Read/Write| Redis
Backend -->|Query| PostgreSQL
Backend -->|Vector Search| ChromaDB
Backend -->|Inference| Gateway
Gateway -->|Local| LocalLLM
Gateway -->|Optional provider| Provider
Backend -->|Execute Playbooks| Ansible
Backend -->|Control| Browser
SLM -.->|Deploy / operate / scale| Backend
- Framework: Vue.js 3 with Composition API
- Build: Vite
- UI: Custom component library, xterm.js for terminal views
- Served by: nginx (also handles TLS termination and reverse proxy)
- Communicates with: Backend via REST and WebSocket
- Framework: FastAPI (Python)
- Async: Full async/await with asyncio
- Auth: JWT tokens with multi-user accounts and role-based access control (RBAC) — the platform is multi-user; permissions are enforced per role on every route
-
Key modules:
-
api/— REST route handlers -
agents/— AI agent loop and orchestration -
agent_loop/— Multi-turn conversation management -
advanced_rag_optimizer.py— RAG retrieval and reranking
-
- Engine: PostgreSQL
- Stores: Users, sessions, fleet inventory, knowledge base metadata, workflow definitions, audit logs
- Role: Task queue (Celery/background jobs) + caching layer
- Used for: Chat session state, async job results, rate limiting
- Role: Vector database for RAG (Retrieval Augmented Generation)
- Stores: Embeddings for knowledge base documents, code, runbooks
- Used for: Semantic search across all indexed content
- Role: Single entry point for all model inference; routes each request to the configured backend
-
Providers: Local inference (Ollama) and external providers (
claude_api,openrouter) — you choose per deployment - Features: Provider routing and fallback, so a request can fail over to an alternate provider
- Your data, your choice: run fully local, use a hosted provider, or mix — your data stays on your infrastructure and no provider is contacted unless you configure one
- Role: Local AI inference backend behind the gateway
- Runs: Open-source models (Llama, Mistral, etc.) entirely on your hardware
- Use when: you want inference to stay on-box with no external calls
- Role: Deploys, operates, and scales the AutoBot platform itself
- Handles: installation, upgrades/self-update, and multi-node scaling of the stack
-
Driven by: Ansible playbooks (e.g. the installer's
deploy-slm-manager.yml) - Note: "SLM" always means Service Lifecycle Manager — it is not a language-model term
- Role: Fleet operations engine
- Executes: Playbooks for multi-server deployments, updates, configuration management
- Driven by: Backend API commands triggered from chat or workflows
- Role: Browser automation via Chromium with vision-in-the-loop
- Used for: Automated browsing where the agent captures the page, reasons over what it sees, and drives the next action; also web scraping and screenshot-based troubleshooting
User types message
→ Frontend (WebSocket)
→ Backend: agent_loop receives message
→ Backend: RAG query to ChromaDB (if knowledge base attached)
→ Backend: context assembled (history + retrieved docs)
→ LLM Gateway: inference request routed to the configured provider (local Ollama or an external provider)
→ Backend: streams response tokens back
→ Frontend: renders streaming response to user
User uploads document
→ Backend: file received, stored in PostgreSQL (metadata)
→ Backend: document chunked and embedded
→ ChromaDB: embeddings stored
→ Available for RAG queries immediately
User issues fleet command in chat
→ Backend: intent parsed by agent
→ Backend: Ansible playbook selected or generated
→ Ansible: executes against target nodes
→ Backend: captures stdout/stderr, updates job status
→ Frontend: shows real-time progress
→ PostgreSQL: audit log written
| Layer | Technology |
|---|---|
| Frontend | Vue.js 3, Vite, nginx |
| Backend | FastAPI, Python 3.11+, asyncio |
| Relational DB | PostgreSQL 15+ |
| Cache / Queue | Redis 7+ |
| Vector DB | ChromaDB |
| AI Inference | LLM gateway — local models via Ollama and/or external providers (claude_api, openrouter) |
| Fleet Ops | Ansible |
| Browser | Chromium (headless), vision-in-the-loop |
| Platform lifecycle | Service Lifecycle Manager (SLM) — deploy / operate / scale |
| Containers | Docker Compose |
| Infrastructure | Ansible (deployment) |
- All inter-service traffic stays on the internal Docker network
- Only ports 80/443 are exposed externally
- JWT-based authentication with configurable expiry
- Multi-user accounts with role-based access control (RBAC) enforced on all backend routes
- You choose where inference runs: local-only inference keeps everything on the host, or you can plug in an external provider through the LLM gateway. Your data lives on your infrastructure; nothing is sent to a provider you did not configure
For larger deployments:
- Backend can run multiple replicas behind a load balancer
- Redis handles distributed task queuing
- PostgreSQL supports read replicas for analytics workloads
- ChromaDB can be replaced with a managed vector DB for very large corpora
See Deployment Guide for production topology examples.