LegalIA is a production-ready, domain-specific AI assistant engineered for legal research and statutory compliance within the Colombian legal framework.
Unlike generic LLM wrappers, LegalIA implements an evidence-first RAG architecture guided by a non-negotiable principle:
Legal analysis demands zero tolerance for hallucinations. If the retrieval pipeline cannot locate verifiable normative provisions or binding jurisprudential precedent, the system explicitly refuses to synthesize an answer rather than guessing.
LegalIA adheres to Clean / Hexagonal Architecture principles, decoupling business rules from external AI providers and data stores.
flowchart TD
subgraph Ingestion["Ingestion & Indexing Engine"]
DOC["Colombian Legal Corpus<br/>(Laws, Codes, Sentencias)"] --> SPLIT["Hierarchical LegalSplitter<br/>(Libro > Título > Capítulo > Artículo)"]
SPLIT --> EMB["Dense Vector Embeddings<br/>(1024-d via pgvector)"]
SPLIT --> FTS["Spanish Full-Text Search<br/>(tsvector + GIN Index)"]
end
subgraph Retrieval["Hybrid Search & Reranking"]
QUERY["User Legal Query"] --> VEC["Cosine Vector Search"]
QUERY --> LEX["Lexical Search (Spanish Stemmer)"]
VEC --> RRF["Reciprocal Rank Fusion (RRF)"]
LEX --> RRF
RRF --> RERANK["Cross-Encoder Reranker<br/>(Contextual Precision Scoring)"]
end
subgraph Synthesis["Synthesis & Verification Guardrail"]
RERANK --> THRESH{Evidence Score<br/>>= Threshold?}
THRESH -- Refusal --> REFUSE["Explicit Refusal:<br/>'NO EVIDENCE → NO ANSWER'"]
THRESH -- Pass --> LLM["LLM Synthesis (OpenAI / Anthropic / Mock)"]
LLM --> VERIFY["Secondary Fact Verification Audit<br/>(Cross-checks tokens against source chunks)"]
VERIFY --> STREAM["Streaming Response + Interactive Citations"]
end
┌────────────────────────┐
│ Web Client / SPA │
│ (React 18 + TS + Vite)│
└───────────┬────────────┘
│ HTTPS
▼
┌────────────────────────┐
│ Caddy (TLS & Ingress) │
└─────┬────────────┬─────┘
│ │
/api/* requests│ │ /* (Static assets)
▼ ▼
┌───────────────────────┐ ┌────────────────────────┐
│ LegalIA API (FastAPI) │ │ React SPA (Nginx) │
│ Python 3.12+ │ │ Tailored Design System │
└───────────┬───────────┘ └────────────────────────┘
│
┌──────────────────┼──────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌───────────────┐ ┌─────────────────┐
│ PostgreSQL 16+ │ │ LLM Providers │ │ Ingestion │
│ pgvector │ │ (OpenAI / │ │ Pipeline │
│ Hybrid Indexing │ │ Anthropic) │ │ (LegalSplitter) │
└──────────────────┘ └───────────────┘ └─────────────────┘
- Statutory laws and court rulings cannot be segmented with generic character splitters without destroying normative context.
- LegalIA's structural parser recognizes Colombian normative containers (
LIBRO,TÍTULO,CAPÍTULO,ARTÍCULO,PARÁGRAFO) and court rulings (Considerandos,Resuelve). - Preserves exact character offsets and genealogical path hierarchy (e.g.
TÍTULO II > CAPÍTULO 1 > ARTÍCULO 13).
- Dense Vector Search: Powered by
pgvectorwith 1024-dimensional embeddings (Alibaba text-embedding-v4, BGE-M3, or local models). - Lexical Search: PostgreSQL full-text search (
tsvector/ GIN index) in Spanish to match exact law numerals and Latin maxims (tutela, habeas corpus, non bis in idem). - Reciprocal Rank Fusion: Merges scores to ensure both conceptual and verbatim relevance.
- Semantic Reranking: Re-scores candidate chunks through a cross-encoder prior to synthesis.
- Verification Layer: An independent secondary verification pass audits factual statements against retrieved chunks before streaming output to the user.
- Citation Tracking: Emits interactive citations linking each answer token directly to official sources (SUIN-Juriscol, Diario Oficial, Corte Constitucional).
- Live token quota meter calculating input, output, and verifier token consumption directly from database
UsageLogrecords. - Complete content privacy: operational logs contain zero prompts or legal dossier texts.
| Layer | Technology | Purpose |
|---|---|---|
| Backend | Python 3.12+, FastAPI, SQLAlchemy 2, Alembic | Async REST API, Hexagonal domain services |
| Database | PostgreSQL 16, pgvector | Relational data, vector embeddings, full-text search |
| Frontend | React 18, TypeScript, Vite | Modern SPA, streaming responses, citation modal |
| LLM Adapters | OpenAI-compatible, Anthropic SDK, Mock | Dynamic provider switching (Claude, GPT, Gemini) |
| Embeddings | Alibaba DashScope, BGE-M3, Mock | Multilingual dense vector generation |
| Reranking | Alibaba GTE Reranker, Mock | Cross-encoder contextual precision |
| Infrastructure | Docker, Docker Compose, Caddy, Nginx | Containerized deployment, automatic HTTPS |
LegalIA/
├── backend/ # FastAPI Application (Clean / Hexagonal Architecture)
│ ├── alembic/ # Database schema migrations (pgvector enabled)
│ ├── app/
│ │ ├── api/routes/ # REST API endpoints & SSE streaming handlers
│ │ ├── core/ # Security (JWT, Argon2), configuration & settings
│ │ ├── db/models/ # SQLAlchemy 2 models (chunks, documents, usage)
│ │ ├── providers/ # LLM, Embedding & Reranking adapters (Hexagonal Ports)
│ │ └── services/ # RAG pipeline orchestration, verification & search
│ └── tests/ # 140+ automated unit & integration test suites
├── frontend/ # React 18 + TypeScript + Vite SPA
│ └── src/
│ ├── components/ # UI components (SidePanel, CitationsModal, QuotaMeter)
│ └── styles/ # Modern CSS tokens & accessibility-first styling
├── ingestion/ # Document extraction, structural splitters & indexing CLI
├── evaluation/ # Precision/Recall benchmarks & legal retrieval evaluation
├── corpus/ # Sample Colombian legal corpus (Constitución, Códigos)
├── docs/ # Comprehensive technical documentation & guides
├── scripts/ # Administrative tools & deployment automation
├── docker-compose.yml # Development multi-container orchestration
└── Makefile # Developer workflow & automation commands
- Docker and Docker Compose
- Git
git clone https://github.com/lawxr/LegalIA.git
cd LegalIAcp .env.example .envEdit .env to configure your preferred LLM provider (openai_compatible, anthropic, or mock for local offline testing):
# Example using OpenAI or OpenAI-compatible endpoint
LLM_PROVIDER=openai_compatible
LLM_BASE_URL=https://api.openai.com/v1
LLM_API_KEY=your_api_key_here
LLM_MODEL=gpt-4o
# Database
POSTGRES_DB=legalia
POSTGRES_USER=legalia
POSTGRES_PASSWORD=your_secure_passworddocker compose up -d --buildThe stack will spin up 4 optimized containers:
legalia-caddy: Ingress reverse proxy & SSL termination (http://localhost:80)legalia-frontend: Nginx serving the React SPA (http://localhost/)legalia-api: FastAPI backend (http://localhost:8000/)legalia-postgres: PostgreSQL 16 with pgvector
docker compose exec legalia-api alembic upgrade headdocker compose exec legalia-api python -m ingestion.main --file corpus/constitucion_extracto.txt --status VIGENTELegalIA includes an extensive automated test suite covering authentication, legal chunking, hybrid retrieval, security, and reranking:
# Run tests inside virtual environment or container
pytest backend/tests -v
# Run with test coverage report
pytest backend/tests --cov=app --cov-report=term-missingFrontend typecheck and build validation:
cd frontend
npm install
npm run buildExplore deep-dive technical guides in the docs/ directory:
- 🏛️ System Architecture: Structural overview, clean domain boundaries, and data flow.
- 🔬 RAG Pipeline & Legal Intelligence: Hierarchical chunking, hybrid search, and verification algorithms.
- 🚢 Deployment Guide: Production deployment, SSL configuration, and environment setup.
- 🎨 Design System: Tailored CSS design tokens, typography, and accessibility guidelines.
- ⚖️ Corpus Sources: Primary sources for Colombian statutory laws and constitutional jurisprudence.
- 📊 Ingestion Status & Corpus Metrics: Processing status, corpus coverage, and extraction audit trail.
- Data Minimization: Prompts, answers, and case files are excluded from logging pipelines.
- Authentication: JWT tokens signed with HS256, Argon2 password hashing, and role-based access.
- Network Isolation: PostgreSQL runs inside an isolated internal Docker bridge network without direct host port exposure in production.
This project is licensed under the MIT License.
Disclaimer: LegalIA is an AI-powered legal intelligence tool designed to assist legal professionals. Outputs do not constitute formal legal counsel.