An enterprise-grade Retrieval-Augmented Generation system for PDFs, tables, images, charts, and audio.
โText + Tables + Charts + Audio โ All in One RAG Systemโ ๐
Multi-Modal RAG 2.0 is a Retrieval-Augmented Generation system designed to process and retrieve information from multiple data modalities instead of relying only on plain text.
The system supports:
- ๐ PDF Text
- ๐ Tables
- ๐ผ๏ธ Images and Charts
- ๐๏ธ Audio and Video Transcriptions
Each modality is converted into searchable information and indexed in ChromaDB. Users can then ask questions through a unified interface, and the system retrieves relevant context before generating an answer with source information.
| Modality | Processing | Output |
|---|---|---|
| ๐ PDF Text | PyPDF2 / Unstructured | Extracted document text |
| ๐ Tables | Camelot / Tabula | Structured table content |
| ๐ผ๏ธ Images & Charts | Gemini Vision | Text descriptions and insights |
| ๐๏ธ Audio / Video | Whisper | Speech transcription |
graph TD
A[PDF / Image / Audio / Video] --> B[Multi-Modal Ingestion Pipeline]
B --> C[PDF Processor]
B --> D[Image Processor]
B --> E[Audio Processor]
C --> F[Text + Tables]
D --> G[Image / Chart Descriptions]
E --> H[Audio Transcriptions]
F --> I[Embedding Pipeline]
G --> I
H --> I
I --> J[ChromaDB Vector Database]
K[User Query] --> L[Retriever]
J --> L
L --> M[Relevant Multi-Modal Context]
M --> N[LLM Generator]
N --> O[Answer + Sources]
P[Streamlit UI] --> K
P --> Q[FastAPI Backend]
Q --> L
Upload File
โ
โผ
Identify File Type
โ
โโโ PDF โโโโโโโบ Extract Text + Tables
โ
โโโ Image โโโโโบ Generate Description
โ
โโโ Audio โโโโโบ Generate Transcription
โ
โผ
Create Embeddings
โ
โผ
Store in ChromaDB
โ
โผ
User Question โโโโโโโโบ Retrieve Relevant Context
โ
โผ
LLM Generation
โ
โผ
Final Answer + Sources
| Feature | Description |
|---|---|
| ๐ PDF Text Extraction | Extracts searchable text from PDF documents |
| ๐ Table Extraction | Extracts table content using Camelot or Tabula |
| ๐ผ๏ธ Image Understanding | Uses Gemini Vision to understand images, charts, and diagrams |
| ๐๏ธ Audio Transcription | Converts audio and video speech into searchable text |
| ๐๏ธ Vector Database | Stores embeddings and metadata in ChromaDB |
| ๐ Semantic Retrieval | Retrieves relevant information based on query meaning |
| ๐ค LLM Generation | Generates answers using Groq or OpenAI-compatible models |
| ๐ Source References | Returns relevant source information with answers |
| โก FastAPI Backend | Provides REST APIs for ingestion and querying |
| ๐ฅ๏ธ Streamlit Dashboard | Provides an interactive upload and chat interface |
| Layer | Technology | Purpose |
|---|---|---|
| Language | Python 3.11+ | Core development |
| Backend | FastAPI | REST API |
| Frontend | Streamlit | Interactive dashboard |
| Orchestration | LangChain | RAG and LLM integration |
| Vector Database | ChromaDB | Embedding storage and retrieval |
| Embeddings | Sentence Transformers | Semantic text embeddings |
| PDF Processing | PyPDF2 / Unstructured | Text extraction |
| Table Extraction | Camelot / Tabula | PDF table extraction |
| Vision | Google Gemini | Image and chart understanding |
| Audio Processing | Whisper | Audio transcription |
| LLM | Groq / OpenAI | Answer generation |
multi-modal-rag/
โ
โโโ src/
โ โโโ ingestion/
โ โ โโโ pdf_processor.py
โ โ โโโ image_processor.py
โ โ โโโ audio_processor.py
โ โ โโโ vector_index.py
โ โ โโโ pipeline.py
โ โ
โ โโโ retrieval/
โ โ โโโ hybrid_retriever.py
โ โ
โ โโโ generation/
โ โ โโโ generator.py
โ โ
โ โโโ app.py
โ
โโโ data/
โ โโโ raw/
โ
โโโ chroma_db/
โ
โโโ dashboard.py
โโโ requirements.txt
โโโ .env
โโโ .gitignore
โโโ README.md
Before running the project, make sure you have:
- Python 3.11+
- Conda or virtual environment support
- Groq API Key
- Google Gemini API Key
- OpenAI API Key (optional)
git clone https://github.com/vaibhav07772/multi-modal-rag.git
cd multi-modal-ragconda create -n multi-rag python=3.11 -y
conda activate multi-ragpip install -r requirements.txtCreate a .env file in the project root:
GROQ_API_KEY=gsk_xxxxxxxxxxxxxxxxxxxxxx
GOOGLE_API_KEY=your_google_api_key
OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxxxxxxxx
OPENAI_API_KEYis optional if you are using only Groq and Gemini.
uvicorn src.app:app --host 127.0.0.1 --port 8000 --reloadThe API server will run at:
http://127.0.0.1:8000
Open another terminal, activate the environment, and run:
conda activate multi-rag
streamlit run dashboard.pyThe Streamlit application will open at:
http://localhost:8501
| Resource | URL |
|---|---|
| ๐ฅ๏ธ Streamlit Dashboard | http://localhost:8501 |
| ๐ก FastAPI Server | http://127.0.0.1:8000 |
| ๐ API Documentation | http://127.0.0.1:8000/docs |
| โค๏ธ Health Check | http://127.0.0.1:8000/health |
| Method | Endpoint | Description |
|---|---|---|
POST |
/ingest |
Upload and process PDF, image, or audio files |
POST |
/ask |
Ask questions using multi-modal context |
GET |
/health |
Check application health |
GET |
/docs |
Interactive Swagger API documentation |
POST /ingest
Content-Type: multipart/form-data
file: <binary file>
The system detects the file type and routes it to the appropriate processor.
{
"query": "What is the main topic of the document?",
"top_k": 5
}{
"answer": "The document discusses...",
"sources": [
{
"source": "file.pdf",
"type": "pdf_text"
}
]
}Upload technical PDFs and ask questions about their content.
"What are the main requirements mentioned in this document?"
Extract information from PDF tables.
"What was the highest revenue in the financial table?"
Analyze charts and graphs.
"What trend does this chart show?"
Upload an audio recording and ask:
"What are the main decisions discussed in this meeting?"
Upload multiple documents and media files and ask questions across all of them.
graph LR
A[PDF] --> E[Knowledge Base]
B[Tables] --> E
C[Images / Charts] --> E
D[Audio / Video] --> E
E --> F[ChromaDB]
F --> G[Semantic Retrieval]
G --> H[LLM]
H --> I[Grounded Answer]
Multi-Modal RAG 2.0 Dashboard โ Upload PDFs, images, charts, or audio files, process them through the multi-modal RAG pipeline, and ask questions using AI.
- Multi-modal embeddings using CLIP
- Support for ImageBind
- OCR support for scanned PDFs
- Microsoft Word document support
- Excel and CSV support
- PowerPoint presentation support
- Real-time streaming ingestion
- Hybrid retrieval with BM25 + Vector Search
- Cross-Encoder reranking
- Agentic RAG workflow with LangGraph
- Query rewriting and self-correction
- Docker containerization
- Cloud deployment
- Monitoring and observability
Contributions are welcome!
- Fork the repository
- Create a new branch
git checkout -b feature/your-feature-name- Make your changes
- Commit your changes
git commit -m "Add new feature"- Push the branch
git push origin feature/your-feature-name- Open a Pull Request
- Use
blackfor Python formatting - Use
isortfor import sorting - Add docstrings to important functions
- Keep modules clean and modular
This project is released under the MIT License.
You are free to use, modify, and distribute this project.
Vaibhav Singh
- GitHub: https://github.com/vaibhav07772
- LinkedIn: https://www.linkedin.com/in/vaibhav07772/
- Email: vs9502778@gmail.com
If you find this project useful, please consider giving the repository a star โญ.
Your support helps the project reach more developers and AI enthusiasts!
๐ Multi-Modal RAG 2.0
"Text + Tables + Charts + Audio โ All in One RAG System" ๐
Built with โค๏ธ using Python, LangChain, ChromaDB, FastAPI, Streamlit, Gemini, Whisper, and Groq.
