An AI-powered Enterprise Support Chatbot that uses Retrieval-Augmented Generation (RAG) to answer questions from uploaded PDF documents.
The application allows users to upload one or multiple PDF documents, processes their content, converts the text into semantic embeddings, retrieves the most relevant sections for a user query, and generates a context-aware response using a locally hosted Llama 3 model through Ollama.
Built with Python, LangChain, FAISS, HuggingFace Embeddings, Ollama, and Streamlit.
-
π Multiple PDF Upload
- Upload one or multiple PDF documents directly through the Streamlit interface.
- PDF content is automatically extracted and processed.
-
π Semantic Search
- Converts document chunks into vector embeddings.
- Uses FAISS for efficient similarity-based retrieval.
- Retrieves the top 4 most relevant document chunks for each question.
-
π§ Retrieval-Augmented Generation
- Retrieved document context is passed to the LLM.
- Llama 3 generates answers based on the retrieved information rather than relying solely on model knowledge.
-
π Local AI Processing
- Uses Ollama to run Llama 3 locally.
- No external LLM API is required for answer generation.
-
π¬ Interactive Chat Interface
- Built with Streamlit.
- Supports conversational question-and-answer interaction.
-
π Source Visibility
- Displays the source documents associated with retrieved content.
-
π Document Summarization
- Includes a document summarization component using Llama 3.
βββββββββββββββββββββββ
β PDF Documents β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β PDF Loader β
β PyPDFLoader β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Text Splitter β
β Chunk: 1000 chars β
β Overlap: 200 chars β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β HuggingFace BGE β
β Embeddings β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β FAISS Vector β
β Store β
ββββββββββββ¬βββββββββββ
β
User Question
β
βΌ
βββββββββββββββββββββββ
β Semantic Retrieval β
β Top 4 β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Context β
β + User Query β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Ollama β
β Llama 3 β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Grounded Answer β
β + Sources β
βββββββββββββββββββββββ
The application follows a standard Retrieval-Augmented Generation workflow:
Users upload PDF files through the Streamlit sidebar.
PDF 1
PDF 2
PDF 3
β
PDF Loader
The application supports multiple PDFs in a single session.
PDF documents are loaded using LangChain's PyPDFLoader.
Each PDF is converted into LangChain Document objects containing the extracted text and metadata.
Large documents are divided into smaller chunks using RecursiveCharacterTextSplitter.
Current configuration:
Chunk Size : 1000 characters
Chunk Overlap : 200 characters
The overlap helps preserve contextual information between adjacent chunks.
Each document chunk is converted into a numerical vector using:
BAAI/bge-small-en-v1.5
The embedding model runs on CPU and uses normalized embeddings.
Document Chunk
β
BGE-Small
β
Vector Embedding
The generated embeddings are stored in a FAISS vector index.
FAISS enables fast similarity-based retrieval of relevant document chunks.
When a user asks a question, the application performs semantic similarity search against the vector store.
The current retriever returns:
Top K = 4
most relevant document chunks.
For example:
User:
"What is the company's refund policy?"
β
Semantic Search
β
βββββββββββββββββββββββββββββββ
β Chunk 1 β Refund Policy β
β Chunk 2 β Return Conditions β
β Chunk 3 β Refund Timeline β
β Chunk 4 β Exceptions β
βββββββββββββββββββββββββββββββ
The retrieved chunks are combined into a context and passed to Llama 3 through Ollama.
The model receives:
Retrieved Context
+
User Question
β
Llama 3
β
Generated Answer
The chatbot uses a temperature of 0 to produce more deterministic responses.
| Technology | Purpose |
|---|---|
| Python | Core application |
| Streamlit | Web interface |
| LangChain | RAG pipeline orchestration |
| PyPDFLoader | PDF document extraction |
| RecursiveCharacterTextSplitter | Document chunking |
| HuggingFace | Embedding generation |
| BAAI/bge-small-en-v1.5 | Semantic embeddings |
| FAISS | Vector similarity search |
| Ollama | Local LLM inference |
| Llama 3 | Answer generation |
support-chatbot-rag/
β
βββ data/
β βββ Uploaded PDF documents
β
βββ src/
β βββ __init__.py
β βββ chatbot.py
β β βββ Llama 3 RAG response generation
β β
β βββ embeddings.py
β β βββ HuggingFace embedding model
β β
β βββ pdf_loader.py
β β βββ PDF loading and extraction
β β
β βββ prompts.py
β β βββ RAG prompt configuration
β β
β βββ retriever.py
β β βββ Semantic document retrieval
β β
β βββ summarizer.py
β β βββ Document summarization
β β
β βββ text_splitter.py
β β βββ Document chunking
β β
β βββ vector_store.py
β βββ FAISS vector store management
β
βββ utils/
β βββ Utility modules
β
βββ app.py
β βββ Streamlit application
β
βββ requirements.txt
β βββ Python dependencies
β
βββ README.md
The repository currently follows this high-level structure with data, src, utils, app.py, and requirements.txt.
git clone https://github.com/prathamesh1079/support-chatbot-rag.gitcd support-chatbot-ragpython -m venv venvvenv\Scripts\activatepython3 -m venv venvsource venv/bin/activatepip install -r requirements.txtThis project uses Ollama to run Llama 3 locally.
Install Ollama and then download the required model:
ollama pull llama3Verify that the model is available:
ollama listThe application is configured to use:
Model: llama3
Temperature: 0
Start the Streamlit application:
streamlit run app.pyThe application will open in your browser.
Typically:
http://localhost:8501
Use the sidebar to upload one or multiple PDF files.
Upload PDFs
βββ company_policy.pdf
βββ employee_handbook.pdf
βββ support_documentation.pdf
Enter a question in the chat box.
Example:
What is the refund policy?
The system searches the FAISS vector store and retrieves the four most relevant chunks.
The retrieved context is passed to Llama 3.
The chatbot generates an answer based on the retrieved document content.
The application provides a Sources section showing the source metadata associated with retrieved documents.
Company_Policies.pdf
How many days do I have to request a refund?
Question
β
Query Embedding
β
FAISS Similarity Search
β
Top 4 Relevant Chunks
β
Context Construction
β
Llama 3
β
Answer
According to the company policy, customers can request
a refund within the specified refund period mentioned
in the policy document.
The answer is generated using the retrieved document context rather than directly querying an external knowledge source.
Responsible for loading and parsing one or multiple PDF files.
PDFLoader()It uses LangChain's PyPDFLoader internally.
Breaks documents into manageable chunks.
DocumentSplitter(
chunk_size=1000,
chunk_overlap=200
)This improves retrieval by allowing the vector store to search smaller sections of the documents.
The project uses:
BAAI/bge-small-en-v1.5
for semantic representation of document chunks.
FAISS manages the document embeddings and performs similarity search.
FAISS.from_documents(
documents,
embeddings
)The project also includes functionality to save and load the FAISS index locally.
The retriever performs semantic search and returns the top four relevant documents.
similarity_search(
query,
k=4
)The chatbot uses Ollama with Llama 3:
ChatOllama(
model="llama3",
temperature=0
)The retrieved documents are supplied as context to the model before generating the final response.
Traditional LLM applications can generate answers based on their pretrained knowledge, which may not contain an organization's latest internal information.
RAG addresses this by introducing a retrieval step:
User Question
β
βΌ
Retrieve Relevant
Documents
β
βΌ
Provide Context
β
βΌ
LLM
β
βΌ
Grounded Answer
This makes the chatbot particularly useful for:
- Enterprise documentation
- Internal policies
- Product documentation
- Employee handbooks
- Technical manuals
- Customer support knowledge bases
This project demonstrates practical implementation of:
- Retrieval-Augmented Generation
- Semantic search
- Vector databases
- Document ingestion
- PDF processing
- Text chunking
- HuggingFace embeddings
- FAISS similarity search
- LangChain pipelines
- Local LLM inference
- Prompt engineering
- Streamlit application development
- Source-aware question answering
Potential improvements include:
- Add conversational memory
- Persist uploaded document indexes between sessions
- Add document management and deletion
- Support DOCX, TXT and CSV files
- Add citation-level source references
- Display retrieved text chunks
- Add retrieval confidence scores
- Add hybrid keyword + semantic search
- Add re-ranking for improved retrieval
- Add authentication and user management
- Add chat history persistence
- Add streaming LLM responses
- Deploy using Docker
- Add evaluation metrics for retrieval and answer quality
- The current implementation focuses on PDF documents.
- Llama 3 must be available through a local Ollama installation.
- Embeddings currently run on CPU.
- The vector store is created from the uploaded documents during application processing.
- Retrieval currently returns a fixed top-4 set of documents.
- Answer quality depends on the quality of the uploaded documents and retrieved context.
Prathamesh Tekale
Computer Science Engineering Student & Software Developer
GitHub: @prathamesh1079
π Multi-PDF Document Processing
β
βοΈ Intelligent Text Chunking
β
π§ BGE Semantic Embeddings
β
π FAISS Vector Search
β
π Top-4 Context Retrieval
β
π€ Llama 3 via Ollama
β
π¬ Grounded Support Response
This project is intended for educational and demonstration purposes.