Skip to content

About

AI-powered support chatbot using RAG, FAISS, HuggingFace embeddings, and Llama 3 for context-aware Q&A from PDF documents.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ€– Enterprise Support Chatbot β€” RAG

An AI-powered Enterprise Support Chatbot that uses Retrieval-Augmented Generation (RAG) to answer questions from uploaded PDF documents.

The application allows users to upload one or multiple PDF documents, processes their content, converts the text into semantic embeddings, retrieves the most relevant sections for a user query, and generates a context-aware response using a locally hosted Llama 3 model through Ollama.

Built with Python, LangChain, FAISS, HuggingFace Embeddings, Ollama, and Streamlit.


✨ Features

  • πŸ“„ Multiple PDF Upload

    • Upload one or multiple PDF documents directly through the Streamlit interface.
    • PDF content is automatically extracted and processed.
  • πŸ” Semantic Search

    • Converts document chunks into vector embeddings.
    • Uses FAISS for efficient similarity-based retrieval.
    • Retrieves the top 4 most relevant document chunks for each question.
  • 🧠 Retrieval-Augmented Generation

    • Retrieved document context is passed to the LLM.
    • Llama 3 generates answers based on the retrieved information rather than relying solely on model knowledge.
  • πŸ”’ Local AI Processing

    • Uses Ollama to run Llama 3 locally.
    • No external LLM API is required for answer generation.
  • πŸ’¬ Interactive Chat Interface

    • Built with Streamlit.
    • Supports conversational question-and-answer interaction.
  • πŸ“š Source Visibility

    • Displays the source documents associated with retrieved content.
  • πŸ“ Document Summarization

    • Includes a document summarization component using Llama 3.

πŸ—οΈ Architecture

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚     PDF Documents   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚     PDF Loader      β”‚
                    β”‚     PyPDFLoader     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   Text Splitter     β”‚
                    β”‚ Chunk: 1000 chars   β”‚
                    β”‚ Overlap: 200 chars  β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ HuggingFace BGE     β”‚
                    β”‚ Embeddings          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚    FAISS Vector     β”‚
                    β”‚       Store         β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                         User Question
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ Semantic Retrieval  β”‚
                    β”‚      Top 4          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚      Context        β”‚
                    β”‚   + User Query      β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚      Ollama         β”‚
                    β”‚      Llama 3        β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   Grounded Answer   β”‚
                    β”‚    + Sources        β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ”„ RAG Pipeline

The application follows a standard Retrieval-Augmented Generation workflow:

1. Document Upload

Users upload PDF files through the Streamlit sidebar.

PDF 1
PDF 2
PDF 3
   ↓
PDF Loader

The application supports multiple PDFs in a single session.


2. PDF Extraction

PDF documents are loaded using LangChain's PyPDFLoader.

Each PDF is converted into LangChain Document objects containing the extracted text and metadata.


3. Text Chunking

Large documents are divided into smaller chunks using RecursiveCharacterTextSplitter.

Current configuration:

Chunk Size    : 1000 characters
Chunk Overlap : 200 characters

The overlap helps preserve contextual information between adjacent chunks.


4. Embedding Generation

Each document chunk is converted into a numerical vector using:

BAAI/bge-small-en-v1.5

The embedding model runs on CPU and uses normalized embeddings.

Document Chunk
      ↓
BGE-Small
      ↓
Vector Embedding

5. FAISS Vector Store

The generated embeddings are stored in a FAISS vector index.

FAISS enables fast similarity-based retrieval of relevant document chunks.


6. Semantic Retrieval

When a user asks a question, the application performs semantic similarity search against the vector store.

The current retriever returns:

Top K = 4

most relevant document chunks.

For example:

User:
"What is the company's refund policy?"

                ↓

        Semantic Search

                ↓

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚ Chunk 1 β€” Refund Policy     β”‚
 β”‚ Chunk 2 β€” Return Conditions β”‚
 β”‚ Chunk 3 β€” Refund Timeline   β”‚
 β”‚ Chunk 4 β€” Exceptions        β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

7. LLM Response Generation

The retrieved chunks are combined into a context and passed to Llama 3 through Ollama.

The model receives:

Retrieved Context
        +
User Question
        ↓
      Llama 3
        ↓
Generated Answer

The chatbot uses a temperature of 0 to produce more deterministic responses.


πŸ› οΈ Tech Stack

Technology Purpose
Python Core application
Streamlit Web interface
LangChain RAG pipeline orchestration
PyPDFLoader PDF document extraction
RecursiveCharacterTextSplitter Document chunking
HuggingFace Embedding generation
BAAI/bge-small-en-v1.5 Semantic embeddings
FAISS Vector similarity search
Ollama Local LLM inference
Llama 3 Answer generation

πŸ“‚ Project Structure

support-chatbot-rag/
β”‚
β”œβ”€β”€ data/
β”‚   └── Uploaded PDF documents
β”‚
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ chatbot.py
β”‚   β”‚   └── Llama 3 RAG response generation
β”‚   β”‚
β”‚   β”œβ”€β”€ embeddings.py
β”‚   β”‚   └── HuggingFace embedding model
β”‚   β”‚
β”‚   β”œβ”€β”€ pdf_loader.py
β”‚   β”‚   └── PDF loading and extraction
β”‚   β”‚
β”‚   β”œβ”€β”€ prompts.py
β”‚   β”‚   └── RAG prompt configuration
β”‚   β”‚
β”‚   β”œβ”€β”€ retriever.py
β”‚   β”‚   └── Semantic document retrieval
β”‚   β”‚
β”‚   β”œβ”€β”€ summarizer.py
β”‚   β”‚   └── Document summarization
β”‚   β”‚
β”‚   β”œβ”€β”€ text_splitter.py
β”‚   β”‚   └── Document chunking
β”‚   β”‚
β”‚   └── vector_store.py
β”‚       └── FAISS vector store management
β”‚
β”œβ”€β”€ utils/
β”‚   └── Utility modules
β”‚
β”œβ”€β”€ app.py
β”‚   └── Streamlit application
β”‚
β”œβ”€β”€ requirements.txt
β”‚   └── Python dependencies
β”‚
└── README.md

The repository currently follows this high-level structure with data, src, utils, app.py, and requirements.txt.


βš™οΈ Installation

1. Clone the Repository

git clone https://github.com/prathamesh1079/support-chatbot-rag.git
cd support-chatbot-rag

2. Create a Virtual Environment

Windows

python -m venv venv
venv\Scripts\activate

macOS / Linux

python3 -m venv venv
source venv/bin/activate

3. Install Dependencies

pip install -r requirements.txt

🧠 Install Ollama

This project uses Ollama to run Llama 3 locally.

Install Ollama and then download the required model:

ollama pull llama3

Verify that the model is available:

ollama list

The application is configured to use:

Model: llama3
Temperature: 0

▢️ Run the Application

Start the Streamlit application:

streamlit run app.py

The application will open in your browser.

Typically:

http://localhost:8501

πŸ’¬ How to Use

Step 1 β€” Upload Documents

Use the sidebar to upload one or multiple PDF files.

Upload PDFs
   β”œβ”€β”€ company_policy.pdf
   β”œβ”€β”€ employee_handbook.pdf
   └── support_documentation.pdf

Step 2 β€” Ask a Question

Enter a question in the chat box.

Example:

What is the refund policy?

Step 3 β€” Retrieval

The system searches the FAISS vector store and retrieves the four most relevant chunks.


Step 4 β€” Generate Answer

The retrieved context is passed to Llama 3.

The chatbot generates an answer based on the retrieved document content.


Step 5 β€” View Sources

The application provides a Sources section showing the source metadata associated with retrieved documents.


πŸ“Š Example

Uploaded Document

Company_Policies.pdf

User Query

How many days do I have to request a refund?

RAG Pipeline

Question
   ↓
Query Embedding
   ↓
FAISS Similarity Search
   ↓
Top 4 Relevant Chunks
   ↓
Context Construction
   ↓
Llama 3
   ↓
Answer

Example Response

According to the company policy, customers can request
a refund within the specified refund period mentioned
in the policy document.

The answer is generated using the retrieved document context rather than directly querying an external knowledge source.


🧩 Core Components

PDF Loader

Responsible for loading and parsing one or multiple PDF files.

PDFLoader()

It uses LangChain's PyPDFLoader internally.


Document Splitter

Breaks documents into manageable chunks.

DocumentSplitter(
    chunk_size=1000,
    chunk_overlap=200
)

This improves retrieval by allowing the vector store to search smaller sections of the documents.


Embedding Model

The project uses:

BAAI/bge-small-en-v1.5

for semantic representation of document chunks.


Vector Store

FAISS manages the document embeddings and performs similarity search.

FAISS.from_documents(
    documents,
    embeddings
)

The project also includes functionality to save and load the FAISS index locally.


Retriever

The retriever performs semantic search and returns the top four relevant documents.

similarity_search(
    query,
    k=4
)

LLM

The chatbot uses Ollama with Llama 3:

ChatOllama(
    model="llama3",
    temperature=0
)

The retrieved documents are supplied as context to the model before generating the final response.


πŸ” Why RAG?

Traditional LLM applications can generate answers based on their pretrained knowledge, which may not contain an organization's latest internal information.

RAG addresses this by introducing a retrieval step:

       User Question
             β”‚
             β–Ό
      Retrieve Relevant
        Documents
             β”‚
             β–Ό
       Provide Context
             β”‚
             β–Ό
          LLM
             β”‚
             β–Ό
      Grounded Answer

This makes the chatbot particularly useful for:

  • Enterprise documentation
  • Internal policies
  • Product documentation
  • Employee handbooks
  • Technical manuals
  • Customer support knowledge bases

🎯 Key Learning Outcomes

This project demonstrates practical implementation of:

  • Retrieval-Augmented Generation
  • Semantic search
  • Vector databases
  • Document ingestion
  • PDF processing
  • Text chunking
  • HuggingFace embeddings
  • FAISS similarity search
  • LangChain pipelines
  • Local LLM inference
  • Prompt engineering
  • Streamlit application development
  • Source-aware question answering

πŸš€ Future Improvements

Potential improvements include:

  • Add conversational memory
  • Persist uploaded document indexes between sessions
  • Add document management and deletion
  • Support DOCX, TXT and CSV files
  • Add citation-level source references
  • Display retrieved text chunks
  • Add retrieval confidence scores
  • Add hybrid keyword + semantic search
  • Add re-ranking for improved retrieval
  • Add authentication and user management
  • Add chat history persistence
  • Add streaming LLM responses
  • Deploy using Docker
  • Add evaluation metrics for retrieval and answer quality

πŸ“Œ Limitations

  • The current implementation focuses on PDF documents.
  • Llama 3 must be available through a local Ollama installation.
  • Embeddings currently run on CPU.
  • The vector store is created from the uploaded documents during application processing.
  • Retrieval currently returns a fixed top-4 set of documents.
  • Answer quality depends on the quality of the uploaded documents and retrieved context.

πŸ‘¨β€πŸ’» Author

Prathamesh Tekale

Computer Science Engineering Student & Software Developer

GitHub: @prathamesh1079


⭐ Project Highlights

πŸ“„ Multi-PDF Document Processing
        ↓
βœ‚οΈ Intelligent Text Chunking
        ↓
🧠 BGE Semantic Embeddings
        ↓
πŸ”Ž FAISS Vector Search
        ↓
πŸ“š Top-4 Context Retrieval
        ↓
πŸ€– Llama 3 via Ollama
        ↓
πŸ’¬ Grounded Support Response

πŸ“„ License

This project is intended for educational and demonstration purposes.

About

AI-powered support chatbot using RAG, FAISS, HuggingFace embeddings, and Llama 3 for context-aware Q&A from PDF documents.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages