An AI-powered teaching assistant that uses Retrieval-Augmented Generation (RAG) to answer questions from course video content.
The system processes educational videos, extracts their audio, transcribes and translates the content using OpenAI Whisper, converts the transcript into vector embeddings using Ollama, retrieves the most relevant content using cosine similarity, and generates context-aware answers using Llama 3.2.
The assistant can also identify which course video and timestamp contains the information relevant to a student's question.
-
🎥 Video-to-Audio Processing
- Extracts audio from course videos using FFmpeg.
- Automatically organizes tutorial metadata such as video number and title.
-
🎙️ Speech-to-Text & Translation
- Uses OpenAI Whisper (
large-v2) to transcribe course videos. - Supports Hindi speech and translates the content into English.
- Stores transcript segments along with timestamps.
- Uses OpenAI Whisper (
-
🧩 Semantic Chunking
- Splits course transcripts into timestamped content chunks.
- Preserves video title, tutorial number, start time, end time, and transcript text.
-
🔎 Semantic Retrieval
- Generates embeddings using Ollama's
bge-m3embedding model. - Uses cosine similarity to find the most relevant transcript chunks.
- Retrieves the top 5 matching chunks for each question.
- Generates embeddings using Ollama's
-
🧠 LLM-Powered Answers
- Uses Llama 3.2 through Ollama to generate natural-language responses.
- Answers are grounded in retrieved course content.
-
⏱️ Timestamp-Based Learning
- Identifies the relevant tutorial and approximate timestamp.
- Helps students directly navigate to the part of the course where a concept is explained.
-
🛡️ Course-Specific Responses
- The assistant is instructed to answer questions related to the course content.
- Unrelated questions are rejected instead of generating unsupported answers.
┌──────────────────────┐
│ Course Videos │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ FFmpeg │
│ Video → Audio │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ OpenAI Whisper │
│ Speech → Text │
│ + Translation │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Timestamped Chunks │
│ JSON Transcript │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Ollama BGE-M3 │
│ Embedding Generation │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ embeddings.joblib │
└──────────┬───────────┘
│
User Question
│
▼
┌──────────────────────┐
│ Query Embedding │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Cosine Similarity │
│ Top 5 Retrieval │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Llama 3.2 │
│ Context Generation │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Answer + Video │
│ + Timestamp │
└──────────────────────┘
| Technology | Purpose |
|---|---|
| Python | Core development |
| OpenAI Whisper | Speech recognition and translation |
| Ollama | Local LLM and embedding inference |
| BGE-M3 | Text embeddings |
| Llama 3.2 | Answer generation |
| Scikit-learn | Cosine similarity |
| NumPy | Numerical processing |
| Pandas | Data manipulation |
| Joblib | Embedding persistence |
| FFmpeg | Video-to-audio conversion |
| JSON | Transcript and metadata storage |
RAG_Based_AI_Teaching_Assistant/
│
├── audios/
│ └── Extracted course audio files
│
├── videos/
│ └── Course video files
│
├── jsons/
│ └── Timestamped transcript files
│
├── unused/
│ └── Experimental / unused files
│
├── create_chunks.py
│ └── Transcribes audio using Whisper and creates timestamped chunks
│
├── process_video.py
│ └── Extracts audio from course videos using FFmpeg
│
├── process_incoming.py
│ └── Main RAG query and answer-generation pipeline
│
├── read_chunks.py
│ └── Reads and processes transcript chunks
│
├── sample.py
│ └── Sample / experimental implementation
│
├── embeddings.joblib
│ └── Stored transcript embeddings
│
├── prompt.txt
│ └── Generated prompt supplied to the LLM
│
├── response.txt
│ └── Generated assistant response
│
└── README.md
Course videos are placed inside the videos/ directory.
The process_video.py script identifies the tutorial number and title and uses FFmpeg to extract the audio.
python process_video.pyThe extracted audio files are stored inside the audios/ directory.
The extracted audio is processed using Whisper.
The system uses the large-v2 Whisper model and processes Hindi speech while translating the transcript into English.
model = whisper.load_model("large-v2")Each transcript is divided into timestamped segments containing:
Video Number
Video Title
Start Time
End Time
Transcript
The resulting data is stored as JSON inside the jsons/ directory.
The transcript chunks are converted into numerical vector representations.
The project uses Ollama's bge-m3 embedding model.
Transcript Chunk
↓
BGE-M3
↓
Vector Embedding
The generated embeddings are stored using Joblib.
embeddings.joblib
The user enters a question such as:
Where is Flexbox explained?
The question is converted into an embedding using the same embedding model.
The query embedding is compared with the stored transcript embeddings using cosine similarity.
The system retrieves the top 5 most relevant transcript chunks.
similarities = cosine_similarity(
np.vstack(df['embedding']),
[question_embedding]
).flatten()This allows the system to perform semantic search instead of relying only on keyword matching.
The retrieved transcript chunks are passed to Llama 3.2 together with the user's question.
The model is instructed to:
- Answer using the retrieved course content.
- Identify the relevant video.
- Provide the relevant timestamp.
- Guide the student toward the appropriate section.
- Avoid answering unrelated questions.
git clone https://github.com/prathamesh1079/RAG_Based_AI_Teaching_Assistant.gitcd RAG_Based_AI_Teaching_Assistantpython -m venv venvActivate it on Windows:
venv\Scripts\activateOn macOS/Linux:
source venv/bin/activateInstall the required packages:
pip install pandas numpy scikit-learn joblib requests openai-whisperFFmpeg must also be installed and available in your system PATH.
Install Ollama and make sure the Ollama service is running.
Pull the required models:
ollama pull bge-m3ollama pull llama3.2The application communicates with Ollama locally through:
http://localhost:11434
Place your course videos inside:
videos/
Then run:
python process_video.pyRun:
python create_chunks.pyThis generates timestamped transcript JSON files.
Process the transcript chunks and generate the embedding dataset required by the RAG pipeline.
The resulting embeddings are stored in:
embeddings.joblib
Start the query pipeline:
python process_incoming.pyYou will be prompted:
Ask a Question:
Example:
Ask a Question: Where is JavaScript DOM manipulation explained?
The system retrieves the most relevant course sections and generates a contextual response.
What is Flexbox and where is it taught?
The system searches the transcript embeddings and retrieves the most semantically relevant chunks.
Top Result
-------------------------
Video: CSS Flexbox
Start: 523.4 seconds
End: 601.2 seconds
Similarity: High
Flexbox is explained in the CSS Flexbox tutorial.
You can find the explanation around 8:43 in the video.
I recommend starting from this section because it covers the
basic Flexbox concepts and how the layout works.
This project demonstrates practical implementation of:
- Retrieval-Augmented Generation (RAG)
- Semantic search
- Vector embeddings
- Large Language Models
- Speech-to-text processing
- Text translation
- Context-aware question answering
- Similarity-based information retrieval
- Local AI inference
- Educational AI applications
Potential improvements include:
- Build a Streamlit or web-based chat interface
- Add a proper vector database such as FAISS or Chroma
- Support multiple courses
- Add conversation memory
- Add clickable video timestamps
- Display retrieved source chunks
- Add confidence/relevance scores
- Support PDF and document-based course material
- Add authentication and user profiles
- Deploy the application as a web service
- The current implementation relies on locally running Ollama models.
- Processing large video collections requires significant computational resources.
- Whisper transcription can be time-consuming for long videos.
- The current retrieval mechanism uses cosine similarity over stored embeddings rather than a dedicated vector database.
- The assistant is designed primarily for questions related to the indexed course content.
Prathamesh Tekale
Computer Science Engineering Student & Software Developer
GitHub: @prathamesh1079
This project uses the following open-source technologies:
- OpenAI Whisper
- Ollama
- BGE-M3
- Llama 3.2
- Scikit-learn
- FFmpeg
- Python
This project is intended for educational and learning purposes.