01_Basics/ SLURM, Modules, Miniconda Environment Setup
02_Notebook_Parallelism/ DP, TP, PP concept notebooks
03_LLM_Training_Techniques/ DDP, Tensor, Pipeline training on GPT-2 Small
04_LLM_Finetunning/ FSDP, TorchTune, Unsloth
- Python ≥ 3.11
- PyTorch ≥ 2.1
- CUDA ≥ 11.8
- NVIDIA GPU
- Conda / Miniconda
- Distributed Data Parallel (DDP)
- Tensor Parallelism
- Pipeline Parallelism
- Fully Sharded Data Parallel (FSDP)
- Fine-tuning workflows
git clone https://github.com/kishoryd/LLMTraining.git
cd LLMTrainingThis section is for users running training jobs via Slurm batch scripts:
sbatch <script_name>.sh
squeue --meCheck output and error logs generated in the same directory:
ls *.out *.errThis section is for users replicating the environment on other clusters.
cd /scratch/$USER
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh --no-check-certificateRefer to setup details in 01_Basics for Jupyter and environment configuration.
chmod 755 Miniconda3-latest-Linux-x86_64.sh
./Miniconda3-latest-Linux-x86_64.shconda create --name Training \
python=3.11 \
pytorch-cuda=12.1 \
pytorch cudatoolkit xformers -c pytorch -c nvidia -c xformers \
-ypip install matplotlib
pip install unsloth
pip install torchtunepip install unsloth
pip install torchtunepython -m ipykernel install --user --name jupyter-notebook --display-name "Training"- Miniconda – https://docs.conda.io/en/latest/miniconda.html
- Unsloth – https://github.com/unslothai/unsloth
- TorchTune – https://pytorch.org/torchtune
- PyTorch – https://pytorch.org
- Hugging Face – https://huggingface.co/docs
- Slurm – https://slurm.schedmd.com
-
Distributed Data Parallel (DDP): https://pytorch.org/docs/stable/generated/torch.nn.parallel.DistributedDataParallel.html
-
Tensor Parallelism: https://pytorch.org/docs/stable/distributed.tensor.parallel.html
-
Pipeline Parallelism: https://pytorch.org/docs/stable/distributed.pipeline.sync.html
-
Fully Sharded Data Parallel (FSDP): https://pytorch.org/docs/stable/fsdp.html