This repository contains the experiment code for the MSc thesis:
Prompting and Fine-tuning Large Language Models for Interlinear Glossing
The code is organized by model family and covers the training, inference, retrieval, and evaluation pipelines used in the thesis.
glosslm/— GlossLM fine-tuning, inference, prompt-variant evaluation, and WAV2Gloss cross-dataset evaluation.polygloss/— PolyGloss inference and evaluation on the PolyGloss dataset.llama/— LLaMA prompting, chrF++ retrieval, LoRA fine-tuning, cached few-shot fine-tuning, and PolyGloss cross-dataset evaluation.
The repository intentionally excludes:
- model checkpoints and LoRA adapters
- cached datasets and retrieved-example caches
- prediction files and evaluation outputs
- logs and SLURM job files
- local virtual environments
- Hugging Face caches and tokens
These files are either too large, environment-specific, or should not be committed.
The experiments rely on external datasets and models, including:
- GlossLM corpus / GlossLM split data
- PolyGloss corpus
- WAV2Gloss text data
meta-llama/Llama-3.1-8B-Instruct- GlossLM and PolyGloss checkpoints
Local paths should be supplied by the user through config files or command-line arguments.
The code is intended to document and reproduce the thesis experiment pipelines. Exact reproduction requires access to the same model checkpoints, dataset versions, and local compute environment.