Parameter-efficient fine-tuning of a transformer encoder for 6-class emotion classification
(sadness · joy · love · anger · fear · surprise) on dair-ai/emotion,
using LoRA via Hugging Face peft. Training touches roughly 1% of the model's parameters, runs in minutes on a
single consumer GPU, and produces a reproducible evaluation report plus a ready-to-publish model card.
| full fine-tuning | LoRA (this repo) | |
|---|---|---|
| trainable parameters (distilroberta-base) | ~82 M | ~0.9 M (adapters + classifier head) |
| artifact to store / ship per task | ~330 MB | a few MB |
| GPU memory | full optimizer state | small fraction |
| accuracy on this task | baseline | typically within ~1 pt of full FT |
Low-rank adapters are injected into the attention query/value projections; the base weights stay frozen and can be
shared across many tasks.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 1. train (≈3 epochs; GPU recommended, CPU works with --train-subset for a smoke test)
python train.py --model distilroberta-base --epochs 3 --output-dir outputs/distilroberta-lora
# 2. evaluate on the test split → metrics.json, classification_report.txt, confusion_matrix.png, model card
python evaluate.py --adapter-dir outputs/distilroberta-lora/adapter
# 3. predict
python predict.py --adapter-dir outputs/distilroberta-lora/adapter \
"i cannot believe we actually won" "why does everything have to go wrong today"Smoke test on a laptop CPU (a couple of minutes, low accuracy, just to verify the pipeline):
python train.py --epochs 1 --train-subset 2000 --batch-size 16 --output-dir outputs/smokeoutputs/distilroberta-lora/
├── adapter/ # LoRA weights + tokenizer + README.md (model card) → push this to the Hub
├── checkpoints/ # Trainer checkpoints (best model kept)
├── train_summary.json # hyper-parameters, parameter counts, val/test metrics, wall time
└── report/
├── metrics.json # accuracy, macro / weighted F1, confusion matrix
├── classification_report.txt
└── confusion_matrix.png
Fill this table from report/metrics.json after your run (hardware, seed and epochs matter):
| base model | trainable % | test accuracy | test F1 (macro) | epochs | hardware |
|---|---|---|---|---|---|
| distilroberta-base | 3 |
huggingface-cli login
python train.py ... --push-to-hub <user>/emotion-lora
# or, after evaluate.py has written the model card:
huggingface-cli upload <user>/emotion-lora outputs/distilroberta-lora/adapter .| flag | default | notes |
|---|---|---|
--model |
distilroberta-base |
any encoder with a sequence-classification head (roberta-base, microsoft/deberta-v3-base, xlm-roberta-base for multilingual) |
--target-modules |
query,value |
DeBERTa uses query_proj,value_proj; check model.named_modules() |
--lora-r / --lora-alpha |
16 / 32 | rank and scaling; alpha = 2·r is a common default |
--lr |
3e-4 | LoRA tolerates ~10× the learning rate of full fine-tuning |
--max-length |
128 | the dataset's tweets are short; raise for longer texts |
pip install -r requirements-dev.txt
ruff check .
pytest -qMIT. The dataset has its own terms; see the dair-ai/emotion dataset card.