COMP6246 — Machine Learning Technologies | University of Southampton
End-to-end machine learning pipeline for human activity classification from triaxial wearable accelerometer data. Compares unsupervised clustering, classical supervised learning, and deep learning under a strict subject-wise generalisation protocol designed to reflect real deployment conditions.
| Model | Weighted F1 | Accuracy |
|---|---|---|
| K-Means (baseline, k=7) | 0.608 | 0.713 |
| K-Means (tuned, k=10) | 0.684 | 0.737 |
| Random Forest (baseline) | 0.612 | 0.569 |
| Random Forest (tuned) | 0.622 | 0.584 |
| 1D CNN | 0.931 | 0.933 |
CNN satisfies business deployment constraints: Precision ≥ 75% and Recall ≥ 50% across all activity classes.
Triaxial accelerometer recordings from wearable sensors positioned at the back and thigh of 18 subjects, sampled at 100 Hz.
- Raw: 5,521,186 samples across 18 subjects
- Post-cleaning: 5,059,040 samples (Subject S007 excluded — systematic sensor malfunction)
- Windowed: 50,566 valid 2-second windows → 200 × 6 tensors
- Classes (7): Walking, Running, Shuffling, Stairs, Standing, Sitting, Lying
Dataset is included in the project folder and loaded directly in the notebook.
- Cycling activity removal (out of business scope per dataset specification)
- Stair label merging: ascending (4) + descending (5) → unified Stairs class (9)
- Subject S007 exclusion: flat-line readings and unphysical amplitude spikes confirmed sensor malfunction
- Sliding window segmentation: 2-second windows at 100 Hz, 1-second overlap → 200 × 6 tensors with majority-vote labels
- Subject-wise train/validation split: no subject appears in both sets — prevents leakage through inter-subject motion signatures
10-dimensional ENMO-based statistical feature vector per window: mean, standard deviation, median, IQR, signal energy.
K-Means — Unsupervised baseline. Majority-vote cluster-label mapping. Tuned across k ∈ {7, 8, 9, 10, 11, 12}; optimal k=10.
Random Forest — Classical supervised baseline on engineered features. Hyperparameter search over: n_estimators {200, 400}, max_depth {20, 30, 40, None}, min_samples_leaf {2, 5}, class_weight {balanced, None}.
1D CNN — Three sequential Conv1D blocks (32 → 64 → 128 filters, kernel=5), each followed by BatchNorm and MaxPool. GlobalAveragePool + Dropout → 7-class softmax. ~38,000 trainable parameters. Operates directly on raw 200 × 6 windows.
# Clone the repository
git clone <repo-url>
cd har-activity-recognition
# Install dependencies
pip install torch numpy pandas matplotlib scikit-learn jupyter
# Launch the notebook
jupyter notebook HAR_Pipeline.ipynbThe notebook runs all three pipelines sequentially. Training curves, confusion matrices, and per-class classification reports are generated inline.
M1/Apple Silicon: CNN training uses the legacy Adam implementation for macOS M1 compatibility. No additional configuration required.
- Subject-wise validation split — ensures performance estimates reflect genuine generalisation to unseen individuals, not memorised subject-specific gait patterns
- Weighted F1 as primary metric — accounts for class imbalance; raw accuracy is misleading when postural classes (sitting, standing) dominate the dataset
- Business constraint evaluation — results assessed against minimum precision (≥75%) and recall (≥50%) thresholds for all operational activity classes
Developed as individual coursework for COMP6246 — Machine Learning Technologies, University of Southampton, 2025–26. Implemented and evaluated independently on M1 Mac (Apple Silicon).