From-scratch CPU vs CUDA implementations of 10 ML algorithms - AdamW, Nadam, L-BFGS, GMM-EM, Kernel PCA, MLP, Random Forest and more. No PyTorch, no cuDNN, no cuBLAS. Up to 24.7x GPU speedup.
-
Updated
May 31, 2026 - Cuda
From-scratch CPU vs CUDA implementations of 10 ML algorithms - AdamW, Nadam, L-BFGS, GMM-EM, Kernel PCA, MLP, Random Forest and more. No PyTorch, no cuDNN, no cuBLAS. Up to 24.7x GPU speedup.
Hand-written CUDA kernels for neural-network training primitives — optimizers, losses, activations and matrix multiplication — behind a Python ctypes interface
Add a description, image, and links to the optimizers topic page so that developers can more easily learn about it.
To associate your repository with the optimizers topic, visit your repo's landing page and select "manage topics."