Software-FP8 (E5M2) weight-only GEMV: 4 codes per int32, decoded in-register inside the K-loop — 2x decode bandwidth on GPUs without FP8 hardware
-
Updated
Sep 15, 2026 - Python
Software-FP8 (E5M2) weight-only GEMV: 4 codes per int32, decoded in-register inside the K-loop — 2x decode bandwidth on GPUs without FP8 hardware
Benchmark workbench for MAX (Mojo) LLM decode kernels vs llama.cpp / cuBLAS / FlashInfer and the memory roofline on consumer NVIDIA GPUs (sm_86/sm_89). A public record, not a competing kernel library.
C++/CUDA inference engine for Qwen2.5-Coder-0.5B, written from scratch: INT4 GEMV/GEMM kernels, Flash-Decoding, a 24-layer decode engine and lossless speculative decoding. Verified against Hugging Face; benchmarked against bandwidth ceilings, PyTorch and llama.cpp.
To associate your repository with the gemv topic, visit your repo's landing page and select "manage topics."