Skip to content
#

gemv

Here are 16 public repositories matching this topic...

C++/CUDA inference engine for Qwen2.5-Coder-0.5B, written from scratch: INT4 GEMV/GEMM kernels, Flash-Decoding, a 24-layer decode engine and lossless speculative decoding. Verified against Hugging Face; benchmarked against bandwidth ceilings, PyTorch and llama.cpp.

  • Updated Oct 10, 2026
  • Python

Add this topic to your repo

To associate your repository with the gemv topic, visit your repo's landing page and select "manage topics."

Learn more