Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
-
Updated
Aug 7, 2026 - Python
Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
Expert streaming inference engine for MoE models larger than VRAM — run 235B+ models on consumer GPUs
Experimental llama.cpp runtime for bounded DeepSeek V4 MoE streaming on dual 16 GB GPUs
Add a description, image, and links to the expert-streaming topic page so that developers can more easily learn about it.
To associate your repository with the expert-streaming topic, visit your repo's landing page and select "manage topics."