Local inference platform for K/IQ-quant GGUF models on Apple Silicon
-
Updated
Aug 2, 2026 - Python
Local inference platform for K/IQ-quant GGUF models on Apple Silicon
a GGUF inference runner in Rust
GGUF-Runner - Want to run LLMs locally, use this guide, and run with LLAMA.cpp
A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
Add a description, image, and links to the gguf-runner topic page so that developers can more easily learn about it.
To associate your repository with the gguf-runner topic, visit your repo's landing page and select "manage topics."