ggml
Here are 276 public repositories matching this topic...
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
-
Updated
Aug 19, 2026 - C++
ggml speech-to-text inference for 16+ model families
-
Updated
Aug 24, 2026 - C++
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
-
Updated
Mar 23, 2025 - C++
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
-
Updated
Dec 3, 2024 - JavaScript
This custom_node for ComfyUI adds one-click "Virtual VRAM" for any UNet and CLIP loader as well MultiGPU integration in WanVideoWrapper, managing the offload/Block Swap of layers to DRAM *or* VRAM to maximize the latent space of your card. Also includes nodes for directly loading entire components (UNet, CLIP, VAE) onto the device you choose
-
Updated
May 8, 2026 - Python
KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM
-
Updated
Aug 21, 2026 - C++
Suno AI's Bark model in C/C++ for fast text-to-speech generation
-
Updated
Nov 16, 2024 - C++
Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX + ggml, 70-joint skeleton with hands.
-
Updated
Aug 18, 2026 - C
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
-
Updated
Aug 25, 2026 - C++
Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)
-
Updated
Aug 8, 2023 - C++
CLIP inference in plain C/C++ with no extra dependencies
-
Updated
Aug 24, 2026 - C++
Improve this page
Add a description, image, and links to the ggml topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the ggml topic, visit your repo's landing page and select "manage topics."