Single-binary GGUF model runtime in C — CPU/CUDA/Metal, OpenAI-compatible. Serves, scores, and trains LoRA directly through the quantized weights it deploys, with byte-reproducible adapters. Tool calls survive the token limit; sparse MoE and schema-constrained decoding included.
c metal json-schema cuda inference moe llama granite ai-agents mixture-of-experts edge-ai openai-api constrained-decoding llm local-llm function-calling qwen gguf tool-calling structured-outputs
-
Updated
Sep 1, 2026 - C