MLX
MLX is a NumPy-like array framework designed for efficient and flexible machine learning on Apple silicon, brought to you by Apple machine learning research.
Here are 67 public repositories matching this topic...
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
-
Updated
Sep 13, 2026 - Rust
Open source, terminal-first AI coding agent with fully transparent multi-model routing. Local (Ollama, LM Studio, MLX, llama.cpp) or cloud, your keys, one TUI. No black box.
-
Updated
Aug 25, 2026 - Rust
A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.
-
Updated
Aug 22, 2026 - Rust
Fast, local, simple voice typing for macOS.
-
Updated
Aug 31, 2026 - Rust
-
Updated
Sep 10, 2026 - Rust
ML model ports (ASR · TTS · LLM · VLM · vision · audio codecs) on the rlx multi-backend compiler + runtime — CPU/Metal/MLX/CUDA/wgpu
-
Updated
Aug 12, 2026 - Rust
A rust mlx server, ui, and model router for fast, dependency free inference on apple architecture
-
Updated
Sep 13, 2026 - Rust
One Mac process. Many models. Real speed. Multi-model LLM serving with prefix reuse, MTP acceleration, and OpenAI APIs — built for Apple Silicon, measured against mlx-lm and llama.cpp.
-
Updated
Sep 11, 2026 - Rust
The free, open-source meeting-notes app for your Mac — offline transcription and AI notes, free in the cloud or 100% on-device.
-
Updated
Sep 14, 2026 - Rust
A CLI in Rust to generate synthetic data for MLX friendly training
-
Updated
Jan 13, 2024 - Rust
The local inference server for coding agents. Pure Rust, one binary, Apple Silicon + NVIDIA. Anthropic + OpenAI APIs; the KV cache survives across turns, so turn 20 starts as fast as turn 2. OPD training on the same runtime.
-
Updated
Sep 13, 2026 - Rust
Set one RAM ceiling. Run bigger sparse language and vision models on Apple Silicon — experts stream from SSD when they don't fit, without hidden quantization. https://aosama.github.io/astronomical/ Join Discord Server https://discord.gg/dc4E6r4WD
-
Updated
Sep 14, 2026 - Rust
Local AI model manager
-
Updated
May 21, 2026 - Rust
Native Rust, single-binary MLX inference server for Apple Silicon. OpenAI/Anthropic-compatible LLM serving — text, vision, audio, embeddings — with the widest weight×KV quantization matrix of any MLX server. No Python, no GGUF.
-
Updated
Sep 14, 2026 - Rust