Run a local LLM on a Mac with an NVIDIA eGPU: custom DriverKit driver (no CUDA runtime), RTX 3090 over Thunderbolt — Qwen3.8-27B at 75.81 tok/s decode @100k ctx, bit-exact, OpenAI-compatible API.
macos inference nvidia thunderbolt egpu driverkit apple-silicon openai-api tinygrad llm local-llm qwen speculative-decoding rtx-3090 cuda-alternative
-
Updated
Oct 5, 2026 - Python