[NeurIPS'23] Speculative Decoding with Big Little Decoder
-
Updated
Feb 6, 2024 - Python
[NeurIPS'23] Speculative Decoding with Big Little Decoder
Run AI models too large for your Mac's memory — at near-full speed. Intelligent expert caching, speculative execution, and 15+ research techniques for MoE inference on Apple Silicon.
State-aware hedged requests for serverless GPU inference — return the first valid result and cancel the losers with an audited receipt.
Run large MLX models on Apple Silicon with flash weight streaming, using native precision beyond RAM limits
Two-tier speculative intent architecture — fires on partial input, reconciles with LLM in parallel. Patent pending.
Extension of the ScheduleFlow Simulator to allow speculative request times at submission and during backfill
Add a description, image, and links to the speculative-execution topic page so that developers can more easily learn about it.
To associate your repository with the speculative-execution topic, visit your repo's landing page and select "manage topics."