Stars
TurboQuant: Near-optimal KV cache quantization for LLM serving on AMD GPUs (arXiv: 2504.19874, ICLR 2026)
Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X
Achieve state of the art inference performance with modern accelerators on Kubernetes
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs
how to optimize some algorithm in cuda.
SGLang is a high-performance serving framework for large language models and multimodal models.
hkgood / Ollama_ChatTTS
Forked from jianchang512/ChatTTS-uiLLM voice chat project by Connect ChatTTS with Local Ollama, 连接本地部署的 Ollama 和 ChatTTS,实现和LLM的语音对话
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
🚀 Collection of components for development, training, tuning, and inference of foundation models leveraging PyTorch native components.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
A high-throughput and memory-efficient inference and serving engine for LLMs
ChatOllama is an open-source AI chatbot that brings cutting-edge language models to your fingertips while keeping your data private and secure.
AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术
Robust Speech Recognition via Large-Scale Weak Supervision
AMD Ryzen™ AI Software includes the tools and runtime libraries for optimizing and deploying AI inference on AMD Ryzen™ AI powered PCs.
Making large AI models cheaper, faster and more accessible
Vitis AI is Xilinx’s development stack for AI inference on Xilinx hardware platforms, including both edge devices and Alveo cards.



