Skip to content
View andyluo7's full-sized avatar

Block or report andyluo7

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

TurboQuant: Near-optimal KV cache quantization for LLM serving on AMD GPUs (arXiv: 2504.19874, ICLR 2026)

Python 4 1 Updated Apr 23, 2026

AMD_Robotics_Hackathon_2025

Jupyter Notebook 13 11 Updated Jan 21, 2026

Fast and Furious AMD Kernels

C++ 446 75 Updated Jul 10, 2026

Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.

Python 68,942 6,206 Updated Jul 27, 2026

gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI

Python 20,260 2,126 Updated Jul 24, 2026

Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X

C++ 79 6 Updated Feb 11, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 3,881 638 Updated Jul 26, 2026

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation

Python 4,694 370 Updated Jun 21, 2025

AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs

TypeScript 49,027 4,671 Updated Jul 27, 2026

how to optimize some algorithm in cuda.

Cuda 3,150 285 Updated Jul 24, 2026

MAD (Model Automation and Dashboarding)

Shell 39 50 Updated Jul 23, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 30,776 7,413 Updated Jul 27, 2026

Sample GLM4V + ChatTTS AI assistant

Python 85 9 Updated Jun 6, 2024

LLM voice chat project by Connect ChatTTS with Local Ollama, 连接本地部署的 Ollama 和 ChatTTS,实现和LLM的语音对话

Python 64 9 Updated Aug 9, 2024

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,223 2,609 Updated Jul 27, 2026
Python 33 9 Updated Jul 22, 2026

🚀 Collection of components for development, training, tuning, and inference of foundation models leveraging PyTorch native components.

Python 232 108 Updated Jul 16, 2026

Mirage Persistent Kernel: Compiling LLMs into a MegaKernel

Cuda 2,392 234 Updated Jul 25, 2026

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

Cuda 1,286 244 Updated Jul 27, 2026

Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.

Python 23,350 2,491 Updated Apr 29, 2025

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,261 19,903 Updated Jul 27, 2026

ChatOllama is an open-source AI chatbot that brings cutting-edge language models to your fingertips while keeping your data private and secure.

TypeScript 3,499 560 Updated May 28, 2026

AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术

Jupyter Notebook 17,296 2,426 Updated Sep 3, 2025
Python 619 58 Updated Jul 31, 2024

Robust Speech Recognition via Large-Scale Weak Supervision

Python 28 2 Updated Feb 3, 2025

AMD Ryzen™ AI Software includes the tools and runtime libraries for optimizing and deploying AI inference on AMD Ryzen™ AI powered PCs.

Python 859 134 Updated Jul 23, 2026

Making large AI models cheaper, faster and more accessible

Python 41,425 4,503 Updated Jul 13, 2026
61 12 Updated Sep 15, 2023

Vitis AI is Xilinx’s development stack for AI inference on Xilinx hardware platforms, including both edge devices and Alveo cards.

Python 1,797 674 Updated Feb 24, 2026
Next