Skip to content
#

awq

Here are 35 public repositories matching this topic...

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.

  • Updated Jul 17, 2026
  • Python

Native Windows vLLM 0.25.1—prebuilt Python 3.13/CUDA 12.8 wheels for RTX 30/40/50 GPUs, no WSL or Docker. OpenAI-compatible serving with Triton/FlashAttention, 10 KV-cache compression formats, Multi-TurboQuant, and experimental persistent CPU/NVMe KV offload.

  • Updated Jul 19, 2026
  • Python

Improve this page

Add a description, image, and links to the awq topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the awq topic, visit your repo's landing page and select "manage topics."

Learn more