Skip to content
#

awq

Here are 82 public repositories matching this topic...

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.

  • Updated Sep 2, 2026
  • Python

Native Windows vLLM 0.27.1 wheels: Python 3.13, PyTorch 2.13 + CUDA 13.0, SM 7.5-12.0 for RTX 20/30/40/50, OpenAI-compatible serving, FlashAttention/Rust, 10 KV formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload - no WSL or Docker.

  • Updated Aug 21, 2026
  • Python

Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.

  • Updated Aug 19, 2026

Add this topic to your repo

To associate your repository with the awq topic, visit your repo's landing page and select "manage topics."

Learn more