Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
-
Updated
Aug 1, 2026 - Python
Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
Evidence-led vLLM tuning and qualification notes for NVIDIA CMP 170HX (sm80)
Strict-FP32 CUDA SGEMM with shape-dispatched SM80 cp.async kernels and reproducible A100 evidence
Add a description, image, and links to the sm80 topic page so that developers can more easily learn about it.
To associate your repository with the sm80 topic, visit your repo's landing page and select "manage topics."