Skip to content
#

int8

Here are 48 public repositories matching this topic...

Custom CUDA kernels for KV-cache eviction + INT8 quantized paged attention (vLLM-oriented), PyTorch C++ extensions with Python API: eviction @16k blocks 1707.75us host -> 562us fused (~3x); INT8 attention p50 1.76-14.18ms; ~49.9% KV memory vs FP16; 26/26 pytest passing on T1000.

  • Updated Jul 14, 2026
  • Python

Improve this page

Add a description, image, and links to the int8 topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the int8 topic, visit your repo's landing page and select "manage topics."

Learn more