SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
-
Updated
Jul 31, 2026 - Python
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
a simple pipline of int8 quantization based on tensorrt.
Nodes to run Hunyuan Image 3 locally with BF16 and NF4 quantized options in Comfyui
👀 Apply YOLOv8 exported with ONNX or TensorRT(FP16, INT8) to the Real-time camera
Trustworthy onboard satellite AI in PyTorch→ONNX→INT8 with calibration, telemetry, and a PhiSat-2 EO tile-filter demo.
A LLaMA2-7b chatbot with memory running on CPU, and optimized using smooth quantization, 4-bit quantization or Intel® Extension For PyTorch with bfloat16.
TensorRT Int8 Python version sample. TensorRT Int8 Python 实现例子。TensorRT Int8 Pythonの例です
40x faster AI inference: ONNX to TensorRT optimization with FP16/INT8 quantization, multi-GPU support, and deployment
Domain-adaptive CLIP retrieval system for surgical video keyframe–text matching. Includes adapter training, INT8 acceleration, PQ compression, and Reliability analysis.
FLUX.1-dev on AMD Radeon consumer GPUs — fast, low-VRAM, and shippable. Backport patches + benchmarks for torchao + diffusers group_offload on ROCm.
Qwen3-8B quantization study across vLLM, TensorRT-LLM, AutoRound, INT8, and MXFP4
🗄️ Manage and search large vector datasets efficiently with this pure Python vector database featuring Int8 quantization and lazy deletion.
RDK X5 one-click YOLO .pt to ONNX / bayes-e INT8 .bin converter for Windows, with auto environment checks, OpenExplorer Docker build, deploy configs and reports.
Custom CUDA kernels for KV-cache eviction + INT8 quantized paged attention (vLLM-oriented), PyTorch C++ extensions with Python API: eviction @16k blocks 1707.75us host -> 562us fused (~3x); INT8 attention p50 1.76-14.18ms; ~49.9% KV memory vs FP16; 26/26 pytest passing on T1000.
Pure Python vector database • int8 quantized • ~1100 QPS @ 50k vectors • single file • no compile • MIT
Add a description, image, and links to the int8 topic page so that developers can more easily learn about it.
To associate your repository with the int8 topic, visit your repo's landing page and select "manage topics."