A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
-
Updated
Sep 2, 2026 - Python
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Hybrid Automatic Video Colorizer (HAVC) server that exposes a GPU-accelerated colorization pipeline for B&W images and video frames based on Diffusion Transformer (DiT) models: Qwen-Image-Edit-2511 (with Nunchaku transformer) and LongCat ImageEdit Turbo. Colorized reference frames are propagated using CMNET2.
W4A4 and INT8 KV-cache quantization for Infinity VAR models. Optimized for high-fidelity generative AI deployment on edge GPUs (e.g. NVIDIA Jetson).
Krea 2 Turbo at 4-bit weights and activations on Nunchaku fused kernels. Runtime port + deepcompressor conversion.
To associate your repository with the svdquant topic, visit your repo's landing page and select "manage topics."