Skip to content
@z-lab

Z Lab

Efficient AI. PI: Zhijian Liu

Popular repositories Loading

  1. dflash dflash Public

    DFlash: Block Diffusion for Flash Speculative Decoding

    Python 5.8k 412

  2. paroquant paroquant Public

    [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

    Python 331 33

  3. sparselora sparselora Public

    [ICML 2025] SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity

    Python 78 7

  4. flashdrive flashdrive Public

    Flash Vision-Language-Action Inference for Autonomous Driving

    Python 64 7

  5. flash-colreduce flash-colreduce Public

    Fast, memory-efficient attention column reduction (e.g., sum, mean, max)

    Python 51 3

  6. omlx-fork omlx-fork Public

    Forked from jundot/omlx

    LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

    Python 1 1

Repositories

Showing 10 of 10 repositories

Top languages

Python C++

Most used topics

Loading…