Skip to content
#

low-vram

Here are 73 public repositories matching this topic...

WeeLLM

WeeLLM runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.

  • Updated Oct 1, 2026
  • Python

Hierarchical RAG architecture scaling to 693K chunks on consumer hardware (4GB VRAM). Features 3-address routing, hybrid vector+graph fusion, and SetFit classification.

  • Updated Feb 11, 2026
  • Python
qwen-image2.1-uncensored-comfyui

ComfyUI deployment kit for Qwen-Image-2.1 Uncensored (GGUF): ready-to-use t2i & image-edit workflows, ComfyUI-GGUF architecture patches, and a 14-second lossless Q8_0 quantization repair tool that fixes 'weight shape [136] vs normalized_shape [128]'. Runs on 8 GB VRAM. Weights NOT included.

  • Updated Sep 21, 2026
  • Python

Add this topic to your repo

To associate your repository with the low-vram topic, visit your repo's landing page and select "manage topics."

Learn more