Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
-
Updated
Oct 2, 2026 - Python
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
The most efficient one-page LoRA trainer for Anima 2B. Optimized for 6GB+ VRAM, featuring a smart dataset analyzer and real-time previews.
One-click Windows installer for Z-Image Turbo AI image generation. Optimized for low-VRAM GPUs (4GB+). Features Gradio web UI, automatic setup, and GGUF model support.
WeeLLM runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.
A ComfyUI Workflow for low vram users
Hierarchical RAG architecture scaling to 693K chunks on consumer hardware (4GB VRAM). Features 3-address routing, hybrid vector+graph fusion, and SetFit classification.
Taiwanese Hokkien (Taigi) speech-to-text transcriber - MediaTek Breeze-ASR-26 with faster-whisper, tuned for RTX 3050 4GB low-VRAM GPUs. Gradio UI, CLI, Docker, SRT/VTT/TXT/JSON.
"Adaptive Hybrid Quantization Framework for deploying 7B+ LLMs on low-VRAM devices (e.g., GTX 1050). Features surgical block alignment and Numba-accelerated inference.
Three MiniMax H3 ComfyUI workflows for 8GB laptop GPUs (20 / 8 / 4 steps). Includes the model list, launch flags, prompt-structure pitfalls, the resolution trap, and measured evidence for why you must restart ComfyUI before every run.
SCAIL-2 (Wan 2.1) Low-VRAM motion transfer — endless videos from a single reference image. 8+ GB GPUs.
llama.cpp fork tuned for running modern models (Gemma-4, Qwen3.x) at full context on 12 GB Turing GPUs (RTX 2060/2070/2080, T4). TurboQuant KV cache (KTQ+VTQ, 2.78 bpw f16-quality), SWA-aware KV, MTP+n-gram speculation.
Modly extension for Hunyuan3D 2.1 patched for Windows AMD and modest PCs
Native Hunyuan3D 2.1 Full extension for Modly, optimized for low-VRAM NVIDIA GPUs with INT8, FP8 and FP16 support.
ComfyUI deployment kit for Qwen-Image-2.1 Uncensored (GGUF): ready-to-use t2i & image-edit workflows, ComfyUI-GGUF architecture patches, and a 14-second lossless Q8_0 quantization repair tool that fixes 'weight shape [136] vs normalized_shape [128]'. Runs on 8 GB VRAM. Weights NOT included.
Contains the notebooks and workflows configured to run inference from Wan 2.2 Animate with ComfyUI on Kaggle T4 GPUs smoothly
Lightweight 6GB VRAM Gradio web app with auto-installer for running AuraFlow locally — no cloud, no clutter.
在 RTX 3060 6GB 上复现 SmolVLA × LIBERO,并开展低显存 LoRA 微调、多随机种子评测与遗忘控制实验。Low-VRAM SmolVLA × LIBERO reproduction with LoRA adaptation, fixed multi-seed evaluation, and forgetting controls.
Simple FP16 image upscaler for all GPUs (low-mid end users)
Type a prompt or drop up to 10 photos: Qwen-Image-2.1 renders it on your 12 GB GPU in ~35 s. Desktop app + CLI, honest benchmark vs gpt-image-2.
To associate your repository with the low-vram topic, visit your repo's landing page and select "manage topics."