A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Sep 8, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
[MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
A Next-Generation Training Engine Built for Ultra-Large MoE Models
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers — with support for 300+ large language models (Qwen3-Next, GLM-5.2, Deepseek-V4, MiniMax-2.7, ...) and 200+ multimodal large models (Qwen3.5, Qwen3-Omni, Gemma4, ...).
Deploy open-source LLMs on AWS in minutes — with OpenAI-compatible APIs and a powerful CLI/SDK toolkit.
GGUF Loader with its Agentic Mode, and floating button, ai Models | Open Source & Offline. Mistral, Deepseek, llama, gemma, qwen
What does gpt-oss tell us about OpenAI's training data?
Tiny engine, immense models — run large MoE LLMs (gpt-oss, Mixtral, Qwen3-MoE) on ordinary machines by streaming experts from disk. OpenAI-compatible server with tool calling + hybrid cloud relay; CPU, Apple Silicon & CUDA (MLX).
agentsculptor is an experimental AI-powered development agent designed to analyze, refactor, and extend Python projects automatically. It uses an OpenAI-like planner–executor loop on top of a vLLM backend, combining project context analysis, structured tool calls, and iterative refinement. It has only been tested with gpt-oss-120b via vLLM.
A local RAG + web search pipeline with gpt-oss and other similar scale models powered by llama.cpp
Batch processing for overnight tasks with gpt-oss 20b
Docker service stack for a comprehensive local ai experience, running on a normal home hardware setup with single GPU. Project-Status: Working, Backlog.
[AICI-26] Difficulty-Aware Adaptive Reasoning for Vietnamese VQA with GPT-OSS
Local LLM benchmark on a Lenovo ThinkStation PGX (NVIDIA GB10 / DGX Spark) - 11 models, 10 embedders, Czech RAG, reasoning on/off
Multi‑agent research AI workflow with cloud API and llama.cpp support, OpenAI vector_search or ChromaDB retrieval, Docker stacks (local & NVIDIA DGX Spark), and model benchmarking.
To associate your repository with the gpt-oss topic, visit your repo's landing page and select "manage topics."