Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
-
Updated
Aug 19, 2026 - Jupyter Notebook
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
Production-pattern Red Hat OpenShift AI 3.4.0 platform with bare-metal ESXi, GPU passthrough, KServe RawDeployment, DeepSeek R1 inference at 12–17 tok/s
A privacy-first Slack bot that integrates local LLMs (Ollama/vLLM/LM Studio) with advanced tools like ComfyUI image generation, SearXNG local search, and On-Demand RAG Memory. Analyze files, execute Python code, and generate music - all while keeping your data inside your own network.
Complete self-hosted AI server stack on the Nvidia DGX Spark (arm64)
To associate your repository with the vllm-inference topic, visit your repo's landing page and select "manage topics."