Helm charts for running large language models on Kubernetes: the inference engine itself, the cache-aware router in front of it, and the pieces that keep both healthy.
helm repo add modelsphere https://modelsphere.github.io/helm-charts
helm repo update
helm search repo modelsphere| Chart | What it deploys |
|---|---|
sglang |
An SGLang inference deployment — single-node or multi-node (LeaderWorkerSet) — with optional CART, autoscaling, and a hang watcher |
vllm |
The same, on vLLM |
cart |
CART on its own: routes each request to the replica that already holds the longest matching prompt prefix |
continuation-gateway |
continuation_gateway on its own: when a streamed completion stalls or drops mid-generation, resumes it once from what the client already received, so the client sees one complete stream |
llm-slo-decision-gen |
Turns SLO requirements into replica recommendations: a decision service plus the SLO storage API and its two CRDs |
autoconfig |
Keeps the routing layer in step with what is actually deployed: watches backends and rewrites OpenResty peers and cache-aware-router workers |
llmscaleoperator |
The autoscaler the sglang and vllm charts hand their LLMScaler objects to: scales replicas on KV-cache utilization, queue depth and TPM rather than CPU |
rdma-injector |
A mutating webhook that injects NCCL_IB_HCA and the node's RDMA device list into pods labelled rdma-ib: "true" |
console |
The ModelSphere community portal — identity (users, roles, login) and a federation gateway to Swiss and other backends |
swiss |
swissd, the deploy control plane for the sglang and vllm charts: reads the model catalog, manages releases through its own ServiceAccount, serves the web UI |
sglang and vllm pull in cart as a subchart, gated on cart.enabled.
Installing either of them gives you an engine and a router that already know
about each other.
sglang also pulls in continuation-gateway as a subchart, gated on
continuationGateway.enabled (off by default). Enabled, it sits in front of the
model's cart; see the comment above continuationGateway in
charts/sglang/values.yaml for what that changes
in the route.
helm install my-model modelsphere/sglang \
--set model.name=my-model \
--set model.path=/models/my-model \
--set cart.enabled=trueEvery chart ships a commented values.yaml; start there rather than from this
README, since that is the file that is kept current.
The charts default to public images on Docker Hub under
4pdosc, alongside upstream
lmsysorg/sglang and vllm/vllm-openai for the engines. Point them at your own
registry by overriding the image values — a mirror is worth setting up if your
cluster has no egress.
Some optional features render custom resources that the chart does not define:
modelRoute.enabledneeds theModelRouteCRD from autoconfigscaler.enabledneeds theLLMScalerCRD from the scaling operator
With the CRD absent, enabling the feature fails the install with
no matches for kind .... The charts deliberately carry no copy of either
schema, because the operator that reconciles it owns it.
Charts live under charts/<name>/. Bump the chart's version in Chart.yaml
in the same change — CI publishes exactly those charts whose version moved, so a
change without a bump ships nothing.
Apache License 2.0 — see LICENSE.