The open-model foundry — where the models this ecosystem routes to are trained, gated and published.
Documentation · The Foundry Pipeline · Catalog & GitOps · How a model is made
Nebula is where models are born. It is not a serving layer and not a router: it turns a dataset and a config into a published, versioned, evaluated model, and it refuses to publish one that does not clear its gate. Everything downstream — the runtime's model router, Orbit's inference stack, an Ollama pull — consumes its output, not its internals.
First family: Centinela, Spanish-first models for finance and back-office work in LATAM.
A Spanish financial sentiment classifier (positivo / neutral / negativo), QLoRA
fine-tuned from Qwen/Qwen3-4B, Apache-2.0, with a deterministic output-validation layer that
constrains generation to the three labels. Runs cheap and self-hosted, even on CPU via GGUF.
Status: the pipeline is complete and locally verifiable end to end; the public Hub release is not out.
catalog/centinela.yamlstill carries placeholder revision SHAs and zeroed eval metrics on purpose — the catalog is filled in by a real run, never by hand. Treat theollama runline below as the shape of the consumer story, not as a live artifact.
ollama run hf.co/astromesh/Centinela-Qwen3-4B:Q4_K_MEach stage is a script under scripts/, driven by the Makefile, and each one is independently
runnable — a failed quantize does not cost you the training run.
make scout → survey candidate base models
make dataset → build and split the training corpus
make train → QLoRA fine-tune (HF Jobs, or locally via WSL2 — see below)
make merge → fold the adapter into the base weights
make quantize → GGUF, Q4_K_M
make eval → the gate: per-metric thresholds, macro-F1 and invalid-rate
make publish → push weights + model card to the Hub
make catalog → record the revision in the GitOps catalog
make lock compiles catalog.lock.json, which ships inside the wheel — that is what a consumer
resolves an alias like prod against, so a rollback is a catalog commit rather than a redeploy.
GitHub runners have no GPU, so the heavy stages run elsewhere:
- HF Jobs (
make train, and thereleaseworkflow's full pipeline) — billed, needs a PRO/Team plan and a token with write + create-repo on theastromeshorg. - Local NVIDIA GPU via WSL2 (
scripts/02_train_local.py,scripts/wsl_train_*.sh) — the fallback when HF Jobs is unavailable. Plain transformers + PEFT + bitsandbytes + TRL, no Unsloth: the 2026.6.x Unsloth build globally patches TRL and is incompatible with TRL 0.24 (it injects an out-of-vocab<EOS_TOKEN>sentinel). Tuned for a 12 GB card — fp16 on Turing,paged_adamw_8bit, per-device batch 2 × grad-accum 8.
uv sync
make test # pytest, CPU only
make lint # ruffCI runs eval-gate (ruff + pytest, CPU) on every push and PR. release is manual
(workflow_dispatch): it authenticates to Hugging Face and launches the whole GPU pipeline as a
single HF Job. It accepts stop_before_publish, which runs train → merge → quantize → eval as a
dry run and prints the eval report to the job log without touching the public Hub — the intended
way to validate the first real GPU run before committing a v0.1 anyone can pull.
Required repository configuration:
| Name | Kind | Value / scope |
|---|---|---|
HF_ORG |
Variable | astromesh |
HF_TOKEN |
Secret | HF fine-grained token: write + create-repo on the astromesh org |
GH_PIPELINE_TOKEN |
Secret | GitHub fine-grained PAT, read-only contents on this repo (so the HF Job can clone it) |
| astromesh | The runtime that routes to what Nebula publishes; its model catalog reads the compiled lock. |
| astromesh-orbit | Provisions the inference stack a published model runs on. |
| The ecosystem map | Every other component, and where each one fits. |