Self-hosted LLM model specs, benchmarks, and evaluation notes for Frank Lite deployment candidates.
| File | Model | Params | Context | Fits 14GB? |
|---|---|---|---|---|
| poolside-laguna-xs2.md | Poolside Laguna XS.2 | 33B (3B active) | 131K | |
| xiaomi-mimo-v2.5.md | Xiaomi MiMo-V2.5 | 310B (15B active) | 1M | ❌ Multi-GPU |
| aeon-qwen3.6-27b-ultimate.md | AEON Qwen3.6-27B Ultimate | 27B dense | 128K | |
| ibm-granite-4.1-8b.md | IBM Granite 4.1 8B | 8B dense | Long-ctx | ✅ ~5GB |
| mistral-medium-3.5-128b.md | Mistral Medium 3.5 | 128B dense | 256K | ❌ API only |
| nvidia-nemotron-3-nano-omni-30b.md | NVIDIA Nemotron-3 Nano Omni 30B | 31B (3B active) | 256K |
See ANALYSIS.md for full evaluation and recommendations.
TL;DR:
- Now: IBM Granite 4.1 8B — fits frank-gpu, Apache 2.0, RAG-explicit, benchmark against frank:v3
- After GPU upgrade: NVIDIA Nemotron-3 Nano Omni 30B NVFP4 — document intelligence focus, 256K ctx
- API tier: Mistral Medium 3.5 — best overall, 256K ctx, configurable reasoning
- Current frank-gpu: 14GB VRAM
- Planned upgrade: A40 ~24GB VRAM
- CPU Worker: 32GB RAM (embedding only, no GPU)