Skip to content

About

Adaptive model-capacity orchestration — dynamically select, load, cache, and budget adapters, experts, models, and future model blocks.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

SelectiveLLM

Adaptive Model-Capacity Orchestration

CI Python 3.11+ License: Apache-2.0 Status: Research Prototype

SelectiveLLM explores a simple question: why activate or keep all available model capacity when a request may need only part of it?

Instead of treating a language-model stack as one fixed block of computation, SelectiveLLM treats model capacity as a runtime resource that can be selected, loaded, cached, rejected, evicted, or expanded depending on the request and the available budget.

Today the framework supports real PEFT/LoRA capacity routing, runtime residency management, memory-aware planning, caching, and multiple routing strategies.

The broader research direction is task-conditioned adaptive computation across adapters, experts, models, and eventually hardware-efficient internal model capacity.

                         Prompt
                           │
                           ▼
                  ┌─────────────────┐
                  │  SelectiveLLM   │
                  │    Controller   │
                  └────────┬────────┘
                           │
                What capacity is useful?
                           │
            ┌──────────────┼──────────────┐
            ▼              ▼              ▼
         Adapter         Expert         Model
            │              │              │
            └──────────────┼──────────────┘
                           ▼
                  Budget-aware planner
                           │
               memory / latency / locality
                           │
                           ▼
                  Selected computation
                           │
                           ▼
                       Inference

Long-term goal: execute the minimum sufficient computation required for each request while preserving useful output quality.


Why SelectiveLLM?

Most inference systems begin with a fixed assumption:

Load the available model capacity and run it.

SelectiveLLM investigates a different assumption:

Determine which capacity is likely to help first, then spend memory and computation selectively.

That creates a general runtime optimization problem:

request
   ↓
capacity candidates
   ↓
expected usefulness
   +
memory cost
   +
load latency
   +
cache locality
   +
uncertainty
   ↓
runtime capacity plan

The current implementation focuses on independently loadable adapters and experts because they provide a practical and measurable testbed for this idea.

The architecture is deliberately broader than LoRA routing.

Capacity type Current status
PEFT / LoRA adapters ✅ Implemented and measured
Runtime expert loading ✅ Implemented
Memory-budget planning ✅ Implemented
Dependency-aware cache / eviction ✅ Implemented
Multiple routing policies ✅ Implemented
Accelerator memory telemetry ✅ MPS / CUDA where available
Learned capacity router 🔬 Planned
Calibrated uncertainty / abstention 🔬 Planned
Multi-model capacity routing 🔬 Planned
Utility-aware capacity optimization 🔬 Planned
Dense parameter-block selection 🧪 Research
Physical semantic parameter paging ❌ Not yet demonstrated

30-second demo

Clone the repository and run the offline demo:

git clone https://github.com/AltanCetinCelik/SelectiveLLM.git
cd SelectiveLLM

python3.11 -m venv .venv
source .venv/bin/activate

pip install -e .
selectivellm demo

The demo:

  • analyzes each request,
  • identifies candidate capabilities,
  • routes the request,
  • creates a memory-constrained capacity plan,
  • loads or reuses selected components,
  • reports selected and rejected capacity,
  • reports cache activity,
  • reports stage-level latency,
  • and records memory semantics explicitly.

Example workflow:

Prompt:
"Use Python to simulate an RLC circuit."

Detected capabilities:
  python                  0.83
  electrical_engineering  0.78
  mathematics             0.45

Candidate capacity:
  python_expert
  electronics_expert
  math_expert

Budget:
  850 MB

Selected:
  base
  python_expert
  electronics_expert

Rejected:
  math_expert -> memory_budget

Inference:
  base + selected capacity

The default demo uses the deterministic control backend, so it can run without downloading model weights.


Architecture

flowchart LR
    P[Prompt] --> A[Analyzer]
    A --> R[Router]
    R --> B[Budget Planner]
    B --> C[Capacity Registry]
    C --> M[Runtime Manager]

    M --> G[Accelerator Resident]
    M --> H[Host / Cached]
    M --> D[Disk Available]

    G --> I[Selected Capacity]
    H --> I
    D --> I

    I --> N[Inference]
    N --> X[Metrics + Provenance]
Loading

Each stage is independently replaceable: routing policy, capacity planning, runtime residency, and inference backend can be evaluated without silently changing the others.

This makes routing quality, memory behavior, transfer cost, cache locality, and inference quality measurable as separate concerns.


What works today

SelectiveLLM currently includes:

  • multi-label prompt analysis,
  • static routing,
  • keyword routing,
  • random routing,
  • oracle routing,
  • local embedding routing,
  • threshold routing,
  • top-k routing,
  • hybrid routing,
  • versioned capacity registries,
  • dependency-aware planning,
  • configurable memory budgets,
  • runtime expert loading,
  • LRU-style capacity lifecycle management,
  • load / unload / transfer telemetry,
  • cache hit and miss tracking,
  • Transformers + PEFT inference,
  • deterministic offline controls,
  • CUDA → MPS → CPU device detection,
  • stage-separated latency metrics,
  • accelerator and host memory telemetry,
  • benchmark manifests and fingerprints,
  • reproducible raw results,
  • negative-result preservation,
  • dense causal-capacity research tooling.

The CLI and Python API use the same underlying engine.


Real-model evidence

A pinned real-model experiment was run using:

Base model: Qwen2.5-1.5B-Instruct
Backend:    Transformers + PEFT
Experts:    code / math / science LoRA adapters
Hardware:   Apple M4, 16 GB unified memory, MPS

The routing experiment produced:

Routing F1
──────────────
Random       0.278
Keyword      0.611
Semantic     0.830
Oracle       1.000

Dynamic semantic routing averaged approximately:

3125.7 MB MPS live allocation

versus:

3412.2 MB all-resident

for a measured difference of approximately:

286 MB

in live MPS tensor allocation.

This demonstrates that runtime capacity selection changes real model residency.

It does not demonstrate a reliable answer-quality advantage from the current routing policy.

That distinction is intentional and important.

Full evidence:

Real quality-memory tradeoff


Research status

SelectiveLLM intentionally separates demonstrated system behavior from open hypotheses.

Demonstrated

✅ Real adapter loading and unloading

✅ Runtime capacity residency changes

✅ Memory-budget-aware planning

✅ Dependency-aware cache behavior

✅ Measurable load and swap latency

✅ Semantic routing can outperform keyword and random routing on the current routing labels

✅ Reproducible deterministic and real-model experiment pipelines

Not yet demonstrated

⚠️ Reliable output-quality improvement from semantic adapter routing

⚠️ Robust domain specialization across the current expert pool

⚠️ Generalization to large public held-out workloads

⚠️ Learned calibrated capacity routing

❌ Stable semantic decomposition of dense model parameters

❌ Physical parameter paging based on prompt semantics

The dense Qwen2.5-1.5B and Qwen2.5-3B causal-localization experiments both failed their preregistered stable-specialization gates.

Those negative results are preserved rather than hidden or discarded.

See:


The broader research direction

The current adapter experiments are a proxy for a more general problem.

SelectiveLLM ultimately targets:

state / request
      ↓
adaptive compute controller
      ↓
┌─────────────────────────────┐
│ Which model?                │
│ Which expert?               │
│ Which adapter?              │
│ How much capacity?          │
│ What should stay resident?  │
│ What should be loaded?      │
│ What should be evicted?     │
│ When should we fall back?   │
└─────────────────────────────┘
      ↓
minimum sufficient computation

The desired optimization target is not routing accuracy by itself.

A future controller should optimize something closer to:

expected utility
    =
expected quality gain
    - memory cost
    - loading cost
    - latency cost
    - cache miss cost
    - uncertainty penalty

subject to runtime constraints such as:

memory <= budget
latency <= target
quality >= acceptable threshold

This is the direction planned for the next generation of the framework.


Current routing limitation

The current default analyzer is intentionally transparent and deterministic.

It uses:

  • lexical evidence,
  • feature hashing,
  • local similarity features.

It is not yet a trained semantic encoder.

This keeps the current benchmark reproducible and offline, but it is not intended to be the final routing architecture.

A future learned routing stack will compare:

random
vs
keyword
vs
feature-hash
vs
embedding encoder
vs
learned classifier
vs
oracle

with probability calibration, uncertainty measurement, and abstention.


Quickstart

Python 3.11 or newer is required.

git clone https://github.com/AltanCetinCelik/SelectiveLLM.git
cd SelectiveLLM

python3.11 -m venv .venv
source .venv/bin/activate

pip install -e .
selectivellm demo

Run a prompt:

selectivellm run \
  --prompt "Explain a MOSFET gate driver" \
  --verbose

Run the reproducible benchmark:

selectivellm benchmark \
  --report \
  --repetitions 5 \
  --seed 42

Python API

from selectivellm import SelectiveLLM

engine = SelectiveLLM.from_config("configs/default.yaml")
result = engine.generate("Use Python to simulate an RLC circuit and plot the transient response")

print(result.text)

print("Capabilities:")
print(result.profile.capabilities)

print("Selected capacity:")
print(result.routing.selected)

print("Rejected capacity:")
print(result.plan.rejected)

print("Memory:")
print(result.memory)

print("Metrics:")
print(result.metrics)

Real Transformers + PEFT

Install the optional Hugging Face backend:

pip install -e ".[hf]"

Then configure compatible local or Hugging Face model and adapter paths.

Example:

selectivellm run \
  --config configs/transformers_peft.example.yaml \
  --prompt "Explain a MOSFET gate driver"

The real backend supports:

  • causal language models,
  • named PEFT adapters,
  • adapter activation,
  • adapter eviction where supported,
  • accelerator timing synchronization,
  • real memory telemetry where PyTorch exposes it.

trust_remote_code defaults to false.

Model and adapter licenses remain separate from the SelectiveLLM Apache-2.0 license.


CLI

# Demo
selectivellm demo

# Route and generate
selectivellm run \
  --prompt "Explain a MOSFET gate driver" \
  --verbose

# Benchmark one router
selectivellm benchmark \
  --router semantic \
  --report

# Benchmark the method matrix
selectivellm benchmark \
  --report \
  --repetitions 5

# Inspect runtime
selectivellm inspect

# Profile one request
selectivellm profile \
  --prompt "Write a Python FastAPI endpoint"

# Inspect available capacity
selectivellm registry list
selectivellm registry inspect python_expert

Benchmark philosophy

SelectiveLLM treats routing, memory, latency, and quality as separate measurements.

The main latency decomposition is:

T_total =
    T_route
  + T_plan
  + T_load
  + T_inference
  + T_orchestration

Generated experiment directories include:

results/<run-id>/
├── config.yaml
├── environment.json
├── manifest.json
├── raw_results.jsonl
├── routing_decisions.jsonl
├── summary.json
├── summary.csv
├── report.md
├── routing_failures.md
└── plots/

The deterministic backend is a systems control for validating routing, planning, caching, telemetry, and reporting behavior without downloading model weights.

It is not treated as evidence of real model-quality improvement.

Read docs/benchmarking.md for the full methodology.


Research tracks

SelectiveLLM is best understood as two related but independent research tracks.

Track A — Adaptive Runtime

Near-term engineering and evaluation:

  • learned capacity routing,
  • calibrated probabilities,
  • uncertainty-aware fallback,
  • cost-aware planning,
  • larger expert pools,
  • multi-model routing,
  • cache and prefetch policies,
  • quality-memory-latency Pareto optimization.

This track already builds on working runtime infrastructure.

Track B — Selective Dense Compute

Higher-risk research:

  • causal capacity localization,
  • stable parameter masks,
  • hardware-aligned block selection,
  • contextual sparsity,
  • sparse execution,
  • eventual parameter paging.

Track A does not depend on Track B succeeding.


Roadmap

v0.2 — Adaptive Compute Router

Planned priorities:

  1. learned semantic capacity router,
  2. calibrated routing probabilities,
  3. uncertainty-aware abstention,
  4. safe fallback policies,
  5. quality-cost-aware planning,
  6. larger held-out evaluation sets,
  7. counterbalanced workloads,
  8. public benchmark tasks where appropriate,
  9. quality-memory-latency Pareto reporting,
  10. expanded model and expert capacity types.

Dense semantic parameter paging remains a separate experimental track and will not be claimed until physical savings and retained quality are demonstrated.

See docs/roadmap.md.


Related work

SelectiveLLM overlaps with several research areas:

  • Mixture-of-Experts,
  • model routing,
  • adapter routing,
  • semantic routing,
  • conditional computation,
  • contextual sparsity,
  • heterogeneous memory,
  • model offloading,
  • parameter-efficient fine-tuning,
  • knowledge localization.

The project does not claim that these individual ideas are new.

The intended research framing is:

Treat independently selectable model capacity as a budgeted runtime resource.

The framework evaluates that capacity across:

relevance
quality
memory
load latency
cache locality
runtime residency
uncertainty

The longer-term research question is whether this abstraction can eventually extend from adapters and models to useful internal model capacity.

See docs/related_work.md.


Documentation

Core

Experimental evidence


Development

pip install -e ".[dev]"

ruff check .
ruff format --check selectivellm tests experiments examples
mypy selectivellm

pytest \
  --cov=selectivellm \
  --cov-report=term-missing \
  --cov-fail-under=70

python -m build

The default test suite is network-free and does not download model weights.

Optional real-model integration tests require user-supplied compatible assets.


What SelectiveLLM is not

SelectiveLLM is not:

  • a document RAG framework,
  • a claim that dense models already contain clean page-ready semantic blocks,
  • a newly trained Mixture-of-Experts architecture,
  • a claim that a 70B dense model can currently run with 7B-equivalent memory,
  • a replacement for quantization,
  • a replacement for conventional offloading.

Those techniques can be complementary.

SelectiveLLM focuses specifically on request-conditioned capacity selection and runtime resource orchestration.


Citation

Use CITATION.cff when citing the project.

When citing experimental results, include the relevant benchmark fingerprint and manifest.

SelectiveLLM is licensed under Apache-2.0.

External models and adapters retain their original licenses.


Contributing

Reproducible positive, null, contradictory, and negative results are welcome.

Before contributing, see:

About

Adaptive model-capacity orchestration — dynamically select, load, cache, and budget adapters, experts, models, and future model blocks.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages