Problem
LayerInferenceEngine (#1946) implements layer-by-layer inference for running 70B+ models on 8GB VRAM. However, it's not registered in the LLM provider registry and has no API endpoint. Users cannot trigger layer-by-layer inference through chat or any API call.
Also, the companion modules (meta_eviction #1952, kv_cache #1964, hf_quantizer #1954, attention_backend #1951) are all standalone utilities — none are wired into an end-to-end generation pipeline yet.
Discovered During
Batch 21 implementation of #1946, #1952, #1951, #1954.
Impact
High — Four inference optimization modules exist but form no usable pipeline. Each works in isolation (tested independently) but the integration that chains them (load config → detect quantization → select attention → layer-by-layer forward → meta evict → generate tokens) is not assembled.
Fix
- Register `LayerInferenceEngine` as a provider in the LLM provider registry
- Wire the full pipeline: HfQuantizer detection → attention backend selection → KV cache creation → layer loading with quantized param handling → forward pass → meta eviction → token generation
- Add an API endpoint or configuration flag to select layer-by-layer mode for batch workloads
- Integrate with existing `/api/chat` when model exceeds VRAM budget
Problem
LayerInferenceEngine (#1946) implements layer-by-layer inference for running 70B+ models on 8GB VRAM. However, it's not registered in the LLM provider registry and has no API endpoint. Users cannot trigger layer-by-layer inference through chat or any API call.
Also, the companion modules (meta_eviction #1952, kv_cache #1964, hf_quantizer #1954, attention_backend #1951) are all standalone utilities — none are wired into an end-to-end generation pipeline yet.
Discovered During
Batch 21 implementation of #1946, #1952, #1951, #1954.
Impact
High — Four inference optimization modules exist but form no usable pipeline. Each works in isolation (tested independently) but the integration that chains them (load config → detect quantization → select attention → layer-by-layer forward → meta evict → generate tokens) is not assembled.
Fix