Skip to content

Integration: LayerInferenceEngine not registered as LLM provider — no API access #3104

Description

@mrveiss

Problem

LayerInferenceEngine (#1946) implements layer-by-layer inference for running 70B+ models on 8GB VRAM. However, it's not registered in the LLM provider registry and has no API endpoint. Users cannot trigger layer-by-layer inference through chat or any API call.

Also, the companion modules (meta_eviction #1952, kv_cache #1964, hf_quantizer #1954, attention_backend #1951) are all standalone utilities — none are wired into an end-to-end generation pipeline yet.

Discovered During

Batch 21 implementation of #1946, #1952, #1951, #1954.

Impact

High — Four inference optimization modules exist but form no usable pipeline. Each works in isolation (tested independently) but the integration that chains them (load config → detect quantization → select attention → layer-by-layer forward → meta evict → generate tokens) is not assembled.

Fix

  1. Register `LayerInferenceEngine` as a provider in the LLM provider registry
  2. Wire the full pipeline: HfQuantizer detection → attention backend selection → KV cache creation → layer loading with quantized param handling → forward pass → meta eviction → token generation
  3. Add an API endpoint or configuration flag to select layer-by-layer mode for batch workloads
  4. Integrate with existing `/api/chat` when model exceeds VRAM budget

Activity

  1. mrveiss commented on Jul 30, 2026

    @mrveiss
    OwnerAuthor

    Reopening the record rather than the issue: this was closed without the fix landing.

    The problem stated here — LayerInferenceEngine not registered as an LLM provider, no API access — is still present as of Dev_new_gui today. LayerInferenceAdapter was written (llm_shared/adapters/layer_inference_adapter.py:32) and exported (llm_shared/adapters/__init__.py:23), but _register_llm_adapters() in initialization/lifespan.py:1274-1312 still registers only Ollama, OpenAI, Anthropic and Groq. grep -rn "LayerInferenceAdapter" outside the definition and the export returns nothing, so it never reaches GET /api/adapters.

    Writing the adapter class satisfied step 1 of the fix; wiring it into startup did not happen, and the closure gate did not require evidence that the adapter appears in the API response.

    Superseded by #13035, which carries this plus the duplicate meta-eviction implementation, and is sequenced behind the correctness and streaming fixes in umbrella #13030 — registering an engine that cannot currently produce correct output would turn dormant defects into live ones.

    Leaving this closed and tracking the work on #13035 to avoid two issues for one fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions