Skip to content

tech-debt(optimization): LayerInferenceAdapter never registered + duplicate meta-eviction implementation (#13030) #13035

Description

@mrveiss

Problem

Two wiring/canonicality defects that keep the optimization stack unreachable and duplicated.

D6 — the adapter is never registered. _register_llm_adapters()
(lifespan.py:1274-1312) registers only
Ollama, OpenAI, Anthropic and Groq. LayerInferenceAdapter is defined
(layer_inference_adapter.py:32)
and exported
(adapters/init.py:23) but never passed to
registry.register(). grep -rn "LayerInferenceAdapter" outside its own definition and the export
returns nothing. It is therefore absent from GET /api/adapters and unreachable at runtime.

Note: #3104 ("LayerInferenceEngine not registered as LLM provider — no API access") was closed
2026-04-01 on exactly this problem. The adapter class was written; the registration was not. This is
a premature closure — the closure gate should have required evidence that the adapter appears in
GET /api/adapters.

D7 — duplicate meta-eviction. meta_eviction.py
is the canonical module (public evict_layer_to_meta, clean_memory,
MetaDeviceEvictionManager, quantized-layer handling, accelerate integration). It has zero
production callers — grep -rn "meta_eviction\|MetaDeviceEviction" hits only
meta_eviction_test.py. Meanwhile
layer_inference.py:578-601
carries its own private _move_to_meta / _set_buffer reimplementation, which evict_layer()
(:285-304) calls instead.

The private copy is the weaker one: it has no quantized-layer path, no accelerate per-parameter
handling, and no eviction accounting.

Fix

  1. Replace layer_inference._move_to_meta with a call to meta_eviction.evict_layer_to_meta; delete
    the private duplicate and its _set_buffer helper.
  2. Register LayerInferenceAdapter in _register_llm_adapters(), behind a config flag, defaulting
    off.

Ordering

Registration is gated on the correctness and streaming children of this umbrella. Exposing an
engine that cannot produce correct output, and that loads the whole checkpoint per token, would turn
dormant defects into live ones. Land this one last.

Acceptance

  • GET /api/adapters lists the layer-inference adapter when the flag is enabled.
  • Exactly one meta-eviction implementation exists in the repo.

Part of #13030.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions