Problem
Two wiring/canonicality defects that keep the optimization stack unreachable and duplicated.
D6 — the adapter is never registered. _register_llm_adapters()
(lifespan.py:1274-1312) registers only
Ollama, OpenAI, Anthropic and Groq. LayerInferenceAdapter is defined
(layer_inference_adapter.py:32)
and exported
(adapters/init.py:23) but never passed to
registry.register(). grep -rn "LayerInferenceAdapter" outside its own definition and the export
returns nothing. It is therefore absent from GET /api/adapters and unreachable at runtime.
Note: #3104 ("LayerInferenceEngine not registered as LLM provider — no API access") was closed
2026-04-01 on exactly this problem. The adapter class was written; the registration was not. This is
a premature closure — the closure gate should have required evidence that the adapter appears in
GET /api/adapters.
D7 — duplicate meta-eviction. meta_eviction.py
is the canonical module (public evict_layer_to_meta, clean_memory,
MetaDeviceEvictionManager, quantized-layer handling, accelerate integration). It has zero
production callers — grep -rn "meta_eviction\|MetaDeviceEviction" hits only
meta_eviction_test.py. Meanwhile
layer_inference.py:578-601
carries its own private _move_to_meta / _set_buffer reimplementation, which evict_layer()
(:285-304) calls instead.
The private copy is the weaker one: it has no quantized-layer path, no accelerate per-parameter
handling, and no eviction accounting.
Fix
- Replace
layer_inference._move_to_meta with a call to meta_eviction.evict_layer_to_meta; delete
the private duplicate and its _set_buffer helper.
- Register
LayerInferenceAdapter in _register_llm_adapters(), behind a config flag, defaulting
off.
Ordering
Registration is gated on the correctness and streaming children of this umbrella. Exposing an
engine that cannot produce correct output, and that loads the whole checkpoint per token, would turn
dormant defects into live ones. Land this one last.
Acceptance
GET /api/adapters lists the layer-inference adapter when the flag is enabled.
- Exactly one meta-eviction implementation exists in the repo.
Part of #13030.
Problem
Two wiring/canonicality defects that keep the optimization stack unreachable and duplicated.
D6 — the adapter is never registered.
_register_llm_adapters()(lifespan.py:1274-1312) registers only
Ollama, OpenAI, Anthropic and Groq.
LayerInferenceAdapteris defined(layer_inference_adapter.py:32)
and exported
(adapters/init.py:23) but never passed to
registry.register().grep -rn "LayerInferenceAdapter"outside its own definition and the exportreturns nothing. It is therefore absent from
GET /api/adaptersand unreachable at runtime.Note: #3104 ("LayerInferenceEngine not registered as LLM provider — no API access") was closed
2026-04-01 on exactly this problem. The adapter class was written; the registration was not. This is
a premature closure — the closure gate should have required evidence that the adapter appears in
GET /api/adapters.D7 — duplicate meta-eviction. meta_eviction.py
is the canonical module (public
evict_layer_to_meta,clean_memory,MetaDeviceEvictionManager, quantized-layer handling,accelerateintegration). It has zeroproduction callers —
grep -rn "meta_eviction\|MetaDeviceEviction"hits onlymeta_eviction_test.py. Meanwhilelayer_inference.py:578-601
carries its own private
_move_to_meta/_set_bufferreimplementation, whichevict_layer()(:285-304) calls instead.
The private copy is the weaker one: it has no quantized-layer path, no
accelerateper-parameterhandling, and no eviction accounting.
Fix
layer_inference._move_to_metawith a call tometa_eviction.evict_layer_to_meta; deletethe private duplicate and its
_set_bufferhelper.LayerInferenceAdapterin_register_llm_adapters(), behind a config flag, defaultingoff.
Ordering
Registration is gated on the correctness and streaming children of this umbrella. Exposing an
engine that cannot produce correct output, and that loads the whole checkpoint per token, would turn
dormant defects into live ones. Land this one last.
Acceptance
GET /api/adapterslists the layer-inference adapter when the flag is enabled.Part of #13030.