Skip to content

tech-debt(optimization): consolidate architecture-family, dtype and compression taxonomies + duplicated accelerate loader #13049

Description

@mrveiss

Discovered during

Audit for umbrella #13030 (lean-hardware inference). These are canonical-source problems in the same
package, but they are refactors rather than defects, so they are filed separately from the #13030
defect children.

C1 — Two divergent architecture-family taxonomies

The repo answers "what architecture is this model?" in two incompatible ways.

Canonical (enum-backed).
openvino_dispatch.py:40-46 defines
ArchitectureFamily(str, Enum) — TRANSFORMER, SSM, LINEAR_ATTENTION, HYBRID — introduced by
GH#7347/#7352, keyed on a model's architecture_family field, with an explicit
UnsupportedArchitectureError when a family has no inference path.

Ad-hoc (inline set, silent fallback).
layer_inference.py:522-537
_layer_prefix_for_arch() keys on HuggingFace model_type against a hardcoded literal set:

gpt2_style = {"gpt2", "gptj", "gpt_neo", "gpt_neox"}
if model_type in gpt2_style:
    return "transformer.h."
# LLaMA, Mistral, Falcon, Qwen, Gemma, Phi, etc.
return "model.layers."

The catch-all returns a LLaMA-style prefix for every unrecognised architecture — including SSM and
hybrid families the enum knows about explicitly. There is no error path; an unsupported architecture
silently produces layer names that do not exist in the checkpoint, which surfaces later as the
KeyError that _run_layer_loop
(:507-509) swallows with
logger.warning("Skipping layer …") — so the model quietly runs with layers missing.

Consolidate: one canonical architecture-family enum, shared by the NPU dispatch and the layer
engine, with layer-name prefixes as data hanging off it and an explicit unsupported-family error
rather than a silent catch-all.

C2 — Stringly-typed dtype, with two vocabularies inside one package

llm_shared/optimization/ uses raw strings for element dtype, and sibling modules disagree on the
spelling:

The clash is already being papered over rather than fixed: kv_cache.py's byte-size map carries
both spellings as separate keys —
"fp16": 2 at line 59 and
"float16": 2 at line 62. That is a
workaround for the missing canonical type, and it only covers the spellings someone happened to hit.

Consolidate: one Dtype enum in the package (or reuse an existing shared one if present), with a
single byte-width mapping and explicit conversion at the torch boundary.

C3 — compression validated in one place, bypassed in the other

layer_inference.py:50 defines
_VALID_COMPRESSIONS = frozenset({"4bit", "8bit", "none"}) and enforces it in
LayerInferenceConfig.__post_init__
(:79-80).

pipeline.py:69 declares its own
compression: str = "none" with no validation, then passes it straight through to
LayerInferenceConfig at :273. The
validation still fires, but the error surfaces one layer away from the field the caller actually set,
and PipelineConfig advertises a wider contract than it accepts.

Consolidate: a Compression enum shared by both configs, so the type carries the constraint and
neither dataclass restates it.

C4 — Duplicated lazy-import helper

def _import_accelerate() is defined twice, character-for-character in intent:

Same package, same optional dependency, two copies with slightly different error text. This is the
same class of problem #12714 fixed for the nine forked _get_torch() loaders — the accelerate loader
was missed.

Consolidate: one lazy-import helper alongside the existing shared llm_shared.torch_loader.

Ordering

These are refactors over code that is currently broken (see #13030's children). Land them after
the #13030 P0 defects, so the enum work is done against a module whose behaviour is verifiable.
C4 is independent and can go any time.

Related

Activity

  1. added a commit that references this issue on Jul 30, 2026
  2. added this to the v0.10.0 milestone on Sep 12, 2026
  3. 23 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions