Discovered during
Audit for umbrella #13030 (lean-hardware inference). These are canonical-source problems in the same
package, but they are refactors rather than defects, so they are filed separately from the #13030
defect children.
C1 — Two divergent architecture-family taxonomies
The repo answers "what architecture is this model?" in two incompatible ways.
Canonical (enum-backed).
openvino_dispatch.py:40-46 defines
ArchitectureFamily(str, Enum) — TRANSFORMER, SSM, LINEAR_ATTENTION, HYBRID — introduced by
GH#7347/#7352, keyed on a model's architecture_family field, with an explicit
UnsupportedArchitectureError when a family has no inference path.
Ad-hoc (inline set, silent fallback).
layer_inference.py:522-537
_layer_prefix_for_arch() keys on HuggingFace model_type against a hardcoded literal set:
gpt2_style = {"gpt2", "gptj", "gpt_neo", "gpt_neox"}
if model_type in gpt2_style:
return "transformer.h."
# LLaMA, Mistral, Falcon, Qwen, Gemma, Phi, etc.
return "model.layers."
The catch-all returns a LLaMA-style prefix for every unrecognised architecture — including SSM and
hybrid families the enum knows about explicitly. There is no error path; an unsupported architecture
silently produces layer names that do not exist in the checkpoint, which surfaces later as the
KeyError that _run_layer_loop
(:507-509) swallows with
logger.warning("Skipping layer …") — so the model quietly runs with layers missing.
Consolidate: one canonical architecture-family enum, shared by the NPU dispatch and the layer
engine, with layer-name prefixes as data hanging off it and an explicit unsupported-family error
rather than a silent catch-all.
C2 — Stringly-typed dtype, with two vocabularies inside one package
llm_shared/optimization/ uses raw strings for element dtype, and sibling modules disagree on the
spelling:
The clash is already being papered over rather than fixed: kv_cache.py's byte-size map carries
both spellings as separate keys —
"fp16": 2 at line 59 and
"float16": 2 at line 62. That is a
workaround for the missing canonical type, and it only covers the spellings someone happened to hit.
Consolidate: one Dtype enum in the package (or reuse an existing shared one if present), with a
single byte-width mapping and explicit conversion at the torch boundary.
C3 — compression validated in one place, bypassed in the other
layer_inference.py:50 defines
_VALID_COMPRESSIONS = frozenset({"4bit", "8bit", "none"}) and enforces it in
LayerInferenceConfig.__post_init__
(:79-80).
pipeline.py:69 declares its own
compression: str = "none" with no validation, then passes it straight through to
LayerInferenceConfig at :273. The
validation still fires, but the error surfaces one layer away from the field the caller actually set,
and PipelineConfig advertises a wider contract than it accepts.
Consolidate: a Compression enum shared by both configs, so the type carries the constraint and
neither dataclass restates it.
C4 — Duplicated lazy-import helper
def _import_accelerate() is defined twice, character-for-character in intent:
Same package, same optional dependency, two copies with slightly different error text. This is the
same class of problem #12714 fixed for the nine forked _get_torch() loaders — the accelerate loader
was missed.
Consolidate: one lazy-import helper alongside the existing shared llm_shared.torch_loader.
Ordering
These are refactors over code that is currently broken (see #13030's children). Land them after
the #13030 P0 defects, so the enum work is done against a module whose behaviour is verifiable.
C4 is independent and can go any time.
Related
Discovered during
Audit for umbrella #13030 (lean-hardware inference). These are canonical-source problems in the same
package, but they are refactors rather than defects, so they are filed separately from the #13030
defect children.
C1 — Two divergent architecture-family taxonomies
The repo answers "what architecture is this model?" in two incompatible ways.
Canonical (enum-backed).
openvino_dispatch.py:40-46 defines
ArchitectureFamily(str, Enum)—TRANSFORMER,SSM,LINEAR_ATTENTION,HYBRID— introduced byGH#7347/#7352, keyed on a model's
architecture_familyfield, with an explicitUnsupportedArchitectureErrorwhen a family has no inference path.Ad-hoc (inline set, silent fallback).
layer_inference.py:522-537
_layer_prefix_for_arch()keys on HuggingFacemodel_typeagainst a hardcoded literal set:The catch-all returns a LLaMA-style prefix for every unrecognised architecture — including SSM and
hybrid families the enum knows about explicitly. There is no error path; an unsupported architecture
silently produces layer names that do not exist in the checkpoint, which surfaces later as the
KeyErrorthat_run_layer_loop(:507-509) swallows with
logger.warning("Skipping layer …")— so the model quietly runs with layers missing.Consolidate: one canonical architecture-family enum, shared by the NPU dispatch and the layer
engine, with layer-name prefixes as data hanging off it and an explicit unsupported-family error
rather than a silent catch-all.
C2 — Stringly-typed dtype, with two vocabularies inside one package
llm_shared/optimization/uses raw strings for element dtype, and sibling modules disagree on thespelling:
dtype: str = "fp16", documented at :84 as"fp16","bf16","fp32"torch_dtype: str | None = "float16"torch_dtype: str | None = NoneThe clash is already being papered over rather than fixed:
kv_cache.py's byte-size map carriesboth spellings as separate keys —
"fp16": 2at line 59 and"float16": 2at line 62. That is aworkaround for the missing canonical type, and it only covers the spellings someone happened to hit.
Consolidate: one
Dtypeenum in the package (or reuse an existing shared one if present), with asingle byte-width mapping and explicit conversion at the torch boundary.
C3 —
compressionvalidated in one place, bypassed in the otherlayer_inference.py:50 defines
_VALID_COMPRESSIONS = frozenset({"4bit", "8bit", "none"})and enforces it inLayerInferenceConfig.__post_init__(:79-80).
pipeline.py:69 declares its own
compression: str = "none"with no validation, then passes it straight through toLayerInferenceConfigat :273. Thevalidation still fires, but the error surfaces one layer away from the field the caller actually set,
and
PipelineConfigadvertises a wider contract than it accepts.Consolidate: a
Compressionenum shared by both configs, so the type carries the constraint andneither dataclass restates it.
C4 — Duplicated lazy-import helper
def _import_accelerate()is defined twice, character-for-character in intent:Same package, same optional dependency, two copies with slightly different error text. This is the
same class of problem #12714 fixed for the nine forked
_get_torch()loaders — the accelerate loaderwas missed.
Consolidate: one lazy-import helper alongside the existing shared
llm_shared.torch_loader.Ordering
These are refactors over code that is currently broken (see #13030's children). Land them after
the #13030 P0 defects, so the enum work is done against a module whose behaviour is verifiable.
C4 is independent and can go any time.
Related
_get_torch()unification that established the pattern