Pure-Rust GGUF + Safetensors checkpoint/tensor substrate with optional format-specific helpers and model-analysis modules.
engram-parser ships a format-independent Checkpoint API over GGUF v3 and Safetensors: tensor inventory, metadata, and raw payload access without branching on container format. Format-specific surfaces remain for callers that need them: GGUF wire types and packed K-quant dequant (Q8_0 / Q5_K / Q6_K / IQ3_M block layout), an optional mmap reader, and Safetensors header/manifest/MoE discovery — the latter behind the off-by-default safetensors Cargo feature, which uses the upstream safetensors crate for canonical validation. MoE per-expert raw-weight extraction is a GGUF-specific model-analysis specialization under engram_parser::analysis::moe.
Safetensors status: shipped behind
--features safetensors. Default builds are GGUF-only. Raw payload reads are bounded per-tensor; enable--features safetensors,mmapfor borrowed mmap-backed access. No Hugging Faceconfig.jsonpolicy.mmap / K-quant status: shipped. Default
load_ggufstill usesfs::read. Enable--features mmapforload_gguf_mmap(memmap2). Packed dequant for Q8_0 / Q5_K / Q6_K / IQ3_M is on the default crate. CUDA host-register is out of scope.
For a published 0.3 release, choose one of the following dependency configurations. Available versions are listed on crates.io. GGUF only, with no dependencies enabled by default:
[dependencies]
engram-parser = "0.3"with Safetensors support:
[dependencies]
engram-parser = { version = "0.3", features = ["safetensors"] }or with Safetensors plus borrowed mmap payloads:
[dependencies]
engram-parser = { version = "0.3", features = ["safetensors", "mmap"] }Equivalently, for a published release: cargo add engram-parser@0.3 (add
--features safetensors / mmap as needed).
For a development snapshot, including before the first registry release, pin an exact Git commit. For example:
[dependencies]
engram-parser = { git = "https://github.com/rmems/engram-parser", rev = "485f0445a93c1ec06e8142a2d602c33030d753d7" }The same feature names apply to Git dependencies. See Releases for the publication process.
| Feature | Adds | Enables |
|---|---|---|
| (default) | — | GGUF parsing, Checkpoint API, MoE extraction, K-quant dequant |
mmap |
optional memmap2 |
load_gguf_mmap, open_checkpoint_mmap, page-aligned tensor slices |
safetensors |
optional upstream safetensors crate |
Safetensors header/manifest/discovery, raw payload access |
Format-agnostic checkpoint inventory — the shared entry point, complete and runnable:
use engram_parser::{Checkpoint, open_checkpoint};
fn main() -> Result<(), engram_parser::ParserError> {
// Works for a .gguf file, a .safetensors file, an HF shard index,
// or a checkpoint directory (Safetensors input needs the feature).
let ckpt = open_checkpoint("./model.gguf")?;
println!("format = {:?}", ckpt.format());
for t in ckpt.tensors() {
println!("{} {} {:?} {}B", t.name, t.dtype, t.shape.dims(), t.byte_len);
}
// Read raw, undecoded payload bytes for a tensor discovered above.
if let Some(t) = ckpt.tensors().first() {
let raw = ckpt.tensor_bytes(&t.name)?; // Cow<[u8]>
println!("{}: {} bytes", t.name, raw.len());
}
Ok(())
}GGUF-specific access (when you need GGUF semantics — metadata helpers, ggml_type, dequant):
use engram_parser::{ggml_type_label, load_gguf};
fn main() -> Result<(), engram_parser::ParserError> {
let layout = load_gguf("./model.gguf")?;
println!("architecture = {}", layout.metadata.architecture());
for (name, tensor) in &layout.tensors {
println!("{} {:?} ({})", name, tensor.dims, ggml_type_label(tensor.ggml_type));
}
Ok(())
}MoE expert analysis — a GGUF-specific specialization under engram_parser::analysis::moe:
use engram_parser::analysis::moe::{extract_expert, list_experts};
use engram_parser::load_gguf;
fn main() -> Result<(), engram_parser::ParserError> {
let layout = load_gguf("./moe-model.gguf")?;
for (block, expert) in list_experts(&layout) {
let weights = extract_expert(&layout, block, expert)?;
if let Some(gate) = &weights.gate {
println!(
"blk.{block}.expert{expert}.gate: dims={:?} dtype={:?} bytes={}",
gate.dims,
gate.dtype,
gate.bytes.len()
);
}
}
Ok(())
}Legacy MoE imports still work. engram_parser::moe::{extract_expert, list_experts, MoeExpertWeights, RawTensor} and the crate-root engram_parser::{extract_expert, list_experts, MoeExpertWeights, RawTensor} paths re-export the same implementation and keep compiling in v0.3; engram_parser::analysis::moe is the canonical documented path going forward.
Safetensors (requires --features safetensors):
use engram_parser::safetensors::{inspect_safetensors_checkpoint, open_safetensors_checkpoint};
fn main() -> Result<(), engram_parser::ParserError> {
let manifest = inspect_safetensors_checkpoint("./model.safetensors")?;
println!(
"kind={} tensors={}",
manifest.checkpoint.input_kind, manifest.checkpoint.tensor_count
);
for tensor in &manifest.tensors {
println!("{} {:?} {}", tensor.name, tensor.shape, tensor.dtype);
}
let checkpoint = open_safetensors_checkpoint("./model.safetensors")?;
// "a.weight" is a placeholder — use a name from the manifest above.
let raw: Vec<u8> = checkpoint.tensor_bytes("a.weight")?; // bounded read
println!("payload bytes = {}", raw.len());
Ok(())
}With --features safetensors,mmap, open_safetensors_checkpoint_mmap returns borrowed &[u8] slices instead of owned copies.
- Parses GGUF v3 magic, header, KV metadata, and tensor directory into an in-memory
GgufLayout. - Applies a documented
ParseLimitsbudget (KV/tensor counts, string sizes, array work, tensor rank, metadata bytes) before allocation or loops proportional to file-declared values. Defaults are generous; trusted callers can override viaload_gguf_with_limits/parse_bytes_with_limitswithout weakening the default path. - Enumerates MoE experts discovered in a checkpoint (model analysis under
engram_parser::analysis::moe). - Extracts the raw byte buffers for one expert's
gate,up, anddownprojections. - Supports stacked (
blk.{B}.ffn_{role}_exps.weight) and per-expert (blk.{B}.ffn_{role}.{E}.weight) conventions. - Models packed GGUF dtype sizes without pulling in a numerical runtime.
- Packed dequant to
f32for Q8_0, Q5_K, Q6_K, and the internal IQ3_M block layout (GGML_TYPE_IQ3_M_BLOCK = 0x4949334D). Existingdequantize_f16is unchanged. - Optional
mmapfeature (memmap2):load_gguf_mmapmaps the file instead offs::readinto aVec<u8>. Default builds still read the whole file; no dependencies are enabled by default.
The off-by-default safetensors feature owns reusable support for:
- Safetensors header deserialization (canonical validation via the upstream
safetensorscrate); - deterministic tensor manifests;
- single-file, Hugging Face shard-index, and directory layouts;
- checkpoint-relative shard resolution with path-escape rejection;
- unique tensor ownership and missing-shard diagnostics;
- tensor name/dtype/shape/offset/shard metadata;
- MoE router/expert candidate discovery and grouping;
- raw tensor payload access by name (
open_safetensors_checkpoint, bounded per-tensor reads); - mmap-backed borrowed payload slices with
--features safetensors,mmap(open_safetensors_checkpoint_mmap); - metadata-only layout-family inference.
The safetensors feature adds the upstream safetensors crate as an optional dependency; upstream types (SafeTensors, TensorView, Dtype) never appear in the public API. Engram keeps checkpoint discovery, shard/path policy, duplicate-key detection, stricter contiguous-range validation, and engram-owned dtype/shape/shard types.
- No neural-network math: no
matmul, forward pass, routing execution, or softmax. - No CUDA, GPU, SIMD, or runtime inference engine. mmap exposes page-aligned tensor byte ranges so a later GPU layer can host-register; this crate does not.
- No tokenization or SNN dynamics.
- No full model-family routing policy.
- No Safetensors payload execution/matmul, and no Hugging Face
config.jsoninterpretation: payload access is byte-level only, with no runtime or model-policy concerns pulled in.
Layout-aware parsing (packed byte sizes; plus the K-quant dequant helpers below) supports:
F32,F16,BF16,F64;I8–I64;Q4_0,Q4_1,Q5_0,Q5_1,Q8_0,Q8_1;- K-quants
Q2_K,Q3_K,Q4_K,Q5_K,Q6_K,Q8_K; - IQ packed layouts including
IQ2_XXS,IQ2_XS,IQ2_S,IQ3_XXS,IQ3_S,IQ1_S,IQ1_M,IQ4_NL,IQ4_XS; - Internal
IQ3_M_BLOCK(111 bytes / 256 values) — not wire type 31.
Remaining codes use DType::Other(u32). Unknown packed layouts fail closed when element count cannot be converted to a modeled byte length.
Numeric helpers: dequantize_f16, dequantize_q8_0, dequantize_q5_k, dequantize_q6_k, dequantize_iq3_m. The parser's primary contract is still layout/raw bytes; these unpackers are optional CPU helpers, not a GGML compute path.
GGUF stores each tensor dtype as a ggml_type integer. engram-parser maps those codes to labels and packed byte sizes so payload and MoE slices remain in range. Packed dequant is implemented for Q8_0, Q5_K, Q6_K, and the internal IQ3_M block id — not a GGML runtime.
Historical wire type 31 is Q4_0_4_4, not the Hugging Face preset name "IQ3_M". HuggingFace IQ3_M is a mixed-quant preset; the 111-byte block decoder uses internal GGML_TYPE_IQ3_M_BLOCK (0x4949334D). Wire 31 still has no packed byte-size model (DType::Other(31)), so such a checkpoint is labeled correctly but rejected during layout parsing.
Current GGUF surface includes:
load_gguf,parse_bytes(and*_with_limitsfor an explicitParseLimitspolicy);#[cfg(feature = "mmap")] load_gguf_mmap/load_gguf_mmap_with_limits→GgufLayoutMmap(page-aligned tensor slices viatensor_page_aligned_bytes);GgufLayout,GgufMetadata,Tensor,DType,ParseLimits;dequantize_f16(onTensor),dequantize_q8_0,dequantize_q5_k,dequantize_q6_k,dequantize_iq3_m;extract_expert,list_experts(canonical:engram_parser::analysis::moe);MoeExpertWeights,RawTensor(also atengram_parser::analysis::moe);ParserError,ParseLimitKind,HostSizeField,Result;- public
GGML_TYPE_*/GGUF_VALUE_TYPE_*constants and theggml_type_labellabel function.
Safetensors surface (--features safetensors) includes:
inspect_safetensors_checkpoint,write_safetensors_manifest;open_safetensors_checkpoint→SafetensorsCheckpoint(tensor,tensor_bytes,resolve_tensor_bytes,manifest);#[cfg(feature = "mmap")] open_safetensors_checkpoint_mmap→SafetensorsCheckpointMmap(borrowed&[u8]payload slices);SafetensorsManifest,SafetensorsCheckpointSource,SafetensorsTensorRecord;classify_tensor,discover_candidates;SafetensorsCandidateSummary,SafetensorsRouterCandidate,SafetensorsExpertGroup;dtype_size_bytes;ParserError::DuplicateTensorOwnershipandParserError::MissingShardfor shard-index diagnostics.
engram_parser::checkpoint (re-exported at the crate root, default features) is one contract over GGUF and Safetensors, so code can inventory tensors, read metadata, and fetch raw payloads without branching on format.
- Trait:
Checkpoint(object safe,Send + Sync):source,format,tensors(sorted by name),tensor,metadata,tensor_bytes. - Detection:
open_checkpointtreats a directory as Safetensors, a file withGGUFmagic as GGUF (any extension), and*.safetensors/*.safetensors.index.jsonas Safetensors. Anything else returnsUnsupportedFormat. Safetensors input without the feature returnsParserError::FeatureDisabled; no other backend is substituted. Withmmap,open_checkpoint_mmapmaps GGUF files and Safetensors shards. - Shape order:
TensorShape::dims()is outermost-first for both formats. GGUF dims (innermost-first on disk) are reversed;native_dims()/native_order()recover the on-disk order. - Dtypes: scalar types normalize to shared
TensorDTypevariants. GGML quantized layouts stayTensorDType::GgufPacked { ggml_type }with the exact wire code, andTensorInfo::native_dtypekeeps the source label (Q4_K,BF16, …). - Payloads: bytes are returned exactly as stored, with length equal to
byte_len. They are borrowed for GGUF and all mmap backends, and owned (bounded per-tensor read) for plain Safetensors. - Location:
sourceis relative toCheckpointSource::root,data_offsetis relative to the tensor-data section, andfile_offsetis absolute. - Metadata: GGUF scalars map to
MetadataValue::{String, UInt, Int, F32, F64}; GGUF arrays are not captured (same asGgufMetadata), andgeneral.alignmentis a layout field (layout().alignment), not a metadata entry. Safetensors metadata is the manifest's flattened string map.
| Backend | Wraps | Feature | Payload |
|---|---|---|---|
GgufBackend |
GgufLayout |
default | borrowed |
GgufMmapBackend |
GgufLayoutMmap |
mmap |
borrowed |
SafetensorsBackend |
safetensors::SafetensorsCheckpoint |
safetensors |
owned |
SafetensorsMmapBackend |
safetensors::SafetensorsCheckpointMmap |
safetensors + mmap |
borrowed |
Migration. Existing APIs are unchanged. To go from load_gguf to the shared contract, use GgufBackend::open(path) (or from_layout(layout) / from_bytes). layout() / into_layout() still reach list_experts, extract_expert, find_tensors_with_suffix, and dequant — now canonically at engram_parser::analysis::moe, with engram_parser::moe and the crate-root symbols kept as re-export shims. AnyCheckpoint::as_gguf(), as_safetensors(), and the mmap variants return the format-specific handles. The Safetensors manifest and candidate-discovery APIs are unchanged.
Format-agnostic inventory example: cargo run --example inspect_checkpoint -- <path> (add --features safetensors for Safetensors, --mmap with the mmap feature).
This crate has no dependencies enabled by default. Enable mmap for memmap2. Enable safetensors for header/manifest/discovery plus raw payload access (upstream safetensors crate); combine with mmap for borrowed payload slices.
cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo build
cargo test
cargo clippy --all-targets --all-features -- -D warnings
cargo test --features mmap
cargo test --features safetensors
cargo build --all-features
cargo test --all-features
# Coverage (requires cargo-llvm-cov)
cargo llvm-cov --all-targets --all-features --locked --lcov --output-path lcov.infoload_gguf reads and retains the complete file. For multi-GB checkpoints use cargo test --features mmap / load_gguf_mmap. CI mmap tests include packed Q8_0/Q5_K/Q6_K/IQ3_M dequant from the mapping and a sparse 2 GiB file that is mapped without fs::read. Real on-disk GGUF pilots still require local files (ENGRAM_GGUF) and are #[ignore]. CUDA host-register is still out of scope.
ENGRAM_GGUF=~/.models/gguf/.../model.gguf ENGRAM_EXPECT_MOE=1 \
cargo test --test real_gguf -- --ignored --nocapture
cargo run --example inspect_gguf -- ~/.models/gguf/.../model.ggufFor directory scans, ENGRAM_GGUF_MAX, and MoE extraction sample counts, see the T1 pilot guidance in REVIEW.md.
Opt-in large GGUF and multi-shard Safetensors mmap validation, fixture inputs, memory bounds and reproducible release reports: checkpoint smoke guide.
See REVIEW.md for quality gates.
- GitHub Actions:
.github/workflows/ci.ymlruns default and all-feature tests separately on Linux, Windows, and macOS, plus a Rust 1.97.1 MSRV job. - Coverage: Linux
cargo llvm-covproduces a downloadable LCOV artifact and uploads to Codecov when the repository token is available. - Quality: Qlty checks workflow syntax and reports maintainability; Rust formatting and Clippy remain the authoritative Rust checks. Coverage is reported only through Codecov.
- Security:
.github/workflows/security.yml
Cursor Cloud Agents use the separate .cursor/Dockerfile selected by
.cursor/environment.json to prepare their coding environment. It is not a
published crate artifact or a Docker image build in GitHub Actions.
Version 0.3.0 is the first crates.io publication target; 0.2.0 was an
unpublished development version. See RELEASE.md for the
mandatory package-inspection, rustdoc, packaging, and publish-dry-run gate.
Tags and GitHub Releases are created only after that gate and publication
have succeeded.
MSRV: Rust 1.97.1.
- Declared through
rust-versioninCargo.toml. - Tested in CI alongside the pinned Rust 1.99.0 build toolchain.
- MSRV bumps require justification and are treated as compatibility-significant changes.
See #14.
Not required for installing or using the crate — retained for attribution.
engram-parser was extracted from an internal research codebase as the
canonical reusable implementation of its GGUF and Safetensors parsing
surfaces; the extraction is one-way and this crate has no upstream research
dependencies. Wire-type labels and the Safetensors manifest format were
validated against the original reference implementation.
Historical issue trackers for the modularization work:
- GGUF modularization: engram-parser#7
- Safetensors promotion: engram-parser#10, README consistency: #61
- mmap + K-quant tracking: engram-parser#45
This repository intentionally keeps documentation in version-controlled files such as README.md and REVIEW.md rather than a GitHub Wiki.
Licensed under either of:
- Apache License, Version 2.0 (LICENSE-APACHE-2.0); or
- MIT license (LICENSE-MIT).