Skip to content

parser: adopt/enhance as canonical GGUF parser by extracting from corinth-canal #7

Description

@rmems

Summary

Extract and enhance reusable GGUF deserialization + MoE per-expert extraction logic from the rmems/corinth-canal experimental reference implementation into Limen-Neural/engram-parser (the canonical zero-dependency parser crate under the Limen-Neural org for modular reusable libraries).

This continues the modularization program (see LIM-9) and aligns with the existing positioning of engram-parser as the GGUF/parse + raw expert extractor base (see README Scope/Boundaries and open #5 for traits).

Sibling tracking issues (created together):

Context

  • engram-parser was initially refactored from logic in corinth-canal (see commit 8dd2af2), deliberately omitting routing/CUDA to stay zero-dep and focused.
  • corinth-canal (rmems personal/experimental ref repo) still duplicates a full GGUF parser + mmap + dequants + adapters in src/moe/{checkpoint.rs,ggml.rs,...}.
  • cortex-tensor (sibling Limen-Neural reusable lib) was also extracted from corinth-canal and has open issues Define MoE extraction traits inside this crate #5 (boundary) + parser: adopt/enhance as canonical GGUF parser by extracting from corinth-canal #7 (traits, notes reusability with engram-parser).
  • Goal: extract useful code from the experimental corinth-canal into Limen-Neural org's focused library crates. corinth-canal remains the end-to-end research vehicle.
  • No current GitHub/Linear/beads issue explicitly tracked this parser extraction + coordination (until these).

Related

Goal

Make engram-parser the single source of truth for GGUF v3 layout parsing + MoE expert weight ripping (raw bytes only). Enable future adoption by cortex-tensor (and potentially corinth-canal's higher layers).

Non-goals

  • Do not move routing math, model adapters, full dequant math, Safetensors, CUDA/GPU registration, or SNN orchestration.
  • Do not change corinth-canal's experimental runtime or validation behavior in this issue (that belongs to the companion migration issue).
  • Keep engram-parser zero-dependency.

Acceptance criteria

  • Additional dtypes from corinth-canal's ggml (e.g. IQ3_M=31 as Other or explicit opaque; verify byte_len handling).
  • Port useful pure helpers where they fit zero-dep contract (e.g. ggml_type_label or equivalent; full value type consts for metadata if missing).
  • MoE extraction traits defined (or advanced) inside the crate (ties directly to open Define MoE extraction traits inside this crate #5).
  • Tests cover additional cases/synthetic patterns inspired by corinth-canal.
  • README updated with explicit note on extraction source + cross-links to cortex-tensor coordination issue and (future) corinth migration issue; ecosystem table remains accurate.
  • Crate remains zero-dep; cargo test --all-features + clippy/fmt clean.
  • Issue cross-links to the new cortex-tensor coordination issue and corinth-canal migration issue.

Validation

# In engram-parser
cargo fmt -- --check
cargo clippy --all-targets --all-features -- -D warnings
cargo test --all-features

# Smoke against real GGUF (use paths from corinth-canal configs or $HOME models)
CHECKPOINT_PATH=... cargo run --example ... (or add simple bin if needed)

Suggested branch

feat/extract-gguf-parser-from-corinth

References

Created as part of modularization follow-up to LIM-9 and engram #5.


Siblings (created 2026-07-01/02):

  • corinth-canal#115
  • cortex-tensor#8
  • Linear LIM-88

Activity

  1. self-assigned this
    on Jul 2, 2026
  2. added
    modularizationWork to make repos more modular and overlapping
    apiAPI changes or trait work
    extractionMain extraction coming from rmems/corinth-canal
    on Jul 2, 2026
  3. rmems commented on Jul 2, 2026

    @rmems
    OwnerAuthor

    Beads coordination issue in corinth-canal (via CLI): corinth-canal-930

  4. rmems commented on Jul 2, 2026

    @rmems
    OwnerAuthor

    Safetensors sibling extraction (Phase 2, parallel to this GGUF work): corinth-canal#116 (bootstrap/supporting issue), cortex-tensor#9 (coord), updated Linear LIM-88, beads raulmc-yqj.

    Note (per review): no dep from corinth-canal; one-way copy of code as inspiration. New target crate will be safetensors-parser (GGUF-only scope of engram-parser preserved). See plan for details.

  5. rmems commented on Jul 2, 2026

    @rmems
    OwnerAuthor

    New CI/DX issues (Docker + Azure): #8 (Azure Pipelines cross-platform) and #9 (Docker workflow). These improve reliability and org consistency (mirroring corinth-canal patterns) alongside the extraction/modularization work. See the session plan.md.

  6. rmems commented on Jul 2, 2026

    @rmems
    OwnerAuthor

    New separate issue for Safetensors extraction (sibling/parallel to this GGUF #7, created in engram-parser per user request): #10

    Tracks the Safetensors case (from corinth-canal#116 source). No dep on corinth-canal (copy from inspiration only). See plan and new #10 for details + cross ecosystem links (cortex#9 etc).

  7. rmems commented on Jul 26, 2026

    @rmems
    OwnerAuthor

    Consumer-side status (cortex-tensor): rmems/cortex-tensor#8 is done on the docs/alignment front — see rmems/cortex-tensor#20.

    • cortex's README now carries a ## Scope / Boundaries section mirroring this crate's, naming engram-parser as the parser-layer provider (GGUF v3 header/KV/tensor-directory parse + per-expert raw weight extraction) and cortex as the consumer (mmap access, dequant→f32, routing math, family adapters).
    • cortex's in-crate parser (src/moe/checkpoint.rs, gguf.rs, dequant.rs) is frozen for parser/dtype work until this issue lands, with freeze markers in the source so no overlapping enhancements start there.
    • Both ecosystem tables now agree, so the two READMEs read as one system.

    Delta still owned here (documented in cortex as parked on this issue, deliberately not ported into cortex): additional dtypes beyond cortex's F32/F16/Q8_0/Q5_K — BF16, Q6_K, IQ3_* — a ggml_type_label-style helper, and the "GGUF wire type 31 is the historical Q4_0_4_4 layout, not IQ3_M" discipline from corinth-canal's src/moe/ggml.rs. rmems/corinth-canal#115's 2026-07-23 note confirms src/moe/gguf/ (metadata/map/dequant/cuda_register) as the extractable unit and one-way-copy contract.

    When this lands, cortex integration should be a follow-up issue there (dependency + adapting parse_checkpoint_layout to load_gguf/GgufLayout), not part of #8.

    Refs: rmems/cortex-tensor#8, rmems/cortex-tensor#20, rmems/corinth-canal#115, LIM-88, LIM-9.

  8. moved this from To triage to Done in Rust stackon Jul 26, 2026
  9. rmems commented on Aug 15, 2026

    @rmems
    OwnerAuthor

    Correction to two earlier comments on this issue: the note that "New target crate will be safetensors-parser (GGUF-only scope of engram-parser preserved)" is superseded.

    Safetensors ships inside engram-parser behind an off-by-default safetensors cargo feature; no separate crate is created. The ## Origin / modularization (#7) README section this issue produced is now joined by a parallel ## Origin / modularization (#10) section following the same shape.

    Unchanged: one-way copy from corinth-canal inspiration, no dependency in either direction, and the zero-dependency charter — [dependencies] stays empty in every feature combination, with the upstream safetensors and serde_json crates explicitly forbidden.

    Rationale of record: #10. Docs PR: #57.

    (Housekeeping: the bead id raulmc-yqj cited above does not exist — the real bead is corinth-canal-le53qh.)

  10. added 2 commits that reference this issue on Aug 15, 2026
  11. rmems commented on Sep 15, 2026

    @rmems
    OwnerAuthor

    Cursor agent — cortex-tensor consume follow-up is rmems/cortex-tensor#47 (wrap parse_checkpoint_layout around engram-parser 0.2.0). It is blocked on #45, not on this closed duplicate. Docs retarget: rmems/cortex-tensor#48.

  12. rmems commented on Sep 15, 2026

    @rmems
    OwnerAuthor

    #45 / RM-367 follow-up (option 1 landed).

    The extraction→adoption gap this issue left after close is tracked here, not as a reopen of this ticket:

    • Consumer blocker: #45 (closed via #73) — optional mmap (load_gguf_mmap) + packed Q8_0 / Q5_K / Q6_K / IQ3_M dequant. Default path stays zero-dep fs::read.
    • Parity/CI evidence: #75 — mmap↔owned packed dequant, page-spanning tensor, sparse 2 GiB mapping without fs::read. Real on-disk multi-GB pilots remain #[ignore] behind ENGRAM_GGUF.
    • Corinth adoption is still corinth-canal#115, gated on this crate (not the reverse).

    Linear: RM-367.

  13. rmems commented on Sep 15, 2026

    @rmems
    OwnerAuthor

    Extraction→adoption gap pointer (from #45 acceptance): #7 shipped the extraction; the consumer (rmems/corinth-canal#115) could not follow against v0.2.0 (fs::read + F16-only dequant).

    That gap was tracked as #45 and shipped in #73 (optional mmap + Q8_0/Q5_K/Q6_K/IQ3_M). corinth-canal#115 is now unblocked under option 1 (upstream first, adopt second). CUDA host-register remains out of this crate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

apiAPI changes or trait workextractionMain extraction coming from rmems/corinth-canalmodularizationWork to make repos more modular and overlapping

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions