Skip to content

Feature Request: Add Support for GLM-OCR Model #19335

Description

@huyphuong99

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

Hi team, pinging @ngxson,
I would like to request support for integrating the GLM-OCR model into llama.cpp.
GLM-OCR is an open-source multilingual OCR model published on Hugging Face:
🔗 https://huggingface.co/zai-org/GLM-OCR

The enhancement requested:
Allow loading and running GLM-OCR using llama.cpp or ggml-compatible formats.
Provide support for text extraction (OCR) inference via existing or extended pipelines.
Optional: Include sample usage scripts demonstrating how to run GLM-OCR within the llama.cpp framework.
This feature would significantly extend the usability of llama.cpp beyond language models, enabling OCR tasks using the same efficient CPU/GPU inference backend.

Motivation

GLM-OCR provides strong multilingual OCR capabilities and is useful for applications involving text recognition from images. Adding support in llama.cpp would:
Enable users to run OCR tasks fully offline with high efficiency
Expand the diversity of model types supported by ggml/llama.cpp
Make it easier to build integrated pipelines using both LLMs and OCR in a single unified runtime
Support more community projects that require lightweight OCR models without depending on cloud APIs
This aligns with llama.cpp’s goal of enabling fast, portable AI inference on local devices.

Possible Implementation

Some possible directions:
Convert GLM-OCR into a ggml/gguf-compatible format, similar to how vision or multimodal models are handled.
Extend the model loader to support GLM-OCR’s architecture (if it diverges from existing supported ones).

Activity

  1. notune commented on Feb 16, 2026

    @notune

    i got a basic prototype of an implementation working (in case you already need it), but quality wise i think i cant yet make a PR, especially since all of the code was written by opus 4.6.
    But i tested it with cuda, cpu and bf16 and q8_0. it was slightly faster then the ollama implementation and produced the same output. I also ensured to eliminated touches on shared code, so it shouldnt interfere with other models.
    https://github.com/notune/llama.cpp/tree/model-glm-ocr

  2. TheNexter commented on Feb 19, 2026

    @TheNexter

    Edit : Forget this it's a bug with webui, api work

    @ngxson For me, GLM-OCR is not working, my models.ini file to provide model to llama-server :
    Version of llama-server : 8096 (8a7097355) Linux x86_64 :

    models.ini :

    [GLM-OCR-Q8_0_default]
    hf = ggml-org/GLM-OCR-GGUF
    

    Output :

    Capture.video.du.2026-02-19.11-53-18.mp4
  3. morpheumstreet commented on Mar 27, 2026

    @morpheumstreet
  4. morpheumstreet commented on Mar 27, 2026

    @morpheumstreet

    llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'glmocr'
    llama_model_load_from_file_impl: failed to load model
    llama_params_fit: encountered an error while trying to fit params to free device memory: failed to load model
    llama_params_fit: fitting params to free memory took 0.04 seconds
    llama_model_load_from_file_impl: using device MTL0 (Apple M3 Ultra) (unknown id) - 228064 MiB free

    the latest version still not support glm ocr

  5. ngxson commented on Mar 27, 2026

    @ngxson
    Collaborator
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions