Prerequisites
Feature Description
Hi team, pinging @ngxson,
I would like to request support for integrating the GLM-OCR model into llama.cpp.
GLM-OCR is an open-source multilingual OCR model published on Hugging Face:
🔗 https://huggingface.co/zai-org/GLM-OCR
The enhancement requested:
Allow loading and running GLM-OCR using llama.cpp or ggml-compatible formats.
Provide support for text extraction (OCR) inference via existing or extended pipelines.
Optional: Include sample usage scripts demonstrating how to run GLM-OCR within the llama.cpp framework.
This feature would significantly extend the usability of llama.cpp beyond language models, enabling OCR tasks using the same efficient CPU/GPU inference backend.
Motivation
GLM-OCR provides strong multilingual OCR capabilities and is useful for applications involving text recognition from images. Adding support in llama.cpp would:
Enable users to run OCR tasks fully offline with high efficiency
Expand the diversity of model types supported by ggml/llama.cpp
Make it easier to build integrated pipelines using both LLMs and OCR in a single unified runtime
Support more community projects that require lightweight OCR models without depending on cloud APIs
This aligns with llama.cpp’s goal of enabling fast, portable AI inference on local devices.
Possible Implementation
Some possible directions:
Convert GLM-OCR into a ggml/gguf-compatible format, similar to how vision or multimodal models are handled.
Extend the model loader to support GLM-OCR’s architecture (if it diverges from existing supported ones).
Prerequisites
Feature Description
Hi team, pinging @ngxson,
I would like to request support for integrating the GLM-OCR model into llama.cpp.
GLM-OCR is an open-source multilingual OCR model published on Hugging Face:
🔗 https://huggingface.co/zai-org/GLM-OCR
The enhancement requested:
Allow loading and running GLM-OCR using llama.cpp or ggml-compatible formats.
Provide support for text extraction (OCR) inference via existing or extended pipelines.
Optional: Include sample usage scripts demonstrating how to run GLM-OCR within the llama.cpp framework.
This feature would significantly extend the usability of llama.cpp beyond language models, enabling OCR tasks using the same efficient CPU/GPU inference backend.
Motivation
GLM-OCR provides strong multilingual OCR capabilities and is useful for applications involving text recognition from images. Adding support in llama.cpp would:
Enable users to run OCR tasks fully offline with high efficiency
Expand the diversity of model types supported by ggml/llama.cpp
Make it easier to build integrated pipelines using both LLMs and OCR in a single unified runtime
Support more community projects that require lightweight OCR models without depending on cloud APIs
This aligns with llama.cpp’s goal of enabling fast, portable AI inference on local devices.
Possible Implementation
Some possible directions:
Convert GLM-OCR into a ggml/gguf-compatible format, similar to how vision or multimodal models are handled.
Extend the model loader to support GLM-OCR’s architecture (if it diverges from existing supported ones).