MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
-
Updated
Aug 21, 2026 - Python
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Tool for automating common video key-frame extraction, video compression and Image Auto-crop/Image-resize tasks
[CVPR'25 Highlight] The official implementation of "GG-SSMs: Graph-Generating State Space Models"
OpenNCC Kit
Real-time webcam demo with SmolVLM(mlx-community/SmolVLM-Instruct-4bit) and MLX-VLM
Vision framework which brings a more robust, Deep Learning-based approach to some usual OpenCV use cases
Python scripts for use with Turi Create to output a Xcode compatible mlmodel file for use with machine learning object detection with the CoreML or Vision frameworks.
Research the flying-car vision high-tech
On-device Perceive → Reason pipeline for Apple Silicon: Core ML + Vision for perception, a swappable LanguageModel (Apple Foundation Models or Claude) for reasoning. Python conversion/quantization toolkit plus a SwiftUI reference app.
Python wrapper around the Native OCR engines that ship with macOS and Windows
Turn a scrolling chat screen recording into a speaker-separated, chronologically ordered transcript. macOS-native (AVFoundation + Vision), no ffmpeg or dependencies.
TextLift OCR text extractor
Mathematical Foundations of Recursive Cortical Ignition: A Conditional Theory of Sub-Threshold Pattern Coding in Recurrent Visual Cortex
Add a description, image, and links to the vision-framework topic page so that developers can more easily learn about it.
To associate your repository with the vision-framework topic, visit your repo's landing page and select "manage topics."