Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
-
Updated
Oct 30, 2025 - Python
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
Ultralytics fork of Apple MobileCLIP for fast image-text inference, training, evaluation, and an iOS demo.
Using Segment-Anything and CLIP to generate pixel-aligned semantic features.
Model merging, task-vector rebasin, and fine-tuning for vision and LLM models.
Clipora is a powerful toolkit for fine-tuning OpenCLIP models using Low Rank Adapters (LoRA).
I switched phones and found thousands of WhatsApp photos waiting for me, mostly memes and screenshots. This finds the ones worth keeping: search your pictures by describing them, and bin the rest into a quarantine you can undo.
Text-to-image search with OpenCLIP, Docker, Flask, Faiss, etc. and a basic front-end.
A doctor-assistive AI system that interprets medical knowledge and patient images simultaneously. It utilizes a Dual-Encoder architecture to cross-reference textbook theory with visual pathology, generating clinically grounded diagnoses.
An AI-powered computer vision system that automatically selects the best wedding photos from thousands of images.
Dynamic cluster-based data sampling for efficient and long-tail-aware vision-language model pre-training.
VALORA AI is a Multimodal Pricing Prediction Model that uses textual and visual data to make precise predictions on product prices
Masked Multi-Component Gated Decomposition Architecture
Local-first semantic image explorer powered by OpenCLIP embeddings and Qdrant vector search, enabling natural-language retrieval across screenshots and photos.
CLIPF: Contrastive Language-Image Pre-training with Word Frequency Masking — a frequency-based text masking strategy for efficient VLM pre-training (WACV 2026)
Feed-forward video-to-3D scene reconstruction using VGGT (CVPR 2025) with open-vocabulary semantic labeling via SAM 2.1 + CLIP. One command, no COLMAP, Apple Silicon + CUDA.
A redesigned CLIP architecture replacing the ViT encoder with modern CNN backbones (ConvNeXt V2) to improve efficiency and inference speed while maintaining strong vision-language alignment, enabling real-time use in robotics and edge applications.
An Edge-AI security pipeline for real-time violence detection and automated forensic reporting using YOLO, CoCa, and Qwen2.5-VL.
基于OpenCLIP的图像检索系统,支持文搜图、图搜图,结合FAISS能够通过用户上传图片库快速建立索引来实现图像检索
To associate your repository with the openclip topic, visit your repo's landing page and select "manage topics."