⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
-
Updated
Mar 9, 2026 - Python
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
[CVPR 2025 Oral] OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels
The PyVisionAI Official Repo
This is an official repository for "Harnessing Vision Models for Time Series Analysis: A Survey".
A2Mamba: Attention-augmented State Space Models for Visual Recognition
Guidance on deploying a generative AI document analysis with Amazon Bedrock AgentCore. Auto-classifies, enhances, and aggregates multi-type documents using Gestalt-informed vision prompts. Custom analyzer creation wizard. Scripted CDK deployment. Gradio frontend included.
Implementation of VisionLLaMA from the paper: "VisionLLaMA: A Unified LLaMA Interface for Vision Tasks" in PyTorch and Zeta
A simple to use package to call various model providers such as openai, anthropic, and others with utmost reliability, security, and performance.
DART (Diffusion-Autoregressive Recursive Transformer) is a novel hybrid architecture that combines diffusion-based and autoregressive approaches for text generation.
Dataset for PerfCam: Digital Twinning for Production Lines Using 3D Gaussian Splatting and Vision Models
An implementation of gated MLPs in tinygrad, as an alternative to transformers.
A framework to compute threshold sensitivity of deep networks to visual stimuli.
Decom-Renorm-Merge: Merging deep learning models through shared representation space.
PDF Translator: AI-powered PDF translation tool using OpenAI API and vision models for cross-platform document conversion.
🤖 Ollama Consumer - A Python-based interactive chat interface for Ollama models with advanced model management, comprehensive benchmarking, vision support, and automatic error recovery. Features dynamic model switching, GPU optimization, and intelligent service monitoring for seamless AI model interactions.
Local-first photo culling, scoring, AI review, and export workbench for photographers.
Python-based multi-modal embedding service for social media content, providing image, video, and text embeddings with OCR, NSFW detection, and batch processing support for AI-driven feeds and recommendations.
Pipeline for selecting natural images that separate vision models and testing their brain-model alignment.
Small example of AI interpreting charts
Add a description, image, and links to the vision-models topic page so that developers can more easily learn about it.
To associate your repository with the vision-models topic, visit your repo's landing page and select "manage topics."