RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
-
Updated
Feb 28, 2026 - Python
RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
UI-Venus is a general-purpose foundation GUI agent for mobile apps, web platforms, and desktop operating systems using only screenshots as input.
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
[ACL 2025] The code repository for "Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning" in PyTorch.
🦙 echoOLlama: A real-time voice AI platform powered by local LLMs. Features WebSocket streaming, voice interactions, and OpenAI API compatibility. Built with FastAPI, Redis, and PostgreSQL. Perfect for private AI conversations and custom voice assistants.
Not a neutral survey — a field manual for engineers who build, train, and ship multimodal retrieval at production scale. The C-L-I triangle (Compression · Localization · Instruction), MLLM encoders vs late interaction, MUVERA economics, and falsifiable forecasts through 2030.
Build a simple basic multimodal large model from scratch. 从零搭建一个简单的基础多模态大模型🤖
A comprehensive survey of Vision–Language Models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets
[AAAI'26] Official implementation of CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
[ACMMM 2026] SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
Using MAIRA-2 multimodal transformer designed for the generation of grounded or non-grounded radiology reports from chest X-rays.
Operational tooling & discipline for multi-agent AI engineering fleets — mutation-proven guards, a 364-test fleet monitor, generalized skills, and specs. Every check ships with proof it can fail.
Evaluating ‘Graphical Perception’ with Multimodal Large Language Models
Multi-Modal Healthcare Assistant
Gemma3 Vision - AI Image Analysis & Chat
ElaMath is a smart, voice-enabled math assistant that helps students solve and understand math problems using both spoken questions and images. It’s powered by the powerful multimodal meta-llama/llama-4-scout-17b-16e-instruct model via Groq API, combined with Whisper for speech recognition and ElevenLabs/gTTS for natural voice responses.
Elarova — A smart, multimodal research assistant designed to help students by combining speech, text, and other input modes for efficient academic research and study support. Powered by state-of-the-art speech recognition, text-to-speech, and AI models, including meta-llama/llama-4-scout-17b-16e-instruct, with an easy-to-use Gradio web interface.
Create a tool that uses a multimodal LLM to describe testing instructions for any digital product's features, based on the screenshots.
Multimodel Document Intelliigence for better document understanding and context awareness for Academic Documents
To associate your repository with the multimodel-large-language-model topic, visit your repo's landing page and select "manage topics."