Making large AI models cheaper, faster and more accessible
-
Updated
Oct 10, 2026 - Python
Making large AI models cheaper, faster and more accessible
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Janus-Series: Unified Multimodal Understanding and Generation Models
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.
⚡ TabPFN: Foundation Model for Tabular Data ⚡
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Chronos: Pretrained Models for Time Series Forecasting
LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence https://arxiv.org/abs/2609.17488
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
Images to inference with no labeling (use foundation models to train supervised models).
EVA Series: Visual Representation Fantasies from BAAI
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
Prompt Learning for Vision-Language Models (IJCV'22, CVPR'22)
To associate your repository with the foundation-models topic, visit your repo's landing page and select "manage topics."