Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
-
Updated
Jul 31, 2026 - Markdown
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks. In NeurIPS 2020 workshop.
Fast, lossless LLM inference via dual-view diffusion decoding.
This implements training of popular model architectures, such as AlexNet, ResNet and VGG on the ImageNet dataset(Now we supported alexnet, vgg, resnet, squeezenet, densenet)
Pytorch Imagenet Models Example + Transfer Learning (and fine-tuning)
CVPR 2025 (Highlight)
VGG16 Net implementation from PyTorch Examples scripts for ImageNet dataset
deeplearning.ai Tensorflow advance techniques specialization
CAE-ADMM: Implicit Bitrate Optimization via ADMM-Based Pruning in Compressive Autoencoders
Compose neural architectures as typed, executable graph cards with two-way PyTorch sync, step-by-step local execution, and a constrained AI graph planner.
ElasticModels is a elasticsearch object modeling tool designed to work in and asynchronous environment. Builded for official elasticsearch client library Main inspiration was mongoose project
Chest X-Ray Image classification using PyTorch.
Build and train a mini Kimi K3 from scratch — a pure-PyTorch playground from a custom 88M model to 2T-scale LLM profiles.
PyTorch model architecture diagram generator: neural network diagram, SVG, PNG, TikZ, LaTeX
This Repository contains TensorFlow implementation of different Image Segmentation Architecture on different types of datasets.
Official PyTorch and CVXPY implementation of Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
Efficient transformer with additive knowledge injection, designed using Domain Abstraction Collapse methodology
MoLE (Mix-of-Language-Experts) model architecture for multilingual programming
Identify traffic sign images through Supervised Classification via Deep Learning and Computer Vision using Python, Tensorflow, Jupyter and Anaconda in AWS Cloud.
读懂模型,而不是追逐名词:开源模型结构图鉴
Add a description, image, and links to the model-architecture topic page so that developers can more easily learn about it.
To associate your repository with the model-architecture topic, visit your repo's landing page and select "manage topics."