Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

10 Commits
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Awesome PR's Welcome Visitors arXiv Hugging Face Papers

Awesome Agentic World Modeling

This repository accompanies the position paper "Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents," and curates the works it reviews. The survey argues for a shift from a World Model that passively predicts the next physical state toward an Agent-Centric World Proxy: an environment-grounded interface that returns the information transition an agent needs (a future state, rendered view, execution result, retrieved memory or skill, or verdict on a plan) to plan, learn, and continually improve.

The taxonomy has two orthogonal axes:

  • Six proxy functions (what feedback the proxy returns): Dynamics, Spatial, Execution, Memory / Experience, Skill, and Reward / Verification.
  • Three empowerment levels (how the proxy improves the agent): L1 inference-time guidance, L2 training-time optimization, and L3 Agent-Proxy co-evolution.

Call for Contributions

Agentic world modeling moves fast, and the boundary between a passive world model and an agent-facing world proxy is still being drawn. This list aims to be a living, community-maintained map of that space, and contributions of new work or corrections are very welcome.

To contribute, open an issue or a pull request:

  • Where: add each entry to the section matching its proxy function (or the closest one), and keep every table in chronological order.
  • Format: follow the existing row layout: a short model name in Model (or - if none), the title with an arXiv badge in Paper, and the abbreviated venue with year in Venue.
  • Verify: check the title, authors, arXiv ID, and venue against the primary source before adding, and prefer the formally published version whenever one exists.
  • Fix: corrections to venues, broken links, or mis-categorized entries are equally welcome; flag them in an issue or fix them directly in a pull request.

Table of Contents

Background

Classical World Models & Model-Based RL

Foundations: state-transition world models and model-based RL that predict the next physical state given a state and an action.

Model Paper Venue
- Dyna, an integrated architecture for learning, planning, and reacting SIGART 1991
- arXiv
Recurrent World Models Facilitate Policy Evolution
NeurIPS 2018
- arXiv
When to Trust Your Model: Model-Based Policy Optimization
NeurIPS 2019
- arXiv
Dream to Control: Learning Behaviors by Latent Imagination
ICLR 2020
- arXiv
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Nature 2020
- A Path Towards Autonomous Machine Intelligence OpenReview preprint 2022
- arXiv
Model-Based Reinforcement Learning: A Survey
FTML 2023
- arXiv
Mastering diverse control tasks through world models
Nature 2025

Continual Improvement & Interaction Cost

Why real-environment interaction alone cannot carry continual improvement: cost, safety, latency, and the need for scalable feedback.

Model Paper Venue
MuJoCo MuJoCo: A Physics Engine for Model-Based Control IROS 2012
- arXiv
Isaac Gym: High Performance GPU-Based Physics Simulation for Robot Learning
NeurIPS 2021
DayDreamer arXiv
DayDreamer: World Models for Physical Robot Learning
CoRL 2023
CodeIt arXiv
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
ICML 2024
Voyager arXiv
Voyager: An Open-Ended Embodied Agent with Large Language Models
TMLR 2024
OSWorld arXiv
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NeurIPS 2024
WebArena arXiv
WebArena: A Realistic Web Environment for Building Autonomous Agents
ICLR 2024
CASCADE arXiv
CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution
arXiv 2025
- arXiv
Continual Learning and Catastrophic Forgetting
Learning and Memory: A Comprehensive Reference 2025
- arXiv
Agent Learning via Early Experience
ICML 2026

The Six Proxy Functions

Each proxy function answers a different question for the agent, but all return agent-usable feedback.

Dynamics Proxy

How would the environment change if the agent executed a given action? Video, driving, game, and model-based dynamics predictors.

Model Paper Venue
- Learning Latent Dynamics for Planning from Pixels ICML 2019
FitVid arXiv
FitVid: Overfitting in Pixel-Level Video Prediction
arXiv 2021
VideoGPT arXiv
VideoGPT: Video Generation using VQ-VAE and Transformers
arXiv 2021
- arXiv
Video Diffusion Models
NeurIPS 2022
- arXiv
Diffusion Models for Video Prediction and Infilling
TMLR 2022
MCVD arXiv
MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation
NeurIPS 2022
GAIA-1 arXiv
GAIA-1: A Generative World Model for Autonomous Driving
arXiv 2023
- arXiv
Transformers Are Sample-Efficient World Models
ICLR 2023
STORM STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning NeurIPS 2023
- arXiv
Diffusion for World Modeling: Visual Details Matter in Atari
NeurIPS 2024
- Video generation models as world simulators 2024
Genie arXiv
Genie: Generative Interactive Environments
ICML 2024
Vista arXiv
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
NeurIPS 2024
TD-MPC2 arXiv
TD-MPC2: Scalable, Robust World Models for Continuous Control
ICLR 2024
RoboCasa RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots RSS 2024
DriveDreamer arXiv
DriveDreamer: Towards Real-World-Driven World Models for Autonomous Driving
ECCV 2024
iVideoGPT iVideoGPT: Interactive VideoGPTs are Scalable World Models NeurIPS 2024
- arXiv
Learning Interactive Real-World Simulators
ICLR 2024
DynamicCity arXiv
DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
ICLR 2025
GameGen-X arXiv
GameGen-X: Interactive Open-world Game Video Generation
ICLR 2025
- Genie 3: A New Frontier for World Models 2025
- arXiv
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
arXiv 2025
- arXiv
World Modelling Improves Language Model Agents
arXiv 2025
MineWorld arXiv
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
arXiv 2025
MobileWorldBench arXiv
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
arXiv 2025
- World Model on Million-Length Video And Language With Blockwise RingAttention ICLR 2025
Vision arXiv
Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models
arXiv 2025
- arXiv
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
arXiv 2025
- Diffusion Models Are Real-Time Game Engines ICLR 2025
Marble Marble: A Multimodal World Model 2025
- arXiv
Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving
AAAI 2025
Matrix-Game arXiv
Matrix-Game: Interactive World Foundation Model
arXiv 2025
Language arXiv
Language-Conditioned World Modeling for Visual Navigation
arXiv 2026
- arXiv
Advancing Open-source World Models
arXiv 2026
NavThinker arXiv
NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation
arXiv 2026
LiDARCrafter arXiv
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
AAAI 2026
Self arXiv
Self-Execution Simulation Improves Coding Models
arXiv 2026
- arXiv
Is Your Driving World Model an All-Around Player?
CVPR 2026
U4D arXiv
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
CVPR 2026
AD-R1 arXiv
AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
CVPR 2026
SPIRAL arXiv
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
arXiv 2026

Spatial Proxy

What would the agent observe from another viewpoint or position? NeRF/Gaussian rendering, 3D/4D reconstruction and generation, and visual imagination for navigation.

Model Paper Venue
NeRF arXiv
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
ECCV 2020
Mip-NeRF arXiv
Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
ICCV 2021
PathDreamer arXiv
PathDreamer: A World Model for Indoor Navigation
ICCV 2021
Mip arXiv
Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
CVPR 2022
SceneDreamer arXiv
SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections
TPAMI 2023
SceneScape arXiv
SceneScape: Text-Driven Consistent Scene Generation
NeurIPS 2023
Text2Room arXiv
Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
ICCV 2023
VoxPoser arXiv
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
CoRL 2023
- arXiv
3D Gaussian Splatting for Real-Time Radiance Field Rendering
TOG 2023
Nerfstudio arXiv
Nerfstudio: A Modular Framework for Neural Radiance Field Development
SIGGRAPH Asia 2023
- arXiv
A Survey on 3D Gaussian Splatting
CSUR 2024
- arXiv
3D Gaussian as a New Era: A Survey
TVCG 2024
- arXiv
Grounding Image Matching in 3D with MASt3R
ECCV 2024
DreamScene arXiv
DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
ECCV 2024
Scaffold-GS arXiv
Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
CVPR 2024
Dreamforge arXiv
Dreamforge: Motion-aware autoregressive video generation for multi-view driving scenes
arXiv 2024
DriveWorld DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving CVPR 2024
DimensionX arXiv
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
arXiv 2024
ManiSkill3 arXiv
ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
arXiv 2024
DUSt3R arXiv
DUSt3R: Geometric 3D Vision Made Easy
CVPR 2024
- arXiv
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
CVPR 2024
CityDreamer arXiv
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
CVPR 2024
Mip-Splatting arXiv
Mip-Splatting: Alias-Free 3D Gaussian Splatting
CVPR 2024
Text2NeRF arXiv
Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
TVCG 2024
Copilot4D Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion ICLR 2024
OccWorld arXiv
OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
ECCV 2024
ReCamMaster arXiv
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
arXiv 2025
- arXiv
Navigation world models
CVPR 2025
PhysX-Anything arXiv
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
arXiv 2025
GAF arXiv
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
arXiv 2025
ParticleFormer arXiv
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
arXiv 2025
- arXiv
3D and 4D World Modeling: A Survey
arXiv 2025
GWM GWM: Towards Scalable Gaussian World Models for Robotic Manipulation ICCV 2025
MindCube arXiv
MindCube: Spatial Mental Modeling from Limited Views
arXiv 2025
GAIA-2 arXiv
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
arXiv 2025
- arXiv
3D Reconstruction with Spatial Memory
3DV 2025
- arXiv
Continuous 3D Perception Model with Persistent State
arXiv 2025
Fast3R arXiv
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
arXiv 2025
WonderWorld arXiv
WonderWorld: Interactive 3D Scene Generation from a Single Image
CVPR 2025
- arXiv
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
IJCAI 2025
- arXiv
Occupancy World Model for Robots
arXiv 2025
TesserAct arXiv
TesserAct: Learning 4D Embodied World Models
arXiv 2025
Aether arXiv
Aether: Geometric-Aware Unified World Modeling
ICCV 2025
OccuBench arXiv
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
arXiv 2026
PointWorld arXiv
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
arXiv 2026
CounterScene arXiv
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
arXiv 2026
MVISTA-4D arXiv
MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
arXiv 2026
X-scene arXiv
X-scene: Large-scale driving scene generation with high fidelity and flexible controllability
NeurIPS 2026
Code2World arXiv
Code2World: A GUI World Model via Renderable Code Generation
arXiv 2026

Execution Proxy

What would the digital environment return after an operation? Web, GUI, code, shell, and API/tool execution predictors.

Model Paper Venue
RT-2 arXiv
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
CoRL 2023
- arXiv
Code as Policies: Language Model Programs for Embodied Control
ICRA 2023
Toolformer arXiv
Toolformer: Language Models Can Teach Themselves to Use Tools
NeurIPS 2023
- arXiv
Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search
NeurIPS 2024
AgentBench arXiv
AgentBench: Evaluating LLMs as Agents
ICLR 2024
- arXiv
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
NeurIPS 2024
Math-Shepherd arXiv
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
ACL 2024
- arXiv
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
ICLR 2025
CWM arXiv
CWM: An Open-Weights LLM for Research on Code Generation with World Models
arXiv 2025
- arXiv
Web World Models
arXiv 2025
WebSynthesis arXiv
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
arXiv 2025
- arXiv
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
TMLR 2025
MegaSaM arXiv
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
CVPR 2025
ToolSandbox arXiv
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
NAACL 2025
ViMo arXiv
ViMo: A Generative Visual GUI World Model for App Agents
arXiv 2025
- arXiv
Cosmos World Foundation Model Platform for Physical AI
arXiv 2025
GTM arXiv
GTM: Simulating the World of Tools for AI Agents
arXiv 2025
- arXiv
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
arXiv 2025
GameFactory GameFactory: Creating New Games with Generative Interactive Videos ICCV 2025
ComAct arXiv
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
arXiv 2026
MobileDreamer arXiv
MobileDreamer: Generative Sketch World Model for GUI Agent
arXiv 2026
Computer arXiv
Computer-Using World Model
arXiv 2026
MCP-Cosmos arXiv
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
arXiv 2026
- arXiv
Generative Visual Code Mobile World Models
arXiv 2026
IterCAD arXiv
IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing
arXiv 2026
- arXiv
Debugging Code World Models
arXiv 2026
NeuralOS arXiv
NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
ICLR 2026
World-Model arXiv
World-Model-Augmented Web Agents with Action Correction
arXiv 2026
WebWorld arXiv
WebWorld: A Large-Scale World Model for Web Agent Training
arXiv 2026

Memory, Experience & Skill Proxy

What prior experience, constraints, or reusable skills does the agent need now? Store-retrieve memory systems and reusable skill/behavior libraries.

Model Paper Venue
- arXiv
Reasoning with Language Model is Planning with World Model
EMNLP 2023
- arXiv
Generative Agents: Interactive Simulacra of Human Behavior
UIST 2023
- arXiv
Cognitive Architectures for Language Agents
TMLR 2024
Vismem arXiv
Vismem: Latent vision memory unlocks potential of vision-language models
arXiv 2025
MemHarness arXiv
MemHarness: Memory Is Reconstructed, Not Replayed
arXiv 2026

Reward / Verification Proxy

Is the agent's behavior correct, safe, and feasible, and how should it improve? Reward models, verifiers, critics, preference models, and LLM-as-judge.

Model Paper Venue
- arXiv
Deep Reinforcement Learning from Human Preferences
NeurIPS 2017
Fine arXiv
Fine-Tuning Language Models from Human Preferences
arXiv 2019
- arXiv
Training Verifiers to Solve Math Word Problems
arXiv 2021
- arXiv
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
NeurIPS 2023
- arXiv
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
NeurIPS 2023
- arXiv
A General Theoretical Paradigm to Understand Learning from Human Preferences
AISTATS 2024
- arXiv
A Survey on LLM-as-a-Judge
arXiv 2024
ORPO arXiv
ORPO: Monolithic Preference Optimization without Reference Model
EMNLP 2024
- arXiv
Let's Verify Step by Step
ICLR 2024
- arXiv
LLM Critics Help Catch LLM Bugs
arXiv 2024
SimPO arXiv
SimPO: Simple Preference Optimization with a Reference-Free Reward
NeurIPS 2024
WorldSimBench arXiv
WorldSimBench: Towards Video Generation Models as World Simulators
arXiv 2024
WorldModelBench arXiv
WorldModelBench: Judging Video Generation Models As World Models
arXiv 2025
Computer arXiv
Computer-Use Agents as Judges for Generative User Interface
arXiv 2025
Critic-V arXiv
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
CVPR 2025

Empowerment Levels (L1 / L2 / L3)

How proxies empower agents: L1 inference-time guidance (advise a frozen agent), L2 training-time optimization (generate experience/signal to train the agent), and L3 Agent-Proxy co-evolution (both improve each other in a loop).

Model Paper Venue
- arXiv
Scaling Agent Learning via Experience Synthesis
arXiv 2025
RAGEN arXiv
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
arXiv 2025
VAGEN arXiv
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
NeurIPS 2025
WebEvolver arXiv
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
EMNLP 2025
- arXiv
Aligning Agentic World Models via Knowledgeable Experience Learning
arXiv 2026
- arXiv
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
ICML 2026
SearchGym arXiv
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
arXiv 2026

Evaluation, Fidelity & Safety

Benchmarks and metrics for world models, fidelity and calibration limits, and the safety/attack surface of proxy-driven learning.

Model Paper Venue
G-Eval arXiv
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
EMNLP 2023
Length arXiv
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
COLM 2024
Prometheus arXiv
Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models
ICLR 2024
BEHAVIOR-1K arXiv
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
arXiv 2024
- arXiv
Evaluating the World Model Implicit in a Generative Model
NeurIPS 2024
- arXiv
World Models: The Safety Perspective
arXiv 2024
- arXiv
Understanding World or Predicting Future? A Comprehensive Survey of World Models
CSUR 2025
- arXiv
How Far is Video Generation from World Model: A Physical Law Perspective
ICML 2025
- arXiv
When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
arXiv 2026
- arXiv
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
arXiv 2026
JailWAM arXiv
JailWAM: Jailbreaking World Action Models in Robot Control
arXiv 2026
- arXiv
Current Agents Fail to Leverage World Model as Tool for Foresight
arXiv 2026
CtrlAttack arXiv
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
arXiv 2026
SafeDream arXiv
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
arXiv 2026

Surveys & Related Work

Related surveys and the paper's core position statements.

Model Paper Venue
- arXiv
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
arXiv 2024
- arXiv
A Survey of World Models for Autonomous Driving
arXiv 2025
- arXiv
A Comprehensive Survey on World Models for Embodied AI
arXiv 2025
- arXiv
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
arXiv 2026

Additional References

Further supporting works cited in the survey.

Model Paper Venue
- arXiv
Proximal Policy Optimization Algorithms
arXiv 2017
Habitat Habitat: A Platform for Embodied AI Research ICCV 2019
- Model-Based Reinforcement Learning for Atari ICLR 2020
- arXiv
Learning to Summarize from Human Feedback
NeurIPS 2020
- arXiv
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
CoRL 2020
CLIPort arXiv
CLIPort: What and Where Pathways for Robotic Manipulation
CoRL 2021
- arXiv
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
CoRL 2022
- arXiv
Constitutional AI: Harmlessness from AI Feedback
arXiv 2022
- arXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
arXiv 2022
TensoRF arXiv
TensoRF: Tensorial Radiance Fields
ECCV 2022
- arXiv
Inner Monologue: Embodied Reasoning through Planning with Language Models
CoRL 2022
- arXiv
Large Language Models are Zero-Shot Reasoners
NeurIPS 2022
- arXiv
Instant Neural Graphics Primitives with a Multiresolution Hash Encoding
TOG 2022
- arXiv
Training Language Models to Follow Instructions with Human Feedback
NeurIPS 2022
Perceiver-Actor arXiv
Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
CoRL 2022
- arXiv
Solving Math Word Problems with Process- and Outcome-Based Feedback
arXiv 2022
Chain-of arXiv
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
NeurIPS 2022
Plenoxels arXiv
Plenoxels: Radiance Fields without Neural Networks
CVPR 2022
Self arXiv
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
CVPR 2023
RT-1 arXiv
RT-1: Robotics Transformer for Real-World Control at Scale
RSS 2023
Self-Refine arXiv
Self-Refine: Iterative Refinement with Self-Feedback
NeurIPS 2023
MemGPT arXiv
MemGPT: Towards LLMs as Operating Systems
arXiv 2023
Reflexion arXiv
Reflexion: Language Agents with Verbal Reinforcement Learning
NeurIPS 2023
ReAct arXiv
ReAct: Synergizing Reasoning and Acting in Language Models
ICLR 2023
- arXiv
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
NeurIPS 2023
RRHF arXiv
RRHF: Rank Responses to Align Language Models with Human Feedback without Tears
NeurIPS 2023
KTO arXiv
KTO: Model Alignment as Prospect Theoretic Optimization
ICML 2024
SWE-bench arXiv
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
ICLR 2024
Oasis Oasis: A Universe in a Transformer 2024
- arXiv
Agent Planning with World Knowledge Model
NeurIPS 2024
DeepSeekMath arXiv
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
arXiv 2024
VGGT arXiv
VGGT: Visual Geometry Grounded Transformer
arXiv 2025
MonST3R arXiv
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
ICLR 2025
- arXiv
The latent space: Foundation, evolution, mechanism, ability, and outlook
arXiv 2026

Acknowledgements

This list is maintained alongside the survey. Thanks to all contributors and to the authors of the works listed here.

Citation

If you find this work useful, please consider citing:

@article{yang2026worldproxy,
  title   = {Quo Vadis, World Modeling?},
  author  = {Yu Yang and Xuemeng Yang and Licheng Wen and Lingdong Kong and Xiaobin Hu and Dongyue Lu and Wei Chow and Xiyan Huang and Yuxiang Feng and Yue Liao and Jianbiao Mei and Daocheng Fu and Rong Wu and Pinlong Cai and Ran Yi and Ying Tai and Jiangning Zhang and Botian Shi and Yong Liu and Shuicheng Yan},
  journal = {arXiv preprint arXiv:2608.02713},
  year    = {2026}
}