Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

5 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Latent World Models

Awesome License PRs Welcome Awesome Lint

Latent World Models banner

Learning to predict, reason, and plan in latent space.

A curated reading map for latent world models, latent reasoning, JEPA-style predictive architectures, and latent planning/control. The focus is work where latent representations are used as the substrate for prediction, imagination, planning, control, action-conditioned modeling, or reasoning. Generic embedding papers are intentionally out of scope unless the latent space itself does useful dynamics, planning, or decision-making work.

Last curated: 2026-06-13.

Table of Contents

Start Here

  1. Why latent world models? Read LeCun's position paper, Ha and Schmidhuber's World Models, and PlaNet.
  2. How do agents plan in latent space? Follow Dreamer, TD-MPC, TD-MPC2, MuZero, and EfficientZero.
  3. Why JEPA? Read I-JEPA, V-JEPA, V-JEPA 2, V-JEPA 2.1, and LeWorldModel.
  4. What is new in latent reasoning? Compare COCONUT, recurrent-depth latent reasoning, concept-space models, and visual latent reasoning in VLMs.
  5. Where does embodied control enter? Study latent action policies, LAPA/LAWM, V-JEPA 2-AC, WAMs, and robotics world models.

Latent Meme Corner

World Models registration desk meme

The model said it was thinking, but all we found was a beautifully compressed hidden state.

Surveys and Position Papers

  • A Path Towards Autonomous Machine Intelligence — LeCun's position paper frames JEPA-style predictive world models as a route from perception to planning and common-sense reasoning. OpenReview
  • Foundation Models for Decision Making: Problems, Methods, and Opportunities — Places world models, planning, and control inside the broader foundation-model-for-agents agenda. arXiv
  • A Tutorial on World Models and Physical AI — A recent physical-AI tutorial that distinguishes explicit rollout models from implicit predictive representations. arXiv

Yann LeCun / JEPA Series

  • Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Introduces I-JEPA, predicting target block representations from context blocks without pixel reconstruction. arXiv Code
  • Learning and Leveraging World Models in Visual Representation Learning — Extends JEPA-style prediction beyond masking by learning latent effects of image transformations. arXiv
  • Revisiting Feature Prediction for Learning Visual Representations from Video — Introduces V-JEPA and shows that feature prediction from video can learn strong motion-aware representations. arXiv Code
  • V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Scales video JEPA and includes V-JEPA 2-AC, a latent action-conditioned world model for robot planning. arXiv
  • V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning — Improves dense spatial-temporal features while retaining video understanding and robotics transfer. arXiv
  • LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — Trains a compact JEPA world model end-to-end from pixels for fast latent planning and surprise detection. arXiv
  • A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures — Provides small, reproducible JEPA examples that connect representation learning, temporal prediction, and action-conditioned world models. arXiv Code

Classical Latent World Models

  • Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images — An early control-from-pixels model that constrains dynamics to be locally linear in latent space. arXiv
  • World Models — Compresses observations with a VAE, models temporal dynamics with an RNN, and trains compact policies inside imagined rollouts. arXiv Website
  • Recurrent World Models Facilitate Policy Evolution — Shows that compact policies can exploit recurrent latent world models trained separately from controllers. arXiv Website
  • Learning Latent Dynamics for Planning from Pixels — PlaNet learns stochastic latent dynamics from images and uses online planning directly in latent space. arXiv Code
  • Dream to Control: Learning Behaviors by Latent Imagination — Dreamer learns behavior by backpropagating value gradients through imagined latent trajectories. arXiv Code
  • Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model — SLAC separates representation learning from task learning by performing RL over learned stochastic latent states. arXiv
  • Mastering Atari with Discrete World Models — DreamerV2 uses discrete latent world states to train policies by imagination at Atari scale. arXiv Code
  • Mastering Diverse Domains through World Models — DreamerV3 demonstrates a robust general world-model agent across control, Atari, Minecraft, and more. arXiv Code

Latent Planning and Control

  • Temporal Difference Learning for Model Predictive Control — TD-MPC learns a task-oriented latent dynamics model and plans with value-guided trajectory optimization. arXiv Code
  • TD-MPC2: Scalable, Robust World Models for Continuous Control — Scales decoder-free latent world models to many tasks, embodiments, and action spaces. arXiv Code
  • Hierarchical Planning with Latent World Models — Learns latent world models at multiple temporal scales so long-horizon embodied planning can use coarser abstractions before fine control. arXiv
  • Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — MuZero learns hidden dynamics for search without reconstructing observations or knowing environment rules. arXiv
  • Learning and Planning in Complex Action Spaces — Sampled MuZero extends latent planning to large and continuous action spaces by sampling candidate actions. arXiv
  • Mastering Atari Games with Limited Data — EfficientZero improves MuZero-style latent planning with better sample efficiency in visual RL. arXiv Code
  • EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data — Extends EfficientZero-style model-based planning across visual, low-dimensional, discrete, and continuous domains. arXiv

Latent Reasoning in Language Models

  • Training Large Language Models to Reason in a Continuous Latent Space — COCONUT replaces some verbal chain-of-thought tokens with continuous hidden-state reasoning steps. arXiv
  • Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Studies recurrent-depth language models that spend extra inference compute by iterating hidden states rather than emitting more tokens. arXiv
  • Large Concept Models: Language Modeling in a Sentence Representation Space — Models sequences at a higher-level concept embedding scale instead of only next-token space. arXiv
  • Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space — Allocates computation between tokens and compressed concept latents through learned semantic boundaries. arXiv
  • Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models — Proposes hidden-state transition statistics for probing when reasoning-like computation is happening internally. arXiv

Latent Reasoning in VLMs

  • Latent Chain-of-Thought for Visual Reasoning — Treats LVLM reasoning as posterior inference and learns latent visual CoT with amortized variational inference rather than only supervised textual rationales. arXiv
  • Multimodal Latent Reasoning via Hierarchical Visual Cues Injection — HIVE injects global-to-regional visual cues into recurrent latent reasoning loops for grounded slow thinking in MLLMs. arXiv
  • Multimodal Latent Reasoning via Predictive Embeddings — Pearl uses JEPA-inspired predictive embedding alignment to internalize multimodal tool-use trajectories in latent space. arXiv
  • Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model — SCOLAR addresses information-gain collapse in long visual latent CoT by anchoring auxiliary visual tokens to the original visual space. arXiv
  • ReGuLaR: Relation-Grounded Latent Reasoning for Large Vision-Language Models — Grounds latent reasoning states in question-relevant objects and relations so LVLMs can reason beyond unstructured continuous traces. arXiv

Latent Action Models

  • Learning Latent Plans from Play — Play-LMP learns a latent plan space from unlabeled play data and reuses those plans for goal-conditioned control. arXiv Website
  • Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills — Learns goal-reaching skills from offline robot data by turning a functional world understanding into action. arXiv Website
  • Learning to Act without Actions — LAPO recovers latent actions from videos so policies and world models can be pretrained without labeled actions. arXiv
  • Latent Action Pretraining from Videos — LAPA quantizes frame-to-frame changes into latent actions for VLA pretraining from web and robot videos. arXiv
  • Latent Action Pretraining Through World Modeling — LAWM learns latent actions through world modeling to reduce dependence on teleoperation labels. arXiv
  • OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation — Introduces slot-addressable world-action latents so robot policies can bind instructions to persistent objects. arXiv
  • Light-WAM: Efficient World Action Models with State-Fusion Action Decoding — Makes WAM-style future-video supervision cheaper by decoding actions from compact fused latent states. arXiv
  • RepWAM: World Action Modeling with Representation Visual-Action Tokenizers — Learns semantic visual-action tokens for instruction-following robot world-action modeling. arXiv Code

Video World Models

  • Learning Interactive Real-World Simulators — UniSim combines heterogeneous interaction datasets to learn controllable real-world visual simulation. arXiv Website
  • Genie: Generative Interactive Environments — Learns action-controllable environments from unlabeled videos using a tokenizer, latent action model, and dynamics model. arXiv
  • Learning Interactive Real-Robot Action Simulators — IRASim generates realistic robot-action videos from initial frames and action trajectories for scalable robot learning. arXiv Website
  • Diffusion Models Are Real-Time Game Engines — GameNGen shows a neural video model can act as an interactive game simulator conditioned on actions. arXiv
  • Learning Generative Interactive Environments By Trained Agent Exploration — GenieRedux studies open implementations of Genie-style interactive environments using trained-agent exploration data. arXiv Code
  • Exploration-Driven Generative Interactive Environments — Uses world-model uncertainty to collect better interaction data for Genie-like simulators. arXiv Code

Embodied AI and Robotics

  • XIRL: Cross-embodiment Inverse Reinforcement Learning — Learns task-progress embeddings from cross-embodiment videos and uses latent distances as rewards. arXiv Website
  • RoboDreamer: Learning Compositional World Models for Robot Imagination — Factorizes language-conditioned video imagination to improve compositional robot planning. arXiv
  • ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance — Adds action-tree structure and visual guidance to robotic manipulation world-model videos. arXiv
  • VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model — Brings leakage-free JEPA-style future state prediction to VLA pretraining and robot manipulation. arXiv
  • PlayWorld: Learning Robot World Models from Autonomous Play — Trains robot video world simulators from unsupervised self-play and uses them for policy evaluation and RL. arXiv
  • Grounded World Model for Semantically Generalizable Planning — Scores imagined visuomotor futures in a vision-language latent space rather than against fixed goal images. arXiv
  • Toward Safe Autonomous Robotic Endovascular Interventions using World Models — Applies TD-MPC2-style world models to safety-critical robotic navigation under visual feedback. arXiv

Evaluation, Interpretability, and Benchmarks

  • PHYRE: A New Benchmark for Physical Reasoning — Tests whether agents can solve mechanics puzzles that require compact physical models and intervention planning. arXiv
  • CLEVRER: CoLlision Events for Video REpresentation and Reasoning — Evaluates causal, predictive, and counterfactual video reasoning over simple physical scenes. arXiv
  • IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments — Uses violation-of-expectation videos to test intuitive physics in richer scenes. arXiv
  • A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs — MVPBench pairs minimally changed videos to expose shortcut-based physical reasoning. arXiv
  • PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models — Tests whether VLMs and video world models obey mechanics principles such as center of mass, lever equilibrium, and Newtonian inertia. arXiv
  • WISER Benchmark — Introduced with Grounded World Models to test semantically generalizable visuomotor planning in latent language-vision space. arXiv

Open-Source Codebases

  • facebookresearch/ijepa — Official I-JEPA implementation for image-based joint-embedding predictive learning. Code arXiv
  • facebookresearch/jepa — Official Meta JEPA codebase for video feature prediction experiments. Code arXiv
  • facebookresearch/eb_jepa — Lightweight educational JEPA library covering image, video, and action-conditioned examples. Code arXiv
  • danijar/dreamerv3 — Reference DreamerV3 implementation for world-model reinforcement learning across domains. Code arXiv
  • nicklashansen/tdmpc2 — Official TD-MPC2 implementation for scalable latent model-predictive control. Code arXiv
  • YeWR/EfficientZero — Official EfficientZero code for sample-efficient MuZero-style visual planning. Code arXiv
  • werner-duvaud/muzero-general — Popular open-source MuZero implementation for studying learned hidden dynamics and MCTS. Code arXiv
  • insait-institute/GenieRedux — Open implementation and dataset work for Genie-style generative interactive environments. Code arXiv
  • wdrink/RepWAM — Code and weights location announced for representation-centric world-action modeling. Code arXiv

Related Awesome Lists

  • Awesome World Models — A broad world-model collection; use this repo as the narrower latent-reasoning and latent-control companion. List
  • Awesome Robotics Foundation Models — A robotics foundation-model reading list useful for embodied and VLA context. List

Contributing

Contributions are welcome when they keep the latent-world-model focus sharp. Please read CONTRIBUTING.md before opening a PR or use the paper suggestion issue template for candidates that need discussion.

Short version: explain why the latent representation matters, prefer primary links, avoid dumping generic embedding papers, and add one useful sentence per entry.

License

This list is released under CC0-1.0. Public-domain curation is the easiest latent space to share.

About

Curated Awesome List for latent world models, JEPA-style predictive architectures, latent reasoning, and latent planning/control

Topics

Resources

Contributing

Stars

3 stars

Watchers

1 watching

Forks

Contributors