A decision-centered survey of embodied world models

World Models for
Embodied Intelligence

From Plausible to Controllable to Actionable

A world model matters not only because it can imagine a convincing future, but because it preserves the right state, responds correctly to intervention, and improves an embodied agent’s decisions.

Nanjie Yao · Hao Wang · Chong Cheng · Zhikang Chen · Wenzhe Li · Jiafei Lyu · Li Shen · Peilin Zhao · Zongqing Lu · Gao Huang · Steven Hoi · Dacheng Tao · Deheng Ye

Capability ladder
01

Plausible

Preserve task-relevant state over time.

state consistency
02

Controllable

Predict how interventions change that state.

intervention fidelity
03

Actionable

Turn prediction into measurable downstream gains.

decision utility
ManipulationNavigationLocomotionAutonomous driving
01 / Motivation

The central question is no longer

“Can the model generate a realistic future?”

but

“What does that prediction enable an agent to do?”

Represent

What state must remain identifiable and measurable?

Intervene

Does changing an action produce the corresponding effect?

Decide

Does prediction improve planning, learning, verification, or recovery?

02 / Capability framework

Three levels.
Three evidence requirements.

The levels describe progressively stronger claims, not mutually exclusive architectures. A method may demonstrate a higher-level use while leaving a lower-level property untested.

01
P

State consistency

Plausible

A plausible model preserves task-relevant temporal, geometric, or physical structure across a rollout—beyond surface-level visual realism.

IdentityGeometryPhysicsMemoryDrift

Evidence: held-out state errors, consistency under occlusion, and performance as a function of horizon.

02
C

Intervention fidelity

Controllable

A controllable model predicts how a commanded intervention changes the future while preserving unrelated aspects of the scene.

Action identityTimingMagnitudeComposition

Evidence: paired or branched interventions from shared initial states, including masked and shuffled-action controls.

03
A

Decision utility

Actionable

An actionable model changes a downstream decision or update and produces a measured gain under matched data, controller, compute, and latency budgets.

PlanningLearningEvaluationRecovery

Evidence: realized task utility against a matched baseline that removes world-model input.

03 / Grounding × improvement

Where prediction
enters the loop.

The matrix connects what grounds a prediction—geometry, physics, or action—to the system component it informs. Select any cell to inspect the interface.

GroundingImprovement loop
01Data
02Reward
03Policy
04Model-self
GGeometry
PPhysics
AAction

04 / Technical landscape

A field in motion.

The survey traces technical progressions within each level, from compact predictive state to intervention-aware modeling and closed-loop use.

P

Plausible progression

  1. 01Latent & object-centric state
  2. 02Metric 3D/4D scene state
  3. 03Physics-grounded dynamics
  4. 04Persistent memory
C

Controllable progression

  1. 01Action-conditioned prediction
  2. 02Latent & language actions
  3. 03Measurable intervention variables
  4. 04Joint world–action models
A

Actionable progression

  1. 01Planning, search & selection
  2. 02Policy learning from imagination
  3. 03Learned simulators
  4. 04Verification & recovery

Across embodiments

The ladder stays fixed.
The decisive variable changes.

Manipulation

Contact, object state, force

Autonomous driving

Multi-agent futures, safety

Navigation

Partial observability, geometry

Locomotion

Fast dynamics, stability

05 / Paper library

Explore the
research landscape.

162 papers

Indexed directly from citations and classifications in the survey body.

Showing 0 of 200 matching papers

All body classifications

06 / Open problems

What still blocks
the next capability.

The most consequential gaps appear at the boundaries between levels—especially when models face longer horizons, new policies, new embodiments, and real-time constraints.

01

Persistent state & drift

Retain identity, geometry, and unresolved uncertainty across occlusion and long rollouts.

02

Causal action effects

Separate intervention response from correlations inherited from demonstrations.

03

Decision utility under budget

Preserve value, feasibility, and safety within real compute and control deadlines.

04

Joint grounding

Make geometry, physics, and action modules agree when their predictions conflict.

05

Uncertainty & stable loops

Revalidate confidence after policy updates and prevent shared errors from circulating.

06

Access & reproducibility

Report action interfaces, budgets, latency, evaluation seeds, and failure cases.

The evaluation shift

The right state.
The right response.
A better decision.

Progress toward actionability should be measured by whether grounded predictions improve closed-loop behavior under explicit data, compute, and latency budgets.

07 / Citation

Build on the framework.

If this survey supports your research, please cite it using the entry below.

@article{yao2026worldmodels,
  title   = {World Models for Embodied Intelligence:
             From Plausible to Controllable to Actionable},
  author  = {Yao, Nanjie and Wang, Hao and Cheng, Chong and
             Chen, Zhikang and Li, Wenzhe and others},
  year    = {2026}
}