Plausible
Preserve task-relevant state over time.
A decision-centered survey of embodied world models
From Plausible to Controllable to Actionable
A world model matters not only because it can imagine a convincing future, but because it preserves the right state, responds correctly to intervention, and improves an embodied agent’s decisions.
Preserve task-relevant state over time.
Predict how interventions change that state.
Turn prediction into measurable downstream gains.
The central question is no longer
“Can the model generate a realistic future?”
but
“What does that prediction enable an agent to do?”
What state must remain identifiable and measurable?
Does changing an action produce the corresponding effect?
Does prediction improve planning, learning, verification, or recovery?
02 / Capability framework
The levels describe progressively stronger claims, not mutually exclusive architectures. A method may demonstrate a higher-level use while leaving a lower-level property untested.
State consistency
A plausible model preserves task-relevant temporal, geometric, or physical structure across a rollout—beyond surface-level visual realism.
Evidence: held-out state errors, consistency under occlusion, and performance as a function of horizon.
Intervention fidelity
A controllable model predicts how a commanded intervention changes the future while preserving unrelated aspects of the scene.
Evidence: paired or branched interventions from shared initial states, including masked and shuffled-action controls.
Decision utility
An actionable model changes a downstream decision or update and produces a measured gain under matched data, controller, compute, and latency budgets.
Evidence: realized task utility against a matched baseline that removes world-model input.
03 / Grounding × improvement
The matrix connects what grounds a prediction—geometry, physics, or action—to the system component it informs. Select any cell to inspect the interface.
04 / Technical landscape
The survey traces technical progressions within each level, from compact predictive state to intervention-aware modeling and closed-loop use.
Plausible progression
Controllable progression
Actionable progression
Across embodiments
Contact, object state, force
Multi-agent futures, safety
Partial observability, geometry
Fast dynamics, stability
05 / Paper library
Indexed directly from citations and classifications in the survey body.
Showing 0 of 200 matching papers
All body classificationsTry a broader keyword or clear one of the classification filters.
06 / Open problems
The most consequential gaps appear at the boundaries between levels—especially when models face longer horizons, new policies, new embodiments, and real-time constraints.
Retain identity, geometry, and unresolved uncertainty across occlusion and long rollouts.
Separate intervention response from correlations inherited from demonstrations.
Preserve value, feasibility, and safety within real compute and control deadlines.
Make geometry, physics, and action modules agree when their predictions conflict.
Revalidate confidence after policy updates and prevent shared errors from circulating.
Report action interfaces, budgets, latency, evaluation seeds, and failure cases.
The evaluation shift
Progress toward actionability should be measured by whether grounded predictions improve closed-loop behavior under explicit data, compute, and latency budgets.
07 / Citation
If this survey supports your research, please cite it using the entry below.
@article{yao2026worldmodels,
title = {World Models for Embodied Intelligence:
From Plausible to Controllable to Actionable},
author = {Yao, Nanjie and Wang, Hao and Cheng, Chong and
Chen, Zhikang and Li, Wenzhe and others},
year = {2026}
}