HumanoidGPT · Part of The Humanoid Group
World Models for Robots
A world model is a learned prediction of what happens next: if the robot does this, the scene will look like that. This guide explains how such models are built and used for robot control, and what they cannot yet do.
By Arjun Rao · Updated
What a world model predicts
A world model takes the current state of the world, as seen through the robot's sensors, together with a proposed action, and predicts the state that follows. Some predict compact internal codes that only the robot's software can read; others predict full video frames that a person can watch. Either way, the model is learned from data rather than written by engineers, which sets it apart from a conventional physics simulator. Because it can be rolled forward many steps, the robot can compare several candidate actions in imagination, discard those that end badly and commit to the most promising, much as a person rehearses a move before making it.
Learning control inside the model
The Dreamer family of algorithms shows the idea at work. With DreamerV3, Hafner and colleagues trained agents that learn a model of their environment and improve by imagining future scenarios, and they report that a single configuration outperformed specialised methods across more than 150 tasks and became the first algorithm to collect diamonds in Minecraft from scratch without human data. DayDreamer, from UC Berkeley, applied the same approach to physical machines with no simulator at all, and trained a quadruped to roll off its back, stand up and walk from scratch in about an hour. The appeal for robotics is data efficiency, since imagined practice is cheap and real practice is not.
Video world models as simulators
A newer branch trains large video models on broad collections of footage so they can generate what a scene would look like after a given action or instruction. UniSim, from researchers at Google DeepMind, UC Berkeley and MIT, combined image, robot and navigation datasets into one interactive simulator that responds both to commands such as 'open the drawer' and to low-level motor actions. The team trained high-level and low-level robot policies purely inside it and report that they transferred zero-shot to real robots. Video world models promise practice in settings that are hard to build in a physics engine, such as cluttered homes, but their physics can be no better than the footage they learned from.
Limits and questions for buyers
World model errors compound. A small mistake in one predicted frame becomes a larger one several frames later, so long plans built in imagination can drift away from reality. Video models in particular can produce scenes that look right but break physics, with objects passing through each other or changing shape. For that reason, a sensible design checks plans against live sensor data and keeps safety limits outside the learned model. If a maker says its humanoid uses a world model, ask what it is used for, whether planning, training in imagination, testing policies before release or generating synthetic data, and how its predictions are checked against the real robot.
Sources and further reading
- Mastering Diverse Domains through World Models — Google DeepMind and University of Toronto (arXiv).
DreamerV3 learns a model of its environment and improves by imagining outcomes, across 150+ tasks with one configuration. - DayDreamer: World Models for Physical Robot Learning — UC Berkeley (arXiv).
World-model learning on four real robots with no simulator, including a quadruped that learned to walk within an hour. - Learning Interactive Real-World Simulators — Google DeepMind, UC Berkeley and MIT (arXiv).
UniSim: a video-based simulator learned from mixed datasets, used to train robot policies that then ran on real robots.
Common questions
Are world models used in humanoid robots today?
Mostly in research and in makers' development pipelines, for training and testing policies, rather than as the only controller on a working robot. Ask any maker whether a world model runs on the robot itself or only during training and evaluation.
What data is needed to train a robot world model?
Large amounts of video or sensor recordings paired with the actions that produced them. Footage of people and general video help with appearance, while robot logs supply the link between commands and consequences, which is why data from deployed robots is valued so highly.
People also search for robot world model, learned simulator, video prediction robotics, DreamerV3, DayDreamer and UniSim.