From digital fluency to physical consequences
The generative AI boom has largely taken place in digital settings. Language models can write, summarise, program and converse because their training material is overwhelmingly made of text, images, audio and software available online. A machine that acts in a warehouse, laboratory, hospital or home faces a different challenge: it must connect perception to movement while dealing with objects that have weight, shape, friction, fragility and unpredictable behaviour.
This is the ambition behind physical AI, embodied AI and so-called world models. The terms are sometimes used loosely, but the common objective is clear. Rather than merely describing a scene, an AI system should build a useful internal representation of it, anticipate what may happen next and select actions that work in the real world.
That would be a meaningful change in the role of AI. A mistaken chatbot answer can waste time or mislead a user; an erroneous decision by a robot can damage equipment, disrupt production or injure someone. The promise is therefore not simply more intelligent machines, but machines whose intelligence can be translated into reliable, bounded physical action.
What a world model is meant to do
A world model is not necessarily a complete, scientifically exact simulation of reality. In practical terms, it is a model that captures enough of a setting’s state and dynamics to support prediction and planning. It may estimate where objects are, which surfaces can support a load, whether an item is likely to slip, or how a scene will change if a robot reaches, pushes or lifts.
This approach differs from a conventional vision system that identifies objects in a single image. Recognition can tell a robot that a cup is on a table. Acting safely requires additional judgements: whether the cup is full, whether its handle is accessible, whether another person is moving nearby and how much force is appropriate.
Recent research reflects a shift towards learning these capabilities from video and interaction data. Meta’s V-JEPA 2 research, for example, used large-scale video learning and subsequent robot data to demonstrate planning for pick-and-place tasks in unfamiliar laboratory settings. Google DeepMind has also developed models that combine visual input, natural-language instructions and robot actions, as well as a reasoning-oriented system intended to support spatial interpretation and task planning.
The important qualification is that demonstrations in controlled environments do not amount to general physical understanding. A model may perform well when objects, lighting, camera positions and task goals resemble its training and evaluation conditions. Real environments are messier: tools are misplaced, packages are deformed, floors are wet and people behave in ways no training set can exhaustively capture.
Why the potential is so large
If world models become robust, their first major impact is likely to be in environments that are structured enough to measure and monitor, but variable enough that fixed automation is expensive. Manufacturing, logistics, agriculture, inspection and laboratory work are obvious candidates.
Today, industrial automation is often highly capable but narrowly engineered. Machines excel at repeated motions in carefully designed workcells. A more adaptable system could potentially be taught new tasks through demonstrations and high-level instructions, while using visual and spatial reasoning to handle routine variation. This could shorten deployment cycles for smaller production runs and make automation viable for tasks that have not justified extensive custom programming.
Simulation is another possible multiplier. Training a robot exclusively through real-world trial and error is slow, costly and sometimes unsafe. World models could generate or enrich virtual scenarios for rehearsal, including uncommon but consequential failures. They could also help engineers test whether a planned action remains safe when the environment changes.
The same principle could matter beyond robots with arms. Autonomous vehicles, drones, assistive devices and scientific instruments all need to make decisions based on changing physical conditions. Better predictive models could improve planning and reduce the amount of real-world data needed for every new environment. In scientific work, systems that model physical processes may help propose experiments or identify promising designs, though their outputs would still require experimental validation.
The data problem is more difficult than it looks
The internet supplied language models with an enormous corpus of human expression. There is no equivalent public record of high-quality physical interaction. A video can reveal that a person opens a drawer, but it may not show the force applied, the resistance of the rail, the grip used or the corrective movements made after a slight snag.
Robotics therefore needs information from cameras, depth sensors, tactile sensors, force measurements and a machine’s own joint positions. It also needs data that pairs observations with actions and outcomes. Gathering such data at scale can be labour-intensive and risks embedding narrow assumptions about particular equipment, facilities or users.
Synthetic data and physics simulators can ease the shortage, but they introduce a persistent sim-to-real problem. A virtual environment can approximate gravity and rigid objects well while still missing the behaviour of cables, cloth, transparent surfaces, fluids, clutter or worn machinery. Models trained in simulation must be tested against the differences that matter in deployment, not merely against polished benchmark tasks.
Reliability will determine whether the technology changes work
A credible physical AI system needs more than impressive task videos. It needs clear operating limits, detection of uncertainty, robust fallbacks and a way for people to intervene. In many settings, the safest response to ambiguity will be to stop, ask for help or hand control back to a conventional system.
Evaluation must also move beyond average success rates. Operators and regulators will need to know how a system behaves when sensors fail, objects are unusual, instructions conflict or a person unexpectedly enters its workspace. Cybersecurity and privacy become more important when cameras, microphones and connected robots operate in workplaces and homes.
This is why the near-term story is likely to be augmentation rather than a sudden replacement of human labour. The most valuable systems may combine learned perception and planning with traditional controls, physical safeguards and tightly defined roles. A robot that reliably handles a limited range of tedious, hazardous or precision tasks can have substantial economic value without being a general-purpose machine.
A second AI transformation, with harder constraints
AI that understands enough physical reality to act safely could be as consequential as AI that understands language, because it would bring computation into contact with production, transport, care and infrastructure. Yet the comparison should not obscure the difference. The physical world does not forgive plausible but wrong outputs.
The decisive breakthroughs may therefore be unglamorous: better sensors, richer interaction data, more realistic simulation, rigorous safety cases and operational standards that reveal when a model should not act. World models could change how machines are trained and deployed. Whether they change the world again will depend on proving that their predictions remain dependable when reality refuses to follow the script.
Sources
- Will a new kind of AI that understands physical reality change the world again? — New Scientist
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Meta AI
- Gemini Robotics brings AI into the physical world — Google DeepMind
- Physical AI and Data Generation for Robotics — National Institute of Standards and Technology



