
Do AIs Really Understand the World—or Only Imitate?
A plain-language take on a hard question behind AGI timelines: if today’s systems only copy patterns in data, scaling alone may never produce real understanding.
Many people in the industry talk as if AGI might arrive within a few years. That optimism is easy to feel when models write, draw, and plan with startling fluency. But a more basic question still sits underneath the hype: do the systems running in data centers actually understand the world—or are they staging a very convincing imitation of understanding?
This is not a word game for philosophers. It shapes the next bet for AI. If current systems already have real understanding, then keeping the current path—more data, more compute, larger models—may be enough. If they do not, and they mainly reproduce patterns found in human-made datasets, then size alone may never deliver general intelligence.
The familiar picture: a copy of the world inside the machine
Most modern AI rests on a common-sense story about cognition. Information enters through sensors. The system builds an internal representation—a kind of copy of the outside world. Then it reasons, decides, and finally acts. In that frame, perception comes first and is mostly passive: see the object, label it, then choose what to do.
Computer vision fits this picture neatly: train on millions of labeled images, extract features such as pointed ears and whiskers, and later output “cat.” Large language models fit it too: absorb vast text, form an internal map of language and world knowledge, then generate answers. Intelligence, in this view, is how accurate and complete that internal copy becomes.
Another view: understanding is what you can do
Sutton and Rafiee draw on enactive cognition, which pushes back hard on that copy-first story. Understanding is not mainly a static model sitting in memory. It is the ability to stay in useful contact with the world through action. Meaning is created in interaction, not waiting fully formed inside an object for a labeler to discover.
A chair makes the contrast easy. Representational thinking says you know it is a chair because it matches an internal “chair” template. The enactive view says you know it because you know you can sit on it, move it, stand on it to reach a high shelf, or use it as a temporary table. Without those possibilities for action, the word “chair” is thin.
Why reading about swimming is not swimming
Supervised and self-supervised learning mostly train on traces of other people’s experience: labeled photos, scraped text, recorded video. Those traces are valuable, but they are not the system’s own lived contact with the world. You can watch every swimming tutorial online and still not know buoyancy, choking on water, or how your limbs keep balance.
The same gap appears with everyday objects. A model that has seen every caption and photo of a cup still has never held one, drunk from one, or dropped one. It can talk about cups with remarkable confidence. That fluency can look like understanding while still being a polished form of imitation.
- A fixed training set is a freeze-frame of someone else’s contact with the world.
- Deploying a finished model assumes learning mostly ends before use; lived understanding does not.
- Impressive answers can still be pattern completion without the ability to intervene when the pattern breaks.
Why the question matters for the next decade
If imitation-plus-scale is enough, the industry’s current playbook stays central. If understanding requires agents that generate their own experience through ongoing interaction, then the bottleneck shifts: not only larger models, but systems that can act, receive feedback, and revise what the world means to them.
Sutton and Rafiee’s paper does not hand over a finished AGI recipe or a launch date. It reframes the destination: intelligence as a process of living contact with a world that is always larger than any internal model. For readers watching the AGI debate, that is the practical takeaway—fluency is real progress, but fluency alone may not settle whether the machine understands.