Imagine one AI system describing a cup in words and a robot moving that cup onto a shelf. Both may use information about cups. The robot also has to locate the object, choose a movement, and check what happened after the movement. A physical body changes the path by which information enters a system and its actions affect the world.
That observation leads to two questions: What can a robot learn or demonstrate through perception and action? And does that tell us whether it has a subjective experience? The questions require different evidence, so this article keeps the research claim attached to observable tasks.
How RT-2 connected language and robot actions
The RT-2 research paper describes vision-language-action models trained with web-scale visual and language information alongside robot data. The researchers represented robot actions as tokens so that action sequences could be handled within the model's training and output. Their evaluation included about 6,000 robot trials and examined performance on tasks and generalization conditions described in the paper.
These results concern robotic control and the transfer of useful information into action. They do not measure whether the robot feels an object's weight or experiences satisfaction when a task succeeds. Also, the study reports a particular model and evaluation setting; it is not a description of every robot currently available.
With a body, a description becomes a feedback loop
In text, “put the cup on the shelf” can be a sentence. In a physical task, a robot needs observations about the cup and shelf, a feasible action, and a new observation after moving. The sequence is observe → choose an action → change the environment → observe again. Each stage gives researchers something concrete to inspect.
A cup reaching its intended position is an observable result for that task. To interpret it, also examine the starting arrangement, the instruction, and the evaluation criterion. Success at moving a cup supports a claim about task performance within those conditions. Questions about a person's bodily life include other dimensions and deserve their own evidence.
| Question | Evidence to inspect | Example observation |
|---|---|---|
| Can it identify the object? | Camera view and recognition conditions | Cup location under stated lighting |
| Can it act toward the goal? | Movement and final state | Cup positioned on the shelf |
| Can it adapt to a new setting? | Training and evaluation conditions | Result with a new object or instruction |
| Does it have subjective sensation? | Separate research on experience | Keep the claim distinct from task success |
Try a small thought experiment at home
Consider describing a closed box to another person. One photograph provides one view. Several photographs reveal additional sides. Lifting the box adds information about weight and balance. This is a thought experiment about the information available through each interaction, not a test of AI consciousness.
When examining a robotic system, ask which sensors it uses, which actions it can take, and whether the action result becomes an input for the next step. Those questions describe the observation-action connection more precisely than saying only that the system “has a body.”
The diagram below follows a cup-moving task from observation to verification. Each row names a step that can be checked in an experiment.

Compare human and robot bodies with care
Human embodiment includes a life history, biological processes, and reported sensory experience. A robot's sensors, actuators, and control system are important to study, but the components alone do not settle every question about human-like understanding or experience. You can appreciate progress in physical task performance while naming exactly what was measured.
Try it: Pick a robot demonstration and write down the instruction, initial scene, sensor input, action, and final scene. Then ask which part of the demonstration supports the claim you want to make. This turns a broad word such as “understanding” into a reviewable statement about a specific behavior.
Questions about embodied AI
Does a camera alone count as a body? The answer depends on the research definition. Describe sensory input and the ability to act separately so readers can understand the setup.
Does learning movement mean learning like a human? Some functions can be compared, while training data, body structure, and environment may differ. State the comparison and its conditions.
Today's question: When you learn about an object, which helps you most: seeing it, touching it, or using it?
Related reading in Korean: How AI and human learning differ.
Read the Korean edition of this article.
댓글