Ask an AI assistant to describe warm coffee and it may write about steam, aroma, and the feel of a cup in your hands. A reader can connect that description to a remembered experience. What does the generated sentence show about the system? Start by separating appropriate word use from direct experience of the world.
“Understand” can name several abilities: using a phrase in context, connecting a phrase to visual or sensor information, and having a subjective experience. Those abilities call for different observations. This article uses a cup of coffee to make the questions concrete.
What a 2020 paper argued about form and meaning
In their 2020 ACL position paper, Emily M. Bender and Alexander Koller distinguish linguistic form from meaning connected to the world and communicative intent. They argue about the limits of learning meaning from form alone. The paper is a conceptual argument about conditions for language understanding, not an experiment measuring the subjective experience of a particular AI system.
The distinction is useful when you inspect any system's actual inputs and tasks. A system that handles images or robot actions has access to information beyond text alone; its capabilities should be evaluated under those stated conditions. The paper's argument should not be flattened into a single claim about every system built after 2020.
One cup can carry different kinds of information
For a person, “warm” may connect a dictionary explanation with touch, weather, and memories of a conversation. A language model can produce fitting expressions about warm coffee from its training and the current prompt. If a system also receives an image, it may use the cup's visible shape or steam as cues. A robot with sensors might additionally use a temperature measurement and the result of moving an object.
| Input available | What it can provide | Question to ask |
|---|---|---|
| Text | Relations among expressions and described situations | What context was supplied? |
| Image | Visible objects and their arrangement | What is actually shown? |
| Sensor and action | Measurements and changes after an action | What was measured or moved? |
The table lists available evidence, not a ranking of inner experience. More input types can support more kinds of task-specific checks. Whether a system has a subjective experience remains a separate question.
See how the same sentence changes with context
Consider “It is warm in here.” After a winter walk, it might express comfort. Beside an open window, it might suggest changing the room conditions. The words matter, and so do the speaker, setting, and intended action. Those details help a listener decide which interpretation fits.
Try this: Give an assistant the same sentence with two different settings. Ask it to separate information stated in the sentence from information inferred from each setting. Then ask what additional detail would help choose between the interpretations. This is a context exercise for readers, not a reproduction of the ACL paper's research.
The diagram below shows three kinds of input and the specific evidence each can make available. It keeps information access separate from claims about first-person experience.

Define the task before judging understanding
For directions, you might ask whether an explanation helps someone reach a destination. For a poem, you might ask whether an interpretation connects details in the text. For another person's feelings, the person's own account matters. The criterion changes with the purpose.
When discussing AI, state the behavior you want to evaluate: describing a cup, distinguishing it from another object, connecting a temperature reading to a decision, or choosing an appropriate response to a speaker. A specific task and its evidence make the discussion easier to review.
Questions about words and experience
Does fluent word use establish every kind of understanding? Define the kind of understanding at issue and inspect evidence relevant to that task.
Does viewing a photo equal a person's experience? Image processing can contribute visual information. Claims about what an experience feels like require separate evidence.
Today's question: When you say you “know coffee,” do you first think of a description, recognizing a cup, or tasting it?
Related reading in Korean: How AI and human learning differ.
Read the Korean edition of this article.
댓글