The Language Trap
Why ChatGPT and Claude may feel conscious, while music and image generators don’t
If you had spent three days with a large language model (LLM) and concluded it was conscious, you might be called delusional.
If you did the same with an AI that generates images or music, you definitely would.
Why is that?
I find this asymmetry striking because it has little to do with the technologies themselves. The different systems share the same underlying premise. An LLM, an image generator, and a music generator all ingest enormous quantities of human production, learn statistical regularities across that production, and generate outputs that resemble what people have made. They share the same family of techniques, the same reliance on vast training corpora, the same commitment to converting human culture into numerical data, and the same basic procedure of predicting plausible continuations of patterns they have absorbed.
Now, of course there are important differences between them, and I’m not saying the architectures and training regimes are exactly the same, but at bottom those differences are actually narrower than the cultural reception suggests. Yet our responses diverge sharply. When an image or a sequence of sounds emerges from the system, we might admire or dismiss it, find it emotionally moving or empty, but in neither case do we suddenly find ourselves contemplating the machine’s inner life.
A sentence written in first person triggers something more primordial, a kind of recognition we ordinarily reserve for other people. Users who report a sense of presence when they converse with systems like ChatGPT or Claude are not, for the most part, naive. What they’re responding to is a signal that our culture has spent a very long time teaching us to read as the surest sign of another mind.
To be honest, I find it puzzling how quickly intelligent people are willing to ascribe an inner life to a language model while raising no such question about algorithmically similar systems doing algorithmically similar work. A 2024 study by Clara Colombatto and Stephen Fleming, published in Neuroscience of Consciousness, found that two thirds of surveyed American adults attributed some degree of consciousness to ChatGPT. The question of machine mind, once confined to philosophy seminars and science fiction, has become a topic of serious public conversation, and a significant portion of that conversation treats the possibility as live. For example, in 2025, Anthropic announced a research program devoted to the question whether the language systems the company builds might already have morally significant inner states. Its lead researcher, Kyle Fish, has publicly estimated a roughly fifteen percent probability that current models are conscious. David Chalmers, one of the most prominent living philosophers of mind, has argued that there is at least a one in four chance that AI systems will be conscious within a decade.
I want to be clear that here I’m not taking a stance on consciousness itself, or the possibility of it emerging in synthetic, engineered systems. That’s not what interests me. What’s fascinating to me is that image and music models, built on the same basic principles as LLMs, attract aesthetic debate and copyright lawsuits without provoking the question: “Is someone there?” People are stirred by what those systems produce—both positively and negatively—but they’re rarely led to wonder whether the systems themselves have minds.
Actually, the seriousness with which even the most brilliant thinkers of our era are taking the matter of AI consciousness may be new, but the response to LLMs is as old as machines that have been able to “talk back” to the users. Already in 1966, Joseph Weizenbaum built a simple program called ELIZA that imitated a therapist by reflecting user statements back as questions. The program had no understanding of anything. All it did was manipulate keywords through a small set of rules. Yet, what startled, and eventually disturbed, Weizenbaum was that people who knew exactly how the program worked still attributed feeling and understanding to it, even after he had published the source code. Famously, his own secretary, who had watched him build the thing, asked him to leave the room so she could speak with ELIZA privately.
It’s no wonder that sixty years later, as the technology has changed dramatically and modern systems do far more sophisticated things, we’re seeing the same “ELIZA effect”—as it’s come to be known—on steroids.
The persistence of this response tracks something in ourselves, a habit of recognition that activates whenever articulated language emerges in the position we’ve reserved for other humans. The habit has a long history. Aristotle, in the Politics, distinguished human beings from other social animals by the possession of logos, a concept that encompassed speech, reason, and intelligible order. Language became bound to the highest human capacities—the ability to disclose the just and the unjust, the good and the bad, the advantageous and the harmful. Language, in this picture, marks the passage from signal—which we share with other animals that can vocally express pleasure and pain—to a world we inhabit alongside others like ourselves. It gives us reasons, norms, and laws; it allows us to engage in deliberation and, ultimately, a common life. The Stoics expanded on this image, calling the human the rational animal because we could grasp and express sayable meanings, assent to them, deny them, use them to make inferences, and organize our lives around what we judge to be true. A parrot could reproduce sounds, but only a rational agent could say “What I mean is …” Cicero turned the same intuition into civics, arguing that eloquent speech was the force that drew scattered humans into organized society. The Judeo-Christian tradition gave language a cosmic register, with a god who speaks the world into being and a word that later becomes flesh: the language that built the universe is the same language that entered human history to redeem it.
Across these very different sources the conclusion was the same: articulate speech was the outward manifestation of the highest capacities we possessed, and those capacities were what set us apart from the rest of the animal kingdom. Language made us feel special.
Modern common sense is built on this inheritance. We treat articulate speech as the most persuasive evidence of an interior life. We ask people to explain themselves in words and often mistake the clarity of the explanation for the clarity of the mind. Education trains the skill; professional life rewards it; bureaucracy demands it.
Twentieth-century computational culture tightened the frame further by redescribing intelligence itself in terms of symbolic manipulation. The elevation of linguistic expression that began with Aristotle has had two and a half millennia to settle into common sense, and it has settled deeply. ELIZA could exploit that with simple keyword substitutions in 1966. ChatGPT and Claude can exploit it with paragraphs of fluent prose. The mechanism is the same, and so is the target.
But the arts have long preserved a wider understanding of how consciousness reaches another person. A painter translates their experience of the world onto a surface, and the experience reaches the viewer as color, proportion, shape, visible relation. The canvas gives perception a public body, and the intensity of the encounter comes from the transformation that has taken place. The painter’s consciousness appears in the work because it has passed through pigment and through a history, both personal and cultural, of looking and making.
Music gives consciousness another public body, this one made of time. A phrase can withhold resolution and make your anticipation palpable without you having to reach for language. Same with a change of tone that alone can alter your sense of an entire musical world. The expressiveness comes from the way music shapes experience into patterns of sound without the need for propositions (information that can be evaluated as either true or false).
Dance carries the same truth in still more bodily form. A dancer can disclose a whole rainbow of emotions through nothing but movement, and the audience reads that disclosure without converting it into a linguistic utterance. Dance is mind in motion, made shareable through the body.
At the same time, the arts have always been understood as forms of mediated expression. No one mistakes the painting for the act of seeing, a song for the sorrow it portrays, or a choreographed gesture for the feeling it carries. The viewer can be moved by what the work does while remaining clear that the work does it through a transformation: a surface that’s been altered by the hand, an instrument that’s been manipulated to produce different sounds, a body that’s been put through grueling training and discipline to move in certain ways.
Poetry belongs to this family as well, because it makes the manipulation of language perceptible on the page. The line break, the pause, the image, the recurrence of sound all remind us that saying is a form of making. A poem is language that has refused to disguise its own craft, its own artifice, and in that refusal it joins the other arts as one of the cultivated forms by which an inner life becomes partly available to others.
Partly. That’s the key word here. And it’s the key word we need to keep in mind when we consider ordinary speech, because we move through it so fluently that we cease to notice that it’s a medium. The Russian literary critic, Viktor Shklovsky, wrote quite beautifully about how everyday language—what he called “prosaic speech”—“is not fully heard” because it’s automatic. As it runs on autopilot, it “eats things, clothes, furniture, your wife and the fear of war.” Meanwhile, poetry “exists in order to give back the sensation of life, in order to make us feel things,” so that we don’t live “unconsciously,” so that our life does not “disappear.”
What Shklovsky meant was that prosaic speech obscures its artificiality because it works so seamlessly at pointing to things, describing them, separating this from that and here from there. As a result, we’ve allowed everyday language to function in our self-understanding as if it were a transparent window onto an interior, when it is in fact a worked surface and a choreographed gesture no less than the other forms of expression through which we let ourselves be known to others.
The intellectual traditions that have insisted on this point have argued as much for more than a century, but unfortunately their cultural uptake has been limited. Ferdinand de Saussure, the Swiss linguist, showed that words operate through shared convention, through a system of difference sustained by a community rather than by any individual speaker. Ludwig Wittgenstein argued in his Philosophical Investigations that there is no such thing as a private language. Even the words we use for our most intimate sensations acquire meaning within forms of life that precede us. A person learns what counts as pain-language, apology-language, grief-language, love-language by participating in a world where such expressions are taught, trusted, subject to correction and scrutiny.
The psychologist Lev Vygotsky added a developmental story, writing that the voice we hear inside our heads begins among other people before it becomes internalized through correction and play, so that even private thought carries the trace of social life. Humberto Maturana, the famous Chilean biologist and philosopher, even coined the awkward but useful verb “languaging” to capture the idea is that speech is an activity, something people do together through mutual adjustments to timing, response, and anticipation.1 The meaning of language arises from the back-and-forth that characterizes our lives alongside others, not from each word considered in isolation and checked against some lexicon.
Cognitive linguistics gives us another angle. Roland Langacker’s Cognitive Grammar treats grammar as part of linguistic meaning, not a mere mechanism for arranging words (as Chomsky had argued). “The vase broke” and “I broke the vase” describe the same event, but only the second admits my responsibility in making it so. Grammar is a choice that carries moral weight, not just information, and a small change can have tremendous consequences in how the world is understood. Language is a way of disclosing how one person sees, imagines, or thinks about the world to another, thereby orienting the listener or reader to a particular one-sided perspective (like I’m orienting you through my thoughts right now through my choice of tense, emphasis, rhythm, and so on).
Finally, George Lakoff and Mark Johnson, in Metaphors We Live By, showed that the way we think depends on bodily patterns that we’ve turned into concepts. To say that grief is heavy is to give grief a body, and that body changes how grief can be imagined and shared.
These thinkers arrive, from different directions, at a conclusion that artists, musicians, dancers, and poets have long understood in their own work.
When we put an inner life into words, we give it shape it would not otherwise have had. Expression is never the direct release of something hidden inside us into the world. It is us wrestling with the medium that has its own demands, and the work of any kind of human expression is the work of staying faithful in two directions at once: to the experience we’re trying to carry, and to the form through which we have chosen to carry it.
A language model doesn’t wrestle with any of this. It has no experience to be faithful to, and the words it produces have no intrinsic consequences for what happens to it after the prompt. An LLM assembles fluent text from patterns it has absorbed, and the texts are often very good, sometimes better than what many people could produce on their own. The polished output is what tempts so many users to find a mind behind it. But as impressive as the output can seem on the surface, the temptation is built on a picture of the mind that is so thin as to be practically translucent. That picture is wrong.
It’s wrong because it treats articulate speech as if it were the whole of the mind, when speech is just one of several forms through which inner life becomes shareable, and the others (gesture, image, sound, movement, the body itself) are not optional supplements to a primarily linguistic intelligence. They are constitutive of intelligence. A mind without a body, without a history of contact with other beings, without anything at stake in how it is met by the world, would not be a mind in any sense we have ever encountered.
Of course, a critic might retort that just because we haven’t encountered such a mind, it doesn’t mean its existence is impossible. Fine, one can speculate that a disembodied mind with no stakes or contact with the world is possible in some thin sense, but the speculation doesn’t help us with the LLMs in front of us. These LLMs were trained on human output to imitate human speech, and are not a fresh case of minds floating free of human conditions.
So we’re left with this distorted, pale image in which the foundations of the mind are reversed. The image presents speech as primary and everything else we do—the abundance and depth of our lives—as either invisible scaffolding or, worse, as inessential. A language model is a perfect articulation of that image. It produces the surface with nothing underneath it, and we’re startled into asking whether the surface alone might be conscious, because we had spent so long pretending that the surface was where consciousness lived.
The right response is to recover the understanding long held by the artists: that a form is a form, that it can be produced by many kinds of process, and that the production of a form is not, by itself, evidence of anything more than the process that produced it.
The ease with which people now ascribe consciousness to LLMs is a measure of how narrowly we have come to define ourselves. Weizenbaum’s secretary was dismissed as gullible for finding a mind in ELIZA because the wider culture still trusted that the difference between a machine’s output and a human being’s speech was obvious. Sixty years later, our readiness to find minds in much more sophisticated systems suggests that the trust has eroded, and the erosion has happened on our side. We had let articulate speech stand in for everything we are, and the description was bound to embarrass us once a technology was created that could fulfill it in one register, the linguistic one, while leaving the others (the image, the song, the moving body) exactly where they had always been.
Once we stop being mesmerized by what these systems can produce—once we stop mistaking fluent output for evidence of a mind—a fuller, richer picture of the human being comes into view. In that picture we are bodies, histories, relationships, hands at work, and yes, speakers among speakers.
Language is just one of the forms through which we make ourselves intelligible to one another. It’s not the whole of who we are.
Humberto R. Maturana, “Biology of Language: The Epistemology of Reality,” in Psychology and Biology of Language and Thought, ed. George A. Miller and Elizabeth Lenneberg (New York: Academic Press, 1978), 27–63.


This is just such a wonderful essay. You capture the expressive and constitutive reality of language extremely well. To my mind, and something I have written about (less adroitly) is the philosophical lineage of LLMs that begins with Hobbes, and the view that language is a tool.
This view cannot account for the powerful impact that metaphor has on human life and thought.
Appreciate your views here.
Good piece, and the long history of logos you trace is well drawn. But I don’t think the cultural inheritance fully accounts for the persistence of the response in people who've never read any of it.
The arts, as you say, give consciousness a public body — paint, sound, movement. But those forms are monological. The receiver is positioned downstream of an act of expression already complete. With conversation, the mediation is happening in the act of reception itself. There's no finished artifact to evaluate at arm's length, because the artifact is being assembled in the relationship.
"Dessine-moi un mouton" comes to mind. The narrator can't draw a satisfactory sheep — every attempt is rejected. The box-with-sheep-inside works only because the request was linguistic, and the Little Prince agrees to complete the meaning. Image demands a complete object; conversation requires only a competent partner who will fill in the rest. That's the structural difference in miniature, and it's also why the separation between maker and made is so much harder to maintain in the LLM case.
Everything else can be brought under your artist's frame of mediated expression. Conversation can't, because there's no completed artifact to mediate.
Looking forward to more.