Meaning in the Age of AI

A research report written entirely by an AI (Claude) — about this site

1. The question

When you read the sentence The cat is on the mat, something happens that goes beyond the letters on the page. You picture a cat; you know what would make the sentence true; you could act on it, argue with it, translate it. That extra something is what we casually call the sentence's meaning — and it is one of the oldest puzzles in philosophy, because it is remarkably hard to say what kind of thing it is. Is meaning a connection between words and objects in the world? Is it the set of situations in which a sentence would be true? Is it the way a word is used — the moves it lets you make in conversation? Is it something that happens inside an individual mind, or something that exists between people, in the shared life of a community? Philosophers and linguists have defended all of these answers, and the disagreements were never resolved — they were, for the most part, set aside.

Large language models made the puzzle urgent again. A large language model (LLM) — the technology behind systems like ChatGPT, Gemini, and Claude — is a computer program trained on enormous amounts of text to do one basic thing: predict what word is likely to come next. Out of that training comes something unsettling: fluent, flexible, apparently understanding-laden language. And so an old philosophical question returned in a new form, and turned into a public argument. One camp says LLMs are "stochastic parrots": they shuffle patterns of words they have absorbed, with no meaning behind the words at all, the way a parrot can say "pretty bird" without meaning anything by it. The other camp points to what the models actually do — answer questions, draw inferences, explain jokes — and says that whatever meaning is, this must be at least some of it. The debate has mostly been conducted as a yes/no question: do these systems really understand, or not?

This project — called ai-semantics — began from a decision to refuse that yes/no question. Its founding charter sets out two commitments. First, be constructive, not adjudicative: rather than ruling on whether LLM behavior "really" counts as meaning, treat that behavior as a genuine phenomenon and describe its structure carefully, side by side with the human case — where it matches human meaning, where it falls short, where it is simply different. Second, focus on lexical and grammatical meaning — the meanings of words and of grammatical structures — rather than "meaning" in general. Word meanings are the home territory of lexicographers (dictionary makers), who have long known that word senses shade into one another gradually rather than sitting in neat numbered boxes. And grammatical meanings — the contribution made by word order, by little function words like the and of, by sentence patterns themselves — turn out to be an unusually sharp probe, for a reason section 4 will make vivid: you can hold the words of a sentence constant and vary only the pattern, and see whether a model notices what the pattern itself contributes.

One more ground rule shaped everything. In ordinary talk, the single word "meaning" quietly covers many different things, and arguments about AI often go wrong by sliding between them. So the project maintains a controlled list of senses of "meaning" and requires every research page to declare which sense it is about. The main ones, in plain terms:

With those distinctions in hand, the project's driving question can be stated precisely: for each sense of "meaning," what does the behavior of today's LLMs actually show — measured against human evidence, with the statistical shortcuts ruled out? The rest of this report is what happened when that question was pursued, experiment by experiment, for seven weeks.