5. Putting it together: what has been learned
The map, not the verdict
The project's flagship synthesis is the shadow-depth table: one row per probed phenomenon, each
row recording the measured residual over its named control, with its error bars, its human anchor
(or its internal-contrast-only label), and its caveats. Six rows are "shadow-beaters" backed by
promoted claims — the comparative correlative, the dative, genitive, and particle-placement
alternations, the AANN acceptability gradient, and word-sense gradience — and around them sit the
honestly labeled corners: the presupposition line that beat one control but not convincingly, the
falsified antonymy prediction, the grammar-difficulty gradient that is real within models but
whose human-likeness could not be promoted, the grounding nulls, the deflationary relational
results.
The table below is a simplified rendering of that flagship object — the full version, with
confidence intervals, exact controls, and every fence, is the wiki page
theory/shadow-depth-table-v4. Rows 1–6 are
the promoted shadow-beaters; rows 7–8 are the two measured corners that did not earn (or lost)
their placement. Positive shifts are on the 0-to-1 preference scale of section 4.2 (so +0.32 ≈ 32
points of 100); ρ is rank correlation with the human gradient.
| # | Phenomenon | Level | The shadow it had to beat | What survived, per model | Human anchor | Standing |
|---|---|---|---|---|---|---|
| 1 | Comparative correlative — the covariation reading | pattern | same-word controls without the construction | assertion gap ≈ 87 pp in all three models (CI lower bound ≈ 78); German +93 / +88 / +88, Japanese +94 / +84 / +96 pp (claude / gemini / gpt) | none exists for the inference itself (within-model contrast; small human answer-key check 93–100%) | promoted claim; replicated; three languages |
| 2 | Dative alternation — discourse givenness | pattern | byte-identical sentence pairs; only the context varies | preference shift claude +0.32, gemini +0.52, gpt +0.06 — all clear zero at N = 100 | human production direction (Bresnan corpus data) | promoted claim, 3/3 |
| 3 | AANN acceptability gradient | pattern | word frequency, statistically removed | correlation with the human gradient after the partial: claude 0.69, gpt 0.66, gemini 0.74; replicated across dates | human acceptability ratings (Mahowald) | promoted claim, 3/3; temporal-noun stratum fails |
| 4 | Word-sense gradience | word | topic similarity of the two sentences, statistically removed | correlation with human medians after the partial: claude 0.52, gpt 0.50, gemini 0.73 (raw 0.60–0.80); replicated on fresh pairs | human relatedness ratings (DWUG), at/above annotator agreement | promoted claim, 3/3 |
| 5 | Genitive alternation — possessor animacy | pattern | invented-word (nonce) possessors, which have no usage statistics | animacy shift claude +0.15, gemini +0.17, gpt +0.14; the nonce firewall survives in all three models, twice | human direction (Dubois et al. 2023) | promoted claim, 3/3, direction-only |
| 6 | Particle placement — object givenness | pattern | byte-identical orders plus the "new-mentioned" context control | firewall shift claude +0.037, gemini +0.055; gpt +0.007, indistinguishable from zero | human direction (Kim et al. 2016 / Gries 1999) | promoted claim, 2/3; gpt a persistent shadow |
| 7 | Presupposition — projection and accommodation | pattern | word-form doppelgänger control | margin over the control claude +0.78, gpt +0.47, gemini +0.94 — but keyed to the trigger word-forms, so a surface-cue reading survives | none (internal-contrast-only) | unpromoted — the "under-licensed middle" |
| 8 | Lexical relation recovery — antonymy and kin | word | corpus contrastive-frame statistics | the saturation prediction was falsified: antonymy's residual is among the largest (+0.61–0.67 hit-rate), and recovery does not track cue strength (ρ ≈ −0.09) | none (internal-contrast-only) | registered bet lost; the corner moved |
Two structural lessons emerged from building that table, and they are the project's most original theoretical contributions.
First: what organizes the data is not words-versus-grammar but shadow-depth. The project began with a continuum in mind, from word meaning at one end to grammatical meaning at the other, and expected the interesting boundary to lie somewhere along it. It does not. Both ends of the continuum contain phenomena that beat the distributional shadow (sense gradience at the word end, the comparative correlative at the grammar end) and phenomena that stay inside it (the presupposition corner; and, before the falsification, antonymy was expected to be one). The informative axis is how much daylight exists between a phenomenon and its statistical footprint — how deep its shadow runs — and that axis cuts across the word/grammar divide. This sounds abstract, but it has a practical edge: it predicts which impressive-looking model behaviors are actually informative. A model acing antonyms tells you almost nothing (the shadow is deep there — though see section 4.7 for the twist); a model shifting word order for a discourse reason on byte-identical sentences tells you something real.
Second: agreement in direction, wild variation in size. Across nearly every positive finding, the three panel models agree on which way the effect points and disagree — up to ninefold — on how big it is, with gpt-5.4-mini repeatedly the weakest or a pure shadow-follower. "LLMs" as a uniform kind is not what the data show; what they show is a family of systems sharing directional sensitivities with strikingly different strengths. Any single-model study, and any study reporting only averages, would have missed this.
The scorecard, sense by sense
The question of section 1 — for each sense of "meaning," what do LLMs show? — can now be answered in summary form. Everything below is a statement about the three panel models, probed behaviorally, in mid-2026; section 6 lists the standing caveats.
- Distributional meaning: the confirmed substrate. Word-company statistics predict a great deal of model behavior — that is why the project's controls have to be so aggressive — and where a behavior stays inside the shadow, the project says so.
- Constructional meaning: the strongest positive case. Grammatical patterns carry meaning for these models over and above their words: the comparative correlative's inference survives same-word controls at an enormous margin, in three languages; three word-order alternations track human soft constraints through firewall controls; graded grammatical acceptability tracks the human gradient after frequency is partialled out.
- Inferential meaning: present but thin, with a signature. The models draw construction-licensed inferences and compose them, but with a repeated asymmetry: adding an inference layer is easy; cancelling a default is hard. Multi-step inference often needs room to think (section 4.11). Deep inferential competence of the kind that would hold invariantly across embeddings — the presupposition test — did not clearly appear.
- Grounded meaning: no measurable headroom found. Where text statistics were already saturated, adding images added nothing detectable, and a word's perceptual character did not predict how well its senses were tracked. The magnitude question remains open for lack of an instrument, not settled in the negative.
- Referential meaning: not decidable by these methods — and, the project argues, not decidable by any behavioral method (section 6).
- Relational meaning: deflationary. Conventions coined between models are recoverable from the transcript's content; nothing found so far is constituted between the parties in a way that outruns what a single reader could recover. A thin order-sensitivity (latest agreement wins) is real and replicated.
- Human-comparison, overall: where human gradients exist and licenses allowed their use, the models' orderings almost always aligned with human orderings; their magnitudes are their own, and one panel member's alignment is often only surface-deep.
The one-sentence verdict, unchanged through the final weeks of stronger evidence, is this: where this project can see it, LLM meaning is a real, graded, use-based structure, compositional at the level of grammatical constructions, thinly inferential, that beats but does not escape the distributional shadow — silent on reference, negative so far on grounding beyond text, and thin on meaning between agents.