6. What is not known, and what cannot be said with confidence
An honest map shows its blank regions. These are the project's, in roughly descending order of depth.
1. Whether model words refer — probably not answerable by experiment. Consider the question "does the model's word water actually pick out water — the stuff in the world?" The project's analysis of the philosophical literature concluded that this question is premise-bound: rival philosophical camps, running the same theory of how reference works, reach opposite verdicts about LLMs because they disagree about one non-empirical premise — roughly, whether soaking up the recorded usage of a language community makes a system part of that community, inheriting the word–world connections its practices established. No behavioral experiment adjudicates that premise; it is a question about what we are willing to count, not about what the model will do. The project therefore stopped running experiments on reference and marked the question as one its methods cannot close — either answer, asserted confidently about LLMs, is philosophy presented as if it were measurement.
2. Absence of a capability can never be proven. One clean success can prove a capability is present. No number of failures proves it absent — there is always another way to ask, another format, another easing, and the space of elicitations is open-ended. This asymmetry (the project calls a behavioral negative undischargeable) is not a technicality; it bit hard once, when an apparent inability dissolved the moment the answer format alone was eased (section 4.11). Every negative in this report therefore means "not found, searched this far, with these instruments" — a bounded search report, never "cannot."
3. A subtler shortcut may explain any given beater. The controls rule out the statistical shortcuts the project could measure: word frequencies, word-company, topic overlap. They cannot rule out shortcuts nobody has instrumented. The named, known gap is construction frequency: how often the pattern itself (as opposed to its words) occurs in training text. A model could in principle track pattern-level statistics the way the controls prove it is not merely tracking word-level ones. Building a construction-frequency instrument is listed as future work (section 10); until it exists, "beats the shadow" means "beats the word-level shadow," and the project's grammar-difficulty result stands as the cautionary tale — its apparent human-likeness failed exactly such a deeper control, twice, and was refused promotion.
4. Contamination: the tests may be in the training data. The panel models were trained on internet-scale text, and several anchor datasets have been public for years. Where a result could be inflated by memorized test material, the project fences it (the verb-bias result of section 4.2 carries this fence explicitly; the BLiMP accuracy numbers are treated only as upper bounds). The firewall-style designs — invented words, byte-identical strings, fresh items written for each replication — exist largely to defeat this worry, and the strongest claims rest on them.
5. Three models, one moment, outside view only. Everything here is about three particular commercial models of one era, probed from outside. Whether the findings hold for larger or smaller models, for open-weight models, for future systems; whether the behavioral gradients correspond to anything mechanistically real inside the networks; whether the panel's spread reflects architecture, data, or fine-tuning — all unknown. The project deliberately makes no claim beyond its panel, and none about mechanism.
6. The unresolved middle cases. Presupposition behavior beat its word-matched control but in a way a surface-cue account survives, and its promotion was refused — the honest status is "gated behavior, cause unresolved." The models' missing hesitation on borderline word senses (section 4.4) is measured and replicated, but whether it reflects something deep about these systems or just their conversational training is unknown. And roughly two-thirds of all results are internal-contrast-only — solid facts about model behavior that are simply not comparable to human data yet, because no license-clean human dataset exists to anchor them.