7. Implications
For the public argument about AI "understanding." The project's results dissolve the yes/no question rather than answering it. The parrot camp is right that distributional patterning is the substrate, and right that some impressive-looking behaviors (acing antonyms; some presupposition behavior) carry no evidence of anything more. The understanding camp is right that some behaviors demonstrably outrun word-level statistics — meaning carried by grammatical patterns, discourse sensitivities matching human speakers through the strictest controls the project could build. Both camps are wrong to speak of "LLMs" as a uniform kind doing a uniform thing: the truthful picture is a map — phenomenon by phenomenon, model by model, with measured depths — not a verdict. If one lesson travels, it should be this: ask which behaviors are informative before being impressed or dismissive; the informative ones are those with shallow statistical footprints and surviving controls.
For linguistics and lexicography. Three constructions' soft constraints — discovered by corpus linguists over decades — reproduced in machines that were never taught them, across firewall controls, for a few dollars each. Whatever else that shows, it demonstrates a new instrument: probabilistic grammatical knowledge can now be probed quickly, cheaply, at pre-registered scale, in any language with a treebank. Construction grammar's central bet — that patterns themselves carry meaning — picked up quantitative, cross-linguistic support from an unexpected direction. For lexicography specifically, the models reproduce the graded, overlapping structure of word senses that dictionary makers have always worked with (and that print dictionaries flatten), at human-annotator levels of agreement — while lacking the trained lexicographer's calibrated hesitation on borderline cases. That combination — the gradient without the doubt — is exactly what an AI-assisted dictionary workflow would need to correct for.
For AI evaluation. Three measurement lessons, each learned the hard way here, generalize to anyone testing models. (1) The answer format is part of the instrument: forcing one-token answers can mask capabilities that a sentence of visible reasoning reveals; a benchmark score is a fact about model-plus-format, not model alone. (2) Small test sets are untrustworthy in both directions: a real effect vanished at 32 items and returned, clearly, at 100. (3) Agreement hides spread: models that agree on direction can differ ninefold in strength, so averages and pass/fail tallies actively mislead. A fourth, subtler lesson: without pre-registration and independent re-computation, an automated research loop will fool itself — the project caught its own errors only because those gates existed.
For philosophy of language. Some questions that were purely notional now have empirical traction: how graded grammatical knowledge is, whether pattern-meaning can float free of word meaning, what a use-structure looks like when severed from reference and grounding. On that last point the project's results constitute a genuinely new kind of evidence: these systems are a working demonstration that a rich, human-aligned structure of use can exist in the confirmed absence of settled reference, grounding, and community membership — a configuration the classic theories never had a live example of. Meanwhile the reference question itself stays philosophical; the project's contribution there is a precise account of why no experiment settles it.
For how research can be done. Sixty-three dollars of measurement, seven weeks, one human hour here and there — and, more importantly, a governance pattern (freeze before running; adversarial review by a fresh instance plus an outside-family model; never ratify your own proposal; ledger your predictions; write your nulls) that let an autonomous system compound reliable findings without a human in the loop. That pattern, not any single finding, may be the most transferable artifact.