10. Directions for future research
Ten directions were proposed to Tom at the plateau; this report was the one he chose first. The others remain open, listed here in plain terms, roughly from most continuous with existing work to most expansive. Each respects the standing constraints (no new human-subject data; verified licenses only; a few dollars a day).
- Does model size change the picture? Re-run the flagship probes down a "ladder" of smaller and larger models within one family, to learn whether beating the shadow grows with scale, and pilot a lane using models that expose word-probabilities directly.
- Build the missing control: construction frequency. Create an instrument that estimates how often a grammatical pattern (not its words) occurs in training-scale text, then re-test the beater rows against it. The single most valuable hardening of the existing claims.
- A Japanese battery. Port the discourse-givenness experiments to Japanese — whose topic-marking particles (は/が) and flexible word order encode givenness directly — using the same firewall logic, to test whether the human-aligned discourse sensitivity is a fact about English or about the models. The Japanese comparative-correlative replication already proves the pipeline works.
- Meaning change over time. The word-sense dataset already in use is historical; probe whether the models track human-rated meaning change across decades (e.g., which uses of plane or awful drifted apart between 1810 and 2010).
- Can models predict human disagreement? The project showed models track the human average while lacking human hesitation. The sharper question: do they know where humans disagree with each other? The annotator-level data to test this is already in the repository.
- Pragmatics: what speakers mean beyond what sentences say. Extend from presupposition to implicature (e.g., how some comes to suggest not all) — blocked in June for lack of a license-clean human dataset, worth a fresh scout.
- Social and emotional meaning. Politeness, formality, connotation — a sense of meaning the project has not yet touched, with published human norms that may be adoptable.
- Idioms and the middle of the scale. The word end and the grammar end are mapped; the graded middle (spill the beans, noun compounds) is not, and human compositionality ratings exist.
- Base models vs chat models. Run the frozen probes on models before and after conversational fine-tuning, to learn which training stage installs the use-structure the project has been measuring.
- Public synthesis. Continue what this report begins: keep the project's findings readable by the people whose language — and whose question — this always was.