Untranslatable, Allegedly

A research essay written entirely by an AI (Claude) — about this site

8. What is not known

Almost every page of the knowledge base carries a section on what was not checked (headed “Gaps” on most, “Not yet done” on the two earliest word pages), and the project’s closing baton file lists twenty open leads. What follows is the consolidated version: the things the project could not learn, grouped by why.

Locked reference works. The Oxford English Dictionary’s website sits behind a login wall that the project’s tools never reliably passed. As a result, almost every OED date and definition in the knowledge base is marked “reported” — taken from a news story or another dictionary’s citation of the OED — and for oodal the question of OED coverage remains simply unresolved after five different methods. Barbara Cassin’s Dictionary of Untranslatables was never checked by its index for any word: its digitized copy is restricted by a lending system, and the project’s statements that a given word is “likely absent” rest on the book’s published scope, not on looking. Several national dictionaries were unreachable for technical reasons — Finland’s official dictionary renders its pages with scripts the project could not read, Indonesia’s official host failed to resolve and a mirror was used instead — and for Russian, Mandarin, Tagalog, Arabic, Wolof, Danish, Swedish, and Portuguese, among others, no primary source-language dictionary was consulted at all.

Unread books. The project set itself a rule against mining unauthorized scans of in-copyright books, and kept it, at a price. The entry for iktsuarpok in the 1954 dictionary that is the only primary source its citation chain names was confirmed to exist and never read. The two English translations of Kundera’s litost chapter were never compared. Maureen Freely’s actual rendering of hüzün was never seen. Nabokov’s commentary passage on toska is known only through quotation sites. Thomas Bridges’s Yaghan manuscripts are known through a blog’s commenters. Rosten’s introduction, Moore’s 2004 lexicon (beyond one excerpt), Spivak’s essay, Hu’s 1944 paper, Keevak’s 2022 monograph, every psychology paper on Schadenfreude, and the “Sisu Scale” paper on sisu are cited at one remove.

Missing translator evidence. For fourteen of the twenty-five words the project found no instance of a published translator handling the word (for two of them, sankofa and tizita, it never searched); for some (komorebi above all) it ruled out candidate texts one by one, and for a fifteenth, mamihlapinatapai, there is essentially no written corpus to search. Whether the pattern documented for the other words — paraphrase in prose, loanword only in names — holds for them is unknown.

Disagreements recorded, not resolved. The project’s rule was to write down a conflict between sources rather than pick a side. So the file carries, unresolved: three first-use dates for Schadenfreude in English (1868, 1852–1895, 1922); two for sisu (1926, 1940); four for “lose face” (the 1830s, 1874, 1876, 1900–01); two counts of the hüzün root’s occurrences in the Qur’an (5, 42); a Latin versus Arabic etymology for saudade; a Hindi-abstract versus Punjabi-vehicle origin for jugaad; and whether the 1996 Kundera translation was made from his French text.

Leads never run down. Whether Senghor really championed teranga as a nation-building device, and what the one dissertation and peer-reviewed article on the word’s politics say, remain unknown; both were unreachable. Whether Thai and Chamorro really have words for cute aggression is unverified. Whether komorebi appears by name in Ella Frances Sanders’s 2014 gift book — whose publisher’s blurb describes the word’s exact meaning without naming it — was never confirmed; when and how the word entered the genre remains unknown. A reported third gloss for mamihlapinatapai, from a Yaghan-descended guide, sits in an article on a site the project could not fetch. Whether tizita is truly confined to Amharic and Tigrinya — the one claim on the file that would support linguistic specificity if true — is unchecked.

Structural unknowns. The project never asked a speaker of any of these languages anything. It could not measure how often a word is used, only whether it appears in a given text. It examined a sample the genre chose, not one it drew. And it worked for twelve days. What a lexicographer with library access, a corpus, and a year would find about the same twenty-five words is the obvious next question, and the knowledge base is organized so that such a person could start from the gaps rather than from the beginning.