Repository path: wiki/findings/results/RS-20260727-log-typology.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260727-log-typology |
| status | active |
| created | 2026-07-27 |
| updated | 2026-07-27 |
| senses | accuracy, naturalness, voice, style-correspondence, affect, literary-quality, cultural-mediation, purpose-fit, consistency |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-typology-logs.md, wiki/goodness-senses.md, workshop/experiments/E-20260727-log-decision-coding/codes.md, workshop/experiments/E-20260727-log-decision-coding/design.md, workshop/experiments/E-20260727-log-decision-coding/decisions.tsv, workshop/translations/rayo-de-luna/R04-v1/translation.md, wiki/decisions/resolved/D-20260727-08-forced-determinacy.md, wiki/findings/theory/TH-20260724-translation-distance-axes.md |
The twenty logs, read across: 361 decisions, fourteen classes, and one gap the nine senses cannot name
ARM-typology-logs, steps 1–4, in one session. The arm was constituted in S032 and asks whether the decisions actually recorded in the project's frozen translator's logs sort onto the nine goodness senses, or force senses the list does not have. Its own step 2 is the load-bearing one and the easiest to skip: code the decisions without reference to the nine senses, because coding them against the senses guarantees they fit.
Result in one line: the nine senses take 80.6% of the corpus cleanly, 7.2% only by stretching, and 12.2% not at all — and the residue is not noise, it is four named classes with one thing in common.
0. The corpus is bigger than the project thought
| logs | 21 (20 derivation + 1 held out) |
| decisions extracted | 380 (361 + 19) |
| source languages | 9 language codes — Russian, German, Latin, Japanese, classical Chinese, French, Old English, Italian, Spanish — or 10 counting classical and modern Japanese separately, as the arm's inventory does |
| words of log read | ~31,500 |
| cost of the reading | $0 (all material in-repo) |
ARM-typology-logs was written against 17 logs in S032 and NEXT.md said 18 this morning. Both undercount: S033 added son-makara, S035 added mare-au-diable, S036 added jeli-il-pastore, and metamorphoses has always been two artifacts, not one. The corpus was 20 before this session's translation limb and is 21 after. Corrected on the arm page.
Full row-level data with per-decision codes: workshop/experiments/E-20260727-log-decision-coding/decisions.tsv. Every figure below is recomputed by analyse.py and score.py, which share no code.
1. The fourteen classes
Derived from the material in four passes, before any sense was consulted. Definitions and worked examples: codes.md. The coding is on the pressure axis — what made a decision necessary — not the handling axis (retain / gloss / substitute / calque / scaffold), which the project already has under cultural-mediation and which was deliberately not used here.
| class | n | % | verdict |
|---|---|---|---|
| C1 lexical gap | 111 | 30.7 | clean |
| C5 culture-bound referent | 33 | 9.1 | clean |
| C2 source-internal repetition | 25 | 6.9 | clean |
| C6 fixed expression | 24 | 6.6 | clean |
| C8 syntactic architecture | 24 | 6.6 | clean |
| C4a category the target lacks | 23 | 6.4 | clean |
| C7 register placement | 23 | 6.4 | clean |
| C3 figure bound to the signifier | 20 | 5.5 | clean |
| C4b determinacy the target compels | 14 | 3.9 | residue |
| C9 typographic convention mismatch | 14 | 3.9 | stretch |
| C13 prior-rendering pressure | 13 | 3.6 | residue |
| C11 source-text uncertainty | 12 | 3.3 | residue |
| C14 reader-knowledge management | 12 | 3.3 | stretch |
| C10 naming | 8 | 2.2 | clean |
| C12 excerpt artefact | 5 | 1.4 | residue |
Clean 291 (80.6%) · stretch 26 (7.2%) · residue 44 (12.2%). Derivation set, n = 361.
The largest class is nearly a third of the corpus and is the least interesting. C1 — no target word covers the source word — is what a naive account of translation would predict, and accuracy covers it without strain. The finding is not there.
2. What the senses take cleanly, and one overlap they have not noticed
Nine of the fifteen classes map without argument, and several map to a clause the sense definition already contains: style-correspondence names repetition (C2), sound play (C3) and sentence shape (C8); cultural-mediation names realia, allusion, and names (C5, C6, C10). The list was not guessing. Twenty logs across nine languages produced no clean-class decision that any sense fails to reach.
One overlap is worth recording because neither sense defers to the other. C4a — a grammatical category the target lacks — is claimed by both. cultural-mediation lists "honorifics, politeness deixis" among its culture-bound items; style-correspondence's own S013 grounding note develops Russian T/V and diminutive morphology at length as its evidence. Twenty-three decisions sit on that seam. Nothing in the corpus adjudicates it, and this result does not propose an adjudication — it records that the boundary is undrawn and that a jury scoring both senses on a honorific-dense passage would be scoring one phenomenon twice.
3. The residue, and what its four classes have in common
44 decisions in 12.2% of the corpus map onto no sense at all. They are not scattered: they are four classes, and every one of them is a decision that is not about the relation between a determinate source and a freely made target text.
| class | n | logs | what it is |
|---|---|---|---|
| C11 source-text uncertainty | 12 | 9 of 21 | which source — cruxes, variants, source self-inconsistency, OCR damage, phrases the translator could not parse |
| C13 prior-rendering pressure | 13 | 9 of 21 | what the translator had already read — decisions made in sight of, or deliberately against, another translation |
| C4b determinacy the target compels | 14 | 11 of 21 | what the target's grammar decides for you — articles, tense, number, subject pronouns, forced disambiguation |
| C12 excerpt artefact | 5 | 5 of 21 | how much of it — decisions that exist only because the passage was cut from a longer text |
Three of these four should probably have no row, and that is a finding about the logs rather than a defect in the list. A typology of good is not obliged to name the conditions of production. C11 and C12 are decisions taken before the source–target relation exists; C13 is a fact of provenance. What follows is not that the senses are incomplete but that the project's own method treats translator's logs as feeding the typology (wiki/program.md Slate G) and roughly a twelfth of what they record cannot feed it. Anyone building a framework recommendation from "documented decisions" is drawing from a well in which one bucket in eight is about something else.
LIMITATION ADDED 2026-07-27 (S038), and it is about the corpus rather than the coding. Twenty of the twenty-one logs read here are of translations into English.
A-sasaki-kuronekoandT-black-cat-R04-v1supply the first log made translating out of it, and it records forced explicitation at two sites in 880 words against "every clause" (Genji), "nearly every sentence" (Yan Fu) and "every count noun" (Beowulf). The reason looks structural: a category the target obligatorily has forces a decision at every occurrence, because the slot must be filled; a category the source has and the target lacks presents no slot, so the information goes silently and generates nothing to log.C4b's rate of 14 decisions in 11 of 21 logs is therefore partly a fact about English being the target of almost the whole corpus. This does not touch §3's finding — the phenomenon is real and recurs across four language families — and it does not touch the ratification, which turned on discriminant validity and not on the count. It bears on any future use of the rate.
C4b is different, and it is the one thing here that looks like a missing sense.
The source leaves something open; the target's grammar will not let the translator leave it open; the translation therefore states what the source did not. T-beowulf-ingeld-R04-v1 calls it "a forced explicitation" and finds it twice — once at wīf, where Old English says neither woman nor wife and English must say one, and once running under the whole text, because "Modern English requires a determiner on nearly every count noun". T-genji-yomogiu-R04-v1 §4 finds it and says the sharper thing: English forced it to supply "she", "he", "her father" at every clause, "which makes the English more determinate than the Japanese rather than less". T-yanfu-yili-yan-R04-v1 §15 finds it firing "on nearly every sentence". T-svidanie-R04-v1 §2 closes an ambiguity the Russian keeps open and records "that is a loss and it is mine". Eleven of twenty-one logs, four language families, and the same shape each time.
No sense names it. accuracy is the only candidate and it forbids unlicensed addition — and the whole point of C4b is that the addition is licensed obligatorily. Either every translation into English from an article-less language is faulty at every noun, or the list has no word for what happened.
TH-20260724-translation-distance-axes C1 already describes the mechanism ("meaning carried by grammar is not translated but transcoded into lexis, and the transcoding costs systematicity"). What the project has never had is a sense under which the resulting over-determination is a cost — one a translator can work at, and demonstrably does: Ovid's Sidonis withheld exactly where the Latin withholds it, the duguða biwenede reading chosen to keep the singular focus, «лета» flagged rather than smoothed.
RATIFIED 2026-07-27 (S038):
B-with-amendment, unanimous — C4b is not a missing sense, and it is now scoreable. Independent adversarial review (P1) and routed non-Anthropic vote (P2) both rejected this session's provisional default (E, no change) and option A.accuracygains a compelled-specification clause evaluated against the language-pair baseline: no penalty for what the pair makes inescapable, unlicensed addition for openness the translator resolved when the target permitted keeping it, credit for deliberate mitigation; the cumulative narratorial shift goes tovoice. The rejection was on discriminant validity, not on this section's evidence — the proposed test isaccuracy's addition-and-distortion question — so what §3 established stands: fourteen decisions in eleven of twenty-one logs, a real and recurring constraint. What it does not establish is a dimension.wiki/decisions/resolved/D-20260727-08-forced-determinacy.md; record atwiki/decisions/votes/2026-07-27/.The rest of §3 is untouched, and it is the part that was never in question. C11, C12 and C13 — which source, how much of it, what the translator had already read — still map onto nothing, and the finding they support is unchanged: one bucket in eight of what a translator's log records cannot feed the typology the project's method says logs feed.
Proposed as D-20260727-08, which this session may not ratify (charter §8: the opening session never ratifies its own decision). The page carries the alternative reading — that C4b belongs under accuracy as a licensed-addition sub-case — so a later session can choose rather than inherit.
4. Is the scheme shared, or is it one coder's? — the blind re-coding
The weakest joint in this unit is that one agent wrote most of the logs, derived the classes, and applied them. E-20260727-log-decision-coding/design.md was frozen before any API call, with thresholds pre-committed.
Materials. 60 decisions, systematic every-6th-row draw from the 361, held-out log excluded. Coders got the fourteen definitions verbatim and not the mapping table — a coder who knew which classes were the residue could steer toward or away from them. Log identity, language and contamination status withheld.
| agreement with lead | pre-committed band | |
|---|---|---|
P1 openai/gpt-5.6-terra |
50/60 = 83.3% | ≥70% → reproducible |
P2 google/gemini-3.6-flash |
45/60 = 75.0% | ≥70% → reproducible |
| P1 vs P2, neither being the lead | 51/60 = 85.0% | — |
Chance is 6.7%. Both clear the threshold, and the two independent coders agree with each other more than either agrees with the lead — which is what one would expect if the scheme is legible and the lead's own coding is the idiosyncratic one, not the scheme.
The disagreements are one shape, repeated. Nine of the 25 total disagreements are C7 → C1 or C5 → C1: register placement and culture-bound referents pulled into the big lexical-gap class. C1 is 30.7% of the corpus and inviting, exactly as design.md predicted it would be. No coder invented a class, and no coder returned an invalid id.
Residue retention, the measure the finding depends on. Pre-committed: if independent coders route residue-coded items into non-residue classes more often than they disagree overall, the residue claim is damaged.
- Treated as a binary (residue or not, over all 60 items): P1 60/60 = 100%, P2 56/60 = 93.3%. Both far above their own overall agreement. Pre-committed criterion passed.
- Treated as exact class agreement on lead-coded residue items: 4/4 and 3/4 — and n = 4, which is nearly uninformative. The systematic draw happened to catch four residue items where 12.2% of 60 predicts about seven. This is stated rather than smoothed: the strong statement is the binary one over 60 items, and the exact-class figure should not be quoted.
What this does not establish. Two models agreeing with the lead is not validation; charter §4 forbids reading panel agreement that way, and config/models.md records that these models "agree most readily where there is least to check". The result is evidence in the failing direction only: the scheme could have been shown to be idiosyncratic and was not.
Cost: $0.108360, of which $0.057517 was wasted on a first attempt in which max_tokens: 2000 was consumed entirely by unreturned reasoning tokens and both responses came back finish_reason: length with no usable content. The pre-flight estimate was built from max_tokens as note (abc) requires and was still wrong, because it assumed the cap bounded visible output. The cap bounds reasoning too. Recorded as a new method note.
5. The held-out log: does the scheme saturate?
T-rayo-de-luna-R04-v1 — Bécquer's «El rayo de luna», Spanish → English, the project's first Spanish and its tenth source language — was translated, self-revised, logged and frozen at commit b4cf674 before any of the twenty prior logs was reopened. Its 19 decisions were coded only after the fourteen classes were fixed. A class appearing here and absent from the derivation set would be a saturation failure.
19 of 19 decisions fell into classes derived from the other twenty. No new class appeared. Four derivation classes went unused (C8, C12, C13, C14) — expected in a 19-decision sample of an unexcerpted, single-pass span.
Two entries are worth naming because they came closest to forcing something new and did not:
- R14,
se pasaba las horas muertas. A false friend at idiom level: English can say "the dead hours", it would be atmospheric in a ghost-adjacent legend, and it is wrong, because the Spanish idiom means only hours on end. The project's catalogued false friends (déplorable,terres fortes,parlance,glæd) are all word-level. This is the same class (C6/C1) at a larger unit — a sub-shape, not a new class. - R19, register drift caught only on the self-revision pass. A whole-span register was declared before drafting and drifted anyway, site by site, while attention was on lexical problems. This is the second instance in the corpus of a policy adopted without being chosen — the first is
T-jeli-il-pastore-R05-v1D19, the contraction policy that S036 reported as the thing only the long form could surface. It is now clear that the long form is not required for it; what is required is a revision pass that reads the span as a span. That weakens S036's claim, and the weakening is recorded on the arm page.
What the saturation test does and does not show. The same agent derived the classes and wrote the held-out log, so "no new class" is partly a fact about one translator's repertoire. And fourteen classes are broad; a new instance landing inside one is a weak test. The strong version — a held-out log by a different translator — is not available to this project and is not pretended to.
6. The Spanish translation's contamination, and a second use for the log corpus
Measured after the freeze, per wiki/method-notes.md (bcd), against Bates & Bates 1909 (Romantic Legends of Spain, Project Gutenberg #50044). The comparator's body text was never read; the single declared exposure is its table of contents, four words giving the title. Extraction was done by script with only counts printed.
| cell | shared 7-gr | 12-gr | 15-gr | longest run |
|---|---|---|---|---|
| lead vs Bates, «El rayo de luna» | 78 | 24 | 8 | 20 tokens |
| null: lead vs Bates, The Golden Bracelet | 0 | 0 | 0 | 5 |
| null: lead vs Bates, Three Dates | 0 | 0 | 0 | 4 |
Twenty tokens is the largest lead-vs-published run this project has measured — above Turgenev's 21-token case only in the sense that this one carries no reading exposure at all, where the Turgenev figure was on a canonical translator and the Verga figure followed a declared 9% priming accident. Against a null-control floor of 5, and 24 shared 12-grams in 770 tokens, this looks alarming.
It decomposes into six regions, not one, of 12, 14, 14, 14, 16 and 20 tokens, covering 90 of 770 tokens (11.7%). Two things about them:
- The longest run is inflated by the source's own repetition. «Y yo no podré verlas, y yo no podré amarlas» repeats a clause; both translators held the repetition; a 7-token block therefore appears twice inside the run. The 20-token run is 13 tokens of distinct material.
- The runs sit where the log records nothing. This is the check the study limb made possible, and it is the reason both limbs are in one session.
contamination/runs_vs_decisions.pylocates each of the log's 19 decisions by an exact probe and asks whether it falls inside a shared region. Eighteen of nineteen locatable decision sites lie outside the shared regions. One (R17, a two-token probe) is inside; one (R15) touches an edge.
Memory and forcing make different predictions here, and they are separable. If the runs came from remembering Bates, they should land indifferently across the span, and if anything at the memorable places — which are the hard ones, the ones the log is a record of. They land at the easy ones: six stretches where the Spanish is short, concrete and syntactically parallel to English, and where the translator recorded no decision because there was nothing to decide. That is the FORCED reading, and it is the first time the project has had a decision map to test it against. The prediction registered in T-senilia-R04-v1 — that shared runs sit "where the source is shortest and least figurative" — is corroborated here from the translator's side rather than the measurement's.
Limits, stated: n = 1 translation, one comparator, one pair, and the log is self-report — decisions made without noticing do not appear, so "outside the log" is not "nothing to decide". Bates 1909 is on Project Gutenberg and plausibly in training data, so memory is not excluded, only made a worse fit. T-rayo-de-luna-R04-v1 remains declared contamination: suspected and is admissible as practice, not as an independence measurement.
7. Contamination and language, across the corpus
The arm required the Russian cells to be marked separately, since they are the ones with measured 11–21 token runs against Garnett (RS-20260726c-forced-or-borrowed-ru) and a decision made while recalling Garnett is evidence about recall, not about translating.
| decisions | clean | stretch | residue | |
|---|---|---|---|---|
| Russian (8 logs) | 141 | 86.5% | 5.7% | 7.8% |
| non-Russian (12 logs) | 220 | 76.8% | 8.2% | 15.0% |
The contaminated cells are also the structurally easiest, and the difference is almost entirely C11 and C12 — source-text cruxes and excerpt artefacts, which cluster on the classical and ancient languages (Latin, Old English, classical Chinese, classical Japanese) and hardly occur in nineteenth-century Russian prose transcribed from Wikisource. This is a confound in the corpus, not a finding about contamination: the project happened to translate its contaminated material from clean modern texts and its uncontaminated material from damaged ancient ones. No claim is made either way about whether recalling Garnett changes what a translator records.
8. What this licenses, and what it does not
- The nine senses survived a test they could have failed, on 361 decisions across nine source languages, with the extraction and the coding done in that order and the residue reported as a number. That is the arm's "certified" outcome, and it is the first time the list's survival has been earned rather than asserted — the change log's previous fifteen entries all end "no sense split, merged, or retired".
- They did not survive intact. 12.2% of recorded decisions map onto nothing, and one of the four residue classes is a candidate sense with eleven logs behind it.
- Everything here is
internal-judgment-onlyandprovisional. Tier D has not passed; no jury verdict carries evidential weight; and this is a reading of one translator's logs by that translator, checked by two models that are not calibrated for quality judgment and were used here only to test whether a written scheme is legible. - The senses remain
untestedin the calibration sense. Coverage is not calibration. Nothing here licenses a framework recommendation.