Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260804f-whose-mind/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260804f-whose-mind
statusfrozen
created2026-08-04
updated2026-08-04
sensesvoice, style-correspondence, accuracy
internal-judgment-onlytrue
provisionaltrue
trackT1
linksworkshop/translations/koyhaa-kansaa/R05-v1/translation.md, workshop/translations/koyhaa-kansaa/register.md, wiki/arms/ARM-atelier-cycle.md, config/models.md, wiki/goodness-senses.md

E-20260804f — whose mind is this? The free-indirect layer of span 5 in English

Frozen 2026-08-04, before dispatch and after the translation limb was committed at 65e5f8a. Study limb of ARM-atelier-cycle step 5. The translation limb is span 5 of T-koyhaa-kansaa-R05-v1 (¶296–375, the night).

1. The wire, in one sentence

Span 5 is the novella's sustained double-interior passage — ten stretches of untagged free indirect discourse alternating between two consciousnesses — and the register has already declared away most of the devices Finnish uses to mark whose mind a paragraph is in, so this experiment asks whether the English still delivers the attribution.

2. Question

Canth marks free indirect discourse with morphemes English has no equivalent of: the clitics -hän and -pä, the spoken determiner se before a name, evaluative lexis in the character's own register, and (at ¶167, ¶354) plural subject with singular verb. D93 records the loss explicitly at ¶366, where Eihän se Mari becomes "Why, Mari"; V3 forbids carrying the colloquial layer as English dialect; V22 forbids the flat rendering of se.

Can a reader with no access to the Finnish say whose consciousness an untagged passage is in?

Two outcomes are both informative. If yes, the lost morphemes were redundant with what English does carry — context, lexis, deixis — and the register's declared losses cost less than they look. If no, the loss is real, it is concentrated where the translation could not compensate, and the framework should say so about Finnish→English narrative prose.

3. Materials

materials/loci.json, built by build_materials.py from the committed artifact. Sixteen loci cut from the span-5 English, each 1–3 sentences, each keyed to a verbatim opening string the builder asserts is unique.

class n ground truth role
TAG 3 the subject of an explicit verb of thinking in the source (hän ajatteli, arveli hän itsekseen, Holpainen ihmetteli) positive control — the English keeps the tag, so seats must get these
NARR 3 narrator — no tag and no character-indexical morpheme in the source; past-tense report of observable events specificity control — seats must not see a mind everywhere
FID 10 the character indexed by the source's clitics / se / evaluative register, cross-checked against the nearest preceding tagged sentence the measurement

Ground truth is fixed by the Finnish, not by the English, and not by taste. A locus is FID only where the source has no tag and does have at least one character-indexical morpheme. The two rules — "nearest preceding tag in the source" and "which character the morphemes index" — agree at all ten FID loci; had they disagreed anywhere, the locus would have been dropped before dispatch.

FID truth is 7 MARI / 3 HOLPAINEN, which is deliberately imbalanced and is why §6 benchmarks against the best constant strategy rather than against chance.

4. Procedure

One dispatch per seat. Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 (config/models.md). Reserve: P5 deepseek/deepseek-v4-pro.

Each seat receives the whole span-5 English — so that context is available, which is the condition a real reader is in — followed by the sixteen passages in fixed order, and answers for each: WHOSE ∈ {MARI, HOLPAINEN, NARRATOR}, QUOTE (verbatim words from the text that decided it), CONF 1–5.

Blinding. Seats see English only. No Finnish, no source, no author, no title, no mention of translation, and no mention of free indirect discourse or of any narratological term — the question is put as an ordinary reader's question about whose thoughts a passage gives. Nothing tells a seat that the passages fall into classes or that any class is a control.

temperature 0, reasoning.effort: low on the first dispatch (note (b)), max_tokens 6,000, usage.include. Raw bytes to disk before any parse.

5. Predictions, registered before dispatch

The lead's own expectation, recorded so it can be wrong: R1 and R2 hold, R3 fails at 3–5 of 9, R4 holds only because of the Mari cells. That is a prediction that the attribution survives for the character the reader is already with and fails at the switch.

6. Benchmarks and how the result is read

A constant guesser is the benchmark, not chance (the S077 lesson: an instrument that loses to a constant guess has measured nothing). Always-MARI scores 8/16 loci overall and 7/10 on FID. The result is reported per class, with the constant-strategy score printed beside it.

Nothing here is a quality claim about the translation. The loci are not scored for goodness; the measurement is whether an attribution recoverable in the source is recoverable in the target. Every evaluative sentence on the result page carries internal-judgment-only and provisional — Tier D is NOT PASSED, and these seats are coders, not jurors.

7. Failure criteria

8. Comparator, after the run, at no cost

Hertzberg's authorized 1886 Swedish (gutenberg.org/ebooks/20518, PD, logged in wiki/base/consulted.md) at the same sixteen loci: does Swedish keep a marker English drops? Swedish has the modal particles ju, nog, väl, which sit closer to Finnish clitics than anything English has. A capability/choice reading by the lead, internal-judgment-only, and not part of the registered predictions.

9. Budget

Declared worst case $0.55: critic $0.18 (P4, max_tokens 12,000, note (abc)); three seats $0.16 at list, doubled to $0.32 for the routing caution config/models.md records at S022. Today's UTC headroom before this run: $3.463398640.


Amendments A1–A9, from the pre-run critic (P4 moonshotai/kimi-k3, critic.md)

Verdict NEEDS-AMENDMENT, nine findings, four BLOCKING. All nine accepted; three of them changed what this experiment is measuring. Applied before any subject dispatch. §§5–7 above are superseded where they conflict with what follows.

A1 (F1, BLOCKING) — the primary prediction was rigged and is replaced. Two of the three Holpainen-FID loci name Mari in the third person (L13 "Mari always took the smallest share", L14 "Why, Mari had not eaten that soup"), so 6 of R3's 9 cells are answerable by the rule the person named is not the thinker, with no sensitivity to anything this experiment is about. R3's threshold was reachable on those two loci alone. The critic is right, and it is worse than it says: L14 was the design's showcase because D93 records the loss there — and D93's own compensation, "Why, Mari", supplies the name that gives the answer away. R3 as written is withdrawn.

A2 (F2, BLOCKING) — presentation order is now per seat, and adjacency is forbidden. Extracts were to be shown in span order, which lets a seat propagate an answer across neighbours (L11→L12, L13→L14) and puts L4 next to L5, a tagged Holpainen locus. Orders are now generated per seat from recorded seeds (105001/2/3), stored in materials/orders.json before dispatch, and the builder asserts that no two adjacent extracts share a ground truth. Extracts are relabelled E1…E16 per seat so the id carries no information; answers are mapped back by the stored order.

A3 (F3, BLOCKING) — the claim in §2 is narrowed to what the run can support. Seats see the whole span, so a correct answer may come from plot context rather than from anything in the extract. The run therefore cannot show that "the lost morphemes were redundant". What it can show is which English device the attribution rides on, which is what A4 makes the new primary.

A4 — the new primary is a mechanism prediction, on all 30 FID cells. Finnish hän is genderless; English must choose she or he. In a two-person scene that is an attribution device the source does not have, and it is available at every locus. R3′ (registered): in ≥ 20 of 30 FID cells, the QUOTE a seat gives as its evidence is a person-index English supplies — a proper name or epithet of the other party, or a gendered pronoun — rather than the diction, the register, or the topic. The lead's expectation, recorded so it can be wrong: R3′ holds, and the English carries the attribution by a device Canth had no access to rather than by anything that answers to her clitics. The old R3 survives only as a descriptive split: R3a the name-cued Holpainen cells (L13, L14), R3b the one non-name-cued Holpainen locus (L4), reported per cell and not a prediction.

A5 (F4, BLOCKING) — the ground-truth cross-check is published, and the design's claim about it was FALSE. §3 asserted that the two rules — nearest preceding tagged sentence in the source and which character the indexing morphemes point at — agree at all ten FID loci. They disagree at L11. ¶357 «Tuhma pökkelö, kun uskoi…» has the nearer thought-tag at ¶355, «Holpainen ihmetteli mielessään», which the mechanical rule would resolve to HOLPAINEN; the epithet rule resolves it to MARI, because a mind that calls Holpainen a stupid blockhead is not his. L11's truth is MARI and rests on the epithet, not on the tag rule, and it is classed as cued (A6) for exactly that reason. The rule is also undefined at L1: ¶296 has no preceding tag inside the span, and the governing tag is ¶295, in span 4, which is Mari's. Both facts are published here rather than left implicit.

locus source ¶ the indexing morpheme in the Finnish nearest preceding thought-tag truth
L1 296 Mitäpä (clitic -pä), evaluative parempi none in span; ¶295 (Mari) MARI
L2 297 tiesipä (-pä), parka, zero person ¶297 «hän sydämessään … nureksi» MARI
L3 299 ne vaan, fronted evaluative Pahoja ¶298 «Sitä ajatellessaan Mari» MARI
L4 316 kuitenkaan + potential mahtaisi ¶316 «Holpainen … koetti sitä karkoittaa» HOLPAINEN
L7 334 evaluative theodicy in Mari's register ¶334 «sen hän varmaan tiesi» MARI
L10 353 pelkäsivätpä (-pä), sentään, taunt register ¶353 «Hän päätti vaieta» MARI
L11 357 epithet Tuhma pökkelö of the other party ¶355 (Holpainen) — rules disagree MARI
L12 361 jokohan (-han), rukka ¶359 «Olisi tehnyt mieli lyödä häntä» MARI
L13 365 -kin, oikeastaan, se Mari nearby ¶365 «hän … tuumaili» HOLPAINEN
L14 366 Eihän (-hän) + se before a name ¶365 «hän … tuumaili» HOLPAINEN

A6 (F8) — every FID locus is classed by what cue the EXTRACT itself carries, and results are reported by class. CUED (a name or epithet of the other party, or a gendered pronoun indexing the thinker): L1, L10, L11, L13, L14. NO-CUE: L2, L3, L4, L7, L12. The no-cue set is 4 MARI to 1 HOLPAINEN and cannot carry a balanced test — the span contains only three untagged Holpainen stretches and two of them name Mari. That asymmetry is itself reportable: in this span Canth's Holpainen-interior is short and anchored to a named object, and her Mari-interior is long and unanchored. It is a fact about the materials and is not fixable by choosing different loci.

A7 (F5, F6) — the run is a qualitative probe and every threshold is descriptive, not inferential. Three seats are not three independent draws; they share training data and may have seen this text's neighbours. No binomial claim is made anywhere. R5 is tightened: the quote must come from the extract itself, be ≥ 3 words, and all 30 FID quotes are audited by hand by the lead, not merely membership-tested. Constant-strategy scores are printed beside every class.

A8 (F7) — FC4 is lowered. A seat is non-discriminating if it gives the same label at ≥ 12 of 16 loci, or the same label at ≥ 8 of 10 FID loci. Both are checked and reported.

A9 (F9) — the seats run at DEFAULT reasoning effort, not low. Handicapping the seats and then using a control to declare the instrument alive builds a self-fulfilling null. Note (b)'s low default is overridden here for this reason and recorded; the budget is re-declared at $0.75 worst case to cover it.