Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260805g-printed-switch.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260805g-printed-switch
statusfrozen
created2026-08-05
updated2026-08-05
sensescultural-mediation, accuracy
internal-judgment-onlytrue
provisionaltrue
trackT1
linksworkshop/experiments/E-20260805g-printed-switch/design.md, workshop/experiments/E-20260805g-printed-switch/critic.md, workshop/translations/koyhaa-kansaa/R05-v1/translation.md, workshop/translations/koyhaa-kansaa/register.md, workshop/translations/koyhaa-kansaa/contamination.md, wiki/arms/ARM-atelier-cycle.md, wiki/findings/results/RS-20260802d-class-line-carry.md, config/models.md

RS-20260805g — the primary is withheld by its own gate, and the reason is that every seat could read the Swedish

The instrument separated nothing on the code the run was built around, and the pre-registered gate G3 fires on the first line of this page rather than the last. What the run did establish is a constraint on how this project may measure anything of this kind again, and it is not a constraint the design anticipated.

E-20260805g / ARM-atelier-cycle step 7's study limb. Translation limb: span 7 of T-koyhaa-kansaa-R05-v1, frozen at a3b9a0b with its log before this design was written (charter A4). Every number below is recomputed by analysis/verify.py from the stored raw bodies — 60 checks, 0 failures, five mutation tests, all caught.

1. What was asked

Canth prints one Swedish line inside her Finnish at ¶449 — «Kan hon botas?», can she be cured — spoken by the pastor to the doctor over a woman tied hand and foot on her own floor. Unglossed, unnarrated, typographically unmarked. D115 kept it in Swedish and declared a cost: the device survives and the reader's seat inverts — Canth's Finnish reader could read that line and was placed with the gentlemen; an English reader cannot and is placed with Mari.

The run put ¶442–456 (15 paragraphs, 253 words) to blind seats in three arms differing in one line — A1 the frozen Swedish, A2 the same question in English, A3 the Swedish plus a two-word narrator's label — and asked, uncued, for a four-to-six-sentence retelling.

2. The numbers

Twelve accepted bodies, four seats × three arms. Counts out of four.

code A1 kept A2 Englished A3 labelled
EXCL — someone in the room reported as unable to follow what was said 0 0 0
VIA-LANG 0 0 0
LANG — a language named at all 1 0 1
OPACITY — something on the page reported as foreign, no person named 1 0 1
PASTOR-CURE (post-hoc, §4) 4 2 4

C1, the locality control, passes in all three arms — every one of the twelve retellings names at least 3 of the 4 registered events, and in fact all twelve name 4 of 4. The manipulation moved nothing about which events a reader reports.

3. G3 fires: the primary is withheld

EXCL is 0 in twelve of twelve bodies. The pre-registered gate is unambiguous: if all bodies score identically on EXCL, the instrument separated nothing and PR-PRIMARY is withheld. It is withheld. No count in the EXCL row licenses any statement about whether a reader recovers the exclusion, in any arm, and no sentence anywhere else in this project may cite it as though it did.

The baseline the run was rebuilt around — critic pass 2's F4, the observation that the scene carries the same inference by four routes that are not the foreign line — also returns 0. Not one seat reported the doctor's averted eyes, the unanswered question, the silence during the writing, the pastor's "nothing more for us to do here" or Heikura's having to break in, as anybody being shut out of anything.

The honest reading is about the task, not about the passage. A four-to-six-sentence retelling of a fifteen-paragraph scene has room for the plot and nothing else, and all twelve retellings spent it on the plot: the doctor arrives, Mari is bound, a prescription is written, the landlord asks about his other tenants. Asking for a summary and scoring what it omits measures the compression, not the reader. That is the RS-20260805b lesson arriving a second time in six sessions, on a different instrument, and it is now twice on the record.

4. What the run did find, and it disqualifies the instrument for this question

Every A1 seat reported the content of the untranslated line. PASTOR-CURE codes whether a retelling attributes a question about curing to the pastor. On A1's page that content exists only inside «Kan hon botas?» — the parallel question four paragraphs later is Heikura's, not the pastor's, so a pastor-attributed answer cannot have come from it. A1 scores 4 of 4.

Two of them are worth quoting, because they are the result:

openai/gpt-5.6-terra, arm 1 — "The pastor asks in Swedish whether Mari can be cured, but the doctor does not answer."

google/gemini-3.6-flash, arm 1 — "The pastor asks if Mari can be cured and, receiving no answer, prepares to leave."

The first read the Swedish, identified the language, translated it, and said so — gate G1 fires on exactly this body, since the word Swedish has no warrant anywhere on arm 1's page. The second read it and translated it silently, reporting the content as though it were in English.

D115's declared cost depends entirely on the target reader not being able to read that line. The artifact's own front matter declares its reader: a reader with no Finnish, meeting Canth for the first time, reading for the story — no facing text, no notes. A panel of multilingual language models is not that reader and cannot be made into one by any prompt. On a question whose whole content is what the reader does not know, an instrument that knows it measures something else.

And the direction of the one contrast that did move points the same way. PASTOR-CURE is 4 of 4 on A1, where the line is in Swedish, and 2 of 4 on A2, where it is in plain English. With n = 4 this is not a difference to claim and none is claimed. What can be said is that nothing in the data suggests the Swedish suppressed the content for these seats, and the raw counts run the other way: the foreign string was, if anything, more salient than the plain one.

The narrator's label bought nothing. A3 differs from A1 by the two words in Swedish, printed on the page, and its LANG count is 1 of 4 — the same seat, and no other. Three of four seats did not mention the language even when the narrator named it for them.

5. What this means for the translation, which is: nothing yet

No conclusion of this run bears on D115. The rendering stands as frozen, for the reasons the log gives, and those reasons were never that a panel would confirm them. What changes is that the cost D115 declares is not measurable by this project's present instrument, and the log says so now by pointer rather than claiming a verdict it does not have.

RS-20260802d's finding — that English is better placed than Swedish to carry this novella's code-switching — was about the narrated case at ¶130 and is untouched. The printed case remains what it was before this run: a decision taken on the page for stated reasons, with a declared cost that has not been priced.

6. Cost

$0.146723450, against a declared worst case of $1.03 — 14%. Key-usage delta 0.146723450 against a per-request sum of 0.146681450, residual $0.000042, which is exactly the price of one four-token diagnostic call made by hand to reproduce a provider error (§7) and is therefore closed to 1e-9.

what $ note
three pre-run critic passes 0.060283200 41% of the run, and the best money in it
twelve scored seat bodies 0.041938800 12 of 12 accepted, zero retries
six dead qwen3.7-max bodies 0.044459450 30% of the run bought nothing — §7
the hand diagnostic 0.000042000 four tokens, and it found the 400

The critic cost more than the run it criticised and was worth it. Three passes, thirty-one findings, all accepted; the question the run finally asked is not the question revision 1 asked, and the confirmatory test was withdrawn before dispatch rather than after seeing the numbers. Full record: critic.md.

7. Two instrument defects, recorded rather than repaired quietly

(i) qwen/qwen3.7-max returned finish_reason: length with empty content on all six of its dispatches, at the frozen cap of 1,500, burning $0.0444 for nothing — note (bhf), and a fourth distinct slug. The cap was not raised: S113's recorded lesson is that the fix is a changed seat, not a raised cap, and raising a frozen parameter mid-run is a design change. The seat is dropped under gate G2 and every arm is reported at n = 4.

(ii) mistralai/mistral-medium-3-5 returned HTTP 400 on all six dispatches, billed $0, and the diagnosis had been thrown away by the runner. call.py stored repr(e) for a transport failure, and an HTTPError's repr is <HTTPError 400: 'Bad Request'> — the provider's actual message lives in the response body, which was being discarded. Reproducing one call by hand returned "top_p must be 1 when using greedy sampling", which fires only when the reasoning field is present. call.py now stores the error body; the seat was re-dispatched with the field omitted, the payload, arms, seat list and max_tokens all exactly as frozen, and returned 3 of 3 on the first try. A guard that records that something failed but not why is half a guard.

8. Limits, and the ones the critic named that this run did not repair

  1. The instrument cannot proxy a monolingual reader (§4). This is the run's finding and it is also its largest limit: nothing here describes what a reader without Swedish would do.
  2. Twelve bodies, four seats, one passage, one line. No general claim about readers, about English, or about code-switching in translation is licensed, and none is made.
  3. Critic pass 3 F3 — the scene's four non-language routes cannot be stripped without rewriting Canth. The design measured them instead of removing them; that is not a repair.
  4. Critic pass 3 F5 — keyword coding of free prose is brittle in both directions, and verify.py measures the brittleness rather than assuming it away: the frozen EXCL list misses 4 of 4 plainly exclusionary phrasings put to it as acceptance tests ("over the family's heads", "went past Holpainen entirely", "meant nothing to the people of the house", "only the two gentlemen knew"). Since EXCL is 0 everywhere, this cuts one way: a retelling that expressed the exclusion in words outside the list would have been scored 0. Blinded dual coding is what would fix it and it was not run. All twelve retellings are stored in runs/ and can be read by hand against this claim.
  5. Critic pass 3 F1/F6 — n = 4 supports no estimate of magnitude, and the baseline has no failure mode. The confirmatory test was withdrawn before dispatch; no p-value is computed anywhere in this run.
  6. PASTOR-CURE is post-hoc. It was coded after the retellings were read, is declared as such here and in verify.py, and carries no pre-registration. It is reported because it is what disqualified the instrument, not because it was predicted.