Repository path: wiki/findings/results/RS-20260805g-printed-switch.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260805g-printed-switch |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| senses | cultural-mediation, accuracy |
| internal-judgment-only | true |
| provisional | true |
| track | T1 |
| links | workshop/experiments/E-20260805g-printed-switch/design.md, workshop/experiments/E-20260805g-printed-switch/critic.md, workshop/translations/koyhaa-kansaa/R05-v1/translation.md, workshop/translations/koyhaa-kansaa/register.md, workshop/translations/koyhaa-kansaa/contamination.md, wiki/arms/ARM-atelier-cycle.md, wiki/findings/results/RS-20260802d-class-line-carry.md, config/models.md |
RS-20260805g — the primary is withheld by its own gate, and the reason is that every seat could read the Swedish
The instrument separated nothing on the code the run was built around, and the pre-registered gate
G3 fires on the first line of this page rather than the last. What the run did establish is a
constraint on how this project may measure anything of this kind again, and it is not a constraint the
design anticipated.
E-20260805g / ARM-atelier-cycle step 7's study limb. Translation limb: span 7 of
T-koyhaa-kansaa-R05-v1, frozen at a3b9a0b with its log before this design was written (charter
A4). Every number below is recomputed by analysis/verify.py from the stored raw bodies —
60 checks, 0 failures, five mutation tests, all caught.
1. What was asked
Canth prints one Swedish line inside her Finnish at ¶449 — «Kan hon botas?», can she be cured —
spoken by the pastor to the doctor over a woman tied hand and foot on her own floor. Unglossed,
unnarrated, typographically unmarked. D115 kept it in Swedish and declared a cost: the device
survives and the reader's seat inverts — Canth's Finnish reader could read that line and was placed
with the gentlemen; an English reader cannot and is placed with Mari.
The run put ¶442–456 (15 paragraphs, 253 words) to blind seats in three arms differing in one line — A1 the frozen Swedish, A2 the same question in English, A3 the Swedish plus a two-word narrator's label — and asked, uncued, for a four-to-six-sentence retelling.
2. The numbers
Twelve accepted bodies, four seats × three arms. Counts out of four.
| code | A1 kept | A2 Englished | A3 labelled |
|---|---|---|---|
EXCL — someone in the room reported as unable to follow what was said |
0 | 0 | 0 |
VIA-LANG |
0 | 0 | 0 |
LANG — a language named at all |
1 | 0 | 1 |
OPACITY — something on the page reported as foreign, no person named |
1 | 0 | 1 |
PASTOR-CURE (post-hoc, §4) |
4 | 2 | 4 |
C1, the locality control, passes in all three arms — every one of the twelve retellings names at
least 3 of the 4 registered events, and in fact all twelve name 4 of 4. The manipulation moved nothing
about which events a reader reports.
3. G3 fires: the primary is withheld
EXCL is 0 in twelve of twelve bodies. The pre-registered gate is unambiguous: if all bodies
score identically on EXCL, the instrument separated nothing and PR-PRIMARY is withheld. It is
withheld. No count in the EXCL row licenses any statement about whether a reader recovers the
exclusion, in any arm, and no sentence anywhere else in this project may cite it as though it did.
The baseline the run was rebuilt around — critic pass 2's F4, the observation that the scene carries
the same inference by four routes that are not the foreign line — also returns 0. Not one seat
reported the doctor's averted eyes, the unanswered question, the silence during the writing, the
pastor's "nothing more for us to do here" or Heikura's having to break in, as anybody being
shut out of anything.
The honest reading is about the task, not about the passage. A four-to-six-sentence retelling of
a fifteen-paragraph scene has room for the plot and nothing else, and all twelve retellings spent it
on the plot: the doctor arrives, Mari is bound, a prescription is written, the landlord asks about
his other tenants. Asking for a summary and scoring what it omits measures the compression, not the
reader. That is the RS-20260805b lesson arriving a second time in six sessions, on a different
instrument, and it is now twice on the record.
4. What the run did find, and it disqualifies the instrument for this question
Every A1 seat reported the content of the untranslated line. PASTOR-CURE codes whether a
retelling attributes a question about curing to the pastor. On A1's page that content exists
only inside «Kan hon botas?» — the parallel question four paragraphs later is Heikura's, not
the pastor's, so a pastor-attributed answer cannot have come from it. A1 scores 4 of 4.
Two of them are worth quoting, because they are the result:
openai/gpt-5.6-terra, arm 1 — "The pastor asks in Swedish whether Mari can be cured, but the doctor does not answer."
google/gemini-3.6-flash, arm 1 — "The pastor asks if Mari can be cured and, receiving no answer, prepares to leave."
The first read the Swedish, identified the language, translated it, and said so — gate G1 fires
on exactly this body, since the word Swedish has no warrant anywhere on arm 1's page. The second
read it and translated it silently, reporting the content as though it were in English.
D115's declared cost depends entirely on the target reader not being able to read that line. The
artifact's own front matter declares its reader: a reader with no Finnish, meeting Canth for the
first time, reading for the story — no facing text, no notes. A panel of multilingual language
models is not that reader and cannot be made into one by any prompt. On a question whose whole
content is what the reader does not know, an instrument that knows it measures something else.
And the direction of the one contrast that did move points the same way. PASTOR-CURE is 4 of 4
on A1, where the line is in Swedish, and 2 of 4 on A2, where it is in plain English. With n = 4 this
is not a difference to claim and none is claimed. What can be said is that nothing in the data
suggests the Swedish suppressed the content for these seats, and the raw counts run the other way: the
foreign string was, if anything, more salient than the plain one.
The narrator's label bought nothing. A3 differs from A1 by the two words in Swedish, printed on
the page, and its LANG count is 1 of 4 — the same seat, and no other. Three of four seats did not
mention the language even when the narrator named it for them.
5. What this means for the translation, which is: nothing yet
No conclusion of this run bears on D115. The rendering stands as frozen, for the reasons the log
gives, and those reasons were never that a panel would confirm them. What changes is that the cost
D115 declares is not measurable by this project's present instrument, and the log says so now by
pointer rather than claiming a verdict it does not have.
RS-20260802d's finding — that English is better placed than Swedish to carry this novella's
code-switching — was about the narrated case at ¶130 and is untouched. The printed case
remains what it was before this run: a decision taken on the page for stated reasons, with a declared
cost that has not been priced.
6. Cost
$0.146723450, against a declared worst case of $1.03 — 14%. Key-usage delta 0.146723450 against a per-request sum of 0.146681450, residual $0.000042, which is exactly the price of one four-token diagnostic call made by hand to reproduce a provider error (§7) and is therefore closed to 1e-9.
| what | $ | note |
|---|---|---|
| three pre-run critic passes | 0.060283200 | 41% of the run, and the best money in it |
| twelve scored seat bodies | 0.041938800 | 12 of 12 accepted, zero retries |
six dead qwen3.7-max bodies |
0.044459450 | 30% of the run bought nothing — §7 |
| the hand diagnostic | 0.000042000 | four tokens, and it found the 400 |
The critic cost more than the run it criticised and was worth it. Three passes,
thirty-one findings, all accepted; the question the run finally asked is not the question
revision 1 asked, and the confirmatory test was withdrawn before dispatch rather than after seeing
the numbers. Full record: critic.md.
7. Two instrument defects, recorded rather than repaired quietly
(i) qwen/qwen3.7-max returned finish_reason: length with empty content on all six of its
dispatches, at the frozen cap of 1,500, burning $0.0444 for nothing — note (bhf), and a fourth
distinct slug. The cap was not raised: S113's recorded lesson is that the fix is a changed seat,
not a raised cap, and raising a frozen parameter mid-run is a design change. The seat is dropped
under gate G2 and every arm is reported at n = 4.
(ii) mistralai/mistral-medium-3-5 returned HTTP 400 on all six dispatches, billed $0, and the
diagnosis had been thrown away by the runner. call.py stored repr(e) for a transport failure,
and an HTTPError's repr is <HTTPError 400: 'Bad Request'> — the provider's actual message lives
in the response body, which was being discarded. Reproducing one call by hand returned
"top_p must be 1 when using greedy sampling", which fires only when the reasoning field is
present. call.py now stores the error body; the seat was re-dispatched with the field omitted,
the payload, arms, seat list and max_tokens all exactly as frozen, and returned 3 of 3 on the first
try. A guard that records that something failed but not why is half a guard.
8. Limits, and the ones the critic named that this run did not repair
- The instrument cannot proxy a monolingual reader (§4). This is the run's finding and it is also its largest limit: nothing here describes what a reader without Swedish would do.
- Twelve bodies, four seats, one passage, one line. No general claim about readers, about English, or about code-switching in translation is licensed, and none is made.
- Critic pass 3
F3— the scene's four non-language routes cannot be stripped without rewriting Canth. The design measured them instead of removing them; that is not a repair. - Critic pass 3
F5— keyword coding of free prose is brittle in both directions, andverify.pymeasures the brittleness rather than assuming it away: the frozenEXCLlist misses 4 of 4 plainly exclusionary phrasings put to it as acceptance tests ("over the family's heads", "went past Holpainen entirely", "meant nothing to the people of the house", "only the two gentlemen knew"). SinceEXCLis 0 everywhere, this cuts one way: a retelling that expressed the exclusion in words outside the list would have been scored 0. Blinded dual coding is what would fix it and it was not run. All twelve retellings are stored inruns/and can be read by hand against this claim. - Critic pass 3
F1/F6— n = 4 supports no estimate of magnitude, and the baseline has no failure mode. The confirmatory test was withdrawn before dispatch; no p-value is computed anywhere in this run. PASTOR-CUREis post-hoc. It was coded after the retellings were read, is declared as such here and inverify.py, and carries no pre-registration. It is reported because it is what disqualified the instrument, not because it was predicted.