Repository path: workshop/experiments/E-20260804f-whose-mind/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804f-whose-mind |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | voice, style-correspondence, accuracy |
| internal-judgment-only | true |
| provisional | true |
| track | T1 |
| links | workshop/translations/koyhaa-kansaa/R05-v1/translation.md, workshop/translations/koyhaa-kansaa/register.md, wiki/arms/ARM-atelier-cycle.md, config/models.md, wiki/goodness-senses.md |
E-20260804f — whose mind is this? The free-indirect layer of span 5 in English
Frozen 2026-08-04, before dispatch and after the translation limb was committed at 65e5f8a.
Study limb of ARM-atelier-cycle step 5. The translation limb is span 5 of
T-koyhaa-kansaa-R05-v1 (¶296–375, the night).
1. The wire, in one sentence
Span 5 is the novella's sustained double-interior passage — ten stretches of untagged free indirect discourse alternating between two consciousnesses — and the register has already declared away most of the devices Finnish uses to mark whose mind a paragraph is in, so this experiment asks whether the English still delivers the attribution.
2. Question
Canth marks free indirect discourse with morphemes English has no equivalent of: the clitics -hän
and -pä, the spoken determiner se before a name, evaluative lexis in the character's own
register, and (at ¶167, ¶354) plural subject with singular verb. D93 records the loss explicitly
at ¶366, where Eihän se Mari becomes "Why, Mari"; V3 forbids carrying the colloquial layer as
English dialect; V22 forbids the flat rendering of se.
Can a reader with no access to the Finnish say whose consciousness an untagged passage is in?
Two outcomes are both informative. If yes, the lost morphemes were redundant with what English does carry — context, lexis, deixis — and the register's declared losses cost less than they look. If no, the loss is real, it is concentrated where the translation could not compensate, and the framework should say so about Finnish→English narrative prose.
3. Materials
materials/loci.json, built by build_materials.py from the committed artifact. Sixteen loci
cut from the span-5 English, each 1–3 sentences, each keyed to a verbatim opening string the builder
asserts is unique.
| class | n | ground truth | role |
|---|---|---|---|
| TAG | 3 | the subject of an explicit verb of thinking in the source (hän ajatteli, arveli hän itsekseen, Holpainen ihmetteli) |
positive control — the English keeps the tag, so seats must get these |
| NARR | 3 | narrator — no tag and no character-indexical morpheme in the source; past-tense report of observable events | specificity control — seats must not see a mind everywhere |
| FID | 10 | the character indexed by the source's clitics / se / evaluative register, cross-checked against the nearest preceding tagged sentence |
the measurement |
Ground truth is fixed by the Finnish, not by the English, and not by taste. A locus is FID only where the source has no tag and does have at least one character-indexical morpheme. The two rules — "nearest preceding tag in the source" and "which character the morphemes index" — agree at all ten FID loci; had they disagreed anywhere, the locus would have been dropped before dispatch.
FID truth is 7 MARI / 3 HOLPAINEN, which is deliberately imbalanced and is why §6 benchmarks against the best constant strategy rather than against chance.
4. Procedure
One dispatch per seat. Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5 (config/models.md). Reserve: P5 deepseek/deepseek-v4-pro.
Each seat receives the whole span-5 English — so that context is available, which is the
condition a real reader is in — followed by the sixteen passages in fixed order, and answers for
each: WHOSE ∈ {MARI, HOLPAINEN, NARRATOR}, QUOTE (verbatim words from the text that decided it),
CONF 1–5.
Blinding. Seats see English only. No Finnish, no source, no author, no title, no mention of translation, and no mention of free indirect discourse or of any narratological term — the question is put as an ordinary reader's question about whose thoughts a passage gives. Nothing tells a seat that the passages fall into classes or that any class is a control.
temperature 0, reasoning.effort: low on the first dispatch (note (b)), max_tokens 6,000,
usage.include. Raw bytes to disk before any parse.
5. Predictions, registered before dispatch
- R1 (power). TAG controls: ≥ 8 of 9 cells correct. If R1 fails the instrument is dead and nothing else on this page is read.
- R2 (specificity). NARR controls: ≥ 6 of 9 cells correct.
- R3 (primary). The HOLPAINEN-FID cells (L4, L13, L14 — 9 cells): ≥ 6 of 9 correct. These are the cells the best constant strategy gets wrong, and L14 is D93's own site.
- R4 (secondary). All FID cells: > 21 of 30 correct — 21 is exactly what answering "MARI" to every FID locus scores.
- R5 (evidence). In ≥ 24 of 30 FID cells the
QUOTEis verbatim text present in the span.
The lead's own expectation, recorded so it can be wrong: R1 and R2 hold, R3 fails at 3–5 of 9, R4 holds only because of the Mari cells. That is a prediction that the attribution survives for the character the reader is already with and fails at the switch.
6. Benchmarks and how the result is read
A constant guesser is the benchmark, not chance (the S077 lesson: an instrument that loses to a constant guess has measured nothing). Always-MARI scores 8/16 loci overall and 7/10 on FID. The result is reported per class, with the constant-strategy score printed beside it.
Nothing here is a quality claim about the translation. The loci are not scored for goodness; the
measurement is whether an attribution recoverable in the source is recoverable in the target. Every
evaluative sentence on the result page carries internal-judgment-only and provisional — Tier D is
NOT PASSED, and these seats are coders, not jurors.
7. Failure criteria
- FC1 — a seat returns no parseable body,
finish_reason: length, or fewer than 16 answer lines: one retry atmax_tokens10,000, then the reserve, then the seat is withdrawn and the result reported at N = 2. It is not replaced silently. - FC2 — R1 fails: the instrument cannot do the easy case, the run is a null, and R3/R4 are not reported as findings.
- FC3 — a seat's
QUOTEis not verbatim in the span: that cell is excluded from R5 and retained for R1–R4, and the count is published. - FC4 — a seat answers the same label at ≥ 14 of 16 loci: that seat is not discriminating. Its cells are retained and the fact is reported beside every number it contributes to.
- FC5 — the pre-run critic returns
NEEDS-REDESIGNon the ground truth: the run does not go until the finding is answered in writing.
8. Comparator, after the run, at no cost
Hertzberg's authorized 1886 Swedish (gutenberg.org/ebooks/20518, PD, logged in
wiki/base/consulted.md) at the same sixteen loci: does Swedish keep a marker English drops?
Swedish has the modal particles ju, nog, väl, which sit closer to Finnish clitics than anything
English has. A capability/choice reading by the lead, internal-judgment-only, and not part of
the registered predictions.
9. Budget
Declared worst case $0.55: critic $0.18 (P4, max_tokens 12,000, note (abc)); three seats
$0.16 at list, doubled to $0.32 for the routing caution config/models.md records at S022.
Today's UTC headroom before this run: $3.463398640.
Amendments A1–A9, from the pre-run critic (P4 moonshotai/kimi-k3, critic.md)
Verdict NEEDS-AMENDMENT, nine findings, four BLOCKING. All nine accepted; three of them changed
what this experiment is measuring. Applied before any subject dispatch. §§5–7 above are superseded
where they conflict with what follows.
A1 (F1, BLOCKING) — the primary prediction was rigged and is replaced. Two of the three Holpainen-FID loci name Mari in the third person (L13 "Mari always took the smallest share", L14 "Why, Mari had not eaten that soup"), so 6 of R3's 9 cells are answerable by the rule the person named is not the thinker, with no sensitivity to anything this experiment is about. R3's threshold was reachable on those two loci alone. The critic is right, and it is worse than it says: L14 was the design's showcase because D93 records the loss there — and D93's own compensation, "Why, Mari", supplies the name that gives the answer away. R3 as written is withdrawn.
A2 (F2, BLOCKING) — presentation order is now per seat, and adjacency is forbidden. Extracts were
to be shown in span order, which lets a seat propagate an answer across neighbours (L11→L12,
L13→L14) and puts L4 next to L5, a tagged Holpainen locus. Orders are now generated per seat from
recorded seeds (105001/2/3), stored in materials/orders.json before dispatch, and the builder
asserts that no two adjacent extracts share a ground truth. Extracts are relabelled E1…E16 per seat
so the id carries no information; answers are mapped back by the stored order.
A3 (F3, BLOCKING) — the claim in §2 is narrowed to what the run can support. Seats see the whole span, so a correct answer may come from plot context rather than from anything in the extract. The run therefore cannot show that "the lost morphemes were redundant". What it can show is which English device the attribution rides on, which is what A4 makes the new primary.
A4 — the new primary is a mechanism prediction, on all 30 FID cells. Finnish hän is
genderless; English must choose she or he. In a two-person scene that is an attribution device
the source does not have, and it is available at every locus. R3′ (registered): in ≥ 20 of 30 FID
cells, the QUOTE a seat gives as its evidence is a person-index English supplies — a proper name or
epithet of the other party, or a gendered pronoun — rather than the diction, the register, or the
topic. The lead's expectation, recorded so it can be wrong: R3′ holds, and the English carries the
attribution by a device Canth had no access to rather than by anything that answers to her clitics.
The old R3 survives only as a descriptive split: R3a the name-cued Holpainen cells (L13, L14),
R3b the one non-name-cued Holpainen locus (L4), reported per cell and not a prediction.
A5 (F4, BLOCKING) — the ground-truth cross-check is published, and the design's claim about it was FALSE. §3 asserted that the two rules — nearest preceding tagged sentence in the source and which character the indexing morphemes point at — agree at all ten FID loci. They disagree at L11. ¶357 «Tuhma pökkelö, kun uskoi…» has the nearer thought-tag at ¶355, «Holpainen ihmetteli mielessään», which the mechanical rule would resolve to HOLPAINEN; the epithet rule resolves it to MARI, because a mind that calls Holpainen a stupid blockhead is not his. L11's truth is MARI and rests on the epithet, not on the tag rule, and it is classed as cued (A6) for exactly that reason. The rule is also undefined at L1: ¶296 has no preceding tag inside the span, and the governing tag is ¶295, in span 4, which is Mari's. Both facts are published here rather than left implicit.
| locus | source ¶ | the indexing morpheme in the Finnish | nearest preceding thought-tag | truth |
|---|---|---|---|---|
| L1 | 296 | Mitäpä (clitic -pä), evaluative parempi |
none in span; ¶295 (Mari) | MARI |
| L2 | 297 | tiesipä (-pä), parka, zero person |
¶297 «hän sydämessään … nureksi» | MARI |
| L3 | 299 | ne vaan, fronted evaluative Pahoja |
¶298 «Sitä ajatellessaan Mari» | MARI |
| L4 | 316 | kuitenkaan + potential mahtaisi |
¶316 «Holpainen … koetti sitä karkoittaa» | HOLPAINEN |
| L7 | 334 | evaluative theodicy in Mari's register | ¶334 «sen hän varmaan tiesi» | MARI |
| L10 | 353 | pelkäsivätpä (-pä), sentään, taunt register |
¶353 «Hän päätti vaieta» | MARI |
| L11 | 357 | epithet Tuhma pökkelö of the other party |
¶355 (Holpainen) — rules disagree | MARI |
| L12 | 361 | jokohan (-han), rukka |
¶359 «Olisi tehnyt mieli lyödä häntä» | MARI |
| L13 | 365 | -kin, oikeastaan, se Mari nearby |
¶365 «hän … tuumaili» | HOLPAINEN |
| L14 | 366 | Eihän (-hän) + se before a name |
¶365 «hän … tuumaili» | HOLPAINEN |
A6 (F8) — every FID locus is classed by what cue the EXTRACT itself carries, and results are reported by class. CUED (a name or epithet of the other party, or a gendered pronoun indexing the thinker): L1, L10, L11, L13, L14. NO-CUE: L2, L3, L4, L7, L12. The no-cue set is 4 MARI to 1 HOLPAINEN and cannot carry a balanced test — the span contains only three untagged Holpainen stretches and two of them name Mari. That asymmetry is itself reportable: in this span Canth's Holpainen-interior is short and anchored to a named object, and her Mari-interior is long and unanchored. It is a fact about the materials and is not fixable by choosing different loci.
A7 (F5, F6) — the run is a qualitative probe and every threshold is descriptive, not inferential. Three seats are not three independent draws; they share training data and may have seen this text's neighbours. No binomial claim is made anywhere. R5 is tightened: the quote must come from the extract itself, be ≥ 3 words, and all 30 FID quotes are audited by hand by the lead, not merely membership-tested. Constant-strategy scores are printed beside every class.
A8 (F7) — FC4 is lowered. A seat is non-discriminating if it gives the same label at ≥ 12 of 16 loci, or the same label at ≥ 8 of 10 FID loci. Both are checked and reported.
A9 (F9) — the seats run at DEFAULT reasoning effort, not low. Handicapping the seats and then
using a control to declare the instrument alive builds a self-fulfilling null. Note (b)'s low
default is overridden here for this reason and recorded; the budget is re-declared at $0.75
worst case to cover it.