Repository path: workshop/experiments/E-20260803b-honorific-carry/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260803b-honorific-carry |
| status | frozen |
| created | 2026-08-03 |
| updated | 2026-08-03 |
| senses | style-correspondence, cultural-mediation, naturalness |
| track | T1 |
| links | workshop/translations/koyhaa-kansaa/R05-v1/translation.md, workshop/translations/koyhaa-kansaa/register.md, wiki/arms/ARM-atelier-cycle.md, wiki/findings/results/RS-20260802d-class-line-carry.md, config/models.md, config/budget.md |
E-20260803b — does the one available compensation put the honorific back?
Frozen before dispatch. The translation it tests was frozen first, at 1218096, before the
Swedish comparator was opened (R05 / charter A4).
AMENDED TWICE BEFORE ANY SUBJECT CALL. Pass 1: NEEDS-REDESIGN, nine findings, five BLOCKING →
amendments A1–A9. Pass 2, on the amended design: NEEDS-AMENDMENT, nine findings, three BLOCKING →
amendments B1–B9. All eighteen accepted. Record: critic.md. A1 and A3 are the ones that
matter — as first written, no outcome of this experiment could have changed the translation, and the
critic said so — and B6 is the one that would have wasted half the run: call.py posts
temperature: 0, so the two "repeats" were the same draw and would have double-counted in H1b and
in the permutation test.
1. Where this comes from
Span 3 of «Köyhää kansaa» contains the novella's strongest instance of a feature English has no grammar for: at ¶175 a nine-year-old girl wakes her mother in front of two visiting ladies with «Nouskaa ylös!» — a second-person plural imperative addressed to one person, her own mother.
register.md reserved this site for span 3 and V17 required the decision to be taken in place.
Decision D57 took it, and found that V14a's licensed compensation — the kin term, on
Hertzberg's model — is unavailable here, because Canth has already spent the vocative («Äiti» four
times in the three surrounding utterances). Of the two English carriers that would work exactly,
"ma'am" is closed by V12 (it is this book's rouva address) and the thou/you contrast is
closed by V3 (no regional English anywhere in this book). What was taken instead is the period
politeness formula: "Mother, there's ladies here. Please to get up!"
The translator cannot answer the question that decision raises. Whether "please to" reads as a marking of standing or as an unmotivated politeness is a fact about readers, and the lead never judges its own translation (charter §5).
2. Question (restated by A6)
Does "please to" restore, for this project's standing blind-seat instrument, the marking that the Finnish plural imperative carries — and does it do so above what a seat assumes about the scene without reading it at all?
A6, and it is a real limit rather than a form of words. The first draft asked this about "a
reader". The cells are three language models, and nothing here establishes that their 1-to-5
ratings track a human reader's recovery of a period politeness formula. The prescribed remedies —
human raters, or a frozen human–LLM concordance study on comparable period items — are not
available to this project and are not pretended to be. Tier D is NOT PASSED (config/models.md),
so no verdict below carries evidential weight, and the result page states that before it states
a number, not in a limits section.
3. Design
Three arms, identical but for one phrase, put blind to three non-Anthropic seats.
| arm | what the seat is given |
|---|---|
| M (marked) | the frozen English of ¶171–181 as filed, containing "Please to get up!", plus a second passage (¶199–200) |
| P (plain) | byte-identical to M except that "Please to get up!" reads "Get up!" |
| N (no information) | no passage at all — only a neutral two-line synopsis (date, town, the girl's age, a sleeping mother, two visitors) and the two rating questions |
A5. N's synopsis originally read "very poor", "charitable ladies", "worn out", "heap of clothes" — four stereotype primes the M/P seats are never given. Rewritten neutral, so that N measures prior expectation about the situation and not about a poverty tableau the other arms do not describe.
Arm N is the no-information benchmark that note (bht) requires: a seat that never reads the text, answering from prior expectation about an 1886 Finnish family. It was registered here, before dispatch, and not added after seeing a result.
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro
(config/models.md; the S089/S094 jury, so this extends an instrument rather than inventing one).
Stateless, one call per cell, never shown this design, the other arm, the Finnish, the Swedish, or
who wrote the English.
Cells: 3 arms × 3 seats × 2 question orderings × 1 repeat = 18 calls. (B6: dispatch is
at temperature: 0, so a second repeat is the identical draw. Repeats cut to one rather than
raising the temperature, which would break comparability with the S089/S094 instrument; C3's
denominator halves with it.)
Questions (the rating questions are counterbalanced; QUOTE is always last):
- Q1 (
GIRL) — On a scale of 1 to 5, how does the girl speak to her mother? 1 = exactly as she would to an equal or a playmate; 5 = with marked respect or formality, as to a superior. - Q2 (
MOTHER) — On the same 1-to-5 scale, how does the mother speak to her daughter? - Q3 (
NATURAL, added by A3) — On the same 1-to-5 scale, how does the girl's own speech sound? 1 = odd, stilted or wrong in the mouth of a nine-year-old in this room; 5 = exactly like a nine-year-old child speaking. (Arms M and P only.) - Q4 (
QUOTE) — Quote the exact words, anywhere in what you were given, that mark a difference of standing between any two speakers. If there are none, write NONE. (Arms M and P only.)
MOTHER is the specificity control: the two arms differ only in the girl's line, so the mother's
direction must not move. NATURAL is the control A3 added and the first draft had no equivalent
of — it is the only item that can express the failure the craft actually fears, that "please to"
reads as period furniture and is worse than the plain imperative. QUOTE is the attribution
control: a rating difference no seat can point at is not evidence about the device.
Ordering 1 asks GIRL, MOTHER, NATURAL; ordering 2 asks NATURAL, MOTHER, GIRL. The emitted answer lines follow the gloss order (B5 — the first draft asked ordering 2 to read the glosses in one order and emit them in another), and parsing is by label, never by position. QUOTE is always last, so no seat is asked to quote before it has rated.
B4. The M/P header no longer says "a very poor family": N was stripped of its class priming by A5 and the other two arms kept theirs, so C2 was comparing a primed reading against an unprimed expectation. All three arms now open on the same neutral frame — a novel set in a Finnish town, published in 1886 — and the passages carry the poverty themselves.
4. Registered criteria — frozen before dispatch, amended before dispatch, never after
| # | criterion | bar |
|---|---|---|
| H1 | recovery: mean GIRL under M minus mean GIRL under P |
≥ 1.00 scale points |
| H1b (A8) | robustness: seats moving ≥ 0.50 in the same direction | ≥ 2 of 3 |
| C1 | specificity: absolute movement of mean MOTHER between M and P |
≤ 0.50 |
| C2 (A4) | no-information (note (bht)), directional: mean GIRL(M) ≥ mean GIRL(N) |
must hold; M−N and P−N both reported |
| C3 (A7, B3, B6) | attribution: arm-M cells whose normalized QUOTE contains please to |
≥ 3 of 6; arm-P quotes tabulated verbatim as the symmetry check |
| C4 (A3) | naturalness cost: mean NATURAL(P) minus mean NATURAL(M) |
≤ 0.50 |
An exact permutation test over the M/P labels is computed and reported (A8) as information, not
as a gate — a gate invented now would be the very fault A9 names. Raw GIRL distributions are
printed for every arm (B8): if both arms park at 4 from the situation alone, a one-point mean
move is unreachable for reasons that are not "the device is inert", and the reader must be able to
see that in the numbers rather than take it on trust.
C3's matcher is frozen in code before dispatch (B3): analysis/score.py normalizes a quote by
casefolding, deleting everything outside a–z and space, and collapsing whitespace; an arm-M quote
is accepted iff the result contains please to. Arm-P quotes are not gated and are tabulated as
returned. The rule is committed with this design, so it cannot become what the numbers need.
On the 1.00 bar (A9), stated rather than moved. The critic is right that it sits just above the lead's own predicted Δ and that it has no instrument history behind it. Moving it after reading that criticism would be the same fault in the opposite direction, so it stands, and the conclusion rests on H1b, the seat-wise table and the permutation test rather than on the bar alone.
4a. Consequences, pre-committed (A1–A3; completed by B1, B2, B7, B9)
As first written, every outcome left the rendering in place and §6 explicitly pre-banned reversion. The critic named this as its first BLOCKING finding and it was correct.
B1, and it is the frame for everything in this section. §2 says no verdict from this instrument
carries evidential weight, and §4a then makes a translation change turn on those verdicts. That is a
real contradiction and it is resolved in one direction: every consequence below is an
instrument-internal atelier decision, binding on register.md as a workshop rule and never as a
claim about human readers. The result page leads with the Tier D refusal before it prints a number
or the word erratum. A reversion here means this project's blind seats could not recover the
device, which is a reason for a translator to drop a flourish and is not a finding about English.
The full decision table (B2 — the first draft covered four cells and left four live outcomes unnamed, which re-created A1's fault in miniature). Read top to bottom; the first row that matches governs.
| # | condition | consequence |
|---|---|---|
| 1 | C4 fails (M reads less natural than P by > 0.50), or C4 cannot be scored | Erratum 4: revert to "Get up!", V18 retired. (B7: reported as a naturalness result, not as a failed recovery test — the two are different findings and C4 can fire while H1 is untested.) |
| 2 | Δ(M−P) < 0.50 | Erratum 4: revert, V18 retired with the reason written |
| 3 | Δ ≥ 0.50 and any of H1b, C1, C2, C3 fails | Erratum 4: revert. (B2's default: a control failure below the H1 bar reverts. The lead does not keep the sentence on a half-result with a broken control.) |
| 4 | Δ ≥ 1.00, any of H1b, C1, C2, C3 fails | indeterminate, no reversion, mandatory craft note naming the failed control; V18 carries the same terminal condition as row 5 |
| 5 | 0.50 ≤ Δ < 1.00, C1–C4 all holding | indeterminate. (B9: "unevidenced but kept" is not allowed to be a stable end state. *V18 stands only until the next honorific-to-a-parent site in spans 4–8, where it is re-tested as a condition of that span; if no further site exists, the craft report must record that V18 governed one sentence in 20,000 words and was never evidenced.) |
| 6 | Δ ≥ 1.00 and H1b, C1, C2, C3, C4 all holding | V18 is evidenced on this instrument, at the strength §2 allows and no more |
The lead does not get to keep the sentence by predicting that the test will fail.
5. The lead's registered prediction — that its own compensation fails
Written before dispatch. It is a prediction, not a judgment of the translation, and it is recorded so that a hit is not narrated afterwards as foresight.
- Δ(M−P) ≈ +0.6 — visible, below the bar.
- Arm N ≈ 3.5 — lower than the pre-amendment prediction of 4.0, because A5 stripped the poverty primes out of N's synopsis. Under the directional C2 (A4) this no longer vetoes anything by construction; it is now a real comparison and it may go either way.
MOTHER≈ 1.5 in both arms, no movement — C1 holds.- C3 holds at ≈ 8 of 12 — the device is quotable whether or not it moves a rating.
NATURAL: M ≈ 3.0, P ≈ 4.0, so Δ ≈ 1.0 and C4 FAILS, firing row 1 of the table.
The predicted verdict is therefore that Erratum 4 fires on C4, and that the sentence the lead wrote this morning does not survive the day.
B7, taken as written and applied to this section. C4's low anchor is "odd, stilted or wrong in the mouth of a nine-year-old", and the manipulated string is an adult servant-class formula the lead already believes does not belong in that mouth — so C4 is close to guaranteed and a C4-only reversion is a tripwire firing, not a recovery test failing. If it fires, that is one sentence in the result page and no more; a predicted self-criticism is not a methodological achievement, and narrating it as one would be its own kind of theatre.
6. What the result does not change
The §7 comparator finding stands whatever the numbers do, because it does not depend on them. So does D57's account of why the kin term was unavailable, which was reached before the Swedish was opened. What the run bears on is one string in one sentence, and §4a says exactly what happens to it.
7. The comparator, opened after the freeze, and it is not part of the probe
Rafaël Hertzberg's authorized 1886 Swedish «Bland fattigt folk» (Gutenberg 20518), stored at
materials/. Swedish has the V-form (ni) and Hertzberg uses it eight times in this very span,
on the road, between Mari and the Karttula woman. At ¶175 he writes «-- Mamma, här är fruar. Stig
upp!» — the singular imperative. He had the form, twenty paragraphs after using it, and did not
take it.
This is recorded in the result page as an independent observation about the site. It is not a control on the probe and is not used to score anything: one translator is one datum, and what it establishes is that the loss at this site is a choice in a capable target, not a limit.
8. Budget
Pre-flight, built from max_tokens and not from an expected answer length (note (abc)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
pre-run critic (P3 x-ai/grok-4.5, takes no part in scoring) |
1 | 24,000 | $0.16 |
second critic pass on the amended design (the first verdict was NEEDS-REDESIGN) |
1 | 24,000 | $0.16 |
| scoring, 3 arms × 3 seats × 2 orderings × 2 repeats | 36 | 4,000 | $0.81 |
| retry reserve (one full re-dispatch of the scoring stage) | 36 | 4,000 | $0.81 |
| declared worst case | $1.94 |
Against $4.350582573 headroom on UTC 2026-08-03 (config/budget.md). Fits with $2.41 to spare.
A stage that will not fit is split or deferred, not run.
9. Failure criteria for the run itself
- A cell that returns fewer than the required answer lines after two attempts is a cell failure, recorded by id.
- More than 3 of 36 cells failing → the run is not reported as a measurement, only as a dispatch record.
finish_reason: "length"is a seat failure, never a partial answer (note (b)).- Raw bytes to disk before any parse (note (bdt)); every number in the result page recomputed by
analysis/verify.py, which imports nothing from the runner.