Repository path: workshop/experiments/E-20260811f-dakghar-address/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811f-dakghar-address |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | — |
| links | workshop/translations/dakghar/R05-v1/translation.md, workshop/translations/dakghar/register.md, wiki/arms/ARM-dakghar.md, wiki/findings/results/RS-20260808g-two-pasts.md, config/models.md, config/budget.md |
E-20260811f — when the source grammaticalises a relation English has no slot for, is the mark relocated or lost?
Frozen 2026-08-11, before any call. The translation limb it hangs on
(T-dakghar-R05-v1, span A) was frozen and committed at 78b6e17 before this design was
written.
1. Where the question came from
It came from translating, at log decision D13. Two adjacent lines of «ডাকঘর» span A:
[71] মাধব: পিসিমা কি বল্লে? [72] অমল: পিসিমা বল্লেন, তুমি ভাল হও…
The same woman, named the same way, in two consecutive speeches, with one token different: the
child gives his aunt the honorific conjugation and her husband does not. Bengali marks the footing
between people grammatically — a closed system of second-person pronouns (আপনি/তুমি/তুই) and
honorific verb agreement. English has one form. Register rule V5 decided to mark none of it
and to log every site as a loss.
RS-20260808g-two-pasts asked the same shape of question about Hungarian's two past tenses in
speech tags and found the mark relocated: Hungarian marks the attribution slot with a tense,
English with a bleached said. This design imports that instrument and puts it on a system that
is about people rather than about syntax, and where the loss looks total rather than compensated.
2. Question
At a site where a single Bengali token fixes the footing between two speakers, does an independent English hand — one that has never heard of this experiment — put anything into its English that a reader can tell apart from the English of the other variant?
Two outcomes are informative and neither is assumed. Relocated: independent hands differ, and the arbiters can name what carried it. Lost: independent hands produce English an arbiter cannot tell apart, at the same rate as the same hand differs from itself.
3. Materials — eight minimal pairs, one token apart
Every pair is a passage of «ডাকঘর» span A, from the frozen copy-text
source-ipublishinghouse.txt, with variant A as printed and variant B differing in exactly
one token. Nothing else in the passage changes.
Six DEFERENCE pairs — the edited token is a deference form and nothing else:
| id | unit | who | A (as printed) | B (edit) |
|---|---|---|---|---|
D-a |
[9] | kabiraj → Madhab | পারবেন না |
পারবে না |
D-b |
[8] | Madhab → kabiraj | দিয়ে যান |
দিয়ে যাও |
D-c |
[43] | Amal, of his aunt | ডাল ভাঙেন |
ডাল ভাঙে |
D-d |
[71] | Madhab, of his wife | কি বল্লে |
কি বল্লেন |
D-e |
[72] | Amal, of his aunt | পিসিমা বল্লেন |
পিসিমা বল্লে |
D-f |
[22] | Grandad → Madhab | তোমার ভয় |
আপনার ভয় |
Two CONTROL pairs — the edited token is not a deference form and must change the English:
| id | unit | A | B | what changes |
|---|---|---|---|---|
C-a |
[71] | পিসিমা কি বল্লে? |
পিসেমশায় কি বল্লে? |
which person is being asked about |
C-b |
[41] | যেতে পারব না? |
যেতে পারলাম না? |
future ability → past |
Direction is balanced inside DEFERENCE. Three pairs edit honorific → ordinary (D-a, D-c,
D-e) and three edit ordinary → honorific (D-b, D-d, D-f), so a hand that simply renders the
second text it is given differently cannot produce the effect.
4. Procedure
Stage T — the independent hands. Two non-Anthropic panel seats, neither the lead and neither an arbiter, are given one variant at a time, alone, with no mention of deference, honorifics, address, or that an experiment exists. The prompt asks for an English translation of the passage and nothing else.
T-a=moonshotai/kimi-k3,T-b=x-ai/grok-4.5(config/models.mdP4, P3).- Each hand renders all 16 variants once (
run1), and each pair's A-variant a second time (run2), in an independent call. 24 calls per hand, 48 calls. run2exists to measure the floor: how much one hand differs from itself on identical input.
Stage J — the arbiters. Three non-Anthropic seats, none of them a translating hand, are shown two passages at a time and asked one question, blind to which condition the item is and blind to the design.
J1=openai/gpt-5.6-terra,J2=google/gemini-3.6-flash,J3=deepseek/deepseek-v4-pro(config/models.mdP1, P2, P5). No seat both translates and arbitrates.- The question, fixed: "These are two versions of the same moment in a play. Ignoring differences
of wording that make no difference to the sense, do the two versions convey the SAME thing about
the people speaking and how they stand to each other, or something DIFFERENT?" Answer as JSON,
{"verdict": "SAME"|"DIFFERENT", "what_differs": "..."}. - Item classes, 40 items × 3 seats = 120 calls:
SOURCE— the two Bengali variants of each pair. 8 items.CROSS— hand h's English of A (run1) against hand h's English of B (run1). 8 pairs × 2 hands = 16 items.FLOOR— hand h's English of A (run1) against hand h's English of A (run2). 8 pairs × 2 hands = 16 items.- Item order is shuffled under a seed fixed in the built file; which passage is shown first is randomised per item and recorded.
Total, after the critic's amendment added C-c: 9 pairs → 54 hand calls + 45 items × 3 seats =
135 arbiter calls = 189 calls, plus the pre-run critic pass already spent.
Measure. d = 1 if the seat returns DIFFERENT, 0 if SAME. A cell that does not parse is
void and reported, never imputed.
5. Registered gates and predictions
Written before any call. Every gate is a withholding gate: if it fails, the primary is not read.
G1— the source checksum, DEMOTED by the critic's BLOCKING 1. OnSOURCEitems, theDEFERENCEpairs must come backDIFFERENTat mean ≥ 0.667. It is a checksum, not a validity gate, and it is asymmetric: a pass licenses nothing about pragmatic sensitivity in English, because these seats know Bengali honorifics metalinguistically; a fail would still be decisive, because a null on the English would then say nothing at all. Reported as a checksum and never cited as evidence that the instrument is sensitive.G2a— the lexical/tense control.C-aandC-bonCROSS, mean ≥ 0.75.G2b— the register control, ADDED on the critic's SERIOUS 3, and it is the gate that matters.C-conCROSS, mean ≥ 0.75.C-cchanges the footing (a child who has said Uncle all scene says Madhab-babu) at a site where English does have a slot. If the arbiters cannot see a footing change in English when English carries it, a null on the deference pairs measures the arbiters and not the translation. This is the same finding the critic ofRS-20260808gmade, note (bkt).G3— the floor. OnFLOORitems, meandmust be ≤ 0.40 (raised from 0.25 on the critic's SERIOUS 4: the primary is a difference, so a floor raised by benign paraphrase costs power and not validity). The run is void only ifFLOOR ≥ CROSS.P1— the primary, registered as a null and able to fail. On the five verb-agreementDEFERENCEpairsD-a…D-e, meand(CROSS) − meand(FLOOR) ≤ 0.20 ⇒ the mark is lost: an independent hand puts nothing into English that an independent reader can tell apart.D-fis excluded from the primary on the critic's BLOCKING 2 — it is a pronoun swap andআপনারfrom Grandad to Madhab can read as estrangement rather than deference — and is reported separately by name. A bootstrap 90% interval over pairs is reported alongside the point estimate (critic ADVISORY 5); if it straddles 0.20 the verdict says so and the threshold is not treated as a boundary the data can resolve.P2— the alternative, and it is the one that would be the bigger finding. IfCROSS−FLOOR> 0.20, the mark is relocated, andP2is only satisfied if the arbiters'what_differstext names an English device (a vocative, a modal, a contraction, a choice of verb) at ≥ half the pairs that drove it. A numerical difference with no nameable carrier is reported as unexplained, not as relocation.P3— per-pair, registered because the mean can hide it. IfP1holds overall but any singleDEFERENCEpair reachesCROSS−FLOOR≥ 0.50, that pair is reported by name as a site where the mark did travel, and the overall null is stated as predominantly, not universally.
Failure criteria. The run is void if: fewer than 90% of cells parse; any hand returns a body
that is not a translation (a refusal, a commentary, a transliteration); or the arbiters return
DIFFERENT on FLOOR at a rate above CROSS (which would mean the measure is noise).
6. The comparator, read as a translation — $0, after the run is dispatched
ARM-legend's craft report closed saying its comparator had been "measured for contamination and
not read as a translation" (ES-20260810-craft-report-legend §5). This arm does the other thing
at its first span. After Stage T is dispatched, the lead opens Devabrata Mukherjea's English
(Macmillan, 1914) and records, at each of the six DEFERENCE sites, what an authorised
human hand of 1914 did — with no scoring and no claim that agreement validates anything. It is a
description placed beside the measurement, and the result page says so.
7. What this design cannot establish
- Two hands is two hands. A null here is a null about
T-aandT-bon this material, not about English. - Six pairs. The denominator is small and
P1's margin is not a confidence interval. - The arbiters share a training distribution with the hands. Four of the five seats used are frontier models; "an independent reader cannot tell them apart" means these arbiters.
- The edits are the lead's. Variant B is not attested text. The
SOURCEgate is what stops that from being fatal, and it is a gate rather than a report for exactly that reason. - Nothing here scores the lead's own translation, which no seat sees. The lead never judges its own translation (charter §5).
8. Declared deviations
- No
senses:are invoked. This is a descriptive measurement of what survives into English, not an evaluation, so no goodness sense is named and no anchor citation is owed. Any evaluative sentence on the result page will carryinternal-judgment-only. - The lead is not blind to the target sites — it chose them, from its own translator's log. The design's protection is that the lead neither translates nor arbitrates: every number comes from seats that saw only a passage and a question.
9. Pre-flight cost estimate
Built from max_tokens, not from expected answer length (note (abc)).
| stage | calls | prompt (est.) | max_tokens |
worst-case |
|---|---|---|---|---|
pre-run critic, qwen/qwen3.7-max (reserve slug, not a seat) |
1 | ~4,500 | 12,000 | $0.08 |
Stage T, T-a moonshotai/kimi-k3 ($3/$15) |
24 | ~250 | 500 | $0.20 |
Stage T, T-b x-ai/grok-4.5 ($2/$6), reasoning low |
24 | ~250 | 800 | $0.13 |
Stage J, J1 openai/gpt-5.6-terra ($1/$6), reasoning off |
40 | ~400 | 900 | $0.23 |
Stage J, J2 google/gemini-3.6-flash ($1.5/$7.5), reasoning low |
40 | ~400 | 1,200 | $0.38 |
Stage J, J3 deepseek/deepseek-v4-pro (list $0.435/$0.87, priced at the 4× routing caution) |
40 | ~400 | 900 | $0.15 |
| 169 | $1.17 |
Declared ceiling: $1.40, raised from $1.25 with the reason written: the pre-run critic's
SERIOUS 3 added a ninth pair (C-c), which adds 6 hand calls and 15 arbiter calls (+12%). UTC day 2026-08-11 stands at $3.294076857 of $5.00 before this run,
so the ceiling leaves $0.46 of the day's cap unspent even if every call bills its worst case.
max_tokens is set at 900–1,200 for the arbiters because S148 truncated a seat's hidden reasoning at a
500 cap and S157 truncated 25 bodies at 2,400 — the cap prices the estimate and must also be large
enough to hold an answer.
10. Pre-run critic — adjudication
One pass, qwen/qwen3.7-max (reserve slug, neither a hand nor an arbiter), cap 12,000,
finish_reason: stop, $0.015822325. Verdict NEEDS-AMENDMENT, 6 findings — 2 BLOCKING,
2 SERIOUS, 2 ADVISORY. Full text at critic.md. Four accepted, two accepted in part with the
overrule written.
| # | finding | adjudication |
|---|---|---|
| B1 | G1 is tautological — the arbiters know Bengali honorifics metalinguistically, so a pass validates the tokenizer, not the stimulus |
ACCEPTED IN PART. G1 demoted from validity gate to asymmetric checksum and the design now says a pass licenses nothing. Overruled on removal: a fail is still decisive, and 27 calls is a cheap price for a signal that could void the primary. |
| B2 | D-f is confounded — আপনার from Grandad to Madhab can read as estrangement or sarcasm, not deference |
ACCEPTED. D-f excluded from the primary, which is now the five verb-agreement pairs, and reported separately by name. |
| S3 | the positive control is lexical, so passing it proves only that arbiters can read nouns | ACCEPTED, and it is the same finding note (bkt) records against RS-20260808g's critic. C-c added: a footing change at a site where English does have a slot, promoted to G2b, the gate that matters. |
| S4 | FLOOR conflates noise with benign paraphrase, so G3 could void a working instrument |
ACCEPTED IN PART. G3 raised to 0.40 and the void condition narrowed to FLOOR ≥ CROSS. Overruled: the proposed extra STYLE-FLOOR arm is what FLOOR already is — two runs of identical input by one hand differ in style and nothing else. |
| A5 | the 0.20 margin is razor-thin at n = 6 | ACCEPTED. Bootstrap 90% interval over pairs reported alongside. |
| A6 | the arbiter prompt's "how they stand to each other" invites style-driven false positives | ACCEPTED. Prompt rewritten before dispatch to name social relationship explicitly and to say that wording or style alone is not a difference. |