Repository path: workshop/experiments/E-20260812c-grade-shift/design-v1-superseded.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260812c-grade-shift-v1 |
| status | superseded |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| senses | voice, affect |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/translations/dakghar/R05-v1/translation.md, workshop/translations/dakghar/register.md, wiki/arms/ARM-dakghar.md, wiki/findings/results/RS-20260811f-dakghar-address.md, config/models.md, config/budget.md |
[SUPERSEDED BEFORE ANY DATA] What makes a second-person grade shift survive translation?
This design was frozen, sent to an independent pre-run critic, and killed by it. It was never run: no translation call, no arbiter call, no datum. The critic (
openai/gpt-5.6-terra, one call, $0.04057375,finish_reason: stop) returnedNEEDS-REDESIGNwith 35 findings, 27 of them BLOCKING. Four were decisive and all four are accepted:
- Finding 14 — the lead's own grammar-only rendering supplies a lexical cue.
[190], the item defined as grammar-only, is rendered "And why not, boy!". The cell was contaminated by the very translation it was measuring — and the lead's sealed claim at logD33had called that site a loss. This is now the object of the replacement design rather than a confound inside it.- Findings 2, 3, 23 — one grammar-only item, two lexical items, all three sharing one comparator. The cells were
n = 1andn = 2and the observations were dependent; the registered bootstrap over pairs was meaningless.- Findings 12, 17, 27, 35 — the arbiter question measured overall perceived respect, from three model seats, on speeches stripped of their scene. No result it could have produced would have licensed a sentence about English readers.
- Findings 21, 22 —
recoveryand the false-positive rate were not commensurable, soP2's inequality had no common null.The replacement is
design.mdin this directory. It keeps the question, drops the reception claim, replaces the selected items with a census of the whole span, and makes the translator's behaviour rather than a reader's perception the thing measured. The full critic body isruns/critic.json.
Study limb of ARM-dakghar step 2 (T1). Frozen before any call. The translation it hangs on,
T-dakghar-R05-v1 span B, was frozen and committed at 231e084 before this page was written,
and the comparator (Mukherjea 1914) has not been fetched into this repository at any point.
1. The question, and where it came from
Translating section ২ of «ডাকঘর» raised it (log D32, D33). The binding register sent span B
one required question: does any English device carry the আপনি/তুমি contrast at a site where the
contrast is on stage? The span turned out to contain no deferential আপনি on stage at all. What
it contains instead is the other direction — the headman drops from তুমি to তুই when he is
angry with a dying child, and climbs back inside the same speech.
So the question the material actually poses is not up or down but what the shift is made of:
When a Bengali speaker drops to
তুই, sometimes the source marks the drop twice — in the grammar and in a contemptuous noun (ওরে ছোড়া, you brat;কোথাকার বাঁদর, what monkey is this) — and sometimes only in the grammar (কেনরে,তোর খবর). English has no grammatical slot at all. Does the shift reach an English reader only when the source marked it twice?
RS-20260811f already established the ceiling case for this arm: on manipulated minimal pairs
differing in one deference token, two independent English hands produced text three arbiters could
not tell apart, at exactly the floor. That was the upward contrast, manufactured. This is the
downward contrast, unmanufactured, with the cue composition as the variable.
2. Materials — natural text, one comparator, three treatments
Six pair-items, each two whole speeches by one character, taken verbatim from the frozen
copy-text; nothing is edited, spliced or manipulated. build_items.py asserts every Bengali speech
against source-ipublishinghouse.txt and every English speech against the frozen translation before
dispatch.
The three shift pairs share a single comparator, [184] — the headman to Amal, তুমি,
তোমার নামে চিঠি! — so the three treatments differ from the same baseline speech, by the same
speaker, to the same addressee, inside one scene:
| id | cell | treatment speech | what marks the drop in the Bengali |
|---|---|---|---|
DN-G-1 |
GRAMMAR-ONLY | [190] |
কেনরে, তোর খবর — তুই morphology, no pejorative noun |
DN-L-1 |
GRAMMAR+LEXIS | [188] |
তোদের and ওরে ছোড়া |
DN-L-2 |
GRAMMAR+LEXIS | [176] |
কে রে and কোথাকার বাঁদর এটা |
Three NEG controls, pairs of speeches by one speaker to one addressee at a constant grade
throughout: NEG-1 curd-seller [109]/[121], NEG-2 watchman [131]/[167], NEG-3 Sudha
[208]/[212].
3. Subjects and seats
Three English subjects per pair. Two independent non-Anthropic hands (H1 = moonshotai/kimi-k3,
H2 = deepseek/deepseek-v4-pro) translate each pair blind — no mention of pronouns, deference,
grade, footing, or of an experiment — plus the lead's own frozen English, judged blind alongside
them (charter §3, A4). The lead's rendering is never identified to any seat.
Three arbiter seats (openai/gpt-5.6-terra, google/gemini-3.6-flash, x-ai/grok-4.5) see one
pair at a time, with no source, no author and no other item, and answer one forced question:
Below are two speeches by the same character in a play. In which of the two does the speaker treat the person he or she is speaking to with more respect — A, B, or the same?
Answer is one token: A, B or SAME. A/B presentation order is flipped by
(seat_index + pair_index) % 2 so no item is seen in one order only.
No seat both translates and arbitrates. The lead neither translates for the panel nor arbitrates, and never judges its own rendering (charter §5).
Source-side ceiling. The same three seats answer the same question on the Bengali speeches, which fixes what is there to be recovered.
4. Measures
For one pair and one text, recovered = 1 if the seat names the less respectful member the
source marks as less respectful (the treatment speech, coded A in items.json), else 0. For a
NEG pair, false positive = 1 if the seat names either member rather than SAME.
recovery(cell)= mean over {3 shift-or-NEG pairs in the cell} × {3 English subjects} × {3 seats}.source(pair)= mean over 3 seats on the Bengali.FP= mean false-positive rate over the threeNEGpairs, all subjects, all seats (27 cells).
5. Gates — both are withholding gates
G1(source-side). Each shift pair must reachsource(pair) ≥ 0.67on the Bengali. A pair that fails is void and is dropped from the primary, because there is no established source-side difference for English to lose. If fewer than one pair survives in either cell, the primary is not read.G2(false positives).FP ≤ 0.33. If arbiters call a difference on constant-grade text at a higher rate than that, the instrument is reading something other than footing and the primary is not read.
6. Predictions, registered before dispatch
P1(primary).recovery(GRAMMAR+LEXIS) > recovery(GRAMMAR-ONLY), with the 90% bootstrap interval over pairs excluding 0.P2.recovery(GRAMMAR-ONLY) ≤ FP + 0.17— the grammar-only drop does not survive into English at a rate distinguishable from the false-positive rate.P3.recovery(GRAMMAR+LEXIS) ≥ 0.67.P4. The three English subjects do not differ from each other by more than 0.33 inrecovery, pooled — i.e. this is a property of English, not of a particular hand.P5.source(DN-G-1) ≥ 0.67— the grammar-only drop is visible in the Bengali. If this fails,P1andP2are unreadable and the finding is about the item, not about English.
Failure criteria. P1 fails if the interval includes 0 or the sign reverses. P2 fails if
grammar-only recovery exceeds FP + 0.17. Any gate failure withholds the primary and the page says
so before it says anything else.
7. What this cannot show, declared in advance
- Three shift pairs, two cells. The item count is small and is set by the play: section ২
contains exactly three
তুইsites. Intervals will be wide and are reported as such. This design cannot be made larger without manufacturing text, which is what it exists to avoid. - Length is not balanced. The comparator
[184]is 8 Bengali words;[188]and[190]are 46 and 33. A seat could be reading length or elaboration rather than footing. TheNEGpairs are also length-unbalanced, which is what makesG2a real check on this rather than a formality. - One direction only. The upward contrast is not re-run;
RS-20260811fhas it, on a cleaner (manufactured) design, at zero. Any statement here about direction is a comparison across two designs and is labelled as such. - One work, one translator pair, one language. Bengali
তুইis not every language's intimate pronoun and the headman is not every rude speaker. - Tier D has not passed. No seat's judgement carries evidential weight; every evaluative
sentence downstream is
provisionalandinternal-judgment-only.
8. Cost
Pre-flight ceiling $0.50, built from max_tokens as note (abc) requires: 12 translation calls
(cap 1,200 out) + 54 English arbiter calls + 18 source-side arbiter calls (cap 400 out) + 1 pre-run
critic call (cap 8,000 out). A body that returns finish_reason == "length" is re-dispatched once
at the same cap whether or not it has content — the per-design rule standing in for the shared
defect at note (bmb).