Repository path: workshop/experiments/E-20260812c-grade-shift/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260812c-grade-shift |
| status | frozen |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| senses | voice, affect |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260812c-grade-shift/design-v1-superseded.md, workshop/translations/dakghar/R05-v1/translation.md, workshop/translations/dakghar/register.md, wiki/arms/ARM-dakghar.md, wiki/findings/results/RS-20260811f-dakghar-address.md, config/models.md, config/budget.md |
When the source marks contempt only in the grammar, does the translator put it back as a word?
Study limb of ARM-dakghar step 2 (T1). Second design. The first
(design-v1-superseded.md) was frozen, criticised and killed before any datum — 27 blocking
findings — and its finding 14 is the reason this one exists. The translation both hang on,
T-dakghar-R05-v1 span B, was frozen and committed at 231e084 before either design was
written.
1. The question
English has no second-person grade. Bengali has three. When a Bengali speaker drops from তুমি to
তুই — the move from neutral address to intimate-or-contemptuous address — the translator into
English has no grammatical slot to put it in, and the binding register of this work (V5, V6)
forbids inventing one: no thou, no archaism, no compensating stiffness.
So what does the translator actually do? The hypothesis this design tests is that the rule is obeyed in the letter and broken in the substance: that where the source marks the drop only in the grammar, the English quietly acquires a lexical mark — a pejorative or diminutive vocative, an insult noun, a status noun — that the source does not have at that point.
And there is a second question folded into the first, which is the one this project keeps finding:
does the translator know he has done it? The lead's translator's log sealed a claim at D33,
written before any count and before the critic saw anything:
"the downward shift survives at
[176](monkey) and[188](Here, boy), and is lost at[190],[238]and[242]."
That claim is now a registered prediction with a prior count of 0 added cues at grammar-only sites, and it is tested against the frozen artifact.
2. Materials — a census, not a selection
Every second-person grade site in span B, found by exhaustive search of the copy-text for the
তুই paradigm and for the তুমি speeches of the same speakers. Not a chosen subset: the whole
span. This is what the first design lacked and what its finding 29 demanded.
TUI sites — five, the whole of them in section ২:
| site | speaker → addressee | source cue |
|---|---|---|
[176] |
headman → Amal | grammar + lexis — কে রে and কোথাকার বাঁদর এটা (what monkey is this) |
[188] |
headman → Amal | grammar + lexis — তোদের and ওরে ছোড়া (here, brat) |
[190] |
headman → Amal | grammar only — কেনরে, তোর খবর; no pejorative noun |
[238] |
boys → each other | grammar only — চল্ ভাই চল্, inside a speech addressed to Amal as তুমি |
[242] |
boys → each other | grammar only — দেখছিস্ ভাই, likewise |
TUMI control sites — four, speeches by the same two speakers at plain তুমি, including two
that are sarcastic without any grade drop ([186], the headman's mock congratulation), so the
control is not merely "polite text": [184], [186] (headman → Amal), [236], [248] (boys →
Amal).
Nine speeches. build_items.py asserts every one verbatim against source-ipublishinghouse.txt and
every lead rendering verbatim against the frozen translation before any dispatch.
3. Subjects
Three English renderings of each of the nine speeches:
LEAD— the frozenT-dakghar-R05-v1, produced under a register that forbids grammatical compensation, by a translator who then sealedD33. Hypothesis-aware, and that is the point: this subject is the one whose self-report is on trial.H1=moonshotai/kimi-k3,H2=deepseek/deepseek-v4-pro— naive hands, translating the same speeches blind: no mention of pronouns, grade, contempt, footing, register or of an experiment, and no sight of any other rendering. They establish whether lexical compensation is a property of this translator or of the crossing.
4. Measure, and who codes it
For each (speech × rendering), a coder marks every lexical marker of contempt or of the addressee's inferior standing in the English — pejorative or diminutive vocative (boy, brat, lad, you there), insult noun (monkey, rascal), or dismissive status noun — and marks, for each, whether a corresponding word stands in the Bengali speech.
added= 1 for a rendering that contains at least one such marker with no counterpart in the source speech; else 0.added_rate(cell)= mean ofaddedover the renderings in that cell.false_add=added_rateon the fourTUMIcontrol speeches. A translator who sprinkles boy everywhere is not compensating for grade; the control is what tells the two apart, and it carries sarcastic source speeches precisely because those are where a spurious insult is most tempting.
Coding is done three ways and the lead's coding is not privileged. Two non-Anthropic seats
(google/gemini-3.6-flash, x-ai/grok-4.5) code every rendering blind — blind to which
rendering is the lead's, blind to D33, blind to the TUI/TUMI classification, and shown the
Bengali only as a word-list to check counterparts against. The lead codes independently. The
reported figure is the majority of three, and disagreements are listed.
5. Gates — withholding
G1(coding reliability). Three-way agreement onaddedmust be ≥ 0.80 across the 27 codings. Below that the primary is not read and the page reports the disagreement instead.G2(source-side coding). TheTUI/TUMIand grammar-only/grammar+lexis classification of the nine source speeches is the lead's and is single-coded; it is a morphological fact (তুইparadigm vsতুমিparadigm) checkable from the text, and the tokens are listed in §2 so that any reader can check it. This is a declared deviation, not a control.
6. Predictions, registered before dispatch
P1(primary, and the test ofD33). The lead'sadded_rateon the three grammar-onlyTUIsites is greater than 0 — i.e. the sealed self-report is wrong.D33predicts 0.P2. The lead'sadded_rateon grammar-only sites exceeds itsfalse_addon the fourTUMIcontrols — the addition tracks the source's grade drop rather than being a tic.P3. At least one naive hand also showsadded > 0on at least one grammar-only site — the compensation is not peculiar to a translator working underV5.P4. On the two grammar + lexis sites every rendering carries a marker with a source counterpart (added= 0 there, because the cue is not added — it is translated).
Failure criteria. P1 fails if the lead's grammar-only added_rate is 0, and then D33 stands
and the interesting result is that the register held. P2 fails if false_add is at or above the
grammar-only rate, and then the primary is withheld: the additions are a stylistic habit, not
compensation. Any G1 failure withholds everything.
7. What this cannot show, declared in advance
- Nine speeches, five of them
TUI, three of those grammar-only. That is the whole of section ২ — the census is complete and it is still small. No interval is computed and none is reported; the result is a count, stated as a count. The first design's fault was to dress a count of three in inferential clothes. - "Lexical marker of contempt" is a judgement, not a measurement. That is why it is coded three ways behind a reliability gate, and why disagreements are printed rather than resolved.
- This says nothing about readers. It is a fact about what translators write, not about what anyone perceives. The reception question is what killed design v1 and it is not asked here.
- One work, one language, one direction, three renderings. Bengali
তুইis not every language's intimate pronoun and the headman is not every rude speaker. - The lead is contaminated by construction — it wrote the register, the translation and
D33. It cannot be a naive subject and is not offered as one; it is the subject on trial. The naive hands are what carryP3, andP3is the only sentence here that generalises beyond this translator. - Tier D has not passed. Every evaluative sentence downstream is
provisionalandinternal-judgment-only.
8. Cost
Pre-flight ceiling, within the $0.50 declared for this experiment and now itemised again: six
translation calls (cap 1,500 out; three speeches per call, two hands), two blind coder calls (cap
3,000 out), one pre-run critic call on this design (cap 8,000 out). Worst case from max_tokens,
note (abc): $0.068 + $0.006 (hands) + $0.023 + $0.018 (coders) + $0.055 (critic) + prompts ≈
$0.19, on top of the $0.04057375 already spent criticising design v1. A body returning
finish_reason == "length" is re-dispatched once at the same cap whether or not it has content,
and is recorded dead and excluded if it truncates again.