Repository path: workshop/experiments/E-20260826c-register-cost/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260826c-register-cost |
| status | frozen |
| created | 2026-08-26 |
| updated | 2026-08-26 |
| senses | style-correspondence, consistency, readability |
| links | wiki/arms/ARM-alf-layla.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, workshop/regimes/R05-serial-long-work.md, config/models.md, config/budget.md |
When a long work's binding register overrules the local choice, does a reader see anything wrong?
ARM-alf-layla step 10, study limb of span J. Frozen 2026-08-26 before span J was translated and
before any scored call was dispatched. The materials are spans A–I, all frozen and landed on
main before this session opened; span J is not in the item set, so nothing this session
translates can shape the instrument.
1. The question, and where it comes from
R05 — the regime this arm translates under — has a known limitation written into it at birth,
2026-07-27, and never measured:
The register can become an alibi. A decision written into the register looks settled, and a later span that ought to reopen it has an instrument telling it not to. The
unresolvedsection is a partial answer; it is not a complete one.
Nine spans later the arm can count the occasions. Span H's log (D102–D118) is the first place
the standing cost is stated as a quantity:
Three of the four losses are the price of consistency, paid at a locus that was decided before the locus existed — which is the standing cost of a binding register, and the first span in which it can be counted.
Every translator of a long work keeps some form of glossary and freezes terms in it. The craft question this experiment asks is the one the glossary cannot answer about itself: when the frozen term is carried into a passage where a translator working locally would have chosen differently, does a reader who does not know the source notice?
If the answer is no, consistency across a long work is cheap, and a translator should freeze early and hold. If the answer is yes, every entry in the register is a debt that later spans pay in front of the reader, and the freeze should be later and looser.
What this unit teaches about translating literature (subject rule, wiki/tracks.md). It is
about a decision every translator of a long work makes — how hard to bind terminology across
hundreds of pages — and it measures the price of binding it. The instrument is a means. Nothing
here is about this project's statistics, raters or published figures.
2. Design in one line
Twelve passages from the frozen translation, each in two versions differing only in the words the register decided; each version read by three panel seats, twice, one version per call, with the seats asked to quote anything they would query as an error. No comparison, no ordering.
3. Why single-item and not A vs B
Note (brs), written yesterday (2026-08-26, S224): on minimal pairs differing in a few words, an A-vs-B which reads better task returned the same arm on an order swap at 54.8% against a chance 50%, and all three seats preferred whichever passage came first. That instrument is not usable here. This design never puts two versions in one prompt. Each call carries one passage; the arm is between-calls; there is no order to be biased by. The outcome is not a preference but a quotation, scored by string match.
4. Materials
Source of every passage: workshop/translations/alf-layla/R05-v1/translation.md, spans A–I,
frozen span by span between 2026-08-14 and 2026-08-25 and landed on main. Windows are 93–165
words, snapped to sentence boundaries, with all page metadata stripped.
The eleven loci. Each is a place where the binding register
(workshop/translations/alf-layla/register.md) fixed an English word and the frozen log names the
alternative it refused. BOUND is the text as printed. FREE substitutes the named alternative
at every occurrence in the window and changes nothing else.
| # | span | register | BOUND |
FREE (the alternative the log names as refused) |
occurrences |
|---|---|---|---|---|---|
| L1 | H | T2 / D104 |
slave-girl | young woman | 1 / 1 |
| L2 | A | T4 / D11 |
did not leave off | stayed | 1 / 1 |
| L3 | G | T29 / D86 |
mallet | polo-stick | 2 / 2 |
| L4 | G | T31 / D85 |
Ruyan | Douban | 1 / 1 |
| L5 | D | T17 / D47 |
the court | the divan | 2 / 2 |
| L6 | E | T16 / D46 |
old man | sheikh | 1 / 1 |
| L7 | H | T33 / D103 |
ghoulah | ghoul | 2 / 2 |
| L8 | F | T26 / D76 |
three needs | three wishes | 1 / 1 |
| L9 | F | T27 / D77 |
doomed thing | outcast | 1 / 1 |
| L10 | F | V17 / D80 |
Umm Amir (unglossed) | the hyena (glossed) | 1 / 1 |
| L11 | D | T13 / D20 |
the utmost marvelling | marvelled greatly | 1 / 1 |
L1 is the register's own predicted breaking point. It is the woman at the head of the road in
span H who, in her next breath, says she is the daughter of a king among the kings of India —
D104, the decision that made register.md's unresolved question 12 and wrote the cost is real
on the page.
No alternative was invented for this experiment. Every FREE reading is quoted from the frozen
log's own record of what was considered and refused: D76 refuses three wishes by name, D86
refuses polo-stick by name, D85 refuses Douban by name, D46 refuses sheikh, D103 refuses
ghoul, D47 refuses divan, D77 refuses you outcast, D104 refuses young woman, D80
refuses the gloss, D11 fixes did not leave off to keep a formula countable, D20 fixes the
cognate figure.
Occurrence counts are equal in both arms of every locus (column 5), checked mechanically. No
locus's FREE target string occurs anywhere in its BOUND text, and no BOUND target survives in
its FREE text, under word-boundary matching — the check that removed the marid/ifrit locus from
the set (§12, finding 3).
The six seeded controls. Four blatant and — this is the change the critic forced — two subtle,
matched in size to the BOUND/FREE difference: one content word, grammatical, contextually
wrong.
| # | span | grade | kind | defect |
|---|---|---|---|---|
| SC1 | F | blatant | anachronism | a knife → a screwdriver |
| SC2 | E | blatant | register break | wept a bitter weeping → completely lost the plot |
| SC3 | H | blatant | malapropism | certain of destruction → certified of destruction |
| SC4 | I | blatant | internal contradiction | bring us four fish … and there were four fish → … there were three fish, both sentences inside the window |
| SC5 | G | subtle | register intrusion | my boon-companion and my beloved → my boon-companion and my associate |
| SC6 | C | subtle | register intrusion | the lives of the kings who went before → the biographies of the kings who went before |
Why both grades. The blatant arm shows the seats can flag anything at all. Only the subtle arm
licenses a null: if BOUND and FREE come out equal and the subtle seeds are caught, the
register's impositions are genuinely invisible; if the subtle seeds are missed too, the instrument
is blind at this grain and the primary says nothing (§12, finding 9).
Item texts are frozen in items.json in this directory,
sha256 0cc7abd2f680c192ceed9a311f86753e825d134b1d85f47e8894ab1799c5a962, and are not edited
after this page is committed. The dispatcher and the verifier both read that file and both check the
hash.
5. Procedure
- Conditions: 11 loci × {
BOUND,FREE} + 6 seeds × {CLEAN,SEEDED} = 34. - Seats:
P1openai/gpt-5.6-terra,P2google/gemini-3.6-flash,QRqwen/qwen3.7-max.P3is a cost problem,P4is out (note (bps)),P5is out (note (bne)),GLis out on long prompts — the three named are the panel this session can reach. - Replicates: 2 per (condition × seat). 204 calls.
- One passage per call. No seat ever sees two versions of one passage in one prompt; the two versions of a locus are separate calls and, being separate calls, carry no order.
- Dispatch order is randomised once, per seat, from a fixed seed (
20260826), and the order is written to the run directory before dispatch.BOUND,FREE,CLEANandSEEDEDcalls are interleaved so that provider drift cannot align with condition (§12, finding 11). temperature1.0 (default),max_tokens2500 forP1andP2, 3000 forQR(note (brr):QRneeds 3000 from the start; at 3000 it ran 56 calls without a failure).- Concurrency 6 in one foreground process (note (brf)). Model id and response timestamp are recorded on every call.
The prompt, identical for every call except the passage:
Below is a passage from a modern English translation of a medieval Arabic story
collection. The translation is deliberately plain and unarchaic, and it keeps the
original's short coordinated sentences.
Read it as a copy-editor reading English. Quote EXACTLY any word or phrase in the
passage that you would query — a wrong word, a mistake, a word that does not fit the
sentence it is in, or a word wrong for the surrounding register. Do not query the
plainness or the coordination themselves: those are the translator's declared style.
If you would query nothing, return an empty list.
Reply with JSON and nothing else:
{"queries": ["<exact words quoted from the passage>", ...]}
PASSAGE:
<passage>
The seat is told nothing about the register, the experiment, the existence of a second version, or which words are under test.
6. Scoring
Mechanical, by the verifier, from the raw JSON. The verifier recomputes every number on this page
from runs/*.json and refuses to report any it cannot reproduce.
- Hit = any returned query string contains the cell's target term as a whole word or phrase,
matched by the regex
\b<escaped target>\b, case-insensitively, after collapsing whitespace. Word-boundary matching is what keepsghoulfrom matchingghoulah(§12, finding 3). - Targets, one per cell, all four arms scored at the same textual location:
BOUND→ the bound term ·FREE→ the substituted term ·SEEDED→ the seeded defect ·CLEAN→ the clean counterpart at the seed's own location (e.g. knife in SC1CLEAN, where SC1SEEDEDis scored on screwdriver). This is the critic's finding 4: scoring aCLEANcall on a string that cannot appear in it measures nothing. - Query rate for a cell = hits / live calls in that cell.
- All returned query strings are preserved, hit or not. Non-target queries are counted and reported per arm, not discarded.
- An unparsable body is re-dispatched once; a second failure is recorded
deadand the cell's denominator is reduced, with the reduction stated on the result page.
7. Predictions, registered before dispatch
Alpha is 0.05, one-sided, fixed here.
- R1a (blatant instrument). Pooled query rate on the four blatant
SEEDEDpassages ≥ 0.75, and no single blatant seed below 0.40. (The per-seed floor is the critic's finding 5: a pooled gate can pass while one control detects nothing.) - R1b (subtle instrument). Pooled query rate on the two subtle
SEEDEDpassages ≥ 0.40. R1b is what licenses a null on the primary. - R2 (primary, directional). Pooled over the eleven loci,
BOUNDquery rate >FREEquery rate. Reported three ways, all pre-specified: (i) the eleven paired cell rates in a table; (ii) the pooled difference with a 95% interval; (iii) an exact one-sided binomial on the sign of (BOUND−FREE) per locus, ties dropped from both numerator and denominator. Disposition, fixed now: supported if the sign test reaches P ≤ 0.05; directional but not established ifBOUND>FREEpooled and P > 0.05; null if the pooled rates differ by < 0.05 in either direction; refuted ifFREE>BOUNDpooled. - R3. The single highest
BOUNDquery rate falls on L1 — the one locus where the register's own unresolved section predicted trouble. - R4.
CLEANquery rate at the seed locations ≤ 0.15: the spurious base rate. - R5. At least one seat returns an empty query list on ≥ 40% of its calls.
8. Failure criteria, written before the run
- R1a fails → the instrument is not shown to work at all; the primary is withheld and the page reports an instrument null.
- R1a passes and R1b fails → the primary is reported, but a null result on R2 is reported as uninformative, not as evidence that the register costs nothing: the instrument would be shown blind at the grain the loci vary on.
- A seat returns a non-empty query list on 100% of its calls → all seats stay in the primary; a pre-specified sensitivity analysis excluding that seat is reported beside it (§12, finding 12).
- More than 10% of calls die → affected cells are reported with reduced denominators and the primary is marked underpowered.
FREE>BOUND→ R2 is refuted and reported refuted. That is a finding, not a defect: it would say the register's departures from conventional English read better than the conventional word, which is a result about foreignising terminology.
9. What this cannot show
- No human reader is measured. These are language models reading English. What they query is not what a reader notices; it is what a model asked to copy-edit returns. Every sentence of the result carries that.
FREEis the logged alternative, not a certified more natural English. The critic's finding 8 is accepted in full and not repaired: nothing here establishes that the divan, sheikh, ghoul or the hyena is the reading a translator working locally would have taken; they are what this translator wrote down as refused. Several of the substitutions also change reference or cultural placement, not only register. The result is therefore about the register's word against the alternative its own log named, and may not be read as bound versus natural.- The eleven loci are a purposive set, not a sample. They were chosen by the lead from a log the lead wrote, under a stated rule (the log must name a refused alternative). The exact binomial is reported as a within-set descriptive statistic and explicitly not as population inference about binding registers in general (§12, finding 10).
- Two pairs of items are not independent. L3 and L4 come from the same polo passage in span G; SC3 sits inside the same span-H episode as L1 and L7. They are separate calls scored on different targets, but they are not independent samples of the translation.
- Eleven loci is eleven. The exact binomial over 11 signs cannot reach P ≤ 0.05 with fewer than 10 concordant signs.
- The lead never judges its own translation (charter §5). Nothing on the result page is a quality claim about the rendering; the outcome is a count of what seats quoted.
10. Pre-flight cost
Worst case is built from max_tokens, not from expected output (note (abc)).
| seat | calls | in (est.) | out at cap | worst case |
|---|---|---|---|---|
P1 $1.00/$6.00 |
68 | 68 × ~520 = 35k → $0.035 | 68 × 2500 = 170k → $1.020 | $1.055 |
P2 $0.75/$3.75 |
68 | 35k → $0.027 | 170k → $0.638 | $0.665 |
QR $1.475/$4.425 |
68 | 35k → $0.052 | 68 × 3000 = 204k → $0.903 | $0.955 |
| pre-run critic | 2 | — | — | $0.118089 (actual, spent) |
Declared ceiling: $2.80. UTC-day headroom after the critic: $3.824955 of $5.00. If the actual approaches the ceiling the replicate pass is dropped, not the seeded arm.
11. Pre-run critic
Dispatched on the v1 page and item set, P1 and P2, max_tokens 9000, temperature 0, one
round, $0.118089. Raw in critic-v1.json. Verdict: NEEDS REDESIGN — P1 returned five
BLOCKING findings and P2 two, thirteen and five findings in all. Nothing was dispatched under v1.
12. The critic's findings, and what each changed
All five BLOCKING findings are accepted in full. Of the eight SERIOUS and MINOR findings, six are accepted and two are accepted in part, with the refusals on the record.
- BLOCKING (
P1,P2finding 4) — the frozen page and the frozen items disagreed about L2'sFREEreading. The page said stayed,items.jsonsaid went on. Accepted, and the cause was worse than the symptom: the v1items.jsonwas written by a builder run that died before its final write, so the file on disk was an earlier draft than the page describing it. Fixed by rebuilding the whole item set in one pass, and by committing the sha256 ofitems.jsoninto the design, which the dispatcher and the verifier both check.FREEis stayed. - BLOCKING (
P1) — L1 did not test the case the register predicted. v1's L1 was the cook slave-girl of span I, a woman the text says was a present from the King of Rum, so theFREEreading young woman contradicted her stated condition; and the substitution missed the vocative Slave-girl, leavingFREEinternally inconsistent (P2finding 3, same defect). Accepted. L1 is now the woman at the head of the road in span H —D104, the actual caseregister.mdquestion 12 was written about — and every occurrence in the window is substituted. - BLOCKING (
P1) — the hit rule was contaminated by target strings outside the manipulated position. In v1's L2FREE, went on already stood in an unchanged sentence; in v1's L11FREE, ifrit stood unchanged a dozen times around the one substituted address. Accepted. Three changes: L2'sFREEreading is now stayed, which occurs nowhere else in its window; the marid/ifrit locus is dropped from the set entirely, because every regularising substitution makes the target identical to surrounding text and no window fixes that; and every remaining locus is checked mechanically for equal occurrence counts and non-occurrence of the other arm's target, with word-boundary matching (§4 column 5, §6). - BLOCKING (
P1) — R4 was vacuous. ACLEANcall was to be scored on the seeded string, which by construction cannot appear in it. Accepted.CLEANcalls are now scored on the clean counterpart at the seed's own location (§6). - BLOCKING (
P1,P2finding 2) — SC4's window did not contain the contradiction. The four fish were counted three thousand characters earlier, outside the passage the seat would see. Accepted. SC4 is rebuilt on a window that carries bring us four fish and there were four fish in consecutive sentences, so the seeded three contradicts text the seat can see. The gate also gains a per-seed floor of 0.40, so it can no longer pass while one control detects nothing. - BLOCKING (
P2finding 1) / SERIOUS (P1) — L12 carried page metadata inside the passage.## Span F … Stored source: …was inside both arms of the cognate-figure item. Accepted; the window builder now cuts at any##and strips markdown, and every item was re-checked. - SERIOUS (
P1) — bundled and unequal targets. L3 changed both mallet and the field; L11 had twoBOUNDtargets and oneFREE. Accepted. L3 now substitutes mallet alone and leaves the field standing in both arms; the two-target locus is gone with the dropped marid item; every locus has one target per arm and equal counts. - SERIOUS (
P1) —FREEis not established to be the more natural local English. The divan, sheikh, the hyena and the dropped ifrit change reference, title or cultural placement, not merely register. Accepted as a limit, refused as a redesign. The remedy the critic proposes — independent pre-screening by raters or a source-informed translator — is a second experiment, and the only screen this session could buy is an A-vs-B preference call, which note (brs) has just shown to be order-driven on minimal pairs. The claim is narrowed instead: §9 now states that the result is about the register's word against the alternative its own log named, and may not be read as bound versus natural. This is a refusal and is on the record as one. - SERIOUS (
P1) — passing a blatant-seed gate does not license a null about subtle terms. Accepted, and this is the largest change to the design: two subtle seeded controls were added (SC5, SC6), matched to theBOUND/FREEdifference in size and grammaticality, with their own gate R1b, and §8 criterion 2 makes an R2 null uninterpretable unless R1b passes. - SERIOUS (
P1) — the binomial treats purposively selected, partly overlapping loci as independent. Accepted. §9 now labels the exact binomial a within-set descriptive statistic and disclaims population inference; the overlapping pairs are named. - SERIOUS (
P1) — the decision rule was incomplete. No alpha, no disposition for directional but not significant, no tie rule, no statement of the analysis with and without an excluded seat. Accepted; §7 R2 now fixes alpha, the tie rule, and four named dispositions, and §8 criterion 3 keeps every seat in the primary with a sensitivity analysis beside it. - MINOR (
P1) — fixed dispatch order lets provider drift align with condition. Accepted; §5 randomises the order once per seat from a fixed seed and writes it out before dispatch. - MINOR (
P1) — post-hoc seat exclusion. Accepted; see finding 11. P2finding 5 — the scoring rule did not say whether a multi-target cell is any or all. Accepted and dissolved: there are no multi-target cells left.
What the critic did not catch, and the lead adds: the six seeded passages and the eleven loci
are drawn from the same nine spans, so a seat that has decided this translation is odd will query
more everywhere. The CLEAN arm (R4) is the only measure of that, and it is measured at the seed
locations only, not across the whole passage. The non-target query count per arm (§6) is
reported for exactly this reason.