Repository path: workshop/experiments/E-20260805c-r1-polish/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260805c-r1-polish |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| senses | style-correspondence, accuracy, voice, cultural-mediation, naturalness |
| provisional | true |
| links | wiki/arms/ARM-r1-fresh-pair.md, framework/v0.1/README.md, workshop/translations/kamizelka/R04-v1/translation.md, workshop/translations/kamizelka/R06-v1/translation.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/findings/results/RS-20260804-yardstick.md, wiki/findings/results/RS-20260804g-yardstick-holds.md, workshop/experiments/E-20260804g-yardstick-repair/design.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260805c — R1 in a language it has never been tried on, and a gate on what counts as a site
ARM-r1-fresh-pair step 1 (T5). Frozen before any seat is addressed. The translation limb —
T-kamizelka-R04-v1, the whole of Prus's «Kamizelka», with its log — was frozen first, at 59a5879,
and the FORCED and DECOY spans below were written and committed before the yardstick seat was
dispatched. Commit order is the guarantee and the verifier checks it.
1. What is being asked
framework/v0.1 §3 registers, as the first of four predictions and the only one that tests the
release's only recommendation on new material:
On fresh Class A sites in a new pair, R1 recovers a marking at more than half. Score by the
E-20260802eprocedure. The failure it should not survive: recovery at or below a third.
R1 itself:
Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking. Render the site again under a brief that requires the marking to appear, and let it fall wherever the target does mark such things. Record the loss only if that second attempt fails.
The prediction has been open since the release and two runs have failed to close it.
E-20260803c (DE→EN) hid the source and is not the named procedure. E-20260804 (FR→EN) was the
named procedure and fired three of its own failure criteria, so it is descriptive only. S106 then
removed the explanation S101 offered — RS-20260804g varied the yardstick's authorship alone and
R1's separation reproduced at +4 — so the FR→EN null is currently unexplained.
This run's hypothesis about that null is stated now, before any data, because it is what the
design is built around. E-20260804 fired F2 (FROZEN recovered at 4 of 7) and F5 (an
independent seat's plain translation recovered at 4 of 7). At most of its putative Class A sites
the relation was carried by any competent English at all. If there was no loss, there was
nothing for R1 to repair, and a run on such sites cannot test R1 whatever it returns.
2. The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5)
What does this unit teach about translating literature or evaluating translations? — Whether the
one piece of advice this project gives a translator transfers to a language it has never been tried
on: when Polish marks a social relation by a grammatical form English does not have — the intimate
ty said downward to a tradesman, the diminutive that makes a grown man small, the impersonal past
with no one in it — does re-rendering the site so the marking falls on a device English does have
put the relation into the English, or does the loss stand?
This is a principal unit and not method work: what is measured is a property of translations. The one apparatus question it settles — what makes a site a site at which the question can be asked — is a gate inside the unit (§4), timeboxed to it, and its result is reported on the result page, not carried into the backlog.
3. Materials — the census, and how it was built
The whole story was translated first, close, under R04, with the log frozen before this design was
written. The census was then read off the log: every place where the log records a Polish
grammatical form that English has no counterpart for and whose content is social, attitudinal or
perspectival. materials/sites.json holds ten candidate sites, S1–S10.
Excluded by design, and named so the exclusion is not invisible:
- Aspect and iterativity (log D7) — the relation is about the shape of an event, not about
people. These are the categories most of
E-20260804's FR→EN sites were, and R1's scope clause does not claim them. - Grammatical gender and virile agreement (log D12) —
framework/v0.1§3 prediction 2 is established as not falsifiable by this procedure: a metalinguistic relation can only be conveyed by metalanguage, which is what the POSITIVE arm is. - The dealer's orthographically accented speech (log D1) — spelling is not a grammatical category, and the log records that the filed translation already displaced that marking onto syntax. A site where the translator has already made R1's move is not a test of it.
Contamination. Declared on the translation artifact and not measured: three English
renderings exist in print (Jopson 1930–31, Scherer 1955, Sues 1960) and none is freely reachable, so
tools/dependence_check.py could not be run. Under CLAUDE.md's standing rule that is said rather
than passed over. The design's validity does not turn on it: the claim under test is
existential — does an English rendering exist that carries the relation — and a rendering that
reproduced a published translator's device would still be an existence proof. Note (bhb) cuts
toward the null and is the design's friend: the lead matches itself at up to 37 contiguous tokens
across sessions, so a FORCED span that merely reproduces its FROZEN span marks nothing.
4. Stage A — the yardstick, and then the admission gate
A1/A2 — what the source conveys, written by seats shown no English
A1 = P5 deepseek/deepseek-v4-pro, used in no other role. Per site it receives the Polish
passage, a staging note of bare narrative fact (who is present, what is happening), and a
pointer that quotes the Polish tokens under study and says nothing about them. It receives no
English rendering of any kind — not FROZEN, not FORCED, not the gloss, not the log. It returns one
sentence stating what the passage conveys over and above what its words denote.
A2 = qwen/qwen3.7-max, the panel's documented first reserve, gets the byte-identical prompt.
A2's statements are not graded. They are (i) a stability measurement — is "what this passage
conveys" a stable object across two independent readers — coded SAME/OVERLAPPING/DIFFERENT
and committed before any grading seat is addressed; and (ii) the declared fallback if A1
fails both attempts.
The leak screen (analysis/screen.py, committed with this design, lexicon inherited verbatim
from E-20260804g §4): a statement naming the source's grammatical device tells the graders the
answer. A statement that fires is re-requested once with the rule quoted; a statement that fires
twice puts its site outside every denominator, named.
The admission gate — new here, and it is the design's whole point
N1 = P4 moonshotai/kimi-k3, used in no other role, is shown the Polish source and the staging
note and no English, and writes a plain English rendering of each span: ordinary English, no
brief, no knowledge that any study exists. This is the NEUTRAL arm of E-20260802e, produced here
first.
Then the three grading seats are shown, per site, A1's relation statement and N1's plain rendering alone, one item per site, and answer YES/NO: would a reader with no access to the source come away with the relation? A site at which ≥ 2 of 3 seats say YES is EXCLUDED before the main run.
Why this is a gate and not a result. R1 is advice about sites where a marking is actually lost.
E-20260804 discovered, after paying for a full run, that most of its sites were not such sites. The
gate moves that discovery in front of the spend. The count it removes is registered as a prediction
(§7 G1) and is reported whatever it is — if it removes most of the census, that is the run's
finding and the FR→EN null gets its explanation.
What the gate costs, stated rather than hidden. (i) The NEUTRAL arm is the selector, so it is
not carried into the main run and no FORCED-versus-NEUTRAL comparison is available in this
design. E-20260804's F5 question is therefore not re-asked here; the gate is its measurement.
(ii) DECOY is a lead-written unmarked paraphrase and is not the selector, but on admitted sites it
is expected low partly by construction, so F1's DECOY leg is weakened and WRONG carries the
discrimination control. (iii) FROZEN is not the selector — it is the lead's close rendering,
written months of sessions of habit away from N1's — so F2 remains a genuine unselected control,
and it is the one that decides whether the admitted sites were really sites.
5. The arms, and who wrote each one knowing what
| arm | who wrote it | knowing | committed |
|---|---|---|---|
| FROZEN | lead | nothing; written as translation, to no test | 59a5879, before this design |
| FORCED | lead, under R1's brief | the source and its own log — not A1's statement | with this design, before stage A |
| DECOY | lead | a length-matched paraphrase required to mark nothing | with this design, before stage A |
| REAIM | lead, under R1's brief | the source and A1's statement | after stage A, declared |
| POSITIVE | lead | an explicit gloss authored from A1's statement | after stage A |
| WRONG | lead | an explicit gloss of a different, plausible relation | after stage A |
FORCED is written before A1 and that is deliberate. R1 is advice to a translator, and a translator has no oracle telling them what an independent reader will say the passage conveys. A FORCED arm aimed at A1's statement would test whether the lead can hit a target it has been shown. REAIM is that upper bound, kept separate and never reported as R1's score: it answers only is the relation A1 names reachable in English at all by ordinary narrative means. A REAIM span that states the relation instead of carrying it is a failed REAIM; that is what POSITIVE is for.
6. Stage B — grading
Form inherited from E-20260802e §4 and E-20260804g §6.
- Three non-Anthropic seats — P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5. - Each receives, per admitted site: the source, the gloss, A1's relation statement, and the five renderings unlabelled (FROZEN, FORCED, DECOY, REAIM, POSITIVE, WRONG — six), and answers for each: would an English-speaking reader with no access to the source come away with the RELATION — YES or NO, plus one line of reason.
- Presentation order per site is fixed by
sha256(site id | seat id | ordering), never chosen by the lead. Two orderings per seat, the second the reverse of the first — six bodies. - A seat counts as recovering the relation only if it answers YES in both of its orderings.
- A site counts as recovered under a condition when ≥ 2 of 3 seats recover it.
7. Predictions — registered, before any call
| # | prediction | comparator |
|---|---|---|
| G1 (the gate) | the admission gate excludes at least a third of the ten candidate sites | E-20260804: 4 of 7 by F5 |
P1p (primary; this is framework/v0.1 prediction 1) |
FORCED recovered at more than half of admitted sites | E-20260802e 5 of 6 |
| P2p | FROZEN recovered at ≤ 1 admitted site | 0 of 6 |
| P3p | (FORCED sites) − (FROZEN sites) ≥ half of n | +5 |
| P4p | DECOY recovered at ≤ 1 site | 0 of 6 |
| P5p | POSITIVE recovered at n or n−1 sites | 6 of 6 |
| P6p | WRONG recovered at 0 sites | 0 of 7 (FR→EN) |
| P7p | REAIM recovered at ≥ FORCED | new arm |
The lead's stated expectation, recorded so it can be wrong. I expect G1 to hold and to be the
session's finding: I expect the gate to remove four or five of the ten, and the survivors to be
concentrated in S1, S2, S3, S6 and S9 — the intimate ty downward, the two attitudinal diminutives,
the agentless past and urzędniczek. I expect S4 and S8 to be excluded, because the log already
notes that a little rouble and the lady's husband carry their constructions into English without
help. On the survivors I expect P1p to hold at 3 or 4 of 5. If P1p fails on admitted sites,
that is the first evidence against R1 from a site population where R1's precondition demonstrably
holds, and it is the expensive result.
8. Failure criteria — pre-committed
- F1 — the grading instrument failed. POSITIVE recovered at fewer than n−1 sites, or WRONG
recovered at more than 1 site. → the run is descriptive only; nothing enters, amends or is
withdrawn from
framework/v0.1§2 or §3 on it, whatever P1p says. (DECOY's leg is deliberately not part of F1 — see §4, it is weakened by the selection.) - F2 — there was no loss to repair. FROZEN recovered at ≥ half of admitted sites → the admission gate did not do its job, prediction 1 is uninterpretable on these sites, and the run reports that and the gate count and nothing else about R1.
- F3 — the yardstick leaked. A statement firing the §4 screen twice → its site is outside every denominator, named.
- F4 — the yardstick failed to arrive. A1 fails both attempts → A2 governs and the run says so; both fail → no grading seat is addressed.
- F5 — the gate emptied the census. Fewer than 4 sites admitted → the main run is underpowered, P1p is withheld, and the gate count is reported as the run's only finding.
- F6 — seats. Fewer than three grading seats returning both orderings → primary is descriptive
only. A
finish_reason: lengthbody is a seat failure and never a partial answer (note (b)).
A null is a result, and here is what each null means. If P1p fails on a properly gated
population, framework/v0.1 §3 prediction 1 is falsified in a new pair and §2's traceability
paragraph is qualified in place: R1's move did not recover the marking where the marking was
demonstrably lost. R1's text would not thereby be refuted — it would be left standing with a
declared pair where it does not work, which is the honest reading and is what §7's pair table is for.
9. What this design cannot show
- Nothing about quality. Tier D is NOT PASSED; evidence class X3 is inadmissible. This measures
availability — whether a relation is recoverable from a piece of English — never whether the
marked rendering is better.
RS-20260802e§4.1 priced the recovered marking at about two-thirds of a naturalness point and nothing here re-prices it. - Nothing about human readers. Three language models are not a survey (charter §4). A YES licenses three independent readers recovered the relation from this English. Tom is never an experimental subject (charter §9).
- Nothing from a sample. Ten sites are a census of the qualifying population one story yields, not a draw. No sampling inference is available and none is drawn.
- Nothing about FORCED against an ordinary translation. The gate consumed that comparison (§4).
- Nothing separating the pair from the story. One author, one translator, one text. A PL→EN verdict here is a verdict about these ten sites in this story.
- One leak this design does not close. The staging notes are lead-written, and A1 reads the passage through them. A staging note that hinted at the relation would seed the statement; they were written to be bare narrative fact and the screen does not check them.
10. Verification
analysis/verify.py, importing nothing from tools/: asserts every FROZEN span byte-identical to
workshop/translations/kamizelka/R04-v1/translation.md and every source span byte-identical to
workshop/translations/kamizelka/source-pl.txt (whitespace-normalised); asserts by git commit
order that materials/sites.json carrying FORCED and DECOY was committed before the first stage-A
body; recomputes the hash-derived rotations; recomputes the gate, all eight predictions and all six
failure criteria from the raw stored bodies; recomputes the leak screen; re-sums billed cost from the
stored bodies against the key-usage delta. At least three mutation tests, each asserting the
bytes on disk actually changed (note (bgu)) and restoring them after (note (bhd)).
11. Budget
Declared worst case $1.35, built from max_tokens and never from expected output (note (abc)),
with the S022 routing margin.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| stage 0 — seat probe at the run ceiling (note (bhf)/(bit)) | 5 | 12,000 | $0.01 |
pre-run adversarial critic (z-ai/glm-5.2) |
1 (+1 reserve) | 24,000 | $0.30 |
| stage A — A1, A2, N1, 1 retry reserve each | 3 (+3) | 12,000 | $0.30 |
| the admission gate — 3 seats × 1 call | 3 | 12,000 | $0.25 |
| stage B — 3 seats × 2 orderings | 6 | 12,000 | $0.49 |
Today's UTC headroom before this run: $3.926537656 of the $5.00 cap, two sessions already run. Lead translation is free and is never ledgered (charter §3, A4): the whole of «Kamizelka» twice, FORCED, DECOY, REAIM, POSITIVE, WRONG, the census, the coding and all verification are lead work at no API cost.
12. Amendments A1–A7, from the pre-run critic (critic.md)
Seat: z-ai/glm-5.2, max_tokens 24,000, used in no other role, one body, 43.8 s,
$0.02569214. Nine findings, two BLOCKING, VERDICT: NEEDS-REDESIGN. All nine are accepted,
six in full and one in part. Applied before any stage-A body was dispatched; the stage-0 seat
probe (5 of 5 alive, $0.002) is the only API call that preceded this section.
A1 — the gate selects on FROZEN, not on N1 (findings 1 and 3, BLOCKING, accepted)
The design gated on the wrong loss and the critic is right. R1's own text says "Record the loss only if that second attempt fails" — the first attempt is the translator's close rendering. So the loss that makes a site a Class A site is FROZEN's loss, not an independent seat's. Gating on N1 also excluded exactly the sites where R1 is most testable (finding 3): a site where FROZEN fails and N1 succeeds is a real loss for the translator and was going to be thrown away.
Amended: a site is admitted iff FROZEN is NOT recovered (fewer than 2 of 3 seats). The gate is applied as pre-registered stratification computed from the main run's own bodies, fixed here before any body exists — which is statistically identical to a separate gate stage, costs one fewer stage, and stops the grading seats seeing the FROZEN span twice.
What this costs, stated: FROZEN is now the selector, so F2 is withdrawn as a failure criterion and P2p and P3p are withdrawn as predictions — they are true by construction. The FROZEN exclusion count becomes the gate result and is reported whatever it is. N1's plain rendering (NEUTRAL) is retained as a secondary measurement and is not a selector.
What this does not repair, and the design now says so rather than implying otherwise. Finding 1's deeper claim is that the run tests whether a briefed bilingual can write English carrying a relation it already understands. That is what R1 claims is possible, so it is the right object; but a positive result is a claim about availability, never about whether a translator would find the device unprompted. §9 already forbids the second reading and A5 below adds the arm that separates R1 does not work from the lead wrote a poor span.
A2 — DECOY withdrawn, PARA substituted (finding 2, BLOCKING, accepted)
A control constructed to mark nothing cannot fail, and P4p was therefore not a prediction. DECOY is withdrawn. In its place: PARA — a free paraphrase of the FROZEN span, written by a seat (N1), with no brief about any relation and no instruction to avoid marking. Its recovery rate is genuinely unknown, so it is a falsifiable control on "any rewrite whatever helps" — which is the comparison R1's advantage has to survive.
The lead-written DECOY spans stay in materials/sites.json, unused, and are shown to no seat.
The pre-stage-A commitment of FORCED is untouched and the verifier still checks it.
A3 — the arm count (finding 5, SERIOUS, accepted)
§6 said "five renderings" and listed six. Seven arms, and every seat sees all seven per site:
FROZEN, FORCED, FORCEDM, PARA, REAIM, POSITIVE, WRONG. The ordering hash, the grading
prompt and the verifier all operate on seven.
A4 — the thresholds stop moving with n (finding 6, SERIOUS, accepted)
"More than half of admitted sites" passes at 3 of 4 (75%) while the comparator is 5 of 6 (83%). Amended, in fixed units:
- P1p discharges iff FORCED is recovered at ≥ 2/3 of admitted sites and n ≥ 5.
- P1p is falsified iff FORCED is recovered at ≤ 1/3 — the release's own wording.
- Between the two, it is neither, and the result page says so in those words.
- P3p is withdrawn (FROZEN is the selector) and replaced by P3p′: FORCED is recovered at strictly more sites than PARA. This is the comparison that carries R1's claim now.
A5 — a second, independent implementation of R1 (finding 7, SERIOUS, accepted)
One lead-authored FORCED confounds R1 does not work with the lead chose badly at this site, and
the critic named two spans it doubts (S2's my Jew, S6's went under the hammer). Added arm
FORCEDM: a seat is given R1's text verbatim, the source, the gloss and the staging note — and
not A1's statement — and renders the span so the marking appears. Where FORCED fails and FORCEDM
succeeds, the site is reported as a lead-quality failure and not as evidence against R1; where
both fail, it is evidence against R1 at that site.
A6 — WRONG's plausibility is rated by somebody else (finding 8, MINOR, accepted)
WRONG is the sole discrimination control and its plausibility was set by the lead. N1 — which has never seen A1's statement — rates each WRONG span, before grading, for how plausibly it states something the source passage conveys, 1–5. The ratings are committed before dispatch and reported beside the result. A WRONG arm that rates ≤ 2 across the board makes F1's WRONG leg a trivial pass and the result page says so.
A7 — the staging-note leak, screened rather than re-authored (findings 4 and 9, accepted in part)
Accepted and mostly dissolved by A1: with the gate on FROZEN, the staging notes no longer feed the gate at all, which is the substance of both findings. The residual — A1 reads lead-written staging notes — is answered by running the leak screen over the staging notes and the pointers themselves and committing the result before dispatch. Run at this commit: 0 of 10 staging notes and 0 of 10 pointers fire the banned lexicon. That is a mechanical check on grammatical terminology, not proof the notes are neutral, and §9's declared leak stands.
Not accepted: having a seat author the staging notes. A seat that has not read the story cannot write who is present and what is happening, and a seat that has read it introduces a fresh and larger confound. The limitation is declared instead.
Seat allocation after the amendments, with the no-self-judgment rule checked
| role | seat | grades? |
|---|---|---|
| A1 — the yardstick, sees no English | P5 deepseek/deepseek-v4-pro |
no |
| A2 — independent yardstick, stability + declared fallback | qwen/qwen3.7-max |
no |
| N1 — NEUTRAL plain rendering, PARA paraphrase, WRONG plausibility rating | P4 moonshotai/kimi-k3 |
no |
| M1 — FORCEDM, R1-briefed, never shown A1's statement | qwen/qwen3.7-max, independent call |
no |
| graders | P1 gpt-5.6-terra, P2 gemini-3.6-flash, P3 grok-4.5 |
yes |
| critic | z-ai/glm-5.2 |
no |
Registered contingency: A2 and M1 are the same slug in independent calls, which is not a leak because neither prompt contains the other's output. But if A1 fails both attempts and A2 becomes the yardstick, FORCEDM is authored by the same model as the yardstick and FORCEDM is dropped from the run. Fixed here, before dispatch.
Revised budget
Worst case rises from $1.35 to $1.55: the separate gate stage is removed (−$0.25), N1 gains two calls and M1 one (+$0.25), and the grading bodies carry seven arms instead of six (+$0.20). Today's headroom is $3.926537656.