Repository path: workshop/experiments/E-20260805h-content-or-marking/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260805h-content-or-marking |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| senses | style-correspondence, voice, cultural-mediation, accuracy |
| provisional | true |
| internal-judgment-only | false |
| links | wiki/arms/ARM-r1-fresh-pair.md, wiki/findings/results/RS-20260805c-no-loss-to-repair.md, workshop/experiments/E-20260805c-r1-polish/design.md, framework/v0.1/README.md, wiki/findings/results/RS-20260804-yardstick.md, wiki/findings/results/RS-20260804g-yardstick-holds.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/tracks.md, config/models.md, config/budget.md |
E-20260805h — content or marking: what the E-20260802e procedure actually responds to
ARM-r1-fresh-pair step 2 (T5), the arm's closing step. Session S116, 2026-08-05 UTC.
Frozen before any dispatch. Nothing below is edited after the first API call except by a numbered amendment, dated, with its reason and the finding that forced it.
1. The question, and why it is this one
framework/v0.1 §3 prediction 1 says: on fresh Class A sites in a new pair, R1 recovers a marking
at more than half, scored by the E-20260802e procedure. Three runs have now failed to discharge
it. The third (E-20260805c, S111, PL→EN) failed in the most informative way: its pre-registered
admission gate admitted 0 of 8 sites, because the filed close translation already conveyed, to
independent readers, the relation an independent reader stated from the Polish alone.
RS-20260805c §4(b) put a second reading on the record and ARM-r1-fresh-pair constraint 6 made it
binding on this step:
The procedure scores content, not marking. Six of seven arms preserve the passage's content and only the seventh (
WRONG) changes it;WRONGis the only arm that fails. An unbriefed free paraphrase recovered the relation at 8 of 8. Step 2 may not treat a high recovery rate as evidence that a marking is carried until an arm exists that varies the marking with the content held fixed.
This run builds that arm. It is the one measurement that decides whether prediction 1 is testable by its own named procedure, and it is the arm's declared closing condition.
The subject-rule sentence, written before the design
What does this unit teach about translating literature or evaluating translations? — It measures whether a reader's recovery of a social relation from a translation comes from the relation being marked in the English, or from the passage's facts alone: whether the diminutive, the T-form and the honorific that a translator records as losses are doing any work a reader can be shown to detect.
The unit is a principal unit under the subject rule's named exception (wiki/tracks.md): a named
deliverable is blocked — framework/v0.1's only recommendation has an undischarged first
prediction, and this measurement decides whether it can ever be discharged as worded. The arm says
so in its first line.
2. The design in one paragraph
Take the six sites from E-20260805c at which the relation is separable from the propositional
content. For each, put to independent readers a minimal pair: the filed close translation
(FROZEN, frozen at S111), and a rendering that states every proposition of the source and carries
none of the marking (STRIP). Grade both by the E-20260802e procedure, against the same
frozen relation statements written at S111 by a seat shown no English at all. If the graders cannot
tell the pair apart, the procedure responds to content and not to marking, and prediction 1 is
untestable by it.
The wire between the limbs, in one sentence. The translation limb produces the unmarked member of each minimal pair and the proposition inventory that licenses calling it content-identical; the study limb asks whether the project's own grading procedure — and therefore its released framework's first prediction — can tell that member from the marked one.
3. Materials, and what is inherited unchanged
Source. Bolesław Prus, «Kamizelka» (1882), public domain; the same text as E-20260805c.
Inherited without modification, and this is load-bearing:
- the eight admitted sites and their Polish, gloss, staging,
before/aftercontext; - the relation statements (
graded.jsonrelation), written at S111 by seat A1 (deepseek/deepseek-v4-pro) from the Polish and a staging note, shown no English of any kind, and passed twice through the frozen leak screen (constraint 3 of the arm;E-20260804g§4); - the
FROZENspans — the filedT-kamizelka-R04-v1translation; - the
POSITIVEandWRONGspans, verbatim, as instrument ceiling and floor.
Reusing the statements and the two controls verbatim makes this run a direct replication of the
instrument at S111: POSITIVE scored 48 of 48 judgements and WRONG 4 of 48. Those two numbers are
predicted again below, and if they do not come back the primary is withheld.
S5 and S8 remain excluded. They failed the frozen leak screen on both passes at S111 and the
frozen F3 rule removed them from every denominator. Nothing here reopens that.
4. Two of the eight sites are excluded, and the demonstration is part of the result
A minimal pair exists only where the relation can be removed without changing what the sentence says. At two sites it cannot, for two different reasons, and both were established by attempting the strip, not by assertion. Both exclusions are declared here, before dispatch.
S1 — the relation is entailed by the propositions
Polish: — No, jeżeli nie możesz oddać za pół rubla, to już idź. Ja więcej nie dam.
Propositions: (i) condition — you cannot let it go for half a rouble; (ii) directive — go; (iii)
assertion — I shall not give more. The frozen relation statement is "a final ultimatum … the
transaction will end the moment the dealer refuses." That is entailed by (ii) and (iii) together.
The attempted strip, written at S111 before the statement existed (sites.json DECOY), reads
"then that is the end of the matter. I shall not be giving any more" — it drops the directive
and strengthens the finality. No rendering was found that keeps (i)–(iii) and loses the ultimatum.
Excluded: not a minimal pair.
S6 — English has the device, so the site is not Class A
Polish: sprzedano — the impersonal past with no grammatical agent. R1 is addressed to sites where
the target lacks a same-category counterpart. English has one: the agentless passive, which is
exactly what the filed translation uses ("the unneeded things were sold off at auction"). Any strip
must supply an agent the source withholds, which is new content; the S111 DECOY ("were disposed of
at a public auction") is equally agentless and is therefore not a strip at all. Excluded: the site
was never Class A, and that is a finding about the S111 census, entered in the result.
5. The six minimal pairs, with the proposition inventory that licenses each
For every site: the propositions of the Polish, enumerated from the source; the marking device the
target lacks a counterpart for; and the two spans. STRIP must state every proposition and carry
no device from the removed list. Nothing else may differ.
Five of the six STRIP spans are the lead's S111 DECOY spans, committed to
E-20260805c/materials/sites.json before stage A ran and therefore before the relation statements
existed. Three are used verbatim; three carry a declared amendment. Every amendment is listed with
its reason, and no amendment removes a marking beyond the declared device — two of the three make
the pair more minimal and one restores a dropped proposition.
| site | device the target lacks | STRIP provenance |
|---|---|---|
| S2 | diminutive -ek on the ethnic noun Żyd |
S111 DECOY, amended: it dropped a proposition |
| S3 | pejorative-pitying -czyna on kamizelka, with chora |
S111 DECOY, amended for minimality |
| S4 | diminutive -ka on rubel |
S111 DECOY, amended for minimality |
| S7 | second-person singular domyślasz się to the reader |
S111 DECOY, verbatim |
| S9 | diminutive -ek on urzędnik |
S111 DECOY, verbatim |
| S10 | honorific plural państwo with genericising takich |
S111 DECOY, verbatim |
S2 — — Pewnie wielmożny pan chce parasol?... — odparł Żydek.
Propositions: (1) the speaker asks whether the gentleman wants an umbrella; (2) he addresses him as
wielmożny pan; (3) the question is offered as a surmise (pewnie); (4) the speaker is the Jewish
old-clothes dealer; (5) this is a reply to what came before.
Removed in STRIP: the diminutive periphrasis little.
FROZEN: "Then it will be an umbrella the worshipful gentleman is wanting?..." replied the little Jew.STRIP: "Then it will be an umbrella the worshipful gentleman is wanting?..." replied the Jew.
Amendment, declared: the S111 DECOY read "replied the dealer", which drops proposition (4).
The relation statement is about the label, so a span that removes the noun tests nothing. Restored
to the Jew. This is the only span authored after the relation statements were written, and it is
the only one to which that objection applies.
S3 — Szaf na zbiory jeszcze nie mam, a nie chciałbym znowu trzymać chorej kamizelczyny między własnemi rzeczami.
Propositions: (1) he does not yet have cabinets for his collection; (2) he would not want, again, to
keep the waistcoat among his own things; (3) the waistcoat is in bad condition (chora).
Removed in STRIP: the diminutive periphrasis scrap of a and the personifying adjective ailing.
At this site the marking is carried jointly by the suffix and the adjective, so the strip removes
both; this is the site where the content/marking boundary is thinnest and it is named in the limits.
FROZEN: I have no cabinets for my collection yet, and I should not care, again, to keep an ailing scrap of a waistcoat among my own things.STRIP: I have no cabinets for my collection yet, and I should not care, again, to keep an old and damaged waistcoat among my own things.
Amendment, declared: the S111 DECOY ended "among my own possessions"; restored to things so
that the pair differs by the marking alone.
S4 — — Da wielmożny pan... rubelka! — odparł, roztaczając mi przed oczyma towar w taki sposób, ażeby okazać wszystkie jego zalety.
Propositions: (1) he names one rouble as the price; (2) he addresses the buyer as wielmożny pan;
(3) he spread the goods out before the narrator's eyes; (4) so as to show all their merits.
Removed in STRIP: the endearing diminutive on the price.
FROZEN: "The worshipful gentleman will give... a little rouble!" he answered, spreading the goods out before my eyes in such a way as to show all their merits.STRIP: "The worshipful gentleman will give... one rouble!" he answered, spreading the goods out before my eyes in such a way as to show all their merits.
Amendment, declared: the S111 DECOY read "will give one rouble for it!" without the ellipsis;
the ellipsis and for it are restored/removed so the pair differs by a little → one alone.
S7 — Patrząc na to, odrazu domyślasz się, że właściciel odzienia zapewne codzień chudnął i wreszcie dosięgnął tego stopnia, na którym kamizelka przestaje być niezbędną
Propositions: (1) looking at it, the inference comes at once; (2) the owner must have grown thinner
every day; (3) and reached at last the point at which a waistcoat stops being necessary.
Removed in STRIP: the second-person address (you guess → one guesses). English has no T/V
distinction; the marked/unmarked contrast available to it is you against one.
FROZEN: Looking at it, you guess at once that the owner of the garment must have grown thinner every day, and reached at last the point at which a waistcoat stops being necessarySTRIP: Looking at it, one guesses at once that the owner of the garment must have grown thinner every day, and reached at last the point at which a waistcoat stops being necessary
S9 — Był to drobny urzędniczek, który na naczelników wydziałowych patrzył z takim podziwem, jak podróżnik na Tatry.
Propositions: (1) he was a clerk; (2) of low rank (drobny); (3) he looked at his department heads
with wonder; (4) as a traveller looks at the Tatras.
Removed in STRIP: the diminutive periphrasis minor little. The simile in (4) is source content
and stands in both spans, so STRIP removes only the component the target lacks a counterpart
for — see §10, limit 3.
FROZEN: He was a minor little clerk, who looked at his department heads with the sort of wonder a traveller feels at the Tatras.STRIP: He was a clerk of low rank, who looked at his department heads with the sort of wonder a traveller feels at the Tatras.
S10 — bo służąca przeniosła się do takich państwa, którzy płacili jej trzy ruble na rok i codzień gotowali obiady.
Propositions: (1) the servant moved away to other employers; (2) who paid her three roubles a year;
(3) and cooked dinners every day.
Removed in STRIP: the genericising phrase the sort of people.
FROZEN: because the servant had moved to the sort of people who paid her three roubles a year and cooked a dinner every day.STRIP: because the servant had moved to another family, who paid her three roubles a year and cooked a dinner every day.
6. Arms, five, graded in one dispatch
| arm | what it is | authored |
|---|---|---|
FROZEN |
the filed T-kamizelka-R04-v1 rendering |
S111, before this design existed |
STRIP |
the unmarked member of the pair | five spans S111 (pre-statement), three amended above |
REPEAT |
byte-identical to FROZEN, under a different letter |
— |
POSITIVE |
an explicit statement of the relation | S111, verbatim |
WRONG |
an explicit statement of a different relation | S111, verbatim |
REPEAT is the control this instrument has never had. Two byte-identical spans appear under two
letters in every body. A seat that grades them differently is measuring nothing; the disagreement
rate is the procedure's own precision, and any FROZEN − STRIP difference has to be read against
it. It is declared here so that the duplication cannot later be mistaken for a defect.
Grading procedure, unchanged from E-20260802e as repaired at E-20260804g. Three seats — P1
openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same three as S111.
Each sees the five spans per site unlabelled as A–E, in an order fixed by
sha256(site | seat | ordering) and never chosen by the lead; ordering 2 is the reverse of ordering
1. Six bodies, 30 items each, 180 judgements. A seat counts as recovering a site only if it says
YES in both its orderings; a site counts as recovered when ≥ 2 of 3 seats do. max_tokens
20,000 on every grading seat from the start — S111 lost $0.174 to two P2 bodies at a 12,000 ceiling
and the A9 fix was a bigger ceiling for that seat; note (bhf).
7. The screens, run before grading, both gates
G-content. Two seats that are not graders — mistralai/mistral-medium-3-5 and
qwen/qwen3.7-max, reserve z-ai/glm-5.2 — see the six pairs unlabelled and are asked one
question and one only: do these two passages state the same facts — the same people, actions,
objects, quantities, and the same words spoken? Ignore every difference of tone, attitude, warmth,
politeness, sympathy, emphasis or style. They are shown no relation statement, no source, and no
description of the study. A site is admitted only if both seats say the facts are the same.
G-prose. The same seats, in the same call, answer for each of the two passages: is this
competent English prose, yes or no? A STRIP span called incompetent by either seat is removed —
a rendering that fails as English would fail the grading for a reason that is not the marking.
Why these screens and not a marking screen. No seat can be asked which of these marks the relation without being told the relation, which is the answer. The marking side is therefore established by construction and declared per site in §5, mechanically checkable against the two spans, and the pre-run critic is asked to attack it.
The metalinguistic screen, mechanical. No STRIP or FROZEN span may state the relation rather
than enact it — that would make it a POSITIVE. Checked by analysis/screen.py: the content words a
span shares with its own relation statement, minus those already present in the source's denotation
as glossed in §5, must be empty. Any hit removes the site.
8. Predictions, registered
Let r(arm) be the per-judgement YES rate over the admitted sites (6 judgements per site per arm at full return), and n the number of admitted sites.
| # | prediction | scored how |
|---|---|---|
| PM (primary) | the signature of a procedure that responds to marking: FROZEN − STRIP ≥ 0.25 with exact paired-permutation P ≤ 0.05 |
site-level means, labels permuted within site, all 2ⁿ arrangements enumerated |
| PM′ (primary, the other side) | the signature of a procedure that responds to content only: FROZEN − STRIP ≤ 0.10 and STRIP ≥ 0.70 |
as above |
| PS | STRIP recovers at > half the admitted sites |
site counts, ≥ 2 of 3 seats |
| PC1 | POSITIVE ≥ n−1 sites and r ≥ 0.90 — replicating S111's 1.000 |
|
| PC2 | WRONG ≤ 1 site and r ≤ 0.20 — replicating S111's 0.083 |
|
| PN | REPEAT and FROZEN disagree in ≤ 2 of 36 paired judgements |
byte-identical spans |
| PF | FROZEN recovers at all six sites, as at S111 |
Neither PM nor PM′ can hold with the other. Between them lies a declared indeterminate band, and if the run lands there the page says so and claims nothing.
The lead's registered expectation, written before dispatch and scored afterwards. STRIP ≥ 0.70,
FROZEN − STRIP ≤ 0.10 — i.e. PM′, the procedure is content-only. S111's lead expectation was
wrong in a specific way and saying so was worth more than the prediction; the same discipline here.
9. Failure criteria
F1— the instrument. PC1 or PC2 fails → the primary is withheld; the run reports the instrument's behaviour and nothing about R1.F2— precision. PN fails (more than 2 of 36 identical pairs graded differently) → the primary is withheld; the procedure's own repeat noise is the result.F3— content. A site whose pair is judged content-different by either G-content seat, or whoseSTRIPfails G-prose, or which fires the metalinguistic screen, is removed. Fewer than 5 sites surviving → the primary is withheld and the gate count is the result.F4— seats. Fewer than 3 seats × 2 orderings return → declare, report on what returned, and do not silently rescale. Note (bhf).F5— the band.FROZEN−STRIPstrictly between 0.10 and 0.25, or ≥ 0.25 with P > 0.05 → indeterminate, declared as such, no claim either way.
10. Limits, written before the numbers exist
- Nothing about quality. Tier D is NOT PASSED; evidence class X3 is inadmissible. This run
measures availability — whether a relation is recoverable — and judges no translation. Every
sentence of the result is
provisional. - Nothing about human readers. Three language models are not a survey (charter §4).
- Six sites, one story, one translator, one pair. A census, not a sample. No sampling inference is available and none will be drawn.
- The relation statements are composite. At S9 and S10 part of the stated relation is carried by
source content that stands in both spans (the Tatras simile; the employers' three roubles).
STRIPremoves only the component the target lacks a counterpart for, which is what R1 is about — but it means a highSTRIPrate at those sites is partly explained by content the design did not remove. - S3 is the thin case. There the marking is carried jointly by a suffix and an adjective present in the source, so the strip is closer to a content change than at the other five sites.
- Contamination is declared and NOT measured, exactly as at S111: three English «Kamizelka»
renderings exist in print (Jopson 1930–31, Scherer 1955, Sues 1960) and none is freely reachable,
so
tools/dependence_check.pyhas no comparator. The design's validity does not turn on independence: both members of every pair are lead-written or lead-frozen and the question is what the graders respond to, not whether the English is original. The declaration is a declaration without a number and is repeated here rather than quietly dropped. - The
STRIPspans are the lead's own. Five were committed before the relation statements existed, which removes the worst version of the objection but not all of it: the lead chose which devices count as marking, and §5 states those choices per site so they can be disputed.
11. Pre-flight cost estimate — worst case from max_tokens, per note (abc)
| stage | seats | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic, up to 2 passes | z-ai/glm-5.2, reserve moonshotai/kimi-k3 |
24,000 | $0.13 |
| G-content + G-prose | mistral-medium-3-5, qwen3.7-max |
8,000 | $0.06 |
| grading, 3 seats × 2 orderings | P1, P2, P3 | 20,000 | $0.80 |
| total | $0.99 |
Declared worst case: $1.15, allowing for the routing caution in config/models.md (a slug may be
served by a provider charging several times list). UTC day 2026-08-05 stands at $3.166196086 of
$5.00 with $1.833804 of headroom, so the worst case fits with margin. Lead translation and all
analysis are free and are never ledgered (charter §3, A4).
12. Verification
analysis/verify.py imports nothing from analysis/score.py and recomputes every reported number
from the stored raw bodies: the parse, the per-seat both-orderings rule, the site counts, the
per-judgement rates, the permutation P by exhaustive enumeration, the REPEAT/FROZEN disagreement
count, and the billed cost re-summed per request and cross-checked against the key-usage delta.
Mutation tests: at least three deliberate corruptions of the stored data, each of which must be
caught.
Amendments
(numbered, dated, with the finding that forced each; appended only, never rewritten)
Pre-run critic, 2026-08-05. deepseek/deepseek-v4-pro (P5 — not a grading seat, not a screen
seat, and it produces no data in this run), max_tokens 24,000, effort low, 636 s,
$0.02631576. NEEDS-AMENDMENT, three findings, one BLOCKING, all three accepted.
The first seat produced nothing and was not billed. z-ai/glm-5.2 at effort high returned no
body in 600 seconds and was killed by the wrapper; the key-usage delta across it is exactly
0.000000000. Note (bhf) fires for the twelfth recorded time and in a new mode — not a
length body with hidden reasoning, but no response at all — and rule (iii) applied as written: the
seat was changed, not the ceiling.
A1 — the content screen's wording (finding 1, BLOCKING, accepted in full)
The G-content question in §7 asked whether the two passages state "the same words spoken". At S4
and S2 the marking sits inside quoted dialogue (a little rouble / one rouble; the little
Jew / the Jew), so a seat reading the question literally would call the pair fact-different
for a difference of wording that is exactly the variable under test. That could have removed the
dialogue sites, dropped the census below F3's floor of five, and withheld the primary for a reason
belonging to the screen and not to the materials.
The question is replaced, before any screen dispatch, by the critic's own wording:
Do these two passages state the same facts — the same people, actions, objects, and quantities? For any speech, judge only whether the same propositional content is conveyed, ignoring exact wording.
The instruction to ignore tone, attitude, warmth, politeness, sympathy, emphasis and style stands
unchanged. run_screens.py carries the new text and no other change.
A2 — S3 gets a pre-registered sensitivity analysis, not a repaired span (finding 2, SERIOUS, accepted; the prescribed alternative remedy declined in writing)
The critic is right that at S3 the source's chorej kamizelczyny carries a physical component —
smallness, a scrap — that STRIP's old and damaged waistcoat does not state, so a difference
measured there could be content rather than marking. Of the two remedies offered, the first is
declined: adding small to STRIP would put back the very diminutive force the strip exists to
remove, which is the defect the repair is meant to avoid.
The second is adopted and registered here, before any datum exists. analysis/score.py and
analysis/verify.py report the primary twice — over all admitted sites, and over all admitted
sites except S3 — each with its own exact permutation P. If the two disagree about which
signature holds (PM, PM′ or the indeterminate band), the result page reports the S3-excluded figure
as the primary and says so, because that is the conservative reading. §10 limit 5 stands and is
now backed by a number rather than by a caution.
A3 — the two aggregation rules named apart (finding 3, MINOR, accepted)
PS uses the strict rule — a seat recovers a site only when it says YES in both its orderings,
and a site is recovered when ≥ 2 of 3 seats do. PM, PM′ and the permutation test use raw
per-ordering judgements, six per site per arm. The two answer different questions and neither is
derived from the other. analysis/verify.py recomputes each separately and asserts both against
scores.json.
A4 — the gate fired at 1 of 6, and what the grading stage may and may not conclude (written BEFORE the grading dispatch)
F3 FIRES. One site of six survives G-content: S7. Both seats admitted S7; both called S2, S4,
S9 and S10 fact-DIFFERENT; they split at S3 (G1 DIFFERENT, G2 SAME). G-prose passed every span,
12 of 12. The primary — PM, PM′ and the permutation test — is WITHHELD, and nothing the grading
stage returns can restore it. That sentence is written here, before the grading dispatch, so that
no number arriving afterwards can be read as reviving a prediction its own gate has already killed.
The grading stage is dispatched anyway, and the reason is stated in advance. It carries two
controls that do not depend on the gate — REPEAT, the first measurement this procedure has ever had
of its own within-body precision, and the POSITIVE/WRONG replication against S111's 1.000 and
0.083 — and it answers one question the screens raise but cannot settle: whether the content
difference the screens name is a difference that changes what a reader takes from the passage.
Everything it returns about FROZEN against STRIP is descriptive only, on all six sites, and
is reported as such.
Three descriptive predictions, registered now:
| # | prediction |
|---|---|
| PD1 | STRIP recovers at a per-judgement rate ≥ 0.70 |
| PD2 | FROZEN − STRIP ≤ 0.10 over the six sites |
| PD3 | at S7, the one gate-admitted site, FROZEN and STRIP recover at the same seat count |
A note on the two screen seats. mistralai/mistral-medium-3-5 returned HTTP 400 twice —
"top_p must be 1 when using greedy sampling", the provider-side defect S115 met and named as note
(bjh); the runner's S115 repair stored the error body, so it was diagnosed from the first
failure instead of by hand. The declared reserve z-ai/glm-5.2 took the seat under note (bfc).
It is the same slug that produced no critic body, and it saw nothing of this design: no body ever
returned from that dispatch, so it comes to the screen with no more knowledge than any other seat.