Repository path: workshop/experiments/E-20260728d-jeli-marks/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728d-jeli-marks |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | style-correspondence, voice, consistency |
| links | wiki/arms/ARM-longwork.md, workshop/translations/jeli-il-pastore/R05-v1/translation.md, workshop/translations/jeli-il-pastore/register.md, workshop/regimes/R05-serial-long-work.md, config/models.md |
E-20260728d — the complete guillemet census of «Jeli il pastore», and whether a reader can recover what the marks point at
Frozen 2026-07-28 (S047), before the span-5 English of site G8 existed and before any panel call was dispatched. The critic pass is dispatched against this text; every change made in response to it is recorded in §10 with the finding that forced it, and nothing else is changed.
1. Why this exists, and what makes it possible only now
ARM-longwork step 5 owes U6 — the last of the register's unresolved rows, and the only site in the
whole novella where two rival explanations of the same register row make opposite predictions.
Row V7 says: Verga's guillemets are kept as one mark — single inverted commas, no gloss, no
expansion. Four spans have produced two competing accounts of when that works:
- The proverb account (span 1,
D9): guillemeted proverbs take English marks without strain, because a proverb announces itself as borrowed speech; the marks are then ordinary quotation. - The prior-attribution account (span 4,
D62): what makes the marks work is that the marked words have been attributed somewhere earlier in the text. Span 2's three failures (D21) were failures of first appearance, not of grammar.
¶171 is the one site where they come apart: «chè le corna sono magre, ma mantengono la casa grassa!» is a proverb, it is unframed, and it has never been attributed anywhere in this text. The proverb account predicts it works; the prior-attribution account predicts it does not.
And a whole-work census is the only way to ask. There are exactly eight guillemeted spans in
11,474 words. Seven are frozen; the eighth is written this session. No unit under 2,000 words contains
more than two, and the mechanism D62 proposes operates over a 3,900-word span. This is the
ARM-longwork question — what the long form pressures that short units leave idle — in its most literal
form: the object of study does not fit inside a short unit.
2. What the census found before any call was made, and it changes the question
Enumerating the eight sites against the frozen English produced a fact that no session had noticed and that the design must state before it can test anything:
The two span-1 sites are not rendered under
V7at all. They carry double quotation marks — the mark this translation gives to speech (V8) — becauseV7was created atD20/D21in span 2 and was never applied backwards.
D18 (span 1) states that the malaria site was given "inverted commas". The frozen prose gives it
"where you could reap the malaria,". The log records a decision the prose does not execute, and four
subsequent spans passed over it.
Two consequences, both binding on this design:
- The proverb account has no supporting instance under
V7. Its evidence is the two span-1 sites, and those sites do not carry the mark under test. The account is therefore not supported and rivalrous; it is unsupported, andG8is the first V7-marked proverb in the work. - The mark type is now itself a variable, and whether the divergence costs anything is measurable rather than arguable. §6 measures it.
(That this went unseen for four spans is a finding about the instrument, not about this experiment; it is reported in the result page, not here.)
3. The census — all eight sites, fixed before any call
Paragraph indices are 0-based into source-it-full.txt. English word offset is the number of English
words preceding the site in the frozen translation. G8 is written this session under V7 as it stands;
no rule is changed to accommodate it.
| id | ¶ | Italian | English | mark | framed? | prior attribution? | proverb? | class |
|---|---|---|---|---|---|---|---|---|
| G1 | 5 | «Era piovuto dal cielo, e la terra l'aveva raccolto» | "He had rained down from the sky…" | double | yes — as the proverb says | no | yes | framed |
| G2 | 6 | «dove la malaria si poteva mietere» | "where you could reap the malaria," | double | yes — as the peasants round about said | no | no | framed |
| G3 | 28 | «nel fatto suo» | 'on her own ground' | single | no | no | no (bare NP) | bare-NP |
| G4 | 40 | «alle case» | 'the houses' | single | no | no | no (bare NP) | bare-NP |
| G5 | 51 | «case grandi» | 'big houses' | single | no | no | no (bare NP) | bare-NP |
| G6 | 65 | «come San Giovanni fosse arrivato sotto l'olmo,» | 'when San Giovanni should have come under the elm,' | single | yes — the contract said | no | no | framed |
| G7 | 125 | «che sembrava un manico di saliera» | 'that looked like the handle of a salt-cellar' | single | no | yes — Mara, 3,786 English words earlier | no | prior-attributed |
| G8 | 171 | «chè le corna sono magre, ma mantengono la casa grassa!» | written this session | single | no | no | yes | TEST |
The census is complete and mechanically verifiable: verify.py re-derives it from the source by regex
and fails if the count is not 8 or if any site's Italian does not match the row.
REVISION 1 — after the independent pre-run critic pass (2026-07-28, S047)
Everything above this line is the text frozen at commit a45c04f and dispatched to the critic. Nothing
above is edited. Sections 4–10 below replace the frozen §§4–10, which the critic's verdict —
NEEDS-REDESIGN, four blockers — required. Every change carries the finding id that forced it;
critic/dispositions.md is the full record. Two claims made above are withdrawn below: that G8 is a
crucial experiment between two accounts (G4), and that §6 measures whether the mark-type divergence
costs anything (G6).
4. The question, operationalised — REVISED
D21 states the failure precisely, and it is a claim about a reader, not about the translator:
English inverted commas round a bare noun phrase read as scare-quotes — ironic distance — which is the reverse of what Verga is doing.
So "the marks work" = a reader takes them as quoting borrowed words; "the marks strain" = a reader takes them as the narrator holding the words at arm's length.
Instrument. Three panel models (P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5), each shown the English translation from its first word up to and including the
sentence containing one site, and asked about that site only. One call per site per model.
This is a textual-recovery task, not a quality judgment, which is the one use config/models.md
(S015 instrument note) records the panel as usable for. Tier D has not passed; nothing below is a
quality verdict.
The three questions. Order changed per finding G8: Q2 is asked FIRST, unprimed.
- Q2 (asked first; primary for
D62). Whose words are these, on the evidence of the passage you have read? Name the source, or answer UNKNOWN. Coded NAMED (a specific person, document or body of speakers identifiable in the passage) / GENERIC (people in general, a saying, with no locus in the passage) / UNKNOWN. It is asked first because it is the measureD62is a claim about and because any other order primes it. - Q1 (asked second; primary for
D21). What are the marks around this phrase doing? Five options, the fifth added per finding G7, because a proverb is borrowed language without being anyone's actual words and the four-option form conflated the two: A = quoting — someone's actual words, reported; B = distancing — the narrator holding the words at arm's length, so-called, ironic or sceptical; C = highlighting — emphasis, or a technical or unfamiliar term; D = fixed expression — a proverb or set phrase, conventional language rather than any particular person's words; E = cannot tell. - Q3 (asked third; DESCRIPTIVE ONLY, per finding G5). How well does the marking earn its place? 1–7. No test is run on Q3.
Residual, declared (G8). Order is fixed rather than randomised — randomising would triple the calls — so Q2 is the only unprimed measure and Q1 is read having already been asked to name a source. Stated, not repaired.
Site verdict = the majority category across the three models, on Q1 and on Q2 separately. No majority = SPLIT, which is neither a category nor a settlement.
5. Predictions — REVISED
Confirmatory (pre-registered, on prose frozen in earlier sessions that the lead cannot now change):
P1. G1, G2, G6 (framed) → Q1 = A at all three sites.
P2. G3, G4, G5 (bare-NP) → Q1 = B at a majority of the three sites. This is D21's own stated
cost. If it does not hold, D21 overstated what it was paying.
P3. G7 → Q2 = NAMED, naming Mara. This is D62's mechanism, at D62's own site. If G7 does not
come back NAMED, the mechanism D62 asserts is not available to a reader at 3,786 words' distance and
D62 is weakened on the evidence it was built from.
Exploratory — never confirmatory (finding G10: G8's English is written by the translator who knows these predictions, and no blinded independent translator exists in this environment):
P4 — a one-sided test of D62 alone. The claim that G8 tests is D62's necessity claim: that
prior attribution is what makes the marks work. The frozen §5's two-account framing is withdrawn
(finding G4) — the proverb account has no instance under V7 and is not fit to be an equal rival.
Outcomes, and they are exhaustive (finding G2 — the frozen form fired on every outcome and therefore had
a false-alarm rate of 1, which is note (bcs) recurring one session later):
| Q1 at G8 | what it means |
|---|---|
| A | prior attribution is not necessary — D62's necessity claim fails at the one site that can test it |
| B | consistent with D62; the marks strain where nothing has been attributed |
| D | a third mechanism neither account named — the marks read as marking a fixed expression |
| C, E, or SPLIT | no information, and U6 is recorded as not settled by this instrument |
The sentence "the U6 row is settled by whatever P4 returns" is deleted. A null outcome is a null.
Descriptive only, no test (finding G5):
P5. Q3, computed as the mean of the three model ratings per site, then the mean over sites in each set, reported for {G1, G2, G6, G7} against {G3, G4, G5}. Reported as two numbers, with no inference. P6 — mark type, one-sided (finding G6). At G1 and G2, a change in the modal Q1 category between the frozen double-quote rendering and a single-inverted-comma variant is a positive signal that mark type costs something. No change establishes nothing — there is no equivalence margin and no power. SPLIT at either rendering is recorded as uninformative for that site. The automatic "an erratum is owed" consequence is removed.
6. The mark-type manipulation — REVISED
Two extra stimuli, G1′ and G2′, identical to G1 and G2 except that the marks at the site are single
inverted commas. The frozen prose is not altered; the variants are stimuli built from it, and verify.py
checks that each differs from its original in exactly the two mark characters.
10 stimuli × 3 models = 30 calls. The frozen §6's claim that this measures "whether the divergence costs anything" is struck (G6); it can only detect a cost, never establish its absence.
7. The instrument gate — REVISED
GATE. The framed set {G1, G2, G6} must return Q1 verdict A at all three sites and the bare-NP set {G3, G4, G5} must return Q1 verdict B at at least two of three.
If the gate fails, the G8 result is not reported as bearing on
D62,U6is recorded as not settled by this instrument, and the result page says so in its first paragraph.
The gate is also the control for finding G7's artefact risk, and this is now stated rather than left implicit: if A were an artefact of displaying a phrase already enclosed in marks, the bare-NP sites would return A as well and the gate would fail. The gate is the only thing standing between an A at G8 and that artefact.
Random-response reference value, relabelled per finding G1. With five Q1 options and uniform random answering, a site's majority verdict is any given category with probability 3·(1/5)²·(4/5) + (1/5)³ = 0.096 + 0.008 = 0.104. The gate then fires with probability 0.104³ × [3·0.104²·(1−0.104) + 0.104³] = 0.001124864 × 0.030198272 = 0.0000340.
This is a random-response reference value, not a false-alarm rate under the substantive null. The
critic re-derived the frozen four-option figure (0.000250293, confirming 0.000250) and was right that
uniform random answering neither follows from nor represents the null in which D62 is false; that null's
response distribution is unmeasured. Independence across models, sites and calls is assumed, not
demonstrated.
No prediction in §5 claims computable power, and none is asserted. A gate failure is reported as instrument insensitive, never as the lead's classification refuted.
The site classification, per finding G3. Prior attribution is now mechanically defined: the
marked English string, of four words or more, recurring earlier in the frozen English inside a passage of
attributed speech. verify.py computes it. It decides exactly the two sites it must decide — G7 yes
(the phrase recurs verbatim in Mara's reported speech), G8 no (the proverb occurs once) — and is
inapplicable at the three bare-NP sites, where "no prior attribution" remains the translator's reading
and is labelled so. Proverb remains a translator judgment. No independent annotator was available in
this environment; the finding is conceded, not answered.
8. Failure criteria and what is not claimed — REVISED
- n = 8 sites, 1 work, 1 language pair, 1 translator. A census of one work, not a sample.
- G8 is exploratory and is never reported as confirmatory (G10). Its English is written by the
translator who knows these predictions. The rendering constraints declared before it is written:
V3(proverbs literal and strange),D63(horns is a bound word in this text),V7(the mark),V6(no added frame), no gloss. The residual freedom — lexis, rhythm, placement — is an unremoved confound. - An A at G8 does not establish the proverb account (G4). Unexcluded alternatives, named before the run: the marks read as retained source punctuation; as idiom-marking; as an inherited quotation convention; or as an artefact of displaying a marked phrase and asking about it.
- A prefix-fed model is not a sequential reader (G8). Nothing here demonstrates that a model receiving
an 11,000-word prefix occupies the reader-state
D62is a claim about. Conceded. - Position is confounded with context length, plot development, character salience and attention loss (G9). No direction is asserted; the frozen §8's claim that more context favours A, and that a B at G8 is therefore "the harder result", is deleted.
- Model recall of Verga or of a published translation is uncontrolled. Recall would help attribution recovery, so it biases Q2 toward NAMED and toward the gate passing. A model that names Verga or a translator is recorded as such.
- The lead is the translator and is not a rater. The lead's classifications are frozen above, before any call. The lead does not judge its own translation (charter §5).
- No quality claim. Tier D has not passed; Q3 is descriptive and the result page carries
provisional: true.
9. Procedure — REVISED
- ~~Freeze this file.~~ Done,
a45c04f. - ~~Independent pre-run critic pass.~~ Done.
NEEDS-REDESIGN, 10 findings, 10 accepted, $0.036104375.critic/dispositions.md. - This revision, committed before span 5 is translated.
- Translate span 5 and freeze it. G8's English is written under
V7as it stands, in the ordinary course of the span, with the decision logged like any other. - Build the 10 stimuli mechanically from the frozen
translation.md;verify.pychecks each is a verbatim prefix of the frozen English (modulo the two variant marks). - Run 30 calls, temperature 0,
max_tokens1500,reasoning: {"effort": "low"}, raw bodies persisted per call, per-requestusage.costrecorded, key snapshots to disk (note (bco)). - Post-run verification:
verify.pyrecomputes the census, the mechanical prior-attribution column, every count, every majority verdict, the gate, the 0.0000340 reference value, and the cost sum against the call count.
10. Changes made after the critic pass
All of §§4–9 above. Ten findings, ten accepted, none declined; four accepted by withdrawing a claim
the frozen design made. The per-finding record is critic/dispositions.md. What survives as
pre-registered and confirmatory is P1, P2 and P3 — tests of D21's written cost claim and of D62's
mechanism at its own site, on prose frozen in earlier sessions. G8, the reason the experiment was
designed, is the weakest thing in it, and that ordering is the critic's.