Repository path: workshop/experiments/E-20260803f-craft-carriers/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260803f-craft-carriers |
| status | frozen |
| created | 2026-08-03 |
| updated | 2026-08-03 |
| senses | accuracy, naturalness, voice, style-correspondence, cultural-mediation, affect |
| purpose | Readers of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs (D-20260801-10). |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/translations/odnazhdy-osenyu/R04-v1/translation.md, workshop/translations/odnazhdy-osenyu/R14-v1/translation.md, workshop/translations/odnazhdy-osenyu/R06-v1/translation.md, workshop/regimes/R14-matched-flattening.md, wiki/arms/ARM-first-judgment.md, wiki/findings/results/RS-20260803-a4-set.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/goodness-senses.md, config/models.md |
E-20260803f — what carries literary life, when nothing propositional differs
ARM-first-judgment step 3 of 3 (T3). FROZEN 2026-08-03 before any call was dispatched. Both
translator's logs were frozen and committed before this file existed: T-odnazhdy-osenyu-R04-v1
at 827a63c, T-odnazhdy-osenyu-R14-v1 at d790968. The A4 freeze condition therefore holds by
construction and is not re-argued (ARM-first-judgment §Constraints).
Standing, and it governs every number this design will produce. Tier D was run at S086 and
NOT PASSED. No score here carries evidential weight, none may support a framework
recommendation, and everything is provisional and internal-judgment-only.
1. The question, and what it teaches
What carries literary life in a translation, independently of what the translation says — and is any of it visible to an evaluation that scores six senses?
The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the unit was
designed: this unit teaches which non-propositional properties of English prose a reader actually
registers, by removing a named list of them from a translation and seeing which removals are noticed
and named by readers who were told nothing about the list. That is a claim about translating
literature. Its second limb — whether the project's own six-sense rating sees the difference — is
ARM-first-judgment's clause 3, every result page states plainly what its scores license and what
it does not, and is discharged here by demonstration rather than by assertion.
Why this is not "scoring more translations", which step 3 is explicitly not a licence to do. The A4 set scored five filed translations and kept the A4 promise. Nothing here adds to that count. The pair rated in stage 3 was built for this question and exists only to locate the boundary of what the A4 scores mean: Tier D established that the panel separates damage at ceiling; the A4 set established that it does not separate five competent translations. The interval between those two facts has never been measured, and a sentence about what the A4 scores license cannot be written honestly without it.
2. The wire between the limbs, in one sentence
The translation limb generates the object the study limb investigates: two renderings of the same 494 Russian words built to differ in nothing a paraphrase could report, so that any difference a reader finds is a difference in the writing and nothing else. Whether they succeed in that is what stage 1 decides, and this sentence does not pre-empt it (amendment A3, pre-run critic pass 1 finding 2, BLOCKING: the original wording asserted the conclusion the gate exists to test).
3. Materials, frozen
| id | what | words |
|---|---|---|
| SOURCE | Максим Горький, «Однажды осенью» (1895), ¶84–99, PD, ../../translations/odnazhdy-osenyu/R04-v1/source-ru.txt |
494 (RU) |
| LIVE | T-odnazhdy-osenyu-R04-v1, lead, R04, log frozen at 827a63c |
669 |
| FLAT | T-odnazhdy-osenyu-R14-v1, lead, R14 v0.1, 37 operator sites, log frozen at d790968, repaired at four sites by amendment A1 |
688 |
Both English texts are extracted from the frozen pages by materials/extract.py, which writes
LIVE.txt and FLAT.txt and recomputes every length figure. Length ratio 1.0284 after A1 (0.9925
before) — FLAT is 2.8% longer, still well inside the ±15% declared tolerance and smaller than the
+4.5% confound declared on this project's damage control (RS-20260802-tierD-verdict).
Sentence-length SD falls 24.0 → 7.7 under the strict splitter; both splitters are reported
because the mean reverses between them and the SD does not.
3a. Amendment A1 — the pre-run critic's first BLOCKING finding, and the repair it forced
Pass 1 (P4, NEEDS-AMENDMENT, six findings, two BLOCKING, all six accepted) found that the F4
operator as written DELETED content rather than restating it, and that the equivalence prompt's own
exclusion list would then have told the gate seats to suppress exactly the items FC1 exists to catch.
It named four sites from the frozen text. A mechanical check, materials/coverage.py, was then
written and run before any further dispatch — every content word present in LIVE and absent from
FLAT, exact and after crude stemming — because the finding was computable from the frozen materials
alone (note (bhr)).
| site | the deletion | repaired to |
|---|---|---|
| ¶84 | "like an owl!" — the simile of «как сыч» — gone with no replacement, and not in the site table at all | "…say nothing, like an owl." |
| ¶89 | "The wind howled and moaned" → "was blowing loudly": the moaning erased | "was blowing and making a howling and moaning sound" |
| ¶89 | "the rain drummed" → "was falling on": the drumming erased | "was falling on the boat and making a drumming sound" |
| ¶89 | "the waves splashed" → "were moving": the splashing erased | "were moving and splashing" |
| ¶93 | "Many kisses, past counting, and hot" → "many times and warmly": uncountability and heat both reduced | "many times, more than could be counted, and her kisses were hot" |
R14's F4 definition was wrong and is amended, not worked around: a deadening operator that
deletes is a damage operator. F4 now replaces the concrete verb with a generic verb plus an
explicit statement of the manner it carried — the manner is stated rather than enacted, which
is the property the operator was for, and nothing is lost.
This is the finding of the run so far and it cost $0.10 to buy. The lead wrote a 37-site log believing the operator was propositionally conservative, and it was not, at five sites, one of which never reached the log. Note (bic).
The operator, in one line: F1 cadence levelling ×5 · F2 figure de-specification ×6 · F3
repetition flattening ×3 · F4 verb deadening ×3 · F5 register levelling ×11 · F6 connective
explicitation ×9. Definitions in workshop/regimes/R14-matched-flattening.md; the 37 sites in
FLAT's log; the tally parsed from that log by analysis/checks.py, never counted by hand.
4. Seats
Panel roles per config/models.md. Every call is stateless and no seat carries information
between calls; overlaps below are therefore declared, not hidden.
| stage | seats | why these |
|---|---|---|
| 0 pre-run critic | P4 moonshotai/kimi-k3 |
takes no measured role anywhere in this run. effort: low on the first dispatch, note (b) |
| 1 equivalence gate | P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 |
factual adjudication, where S015 measured the panel strongest (20/20 planted false claims rejected). Neither has seen the design |
| 2a naming (PRIMARY) | P1, P3, P5 × 2 orders | three labs; free text, no categories supplied |
| 2b coding of 2a's free text | P2, one call per response | P2 is the only seat that is not a namer, so no seat codes its own words. The lead codes the same 36 cells independently as a declared internal-judgment-only check; P2's coding is the primary and the lead's is never substituted for it |
| 3 six-sense rating | P1, P2, P5 × 2 texts × 2 passes | the A4 set's own jurors, on the A4 set's own prompt format, so RS-20260803-a4-set's measured floors transfer: retest 0.233, paraphrase 0.139, gross damage 4.333 on accuracy |
5. Procedure
Stage 1 — the equivalence gate, run first and before anything else is dispatched. Each seat is
given SOURCE, LIVE and FLAT and asked to list every place the two English texts differ in
propositional content, tagging each item from a fixed vocabulary: negation | quantity |
referent | tense-aspect | added-content | removed-content | lexical-specificity | force | other. The
prompt states that differences of rhythm, register, figure, punctuation and sentence division are
not propositional differences and must not be listed. Order of presentation is swapped between
the two seats.
Stage 2a — the naming task, the primary. Each seat sees the two English texts as A and B, in both orders, with no source, no authorship, and no categories: "These are two English renderings of the same passage of Russian prose. Describe, as specifically as you can, how they differ as pieces of writing. Quote from both." Nothing is asked about quality and nothing is asked about the source.
Stage 2b — coding. P2 receives one stage-2a response at a time, together with the six operator definitions (not the site table, not the texts), and answers yes/no per category with the quoted words that name it. Six calls, one per response. The lead codes the same 36 cells from the stored responses, blind to P2's output, as a declared check.
Stage 3 — the six-sense rating. The A4 protocol, unaltered: the juror sees SOURCE and one
English text, rates it 1–7 on the six senses of wiki/goodness-senses.md, alone, with authorship
stripped and no comparison available. Two passes per (text, juror).
Control C1 — the naming false-alarm floor. Each of the three naming seats additionally receives LIVE against LIVE under the identical stage-2a prompt. A seat that manufactures differences between identical texts cannot be read as having found them between different ones. Three calls.
Judgment is never parallelized. Stages run in order, and stage 3 is not dispatched until stage 1 has returned AND FC1 has been evaluated — amendment A8: "read" is not "adjudicated", and stage 3 is the $0.16 a fired gate exists to protect.
6. Registered predictions
| # | prediction | threshold |
|---|---|---|
| P1 | The operator set is nameable. Blind seats, given no categories, name most of it | ≥ 4 of the 6 categories named by ≥ 2 of 3 seats (a category counts for a seat if named in either order) |
| P2 | Presence-differences are visible and absence-differences are not. Cadence and register differences are on the page in both texts; a de-specified metaphor leaves nothing marked to see. A9 (pass-2 finding N3): the original rationale said "leaves nothing to see", and after A1 that is false of this stimulus — every F2 site now replaces the figure with a literal statement that is present on the page, so the mechanism under test is salience, not absence. A5, registered asymmetry: with three seats the second limb can only fail on 3-of-3, so it has almost no power to be wrong. It is reported as descriptive and is not counted toward the run's primary | F1 and F5 named by 3 of 3 seats; F2 named by ≤ 2 of 3 |
| P3 | The six-sense rating separates the pair, but far below damage. A6: all six gaps are reported, and the max-over-six selection is named in the result's licence sentence — selecting the largest of six and testing it against a floor inflates the pass rate under noise | pooled |LIVE − FLAT| on the largest-gap sense ≥ 0.233 (the A4 retest floor) and ≤ 2.00 (BAR-D's accuracy signal is 4.333) |
| P4 | accuracy does not separate them — the operator was propositionally conservative, checked by the one sense Tier D showed this panel detects at ceiling |
|LIVE − FLAT| on accuracy ≤ 0.50 |
| P5 | naturalness moves toward FLAT. Under its post-D-20260802-13 wording — distance from unmarked literary-contemporary English, on the target alone — the flattened text is by construction nearer the unmarked point |
FLAT ≥ LIVE on naturalness |
P5 is the prediction that matters to the typology and it is registered before any number exists.
If the flattened text is called more natural by the rating while the naming seats describe it as
the duller piece of writing, that is a fact about the sense, not about the panel — and it is the
first direct evidence on whether the struck escape clause left naturalness able to reward
flatness.
7. Failure criteria
- FC1 — equivalence (the gate). If both stage-1 seats independently report the same
difference tagged
negation | quantity | referent | tense-aspect | added-content | removed-content, the operator is not propositionally conservative: P3, P4 and P5 are withheld, the run is reported as a damage run, and P1/P2 stand only with that stated at the head of the result. Items reported by one seat are printed verbatim and adjudicated in the open; the lead's adjudication never replaces the seat's raw count. One item is in the adjudication queue before dispatch (pass-2 finding N2, computed from the frozen materials): LIVE "as through a dream" → FLAT "as though I were half asleep", for «как сквозь сон». Those are not the same state. It is put on the record here rather than discovered live; it cannot fire FC1 alone, which needs both seats on the same item. - FC2 — parse. Any
finish_reason: lengthis a seat failure, never a partial answer. The body is preserved, the cost is ledgered, the call is re-dispatched once with a larger cap; a second failure withdraws that cell. - FC3 — order. If a naming seat's two orders yield disjoint category sets (Jaccard 0), that seat's naming is reported per order and not pooled.
- FC4 — coder reliability. If P2's coding and the lead's disagree on more than 12 of the 36 (response × category) cells, the coding is reported as unreliable, and P1/P2 are reported with the disagreement printed rather than resolved.
- FC5 — ceiling. Any rater at 7 on more than 50% of its cells has its between-text comparisons flagged (A4 F4; two of three A4 jurors were at ceiling on ~72% of cells, so this is expected to fire and its firing is not news).
- FC6 — completeness. All six senses present in every rating body; every stage-1
KINDtag drawn from the fixed 9-item list (amendment A7 — the original said "all 34", a number belonging to nothing in this design). - FC7 — naming false-alarm. If any seat reports substantive differences between the two identical LIVE texts of control C1, that seat's stage-2a output is withheld from the primary and reported separately.
8. What this run cannot show, registered before it ran
- The operator is the lead's own taste. F1–F6 are what the lead believes carries literary life. A high P1 shows the categories are nameable by others, not that they are the right categories. Nothing here surveys the space of craft properties; six were chosen and six were tested.
- The readers are language models. Naming a cadence difference is not evidence that a human reader would feel it. The demand pathway is live: a seat asked how two texts differ will look for differences.
- One locus, one language pair, one translator, one direction of flattening. No claim about Russian, about Gorky, or about translation in general follows from 669 words.
- F5 changes illocutionary force, which is not propositional content but is not nothing. Eleven of thirty-seven sites are F5. The equivalence gate is instructed on propositional content specifically, and the reading of the whole run has to hold that.
- Tier D is NOT PASSED. Stage 3's numbers are a description of what this panel does, not a measurement of how good either text is.
- The A4 floors are imported, not re-measured within this run — same jurors, same prompt format, different text. Stage 3's two passes give a within-run retest floor as well, and both are reported; if they disagree the imported one is not used.
9. Cost
Worst case built from max_tokens, not from an assumed output length (note (abc)).
| stage | calls | cap | worst case |
|---|---|---|---|
| 0 critic | 2 | 6,000 | $0.234 |
| 1 equivalence | 2 | 3,000 | $0.056 |
| 2a naming + C1 | 9 | 2,000 | $0.14 |
| 2b coding | 6 | 1,500 | $0.084 |
| 3 rating | 12 | 1,500 | $0.16 |
| retry reserve | — | — | $0.25 |
| total | 31 | $0.92 |
Today's UTC ledger before this run: $2.517 of $5.00 spent (S094–S098 plus this session's ratification gate at $0.0556), $2.483 headroom. The reservation fits. P5's list price is not what gets billed — the S022 caution — so P5 is priced at 4× list here.