Repository path: workshop/experiments/E-20260813b-affect-yardstick/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260813b-affect-yardstick |
| status | frozen |
| created | 2026-08-13 |
| updated | 2026-08-13 |
| senses | affect |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-affect-unprompted.md, wiki/findings/results/RS-20260812f-affect-unprompted.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, workshop/translations/caldura-mare/R06-v1/translation.md, workshop/translations/caldura-mare/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md |
E-20260813b — is "closest to what the original does" a judgment about the original, or a
preference for the prose that reads like the description?
ARM-affect-unprompted step 2. The arm's step 2 is written in the arm page as "write the
verdict into wiki/goodness-senses.md §affect", with the instruction "scope it from the
result, not from here." Scoped from the result, it cannot be done as a writing step, and this
design says exactly why.
Everything here is internal-judgment-only and provisional. affect is untested, Tier D
is NOT PASSED (config/models.md), and no jury verdict in this project carries evidential
weight.
1. Why the writing step is not a writing step
RS-20260812f §5 reported that the two halves of affect — the experience the English produces
in its reader (H1) and how comparable that experience is to the one the source produces in
its own reader (H2) — separate without anyone pointing the translator at either, and that the
separation runs the opposite way to the standard story: the foreignizing arm won H1, the
plain arm won H2 at 14 of 14 segments, and every one of 17 discordant cells moved toward the
plain arm.
Its own limit 5 says the sharpest thing that can be said against that:
H2supplies a yardstick andH1does not, soH2's seats are answering a question with a document in it. A preference for the plainer arm may be a preference for the arm that matches a plainly-written English description.
If that reading is right, the run did not measure a second half of affect at all. It measured
whether a passage of English resembles a passage of English — and then the separation between
the halves, which is what the ratified rule's reversion condition demands, is an artifact of one
question having a document attached and the other not. The discharge and the reversal both
hang on it, so writing either into wiki/goodness-senses.md before testing it would put a
figure into the senses page that a single obvious control could destroy.
Subject rule (wiki/tracks.md, continue-prompt.md §4.5), stated in one sentence. What this
unit teaches about translating literature: whether a translator can be told which of two
renderings comes closer to what the original does to its own reader — the oldest evaluative claim
in translation and the one every "equivalent effect" argument rests on — or whether that judgment
collapses, whenever the judge cannot read the source, into a comparison between the translation
and whatever prose the description happens to be written in. The exception the subject rule
names also applies on its own terms: a named deliverable — wiki/goodness-senses.md §affect,
the arm's declared completion criterion — is blocked by the limit, and the arm says so.
2. Materials
Ion Luca Caragiale, «Căldură mare» (1899), the sketch whole: a man calls at a house in Strada
Pacienței on a 33° afternoon, spends four pages failing to leave a message with a servant who
answers every question exactly and helpfully, discovers he wants Strada Sapienței, and then asks
four more people for Strada Pacienței, which is where he is standing. 120 paragraphs, 916
Romanian words, public domain (Caragiale 1852–1912), read whole. Copy-text: Romanian
Wikisource, fetched 2026-08-13, workshop/translations/caldura-mare/source.txt.
The project's first Romanian source.
Two arms, both lead, both $0, both frozen at 783c87b before this design existed:
| arm | English words | what it is |
|---|---|---|
R06 |
1,142 | T-caldura-mare-R06-v1 — lead single pass, source only, no rule set |
R08 |
1,190 | T-caldura-mare-R08-v1 — + Venuti's ten foreignizing rules, frozen 2026-07-28 |
R07 is not built and is not in this run. RS-20260812f §3's gate found, on two independent
seats, that the fluency programme taken whole is the sense's first half under another name,
and excluded it from the primary. Building it again in order to exclude it again would be paying
for a settled result.
Contamination: none, measured — tools/dependence_check.py, R06 against the whole of
Lucy Byng's Caragiale (Roumanian Stories, 1921, Project Gutenberg #38991, 8,748 words): longest
common run 5 tokens, 0 shared 7-grams, verdict clean. No English rendering of this
sketch exists to compare against; the cross-text figure is reported for what it is on both
translation pages.
Segments: 15, cut at paragraph boundaries chosen on the source's structure by code.py
before any prompt existed, identical for the source and both arms, with reconstruction asserted.
Segment lengths run 31–193 English words.
3. Conditions
All three judging conditions put the same two arms to the same three seats as a forced binary choice, in separate stateless calls, with per-cell label assignment from a fixed seed.
| id | document shown | question |
|---|---|---|
H1 |
none | which of the two does the most to you as a reader of English? |
H2P |
the plain yardstick | which comes closest to doing to its English reader what the original does to its own reader? |
H2M |
the marked yardstick | identical wording to H2P |
H2O |
the ornate yardstick | identical wording to H2P |
The three H2 conditions differ in exactly one thing: the prose style of the document. The
label order for a given (segment, seat) is drawn from the same seed key in all three, so the
conditions are paired cell by cell and nothing but the document moves.
The yardsticks
YP(plain). A source-side seat that judges nothing in this run reads the Romanian segment and writes two fields kept strictly apart: EFFECT — what the passage does to a reader of the Romanian, with every reference to a property of the language forbidden — and CAUSE, where linguistic properties may be named freely. Only EFFECT is ever shown to a judge (RS-20260812f's pre-run critic, BLOCKING 1: a yardstick that names the mechanism turnsH2into feature-matching).YM(marked / estranged). The same seat, in a fresh stateless call, is shown only the EFFECT text — not the Romanian, not the arms, not CAUSE — and rewrites it into markedly non-standard English: dislocated and front-loaded syntax, archaism, unidiomatic literalness, fragments left unrepaired. No claim added, removed or altered.YO(ornate). The same seat, the same input, rewritten into elaborate high-literary English — long balanced periods, latinate diction, rhetorical figure — that stays completely fluent and idiomatic. No claim added, removed or altered. This document is marked relative to plain English but not marked inR08's direction, and it is in the design because the pre-run critic required it (§10).
Holding the writer constant and varying only the style instruction is the point. P2 never
judges in this run, so no seat is judging prose it wrote (charter §5).
4. Predictions, registered before any call
P1— replication, with direction. UnderH1the segment majority prefersR08more often thanR06; underH2Pit prefersR06more often thanR08; and the segments where theH1andH2Pmajorities disagree are one-sided towardR06underH2P. Primary statistic: exact two-sided binomial at p = 0.5 over discordant segments — the same statisticRS-20260812f§5 used, so the two runs are directly comparable.P2— the confound, as a tracking prediction. The confound limit 5 alleges is that anH2judge picks the arm whose English resembles the document's English. That hypothesis makes a sharp, symmetric prediction: theH2choice moves with the style match, whatever the document's style. It is tested in two parts, and both are registered as predictions that can fail.P2a— invariance. TheH2segment majority is the same arm under all three documents at ≥ 11 of 15 segments.P2b— non-tracking. Take every (segment, document-pair) case in which theGSstyle match differs between the two documents. In those cases theH2majority does not differ with it: the proportion that co-move is < 0.5. The confound predicts ≈ 1.0. Exact two-sided binomial over the co-moving count.P3— the manipulation check.GSmust actually move the style match: theGStwo-seat-agreed match differs between at least one document pair at ≥ 8 of 15 segments, and the plain document is matched toR06at ≥ 11 of 15.
What each outcome licenses, written now so it cannot be chosen later:
P3 |
P2 |
what is written into wiki/goodness-senses.md |
|---|---|---|
| holds | both hold | limit 5 is defeated on this material: H2 does not follow the document's style, so the separation of the halves and the direction stand. Discharge D-20260804-16 condition 3, with the residual limit of §9.3 attached. |
| holds | P2b fails |
H2 tracks the document's style. The §5 direction is reported as confounded and is not written into the senses page as a direction; the discharge is not written; the motion the condition implies is opened for a later session. |
| holds | P2a fails, P2b holds |
H2 moves but not with the style: something else in the document is doing the work. Neither the discharge nor the withdrawal is written; the run reports the anomaly and the arm closes on the null. |
| fails | — | P2 is void, not null (F3). The run reports P1 alone and records that the confound remains untested. |
What a P2b failure would and would not establish. It would establish that H2 is
sensitive to the document's prose style. It would not establish that the H2P result of
RS-20260812f was style-driven rather than comparability-driven, because under a plain document
the style match and the comparability answer point the same way and this design cannot separate
them. That asymmetry is stated here, before the numbers, and travels with any citation.
5. Gates, run before the primary is read
GA— arm parity. One seat per segment: do the two arms state the same facts? Wording, register, vocabulary, grammar and word order are explicitly not differences of fact. Note (bmh) applies: a flag is read before it is counted.GY— yardstick content parity. One seat per segment per rewrite (YPvsYM,YPvsYO), shown unlabelled: do these two say the same things about the passage, ignoring style entirely?GS— the style manipulation check. Two seats × 3 styles × 15 segments, shown one yardstick and both arms: which of the two English versions is written in a style more like this document's? The prompt mentions neither the source, nor effect, nor comparability.
6. Failure criteria, registered
F1—GAreturns a genuine fact difference at more than 3 of 15 segments → the arms are not parallel and the primary is void.F2—GYreturns a content difference for a rewrite at a segment → that (segment, document) drops from theP2pool. More than 5 of 15 dropping for a document voids that document's arm ofP2, and if both rewrites drop that farP2is void and the run reports the manipulation as unbuildable.F3—P3fails →P2is void and is reported as void, not as a null. A style manipulation that did not happen cannot test a style confound. Registered because it is the way this run could most easily fool itself.F4— a yardstick body that is empty, that leaks the CAUSE field, or that contains a Romanian-diacritic character or a source word → that segment's yardstick is void and the segment drops from everyH2condition that would have shown it.F5— more than 10% of the bodies at any judging stage come back empty orfinish_reason: lengthafter the note (bmb) remedy → that stage is void.F6— a seat refuses or returns an unparseable choice at a cell → the cell is excluded and counted; a stage losing more than 10 of its 45 cells is void.
7. Seats, caps, and the money
Roles from config/models.md; slugs are logged as provenance by the runner.
| role | seat | used for | cap |
|---|---|---|---|
| source-side, never judges | P2 google/gemini-3.6-flash, effort low |
YP, YM, YO |
800 |
| judge | P1 openai/gpt-5.6-terra |
H1, H2P, H2M, H2O, GA, GY, GS |
900 / 400 |
| judge | P3 x-ai/grok-4.5 |
H1, H2P, H2M, H2O |
900 |
| judge | P5 deepseek/deepseek-v4-pro, effort low |
H1, H2P, H2M, H2O, GS |
4000 / 2000 |
| pre-run critic | qwen/qwen3.7-max |
the critic passes | 16000 |
Note (bmb) is applied in the FIRST dispatch, not after it fires. Both reasoning-capable seats
have their effort pinned and their caps sized: P5 at 4,000 on the judging stages and 2,000 on
GS; P3 at 900, which RS-20260812f measured as sufficient on this slug. P4
moonshotai/kimi-k3 is not used: the 4,000-token cap the note prescribes prices its calls out
of any ceiling this run could declare, which NEXT.md carries as a standing fact.
360 calls. Pre-flight worst case is computed by run.py --dry-run from the exact prompts and
the caps, per note (abc), and recorded in §7.1 below before dispatch. P5's billed rate is
priced at the worst plausible provider, not the list rate (config/models.md, the 2026-07-25
routing caution).
7.1 Pre-flight, from --dry-run before dispatch
Worst case built from the caps the requests actually permit, per note (abc), not from an assumed output length.
| calls | 360 — yp 15, ym 15, yo 15, gy 30, ga 15, gs 90, h1 45, h2p 45, h2m 45, h2o 45 |
P1 openai/gpt-5.6-terra |
$0.620241 |
P2 google/gemini-3.6-flash |
$0.296984 |
P3 x-ai/grok-4.5 |
$0.384875 |
P5 deepseek/deepseek-v4-pro (list) |
$0.311454 |
| study worst case (list rates) | $1.613553 |
P5 at 4× routing (config/models.md caution) |
+$0.934361 |
| pre-run critic worst case (one pass) | $0.097800 |
| re-dispatch contingency (10% at 2×) | $0.161355 |
| TOTAL WORST CASE | $2.807070 |
| first critic pass, already billed | $0.050053 |
| declared ceiling | $3.00 |
Ceiling history, recorded rather than adjusted afterwards. $2.00 while the design was being
written; $2.20 when the first pre-flight came back at $2.041076 on 255 calls; $3.00 when
the pre-run critic's BLOCKING finding was accepted and the third document added, taking the
pre-flight to $2.807070 on 360 calls. Today's ledger stands at $1.319019290 of $5.00, so a $3.00
ceiling leaves $0.68 of the day's cap for any later session — the cost of taking the critic's
finding seriously, and it is declared here rather than discovered later. The alternative
considered and rejected was to cut P5's cap below the 4,000 that note (bmb)'s remedy
prescribes.
Lead translation of both arms is $0 and is not ledgered.
8. Verification
verify.py recomputes every reported number from the stored bodies, imports nothing from
analyse.py, and asserts the design's own invariants, among them:
- No
H1prompt contains a yardstick. - No
H2PorH2Mprompt contains thecausefield of any yardstick. - No judge prompt contains a Cyrillic or Romanian-diacritic character, and none contains the
arm labels
R06orR08. - For every (segment, seat), the
H2P,H2MandH2Oprompts differ only in the yardstick block — the arms, their order and the question wording are byte-identical. - Each arm's text sits behind the label the seed says it should.
GSprompts mention neither the source nor effect nor comparability.- Every reported count is recomputed from
runs/*.json, and the binomial probabilities are recomputed by exhaustive enumeration rather than by calling a library.
Mutation tests: the verifier is run against deliberately corrupted copies of the record and must catch each corruption.
tools/metric_a.py's clopper_pearson was repaired at S172 (RS-20260813a §5, note (bmy));
this run's proportions are at n = 15 and n = 45, where the repaired path is the one exercised.
9. What this run cannot establish
Written before the numbers exist.
- Model seats, not readers. Tier D is NOT PASSED. A finding here is a finding about how three seats behave, not about human readers of English or of Romanian.
- One more work, one more pair, one hand.
R06andR08are the same translator in one session in that order, soR08is downstream ofR06. - A
P2that holds defeats the style reading of limit 5 and nothing else. A yardstick is still a document, and having a document at all — as againstH1's having none — remains a difference between the two halves that this design does not remove. It cannot: a comparability question with no description of the original in it is not a comparability question. GSmeasures what seats call stylistic similarity, which is itself a model judgment and is not anchored.- 15 segments. A one-sided 11-of-15 is P = 0.118 by the two-sided sign test; the run is powered to detect a strong effect and not a moderate one, and the registered bars are set where they are for that reason.
- The
YMdocument is deliberately the strongest form of the manipulation, and the pre-run critic is right that its instruction names featuresR08's rule set also has. That is whyP2is a tracking prediction across three documents rather than a survival test under one, and whyYOis in the run. A citation ofP2carries §4's asymmetry paragraph.
10. Amendment, before dispatch: the pre-run critic's BLOCKING finding, accepted
The independent pre-run critic (qwen/qwen3.7-max, non-panel, runs/critic.txt) returned
NEEDS-REDESIGN with one BLOCKING finding, and it is right.
Its finding: YM's instruction — dislocated word order, archaism, unidiomatic literalness,
unrepaired fragments, "foreign-sounding, deliberately not fluent" — is a description of what
R08's rule set produces. So H2M would cue R08 by stylistic kinship, P2 as originally
written would be near-impossible to pass, and the design's result→option map turned that
engineered failure into the conclusion "limit 5 is vindicated, withdraw §5's direction". In its
words: it "proves only the trivial fact that if you give judges a yardstick that sounds like R08,
they pick R08."
Accepted in full, and the design was amended in three places before any study call was dispatched:
- A third document,
YO— the critic's requested control: marked relative to plain English, fluent, and not marked inR08's direction. 105 calls and about $0.9 of worst case. P2rewritten as a tracking prediction (P2ainvariance,P2bnon-tracking) rather than asR06-survival under one document. The confound hypothesis predicts co-movement between theGSstyle match and theH2choice whatever the document's style; that prediction is symmetric, is not satisfied by construction, and can fail in either direction.- The result→option map no longer converts a
P2failure into a withdrawal of §5. §4 now states, before the numbers, exactly what a failure would and would not establish.
The critic's own remedy — replace the estranged style with an ornate one — was not taken,
and the reason is on the record: an ornate document is still fluent English, so GS would
plausibly match it to R06 as well, the manipulation would not move, F3 would fire, and the
run would have no power against the confound at all. Keeping the strong manipulation and
adding the neutral one is what makes both the sensitivity and the tracking question answerable.
The amended design was put back to the same critic for a second pass before dispatch
(runs/critic2.txt): PROCEED-WITH-AMENDMENT, one MINOR finding — that F4 and
verification invariant 4 still named only H2P and H2M and had not been extended to H2O.
Both were already repaired in the working copy when the second pass was dispatched, and the
committed text carries the repair; the finding is recorded as discharged on arrival rather
than quietly dropped.
Critic cost: $0.050053 + $0.057559 = $0.107612, both inside the declared ceiling.