Repository path: workshop/experiments/E-20260725-tierD-ladder/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260725-tierD-ladder |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| links | config/models.md, config/budget.md, wiki/program.md, wiki/goodness-senses.md, wiki/base/sources/S-berman-tendances.md, workshop/translations/berman-tendances/R04-v1/translation.md, wiki/base/anchors/A-shaw-spider-thread/A-shaw-spider-thread.md, wiki/base/anchors/A-yosano-yomogiu/A-yosano-yomogiu.md, workshop/translations/genji-yomogiu/R04-v1/translation.md, workshop/experiments/E-20260725-tierD-ladder/critic.md, workshop/experiments/README.md |
Frozen design (v2) — Tier D detection calibration: a pilot, not a gate
v1 was rejected by the independent critic pass (critic.md, verdict NEEDS-REDESIGN, twenty-one blockers). This is the rewrite; 26 of 29 points are implemented and 3 are accepted as stated limitations, with dispositions tabulated in critic.md. A run may proceed only against this version.
No senses: field: this page designs an evaluation of the jury, not of a translation, and asserts no evaluative claim about any translation.
0. What this run can and cannot be — read first
This run cannot pass Tier D. Charter §5 and wiki/program.md Slate A item 4 make a held-out arm — reference against an independent same-quality translation, chance expected — a mandatory control. The critic established by measurement that the project's stored materials cannot supply one: the only candidate pair is Beowulf, where Gummere 1909 is alliterative verse (56 half-lines, 5 archaic tokens), Kirtlan 1913 is prose with 17 archaic tokens, and the lead's is plain modern prose with none. Verse against prose across a century of English is not a same-quality control; it can only produce an uninterpretable number.
v1 responded by quietly rewriting the control's success criterion inside the experiment page that benefited from it. That is exactly the move charter §2.8 forbids. v2 does not relax the control. It declares the control unsatisfiable from stored materials, drops the arm, and declines to claim a Tier D pass at all. config/models.md stays NOT CALIBRATED whatever these arms show, and finding a genuine same-quality pair becomes the named blocker on a real Tier D pass, recorded in NEXT.md.
What the run is: the first empirical characterisation of the instrument Tier D will need — the false-alarm rate against quality-neutral editing, the panel's actual use of a numeric scale, and detection and cross-sense specificity on four operators. Every finding is provisional.
1. Question
On which goodness senses can the non-Anthropic panel detect deliberate, sense-targeted degradation of a translation, and does the damaged sense fall further than the senses that were not damaged?
2. Materials
Two passages, two provenances, two source languages. All texts are public domain or lead-authored and already in the repository; nothing new is fetched and nothing new is stored.
| id | source | reference translation | provenance | source language |
|---|---|---|---|---|
| A | Akutagawa Ryūnosuke, 蜘蛛の糸 §一 (1918), 805 chars — workshop/canon/kumonoito/source.txt |
Glenn W. Shaw, "The Spider's Thread" §I (1930), 314 words — wiki/base/anchors/A-shaw-spider-thread/ |
published | modern Japanese |
| B | Murasaki Shikibu, 『源氏物語』「蓬生」§1‑3 (c. 1008), 494 chars — wiki/base/anchors/A-yosano-yomogiu/ |
the lead's own T-genji-yomogiu-R04-v1 §1‑3, 325 words |
lead | classical Japanese |
Word counts are measured, not estimated (critic A1 corrected v1's 318 and 380).
Why B carries the lead provenance. The design requires references of both provenances. Among the five available lead translations, the Genji §1‑3 passage is the only literary narrative prose into English whose source is stored beside it and whose source is not already committed to another role in this design. The Beowulf lead translation is excluded on a second ground the critic supplied: its own frozen log records its plainness as a deliberate choice to occupy one pole of that study's scale, which makes it an unrepresentative reference. (v1 gave a different and false reason — struck.)
What the jury is shown, and what was stripped. Each item is source + two unattributed English texts. Provenance headers naming Glenn W. Shaw, licence blocks, title lines, footnote markers and chapter headings are all removed by materials/build.py; the critic verified independently that none survives into the built files. Prose is unwrapped to one paragraph per line so that hard-wrap differences cannot fingerprint a text.
Not used, and why: the classical-Japanese caveat stands. B's source is harder for the jury to read than A's, so B's accuracy scores are expected to be noisier. Every metric is reported per passage as well as pooled.
3. Perturbation operators
Three of four come from Berman's tendances déformantes, read in French this session and frozen at T-berman-tendances-R04-v1 before this design existed (freeze commit 0aeeb93; the hash changed from 0111d38 when the branch was rebased to reset commit authorship — same tree, same ordering). The fourth does not, for a reason the frozen mapping establishes.
| op | catalogue source | what is done | target sense(s) | prediction on naturalness |
|---|---|---|---|---|
| O1 | Berman 5, l'appauvrissement qualitatif | content words replaced by same-meaning words lacking the original's sonorous / iconic richness | literary-quality |
flat |
| O2 | Berman 1 + 3, rationalisation, allongement | punctuation and clause order recomposed to canonical English discursive order; long periods split within their paragraph; verbs nominalised; connectives added | style-correspondence |
flat or up |
| O3 | Berman 10 + 11, destruction des réseaux vernaculaires / des locutions et idiotismes | source-marking and period-marked idiom removed; every edit preserves the referent | voice + cultural-mediation |
flat or up |
| O4 | the project's own catalogue — the documented Shaw Christianisation (A-shaw-spider-thread) and Swann (1974) on Turney's lapses |
semantic errors: wrong referent, wrong word sense, invented detail, dropped negation. Style untouched. | accuracy |
flat |
O4 is not from Berman, and that is a finding. The frozen mapping records that none of the twelve tendencies targets accuracy — Berman's quarrel is with the translator who writes well. A Berman-only operator set cannot test the sense a jury is most likely to detect.
Four operators, not twelve. The frozen log records that Berman's twelve collapse to roughly five independent operations. Tendencies 8, 9 and 12 are excluded on a stated ground: a signifier network, a systematism and a superimposition of languages cannot be shown, let alone dosed, inside 320 words.
O3 was rebuilt after the critic pass. v1's O3 silently injected accuracy damage (an omission of "pistils and stamens", 八月 rendered "September", 総角 rendered "local kids", a stereoscope becoming a window) and crashed register ("Hold on now… it's got a soul too", which is Berman's vulgarisation, tendency 4, not 10/11). v2's O3 keeps every referent and stays inside the reference's own register; cultural-mediation is now scored, matching the frozen mapping.
Dose. One dose, 8 edit sites per item. v1's light dose is dropped: the light sets were nested subsets that retained the loudest edit, so they could not have measured a dose effect (critic D24), and the budget could not have afforded them anyway (E26). The dose axis is deferred explicitly, in NEXT.md, rather than by arithmetic accident.
Every edit site is logged verbatim in materials/perturbations.md, committed before the run, with the items.json SHA-256 recorded in the log and in every raw response file.
4. Arms and stages
10 items. The jury never learns that anything was modified, which text is the reference, or who produced either.
| stage | arm | items | purpose |
|---|---|---|---|
| 1 | sham | A-sham, B-sham | dispatched first, standalone. 8 substitutions per item, the same count as a targeted arm, chosen as free variation. Measures (a) the false-alarm rate against quality-neutral editing and (b) — the reason it runs first — how these three jurors actually use a 1–7 scale, which the project has never measured. |
| 2 | targeted | A×{O1,O2,O3,O4}, B×{O1,O2,O3,O4} = 8 | detection and cross-sense specificity |
Order swap is mandatory on every item: each pair is dispatched twice, slots swapped. S010 measured slot preference at 0.63–0.85; an unswapped run is uninterpretable.
Stage 2 runs only if stage 1 clears two pre-registered conditions: 1. Scale usage. Pooled score standard deviation ≥ 0.75 and at least 4 distinct integers used across the six senses. If the jurors compress into two or three adjacent values, a 0.75-point specificity margin is not measurable and stage 2 is not dispatched. 2. Budget. Stage 1's actual billed cost leaves the stage-2 worst case (§8) inside the day's headroom.
On the sham's neutrality — conceded, not assumed. The critic's structural point is right: the reference is a considered text, so any eight substitutions tend to move off a local optimum, and a perfectly neutral sham is not constructible. v1's sham was worse than that — measurably archaizing (upon ×2, every one, about the grounds) and in one place sense-shifting (glanced ≠ looked). Those are removed. What remains is free variation (in bloom→blooming, which→that, until→till, remembered→recalled, stayed on→remained), word- and paragraph- and sentence-count neutral by measurement. The sham is therefore reported as an upper bound on the false-alarm rate, and its direction is a result rather than an assumption. It no longer holds a global veto (v1's [0.30, 0.70] band voided a perfectly neutral run ~15% of the time on 12 votes).
5. Jury and procedure
- Jurors: P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P5deepseek/deepseek-v4-pro, resolved fromconfig/models.mdat run time and logged as provenance. P4 (kimi, measured $0.09089/call in S014 — 6.0× P1, 9.0× P5) and P3 are excluded on cost. Both exclusions are power limitations, not findings about those models, and they mean the US-taste-correlation trigger inconfig/models.mdcannot fire on this run. - Panel currency was verified free against the OpenRouter model list: all five slugs live at unchanged prices, and no panel lab has shipped anything newer than the 2026‑07‑23 selection (
gpt-5.6-terra2026‑07‑09,gemini-3.6-flash2026‑07‑21,grok-4.52026‑07‑08,kimi-k32026‑07‑16,deepseek-v4-pro2026‑04‑24). That reasoning and evidence is written intoconfig/models.mditself, not asserted here (critic F32). No competence claim is made from it — only liveness, pricing and release recency were checked. - Six senses, scored 1–7 for each translation:
accuracy,naturalness,voice,style-correspondence,literary-quality,cultural-mediation. Plus a forced overall preference, no ties. Strict JSON. naturalnessis put to the jury with its register frame named, aswiki/goodness-senses.mdrequires and as v1 failed to do: "Reads as fluent, idiomatic English prose within its own evident register and period … do NOT penalise a text for being of an older or a more recent idiom than the other, and do not reward whichever is closer to present-day usage." S014 measured this sense as register-cued; without the frame the probe measures period preference.naturalnessis the specificity probe, not a neutral control. O2 and O3 make the prose more canonical and more domestic, so a jury that is reading rather than pattern-matching should hold naturalness flat while dropping the target sense.- Judgment is not parallelized (charter §6): strictly sequential dispatch.
max_tokens10,000 (S014's value; S010's lesson is that reasoning models need headroom or they truncate into parse failures). One retry per payload. A discarded attempt is written to<path>.attemptN.jsonand its cost is ledgered — truncation is billed whether or not it is used.- Raw request/response JSON preserved for every call, with the payload SHA-256 and the items-manifest SHA-256 in each file.
6. Pre-registered firing rule
The unit of analysis is (juror × passage), with the two orderings averaged within the unit — not the individual vote. v1 pooled 12 votes as if independent; the calibration-v1 critic raised that, it was accepted, S014 implemented it, and v1 reverted it (critic B7).
For each operator there are 6 units (3 jurors × 2 passages). A unit is scored:
- +1 — the reference is preferred in both orderings
- −1 — the variant is preferred in both
- 0 — split
Under a null of independent coin-flip preferences, P(+1) = 0.25.
Detection fires for operator O iff at least 5 of 6 units are +1 and none is −1. Exact one-sided P = 0.00464; across four operator cells the family-wise rate is ≈ 0.018. (v1's rule — 9 of 12 pooled votes — had P = 0.073 per cell and a 0.365 chance of at least one cell firing by accident.)
Specificity fires for O iff both:
1. the mean drop on O's target sense(s) exceeds the mean drop across the non-target senses excluding naturalness by ≥ 0.75 scale points; and
2. the largest per-sense drop, excluding naturalness, is one of O's targets.
naturalness is excluded from the baseline because the design predicts it will rise for O2 and O3, which would mechanically inflate the very margin it is meant to be independent of (critic B9). Both the excluding and the including figures are computed and reported; only the excluding one fires the rule. Condition 2 is scale-free and carries the actual content of "specificity"; condition 1 is calibrated against the scale usage stage 1 measures.
The naturalness probe has a rule and a consequence (v1 stated a prediction that could not fail — the third occurrence of that defect in this project): for O2 and O3, drop(naturalness) must be ≤ 0.25 scale points. If it exceeds that, the pre-registered reading is "the jury lowers every sense together", and no specificity claim is made for that operator regardless of its margin.
The sham's three-way interpretation rule, on its own 6 units: - ≥5 of 6 units +1 → the jury prefers unedited text as such; every detection result in stage 2 is confounded with edit-presence and no detection claim is licensed from this run. - ≤1 of 6 units +1 → the jury prefers edited text; same conclusion, opposite sign. - otherwise → the false-alarm rate is acceptable at this resolution, and stage 2's detection results stand on their own terms.
Gate status, pre-committed for every outcome (critic F34): Tier D remains NOT PASSED in all cases, because the held-out control is absent (§0). If detection fires and specificity fires on a sense, the recorded claim is "the jury detects S-damage at 8 sites, pending a held-out control". If detection fires and specificity fails, the claim is "detects damage, not sense-calibrated on S". If detection fails, the claim is "no detection at this dose and this power".
7. Predictions, written before the run
accuracy(O4) fires detection. Most likely of the four.literary-quality(O1) shows the weakest detection — the edits preserve meaning and grammar.- Specificity fails on at least two of the four operators: the jury says "worse" broadly. This is the modal outcome in the lead's estimate and is the thing the old Tier P design could never measure.
internal-judgment-only. - Detection is weaker on passage B than on A, because the jury reads classical Japanese less well.
- The sham lands in the middle band. The critic predicts the opposite — that the reference wins the sham arm outright, because any edit of a considered text drifts downward. Both predictions are on the record; the sham arm settles it.
naturalnessdoes not fall by more than 0.25 for O2 and O3.- The pooled scale SD exceeds 0.75 and at least four distinct integers are used.
8. Budget
Headroom, UTC 2026‑07‑25: $1.600774 ($5.00 − $3.399226). Key-usage snapshot at session start: 18.290925792.
The estimate is stated as a range with a worst case, built from S014's measured per-call figures on a per-sense scoring task — means for the central estimate, maxima for the worst case (critic E27; the only run in this project to land inside its estimate is the one that did this).
| stage | items | payloads | calls | central (means, $0.04146/payload) | worst case (maxima, $0.06467/payload) |
|---|---|---|---|---|---|
| 1 — sham | 2 | 4 | 12 | $0.166 | $0.259 |
| 2 — targeted | 8 | 16 | 48 | $0.663 | $1.035 |
| total | 10 | 20 | 60 | $0.829 | $1.293 |
Worst case $1.293 < headroom $1.601, with $0.31 (19%) of margin. Hard abort at $1.40, checked before every dispatch using the worst call seen so far as the margin. Because the sham is stage 1 and standalone, an abort can only ever truncate targeted arms — never the control that gates them (critic E28: v1's dispatch order would have dropped the sham first).
9. Failure criteria — what voids or qualifies the run
- >10% of calls failing or unparseable after one retry → reported as a failed run, not patched.
- Stage 1 scale usage below the §4 threshold → stage 2 is not dispatched, and the run reports the scale finding alone. This is a legitimate and cheap outcome.
- Sham outside the middle band → §6's interpretation rule applies; detection claims are withdrawn, and the sham result is itself reported.
- Post-run verification failing to reproduce a reported number → the number is withdrawn, not corrected in place.
- No order-flip threshold. v1 set one at 0.50, just above P5's measured 0.444 — a rule its most unstable juror could not trigger, on an estimate whose own 95% interval is about ±0.25 at this n. Replaced by a pre-committed P5-excluded sensitivity analysis, reported alongside the primary, per the disposition already accepted in
E-20260725-anchor-verification(C2).
10. Verification
tools/verify_tierD.py recomputes every reported number from runs/ and the items manifest only, re-parsing each model's own text rather than trusting the runner's cached parse. It also checks the perturbation log by reconstruction — rebuilding each variant from its reference plus the logged edits and asserting byte equality, which catches edits that are present but unlogged. v1's check only tested the other direction and, the critic showed, missed exactly that: three edits in the built files were absent from the log at the time of review.
11. Known threats, stated in advance
- No held-out control (§0). This is the reason the run cannot pass Tier D, and it is the largest single weakness.
- The sham is lead-authored, like the perturbations. It controls for edit-presence, not for lead-shapedness. Both come from the same hand.
- O2 is identifiable by counting full stops. Measured: A 14 → 20 sentence-final periods; B 10 → 18. That is the operator, not an artefact — but it means an O2 "worse" verdict is compatible with a jury that never read either text, and the
style-correspondencespecificity claim cannot distinguish reading from counting. - O3 removes all transliterated Japanese from passage A (Sanzu-no-Kawa, Hari-no-Yama). Presence or absence of romaji is a one-glance discriminator. Intrinsic to tendency 10; named so a reader does not have to find it.
- Length is controlled. Measured: A variants 313–328 against a 314-word reference, B variants 316–324 against 325. Paragraph counts identical across every variant and its reference.
- Power. Six units per cell distinguishes near-ceiling from near-chance and nothing finer. No claim of the form "sense X is better detected than sense Y" is licensed unless the gap is the full width of the scale.
- Three jurors, two US labs and one Chinese lab. Narrower than panel v1 by design and by budget.
- Reference quality is uncontrolled. Shaw 1930 is not a perfect translation — S015 documented a Christianisation in this very passage. The jury is asked to prefer a less damaged text, not a good one.
- Detection ≠ judgment. Firing on every sense here would still not license ranking two good translations. That is Tier P, which ran once and failed.