Repository path: workshop/experiments/E-20260821b-matched-heard/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260821b-matched-heard |
| status | frozen |
| created | 2026-08-21 |
| updated | 2026-08-21 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-matched-shape-heard.md, workshop/experiments/E-20260821b-matched-heard/critic-response.md, workshop/regimes/R40-matched-shape-cjk.md, workshop/regimes/R39-matched-shape.md, workshop/translations/maigan/R40-v1/translation.md, workshop/translations/kalila-fanza/R39-v1/translation.md, wiki/method-notes.md, config/budget.md, config/models.md |
E-20260821b — is a formal match built in English found by a reader who sees only the English
v2, frozen before any dispatch of the run. v1 was frozen at 347ad299, put to an independent
adversarial critic, and returned NEEDS-REDESIGN with 8 BLOCKING findings; v1 was never
dispatched. Every finding and its disposition — eleven accepted, one overruled on a stated ground
— is in critic-response.md. What follows is what will actually run.
1. Question
A translator carrying a source's matched shape into English builds a match out of declared resources. Does a reader who sees only the English find it?
v1 asked a second question — is the answer the same for all four resources? — and the critic's MAJOR 10 established that this design cannot answer it: resource (C) is twelve-thirteenths Chinese, so a resource difference and a source difference are the same difference. That question is struck from this experiment and left to the arm.
The estimand is prompted detection (critic MINOR 12): a seat is told to look for formal recurrence. Nothing here measures whether an unprompted reader would notice.
2. Why this can be asked now
R39 (S208) and R40 (this session) both require, at every locus where the hand records a match,
the plain wording it refused at that same span, written at the same sitting as the rendering.
Two whole renderings carry that record: T-kalila-fanza-R39-v1 (Arabic, 24 loci) and
T-maigan-R40-v1 (Chinese, 12 loci). The minimal pairs were frozen before this experiment existed,
for another purpose.
3. Which loci are eligible, and who decides
3.1 Excluded by the regimes themselves — 9
Four Arabic loci are self-flagged MATCHED — free (F57 F81 F91 F92): the refused wording is
identical to the device. Five more are excluded because their frozen refusal had to be bent to
stand inside the sentence's frame (critic BLOCKING 5), and a bent refusal is not a frozen refusal:
F59 F65 F89 (Arabic) and M9 M12 (Chinese). Each carries an adjust field in
materials/substitutions.json saying exactly what was changed.
22 loci go forward to the screen.
3.2 The eligibility screen — bought, not asserted
The critic's BLOCKING 2 and 3 are that many refusals still carry a match, and that the lead's own audit of which ones is inconsistent. So the lead no longer decides eligibility.
P3 x-ai/grok-4.5 — not P1, which wrote the critic pass — is shown, one locus at a time and
blind to everything else here, the members of a span and the sentence they stand in, and asked
whether those members echo each other in form. Both versions of all 22 loci are screened, in
shuffled order, with no indication that there are two versions of anything: 44 calls.
- A locus is ELIGIBLE for the primary iff the screener answers NO on its plain members.
- The screener's answer on the matched members is the manipulation check: a locus where it answers NO to both is a locus where the built match is not visible even when pointed at, and is reported as such.
- The lead's own audit is printed beside the screen in the result and is not used to exclude anything. It is a record of what the hand believed, offered so that hand and screen can be compared.
4. Materials
materials/substitutions.json— 27 loci: spans (each an exact substring of its frozen rendering, with the refusal at that span), annotated members in both versions, the declared resource, andadjustwhere a refusal was bent.materials/build.py— deterministic. Asserts each span occurs exactly once in its rendering and each annotated member lies inside its span; cuts the text into windows of ≥ 100 words, never closing a window while a locus still has a span to come (critic BLOCKING 4); writes both versions. 15 passages, 27 of 27 loci covered, none split, at most four eligible loci in any passage.- The plain arm is conservative. Matched wordings of excluded loci stay in both versions, so the plain passages still contain figures. That can only raise plain-arm recovery.
5. Seats
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5
(config/models.md). P4 and P5 are out by notes (bps) and (bne); GL is out on long
prompts. There is no fourth reading seat available to buy, and the result's limits must say so
rather than imply a panel. P3 screens eligibility, P2 screens parity, and all three read
passages; that a screener also reads is recorded as a limit.
No seat is told that this is a translation, that a source exists, that there are two versions of anything, or what the hypothesis is.
6. Procedure
6.1 (bqm)'s remedy
Note (bqm) fires at every say the same thing without the device design. Its remedy here is §3.2: the replacement is judged by an independent hand, and the lead's judgement of it is published but does no work. v1 proposed instead to buy a second hand to write replacements; the critic's BLOCKING 2 showed the problem is not who wrote them but whether they still carry the match, which is what the screen measures directly.
6.2 The parity screen — a hard gate
P2 receives 26 pairs: 22 real matched/plain span pairs and 4 with a planted content error
(a changed referent, number or polarity), shuffled, and answers SAME or DIFFERENT.
G1— if fewer than 3 of the 4 planted errors are caught, the primary is WITHHELD (critic BLOCKING 6). Not caveated: withheld.G2— length. The two versions' word counts are reported per passage; if any passage's two versions differ by more than 15%, the primary is withheld for that passage's loci.- Any real pair judged
DIFFERENTis excluded from the primary and reported with the reason.
6.3 The main run
15 passages × 2 versions × 3 seats = 90 bodies. One passage-version per call, temperature 0,
order shuffled on a seed fixed in run.py, output cap 900 (S208's costing finding: raise the
cap rather than lean on the doubled-cap retry), one re-dispatch at the doubled cap per note
(bgk).
Each dispatch is an independent completion with no conversation history, which is the ground on which critic BLOCKING 1 is overruled: a seat cannot recall its answer to the other version because it is never shown it. The statistical dependence that remains is handled in §8.
The seat reads the passage and lists every place where two or more stretches are built to echo each
other in form, quoting the stretches and naming what they share. At most twelve groups
(raised from six on critic MAJOR 11); truncation is counted and reported. The exact prompt is
materials/seat_prompt.txt; the word translation does not appear in it.
7. Scoring — frozen, and against annotated members
Normalisation: lowercase, strip all but [a-z0-9 ], collapse whitespace. Stopwords, frozen
here: a an and the of to in on at for from by with is are was were be been it its his her their
them they he she i you we not no nor that this these those as so but or if than then there which
who whom whose what when where how all any each every both other another such same own very can
could may might must shall should will would do does did done have has had.
A returned string m matches annotated member k iff, after normalisation, m ⊆ k or k ⊆ m, and m and k share at least one content token (not in the stopword list).
A locus is RECOVERED by a seat in a version iff one returned group contains ≥ 2 strings that match ≥ 2 distinct annotated members of that locus in that version.
The critic's counter-example — a group ["the", "him"] scoring RECOVERED on F18 — is a mutation
test in verify.py and must score NOT RECOVERED.
The free-text property is reported verbatim and classified descriptively only. No prediction
rests on it (critic MAJOR 9).
8. Predictions, registered
| prediction | |
|---|---|
PR1 |
Matched-arm recovery exceeds plain-arm recovery by at least 0.30 in locus-level proportion of seat cells, over the eligible loci. |
PR2 |
The paired difference is significant by a permutation test with the version label permuted within passage — the unit that was assigned — one-sided, 20,000 permutations, seed fixed in analyse.py, α = 0.05. |
PR3 |
Matched-arm recovery is at least 0.60 of eligible locus×seat cells. A direction result on tiny absolute numbers is a different finding and may not be reported as this one. |
PR3′ |
Plain-arm recovery is at most 0.50. Registered because a plain arm near the matched arm makes a significant difference practically empty (critic BLOCKING 8). |
PR4 |
Plain-arm recovery is not zero. English prose carries incidental parallelism; an exact zero would suggest the seats are answering the wording's oddity rather than reading for form. |
v1's PR4 (resource comparison) and PR5 (property classification) are withdrawn on critic
MAJOR 10 and MAJOR 9.
The lead's expectations, recorded so that they can be wrong: PR1 holds, PR2 holds, PR3
holds, PR3′ fails — the lead expects the plain arm to run high, because the frames that put
members in matched position survive every substitution — and PR4 holds.
9. Failure criteria
F1.G1fails → primary withheld (§6.2).F2′. Fewer than 10 eligible loci survive §3.2 and §6.2 → no pooled primary is reported. The per-locus table is published and the arm's question is answeredNOT ANSWERED ON THIS MATERIAL, with the reason: the subtraction could not be built.F3.PR3fails → the direction may be reported and no statement of the form the match is heard may be written intowiki/goodness-senses.mdorframework/v0.2; the absolute rate is written instead.F4.PR1's effect floor is not met → the primary is reported as not established, whateverPR2returns. Significance without magnitude is not a result here.F5. Any seat returning fewer than 40 of its 60 bodies parsed and non-dead is reported; no seat is dropped after the fact.- No resource-level claim may be written into the framework from this run, per §1.
10. Budget
Ceiling $1.20 for the whole experiment; stop-loss in each runner $1.00, counting the critic call already spent. Spent so far: $0.096713 (critic round 1). Remaining planned bodies: 44 eligibility + 26 parity + 90 main = 160, all short-prompt. Worst case at the caps the requests permit, with a doubled-cap retry on every one (note (abc)): ≈ $0.62, so ≈ $0.72 with the critic — inside the stop-loss.
Today's UTC day carries $2.097425 from S208 of $5.00; $2.902575 remains.
11a. Amendment, written after the screens and before the main run
Timestamp and honesty of this amendment. §3.2's eligibility screen and §6.2's parity screen were bought and read before this paragraph was written; no body of the main run had been dispatched. What follows changes what the main run is for. It registers one new prediction, and that prediction is post-screen, which is said here so that it can never be read as pre-registered.
F2′ has fired. Of the 22 loci screened, an independent seat says the refused plain wording
still echoes in form at 16; only six are eligible. Of those six, four are judged DIFFERENT in
content by the parity seat — which caught 4 of 4 planted errors, so G1 passes and the screen
is informative. Two loci survive both screens, against F2′'s floor of ten.
So the primary is WITHHELD, before the run, by the design's own rule. No pooled matched-versus- plain claim will be made from this material, whatever the run returns.
Why the main run is still dispatched. The withheld primary was the subtraction. The arm's first question — does a reader who sees only the English find the match? — does not need a subtraction: it is answered by the matched arm alone, as an absolute rate. And the plain arm now has a different job: if the screener is right that the refusals still echo, then seats reading whole plain passages should find those figures at rates near the matched arm's, which is a passage-level corroboration of a span-level screen rather than a control.
PR6, registered post-screen and pre-run: matched-arm recovery is ≥ 0.60 of locus×seat cells.PR7, registered post-screen and pre-run: plain-arm recovery is within 0.20 of the matched arm over the 16 loci the screener says still echo — because on the screener's evidence those passages still contain a figure at that place.- Everything in §7 (scoring) and §9 (
F3,F5) stands unchanged.PR1,PR2,PR3′andPR4are withheld outputs: they will be computed and printed for the record, and not reported as results.
11. What this cannot establish
- Nothing about whether either rendering is good. No seat is asked; the lead never judges his own translation (charter §5).
- Nothing about resources (§1).
- Nothing about unprompted reading (§1, the estimand).
- Nothing about human readers. Three language models are not readers of literature.
NOT CALIBRATEDstands (config/models.md). - Nothing about translators other than this one. Both renderings are the lead's, under regimes
the lead wrote, and the hand knew the arm's question while rendering the Chinese
(
T-maigan-R40-v1§3.4).