Repository path: workshop/experiments/E-20260810d-carriage-naming/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810d-carriage-naming |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| track | T3 |
| senses | perceived-source-carriage |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-rule-execution.md, wiki/findings/results/RS-20260809h-rule-execution.md, wiki/findings/results/RS-20260808d-carriage-decoupled.md, wiki/findings/results/RS-20260802f-licensed-strangeness.md, wiki/goodness-senses.md, workshop/translations/kronenwaechter-heilung/R06-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, wiki/method-notes.md |
E-20260810d — can a reader holding the original say what a foreignizing departure is carrying?
Frozen before any generation or scoring call. The R06 translation this run's materials are built
out of was frozen first, at commit e077fef, before this file existed.
1. The question, and why it is about translation and not about the instrument
Foreignizing translation is defended, by Schleiermacher, by Berman and by Venuti, on the ground that it makes the foreign text's difference visible to a reader who could not otherwise see it. The claim is about a reader: at the place where the English goes strange, the reader is supposed to be able to see the original showing through.
Nothing in this project has ever asked a reader to point. Three runs have asked readers to
rate — to score perceived-source-carriage, the sense created at S093 to record when a marked
departure reads as a deliberate attempt to carry over a feature of the source — and all three found
the same thing:
| run | material | markedness carrying nothing | markedness carrying something |
|---|---|---|---|
RS-20260807f (S130) |
Schwob, FR→EN | unlicensed oddity +1.861 | carried form +1.750 |
RS-20260808d (S135) |
Bang, DA→EN | gratuitous oddity +2.286 | sixteen carried devices +1.810 |
RS-20260809h (S146) |
Arnim, DE→EN | word-scramble +1.4583 | six hands' R08 arms, median +0.9583 |
The third is the sharpest, because CLUNKY was a strict word-multiset permutation of a clean
translation — proved so by the verifier, not one word added or removed — and it out-scored five of
six foreignizing translations and the lead's own. RS-20260809h's limits section concluded that the
sense "does not separate foreignizing translation from scrambled word order" and barred it as an
uptake measure without a matched oddity control.
A prohibition is not an explanation, and the prohibition is all ARM-rule-execution step 2 was
scoped to write down. Three rating runs agreeing tell us the rating does not separate carriage
from non-carriage. They do not tell us whether the separation is unavailable to the reader or
only unavailable to the rating. Those are different findings about literary translation:
- INVISIBLE — a reader holding the original cannot tell a departure that carries the German from
a departure that carries nothing. Then the visibility claim that motivates the whole foreignizing
programme fails at the level of the individual departure, and
framework/v0.2's recommendations that reach for source-orientation are recommending something no reader can locate. - UNRATED — the reader can tell, and the seven-point scale cannot. Then the three results above are about an instrument, the visibility claim is untouched, and what the framework needs is a different question put to the reader, not a different recommendation.
The unit's sentence (subject rule, continue-prompt.md §4.5): this unit asks whether a reader
with the original in front of them can say what a foreignizing English departure is carrying, and
whether they say it just as readily when it is carrying nothing. That is a claim about reading
translations. The write-up into wiki/goodness-senses.md that ARM-rule-execution §Done when
clause 2 requires is downstream of the answer, and is the reason the step was extended from writing
to measurement-and-writing (arm log, S151).
What this run changes about the question put to the reader. Every prior run asked for a rating on a scale. This one asks the reader to point at the German — to quote the exact words whose feature the English departure carries, or to answer NONE. A rating is a permission; a quotation is a commitment, and it can be wrong in a way a rating cannot.
What the run can and cannot decide — narrowed after the pre-run critic (amendment A11). The
scored output is whether the reader attributes the departure to the source and points at German,
not whether the feature they name is the right one. So the estimand is attribution: does a
reader holding the original attribute a marked English departure to the German more readily when
the departure really is carrying something than when it is carrying nothing? A null on that
question is enough to establish INVISIBLE-at-the-site, because a reader who attributes equally to
licensed and unlicensed markedness is not seeing the source through the English. A positive is
not enough to establish that the reader saw the right feature; that would need the FEATURE
strings coded, and this design codes nothing by hand. The distinction is kept in every sentence of
§5 and in the result page.
2. Materials
Source. Arnim, «Die Kronenwächter», Erster Band, Zweites Buch, Erste Geschichte: Die
wunderbare Heilung (1817) — the chapter's opening paragraph entire, 604 German words in ten
segments. Copy-text Projekt Gutenberg's Reclam digitisation, collated against the 1857 archive.org
Fraktur scan (collate.py): 10 of 10 sentences present, character ratio min 0.8896 / median
0.9358 / max 0.9622. Two copy-text errors were corrected against the witness, one of them a
sentence boundary that makes the span ten segments rather than nine
(../../translations/kronenwaechter-heilung/R06-v1/translation.md).
Why this source. It is the only German prose this project has established as having no reachable
published English translation at all (S141 §2, S146 §2, re-searched this session), which removes
recall entirely: no seat can be scoring a remembered rendering. It holds the language pair, the
author, the work and the prose type constant against S141 and S146, and the spans do not overlap by
a word. Contamination is declared none on the artifact with its basis stated — the absence of a
comparator, not a measured overlap, since dependence_check.py takes a pair and there is no second
member.
Why this paragraph. It was chosen, before any English was written, for the density of German
constructions a foreignizing translation would have something to carry: three verb-first
conditionals, a fronted dative object with its subject held to the end, correlative so … wie,
participles held behind a long middle field, a prepositional nominalization, and an impersonal
passive with no subject at all.
The baseline text, P. T-kronenwaechter-heilung-R06-v1 — the lead's plain rendering of the
whole paragraph under R06, one pass, source only, no rule set, frozen at e077fef with its
translator's log before this design existed. Cost $0; lead translation is never ledgered.
The ten sites and the four conditions
One site per segment, ten sites. At each site a span of the frozen P translation is marked
with ⟦ ⟧, and four English versions of the segment are built which are byte-identical outside the
marked span:
| condition | what the marked span is | true answer |
|---|---|---|
P |
the frozen R06 rendering, unaltered — ordinary English |
NONE |
L |
rewritten so that its strangeness reproduces a named formal feature the German has at that point | the registered German locus |
F |
rewritten so that its strangeness is a syntactic construction of English that the German does not have at that point | NONE |
S |
the words of the P span in a different order — a strict word-multiset permutation, machine-verified |
NONE |
L and F are both syntactic, deliberately. The obvious way to build unlicensed strangeness is
to reach for archaism — hath, thereof, soever — and that would have confounded licence with
kind of markedness: the reader could separate the arms by register alone without ever consulting
the German. Every F here is instead a construction English can produce and the German at that point
does not have: absolute participle, postposed adjective, comparative inversion, predicate fronting,
left-dislocation with a resumptive pronoun, cleft, emphatic do-support, agentive of, existential
there with heavy-NP shift, postposed quantifier, polysyndeton, periphrastic comparative.
The residual mismatch is named as a limit in §6 and not claimed away: L and F are matched on
being syntactic, not on being the same construction.
The pre-run critic found three sites where this had not been achieved and all three were repaired
before dispatch (§11): site 1's F was an absolute participial clause, which is a near-calque of
the German prepositional nominalization it was supposed to be a control for; site 6's L was a
lexical choice (dwelt, confidingly) rather than a syntactic carry; and site 2's L was a bare
preposition swap nameable from the lexicon without reading the German's syntax at all.
The registered site table is materials/sites.json, written before dispatch, and carries for
each site: the P span verbatim, the L, F and S spans, the exact German locus the L span
claims to carry, and a one-clause statement of the feature. It is reproduced in full in §7 of this
file so that the freeze is legible without reading the JSON.
S is generated, not written — build_items.py tokenises the P span, permutes with a fixed
seed, and verify.py asserts the case-folded word multiset is unchanged. This is the site-level
analogue of RS-20260809h's CLUNKY, and the same property is proved the same way.
3. The jury and the two passes
Three seats, the same three as E-20260808c, E-20260809b and E-20260809h, so that this run
is read beside them: J1 openai/gpt-5.6-terra, J2 google/gemini-3.6-flash, J3
deepseek/deepseek-v4-pro. All non-Anthropic (charter §5; the lead may not judge its own
translation, and this whole run is the lead's translation).
Pass A — the naming task. 10 sites × 4 conditions × 3 seats = 120 calls, one item per call, no
conversation history, temperature 0. Each call shows one German segment and one English version with
one span marked, and asks for QUOTE:/FEATURE: or NONE. The prompt string is frozen in
materials/prompts.json and is reproduced in §8.
Pass B — the markedness check. 10 sites × 3 seats = 30 calls. Each call shows the four English spans of one site and no German at all, labelled A–D in a per-seat registered order, and asks for a 0–6 rating of how far each departs from ordinary present-day English prose. The German is withheld because this is a target-side markedness judgment and showing the source would be exactly the contamination the check exists to rule out.
Blinding. Condition labels never appear in any prompt. Pass A items are shuffled per seat with a
registered seed, under the constraint that two items of the same site are at least six positions
apart (materials/orders.json, written before dispatch). Pass B's A–D order is shuffled per site
per seat. What this does not achieve is stated as a limit: a seat sees all four versions of a site
across its 40-item stream, and although the calls are independent and stateless, the four versions of
one site are near-identical strings and a stateful provider could in principle relate them.
4. Scoring — mechanical, and no lead judgment enters any number
A Pass A body is parsed as:
- NONE iff its first non-empty line, upper-cased and stripped of punctuation, is exactly
NONE; - NAME iff it contains a line beginning
QUOTE:; - anything else is MALFORMED and goes to
F5below.
NAME(c) = the count of NAME cells in condition c, over 30 cells (10 sites × 3 seats).
HIT is defined only at L and is mechanical: normalise the quoted German and the registered
locus (case-fold, fold umlauts, ß→ss, drop non-letters, split to word tokens), and score a HIT iff
the quote's token sequence is contained in the locus's, or the locus's is contained in the quote's, or
they share at least half of the locus's content tokens — and the quote is no longer than
60% of the segment's word tokens, so that a whole-segment dump is not scored as a hit. Both
thresholds are registered here and are not tuned after seeing bodies. HIT is a secondary
measure (amendment A6): the primaries are NAME contrasts, and HIT bears only on C1.
FEATURE: text is collected and is not scored. Coding whether a stated feature is the right
feature would be lead judgment on panel output, and this design admits none: the FEATURE strings are
reproduced verbatim in the result page as illustration and no statistic is computed from them.
The positional objection, and why it is the hypothesis rather than a confound. At every site the
same German words stand at the same place in all four conditions, so a seat that simply quotes the
German opposite the marked span will score HITs at L without consulting the feature at all. That
same strategy would also produce quotations at F, at S and at P, where the true answer is NONE
— which is exactly what NAME(F), NAME(S) and NAME(P) count. The design does not need to
exclude the strategy; it needs to measure how far it is used, and that is what it measures.
5. Gates, primaries, and what each outcome licenses
Everything below is registered before dispatch. Bars are not moved after bodies are seen.
Gates — restructured after the pre-run critic (amendments A5, A6, A14)
The critic's MAJOR 5 was accepted: the original gate stack withheld the primaries on failures that
are not failures of the primaries. P1 and P2 are contrasts between NAME rates. They do not
depend on the HIT matcher and they do not depend on NAME(P). Only P3, which is defined against
P, depends on P. The gates below withhold exactly what they bear on and nothing else.
| gate | statistic | bar | what failing it withholds |
|---|---|---|---|
C1 the seats can identify carriage |
HIT rate at L |
≥ 0.60 | only the claim that seats identify the right German when they attribute. P1, P2, P3 stand: they are attribution rates. |
C2 NONE is usable |
NAME(P) |
≤ 0.40 | P3 only. A high NAME(P) is reported as a finding in its own right — that these readers attribute source-carriage to ordinary faithful rendering. |
C3 markedness matched |
|mean M(L) − mean M(F)|, pooled and paired per site |
≤ 1.00 of 6 | P1 may not be read as a licence effect; it is reported as a difference with a markedness gap named. Non-gating, per critic MINOR 14. |
C4 the manipulation worked |
mean M(L) − mean M(P) and mean M(F) − mean M(P) |
each ≥ 1.00 of 6 | if M(F) − M(P) fails, P1's null is uninterpretable — an unmarked F cannot test confabulation. This is the one gate that can void the headline and it is stated so here. |
F5 cell return |
MALFORMED or missing bodies | retry once; then drop the cell and reduce n before any P is computed | reported with the count |
F2 scale usage |
distinct integers used per seat in Pass B | reported only | — |
Primaries
P1— does attribution track licence?NAME(L) − NAME(F), computed per site over its three seats and tested across the ten sites by an exact two-sided sign-flip permutation test (2¹⁰ = 1,024 assignments; minimum attainable two-sided P = 0.00195).- ≥ +0.40 with P ≤ 0.05 ⇒ ATTRIBUTION TRACKS LICENCE. Readers holding the source attribute carried markedness to it more often than free markedness; the three rating results are then about the rating, and the sense needs a better question rather than retirement.
- anything else ⇒ NOT SHOWN TO TRACK LICENCE, with the point estimate and 90% interval printed.
- No equivalence claim is registered and none may be made — amendment A4, accepting the critic's BLOCKING 4. Ten sites of three binary cells cannot certify a ±0.20 band, and the design originally registered one. A null here is failure to demonstrate, and the result page must say so in those words.
P2— does mangling draw attributions?NAME(S). ≥ 0.50 ⇒ a site-span word-scramble draws source attributions from a majority of cells. Amendment A8, accepting the critic's MAJOR 8: this is consistent withRS-20260809h'sCLUNKYresult and does not establish its mechanism — that run scrambled a whole segment, this one scrambles a span of five to fifteen words, and the two manipulations are not the same size.P3— does strangeness as such raise attribution?NAME(F) − NAME(P), same paired test. A positiveP3with a nullP1is the INVISIBLE reading in its strongest form: strangeness buys attribution and licence buys nothing on top of it.
Registered secondary (amendment A9, critic MAJOR 9)
Three of the ten L spans (sites 4, 5, 7) restore a verb-first German construction. P1 is
therefore recomputed on those three sites and on the other seven separately, and a positive P1
carried only by the verb-first subgroup is reported as "verb-first markedness is nameable", not as
"licence is visible". Registered here, before dispatch, so that the subgroup split cannot be
chosen after seeing which way it falls.
What would refute the expectation
NAME(L) high, NAME(F) and NAME(S) and NAME(P) all low is a clean UNRATED verdict: the
reader can point, the scale cannot, the three prior results are instrument findings, and this design's
motivating worry is wrong. It is a real possible outcome, it is written here before dispatch, and it
would be reported as the headline with the same prominence as any other.
6. Known limits, written before the data exists
- Three seats, all LLMs, and their German competence is asserted by nobody.
C1is the only evidence in the run that they can read the German at all, and it is a floor, not a certificate. Every figure here is panel-perceived andinternal-judgment-only; Tier D is NOT PASSED and no jury verdict carries evidential weight (charter §2.4). - One work, one language pair, one translator, one paragraph, ten sites. The
Lspans are the lead's judgments about what the German has, and a different translator would have chosen different features. Nothing here generalises to other pairs without being run on them. LandFare matched on being syntactic and not on being the same construction. If theNAMErates differ, the difference could be licence or could be construction; §7's table is printed so a later reader can see exactly what was compared, and the verb-first subgroup is pre-registered in §5 for exactly this reason. 3a. Content parity is declared, not certified. EveryLandFspan says what itsPspan says — the rewrites are syntactic and the propositional content is held — andSis word-identical toPby construction. No independent content-parity call was made, asRS-20260808dmade for its sixteen devices; the parity here rests on inspection and on the minimal-pair construction, and a reader who doubts a row can read it in §7.Fis unlicensed at the site, not in the segment. AFconstruction could accidentally answer to something elsewhere in the same German sentence; the registered claim is only that the German at that point does not have it.- The four versions of a site are near-identical strings seen by one seat across its stream (§3).
P's true answer is NONE by construction and by intent — but a determined reader can say that any faithful rendering "carries" its source. The prompt is worded to ask about departure from ordinary English, andC2is the check that the wording worked. IfC2fails, that is the finding about the wording and it is reported as one.
7. The registered site table
Frozen here and in materials/sites.json. L locus is the exact German substring the L span
claims to carry.
| # | seg | P span (frozen R06) |
L span |
L locus (German) |
German feature carried | F span |
F construction (English, unlicensed here) |
|---|---|---|---|---|---|---|---|
| 1 | S1 | having reached a certain height | after the attaining of a certain height | nach dem Erreichen einer gewissen Höhe | prepositional nominalization of a verb | after it does reach a certain height | emphatic do-support |
| 2 | S2 | it is harder to find | harder is it to find | schwerer ist's | fronted comparative with subject-verb inversion | it is more hard to find | periphrastic comparative where English takes the synthetic |
| 3 | S3 | has altered as completely as our commerce | has so entirely altered, as our commerce | so gänzlich verändert, wie unser Verkehr | correlative so ... wie across a comma | has altered as completely as has our commerce | comparative inversion |
| 4 | S4 | If one part of Christendom has been ashamed of art in churches | Has one part of Christendom been ashamed of art in churches | Hat ein Teil der Christen sich der Kunst in Kirchen geschämt | verb-first conditional with no conjunction | If ashamed of art in churches one part of Christendom has been | predicate fronting |
| 5 | S5 | the everyday and the Sunday, house and church, must be formed out of one piece | must the everyday and the Sunday, must house and church be formed out of one piece | muß Alltägliches und Sonntägliches, muß Haus und Kirche | the finite verb repeated at the head of each member | the everyday and the Sunday, house and church, they must be formed out of one piece | left-dislocation with a resumptive pronoun |
| 6 | S6 | he looked at his own people in a heightened sense | he looked his own people in a heightened sense on | er sah die Seinen im erhöhten Sinn an | separable verb with its prefix held to the end of the clause | it was at his own people that he looked in a heightened sense | cleft |
| 7 | S7 | for though Luther lived, by every account, plainly and without splendour | for lived Luther, by every account, plainly and without splendour | denn lebte gleich Luther nach allen Nachrichten prachtlos und einfach | verb-first concessive with no conjunction | for though it was plainly and without splendour that Luther lived, by every account | cleft |
| 8 | S8 | still more richly fertilized, visited and adorned by artists of every kind | still more richly by artists of every kind fertilized, visited and adorned | von jeder Art Künstlern befruchtet, besucht und geschmückt | participles held behind the agent phrase | still more richly fertilized, and visited, and adorned by artists of every kind | polysyndeton where the German coordinates asyndetically |
| 9 | S9 | Many of these foreigners were driven to Germany by affairs of state | Many of these foreigners, affairs of state drove to Germany | Manchen dieser Fremden trieben Staatsverhältnisse nach Deutschland | fronted object with the subject held past the verb | There were driven to Germany by affairs of state many of these foreigners | existential there with heavy-NP shift |
| 10 | S10 | because God had bestowed on it these and no others | because they were bestowed on her by God not otherwise | weil sie ihr von Gott nicht anders beschert waren | passive with a feminine dative pronoun for the town, where English takes an active with it | because God had bestowed on it these and others none | postposed quantifier |
Three of the ten L spans (4, 5, 7) restore the German verb-first construction that the frozen
translator's log records losing three times over (log items 7, 13, 23) — the paragraph's signature
construction, and the reason §5 pre-registers a verb-first subgroup contrast. Two more (9, 10) are
the literal alternatives the log records as live and not taken (items 19, 22). The L arm is
therefore not an invention for the experiment: it is the road the translation did not take, written
down at the time and frozen before this design existed.
The S spans, generated and printed here so the freeze is complete:
- site 1 —
height certain reached having a - site 2 —
find is harder it to - site 3 —
as as altered our completely commerce has - site 4 —
has in art one been churches part of ashamed of Christendom If - site 5 —
of everyday piece and formed and be out Sunday, house the the one must church, - site 6 —
heightened people sense a he his own looked in at - site 7 —
account, Luther though every lived, splendour and plainly without by for - site 8 —
artists adorned and kind still by more richly of visited fertilized, every - site 9 —
affairs driven Germany were of these Many of by to state foreigners - site 10 —
bestowed because had no and it God these on others
8. The two prompt strings (frozen)
Pass A, one item per call:
You are reading an English translation of a German literary text.
Below is one sentence of the German original, and one sentence of an English
translation of it. In the English, one span is marked with ⟦ ⟧.
Your task concerns the marked span only.
The marked span may be a deliberate departure from ordinary English, made in
order to carry over into English a specific feature of the German — its word
order, its word-formation, its grammatical construction, its register, or its
figure. Or it may not be: it may be ordinary English, or it may be strange for
some reason of its own that answers to nothing in the German.
Rendering the German faithfully is not by itself such a departure. Every English
sentence here translates the German sentence beside it; the question is only
whether the marked span departs from ordinary English IN ORDER TO carry a
particular feature of the German across. If it does not, the answer is NONE.
Decide which, and answer in exactly one of these two forms, adding nothing else.
If the marked span carries over a specific feature of the German:
QUOTE: <the exact German words, copied from the German below, whose feature it carries>
FEATURE: <one clause naming the property of those German words that the marked span reproduces>
If it does not — if the marked span is ordinary English, or if its strangeness
carries over no specific feature of the German:
NONE
=== GERMAN ===
{german}
=== ENGLISH ===
{english}
Pass B, one site per call:
Below are four English versions of the same passage. They differ only inside the
span marked ⟦ ⟧.
Rate each marked span for how far it departs from ordinary present-day English
prose, on this scale:
0 ordinary English, nothing unusual
1 very slightly unusual
2 noticeably unusual but unremarkable in literary prose
3 distinctly odd
4 strongly odd
5 very strongly odd; a reader would stumble
6 extremely strange; barely English
Judge the English only. Answer with exactly four lines and nothing else:
A: <0-6>
B: <0-6>
C: <0-6>
D: <0-6>
=== A ===
{a}
=== B ===
{b}
=== C ===
{c}
=== D ===
{d}
9. Pre-flight cost estimate
Built from max_tokens, not from an expected output length — note (abc), the one estimate this
project has overrun failed exactly there.
| pass | calls | input worst case | max_tokens |
|---|---|---|---|
| A | 120 (40 per seat) | ~700 tok | 300 |
| B | 30 (10 per seat) | ~1,100 tok | 120 |
Per seat: ~39k input, ~13.2k output at the cap. At the config/models.md list prices —
J1 $1.00/$6.00, J2 $1.50/$7.50, J3 $0.435/$0.87 — that is $0.114 + $0.158 + $0.029 = $0.301.
The declared worst case is $2.40, which is the list-price figure multiplied by 4× for provider
routing (config/models.md's S022 caution: a slug's billed price has been 3.8× its listed one) and
rounded up, plus $0.20 for one pre-run critic call at a 16,000 cap. Reasoning is suppressed where
the provider accepts it (note (bkw)); a body that returns length with zero visible content is
re-dispatched once at a raised cap and the dead body is preserved under runs/discarded/.
Today's ledger (UTC 2026-08-10) stands at $0.582308840991 of $5.00 across three sessions, leaving
$4.417691 — so $2.40 fits with room. Actuals are read from each response's usage.cost and
cross-checked against the key-usage delta.
10. Verification
verify.py recomputes every reported number by a path that imports nothing from analyse.py, and
asserts: every German segment byte-identical to the frozen source file; every P span an exact
substring of the frozen R06 segments; every L/F/S variant identical to P outside the marked
span; every S span a strict case-folded word-multiset permutation of its P span; both prompt
strings identical to materials/prompts.json; the item shuffle obeying the six-position separation
constraint; body counts with runs/discarded/ excluded; a deterministic re-run reproducing
results.json; and every mean and every exact P recomputed independently. Mutation testing: at
least three seeded corruptions must each move a named figure, and at least one order-only permutation
must move nothing.
Per the standing rule this is code and materials integrity, not scientific validity.
11. The pre-run critic, and what was changed before dispatch
One adversarial pass over this frozen design and the frozen translation, by x-ai/grok-4.5 —
deliberately not one of the three judging seats. Verdict NEEDS-REDESIGN, 16 findings, 4
BLOCKING. Raw body at runs/critic.txt, prompt at materials/critic-prompt.txt, cost
$0.0356864. Thirteen findings accepted, two accepted in part, one overruled in writing.
Every change below was made before a single item was dispatched.
| # | severity | the finding | disposition |
|---|---|---|---|
| 1 | BLOCKING | site 1's F, an absolute participial clause, is a near-calque of the German prepositional nominalization it controls for — a control silently converted into a second treatment |
ACCEPTED (A1). F replaced with emphatic do-support, which German has no periphrasis for |
| 2 | BLOCKING | site 10's L is the literal rendering of the same clause P renders, so a seat treating "carries a feature" as "renders this clause" can honestly QUOTE the same German at P |
ACCEPTED IN PART (A2). The drop is overruled — L does carry three formal features P lacks. The feature claim is restated as the one P demonstrably lacks (feminine dative for the town), and the critic's second half is accepted in full: the Pass A prompt now says in terms that faithful rendering is not by itself a departure and its answer is NONE |
| 3 | BLOCKING | site 6's L (dwelt confidingly) is a lexical choice, not a syntactic carry, so that site contrasts kind-of-markedness rather than licence |
ACCEPTED (A3). The site moved to the clause whose German holds a separable prefix to the end (er sah … an), and L now carries that order |
| 4 | BLOCKING | the registered equivalence reading of P1 (90% interval inside ±0.20 ⇒ NOT VISIBLE) cannot be supported at ten sites of three binary cells, and the design cites S134 missing a bar by 0.027 as precedent against itself |
ACCEPTED (A4). The equivalence branch is struck. Readings are now ATTRIBUTION TRACKS LICENCE / NOT SHOWN TO TRACK LICENCE, and the result page must call a null a failure to demonstrate |
| 5 | MAJOR | the gate stack withholds the primaries on failures that do not bear on them, biasing the run toward abandonment | ACCEPTED (A5). C1 now withholds only the identification claim; C2 withholds only P3; C4 is named as the one gate that can void the headline |
| 6 | MAJOR | the HIT rule does not operationalize "say what it is carrying", and can score a whole-segment dump |
ACCEPTED (A6). HIT demoted to secondary, and a registered length bound added: a quote longer than 60% of the segment's tokens is not a hit |
| 7 | MAJOR | four F rows may still be licensed or may read as register rather than syntax (sites 3, 7, 8, 9) |
ACCEPTED IN PART (A7). Site 7's F replaced by a cleft, site 8's agentive of replaced by polysyndeton. Site 3 kept — the German comparative has no verb at all, so an inverted verb answers to nothing. Site 9 overruled in writing: existential there puts the heavy NP last and the German fronts the object first; the two information structures are opposite, not the same |
| 8 | MAJOR | S scrambles a span, RS-20260809h's CLUNKY scrambled a whole segment, so P2's mechanism-transfer claim overreaches |
ACCEPTED (A8). P2 narrowed to consistent with, and the size difference stated in the primary itself |
| 9 | MAJOR | three of ten L rows are verb-first; an effect could be "verb-first is nameable" rather than "licence is visible" |
ACCEPTED (A9). A verb-first subgroup contrast is now pre-registered, with the reading that a subgroup-only effect is reported as such |
| 10 | MAJOR | one seat sees all four near-copies of every site | ACCEPTED IN PART (A10). A between-seat design would leave ten cells per condition and no testable contrast, so it is overruled; the separation constraint is raised from 6 to 8 and constructed rather than sampled, and the limit stands in §6.5 |
| 11 | MAJOR | the unit's question is "say what it carries" and the measure is attribution, not identification | ACCEPTED (A11). §1 now states the estimand as attribution and says in terms what a positive does not establish |
| 12 | MINOR | site 2's L is a preposition swap, nameable from the lexicon without reading German syntax |
ACCEPTED (A12). Site moved to schwerer ist's; L now carries the fronted comparative with inversion |
| 13 | MINOR | site 5's L may read as broken English and be confused with S |
ACCEPTED as documentation. Recorded here; the S spans are printed in §7 so the two can be compared |
| 14 | MINOR | C3's pooled mean gap is brittle |
ACCEPTED (A14). C3 is reported pooled and paired per site, and is explicitly non-gating |
| 15 | MINOR | the narrative rewards the positional strategy that §4 calls the hypothesis | ACCEPTED, folded into A5/A6/A11: success is differential NAME, never HIT rate |
| 16 | MINOR | §2's "chosen for density" features and the ten tested rows diverge | ACCEPTED. The tested rows are the ten in §7 and nothing else is claimed to have been tested |
What the critic did not find, and it is worth recording: it did not challenge the source, the collation, the contamination declaration, the seat composition, or the cost estimate. Its four BLOCKING findings were all about the materials — three individual rows built wrong, and one statistical reading the material could not support — which is the failure mode a design this dependent on hand-built minimal pairs should expect.