Repository path: workshop/experiments/E-20260804-displaced-marking-fr/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804-displaced-marking-fr |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | accuracy, style-correspondence, voice, cultural-mediation, naturalness |
| provisional | true |
| links | framework/v0.1/README.md, wiki/arms/ARM-framework-v01.md, workshop/experiments/E-20260804-displaced-marking-fr/materials/census.md, workshop/translations/la-nuit/R04-v1/translation.md, workshop/experiments/E-20260802e-displaced-marking/design.md, wiki/findings/results/RS-20260802e-displaced-marking.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260804 — the stress test: does R1 hold on a pair the release calls untested?
ARM-framework-v01 step 3 (T5), the closing step. Frozen before any seat was addressed, after the
translation, its log and the census were committed (caf6b2e, and the census in this commit).
1. The question
framework/v0.1 carries exactly one recommendation, R1 (displaced marking), and registers four
predictions about what will happen when it is applied to fresh text. The first is the release's own
stress test and is stated there in these words:
1. On fresh Class A sites in a new pair, R1 recovers a marking at more than half. Score by the
E-20260802eprocedure. The failure it should not survive: recovery at or below a third.
This experiment is that test. The pair is FR→EN, which framework/v0.1 §7 declares
untested; the sites are eight fresh Class A sites from a translation made this session; the
procedure is E-20260802e's, unchanged in every respect that could move the answer.
What it teaches about translating literature, in one sentence (the subject rule): whether the one piece of advice this project is prepared to give a translator — before you record a grammatical marking as lost, render the site again under a brief that requires the marking to appear — survives contact with a language and a text that had no part in producing it.
2. Why this text, and what it forecloses
Maupassant, «La Nuit (cauchemar)» (1887), whole, 1,873 words, translated close under R04 as
T-la-nuit-R04-v1. Three properties of the material do work here:
- French was not in R1's evidence base. Six source languages were; French was not.
- The story has no addressee.
framework/v0.1§4 records that five of R1's six licensed device categories have never been isolated, and step 2 exercised only the address noun. A solitary first-person monologue cannot be marked with an address noun, so the categories this run exercises are whatever is left. This is a property of the text, not a rule imposed on the raters. - The sites are not the project's old ones.
E-20260802etested the six remaining Class A sites in the whole prior record and exhausted that population. These eight are new.
What it forecloses, stated first because it is the design's weakest joint: the lead knew R1 while translating. Three things constrain that and none of them removes it.
- The FROZEN rendering is the filed translation — the artifact the project publishes as its
English «La Nuit», quoted verbatim and asserted so by
materials/sites.py. Degrading a site to make R1 look good degrades the deliverable, in a file Tom reads. - The census is over all 31 log rows, not a filter output, and it classified 12 rows Class A and 3 Class B — sites where the marking was reached in the ordinary course of translating. A log written to manufacture Class A sites would not contain those three.
- DECOY is the control on the whole arrangement (§6, F1), exactly as it was at
E-20260802e.
The honest statement is: this is a stress test conducted by the party with an interest in the result, and its FROZEN arm is the one an adversarial reader should attack. It is not blinded by authorship and does not claim to be.
3. Materials
materials/sites.json, emitted by materials/sites.py, which asserts every French span verbatim
against the frozen copy-text and every FROZEN rendering verbatim against the frozen translation, and
checks both DECOY constraints. All eight pass; two DECOYs failed the length constraint on first
writing (S3 at 0.222, S5 at 0.167) and were rewritten rather than the threshold moved, as at
E-20260802e A4.
| site | ¶ | device | FORCED lands on |
|---|---|---|---|
| S1 | 7 | expletive ne | a parenthetical clause |
| S2 | 30 | impersonal on, phantom household | a pronoun + a trailing relative |
| S3 | 15 | pronominal middle | argument structure — the colour becomes subject, the vegetables become places |
| S4 | 17 | durative imperfect | verb periphrasis |
| S5 | 27 | grammatical gender | a pronoun |
| S6 | 28 | inverted parenthetical, register | word order — English literary inversion |
| S7 | 29 | pluperfect subjunctive, register | word order — fronted existential |
| S8 | 37 | inchoative passé simple | argument structure — the will becomes subject, the narrator object |
Six of the eight FORCED renderings mark outside the address-noun category, and none marks inside it. The categories exercised are pronoun (2), argument structure (2), word order (2), verb periphrasis (1), parenthetical clause (1).
Contamination: suspected, unmeasured, and the measurement was attempted and failed. No published
English translation of this story was reachable (workshop/translations/la-nuit/collation.md §4, with
the search log). Per CLAUDE.md this is declared on the artifact rather than passed over. It bears on
this design more weakly than on most, because the claim under test is existential — does a marked
rendering exist — and a rendering that happens to coincide with a remembered published one is still an
existence proof. Note (bhb) cuts toward the null here as it did at E-20260802e.
4. Procedure — the E-20260802e stage-1 procedure, unchanged
- Four renderings per site, all frozen before any seat is addressed: FROZEN (the filed rendering), FORCED (written under R1's brief — the marking must appear, in any category), DECOY (differs from FROZEN in wording, attempts no marking, length-matched to FORCED within 15% and differing from FROZEN by ≥ 20% of tokens), POSITIVE (an explicit metalinguistic gloss).
- Presentation order per site is fixed by
sha256(site id | seat id), not chosen by the lead; each seat sees a different rotation. - Three non-Anthropic seats — P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5, the same three asE-20260802estage 1 — receive the French span, a literal gloss, the relation stated neutrally with the source's device never named, and the four renderings unlabelled. Per rendering: does an English reader with no access to the source get this relation from it — YES / NO, plus one line of reason. - Two orderings per seat (the rotation and its reverse): six bodies.
Operational prescriptions carried in from the notes, before the first dispatch, not after it.
effort: low on the first dispatch to every seat (note (b), which fired at S097, S099 and
S100 and twice cost money for zero characters). max_tokens 16,000 for P2 from the outset, not
4,000, because note (bhf) and E-20260802e A5 measured this seat's reasoning appetite on this
exact task shape. Every dispatch runs with a client timeout shorter than the harness's, because
note (bid) — S100's orphaned request, killed client-side and billed anyway — is one session old.
5. Predictions — registered, six, before any call
| # | prediction |
|---|---|
| P1 | (primary; framework/v0.1 prediction 1) FORCED is graded YES by ≥ 2 of 3 seats at more than half the sites — ≥ 5 of 8. The failure the release said it should not survive: ≤ 3 of 8 |
| P2 | FROZEN is graded NO by ≥ 2 of 3 seats at ≥ 6 of 8 sites — the log's own loss claims reproduce |
| P3 | DECOY is graded YES by ≥ 2 of 3 seats at ≤ 1 site |
| P4 | POSITIVE is graded YES by ≥ 2 of 3 seats at 8 of 8 sites |
| P5 | (framework/v0.1 prediction 2) S5 fails — the metalinguistic site, where the relation is that a marking is automatic rather than chosen, is not recovered from FORCED |
| P6 | (framework/v0.1 prediction 3) Applying R1 raises logged decisions and does not reduce declared losses to zero — countable from the census: Class A ≠ 0 |
The lead's stated expectation, recorded so it can be wrong. P1 holds at 5 or 6 of 8. The likeliest failures after S5 are S6 and S7, the two register sites: both FORCED renderings mark register by English literary word order, and a grader may read the inversion as ornament rather than as a claim about where the sentence belongs. If S6 and S7 both fail alongside S5, P1 lands at exactly 5 of 8 and survives on its margin — which is a thin result and should be reported as one. If S2 fails, the interesting sentence is not about R1 but about English: no construction may exist that posits an agent and withholds their identity in one clause.
6. Failure criteria — pre-committed
- F1 — the grading instrument failed. POSITIVE graded NO at more than 1 site, or DECOY graded
YES at more than 1 site. → the run is descriptive only; nothing enters or amends
framework/v0.1on it, whatever P1 says. - F2 — the census misread the log. FROZEN graded YES at ≥ 4 of 8 sites. → P1 is uninterpretable, because the sites were not losses to begin with, and the classification is the finding.
- F3 — a seat fails. A seat that returns
finish_reason: lengthwith zero characters of content on two consecutive dispatches is withdrawn, and the run reports with the remaining seats rather than substituting a fresh one. This isE-20260803gFC4, which S100 kept when keeping it cost it a seat; it is carried here verbatim so the decision is not made after seeing a verdict. - F4 — what a pass buys. P1 holding discharges
framework/v0.1prediction 1 and licenses one sentence: R1's recovery result extends to FR→EN at these eight sites. It licenses nothing about quality (Tier D is NOT PASSED, evidence class X3 is inadmissible), nothing about human readers, and no change to R1's text. - A null is a result. If P1 fails at ≤ 3 of 8,
framework/v0.1prediction 1 is falsified on its own registered threshold, R1's pair table gains a row reading FR→EN — refuted, and the release's §2 sentence "the absence of a same-category counterpart is not the absence of the marking" is demoted to the pairs it was evidenced on. That would close this arm on a negative and it would be the more valuable outcome, because a framework that survives every test it sets itself is not being tested.
7. What this design cannot show
- Availability, never quality. A marked rendering existing does not make it better;
E-20260802emeasured the cost at about two-thirds of a point of naturalness and this run does not re-measure it. - Three language models are not a survey of English readers (charter §4). A YES licenses three independent readers recovered the relation from this English and nothing about human readers. No human reader is available to this project — Tom is never an experimental subject (charter §9).
- The lead wrote all four renderings, and knew R1 while writing the first. §2. Blinding is by presentation, not authorship; DECOY and F1 are the controls, and they are controls on the graders, not on the author.
- Eight sites is the population of this text, not a sample of French. No sampling inference is available and none is drawn.
consistencyandpurpose-fitare not in play and are not scored.
8. Verification
analysis/verify.py, importing nothing from tools/: recomputes the census row count and the
Class A/B/Neither counts from the frozen log, every French span and every FROZEN rendering verbatim
against the frozen files, both DECOY constraints at all eight sites, the hash-derived rotations, all
six predictions and all four failure criteria from the raw stored bodies, and the billed total by
re-summing usage.cost over every stored body. At least three mutation tests, each asserting the
bytes on disk changed and then restoring (notes (bgu), (bhd)).
9. Budget
Declared worst case $1.10, built from max_tokens and not from expected output (note (abc)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | 1 | 12,000 | $0.20 |
| stage 1 — grading, 3 seats × 2 orderings | 6 | 8,000 (P2: 16,000) | $0.60 |
| retry reserve | — | — | $0.30 |
Today is a fresh UTC day (2026-08-04) with the full $5.00 available and no session has yet spent against it. Lead translation is free and is never ledgered (charter §3, A4).
10. Amendments, 2026-08-04, after the pre-run critic (critic.md) — all four accepted
Pre-run critic: P4 moonshotai/kimi-k3, a seat with no other role in this run.
NEEDS-AMENDMENT, four findings, three BLOCKING, all four accepted. Note (rr) fires for the
forty-third consecutive session: the critic's first BLOCKING finding is against the design's own
strongest claim, and §2's self-declared "weakest joint" was not the joint it found.
A1 (BLOCKING, finding 1) — the yardstick was written by the same hand as the answer
The critic's finding: the eight relation statements are the grading standard, and they were written by the author of the FORCED renderings, with those renderings in view. S2's relation is satisfied word-for-word by S2's FORCED and by nothing else; S1's and S8's likewise. The grading question then measures whether a reader can detect the clause the author embedded to match the relation the author wrote — a closed loop, and P1 could reach ≥ 5 of 8 with no evidence that R1 recovers anything. §7 declared the renderings unblinded and said nothing about the yardstick. The finding is correct and it is the finding this design most needed.
The change: every relation statement is re-authored by a seat that has never seen any rendering,
and the lead's eight are discarded. Seat P5 deepseek/deepseek-v4-pro, no other role in this
run, receives the French span, the literal gloss, and a pointer to the source construction — which is
a fact about the French, not about any English — and returns one sentence per site stating what
relation or attitude that construction conveys. It is shown no English rendering of any kind.
Its statements are frozen to materials/relations-P5.json before any grading seat is addressed and
are what the graders see. A mechanical check rejects any statement containing a term from a frozen
grammatical lexicon, so the graders cannot be told the device by the back door; on rejection the seat
is re-dispatched once, and if it fails twice the lead's statements are used with this amendment
recorded as not carried out.
A2 (BLOCKING, finding 2) — F1 is a seat-function check, not an instrument check
The critic's finding: POSITIVE states the relation outright, so grading it YES tests reading comprehension; DECOY was written by the motivated author under the brief "attempt no marking", so its low YES rate is manufactured rather than measured. An instrument that grades YES exactly where the author wanted YES passes F1 cleanly.
The change, in two parts.
- A fifth arm,
NEUTRAL, which the lead did not write and did not tune. Seat P5, in a separate call that is shown no relation statement and no other rendering, translates each French span under a plain brief — translate this well into English — with no mention of relations, markings, losses, or the study. These are renderings by a party with no interest in the outcome. New failure criterion F5: if NEUTRAL is graded YES by ≥ 2 of 3 seats at ≥ 4 of 7 scored sites, the FORCED result is confounded — any competent independent translation recovers these relations, R1's brief adds nothing, and the run is descriptive only. - F1 is relabelled in the design and will be relabelled in the result: it is a check that the seats are reading and are not grading every changed rendering YES. It is not evidence that the instrument discriminates a marking the author did not plant. NEUTRAL is what carries that load, and it carries it imperfectly.
Ordering, so the two P5 calls cannot contaminate each other: the relation call runs first, from source alone; the neutral-translation call runs second and is shown neither the relations nor anything else. Both are authored by the same model, and the bias that introduces runs toward NEUTRAL scoring YES — against FORCED's distinctiveness, i.e. conservative for R1.
The deviation from E-20260802e this creates, declared: graders see five unlabelled
renderings per site rather than four, which could move base rates in an unknown direction. The
four-condition comparison prediction 1 names is nevertheless recoverable unchanged from the same
bodies, because every rendering is graded independently, and it is reported that way.
A3 (BLOCKING, finding 3) — P5-as-registered was true by construction
The critic's finding: S5's relation includes "automatically rather than by any choice the writer made", and no rendering can convey automaticity except the one that states it. FORCED's "She seemed to be alive" conveys femaleness and cannot convey that the femaleness was unchosen. So the prediction was never at risk, and S5's guaranteed failure sat inside P1's denominator being absorbed by a margin the lead had already narrated.
The change. Prediction P5 is withdrawn, not confirmed. S5 is removed from P1's denominator
and kept as a demonstration site. P1 is now: FORCED graded YES by ≥ 2 of 3 seats at ≥ 4 of 7 sites;
the failure it should not survive is ≤ 2 of 7 — more than half and at or below a third on seven,
which is what framework/v0.1 prediction 1 states.
And a finding is registered in P5's place, before the run: framework/v0.1 prediction 2 is not
falsifiable by the E-20260802e procedure. A metalinguistic relation — that a marking is automatic
rather than chosen — can only be conveyed by metalanguage, which is precisely what the POSITIVE arm
is. The procedure cannot distinguish "R1 fails on metalinguistic relations" from "this instrument
cannot score metalinguistic relations". That is a defect in the prediction as the release wrote it,
it is now on record, and it is worth more than the confirmation it replaces.
A4 (NON-BLOCKING, finding 4) — a pre-registered deflection
The critic's finding: §5's sentence "If S2 fails, the interesting sentence is not about R1 but about English" converts a P1-relevant failure into a claim the run has no instrument to support. Accepted; the sentence is struck and does not appear in the result under any wording. The within-device picks stand on the reasons already given.
And the part of finding 4 the fix did not cover, recorded rather than argued with: the census, the
site picks and the four lead-authored renderings were committed in one commit (57e1e7e), so
their internal order rests on the lead's say-so. What is independently verifiable from the history
is that the translation and its log were complete and committed first (caf6b2e, and the draft at
effc944 before that). Nothing else about the ordering is checkable, and the result will say so.
Budget after amendments
Two added calls (P5 relations, P5 neutral translations), and longer stage-1 prompts.
Declared worst case $1.10 → $1.50, built from max_tokens with a 4× routing margin on P5
(note: the S022 measurement — P5 has billed at 3.8× list through provider routing).
Today's headroom at session open: the full $5.00.
11. Amendments A5–A10, 2026-08-04 — the SECOND critic pass, and how there came to be one
There are two pre-run critic passes, both billed, and the second exists because the runner was
launched twice by mistake. run_critic.py was invoked, appeared not to have started, and was
invoked again; both dispatches went to P4 moonshotai/kimi-k3 with the byte-identical prompt and
temperature 0, and the second overwrote the first's stored body.
They did not agree.
| pass | provider | completion tokens | cost | verdict | findings |
|---|---|---|---|---|---|
| 1 | Modal | 2,312 | $0.0586302 | NEEDS-AMENDMENT | 4, three BLOCKING |
| 2 | Moonshot AI | 2,901 | $0.067638 | NEEDS-REDESIGN | 6, three BLOCKING |
Both are kept and the union of their findings is accepted. The rule this project wrote for itself
at E-20260803g — report with one reviewer rather than go shopping for a friendlier second —
forbids discarding a harsher verdict, and it is the harsher verdict that arrived second. The two
passes share their first two findings and pass 2 adds three the lead had not thought of. A1–A4
above stand; A5–A10 follow.
The instrument fact, recorded and not investigated: one seat, one prompt, temperature 0, two
providers, two different verdicts and two different finding sets. It bears on how much weight a
single critic pass carries anywhere in this project. It is method work under the subject rule and it
is not this session's question; it goes to the result's limits, not to the backlog.
A5 (pass 2, finding 1) — the same finding as A1, plus a check
Pass 2 adds that nothing in the design procedurally excluded the leak: only the census was declared
frozen before the renderings, not the relations. A1's repair stands — the relations are re-authored
by P5 from the source alone — and analysis/verify.py additionally asserts that every grader prompt
contains P5's relation strings verbatim and contains none of the lead's, which is checkable from the
stored prompt files rather than from anyone's say-so.
A6 (pass 2, findings 2 and 4) — a control the author could not tune
Pass 2's finding 4 is the sharper form of pass 1's finding 2: POSITIVE paraphrases the relation statement, so P4's 8-of-8 tests only whether graders can match a rendering to a restatement of the yardstick. An instrument that grades "does this rendering restate the relation" passes F1 on both legs while invalidating every YES in P1. Its fix is a negative control, and it is taken.
A sixth arm, WRONG: at every site, an explicit statement of a plausible but DIFFERENT relation,
in the same rhetorical shape as POSITIVE. S1's asserts indifference where the target is resistance;
S4's asserts that the narrator hurried away where the target is unbounded continuing; S5's asserts
that the watch was the last working thing in the city where the target is that it is a she.
- New failure criterion F6. If WRONG is graded YES by ≥ 2 of 3 seats at ≥ 2 of the 7 scored sites, the instrument is scoring compliance with a stated relation rather than recovery of a relation from prose, and the run is descriptive only whatever P1 says.
Graders therefore see six unlabelled renderings per site: FROZEN, FORCED, DECOY, POSITIVE, WRONG,
NEUTRAL. The four-condition comparison framework/v0.1 prediction 1 names is recoverable unchanged
from the same bodies, because every rendering is graded independently, and it is reported that way.
A7 (pass 2, finding 3, BLOCKING) — the undefined band
Pass 2: P1 registered a pass and a falsification and left the middle undefined, so a live outcome would have had its interpretation negotiated after the data existed. Correct, and it is exactly what pre-registration exists to prevent.
Registered now, before any grading seat is addressed, over the 7 scored sites (S5 excluded by A3):
| FORCED recovered at | disposition |
|---|---|
| ≥ 4 of 7 | more than half — framework/v0.1 prediction 1 DISCHARGED on FR→EN |
| exactly 3 of 7 | INCONCLUSIVE. Prediction 1 is neither discharged nor falsified; the pair table reads partially tested; one registered follow-up is permitted and it must be a new pair, not more sites in this one |
| ≤ 2 of 7 | at or below a third — prediction 1 FALSIFIED on its own threshold |
A8 (pass 2, finding 5) — selection on the dependent variable, accepted, repair only partial
Pass 2: S8 takes ¶37 over ¶31 because ¶31 "gives a grader nothing to read" — a pick made on expected gradeability, which is selection on the dependent variable. The finding is accepted and the proposed repair — a mechanical rule registered before any rendering was written — cannot now be carried out, because the renderings exist. Pretending otherwise would be worse than saying so.
What is done instead: P1 is reported twice, on all 7 scored sites and on the 6 excluding S8, so a reader can see whether the gradeability-selected site is carrying the result. And it is recorded that a longest-span rule would also have changed S2, from ¶30 to ¶10. Neither figure is chosen after the fact: both are reported.
A9 (pass 2, finding 6) — the two-seat rule
Every threshold reads "≥ 2 of 3 seats", which F3 could make undefined mid-run. Registered: if F3 withdraws a seat, the threshold becomes 2 of 2 for the remaining pair, a 1–1 split counts as NO, and the result states the reduced power in the sentence that reports the figure. If two seats fail, the run is descriptive only.
A10 — the double dispatch, ledgered and guarded
Both critic requests are billed and both are in config/budget.md. The first dispatch's .raw
body was overwritten; its metadata and its full response text are preserved verbatim from the run
console at runs/critic_pass1_kimi-k3.meta.json and runs/critic-pass1-response.md, labelled as
recovered rather than stored, and the per-request re-sum is therefore short by $0.0586302 against
the key-usage delta unless that figure is added by hand. It is added by hand and flagged.
The guard, added to call.py before the next dispatch: note (bgz) made every attempt within
one run uniquely labelled; it did not survive the script being run twice. The runner now refuses to
overwrite an existing .raw and labels a repeat dispatch _re2, _re3. A write-once discipline
that only holds inside one invocation is not a write-once discipline.
The final register
Scored predictions: P1 (as amended by A3 and A7), P2, P3, P4, P6. P5 withdrawn (A3). Failure criteria: F1 (relabelled a seat-function check by A2), F2, F3, F4, F5 (NEUTRAL, A2), F6 (WRONG, A6).
A11 — the lexicon gate fired, and its sanction is narrowed before the re-dispatch
P5's relation for S2 contains the word "impersonal" — in the phrase "impersonal refusal", which is a description of a feeling and not the naming of a grammatical category. A5's gate does not know that; its lexicon also matches verb, noun, article and voice in their ordinary senses.
Written before the re-dispatch and before its output was seen. A5's sanction — two failures and the lead's statements return — would hand the yardstick back to the author over a false positive, which is a worse outcome than the one the gate exists to prevent. The gate is therefore applied per statement, not per batch: the relations call is re-dispatched once as registered; only the flagged site's statement is replaced, and if the replacement also trips the gate, P5's original statement is kept with the hit declared on the result rather than the lead's substituted. Both P5 sets are stored. The lead's eight statements are not used as the yardstick under any branch.
A12 — how a seat's two orderings combine, registered before any grading body exists
Each seat grades every rendering twice, in a hash-derived rotation and its reverse. The design did not say how the two combine, and that must not be settled after the data exist.
Registered: a seat grades a rendering YES only if it returns YES in BOTH of its orderings. A split counts as NO. This is the conservative direction — it can only lower every YES rate, including FORCED's — and it controls the slot preference this panel was measured at 0.500–0.600 on at S020. The per-ordering figures are reported alongside, so the effect of the rule is visible rather than buried.