Repository path: wiki/findings/results/RS-20260728j-classb-marking.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260728j-classb-marking |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | accuracy, voice, style-correspondence, naturalness, consistency, affect, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-longwork.md, wiki/findings/results/RS-20260727-jeli-fid.md, wiki/findings/results/RS-20260727d-criterion-widening.md, wiki/findings/results/RS-20260728d-jeli-marks.md, workshop/experiments/E-20260728j-classb-marking/design.md, workshop/experiments/E-20260728j-classb-marking/forced/renderings.md, workshop/experiments/E-20260728j-classb-marking/dole/sites.md, workshop/translations/jeli-il-pastore/R05-v1/translation.md, workshop/translations/jeli-il-pastore/R05-v1/span2-fid-prereg.md, wiki/findings/theory/TH-20260724-translation-distance-axes.md, wiki/method-notes.md, wiki/base/consulted.md |
RS-20260728j — the marking was available all along, in a different grammatical category; and ARM-longwork closes
The unit. S052, principal unit ARM-longwork step 6 (T1), the arm's last step, taken on the
ledger's own selection (T1 at 4, the tool's named track). Two limbs, wired in one sentence:
the study limb reads Dole 1896 at span 2's nine Class B sites to ask whether the frozen log's
"nothing to repair with" is a fact about English or about the translator, and the translation limb is a
forced re-translation of those sites, frozen before the comparator was opened, which asks the same
question from the inside.
The discipline, and it is why this page can say anything. Design frozen at 7968abe → independent
pre-run critic (NEEDS-REDESIGN, ten findings) and amendment at 1582101 → every lead
judgment, prose and gradings alike, frozen at e923abc → only then fetch_dole.py.
verify.py: 91 checks, 0 failures. Cost $0.024405625, one critic call, 17% of its worst case.
1. The claim, and what happened to it
Log decision D22, frozen at 215434c and repeated as RS-20260727-jeli-fid §2's headline:
The other three are conditionals … here the English would is the ordinary unmarked form, so the Italian's out-of-sequence marking simply does not appear … there is nothing to repair with — English has no marked conditional to reach for. Six of nine carried; three of nine flattened with no available alternative.
Two clauses, and only one of them is true.
| id | kind | frozen (S039) | forced (S052) | Dole 1896 |
|---|---|---|---|---|
| B1 | gnomic | MARKED m1 | — | MARKED m1 |
| B2 | gnomic | MARKED m1 | — | MARKED m1 |
| B3 | gnomic | MARKED m1 | — | AMBIGUOUS |
| B4 | conditional | UNMARKED | MARKED m1 | UNMARKED |
| B5 | gnomic | MARKED m1 | — | MARKED m1 |
| B6 | conditional | UNMARKED | MARKED m1 | MARKED m1 |
| B7 | conditional | UNMARKED | MARKED m1 | MARKED m1 |
| B8 | gnomic | MARKED m1 | — | MARKED m1 |
| B9 | modal | MARKED m2 | — | UNMARKED (see §5) |
"English has no marked conditional to reach for" is true and stays true: nothing in either the forced renderings or Dole marks by producing an English conditional. "Three of nine flattened with no available alternative" is false. All three sites admit a marked rendering; two of the three are marked by a published translator of 1896; and every successful marking, mine and Dole's alike, works by changing grammatical category — conditional → present tense, or conditional → were-subjunctive.
The general form, and it is what the arm hands on. A marked source form need not be matched by a marked target form of the same category. A translator searching category-for-category — looking for an English conditional to answer an Italian one — will correctly report an absence, and the absence will not be there. The frozen log named the right gap and drew the wrong conclusion from it, and it did so in a sentence that read as a fact about the language pair.
2. Predictions
Six were registered. Four are falsified, and the two confirmed ones are the two worth least — P1 was declared pre-seen on the design page before anything was rendered.
| # | prediction | outcome |
|---|---|---|
| P1 | ≥1 conditional admits a MARKED rendering, ≥1 grammatically | confirmed — declared pre-seen; carries nothing |
| P2 | all three conditionals admit one | confirmed |
| P3′ | grammatical marking changes NP definiteness | FALSIFIED at B4 and B7 |
| P4 | Dole marks ≥4 of 5 gnomic sites and 0 of 3 conditionals | FALSIFIED — 4 of 5 holds, 2 of 3 conditionals marked |
| P5 | forced and Dole never coincide at a conditional | FALSIFIED at both B6 and B7 |
| P6′ | ≥2 closed-list particles in the forced ¶42 | FALSIFIED on the count (N = 1), confirmed on clustering |
P5's falsification is the strongest single result on this page. The repair the lead invented under a forced brief, from the Italian alone, with Dole unopened, is the repair Dole actually used — at both sites, in the same class, in the same construction:
B6 —
col solfato si sarebbe guarita subitoforced: "sulphate cures it at once" · Dole 1896: "a little quinine cures it quickly"B7 —
quando fosse rimasto soloforced: "were he left alone" · Dole 1896: "in case he were left alone in the world"
Two translators, 130 years apart, one working from the Italian with the other's text closed, reach the
same construction at the same two sites. That is the answer RS-20260727-jeli-fid §5 asked for and
could not have.
P3′'s falsification is worth more than P3′ was. The design's original P3 said the cost of marking was
episodic → generic; the critic replaced it with a determiner test readable off the page; the determiner
test fails at two sites of three. The cost is not referential — it is modal and temporal: B4 turns a
hypothetical into an assertion (saprebbe → knows), B6 turns a counterfactual into a general truth
(si sarebbe guarita → cures it), B7 turns a certainty into a supposition (quando fosse rimasto solo
is when, not if, because the father is dying). Three sites, three different costs, none of them
visible in a determiner — and the loose original P3 would have been confirmed on exactly this data.
One caveat on B6, and the critic named it before the run. Dole's cures sits in a sentence he has
already shifted to the present at a point where the Italian has an imperfect ("the disease is one of
those clear and evident ones", for la malattia **era**). So B6's mark is real by the frozen definition
but is not clean evidence that Dole was responding to the conditional — he may have been generalising
the whole sentence. B7 has no such confound and carries the result on its own.
3. The frozen anchor procedure failed at exactly the two sites that falsify the prediction
Amendment B2 replaced a discretionary NO-COUNTERPART rule with lexical anchors frozen before Dole was opened, precisely so the lead could not rescue P4 by declaring an inconvenient site non-corresponding. Seven of nine anchors hit. The two that missed are B6 and B7.
- B6's anchor required
sulphate; Dole writes quinine — a realia substitution, strategy three oncultural-mediation's own list. - B7's anchor required one of
help / shift / manage / take care / care of; Dole writes "do for himself".
Applied mechanically, the control written to stop the lead protecting its prediction would have protected it — NO-COUNTERPART at both sites, P4 and P5 confirmed, the page reporting that a published translator declines the mark exactly as the lead's policy did. The sites were found by reading the paragraph around them and are coded here against the lead's prediction.
The mechanism is general. An anchor written in the lead's own target vocabulary is a list of the words the lead chose. Where the comparator did something different the wording diverges most, so an anchor set built that way is least likely to match exactly where disagreement lives. It is biased toward finding agreement — the same shape as a control that can only confirm. Method note (bdi).
4. The pre-declared secondary measurement: strategic convergence is invisible to the instrument
Declared in Stage 1 with no prediction attached. Both renderings of ¶42 measured against Dole's ¶42
(255 words) with tools/dependence_check.py:
| words | shared 7-grams | shared 12-grams | longest run | |
|---|---|---|---|---|
| frozen ¶42 vs Dole | 249 | 3 | 0 | 8 |
| forced ¶42 vs Dole | 239 | 2 | 0 | 8 |
The two translators agree on the construction at two of four Class B sites in this paragraph, and the
longest shared run does not move by one token — it is the same eight-token string in both
("with more wood for the night and the"), at a site with no Class B mark in it at all. The
longest-run instrument measures lexical coincidence and not methodological convergence, and here the two
came apart in a measured case. That bears on ARM-forced-defence and on the S050 backlog row about
self-revision: a null run length is not evidence that two translators did different things.
5. Two defects in the frozen instrument, reported rather than patched
- The m2 closed list is too narrow. It contains
must haveand notmust be, so Dole's "Mara also must be grown tall" grades UNMARKED where the lead's "must have grown up" grades MARKED, on the same construction. B9 is not in P4's denominator and no reported figure moves. The list stays as frozen. - The verifier caught its own implementation.
verify.py's first run graded "would know" as a present-tense finite verb at three cells and disagreed with the stored table. The frozen §3 rule says "a finite verb in the present tense" and is correct as written; a token lookup needs a side-condition (not after a modal orto) to implement it, and needs the were-subjunctive detected uninverted as well as inverted. Both were added to the verifier and neither changed a datum — the disagreement was the verifier's, and it is recorded because a verifier that has never disagreed with the analysis has not been shown to be capable of it.
6. The paragraph-bounded recount — ARM-longwork step 6, absorbed item (b)
S047 measured a three-token inflation on this work's largest figure because the frozen metric does not
break runs at paragraph boundaries, and scheduled the same recount over the project's other published
figures onto this step. Done: 23 lead-vs-published pairs — the seven cells of
E-20260725c-contamination-sweep (15 pairs) and the eight Turgenev prose-poem loci of
E-20260726c / E-20260728b (8).
Two of 23 are inflated, both by 2–3 tokens, and neither carries a claim.
| pair | whole | paragraph-bounded | inflation |
|---|---|---|---|
posle-teatra ~ Garnett 1920 |
19 | 17 | +2 |
bezhin-lug ~ Garnett 1895 |
14 | 11 | +3 |
| all 21 others | — | unchanged | 0 |
And the figure that mattered most is unmoved: all eight Turgenev loci are paragraph-bounded exactly as
published — 21, 18, 17, 16, 16, 15, 13, 11. The 21-token run quoted in CLAUDE.md as the ceiling
of the project's measured contamination range lies inside one paragraph of both texts and stands at
21. The standing rule built on it is unaffected.
What this recount does not cover, stated because it is not everything: the per-artifact contamination
measurements filed on individual translation pages since S036 (Jeli spans 1–5, recounted at S047;
wang-liulang, takasebune, bargamot, yingyi-jiejixing, osso-di-morto, genealogie,
metamorphoses). Those are not a published set and no cross-artifact figure rests on them.
posle-teatra's and bezhin-lug's corrected values are recorded here; no published claim changes,
because neither figure is quoted anywhere as a threshold.
7. ARM-longwork step 6 — the closing report, per goodness sense
The arm's completion criterion: "the work is translated whole, with a cumulative log, and the arm reports
what it surfaced that short units did not — per goodness sense, naming any sense the long form pressured
that short units left idle." Five spans, 11,474 Italian words → 13,049 English, 195 paragraphs to 195,
log D1–D84, one binding register, seven errata. Every claim below is internal-judgment-only and
provisional, and every one of them is about what the form pressured, not about how good the result
is — the lead does not judge its own translation (charter §5).
consistency — the charter's named test bed, and the sense the arm actually moved
Four findings, and the second is the most transferable thing the arm produced.
- A policy can be adopted by drift and bind 11,000 further words.
D19's contraction split (contractions in speech, none in narration) was never decided; it was noticed at the first freeze as something already done. A short unit tidies such policies before it freezes and so never has to see them, which is why seventeen prior translations produced no instance. - The register catches forward drift and is structurally blind backwards. S047's mechanical census
found two rows wrong about prose they describe —
V7was never applied backwards to span 1's two guillemet sites, andV10's grammatical scope is contradicted bybianco biancoin span 2, three spans before the rule existed (erratum E6). One failure mode, found twice in one session, and it is a property of the instrument rather than of either row. What a long work needs is not a better memory but a register that is re-applied backwards at every freeze. - A term row can break on a referent it was never tested against.
N15setmandra→ herd atD7when every referent was equine; span 4 puts Jeli among sheep and English does not permit a herd of sheep (erratum E5). Four spans of latency between the decision and the collision. - The word most in need of a row never got one, because it did not look like terminology.
roba— Verga's great word — six occurrences, four English renderings, entered at S047 as irreducible. The register has rows for realia and for characters; it has no mechanism for an ordinary word that is doing structural work.
voice — pressured by frequency, which is precisely what a passage hides
Span 1 had two free-indirect sites and no policy could be built on them. Span 2 had 21, in three classes, one of which (bare-NP guillemets) does not occur in span 1 at all. The class that breaks a policy is often absent from the first two thousand words. This session adds the other half of the same point: the repair is also absent from a short unit's view. Nine instances of one feature are what made it visible that the failures were not random — that all three failures were conditionals, that the six successes were tense shifts, and therefore that the answer was a tense and not a mood. One or two instances would have shown a translator a gap and given no way to see its shape.
style-correspondence — the arm's sharpest result, and §1 changed it
The frozen claim was that a source marking is unrecoverable in this pair. It is recoverable, in a
different grammatical category, and an independent published translator recovers it at two of three
sites. This sharpens TH-20260724-translation-distance-axes C1 rather than contradicting it: C1 says
grammar-borne meaning is transcoded into lexis at a cost in systematicity; here it is transcoded into
other grammar, at a cost in modality and temporality and at no cost in systematicity at all. C1's
mechanism has a second branch and the theory page did not have it. (Recorded on this page; not applied
to TH-20260724, which is a cross-pair theory page and this is one pair, two translators, three sites.)
accuracy — the cost of the recovered marking, and it is licensed by the ratified clause
Under D-20260727-08's compelled-specification rule, an avoidable resolution of source openness scores
as unlicensed addition. All three markings are avoidable — the frozen prose demonstrates the alternative —
so all three score here: hypothetical → assertion, counterfactual → general truth, certainty →
supposition. The long form is what made this countable. Three conditionals inside one paragraph is
enough to see that the three costs are three different costs; one conditional would have produced one
anecdote.
naturalness — a null, and recorded as one
Across five spans and 84 logged decisions the arm produced no case where the long form pressured
naturalness in a way a short unit would not. In this session, no forced rendering was rejected as
unidiomatic; the single rejection (a present tense under the participial tag thinking) was
ungrammatical, not unnatural. The sense's problems here are the ordinary ones.
affect — a null, and the one that costs the arm something
Charter §3 names three senses defined over spans longer than anything the project had translated:
consistency, voice, and arguably affect. The arm moved the first two and produced nothing at
all on the third. 13,049 English words of a story whose whole force is cumulative — a boy's slow
isolation ending in a killing — and not one logged decision, in five spans, turns on an effect that
accumulates. The honest reading is that the arm was not designed to see it: a translator's log records
choices, and affect is a property of reception. Nothing in a serial-translation regime makes a
cumulative effect visible to the person producing it. Recorded as an unmet part of the charter's own
reason for wanting a long work.
cultural-mediation — worked hard, and one length-independent observation with a method consequence
34 term rows, onze/tumoli untranslated, gnà/compare/zio kept, Corna d'oro translated because
the horns are a thread. None of that needed a long work; the fork is the same fork a short unit meets.
The one thing worth carrying is §3's: Dole's solfato → quinine is a substitution, and it broke a
frozen control. A methodological instrument built on source-side realia terms is unsafe whenever the
strategy set includes substitution — which it always does.
literary-quality and purpose-fit — not reached
No claim. literary-quality would require judging the lead's own prose, which charter §5 forbids;
purpose-fit was fixed by R05 at the outset and never varied.
8. What this page does not license
Nothing here is evidence about C1, about English–Italian as a pair, or about any language pair in general. Two translators at nine sites in one story is a case. The design's §7 originally said a confirmed P4 would show that "the pair genuinely lacks the resource"; that sentence was withdrawn by amendment A4 before the run, and the limit stated there applies to the falsification just as much: what this shows is that one published English «Jeli» marked at two of three conditionals, and that the lead, working blind, found the same two constructions.
Every grade on this page is the lead's, applied by a mechanical definition frozen before any rendering
existed and re-derived independently by verify.py, which can check reproducibility and cannot check
correctness (amendment D). There is no independent grader. Every cell is quoted in full so a later
reader can regrade the whole table without re-running anything.
The frozen span-2 prose is not revised. R05 is append-only; erratum E7 records D22's false
clause on the translation page, and the prose stands as it was translated.