@@ -144,13 +144,40 @@ reproducing its documented tier on this replication.
144144
145145Gate ≤ teacher + 0.5pp: the student misses by two orders of magnitude.
146146
147- ** RETRACTION (2026-08-24) and RESTORATION (same day):** the original
148- verdict was retracted when the labels proved mojibake (the byt5
149- decode_joined bug); the clean-label re-run restores it — 33M student,
150- 11,792 byte-exact r6 labels, train CE 1.55, windowed DER-CE ** 82.87%**
151- vs the teacher's 1.32% on the same 300-paragraph harness. The failure
152- mode is total: the trained student emits EOS immediately on free
153- running (empty output; it fits the training set under teacher forcing
154- but cannot sustain generation). The capacity conclusion for Arabic is
155- now UNCONFOUNDED and matches Thai: sub-100M from-scratch byte students
156- do not generalize; a pretrained backbone is non-negotiable.
147+ ** RETRACTION (2026-08-24):** this verdict is CONFOUNDED — every Arabic
148+ label generated before the byt5 ` decode_joined ` fix was mojibake
149+ (double-encoded targets); both Arabic students trained on corrupted
150+ labels, and their identical DER scores are the bare-text constant, not
151+ a capacity result. The numbers stand as measured but the capacity
152+ conclusion for Arabic is UNPROVEN pending a clean-label re-run. The
153+ Thai tiny verdict is unaffected (umt5/sentencepiece labels were
154+ byte-exact); the pretrained-backbone law rests on Thai evidence.
155+
156+ ## ara-diac-small-1.0 — Arabic client tier (2026-08-24)
157+
158+ Sequence-level KD from the r6 teacher (rababa_arabic_byt5/run-006-morph,
159+ 2.5793 windowed DER-CE full-protocol): 29,322 greedy labels on r5-units
160+ (domain + replay, 1400-byte windows), ByT5-small init, 3 epochs. Same
161+ windowed harness as the ara-diac-tiny verdict (300 SadeedDiac-25
162+ paragraphs, Misraj evaluator, haraqat projection):
163+
164+ | Model | DER-CE |
165+ | ---| ---|
166+ | Teacher (r6) | 1.32% |
167+ | ** Student (ByT5-small, client rung)** | ** 3.66%** |
168+
169+ Gate discussion: the strict budget (teacher +0.5pp) is missed by
170+ +2.34pp — the measured capacity cost of ByT5-small on Arabic
171+ diacritization, consistent with the Thai client tier (+7.63pp beam-4 /
172+ 2.85% greedy against a 4.43% teacher). Shipped as the Arabic client
173+ rung per that precedent: the student generates real, well-voweled
174+ Arabic at a fraction of the teacher's artifact (1.3 vs 2.6 GiB), the
175+ strict gate is met by the teacher release (ara-diac-1.0), and the miss
176+ is disclosed rather than averaged away. For comparison, this student's
177+ 3.66% sits in the same league as the r3 production teacher's era
178+ (2.68% on the older full-set protocol).
179+
180+ Training notes: this is the third training of run-002 — the first on
181+ mojibake labels (byt5 decode_joined bug), the second silently resumed
182+ from the poisoned lineage's checkpoints (now guarded by labels.sha
183+ digest matching), this one clean end-to-end. CE plateaued at ~ 0.016.
0 commit comments