Skip to content

Commit 7f522c4

Browse files
author
Ronald Tse
committed
docs(verdicts): run-009-yallamorph teacher negative on both surfaces - r7 stays canonical
1 parent 27676d3 commit 7f522c4

3 files changed

Lines changed: 60 additions & 10 deletions

File tree

‎TODO.sota-2026/01-r8-teacher-yallamorph.md‎

Lines changed: 36 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,6 @@
11
# 01 — r8 teacher: run-009-yallamorph (YallaMorph/CamelMorph aux stream)
22

3-
Status: TRAINING IN FLIGHT (launched 2026-10-01 18:07, app ap-f1R8ChDikMKUGyBuJGO6;
4-
26,289 steps on A100-80GB; supervisor /tmp/r9-supervisor.sh v4 relaunches on
5-
kill-storm deaths — training is checkpoint-resumable and EVAL_DONE-idempotent)
3+
Status: CLOSED — VERIFIED NEGATIVE on both surfaces (2026-10-02); r7 stays canonical
64
Literature basis: YallaMorph (arXiv 2609.10153, EMNLP 2026) — 663,804
75
controlled morphological-generation instances over 4,795 lemmas,
86
constructed from CamelMorph MSA via CAMeL Tools. The GitHub repo ships
@@ -71,12 +69,41 @@ Recipe = r7 verbatim (train_arabic_r7.py) + one new aux stream:
7169
5. [x] Launch `modal run --detach` (+ supervisor; one double-launch
7270
incident from `modal app list` name truncation — grep prefix
7371
"rababa-ara", dupes stopped, volume verified clean).
74-
6. [ ] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
75-
≤ 2.389 (r7 2.2864 + 0.1 tolerance).
76-
7. [ ] OOD gate: eval_wikinews_multiref improves over 17.3794/11.8273.
77-
8. [ ] Canonical replacement only if ID improves outright; else record
78-
as ablation. Student distillation from r9 only after canonical
79-
call.
72+
6. [x] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
73+
≤ 2.389 (r7 2.2864 + 0.1 tolerance) — **FAILED: 2.4895**.
74+
7. [x] OOD gate: WikiNews-2024 multi-ref — **FAILED: 17.4265/12.1093**
75+
vs r7's 17.3794/11.8273 (worse on both axes; no trade).
76+
8. [x] Canonical call: **r7 stays canonical.** run-009 recorded as
77+
ablation. No student distillation from r9.
78+
79+
## Result (measured 2026-10-02) — NEGATIVE, first negative teacher-side result
80+
81+
| surface | r9 | r7 | verdict |
82+
|---|---|---|---|
83+
| SadeedDiac-25 Total DER (CE) | 2.4895 | 2.2864 | +0.20pp worse — gate fail |
84+
| Morph DER | 1.5054 | 1.3343 | worse |
85+
| WikiNews-2024 WER / DER | 17.4265 / 12.1093 | 17.3794 / 11.8273 | worse on both |
86+
87+
**Mechanism (data-backed):** the aux corpus's vocalization convention
88+
is far denser than the benchmark's — marks per 100 letters: fatha
89+
43.0 vs 28.4 (1.5×), damma 11.9 vs 7.1 (1.7×), shadda 8.6 vs 4.3
90+
(2.0×), tanwīn 0.3 vs 2.0 (nearly absent). CamelMorph paradigm forms
91+
are fully-explicit isolated words; SadeedDiac-style text is partially
92+
vocalized running text. At 25% dose, the aux stream taught convention
93+
drift, not morphology: the teacher over-marks the plain stream.
94+
95+
**Lesson (refines the knowledge-injection template):** r6's qalsadi
96+
aux worked because it was running text IN benchmark convention; r9's
97+
paradigm tables were out-of-context AND differently-conventioned.
98+
Knowledge injection is delivery-vehicle-sensitive. A future teacher
99+
lever from lexical resources must be rendered into benchmark-
100+
convention running text (e.g. lexically-guided text selection from
101+
the corpus, not paradigm tables).
102+
103+
The lineage: r5→r6→r7 all-positive; r9 first negative. Teacher-side
104+
levers are NOT closed (r7's news mix remains the proven axis), but
105+
this instantiation is dead: no re-run, no dose tuning, no convention
106+
normalization retry — recorded and closed.
80107

81108
## Gates / risks
82109

‎docs/RESULTS.md‎

Lines changed: 23 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1002,3 +1002,26 @@ Sinkhorn embeddings, engram memory, and representation displacement
10021002
(the only arm with a measured mechanistic premise) have all been run
10031003
to verdict. The frontier mover remains teacher-side data only —
10041004
run-009-yallamorph (TODO.sota-2026/01) is the active lever.
1005+
1006+
## run-009-yallamorph teacher — NEGATIVE on both surfaces; r7 stays canonical (2026-10-02)
1007+
1008+
The 2026-09-30 sweep's identified teacher-side lever (YallaMorph/
1009+
CamelMorph morphological paradigm aux, 25% dose, r7-init) measured
1010+
**worse on both surfaces**: SadeedDiac-25 windowed zero-skip DER
1011+
2.4895 (r7: 2.2864 — gate bar 2.389 failed) and WikiNews-2024
1012+
multi-ref 17.4265/12.1093 WER/DER (r7: 17.3794/11.8273) — no ID
1013+
improvement and no OOD trade. **r7 remains canonical**; run-009 is
1014+
recorded as the lineage's first negative teacher-side result.
1015+
1016+
Mechanism, data-backed: the paradigm corpus's vocalization convention
1017+
is 1.5–2.0× denser than benchmark text (fatha 43.0 vs 28.4, damma
1018+
11.9 vs 7.1, shadda 8.6 vs 4.3 marks per 100 letters; tanwīn nearly
1019+
absent) and its forms are isolated words rather than running text.
1020+
At 25% dose the aux stream injected convention drift — the teacher
1021+
over-marks the plain stream — instead of morphology. The knowledge-
1022+
injection template survives with a sharper edge: r6's aux worked as
1023+
running text in benchmark convention; delivery vehicle matters as
1024+
much as the knowledge. Paradigm-table aux is closed (no dose tuning,
1025+
no convention-normalization retry recorded as future work; a lexical
1026+
lever, if ever revisited, must be rendered into benchmark-convention
1027+
running text). The frontier mover on record remains the r7 news mix.

‎docs/paper.adoc‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -303,7 +303,7 @@ zip through the runtime itself — sha-pinned to the index) agree on
303303
re-derives exactly by the same tooling from the published prediction
304304
files, which is how the discrepancy was caught.
305305

306-
Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). One measurement caveat travels with all such arm verdicts: each is a single training seed, and a recent small-model distillation audit shows per-seed variance large enough to swallow sub-point deltas — with bimodal collapse in some KD variants — so our paired bootstrap's protection extends to prediction resampling but not the seed axis (Sumit et al. 2026, arXiv 2608.27729); future arms run multi-seed or carry this caveat. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. A final arm tested the residual in representation space, on the one mechanism whose premise we could measure before spending training compute: the teacher's news-domain fine-tune shifts its encoder representations in a direction that is domain-general in early and mid layers (cosine 0.59–0.95 between the shift measured on classical vs news text, layers 0–8), so we regressed the student's encoder hidden states toward ridge-projected, extrapolated teacher targets along that direction — and it was the *most* harmful lever of all (5.8627, +1.3926pp [1.156, 1.646] worse than its vanilla peer), despite healthy training and a pre-registered loss budget. A domain-general direction is necessary but not sufficient: pulling a 300M byte student's representations toward projected 580M targets displaces what its decoder relies on. With that, the residual is closed on evidence rather than exhaustion. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. The next teacher rung is in flight on exactly that axis: systematic morphological paradigm coverage from the YallaMorph/CamelMorph resource (Reda et al. 2026, arXiv 2609.10153 — 663,804 controlled morphological-generation instances) added as an auxiliary stream to the news-domain mix that produced the current rung.
306+
Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). One measurement caveat travels with all such arm verdicts: each is a single training seed, and a recent small-model distillation audit shows per-seed variance large enough to swallow sub-point deltas — with bimodal collapse in some KD variants — so our paired bootstrap's protection extends to prediction resampling but not the seed axis (Sumit et al. 2026, arXiv 2608.27729); future arms run multi-seed or carry this caveat. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. A final arm tested the residual in representation space, on the one mechanism whose premise we could measure before spending training compute: the teacher's news-domain fine-tune shifts its encoder representations in a direction that is domain-general in early and mid layers (cosine 0.59–0.95 between the shift measured on classical vs news text, layers 0–8), so we regressed the student's encoder hidden states toward ridge-projected, extrapolated teacher targets along that direction — and it was the *most* harmful lever of all (5.8627, +1.3926pp [1.156, 1.646] worse than its vanilla peer), despite healthy training and a pre-registered loss budget. A domain-general direction is necessary but not sufficient: pulling a 300M byte student's representations toward projected 580M targets displaces what its decoder relies on. With that, the residual is closed on evidence rather than exhaustion. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. That axis is not uniformly fertile: the next teacher rung — systematic morphological paradigm coverage from the YallaMorph/CamelMorph resource (Reda et al. 2026, arXiv 2609.10153) added as a 25% auxiliary stream to the news-domain mix — measured *worse* on both surfaces (in-domain 2.4895 vs 2.2864; out-of-domain 17.43/12.11 vs 17.38/11.83), the lineage's first negative teacher-side result. The mechanism is measurable: the paradigm corpus vocalizes 1.5–2.0× more densely than benchmark text (shadda 2.0×, damma 1.7×, tanwīn nearly absent), so a 25% dose injects convention drift rather than morphology. The knowledge-injection lever survives with a sharper edge — the r6 auxiliary stream worked as running text in benchmark convention; paradigm tables are out-of-context and differently-conventioned, and the delivery vehicle matters as much as the knowledge.
307307

308308
== The decode protocol is part of the measurement
309309
[[section-decode]]

0 commit comments

Comments
 (0)