Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 36 additions & 9 deletions TODO.sota-2026/01-r8-teacher-yallamorph.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,6 @@
# 01 — r8 teacher: run-009-yallamorph (YallaMorph/CamelMorph aux stream)

Status: TRAINING IN FLIGHT (launched 2026-10-01 18:07, app ap-f1R8ChDikMKUGyBuJGO6;
26,289 steps on A100-80GB; supervisor /tmp/r9-supervisor.sh v4 relaunches on
kill-storm deaths — training is checkpoint-resumable and EVAL_DONE-idempotent)
Status: CLOSED — VERIFIED NEGATIVE on both surfaces (2026-10-02); r7 stays canonical
Literature basis: YallaMorph (arXiv 2609.10153, EMNLP 2026) — 663,804
controlled morphological-generation instances over 4,795 lemmas,
constructed from CamelMorph MSA via CAMeL Tools. The GitHub repo ships
Expand Down Expand Up @@ -71,12 +69,41 @@ Recipe = r7 verbatim (train_arabic_r7.py) + one new aux stream:
5. [x] Launch `modal run --detach` (+ supervisor; one double-launch
incident from `modal app list` name truncation — grep prefix
"rababa-ara", dupes stopped, volume verified clean).
6. [ ] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
≤ 2.389 (r7 2.2864 + 0.1 tolerance).
7. [ ] OOD gate: eval_wikinews_multiref improves over 17.3794/11.8273.
8. [ ] Canonical replacement only if ID improves outright; else record
as ablation. Student distillation from r9 only after canonical
call.
6. [x] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
≤ 2.389 (r7 2.2864 + 0.1 tolerance) — **FAILED: 2.4895**.
7. [x] OOD gate: WikiNews-2024 multi-ref — **FAILED: 17.4265/12.1093**
vs r7's 17.3794/11.8273 (worse on both axes; no trade).
8. [x] Canonical call: **r7 stays canonical.** run-009 recorded as
ablation. No student distillation from r9.

## Result (measured 2026-10-02) — NEGATIVE, first negative teacher-side result

| surface | r9 | r7 | verdict |
|---|---|---|---|
| SadeedDiac-25 Total DER (CE) | 2.4895 | 2.2864 | +0.20pp worse — gate fail |
| Morph DER | 1.5054 | 1.3343 | worse |
| WikiNews-2024 WER / DER | 17.4265 / 12.1093 | 17.3794 / 11.8273 | worse on both |

**Mechanism (data-backed):** the aux corpus's vocalization convention
is far denser than the benchmark's — marks per 100 letters: fatha
43.0 vs 28.4 (1.5×), damma 11.9 vs 7.1 (1.7×), shadda 8.6 vs 4.3
(2.0×), tanwīn 0.3 vs 2.0 (nearly absent). CamelMorph paradigm forms
are fully-explicit isolated words; SadeedDiac-style text is partially
vocalized running text. At 25% dose, the aux stream taught convention
drift, not morphology: the teacher over-marks the plain stream.

**Lesson (refines the knowledge-injection template):** r6's qalsadi
aux worked because it was running text IN benchmark convention; r9's
paradigm tables were out-of-context AND differently-conventioned.
Knowledge injection is delivery-vehicle-sensitive. A future teacher
lever from lexical resources must be rendered into benchmark-
convention running text (e.g. lexically-guided text selection from
the corpus, not paradigm tables).

The lineage: r5→r6→r7 all-positive; r9 first negative. Teacher-side
levers are NOT closed (r7's news mix remains the proven axis), but
this instantiation is dead: no re-run, no dose tuning, no convention
normalization retry — recorded and closed.

## Gates / risks

Expand Down
23 changes: 23 additions & 0 deletions docs/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -1002,3 +1002,26 @@ Sinkhorn embeddings, engram memory, and representation displacement
(the only arm with a measured mechanistic premise) have all been run
to verdict. The frontier mover remains teacher-side data only —
run-009-yallamorph (TODO.sota-2026/01) is the active lever.

## run-009-yallamorph teacher — NEGATIVE on both surfaces; r7 stays canonical (2026-10-02)

The 2026-09-30 sweep's identified teacher-side lever (YallaMorph/
CamelMorph morphological paradigm aux, 25% dose, r7-init) measured
**worse on both surfaces**: SadeedDiac-25 windowed zero-skip DER
2.4895 (r7: 2.2864 — gate bar 2.389 failed) and WikiNews-2024
multi-ref 17.4265/12.1093 WER/DER (r7: 17.3794/11.8273) — no ID
improvement and no OOD trade. **r7 remains canonical**; run-009 is
recorded as the lineage's first negative teacher-side result.

Mechanism, data-backed: the paradigm corpus's vocalization convention
is 1.5–2.0× denser than benchmark text (fatha 43.0 vs 28.4, damma
11.9 vs 7.1, shadda 8.6 vs 4.3 marks per 100 letters; tanwīn nearly
absent) and its forms are isolated words rather than running text.
At 25% dose the aux stream injected convention drift — the teacher
over-marks the plain stream — instead of morphology. The knowledge-
injection template survives with a sharper edge: r6's aux worked as
running text in benchmark convention; delivery vehicle matters as
much as the knowledge. Paradigm-table aux is closed (no dose tuning,
no convention-normalization retry recorded as future work; a lexical
lever, if ever revisited, must be rendered into benchmark-convention
running text). The frontier mover on record remains the r7 news mix.
2 changes: 1 addition & 1 deletion docs/paper.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -303,7 +303,7 @@ zip through the runtime itself — sha-pinned to the index) agree on
re-derives exactly by the same tooling from the published prediction
files, which is how the discrepancy was caught.

Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). One measurement caveat travels with all such arm verdicts: each is a single training seed, and a recent small-model distillation audit shows per-seed variance large enough to swallow sub-point deltas — with bimodal collapse in some KD variants — so our paired bootstrap's protection extends to prediction resampling but not the seed axis (Sumit et al. 2026, arXiv 2608.27729); future arms run multi-seed or carry this caveat. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. A final arm tested the residual in representation space, on the one mechanism whose premise we could measure before spending training compute: the teacher's news-domain fine-tune shifts its encoder representations in a direction that is domain-general in early and mid layers (cosine 0.59–0.95 between the shift measured on classical vs news text, layers 0–8), so we regressed the student's encoder hidden states toward ridge-projected, extrapolated teacher targets along that direction — and it was the *most* harmful lever of all (5.8627, +1.3926pp [1.156, 1.646] worse than its vanilla peer), despite healthy training and a pre-registered loss budget. A domain-general direction is necessary but not sufficient: pulling a 300M byte student's representations toward projected 580M targets displaces what its decoder relies on. With that, the residual is closed on evidence rather than exhaustion. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. The next teacher rung is in flight on exactly that axis: systematic morphological paradigm coverage from the YallaMorph/CamelMorph resource (Reda et al. 2026, arXiv 2609.10153 — 663,804 controlled morphological-generation instances) added as an auxiliary stream to the news-domain mix that produced the current rung.
Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). One measurement caveat travels with all such arm verdicts: each is a single training seed, and a recent small-model distillation audit shows per-seed variance large enough to swallow sub-point deltas — with bimodal collapse in some KD variants — so our paired bootstrap's protection extends to prediction resampling but not the seed axis (Sumit et al. 2026, arXiv 2608.27729); future arms run multi-seed or carry this caveat. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. A final arm tested the residual in representation space, on the one mechanism whose premise we could measure before spending training compute: the teacher's news-domain fine-tune shifts its encoder representations in a direction that is domain-general in early and mid layers (cosine 0.59–0.95 between the shift measured on classical vs news text, layers 0–8), so we regressed the student's encoder hidden states toward ridge-projected, extrapolated teacher targets along that direction — and it was the *most* harmful lever of all (5.8627, +1.3926pp [1.156, 1.646] worse than its vanilla peer), despite healthy training and a pre-registered loss budget. A domain-general direction is necessary but not sufficient: pulling a 300M byte student's representations toward projected 580M targets displaces what its decoder relies on. With that, the residual is closed on evidence rather than exhaustion. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. That axis is not uniformly fertile: the next teacher rung — systematic morphological paradigm coverage from the YallaMorph/CamelMorph resource (Reda et al. 2026, arXiv 2609.10153) added as a 25% auxiliary stream to the news-domain mix — measured *worse* on both surfaces (in-domain 2.4895 vs 2.2864; out-of-domain 17.43/12.11 vs 17.38/11.83), the lineage's first negative teacher-side result. The mechanism is measurable: the paradigm corpus vocalizes 1.5–2.0× more densely than benchmark text (shadda 2.0×, damma 1.7×, tanwīn nearly absent), so a 25% dose injects convention drift rather than morphology. The knowledge-injection lever survives with a sharper edge — the r6 auxiliary stream worked as running text in benchmark convention; paradigm tables are out-of-context and differently-conventioned, and the delivery vehicle matters as much as the knowledge.

== The decode protocol is part of the measurement
[[section-decode]]
Expand Down
Loading