diff --git a/TODO.sota-2026/01-r8-teacher-yallamorph.md b/TODO.sota-2026/01-r8-teacher-yallamorph.md new file mode 100644 index 0000000..2e5d812 --- /dev/null +++ b/TODO.sota-2026/01-r8-teacher-yallamorph.md @@ -0,0 +1,76 @@ +# 01 — r8 teacher: run-009-yallamorph (YallaMorph/CamelMorph aux stream) + +Status: SPECIFIED (2026-10-01) — not launched +Literature basis: YallaMorph (arXiv 2609.10153, EMNLP 2026) — 663,804 +controlled morphological-generation instances over 4,795 lemmas, +constructed from CamelMorph MSA via CAMeL Tools. The GitHub repo ships +**samples only**; the full benchmark is not downloadable as a corpus. +Since the underlying resource (CamelMorph MSA + CAMeL Tools) is public, +we regenerate the paradigm pairs ourselves — a deterministic +dictionary/morphology resource, NOT LLM-generated labels (guardrail +compliant, see [[no-llm-teacher-distillation]]). + +## Thesis + +Every frontier move on the Arabic teacher came from data-side levers +(r5 2.6775 → r6-morph-aux 2.5793 → r7-news 2.2864). Morph coverage was +the r6 lever; YallaMorph/CamelMorph gives systematic, +feature-complete morphological coverage (incl. cliticized forms — +exactly the iʿrāb/clitic band where the residual concentrates) on top +of r7's news mix. + +Naming: the `run-008` slot was consumed by the IPA-ablation +(train_arabic_r8.py → run-008-ipa). This teacher is +**run-009-yallamorph**; prose alias "r8-generation teacher". + +## Design + +Recipe = r7 verbatim (train_arabic_r7.py) + one new aux stream: + +- Stream A (plain): r7's mix unchanged — cached r5-units (anchor) + + news×3 + WikiNews-2014-gold×4. +- Stream M (aux, NEW): `MORPH: ` ASCII prefix (byte-distinct from + Arabic, mirrors the proven `TAG: ` mechanism), input = + `MORPH: | `, target = + ``. Features rendered as ASCII + key:value tokens (pos, aspect, state, case, ...). Multi-form + instances: one line per valid form. IMPOSSIBLE configs: excluded + (no target). +- Aux upsampled to ~25% of the mix (r6-proven dose). +- Init: run-007-news/best. Batch 2/accum 15, A100-80GB, bf16, 1 epoch, + save every 300 steps, volume commit on save (r5/r6/r7-proven). +- Inference contract unchanged: no prefix at inference = plain + diacritization. + +## Steps + +1. [ ] Clone CAMeL-Lab/YallaMorph; extract xlsx samples as the + validation set for our generated forms. +2. [ ] Data build: camel-tools + camel_data MSA; generate paradigm + pairs; validate forms against YallaMorph samples (match rate + reported); dedupe; cap 300k aux lines; volume put to + /datasets/yallamorph-aux/lines.txt. +3. [ ] TDD the line builder (pure function: feature dict → line pair). +4. [ ] train_arabic_r9_yallamorph.py (r7 copy + stream M + gates). +5. [ ] Launch `modal run --detach` (retry-loop supervisor per + [[modal-always-detach]]). +6. [ ] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER + ≤ 2.389 (r7 2.2864 + 0.1 tolerance). +7. [ ] OOD gate: eval_wikinews_multiref improves over 17.3794/11.8273. +8. [ ] Canonical replacement only if ID improves outright; else record + as ablation. Student distillation from r9 only after canonical + call. + +## Gates / risks + +- Aux format interference with the plain stream: mitigated by the + ASCII prefix mechanism (r6 evidence: deterministic at inference). +- CamelMorph license: record the underlying resource license in the + aux data dir README before any redistribution (training-internal use + only for now). +- Kill criterion: if ID gate fails after a clean run, close the lever + and keep r7 canonical (news mix remains the last mover). + +## Result + +(to be written only from measured numbers) diff --git a/TODO.sota-2026/02-paper-sweep-citations.md b/TODO.sota-2026/02-paper-sweep-citations.md new file mode 100644 index 0000000..628fb57 --- /dev/null +++ b/TODO.sota-2026/02-paper-sweep-citations.md @@ -0,0 +1,58 @@ +# 02 — Paper/docs: 2026-09-30 arXiv sweep citations + +Status: SPECIFIED (2026-10-01) + +Sweep window Aug 26 → Sep 30, 2026 (cs.CL/cs.LG: distillation, +byte-level, diacritization, Muon). Four findings change what the +papers should cite. Protocol rule stands: never quote cross-protocol +numbers. + +## Citations to add (paper.adoc + RESULTS.md) + +1. **arXiv 2609.12303 — "Breaking the Token Ceiling: Distilling + Smaller, Stronger Byte Models"** (Meta/FAIR). First large-scale + distillation × tokenization study (~1B params, up to 1T bytes): + distilled byte students start worse but surpass token students with + compute — higher asymptote (+4% predicted), 6× data efficiency, + 256-vocab avoids top-k logit truncation. External scaling-law + validation of our byte-student lineage (ByT5-style, byte+3 table). + Where: background/related work + one sentence in the frontier + discussion. +2. **arXiv 2609.37510 — MAESTRO ("From Dissonance to Orchestration")**. + Teacher intervention in on-policy distillation adds off-policy + load; always-on intervention yields diminishing returns; adaptive + disagreement-gated takeover is the fix. Cite as the mechanistic + account of our measured GKD negative (6.0036) — our always-on GKD + is the maximum-intervention point of their axis. Do NOT re-run + (ledger: student-side levers closed). +3. **arXiv 2608.27729 — "Below the Noise Floor"** (bimodal seed + collapse in small-model KD). Per-seed σ 2.8–48.7pp; single-seed KD + gains <5pp are unresolvable; 3/7 KD variants collapse bimodally. + Cite in the evaluation-methodology discussion: strengthens our + paired-bootstrap/full-set discipline AND adds the honest caveat + that our single-run arm verdicts (e.g. headwise Muon +0.2267pp) + measure prediction-resampled variance, not seed variance. +4. **arXiv 2609.10153 — YallaMorph** (EMNLP 2026). Cite as the r8 + teacher's data lever (morphological aux coverage) and as evidence + the field's Arabic-morphology energy moved to LLM evaluation, not + text-diacritization SOTA. + +## Also record (no citation needed) + +- No new text-only SadeedDiac-25 competitor in the window; KSAA-2026 + speech-diacritization winner (2605.25928, 23.26% WER) is a speech + modality — not protocol-comparable; noted in RESULTS.md sweep entry. + +## Steps + +1. [ ] RESULTS.md: "2026-09-30 sweep" entry (4 citations + the + no-new-competitor note). +2. [ ] paper.adoc: related-work sentences for 2609.12303 + MAESTRO; + methodology caveat sentence citing 2608.27729; YallaMorph in + the discussion where the r8 teacher lever is named. +3. [ ] PR to interscript/interscript-ml (branch off default, no AI + attribution, explicit-path staging). + +## Result + +(to be written only after merge) diff --git a/TODO.sota-2026/03-student-side-closed.md b/TODO.sota-2026/03-student-side-closed.md new file mode 100644 index 0000000..e563a7a --- /dev/null +++ b/TODO.sota-2026/03-student-side-closed.md @@ -0,0 +1,53 @@ +# 03 — Student-side lever family: CLOSED (record + caveat) + +Status: SPECIFIED (2026-10-01) + +## The record + +The student–teacher residual (2.1 student 4.5701 vs r7 teacher 2.2890) +has now resisted every lever measured: + +| lever | result | +|---|---| +| corpus scale / register mix (both directions) | negative / flat-negative | +| on-policy distillation (GKD) | negative (6.0036) | +| product-key memory layers | real but small (−0.70pp) | +| epochs 3→6 | −0.25pp (secondary) | +| Muon optimizer | −2.96pp (adopted — shipped in 2.0/2.1) | +| headwise Muon (671B-scale recipe) | separated-negative (+0.2267pp) | +| Sinkhorn embeddings | flat | +| lexical memory (engram) | flat | + +The 2026-09-30 literature sweep independently corroborates the +closure: the OPD wave (MAESTRO 2609.37510; RIDE 2609.36484; Fisher +sparsity 2609.36262; sparse supervision 2609.04565) targets reasoning +trajectories with distribution-shift mechanisms our deterministic +dense-label task does not have. No new student-side method in the +window contradicts the verdict. + +**Standing rule: no further GPU spend on student-side levers without a +pre-registered mechanism novel to the ledger.** The frontier mover on +record is teacher-side data (r5→r6→r7). Next frontier experiments: +01 (YallaMorph aux) and 05 (RIDE probe — the only student-side item +with a cheap kill-gated probe). + +## Caveat to publish (methodology honesty) + +Per arXiv 2608.27729: our paired between-students bootstrap measures +prediction-resampled variance, NOT training-seed variance. All arm +verdicts to date are single-seed runs. The headwise-Muon +separated-negative (+0.2267pp, p=0.017) is directionally consistent +with its size class, but the seed axis is unmeasured; future arms run +multi-seed or state the caveat. Ship decisions are unaffected (the +base recipe shipped regardless). + +## Steps + +1. [ ] RESULTS.md: lever-ledger closure entry + seed-variance caveat + (same PR as 02, separate commit). +2. [ ] paper.adoc: one sentence attaching the caveat to the + optimizer-recipe arm paragraph. + +## Result + +(to be written only after merge) diff --git a/TODO.sota-2026/04-stoicheia-diffusion-wo.md b/TODO.sota-2026/04-stoicheia-diffusion-wo.md new file mode 100644 index 0000000..7fccd4b --- /dev/null +++ b/TODO.sota-2026/04-stoicheia-diffusion-wo.md @@ -0,0 +1,42 @@ +# 04 — WO (queued, unscheduled): plane-factorized char-level masked diffusion + +Status: QUEUED (2026-10-01) — entry criteria at the bottom +Literature basis: Stoicheia (arXiv 2608.07249) — 405M character-level +masked-diffusion encoder for Ancient Greek; input factors into five +aligned, independently maskable planes (letters, boundaries, +**diacritics**, capitalization, punctuation). One backbone restores +lacunae / re-segments / **accentuates** / punctuates. Beats Ithaca +24.6 → 15.5 CER with matched random-init controls (methodology kin). + +## Why it is the highest-upside architecture item + +- Diacritics as an explicit maskable plane IS our task's factorization + (base letters given; only harakat positions are unpredictable). +- Non-autoregressive parallel decode = potential large CPU-latency win + over our KV-cache seq2seq (our runtime bottleneck). +- Independent planes compose tasks without retokenization — one model + could serve diacritization + our other normalization tasks. + +## Why it is queued, not scheduled + +- Ledger: architecture transfers measured negative at our scale + (depth cut, lexical memory, PKM; Hebrew depth-cut catastrophic). + Stoicheia's evidence is restoration/scansion at 405M with heavy + pretraining (380M words) — not a matched transfer case. +- New runtime contract: non-AR iterative decode does not fit IMF v1 + KV-cache graphs; needs its own export + parity path. +- Diffusion decode needs step-count/quality calibration per language. + +## Spec (if entered) + +1. Corpus: reuse Arabic combined + news + YallaMorph-aux; planes = + (base letters, harakat, word boundaries). Letters plane held + (skeleton-preserving, as our students already do). +2. Backbone: ~300M encoder, plane-aligned embeddings; pretrain + masked-plane objective, then SFT on diacritization. +3. Gates: same windowed zero-skip SadeedDiac harness; must beat + 4.5701 (student rung) AND 2.2864 (teacher rung) to matter; CPU + decode latency benchmarked against ara-diac-small-int8static-2.1. +4. Entry criteria: (a) 01 (run-009) lands and re-ranks the teacher + frontier; (b) owner authorizes a new architecture line; (c) a + decode-parity design exists for non-AR models (IMF v2 question). diff --git a/TODO.sota-2026/05-ride-sft-residual-extrapolation.md b/TODO.sota-2026/05-ride-sft-residual-extrapolation.md new file mode 100644 index 0000000..d72805c --- /dev/null +++ b/TODO.sota-2026/05-ride-sft-residual-extrapolation.md @@ -0,0 +1,53 @@ +# 05 — RIDE-style SFT-residual extrapolation: probe-first arm + +Status: SPECIFIED (2026-10-01) — probe only; training arm gated on probe +Literature basis: RIDE (arXiv 2609.36484) — extrapolate the +teacher-over-base residual directly in representation space: +student hidden states regressed toward +`h_target = h_teacher + λ·(h_teacher − h_base)`; approaches or exceeds +the teacher across four base/RL-teacher pairs. + +## Scope correction (user-confirmed 2026-10-01) + +The mechanism does NOT require an RL teacher. It needs any +(base, improved) checkpoint pair; the residual direction +`d = improved − base` is what is extrapolated. Our **r6→r7** SFT pair +(run-006-morph → run-007-news) qualifies. What remains forbidden is RL +*training* ([[rl-negative-diacritization]] — measured flat 3×), not +residual extrapolation of an SFT delta. + +## Why probe-first + +- Student-side lever (ledger: 8 negatives) — do not spend GPU on a + training arm before the direction is shown to transfer. +- The r7 delta is small (−0.29pp ID) and domain-shaped (news mix); + extrapolating a domain-idiosyncratic direction would amplify news + specialization, not general diacritization competence. + +## Probe (cheap: forward passes only, no training) + +1. Load run-006-morph/best and run-007-news/best (580M ByT5 each). +2. Forward N=200 units from two domains: SadeedDiac val paragraphs + (classical) + WikiNews-2024 text (news). Capture per-layer + mean-pooled encoder hidden states. +3. Per layer: d_classical = mean(h_r7) − mean(h_r6) on classical; + d_news likewise on news. Compute cos(d_classical, d_news). +4. **Kill criterion: max-layer cosine < 0.5 ⇒ direction is + domain-idiosyncratic ⇒ close the arm, record in RESULTS.md.** +5. Pass ⇒ full arm: hidden-state distillation with displacement + (λ ∈ {0.5, 1.0}) as an aux loss on the student trainer — requires + a new feature-regression path in modal_distill (spec before code; + TDD the loss on synthetic tensors). + +## Steps + +1. [ ] TDD pure computation: `residual_directions(h_base, h_teacher)` + and cosine sim on synthetic tensors (tests first, watch fail). +2. [ ] Modal probe script (two models × 200 units × 2 domains; + A100 minutes, not hours). +3. [ ] Run probe; write verdict + per-layer cosine table here. +4. [ ] Gate decision: close, or spec the training arm separately. + +## Result + +(to be written only from measured numbers) diff --git a/docs/RESULTS.md b/docs/RESULTS.md index b29284f..5845a7b 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -911,3 +911,64 @@ scale-boundary data point. The Sinkhorn-balanced embedding update remains the frontier student.** The TODO.impl recipe ledger is now fully measured: every optimizer/architecture lever is closed; the only frontier mover on record is data-side (teacher r5→r6→r7). + +## 2026-09-30 arXiv sweep — no new competitor; our premises externally validated + +Monthly sweep (window Aug 26 → Sep 30, 2026: distillation, byte-level +modeling, diacritization, optimizer literature). Competitive position +unchanged: **no new text-only Arabic diacritization system appeared on +SadeedDiac-25** — the field's Arabic-diacritization energy moved to the +speech modality (KSAA-2026 Task 2 winner, 23.26% WER, speech input: +not protocol-comparable). r7 (2.2864) remains the best dedicated model +measured under our protocol; only Claude-3.7-Sonnet's published 1.3941 +sits above it. + +Four findings enter the record: + +- **arXiv 2609.12303 (Meta/FAIR), "Breaking the Token Ceiling"** — + first large-scale distillation × tokenization study (~1B params, up + to 1T bytes): distilled *byte* students start worse but surpass + token students with compute (predicted +4% asymptote, 6× data + efficiency, 256-symbol vocab eliminates top-k logit truncation). + Independent scaling-law validation of the byte-student lineage we + ship. +- **arXiv 2609.37510, MAESTRO** — teacher intervention in on-policy + distillation injects off-policy load; always-on intervention is the + worst point of the axis. Mechanistic account of our measured GKD + negative (6.0036); corroborates closing the on-policy lever without + a re-run. +- **arXiv 2608.27729, "Below the Noise Floor"** — per-seed σ + 2.8–48.7pp in small-model KD; single-seed gains below ~5pp are + unresolvable; 3/7 KD variants collapse bimodally. Validates our + full-set + paired-bootstrap discipline and motivates the + seed-variance caveat recorded in the next entry. +- **arXiv 2609.10153, YallaMorph (EMNLP 2026)** — 663,804 controlled + Arabic morphological-generation instances (CamelMorph MSA). The + concrete teacher-side data lever for the next teacher rung + (TODO.sota-2026/01); also confirms the field's morphology work + targets LLM evaluation rather than text-diacritization SOTA. + +## Student-side lever family closed; seed-variance caveat recorded (2026-10-01) + +The residual ledger, complete: corpus scale ✗, register mix ✗ (both +directions), on-policy GKD ✗ (6.0036), PKM memory (real, −0.70pp), +epochs (−0.25pp), headwise Muon ✗ (separated-negative), Sinkhorn +embeddings ✗ (flat), engram lexical memory ✗ (flat). The 2026-09-30 +literature sweep surfaced no student-side method that escapes the +closure — the current on-policy wave (MAESTRO 2609.37510, RIDE +2609.36484, Fisher-sparsity 2609.36262, sparse supervision 2609.04565) +targets reasoning-trajectory distribution shift that a deterministic +dense-label task does not have. + +Standing rule: **no further GPU spend on student-side levers without a +pre-registered mechanism novel to this ledger.** Frontier experiments +continue teacher-side (TODO.sota-2026/01) and via the kill-gated RIDE +direction probe (TODO.sota-2026/05) — the only student-side item with +a cheap probe before any training compute. + +Caveat (per 2608.27729): every arm verdict above is a single training +seed; the paired between-students bootstrap resamples predictions, not +seeds. The headwise-Muon separated-negative (+0.2267pp, p=0.017) is +directionally consistent for its size class, but the seed axis is +unmeasured. Future arms run multi-seed or carry this caveat. Ship +decisions are unaffected — the base recipe shipped on its own merits. diff --git a/docs/paper.adoc b/docs/paper.adoc index 052034b..7bb9e0c 100644 --- a/docs/paper.adoc +++ b/docs/paper.adoc @@ -57,7 +57,7 @@ The OGC Abstract Standard for Interoperable Script Conversion Systems specifies === Sequence-to-sequence models for G2P and diacritization -Grapheme-to-phoneme conversion and diacritization are canonical seq2seq tasks. Character-level transformer models dominate recent literature; for the languages considered here, strong results have been reported with mT5-family encoders (e.g., umt5-based Thai G2P at 6.37% published PER) and ByT5 — the byte-level T5 variant that removes the tokenizer entirely, processing raw UTF-8 bytes. ByT5 is attractive for an artifact contract precisely because it eliminates vocabulary drift: the tokenizer is fixed by construction, so a model zip never needs a vocab file, and any runtime can encode input by table lookup (token id = byte + 3). +Grapheme-to-phoneme conversion and diacritization are canonical seq2seq tasks. Character-level transformer models dominate recent literature; for the languages considered here, strong results have been reported with mT5-family encoders (e.g., umt5-based Thai G2P at 6.37% published PER) and ByT5 — the byte-level T5 variant that removes the tokenizer entirely, processing raw UTF-8 bytes. ByT5 is attractive for an artifact contract precisely because it eliminates vocabulary drift: the tokenizer is fixed by construction, so a model zip never needs a vocab file, and any runtime can encode input by table lookup (token id = byte + 3). Recent scaling evidence supports the choice on quality grounds as well: the first large-scale distillation study over matched tokenization schemes finds that distilled *byte* students start behind token students at low compute but reach a higher asymptotic ceiling — up to +4% predicted on averaged downstream tasks at roughly one-sixth the training data — with the 256-symbol vocabulary eliminating top-k logit truncation entirely (Marathe et al. 2026, arXiv 2609.12303). The 2024–2025 landscape consolidated around two directions this work deliberately does not follow. First, *LLM prompting as the conversion engine*: Fetrat et al. (2024) show prompted LLMs with post-processing reach 8.30% PER on their Persian G2P benchmark, and subsequent work constructs intermediate transliteration languages trained on LLM-generated data (Bertina et al. 2025). We exclude LLM-generated supervision on reliability grounds — language models hallucinate diacritics — but note these systems as evidence that the tasks are reachable by many architectures. Second, *benchmark consolidation and protocol sensitivity*: SadeedDiac-25 (Aldallal et al. 2025) was introduced precisely because leaderboard numbers were not comparable across papers; Mohamed & Mubarak (2025) propose multi-reference evaluation (WikiNews-2024) for the same reason. Our methodology — one harness per language, protocol-matched external comparisons only, zero skipped examples — adopts the same discipline at the artifact level. On the model side, small decoder-only LMs (Sadeed, 1.5B) and LSTM/BERT hybrids (D-Nikud for Hebrew; Rosenthal & Shaked 2024) report strong in-domain results; none ship as portable artifacts, which is the gap this paper fills. @@ -65,7 +65,7 @@ The 2024–2025 landscape consolidated around two directions this work deliberat We apply two standard distillation regimes: *sequence-level* distillation, where the teacher generates target strings consumed by the student's standard cross-entropy objective, and *logit* distillation (KL between teacher and student distributions plus ground-truth CE). Sequence-level KD is the natural fit when teacher and student occupy different vocabulary spaces (a sentencepiece umt5 teacher into a byte-level student); logit KD when both are byte-level and share the tokenizer. -Both regimes are *off-policy*: the training distribution is the teacher's (or the corpus's), never the student's own. The on-policy line — MiniLLM's reverse KL on student-sampled sequences (Gu et al. 2024) and Generalized Knowledge Distillation (Agarwal et al. 2024), which trains on student-generated mistakes scored by the teacher — is designed to fix exactly the pathology we measure in <>: students whose distributions are too flat to rank completions by likelihood. Sharpening the student on its own errors, rather than only imitating teacher output, is the principled next lever for both the shrink cost and the decode behavior; its measured outcome is recorded in <>. +Both regimes are *off-policy*: the training distribution is the teacher's (or the corpus's), never the student's own. The on-policy line — MiniLLM's reverse KL on student-sampled sequences (Gu et al. 2024) and Generalized Knowledge Distillation (Agarwal et al. 2024), which trains on student-generated mistakes scored by the teacher — is designed to fix exactly the pathology we measure in <>: students whose distributions are too flat to rank completions by likelihood. Sharpening the student on its own errors, rather than only imitating teacher output, is the principled next lever for both the shrink cost and the decode behavior; its measured outcome is recorded in <>. Subsequent controlled analysis explains why that outcome was to be expected: teacher intervention in on-policy distillation trades rollout improvement against off-policy load, with always-on intervention the worst point of the axis (Wang et al. 2026, arXiv 2609.37510) — and our setting, where the teacher's supervision is deterministic at every position, sits at exactly that maximum-intervention end. === Quantization and deployment-aware inference @@ -303,7 +303,7 @@ zip through the runtime itself — sha-pinned to the index) agree on re-derives exactly by the same tooling from the published prediction files, which is how the discrepancy was caught. -Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. +Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). One measurement caveat travels with all such arm verdicts: each is a single training seed, and a recent small-model distillation audit shows per-seed variance large enough to swallow sub-point deltas — with bimodal collapse in some KD variants — so our paired bootstrap's protection extends to prediction resampling but not the seed axis (Sumit et al. 2026, arXiv 2608.27729); future arms run multi-seed or carry this caveat. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. The next teacher rung is in flight on exactly that axis: systematic morphological paradigm coverage from the YallaMorph/CamelMorph resource (Reda et al. 2026, arXiv 2609.10153 — 663,804 controlled morphological-generation instances) added as an auxiliary stream to the news-domain mix that produced the current rung. == The decode protocol is part of the measurement [[section-decode]]