Skip to content

Commit 4808258

Browse files
author
Ronald Tse
committed
docs(todos): TODO.sota-2026 - r9-yallamorph teacher spec, sweep citations record, closure rule, diffusion WO queue, RIDE probe spec
1 parent 290ad81 commit 4808258

5 files changed

Lines changed: 282 additions & 0 deletions
Lines changed: 76 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,76 @@
1+
# 01 — r8 teacher: run-009-yallamorph (YallaMorph/CamelMorph aux stream)
2+
3+
Status: SPECIFIED (2026-10-01) — not launched
4+
Literature basis: YallaMorph (arXiv 2609.10153, EMNLP 2026) — 663,804
5+
controlled morphological-generation instances over 4,795 lemmas,
6+
constructed from CamelMorph MSA via CAMeL Tools. The GitHub repo ships
7+
**samples only**; the full benchmark is not downloadable as a corpus.
8+
Since the underlying resource (CamelMorph MSA + CAMeL Tools) is public,
9+
we regenerate the paradigm pairs ourselves — a deterministic
10+
dictionary/morphology resource, NOT LLM-generated labels (guardrail
11+
compliant, see [[no-llm-teacher-distillation]]).
12+
13+
## Thesis
14+
15+
Every frontier move on the Arabic teacher came from data-side levers
16+
(r5 2.6775 → r6-morph-aux 2.5793 → r7-news 2.2864). Morph coverage was
17+
the r6 lever; YallaMorph/CamelMorph gives systematic,
18+
feature-complete morphological coverage (incl. cliticized forms —
19+
exactly the iʿrāb/clitic band where the residual concentrates) on top
20+
of r7's news mix.
21+
22+
Naming: the `run-008` slot was consumed by the IPA-ablation
23+
(train_arabic_r8.py → run-008-ipa). This teacher is
24+
**run-009-yallamorph**; prose alias "r8-generation teacher".
25+
26+
## Design
27+
28+
Recipe = r7 verbatim (train_arabic_r7.py) + one new aux stream:
29+
30+
- Stream A (plain): r7's mix unchanged — cached r5-units (anchor) +
31+
news×3 + WikiNews-2014-gold×4.
32+
- Stream M (aux, NEW): `MORPH: ` ASCII prefix (byte-distinct from
33+
Arabic, mirrors the proven `TAG: ` mechanism), input =
34+
`MORPH: <features-as-ascii> | <undiacritized form>`, target =
35+
`<fully diacritized form>`. Features rendered as ASCII
36+
key:value tokens (pos, aspect, state, case, ...). Multi-form
37+
instances: one line per valid form. IMPOSSIBLE configs: excluded
38+
(no target).
39+
- Aux upsampled to ~25% of the mix (r6-proven dose).
40+
- Init: run-007-news/best. Batch 2/accum 15, A100-80GB, bf16, 1 epoch,
41+
save every 300 steps, volume commit on save (r5/r6/r7-proven).
42+
- Inference contract unchanged: no prefix at inference = plain
43+
diacritization.
44+
45+
## Steps
46+
47+
1. [ ] Clone CAMeL-Lab/YallaMorph; extract xlsx samples as the
48+
validation set for our generated forms.
49+
2. [ ] Data build: camel-tools + camel_data MSA; generate paradigm
50+
pairs; validate forms against YallaMorph samples (match rate
51+
reported); dedupe; cap 300k aux lines; volume put to
52+
/datasets/yallamorph-aux/lines.txt.
53+
3. [ ] TDD the line builder (pure function: feature dict → line pair).
54+
4. [ ] train_arabic_r9_yallamorph.py (r7 copy + stream M + gates).
55+
5. [ ] Launch `modal run --detach` (retry-loop supervisor per
56+
[[modal-always-detach]]).
57+
6. [ ] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
58+
≤ 2.389 (r7 2.2864 + 0.1 tolerance).
59+
7. [ ] OOD gate: eval_wikinews_multiref improves over 17.3794/11.8273.
60+
8. [ ] Canonical replacement only if ID improves outright; else record
61+
as ablation. Student distillation from r9 only after canonical
62+
call.
63+
64+
## Gates / risks
65+
66+
- Aux format interference with the plain stream: mitigated by the
67+
ASCII prefix mechanism (r6 evidence: deterministic at inference).
68+
- CamelMorph license: record the underlying resource license in the
69+
aux data dir README before any redistribution (training-internal use
70+
only for now).
71+
- Kill criterion: if ID gate fails after a clean run, close the lever
72+
and keep r7 canonical (news mix remains the last mover).
73+
74+
## Result
75+
76+
(to be written only from measured numbers)
Lines changed: 58 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,58 @@
1+
# 02 — Paper/docs: 2026-09-30 arXiv sweep citations
2+
3+
Status: SPECIFIED (2026-10-01)
4+
5+
Sweep window Aug 26 → Sep 30, 2026 (cs.CL/cs.LG: distillation,
6+
byte-level, diacritization, Muon). Four findings change what the
7+
papers should cite. Protocol rule stands: never quote cross-protocol
8+
numbers.
9+
10+
## Citations to add (paper.adoc + RESULTS.md)
11+
12+
1. **arXiv 2609.12303 — "Breaking the Token Ceiling: Distilling
13+
Smaller, Stronger Byte Models"** (Meta/FAIR). First large-scale
14+
distillation × tokenization study (~1B params, up to 1T bytes):
15+
distilled byte students start worse but surpass token students with
16+
compute — higher asymptote (+4% predicted), 6× data efficiency,
17+
256-vocab avoids top-k logit truncation. External scaling-law
18+
validation of our byte-student lineage (ByT5-style, byte+3 table).
19+
Where: background/related work + one sentence in the frontier
20+
discussion.
21+
2. **arXiv 2609.37510 — MAESTRO ("From Dissonance to Orchestration")**.
22+
Teacher intervention in on-policy distillation adds off-policy
23+
load; always-on intervention yields diminishing returns; adaptive
24+
disagreement-gated takeover is the fix. Cite as the mechanistic
25+
account of our measured GKD negative (6.0036) — our always-on GKD
26+
is the maximum-intervention point of their axis. Do NOT re-run
27+
(ledger: student-side levers closed).
28+
3. **arXiv 2608.27729 — "Below the Noise Floor"** (bimodal seed
29+
collapse in small-model KD). Per-seed σ 2.8–48.7pp; single-seed KD
30+
gains <5pp are unresolvable; 3/7 KD variants collapse bimodally.
31+
Cite in the evaluation-methodology discussion: strengthens our
32+
paired-bootstrap/full-set discipline AND adds the honest caveat
33+
that our single-run arm verdicts (e.g. headwise Muon +0.2267pp)
34+
measure prediction-resampled variance, not seed variance.
35+
4. **arXiv 2609.10153 — YallaMorph** (EMNLP 2026). Cite as the r8
36+
teacher's data lever (morphological aux coverage) and as evidence
37+
the field's Arabic-morphology energy moved to LLM evaluation, not
38+
text-diacritization SOTA.
39+
40+
## Also record (no citation needed)
41+
42+
- No new text-only SadeedDiac-25 competitor in the window; KSAA-2026
43+
speech-diacritization winner (2605.25928, 23.26% WER) is a speech
44+
modality — not protocol-comparable; noted in RESULTS.md sweep entry.
45+
46+
## Steps
47+
48+
1. [ ] RESULTS.md: "2026-09-30 sweep" entry (4 citations + the
49+
no-new-competitor note).
50+
2. [ ] paper.adoc: related-work sentences for 2609.12303 + MAESTRO;
51+
methodology caveat sentence citing 2608.27729; YallaMorph in
52+
the discussion where the r8 teacher lever is named.
53+
3. [ ] PR to interscript/interscript-ml (branch off default, no AI
54+
attribution, explicit-path staging).
55+
56+
## Result
57+
58+
(to be written only after merge)
Lines changed: 53 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,53 @@
1+
# 03 — Student-side lever family: CLOSED (record + caveat)
2+
3+
Status: SPECIFIED (2026-10-01)
4+
5+
## The record
6+
7+
The student–teacher residual (2.1 student 4.5701 vs r7 teacher 2.2890)
8+
has now resisted every lever measured:
9+
10+
| lever | result |
11+
|---|---|
12+
| corpus scale / register mix (both directions) | negative / flat-negative |
13+
| on-policy distillation (GKD) | negative (6.0036) |
14+
| product-key memory layers | real but small (−0.70pp) |
15+
| epochs 3→6 | −0.25pp (secondary) |
16+
| Muon optimizer | −2.96pp (adopted — shipped in 2.0/2.1) |
17+
| headwise Muon (671B-scale recipe) | separated-negative (+0.2267pp) |
18+
| Sinkhorn embeddings | flat |
19+
| lexical memory (engram) | flat |
20+
21+
The 2026-09-30 literature sweep independently corroborates the
22+
closure: the OPD wave (MAESTRO 2609.37510; RIDE 2609.36484; Fisher
23+
sparsity 2609.36262; sparse supervision 2609.04565) targets reasoning
24+
trajectories with distribution-shift mechanisms our deterministic
25+
dense-label task does not have. No new student-side method in the
26+
window contradicts the verdict.
27+
28+
**Standing rule: no further GPU spend on student-side levers without a
29+
pre-registered mechanism novel to the ledger.** The frontier mover on
30+
record is teacher-side data (r5→r6→r7). Next frontier experiments:
31+
01 (YallaMorph aux) and 05 (RIDE probe — the only student-side item
32+
with a cheap kill-gated probe).
33+
34+
## Caveat to publish (methodology honesty)
35+
36+
Per arXiv 2608.27729: our paired between-students bootstrap measures
37+
prediction-resampled variance, NOT training-seed variance. All arm
38+
verdicts to date are single-seed runs. The headwise-Muon
39+
separated-negative (+0.2267pp, p=0.017) is directionally consistent
40+
with its size class, but the seed axis is unmeasured; future arms run
41+
multi-seed or state the caveat. Ship decisions are unaffected (the
42+
base recipe shipped regardless).
43+
44+
## Steps
45+
46+
1. [ ] RESULTS.md: lever-ledger closure entry + seed-variance caveat
47+
(same PR as 02, separate commit).
48+
2. [ ] paper.adoc: one sentence attaching the caveat to the
49+
optimizer-recipe arm paragraph.
50+
51+
## Result
52+
53+
(to be written only after merge)
Lines changed: 42 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,42 @@
1+
# 04 — WO (queued, unscheduled): plane-factorized char-level masked diffusion
2+
3+
Status: QUEUED (2026-10-01) — entry criteria at the bottom
4+
Literature basis: Stoicheia (arXiv 2608.07249) — 405M character-level
5+
masked-diffusion encoder for Ancient Greek; input factors into five
6+
aligned, independently maskable planes (letters, boundaries,
7+
**diacritics**, capitalization, punctuation). One backbone restores
8+
lacunae / re-segments / **accentuates** / punctuates. Beats Ithaca
9+
24.6 → 15.5 CER with matched random-init controls (methodology kin).
10+
11+
## Why it is the highest-upside architecture item
12+
13+
- Diacritics as an explicit maskable plane IS our task's factorization
14+
(base letters given; only harakat positions are unpredictable).
15+
- Non-autoregressive parallel decode = potential large CPU-latency win
16+
over our KV-cache seq2seq (our runtime bottleneck).
17+
- Independent planes compose tasks without retokenization — one model
18+
could serve diacritization + our other normalization tasks.
19+
20+
## Why it is queued, not scheduled
21+
22+
- Ledger: architecture transfers measured negative at our scale
23+
(depth cut, lexical memory, PKM; Hebrew depth-cut catastrophic).
24+
Stoicheia's evidence is restoration/scansion at 405M with heavy
25+
pretraining (380M words) — not a matched transfer case.
26+
- New runtime contract: non-AR iterative decode does not fit IMF v1
27+
KV-cache graphs; needs its own export + parity path.
28+
- Diffusion decode needs step-count/quality calibration per language.
29+
30+
## Spec (if entered)
31+
32+
1. Corpus: reuse Arabic combined + news + YallaMorph-aux; planes =
33+
(base letters, harakat, word boundaries). Letters plane held
34+
(skeleton-preserving, as our students already do).
35+
2. Backbone: ~300M encoder, plane-aligned embeddings; pretrain
36+
masked-plane objective, then SFT on diacritization.
37+
3. Gates: same windowed zero-skip SadeedDiac harness; must beat
38+
4.5701 (student rung) AND 2.2864 (teacher rung) to matter; CPU
39+
decode latency benchmarked against ara-diac-small-int8static-2.1.
40+
4. Entry criteria: (a) 01 (run-009) lands and re-ranks the teacher
41+
frontier; (b) owner authorizes a new architecture line; (c) a
42+
decode-parity design exists for non-AR models (IMF v2 question).
Lines changed: 53 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,53 @@
1+
# 05 — RIDE-style SFT-residual extrapolation: probe-first arm
2+
3+
Status: SPECIFIED (2026-10-01) — probe only; training arm gated on probe
4+
Literature basis: RIDE (arXiv 2609.36484) — extrapolate the
5+
teacher-over-base residual directly in representation space:
6+
student hidden states regressed toward
7+
`h_target = h_teacher + λ·(h_teacher − h_base)`; approaches or exceeds
8+
the teacher across four base/RL-teacher pairs.
9+
10+
## Scope correction (user-confirmed 2026-10-01)
11+
12+
The mechanism does NOT require an RL teacher. It needs any
13+
(base, improved) checkpoint pair; the residual direction
14+
`d = improved − base` is what is extrapolated. Our **r6→r7** SFT pair
15+
(run-006-morph → run-007-news) qualifies. What remains forbidden is RL
16+
*training* ([[rl-negative-diacritization]] — measured flat 3×), not
17+
residual extrapolation of an SFT delta.
18+
19+
## Why probe-first
20+
21+
- Student-side lever (ledger: 8 negatives) — do not spend GPU on a
22+
training arm before the direction is shown to transfer.
23+
- The r7 delta is small (−0.29pp ID) and domain-shaped (news mix);
24+
extrapolating a domain-idiosyncratic direction would amplify news
25+
specialization, not general diacritization competence.
26+
27+
## Probe (cheap: forward passes only, no training)
28+
29+
1. Load run-006-morph/best and run-007-news/best (580M ByT5 each).
30+
2. Forward N=200 units from two domains: SadeedDiac val paragraphs
31+
(classical) + WikiNews-2024 text (news). Capture per-layer
32+
mean-pooled encoder hidden states.
33+
3. Per layer: d_classical = mean(h_r7) − mean(h_r6) on classical;
34+
d_news likewise on news. Compute cos(d_classical, d_news).
35+
4. **Kill criterion: max-layer cosine < 0.5 ⇒ direction is
36+
domain-idiosyncratic ⇒ close the arm, record in RESULTS.md.**
37+
5. Pass ⇒ full arm: hidden-state distillation with displacement
38+
(λ ∈ {0.5, 1.0}) as an aux loss on the student trainer — requires
39+
a new feature-regression path in modal_distill (spec before code;
40+
TDD the loss on synthetic tensors).
41+
42+
## Steps
43+
44+
1. [ ] TDD pure computation: `residual_directions(h_base, h_teacher)`
45+
and cosine sim on synthetic tensors (tests first, watch fail).
46+
2. [ ] Modal probe script (two models × 200 units × 2 domains;
47+
A100 minutes, not hours).
48+
3. [ ] Run probe; write verdict + per-layer cosine table here.
49+
4. [ ] Gate decision: close, or spec the training arm separately.
50+
51+
## Result
52+
53+
(to be written only from measured numbers)

0 commit comments

Comments
 (0)