Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
76 changes: 76 additions & 0 deletions TODO.sota-2026/01-r8-teacher-yallamorph.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
# 01 — r8 teacher: run-009-yallamorph (YallaMorph/CamelMorph aux stream)

Status: SPECIFIED (2026-10-01) — not launched
Literature basis: YallaMorph (arXiv 2609.10153, EMNLP 2026) — 663,804
controlled morphological-generation instances over 4,795 lemmas,
constructed from CamelMorph MSA via CAMeL Tools. The GitHub repo ships
**samples only**; the full benchmark is not downloadable as a corpus.
Since the underlying resource (CamelMorph MSA + CAMeL Tools) is public,
we regenerate the paradigm pairs ourselves — a deterministic
dictionary/morphology resource, NOT LLM-generated labels (guardrail
compliant, see [[no-llm-teacher-distillation]]).

## Thesis

Every frontier move on the Arabic teacher came from data-side levers
(r5 2.6775 → r6-morph-aux 2.5793 → r7-news 2.2864). Morph coverage was
the r6 lever; YallaMorph/CamelMorph gives systematic,
feature-complete morphological coverage (incl. cliticized forms —
exactly the iʿrāb/clitic band where the residual concentrates) on top
of r7's news mix.

Naming: the `run-008` slot was consumed by the IPA-ablation
(train_arabic_r8.py → run-008-ipa). This teacher is
**run-009-yallamorph**; prose alias "r8-generation teacher".

## Design

Recipe = r7 verbatim (train_arabic_r7.py) + one new aux stream:

- Stream A (plain): r7's mix unchanged — cached r5-units (anchor) +
news×3 + WikiNews-2014-gold×4.
- Stream M (aux, NEW): `MORPH: ` ASCII prefix (byte-distinct from
Arabic, mirrors the proven `TAG: ` mechanism), input =
`MORPH: <features-as-ascii> | <undiacritized form>`, target =
`<fully diacritized form>`. Features rendered as ASCII
key:value tokens (pos, aspect, state, case, ...). Multi-form
instances: one line per valid form. IMPOSSIBLE configs: excluded
(no target).
- Aux upsampled to ~25% of the mix (r6-proven dose).
- Init: run-007-news/best. Batch 2/accum 15, A100-80GB, bf16, 1 epoch,
save every 300 steps, volume commit on save (r5/r6/r7-proven).
- Inference contract unchanged: no prefix at inference = plain
diacritization.

## Steps

1. [ ] Clone CAMeL-Lab/YallaMorph; extract xlsx samples as the
validation set for our generated forms.
2. [ ] Data build: camel-tools + camel_data MSA; generate paradigm
pairs; validate forms against YallaMorph samples (match rate
reported); dedupe; cap 300k aux lines; volume put to
/datasets/yallamorph-aux/lines.txt.
3. [ ] TDD the line builder (pure function: feature dict → line pair).
4. [ ] train_arabic_r9_yallamorph.py (r7 copy + stream M + gates).
5. [ ] Launch `modal run --detach` (retry-loop supervisor per
[[modal-always-detach]]).
6. [ ] ID gate: windowed zero-skip SadeedDiac-25 full 1,200-para DER
≤ 2.389 (r7 2.2864 + 0.1 tolerance).
7. [ ] OOD gate: eval_wikinews_multiref improves over 17.3794/11.8273.
8. [ ] Canonical replacement only if ID improves outright; else record
as ablation. Student distillation from r9 only after canonical
call.

## Gates / risks

- Aux format interference with the plain stream: mitigated by the
ASCII prefix mechanism (r6 evidence: deterministic at inference).
- CamelMorph license: record the underlying resource license in the
aux data dir README before any redistribution (training-internal use
only for now).
- Kill criterion: if ID gate fails after a clean run, close the lever
and keep r7 canonical (news mix remains the last mover).

## Result

(to be written only from measured numbers)
58 changes: 58 additions & 0 deletions TODO.sota-2026/02-paper-sweep-citations.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# 02 — Paper/docs: 2026-09-30 arXiv sweep citations

Status: SPECIFIED (2026-10-01)

Sweep window Aug 26 → Sep 30, 2026 (cs.CL/cs.LG: distillation,
byte-level, diacritization, Muon). Four findings change what the
papers should cite. Protocol rule stands: never quote cross-protocol
numbers.

## Citations to add (paper.adoc + RESULTS.md)

1. **arXiv 2609.12303 — "Breaking the Token Ceiling: Distilling
Smaller, Stronger Byte Models"** (Meta/FAIR). First large-scale
distillation × tokenization study (~1B params, up to 1T bytes):
distilled byte students start worse but surpass token students with
compute — higher asymptote (+4% predicted), 6× data efficiency,
256-vocab avoids top-k logit truncation. External scaling-law
validation of our byte-student lineage (ByT5-style, byte+3 table).
Where: background/related work + one sentence in the frontier
discussion.
2. **arXiv 2609.37510 — MAESTRO ("From Dissonance to Orchestration")**.
Teacher intervention in on-policy distillation adds off-policy
load; always-on intervention yields diminishing returns; adaptive
disagreement-gated takeover is the fix. Cite as the mechanistic
account of our measured GKD negative (6.0036) — our always-on GKD
is the maximum-intervention point of their axis. Do NOT re-run
(ledger: student-side levers closed).
3. **arXiv 2608.27729 — "Below the Noise Floor"** (bimodal seed
collapse in small-model KD). Per-seed σ 2.8–48.7pp; single-seed KD
gains <5pp are unresolvable; 3/7 KD variants collapse bimodally.
Cite in the evaluation-methodology discussion: strengthens our
paired-bootstrap/full-set discipline AND adds the honest caveat
that our single-run arm verdicts (e.g. headwise Muon +0.2267pp)
measure prediction-resampled variance, not seed variance.
4. **arXiv 2609.10153 — YallaMorph** (EMNLP 2026). Cite as the r8
teacher's data lever (morphological aux coverage) and as evidence
the field's Arabic-morphology energy moved to LLM evaluation, not
text-diacritization SOTA.

## Also record (no citation needed)

- No new text-only SadeedDiac-25 competitor in the window; KSAA-2026
speech-diacritization winner (2605.25928, 23.26% WER) is a speech
modality — not protocol-comparable; noted in RESULTS.md sweep entry.

## Steps

1. [ ] RESULTS.md: "2026-09-30 sweep" entry (4 citations + the
no-new-competitor note).
2. [ ] paper.adoc: related-work sentences for 2609.12303 + MAESTRO;
methodology caveat sentence citing 2608.27729; YallaMorph in
the discussion where the r8 teacher lever is named.
3. [ ] PR to interscript/interscript-ml (branch off default, no AI
attribution, explicit-path staging).

## Result

(to be written only after merge)
53 changes: 53 additions & 0 deletions TODO.sota-2026/03-student-side-closed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# 03 — Student-side lever family: CLOSED (record + caveat)

Status: SPECIFIED (2026-10-01)

## The record

The student–teacher residual (2.1 student 4.5701 vs r7 teacher 2.2890)
has now resisted every lever measured:

| lever | result |
|---|---|
| corpus scale / register mix (both directions) | negative / flat-negative |
| on-policy distillation (GKD) | negative (6.0036) |
| product-key memory layers | real but small (−0.70pp) |
| epochs 3→6 | −0.25pp (secondary) |
| Muon optimizer | −2.96pp (adopted — shipped in 2.0/2.1) |
| headwise Muon (671B-scale recipe) | separated-negative (+0.2267pp) |
| Sinkhorn embeddings | flat |
| lexical memory (engram) | flat |

The 2026-09-30 literature sweep independently corroborates the
closure: the OPD wave (MAESTRO 2609.37510; RIDE 2609.36484; Fisher
sparsity 2609.36262; sparse supervision 2609.04565) targets reasoning
trajectories with distribution-shift mechanisms our deterministic
dense-label task does not have. No new student-side method in the
window contradicts the verdict.

**Standing rule: no further GPU spend on student-side levers without a
pre-registered mechanism novel to the ledger.** The frontier mover on
record is teacher-side data (r5→r6→r7). Next frontier experiments:
01 (YallaMorph aux) and 05 (RIDE probe — the only student-side item
with a cheap kill-gated probe).

## Caveat to publish (methodology honesty)

Per arXiv 2608.27729: our paired between-students bootstrap measures
prediction-resampled variance, NOT training-seed variance. All arm
verdicts to date are single-seed runs. The headwise-Muon
separated-negative (+0.2267pp, p=0.017) is directionally consistent
with its size class, but the seed axis is unmeasured; future arms run
multi-seed or state the caveat. Ship decisions are unaffected (the
base recipe shipped regardless).

## Steps

1. [ ] RESULTS.md: lever-ledger closure entry + seed-variance caveat
(same PR as 02, separate commit).
2. [ ] paper.adoc: one sentence attaching the caveat to the
optimizer-recipe arm paragraph.

## Result

(to be written only after merge)
42 changes: 42 additions & 0 deletions TODO.sota-2026/04-stoicheia-diffusion-wo.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# 04 — WO (queued, unscheduled): plane-factorized char-level masked diffusion

Status: QUEUED (2026-10-01) — entry criteria at the bottom
Literature basis: Stoicheia (arXiv 2608.07249) — 405M character-level
masked-diffusion encoder for Ancient Greek; input factors into five
aligned, independently maskable planes (letters, boundaries,
**diacritics**, capitalization, punctuation). One backbone restores
lacunae / re-segments / **accentuates** / punctuates. Beats Ithaca
24.6 → 15.5 CER with matched random-init controls (methodology kin).

## Why it is the highest-upside architecture item

- Diacritics as an explicit maskable plane IS our task's factorization
(base letters given; only harakat positions are unpredictable).
- Non-autoregressive parallel decode = potential large CPU-latency win
over our KV-cache seq2seq (our runtime bottleneck).
- Independent planes compose tasks without retokenization — one model
could serve diacritization + our other normalization tasks.

## Why it is queued, not scheduled

- Ledger: architecture transfers measured negative at our scale
(depth cut, lexical memory, PKM; Hebrew depth-cut catastrophic).
Stoicheia's evidence is restoration/scansion at 405M with heavy
pretraining (380M words) — not a matched transfer case.
- New runtime contract: non-AR iterative decode does not fit IMF v1
KV-cache graphs; needs its own export + parity path.
- Diffusion decode needs step-count/quality calibration per language.

## Spec (if entered)

1. Corpus: reuse Arabic combined + news + YallaMorph-aux; planes =
(base letters, harakat, word boundaries). Letters plane held
(skeleton-preserving, as our students already do).
2. Backbone: ~300M encoder, plane-aligned embeddings; pretrain
masked-plane objective, then SFT on diacritization.
3. Gates: same windowed zero-skip SadeedDiac harness; must beat
4.5701 (student rung) AND 2.2864 (teacher rung) to matter; CPU
decode latency benchmarked against ara-diac-small-int8static-2.1.
4. Entry criteria: (a) 01 (run-009) lands and re-ranks the teacher
frontier; (b) owner authorizes a new architecture line; (c) a
decode-parity design exists for non-AR models (IMF v2 question).
53 changes: 53 additions & 0 deletions TODO.sota-2026/05-ride-sft-residual-extrapolation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# 05 — RIDE-style SFT-residual extrapolation: probe-first arm

Status: SPECIFIED (2026-10-01) — probe only; training arm gated on probe
Literature basis: RIDE (arXiv 2609.36484) — extrapolate the
teacher-over-base residual directly in representation space:
student hidden states regressed toward
`h_target = h_teacher + λ·(h_teacher − h_base)`; approaches or exceeds
the teacher across four base/RL-teacher pairs.

## Scope correction (user-confirmed 2026-10-01)

The mechanism does NOT require an RL teacher. It needs any
(base, improved) checkpoint pair; the residual direction
`d = improved − base` is what is extrapolated. Our **r6→r7** SFT pair
(run-006-morph → run-007-news) qualifies. What remains forbidden is RL
*training* ([[rl-negative-diacritization]] — measured flat 3×), not
residual extrapolation of an SFT delta.

## Why probe-first

- Student-side lever (ledger: 8 negatives) — do not spend GPU on a
training arm before the direction is shown to transfer.
- The r7 delta is small (−0.29pp ID) and domain-shaped (news mix);
extrapolating a domain-idiosyncratic direction would amplify news
specialization, not general diacritization competence.

## Probe (cheap: forward passes only, no training)

1. Load run-006-morph/best and run-007-news/best (580M ByT5 each).
2. Forward N=200 units from two domains: SadeedDiac val paragraphs
(classical) + WikiNews-2024 text (news). Capture per-layer
mean-pooled encoder hidden states.
3. Per layer: d_classical = mean(h_r7) − mean(h_r6) on classical;
d_news likewise on news. Compute cos(d_classical, d_news).
4. **Kill criterion: max-layer cosine < 0.5 ⇒ direction is
domain-idiosyncratic ⇒ close the arm, record in RESULTS.md.**
5. Pass ⇒ full arm: hidden-state distillation with displacement
(λ ∈ {0.5, 1.0}) as an aux loss on the student trainer — requires
a new feature-regression path in modal_distill (spec before code;
TDD the loss on synthetic tensors).

## Steps

1. [ ] TDD pure computation: `residual_directions(h_base, h_teacher)`
and cosine sim on synthetic tensors (tests first, watch fail).
2. [ ] Modal probe script (two models × 200 units × 2 domains;
A100 minutes, not hours).
3. [ ] Run probe; write verdict + per-layer cosine table here.
4. [ ] Gate decision: close, or spec the training arm separately.

## Result

(to be written only from measured numbers)
61 changes: 61 additions & 0 deletions docs/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -911,3 +911,64 @@ scale-boundary data point. The Sinkhorn-balanced embedding update
remains the frontier student.** The TODO.impl recipe ledger is now
fully measured: every optimizer/architecture lever is closed; the only
frontier mover on record is data-side (teacher r5→r6→r7).

## 2026-09-30 arXiv sweep — no new competitor; our premises externally validated

Monthly sweep (window Aug 26 → Sep 30, 2026: distillation, byte-level
modeling, diacritization, optimizer literature). Competitive position
unchanged: **no new text-only Arabic diacritization system appeared on
SadeedDiac-25** — the field's Arabic-diacritization energy moved to the
speech modality (KSAA-2026 Task 2 winner, 23.26% WER, speech input:
not protocol-comparable). r7 (2.2864) remains the best dedicated model
measured under our protocol; only Claude-3.7-Sonnet's published 1.3941
sits above it.

Four findings enter the record:

- **arXiv 2609.12303 (Meta/FAIR), "Breaking the Token Ceiling"** —
first large-scale distillation × tokenization study (~1B params, up
to 1T bytes): distilled *byte* students start worse but surpass
token students with compute (predicted +4% asymptote, 6× data
efficiency, 256-symbol vocab eliminates top-k logit truncation).
Independent scaling-law validation of the byte-student lineage we
ship.
- **arXiv 2609.37510, MAESTRO** — teacher intervention in on-policy
distillation injects off-policy load; always-on intervention is the
worst point of the axis. Mechanistic account of our measured GKD
negative (6.0036); corroborates closing the on-policy lever without
a re-run.
- **arXiv 2608.27729, "Below the Noise Floor"** — per-seed σ
2.8–48.7pp in small-model KD; single-seed gains below ~5pp are
unresolvable; 3/7 KD variants collapse bimodally. Validates our
full-set + paired-bootstrap discipline and motivates the
seed-variance caveat recorded in the next entry.
- **arXiv 2609.10153, YallaMorph (EMNLP 2026)** — 663,804 controlled
Arabic morphological-generation instances (CamelMorph MSA). The
concrete teacher-side data lever for the next teacher rung
(TODO.sota-2026/01); also confirms the field's morphology work
targets LLM evaluation rather than text-diacritization SOTA.

## Student-side lever family closed; seed-variance caveat recorded (2026-10-01)

The residual ledger, complete: corpus scale ✗, register mix ✗ (both
directions), on-policy GKD ✗ (6.0036), PKM memory (real, −0.70pp),
epochs (−0.25pp), headwise Muon ✗ (separated-negative), Sinkhorn
embeddings ✗ (flat), engram lexical memory ✗ (flat). The 2026-09-30
literature sweep surfaced no student-side method that escapes the
closure — the current on-policy wave (MAESTRO 2609.37510, RIDE
2609.36484, Fisher-sparsity 2609.36262, sparse supervision 2609.04565)
targets reasoning-trajectory distribution shift that a deterministic
dense-label task does not have.

Standing rule: **no further GPU spend on student-side levers without a
pre-registered mechanism novel to this ledger.** Frontier experiments
continue teacher-side (TODO.sota-2026/01) and via the kill-gated RIDE
direction probe (TODO.sota-2026/05) — the only student-side item with
a cheap probe before any training compute.

Caveat (per 2608.27729): every arm verdict above is a single training
seed; the paired between-students bootstrap resamples predictions, not
seeds. The headwise-Muon separated-negative (+0.2267pp, p=0.017) is
directionally consistent for its size class, but the seed axis is
unmeasured. Future arms run multi-seed or carry this caveat. Ship
decisions are unaffected — the base recipe shipped on its own merits.
Loading
Loading