@@ -60,6 +60,26 @@ IMF test split (1,864 long sentences), same harness for both models
6060Shrink cost +5.58pp — inside the ~ 5.6pp budget pre-accepted for this
6161pair (rababa docs/DISTILL-SOURCE-PROMPT.md section 2).
6262
63+ ## tha-g2p-small-1.0 — Thai G2P client tier (2026-08-22)
64+
65+ The client-tier release of the Thai G2P distillation: run-003,
66+ ByT5-small student on the full label set (48,757 usable beam-4 labels
67+ from the B-K/umt5-thai-g2p-v2-0.5k teacher). Same harness as
68+ tha-g2p-base-1.0 (beam-4, corpus-level PER, 1,219 held-out Kaikki Thai
69+ test sentences, ` src/gpu/modal_distill.py::evaluate_per ` ; checkpoint
70+ re-measured 2026-08-22 for this release).
71+
72+ | Model | PER | Exact match |
73+ | ---| ---| ---|
74+ | Teacher (B-K/umt5 hub base) | 4.43% | 95.57% |
75+ | ** Student (ByT5-small, client rung)** | ** 12.06%** | 87.94% |
76+
77+ Shrink cost +7.63pp — outside the +5pp server-tier gate (that gate is
78+ met by tha-g2p-base-1.0 at 9.19%): shipped anyway per the frontier
79+ below, as the smallest artifact that does not collapse. Exported at
80+ int8 (~ 300MB); see the frontier table for why no smaller rung exists
81+ today.
82+
6383## Client-tier size–quality frontier (2026-08-22)
6484
6585Thai G2P, same harness (beam-4 corpus PER, 1,219 Kaikki sentences; teacher
@@ -70,7 +90,7 @@ B-K umt5 4.43%):
7090| custom 8+8 d384 | random | 33M | ~ 30MB | 75.80 (collapsed) |
7191| custom 8+8 d384 + bridges | random | 33M | ~ 30MB | 71.12 |
7292| custom 10+10 d512 + bridges | random | 70M | ~ 70MB | 78.51 |
73- | ByT5-small | pretrained | 300M | ~ 300MB | 12.63 |
93+ | ByT5-small | pretrained | 300M | ~ 300MB | 12.06 |
7494| ByT5-base (server tier) | pretrained | 580M | 1.2GB fp32 | 9.19 |
7595
7696Findings: (1) random-init byte-level seq2seq collapses regardless of
@@ -79,7 +99,7 @@ capacity at this scale — the microkimi bridges improve structure (75.8 →
7999not help (70M = 78.5). (2) ByT5-small's width (d=1472) dominates its
80100parameter count — depth-pruning yields no useful intermediate rung
81101(263M). (3) The pretrained rung is the whole quality cliff: 300M at
82- 12.63% vs 70M at 78.5%.
102+ 12.06% (run-003, full labels; 12. 63% on the 23K subset) vs 70M at 78.5%.
83103
84104Conclusion: G2P client tier ships at the ByT5-small rung (~ 300MB int8)
85105today; a 30–70MB G2P tier requires byte-level pretraining of the small
0 commit comments