You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/paper.adoc
+63-3Lines changed: 63 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -67,9 +67,24 @@ We apply two standard distillation regimes: *sequence-level* distillation, where
67
67
68
68
Post-training weight quantization of transformer MatMuls (8-bit dynamic; 4-bit blockwise) is standard deployment practice for large language models. Two aspects are, to our knowledge, under-treated in the literature and addressed by our contract: (a) *parity gates must be precision-aware* — a flat torch-vs-ONNX error bound that passes at fp32 will spuriously fail quantized artifacts (we measure fp16 ≈ 0.43pp and int8 ≈ 0.84pp character-error deltas on an unquantized-verified model); and (b) quantized graphs may use operators unavailable to old consumer runtimes (our 4-bit artifacts require MatMulNBits, absent from the oldest Ruby consumer), so precision is a first-class, advertised property of each artifact rather than a silent implementation detail.
69
69
70
-
=== Beam search and length normalization
71
-
72
-
Length-normalized beam search is the conventional decode for seq2seq evaluation. Murray & Chiang and others have analyzed its length biases; our contribution is a concrete, deployment-scale demonstration that on *distilled byte-level students* — whose per-token distributions are far flatter than their teachers' — likelihood-ranked completion is pathological in both directions (length normalization prefers long garbage; raw cumulative scores prefer short ones), while greedy decoding rides the argmax margins to a 4.2x lower error rate. Published evaluations that decode such models with beams risk reporting the decode, not the model.
70
+
=== Beam search and its pathologies
71
+
72
+
Length-normalized beam search is the conventional decode for seq2seq
73
+
evaluation, and its failure modes are well studied: Murray & Chiang
74
+
(2018) analyze length biases and their corrections; Meister, Cotterell &
75
+
Vieira (2020) show that beam search's errors are not merely scoring
76
+
artifacts but reflect which completions the underlying model actually
77
+
prefers. Our result sits at their intersection, on a model class they
78
+
did not consider — *distilled byte-level students*, whose per-token
79
+
distributions are far flatter than their teachers' — and at deployment
80
+
scale: likelihood-ranked completion is pathological in both directions
81
+
(length normalization prefers long garbage; raw cumulative scores prefer
82
+
short ones), while greedy decoding rides the argmax margins to a 4.2x
83
+
lower error rate. Published evaluations that decode such models with
84
+
beams risk reporting the decode, not the model.
85
+
86
+
image:paper-assets/architecture.png[The IMF v1 pipeline,width=580,align=center]
87
+
.Caption: The phonological-layer pipeline. Teachers are frozen before any student trains; students pass pre-registered error budgets and a precision-aware parity gate before the artifact enters the index; three runtimes resolve the same checksummed zip and produce byte-identical output.
73
88
74
89
== The artifact contract (IMF v1)
75
90
[[section-imf]]
@@ -120,6 +135,21 @@ Every language defines one harness, used for teacher and student identically:
120
135
* *Arabic diacritization:* windowed DER-CE — inputs longer than 1,400 bytes split at word boundaries, decoded greedily per window, stitched, and the teacher's haraqat projected onto the input's letters through letter-level alignment, scored by the Misraj evaluator over SadeedDiac-25 paragraphs. The windowed protocol removes truncation losses on long hadith paragraphs.
121
136
* *Persian G2P:* test-split CER plus a SentenceBench homograph accuracy (ezafe-normalized).
|Labeling |teacher greedy/beam-4 generation, resumable, ASCII-escaped transport
149
+
|Checkpoints |every 500 steps, with labels-file digest for lineage-safe resume
150
+
|Hardware |single A10G (24 GB) per training; exports and parity on CPU
151
+
|===
152
+
123
153
=== Training regimes
124
154
125
155
Students train either by sequence-level KD (teacher-generated targets; cross-entropy) or logit KD (KL + CE on gold). Client-tier students initialize from the ByT5-small pretrained backbone; we additionally constructed from-scratch custom byte-level architectures (33M and 70M parameters, with and without linear "bridge" projections that map teacher activations into the student's geometry via closed-form ridge regression) to test whether small, task-specific backbones could replace pretrained ones. They could not (<<section-frontier>>).
@@ -171,6 +201,9 @@ The campaign's central engineering question: how small can a client artifact be?
.Caption: The client-tier frontier on Thai G2P (identical harness). Randomly initialized students collapse at every capacity tested; the pretrained rung (green) is the whole quality cliff. The hollow point marks the same ByT5-small student scored under the beam-4 evaluation convention — see <<section-decode>>.
206
+
174
207
Three findings:
175
208
176
209
1. **Random initialization collapses regardless of capacity or structure.** The bridge projections (teacher activations ridge-regressed into the student's geometry) measurably improve internal structure — 75.80 → 71.12 — but cannot rescue task accuracy. Enlarging without pretraining does not help (70M scores *worse* than 33M).
@@ -198,6 +231,9 @@ Porting beam search into the runtimes forced an apples-to-apples measurement on
198
231
199
232
Exact match barely moves; edit distance collapses by 4.2x. The signature is diagnostic: beam decoding is mangling precisely the sentences the model did not get exactly right.
200
233
234
+
image:paper-assets/decode.png[Decode distributions and alpha sweep,width=620,align=center]
235
+
.Caption: (a) Top-1 log-probabilities at successive positions of a representative greedy decode, against the uniform baseline ln(259); the argmax is barely distinguished from the field. (b) PER by decode on a 150-sentence subset; no normalization exponent recovers greedy quality.
236
+
201
237
=== The mechanism
202
238
203
239
Distilled byte-level students have *flat* per-token distributions. On the measured artifact, the top-1 log-probability at a typical decoding position is ≈ −4.6 against a uniform baseline of ln(259) ≈ −5.6 — the argmax is barely distinguished from the field. Greedy decoding rides these tiny margins consistently and produces coherent output.
@@ -208,6 +244,17 @@ Beam search ranks *completions* by accumulated likelihood. On flat distributions
208
244
* with raw cumulative scores, short completions win — the maximum-likelihood complete hypothesis on this model genuinely is a truncation (`sa˨˩` for "สวัสดี");
209
245
* an α sweep (0, 0.3, 0.6, 1.0) produced PERs of 69.8, 70.1, 70.4, and 45.9 on a 150-sentence subset where greedy scores 3.20 — *no* normalization exponent recovers greedy quality.
210
246
247
+
A concrete example from the shipped artifact, input Thai for "hello":
The published beam-4 figures were not wrong measurements; they were measurements of the decode. Under a protocol users never execute, they overstated error 4.2x.
212
259
213
260
=== Implications
@@ -252,6 +299,19 @@ The phonological layer ships: eleven models, four languages, two tiers, one arti
252
299
253
300
== References
254
301
302
+
* [%hardbreaks]
303
+
Yao, K., Zweig, G. _Sequence-to-sequence neural net models for grapheme-to-phoneme conversion._ INTERSPEECH 2015.
304
+
+
305
+
* Rao, K., Peng, F., Sak, H., Beaufays, F. _Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks._ ICASSP 2015.
306
+
+
307
+
* Yolchuyeva, B., Németh, G., Gyires-Tóth, B. _Grapheme-to-phoneme conversion with deep learning._ Speech Communication 114 (2019).
308
+
+
309
+
* Meister, I., Cotterell, R., Vieira, T. _If beam search is the answer, what was the question?_ EMNLP 2020.
310
+
+
311
+
* Dettmers, T., Lewis, M., Shleifer, S., Zettlemoyer, L. _LLM.int8(): 8-bit matrix multiplication for transformers at scale._ NeurIPS 2022.
312
+
+
313
+
* Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D. _GPTQ: Accurate post-training quantization for generative pre-trained transformers._ ICLR 2023.
0 commit comments