From 0e2a9e51ab22d92344469d50a1b17a442bbb383a Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Thu, 1 Oct 2026 12:32:41 +0800 Subject: [PATCH] docs(papers): campaign close - frontier table rows, residual closure, static-int8 artifact paper.adoc: three post-2.1 rows added to the frontier table (Engram 4.67, headwise Muon 4.82, Sinkhorn 4.55, all with teacher-delta CIs); the residual paragraph now records the optimizer-recipe arms transferring negatively at 300M (headwise Muon separated-worse +0.2267pp [0.016, 0.411]; Sinkhorn indistinguishable from AdamW); the discussion's on-policy 'open lever' updated to its measured negative (6.0036) and the path restated as teacher-side data. paper-c.adoc: artifacts section records the shipped static-int8 composition (ara-diac-small-int8static-2.1, 491MiB, cer_delta 0.1038pp, zero confident flips). --- docs/paper-c.adoc | 7 ++++++- docs/paper.adoc | 9 ++++++--- 2 files changed, 12 insertions(+), 4 deletions(-) diff --git a/docs/paper-c.adoc b/docs/paper-c.adoc index 580876c..fe93732 100644 --- a/docs/paper-c.adoc +++ b/docs/paper-c.adoc @@ -126,5 +126,10 @@ under one framing do not transfer to another. hardware for E1 == 9. Artifacts -All benchmarks, the index workflow, and the 23-model catalog are +All benchmarks, the index workflow, and the model catalog are public; every number in this paper regenerates from committed scripts. +The static-int8 composition ships in the catalog as +`ara-diac-small-int8static-2.1` (491 MiB): calibrated QUInt8 +activations with an int8-dynamic encoder, gated at cer_delta 0.1038pp +over 2,480 parity samples with zero confident-position flips — the +first artifact to clear the margin budget without residue. diff --git a/docs/paper.adoc b/docs/paper.adoc index 5f2456b..052034b 100644 --- a/docs/paper.adoc +++ b/docs/paper.adoc @@ -65,7 +65,7 @@ The 2024–2025 landscape consolidated around two directions this work deliberat We apply two standard distillation regimes: *sequence-level* distillation, where the teacher generates target strings consumed by the student's standard cross-entropy objective, and *logit* distillation (KL between teacher and student distributions plus ground-truth CE). Sequence-level KD is the natural fit when teacher and student occupy different vocabulary spaces (a sentencepiece umt5 teacher into a byte-level student); logit KD when both are byte-level and share the tokenizer. -Both regimes are *off-policy*: the training distribution is the teacher's (or the corpus's), never the student's own. The on-policy line — MiniLLM's reverse KL on student-sampled sequences (Gu et al. 2024) and Generalized Knowledge Distillation (Agarwal et al. 2024), which trains on student-generated mistakes scored by the teacher — is designed to fix exactly the pathology we measure in <>: students whose distributions are too flat to rank completions by likelihood. Sharpening the student on its own errors, rather than only imitating teacher output, is the principled next lever for both the shrink cost and the decode behavior; we leave it as future work (<>). +Both regimes are *off-policy*: the training distribution is the teacher's (or the corpus's), never the student's own. The on-policy line — MiniLLM's reverse KL on student-sampled sequences (Gu et al. 2024) and Generalized Knowledge Distillation (Agarwal et al. 2024), which trains on student-generated mistakes scored by the teacher — is designed to fix exactly the pathology we measure in <>: students whose distributions are too flat to rank completions by likelihood. Sharpening the student on its own errors, rather than only imitating teacher output, is the principled next lever for both the shrink cost and the decode behavior; its measured outcome is recorded in <>. === Quantization and deployment-aware inference @@ -289,6 +289,9 @@ The Arabic client line has since been extended to a complete, confidence-bracket |2.0 rung (Muon, 3 ep) |300M |5.08 |[2.36, 2.82] (corrected; see below) |2.1 rung (Muon, 6 ep) |300M |4.57 |[1.91, 2.35] |2.1 + full Tashkeela (G2b) |300M |4.82 |[2.19, 2.55] +|2.1 + Engram lexical memory |367M |4.67 |[2.03, 2.47] +|2.1 + headwise Muon (Q/K) |300M |4.82 |[2.15, 2.55] +|2.1 + Sinkhorn embedding update |300M |4.55 |[1.84, 2.23] |teacher r7 |580M |2.29 |— |=== @@ -300,7 +303,7 @@ zip through the runtime itself — sha-pinned to the index) agree on re-derives exactly by the same tooling from the published prediction files, which is how the discrepancy was caught. -Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp) — and we report it as a property of the compression itself rather than a shortfall of any single method. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit. It reframes as a teacher–student interaction the corpus cannot reach — candidate levers are on-policy distillation (in flight) and label-distribution effects, not more in-domain text. +Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. A third lever, on-policy distillation (GKD — training on student-generated mistakes scored by the teacher), also came back negative: 6.0036 [3.109, 3.743], 1.43pp *worse* than the off-policy rung it was meant to improve. The residual has now resisted every lever tested — corpus scale, register mix, on-policy correction, memory layers (real but 0.70pp), epochs (0.25pp), and two optimizer-recipe arms imported from the 2025 frontier-LLM literature — and we report it as a property of the compression itself rather than a shortfall of any single method. The optimizer-recipe arms are instructive because they transfer negatively at our scale: head-wise Muon on Q/K projections (reported positive at 671B scale in DeepSeek-V4.1-Flash) scores 4.8164, separated-worse than its vanilla peer by +0.2267pp [0.016, 0.411] under a paired between-students bootstrap on identical data; the Sinkhorn-balanced embedding update is statistically indistinguishable from AdamW (−0.0903pp [−0.267, 0.069]). Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit, and it is not an optimizer artifact. It reframes as a teacher–student interaction the corpus cannot reach — the remaining lever on our record is teacher-side: every frontier move in this table's upper rows came from the teacher's data, not the student's training. == The decode protocol is part of the measurement [[section-decode]] @@ -388,7 +391,7 @@ The general lesson: when a file "reads corrupt remotely and clean locally," susp *Decode.* The greedy-vs-beam result is established for distilled byte-level students. Whether sharp teachers are best decoded greedily is unmeasured; their beam numbers stand as measurements under a stated protocol. Notably the beam gain is language-dependent at the teacher tier — Hebrew teachers gain ~12 DER points from beam search while the Arabic teacher gains nothing (2.5793 greedy vs 2.5588 beam-4) — consistent with the sharpness account: the pathology tracks distribution flatness, not the task. -*On-policy distillation is the open lever.* Our students train exclusively on off-policy supervision (teacher-generated strings, or teacher logits on corpus targets). The measured flatness of student distributions (<>) is precisely the failure mode on-policy regimes — reverse KL on student-sampled sequences (Gu et al. 2024; Agarwal et al. 2024) — are designed to correct. A GKD-style pass over the shipped students could plausibly reduce shrink costs and restore sane beam behavior at once; it is the first experiment we would run next. +*On-policy distillation was the open lever — it measured negative.* Our students train exclusively on off-policy supervision (teacher-generated strings, or teacher logits on corpus targets), and the measured flatness of student distributions (<>) is precisely the failure mode on-policy regimes — reverse KL on student-sampled sequences (Gu et al. 2024; Agarwal et al. 2024) — are designed to correct. The GKD-style pass over the shipped recipe was run pre-registered: 6.0036 full-set, 1.43pp *worse* than the off-policy rung it was meant to improve. With on-policy correction, corpus composition, optimizer recipes, memory, depth, width, and epochs all measured, the student-side search is closed at this scale; the frontier moves through the teacher's data (the r5→r6→r7 sequence: −0.10pp aux-task supervision, −0.29pp domain mix), and the next experiment is an r8 teacher. *Client tier.* 193 MiB is not 30 MiB. The quantization ladder is exhausted (int4 measured, lower precisions untested); further shrinkage is a pretraining problem, not a compression problem.