@@ -521,3 +521,30 @@ levers: on-policy distillation — GKD in flight; label-distribution
521521effects). G2a (2.1) remains the best student; G2b stands as the
522522closing negative of the data-side program. Provenance: labels sha256
523523b59e2f56 (235.0MB).
524+
525+ ## heb-diac-small-s46-layerdrop — the depth-cut does NOT transfer: 77.48 DER (2026-09-05)
526+
527+ Item 04 (TODO.publish-client): the Arabic width/depth finding
528+ replicated on Hebrew — encoder 12->6 verbatim layer copy from
529+ pretrained ByT5-small, single variable vs run-002-s46 (same s46
530+ teacher, hebrew-v4 corpus, 3 epochs, logit-KD recipe). Training
531+ converged normally (val_loss 0.550); the full Nakdimon gate did not:
532+
533+ | Model | DER (n=1864) |
534+ | ---| ---|
535+ | teacher s46 | 23.72 (reproduces exactly) |
536+ | full-depth student (run-002, 1.1) | 30.38 |
537+ | ** layerdrop student (run-003)** | ** 77.48** — collapse |
538+
539+ Paired delta +53.77pp [ 51.64, 55.92] , p=0.
540+
541+ Verdict: the "depth is the compressible axis" finding does NOT
542+ transfer as-is. Confound, stated honestly: the Arabic rung used
543+ sequence-KD (teacher labels, Muon, 6 epochs); this run used the Hebrew
544+ lineage's logit-KD (alpha-KL + CE) at its native 3 epochs — so the
545+ collapse may be recipe-dependent (depth-cut + logit-KD), not purely
546+ linguistic. Either way the cross-lingual generalization claim is
547+ closed as a negative: depth-compressibility is NOT a universal
548+ property of pretrained ByT5-small; it held under one recipe on one
549+ language. Paper B's depth paragraph is scoped accordingly (this entry
550+ is its counterexample).
0 commit comments