feat(lang,backends,io): BITNET_PLANES encoding + fused lm_head kernel pack + loader requantizer (#1150) - #1171
Merged
Conversation
TensorEncoding.BITNET_PLANES is NeoGPU's lm_head weight format as a plain SKaiNET encoding (#1150): 8 sequentially-packed trit planes + one FP16 scale per row, plane p worth 1/3^p — "16 bits as eight ternary digits", truncation error <= rowScale/(2*3^7). The point is speed, not memory (2 B/weight, FP16-sized): the vendored fused 4-plane LUT kernel reads planes 0-3 in one baseline-NEON pass, and an application can rescore top candidates with planes 4-7 (NeoGPU's two-stage lm_head) through the codec's row accessors. The format stays format-driven end to end, per the #1136 physiology: no new op, layer, or fusion pass anywhere. - TernaryCodec.encodeBitNetPlanes is the Kotlin port of NeoGPU's hs_mlt_lmhead_encode — per-row absmax as FP16 (decomposed against the FP16-rounded scale the decoder will read, not the exact one), repeated round-to-trit with x3 residual; decodeBitNetPlanes/Row + planesRowScale are the reference readers. BitNetPlanesTensorData blocks per ROW. - BITNET_PLANES deliberately has no BlockSpec and no int8 activation hint: its geometry is row-scoped (row-count-free BlockSpec cannot say it), and without the pack dispatch must fall to the decoding reference, never the requantize adapter — pinned by test. - TernaryLmheadNative + TernaryPlanesKernelPack: the native seam computes 4 fused planes per call (skainet_ternary_lmhead_stage1, exported in #1137); the view kernel makes two calls combined as s0 + s4/81, so dispatch's matmul equals the decoded 8-plane matmul EXACTLY — stage-1 truncation is an application decision, never a dispatch surprise. - FFM face NativeTernaryLmheadKernel (row scales = the weight segment sliced at the aligned scale offset); JNI/K-N faces are follow-ups within #1150. - StreamingGgufParametersLoader gains its first requantizer: RequantizeTo(BITNET_PLANES) from F32/F16/BF16/quantized/I2_S sources, OUT_IN orientation required (the scales are per output row), traced as requantize-planes. Every other RequantizeTo still fails eagerly. Tests: codec truncation-bound goldens + plane-0 sign structure; pack contract with a FakeNative (two-call exactness vs the 8-plane reference, nothing registered without native, no requantize adapter without the pack); the REAL vendored kernel behind REAL dispatch at BitNet hidden size; loader requantize round-trip within the bound + eager rejections. Refs #1150, #1136 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
michalharakal
added a commit
that referenced
this pull request
Aug 26, 2026
…heck The binary-compatibility validator's jvmApiCheck wants the .api dump to name what #1171 made public — TernaryCodec's plane codec functions, BitNetPlanesTensorData, and TensorEncoding.BITNET_PLANES. This also heals the check on develop, which went red when #1171 merged without the dump. Refs #1150 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
michalharakal
added a commit
that referenced
this pull request
Aug 26, 2026
…-planes chore(bcv): regenerate the lang-core API dump missed by #1171
This was referenced Aug 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 6 of #1136 — the core of #1150 (JNI/Kotlin-Native faces of the lmhead kernel remain as follow-ups within the issue).
What
TensorEncoding.BITNET_PLANES— 8 sequentially-packed trit planes + FP16 per-row scales (plane p worth 1/3^p; truncation ≤ rowScale/(2·3⁷)). Speed format, not memory (2 B/weight): the fused 4-plane LUT kernel reads planes 0–3 in one baseline-NEON pass; planes 4–7 exist for application-level top-k rescoring. Deliberately no BlockSpec, no int8 activation hint — geometry is row-scoped, and without the pack dispatch falls to the decoding reference, never the requantize adapter (pinned by test).TernaryCodec.encodeBitNetPlanes— Kotlin port of NeoGPU'shs_mlt_lmhead_encode(decomposes against the FP16-rounded scale the decoder will read);decodeBitNetPlanes(Row)+planesRowScale;BitNetPlanesTensorDatablocks per row and carries the exact dispatch key.TernaryLmheadNative+TernaryPlanesKernelPack— native seam = 4 fused planes per call (skainet_ternary_lmhead_stage1, exported since feat(native): vendored NeoGPU ternary f32 LUT kernel + FFM downcall (#1137) #1151); the view kernel combines two calls ass0 + s4/81, so dispatch's matmul equals the decoded 8-plane matmul exactly — the invariant survives, and stage-1-only scoring stays an application decision.NativeTernaryLmheadKernel(+install()); row-scale pointer = weight segment sliced at the (always even) scale offset.RequantizeTo(BITNET_PLANES)from F32/F16/BF16/quantized/I2_S sources,OUT_INorientation required, traced asrequantize-planes. Every otherRequantizeTostill fails eagerly. This is the hook transformers#337 wires foroutput.weight.Tests (all green locally, macOS arm64)
BitNetPlanesCodecTest— truncation-bound reconstruction, plane-0 sign structure, FP16 row scale, geometry helpers, TensorData↔codec agreementTernaryPlanesKernelPackTest— FakeNative two-call exactness vs the 8-plane reference;install(null)registers nothing; decoding reference (not int8 adapter) serves without the packTernaryPlanesFfmTest— the REAL vendoredhs_ml_lmhead_stage1behind REAL dispatch at BitNet hidden size (n=512, k=2560), incl. its internal 4-thread poolPlanesRequantizeLoadTest— F32→planes round-trip within the bound; eager rejection of other RequantizeTo targets and missing OUT_INRemaining in #1150 after this PR: JNI + Kotlin/Native faces of the lmhead kernel (same pattern as #1161), qemu lane for them.
🤖 Generated with Claude Code