Skip to content

feat(lang,backends,io): BITNET_PLANES encoding + fused lm_head kernel pack + loader requantizer (#1150) - #1171

Merged
michalharakal merged 1 commit into
developfrom
feature/1150-bitnet-planes
Aug 26, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1150-bitnet-planes

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Phase 6 of #1136 — the core of #1150 (JNI/Kotlin-Native faces of the lmhead kernel remain as follow-ups within the issue).

What

  • TensorEncoding.BITNET_PLANES — 8 sequentially-packed trit planes + FP16 per-row scales (plane p worth 1/3^p; truncation ≤ rowScale/(2·3⁷)). Speed format, not memory (2 B/weight): the fused 4-plane LUT kernel reads planes 0–3 in one baseline-NEON pass; planes 4–7 exist for application-level top-k rescoring. Deliberately no BlockSpec, no int8 activation hint — geometry is row-scoped, and without the pack dispatch falls to the decoding reference, never the requantize adapter (pinned by test).
  • TernaryCodec.encodeBitNetPlanes — Kotlin port of NeoGPU's hs_mlt_lmhead_encode (decomposes against the FP16-rounded scale the decoder will read); decodeBitNetPlanes(Row) + planesRowScale; BitNetPlanesTensorData blocks per row and carries the exact dispatch key.
  • TernaryLmheadNative + TernaryPlanesKernelPack — native seam = 4 fused planes per call (skainet_ternary_lmhead_stage1, exported since feat(native): vendored NeoGPU ternary f32 LUT kernel + FFM downcall (#1137) #1151); the view kernel combines two calls as s0 + s4/81, so dispatch's matmul equals the decoded 8-plane matmul exactly — the invariant survives, and stage-1-only scoring stays an application decision.
  • FFM face NativeTernaryLmheadKernel (+install()); row-scale pointer = weight segment sliced at the (always even) scale offset.
  • Loader's first requantizer: RequantizeTo(BITNET_PLANES) from F32/F16/BF16/quantized/I2_S sources, OUT_IN orientation required, traced as requantize-planes. Every other RequantizeTo still fails eagerly. This is the hook transformers#337 wires for output.weight.

Tests (all green locally, macOS arm64)

  • BitNetPlanesCodecTest — truncation-bound reconstruction, plane-0 sign structure, FP16 row scale, geometry helpers, TensorData↔codec agreement
  • TernaryPlanesKernelPackTest — FakeNative two-call exactness vs the 8-plane reference; install(null) registers nothing; decoding reference (not int8 adapter) serves without the pack
  • TernaryPlanesFfmTest — the REAL vendored hs_ml_lmhead_stage1 behind REAL dispatch at BitNet hidden size (n=512, k=2560), incl. its internal 4-thread pool
  • PlanesRequantizeLoadTest — F32→planes round-trip within the bound; eager rejection of other RequantizeTo targets and missing OUT_IN
  • Full jvm suites of lang-core / backend-api / backend-native-cpu / io-gguf pass; commonMain cross-compiles (js, linuxX64)

Remaining in #1150 after this PR: JNI + Kotlin/Native faces of the lmhead kernel (same pattern as #1161), qemu lane for them.

🤖 Generated with Claude Code

TensorEncoding.BITNET_PLANES is NeoGPU's lm_head weight format as a plain
SKaiNET encoding (#1150): 8 sequentially-packed trit planes + one FP16
scale per row, plane p worth 1/3^p — "16 bits as eight ternary digits",
truncation error <= rowScale/(2*3^7). The point is speed, not memory
(2 B/weight, FP16-sized): the vendored fused 4-plane LUT kernel reads
planes 0-3 in one baseline-NEON pass, and an application can rescore top
candidates with planes 4-7 (NeoGPU's two-stage lm_head) through the
codec's row accessors.

The format stays format-driven end to end, per the #1136 physiology: no
new op, layer, or fusion pass anywhere.

- TernaryCodec.encodeBitNetPlanes is the Kotlin port of NeoGPU's
  hs_mlt_lmhead_encode — per-row absmax as FP16 (decomposed against the
  FP16-rounded scale the decoder will read, not the exact one), repeated
  round-to-trit with x3 residual; decodeBitNetPlanes/Row + planesRowScale
  are the reference readers. BitNetPlanesTensorData blocks per ROW.
- BITNET_PLANES deliberately has no BlockSpec and no int8 activation
  hint: its geometry is row-scoped (row-count-free BlockSpec cannot say
  it), and without the pack dispatch must fall to the decoding reference,
  never the requantize adapter — pinned by test.
- TernaryLmheadNative + TernaryPlanesKernelPack: the native seam computes
  4 fused planes per call (skainet_ternary_lmhead_stage1, exported in
  #1137); the view kernel makes two calls combined as s0 + s4/81, so
  dispatch's matmul equals the decoded 8-plane matmul EXACTLY — stage-1
  truncation is an application decision, never a dispatch surprise.
- FFM face NativeTernaryLmheadKernel (row scales = the weight segment
  sliced at the aligned scale offset); JNI/K-N faces are follow-ups
  within #1150.
- StreamingGgufParametersLoader gains its first requantizer:
  RequantizeTo(BITNET_PLANES) from F32/F16/BF16/quantized/I2_S sources,
  OUT_IN orientation required (the scales are per output row), traced as
  requantize-planes. Every other RequantizeTo still fails eagerly.

Tests: codec truncation-bound goldens + plane-0 sign structure; pack
contract with a FakeNative (two-call exactness vs the 8-plane reference,
nothing registered without native, no requantize adapter without the
pack); the REAL vendored kernel behind REAL dispatch at BitNet hidden
size; loader requantize round-trip within the bound + eager rejections.

Refs #1150, #1136

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1171 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 2f7dbfb into develop Aug 26, 2026
20 of 21 checks passed
@michalharakal
michalharakal deleted the feature/1150-bitnet-planes branch August 26, 2026 13:38
michalharakal added a commit that referenced this pull request Aug 26, 2026
…heck

The binary-compatibility validator's jvmApiCheck wants the .api dump to
name what #1171 made public — TernaryCodec's plane codec functions,
BitNetPlanesTensorData, and TensorEncoding.BITNET_PLANES. This also heals
the check on develop, which went red when #1171 merged without the dump.

Refs #1150

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
michalharakal added a commit that referenced this pull request Aug 26, 2026
…-planes

chore(bcv): regenerate the lang-core API dump missed by #1171
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant