Skip to content

[bitnet] T1: BitNet architecture — ModelFamily, network def, packed-weight loader #336

Description

@michalharakal

Part of the BitNet tracking issue (T1). Baseline BitNet support through the standard (unfused) path — correctness first, fusion in T2.

Tasks

  • ModelRegistry: add ModelFamily.BITNET + detect() branch for GGUF general.architecture values "bitnet" / "bitnet-25" / "bitnet-b1.58"; wire into the skainet-cli/Main.kt dispatch chain
  • New llm-inference/bitnet/: network def bitnetNetwork(...) via decoderTransformerNetwork (BitNet-2B4T is llama-like with sub-layernorm — verify layer structure against the GGUF metadata), BitNetGgufTensorNames, weight loader accepting packed BITNET_B1_58 tensors from the SKaiNET I2_S loader (SKaiNET#1140)
  • Name resolver entries (output.weight / lm_head.weight) in LLMWeightNameResolvers
  • Baseline correctness: decode through the standard VoidDense lm_head path — the packed ternary weights dispatch to SKaiNET's exact ternary f32 kernel (or its reference fallback) via linearProject → ops.matmul; no fusion yet

Tests

  • Weight-loading test against a small synthetic BitNet GGUF (packed tensors arrive packed, not FP32-widened)
  • Logits parity vs FP32-widened load of the same model
  • CLI smoke: architecture auto-detected, coherent generation

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions