Skip to content
This repository was archived by the owner on Sep 29, 2026. It is now read-only.

Repository files navigation

hnetbit — Hierarchical Ternary Byte-Level Language Models

Benchmark harness for a MatMul-free, byte-level language model with learned dynamic chunking (H-Net-style hierarchy) on ternary {−1, 0, +1} weights, compared against a BPE Transformer and a flat MatMul-free HGRN on 25B bytes of Spanish text.

Three architectures, two sizes (150M / 350M), identical data budget, identical optimizer settings:

Model Description Implementation
transformer Llama-style Transformer, GPT-2 BPE tokenizer, FP32 weights HuggingFace LlamaForCausalLM
matmulfree Flat HGRN recurrent LM, byte-level, ternary weights third_party/matmulfreellm (submodule)
hybrid HNetBit — hierarchical HGRN with dynamic chunking, byte-level, ternary weights hnet_bit/

The hybrid model combines dynamic chunk boundaries (a small routing network learns where to split the byte stream at each hierarchy stage) with a MatMul-free ternary HGRN backbone. See hnet_bit/README.md for the full architecture reference.

Reported results (25B bytes each)

Quality is measured in bits per byte (BPB, lower is better), computed from the final evaluation over 200 held-out batches. Deployment size is the compact ternary export (model_deploy.pt, frozen ternary weights packed at ~2.1 bits/parameter).

Model Size Parameters Non-emb. params BPB Deploy size
matmulfree 150M 113.8M 113.4M 1.4743 28.4 MiB
matmulfree 350M 309.1M 308.6M 1.3698 75.8 MiB
transformer 150M 190.5M 113.3M 1.4837 — (FP32 dump)
transformer 350M 505.6M 402.7M 1.3861 — (FP32 dump)
hybrid (HNetBit) 150M 138.3M 138.0M 1.6221 37.9 MiB
hybrid (HNetBit) 350M 419.4M 419.1M 1.4373 115.4 MiB

The full aggregate is committed at results/results.csv; per-run training statistics, inference profiles, and segmentation analyses are under results/.

Install

git clone --recursive <repository-url>   # the anonymized snapshot exposes its own URL
cd hnetbit
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

On a cloud GPU box, bash scripts/setup_cloud.sh installs everything (PyTorch + CUDA 12.4, Triton, causal-conv1d, tooling) instead.

If you cloned without --recursive:

git submodule update --init --recursive

Data (one-time; downloads ~8.7 GB from HuggingFace)

python train_spanish.py --model hybrid --size 150M --max_steps 1 --batch_size 1
# subsequent runs: add --skip_data_build

The corpus (jhonparra18/spanish_billion_words_clean) is downloaded on first run and cached under data/spanish/. It is not redistributed here.

Smoke test (CPU)

bash test_smoke.sh          # hybrid tiny, CPU
bash test_smoke.sh --gpu    # all models, requires CUDA + HuggingFace login

Train

python train_spanish.py --model {transformer,matmulfree,hybrid} --size {150M,350M}

All runs consume the same 25B underlying text bytes with an effective batch size of 32 (4 × 8 gradient accumulation), AdamW, bf16, and a WSD schedule (1% warmup / 79% stable / 20% cosine decay). Outputs land in runs/spanish/{model}_{size}/.

Resume from a checkpoint

python train_spanish.py --model hybrid --size 150M --skip_data_build \
    --resume_from runs/spanish/hybrid_150M/checkpoint_step_125000.pt

Restores model, optimizer, scheduler, scaler, global step, bytes seen, and the best validation BPB. Intermediate step checkpoints are pruned during training; only checkpoint_best.pt, checkpoint_final.pt, and milestone checkpoints survive. Training and validation logs are appended on resume.

Export / profile

# Compact ternary export (~2.1 bits/param; Transformer exports are plain FP32 dumps)
python export_deployment.py --checkpoint runs/spanish/hybrid_150M/checkpoint_best.pt

# Prefill latency + decode throughput + inference memory
python profile_inference.py --export runs/spanish/hybrid_150M/model_deploy.pt

# Matched-byte contexts for all six deploys (single GPU)
python scripts/profile_all_deploys.py

Segmentation analysis

python scripts/show_segmentation.py --deploy runs/spanish/hybrid_350M/model_deploy.pt
python scripts/segmentation_stats.py --deploy ... --samples-dir ...
python scripts/segmentation_stats_ext.py --deploy ... --out results/segmentation/seg_stats_ext_350M.json
python scripts/segmentation_baselines.py --samples-dir results/segmentation/samples_val

The committed segmentation outputs are under results/segmentation/ (statistics, fixed-grid baselines, bootstrap CIs, annotated samples).

Results

# Aggregate everything under runs/spanish/ into one CSV
python generate_results.py --runs_dir ./runs/spanish --output results.csv

# Or reproduce the committed aggregate from the committed metadata
python generate_results.py --runs_dir ./results --output /tmp/results_regen.csv

generate_results.py reads per-run metadata only — it never modifies run directories. If a checkpoint is present it also (re)generates the compact model_deploy.pt and records its size.

Reproducing the figures

python scripts/make_prefill_scaling_figure.py --runs-dir results --lang en \
    --out figures/prefill_scaling

python scripts/plot_segmentation_example.py \
    --csv results/segmentation/seg_s06_fixed.csv \
    --raw-file results/segmentation/samples/s06_offset45M.txt \
    --positions-json results/segmentation/seg_s06_fixed.json \
    --window 0:129 --lang en --label "HNetBit 350M" \
    --out figures/segmentation_example

Deployment artifacts

The compact ternary exports are published as assets of the deploys-v1 release tag on the GitHub mirror (they are not in Git; the 350M hybrid is

100 MB). If the artifacts are not reachable from the anonymized snapshot, regenerate an export from a checkpoint with export_deployment.py.

Run Release asset Size
hybrid_150M hybrid_150M_model_deploy.pt 39.8 MB
hybrid_350M hybrid_350M_model_deploy.pt 121.0 MB
matmulfree_150M matmulfree_150M_model_deploy.pt 29.8 MB
matmulfree_350M matmulfree_350M_model_deploy.pt 79.5 MB

The Transformer exports are plain FP32 state-dict dumps (no BitLinear layers to pack) and are intentionally not distributed — regenerate them with export_deployment.py if needed.

Loading exports (frozen-ternary quirk)

The ternary exports store frozen weights u = weight_quant(w) with |u| = mean(|w|) on the non-zeros. Because BitLinear.forward re-quantizes self.weight on every call, loading the raw state dict naively shrinks each ternary layer by its nonzero density and degrades the model (MatMul-free 150M would score BPB 3.99 instead of 1.47). The harness handles this automatically: profile_inference.load_model_from_source calls hnet_bit.ops.bitnet.compensate_frozen_ternary_weights for exports containing packed_weights, scaling each stored weight by 1/density so the forward pass reproduces the frozen ternary weights exactly. Custom loaders must call the same helper after load_state_dict. The FP32 Transformer export is unaffected.

Hardware / environment notes

  • All reported runs used a single A100 80 GB; 25B bytes per model.
  • Triton is required for fast training (and is a hard import for hnet_bit.ops). In containers that block Triton's JIT compilation, set HNETBIT_DISABLE_TRITON=1 to fall back to naive PyTorch loops (3–10× slower, identical results).
  • HNetBit's short convolution uses a PyTorch fallback when causal-conv1d is unavailable (functionally equivalent, slightly slower).
  • Gradient checkpointing is functional for transformer / matmulfree but is not implemented in HNetBit's hierarchical forward pass — the hybrid runs trained without activation recomputation.
  • Learning rates differ by family (3e-4 Transformer / 4e-3 at 150M and 2.5e-3 at 350M for the ternary models), following the original MatMul-free and H-Net recipes. Context windows are nearly symmetric (~4,043 bytes for the BPE Transformer vs 4,096 bytes for the byte-level models).

Third-party baseline caveat

The MatMul-free baseline is used as a Git submodule (third_party/matmulfreellm) pinned to upstream commit f24cfe5 (https://github.com/ridgerchu/matmulfreellm, Apache-2.0). The working copy used for the reported runs carried a few additional local edits that adapted its recurrent-cache handling to the transformers version in use; those edits are not redistributed here. Training, BPB evaluation, and export of the baseline are unaffected. If you call model.generate() on the MatMul-free model with a recent transformers release and hit cache or shape errors, apply equivalent fixes locally or pin a transformers version compatible with upstream.

Repository layout

hnet_bit/                  # the proposed model (BitLinear ops, HGRN blocks, dynamic chunking)
third_party/matmulfreellm/ # MatMul-free baseline (submodule)
train_spanish.py           # training entry point
model_factory.py           # builds all three architectures
training_config_spanish.py # config, WSD scheduler, optimizer
data_spanish.py            # corpus download + byte/BPE tokenization
metrics_spanish.py         # BPB, inference memory
export_deployment.py       # checkpoint -> compact ternary export
profile_inference.py       # prefill/decode profiling
generate_results.py        # per-run metadata -> aggregate CSV
test_smoke.sh              # CPU smoke test
scripts/                   # dataset build, profiling, segmentation, figures, cloud setup
configs/                   # per-run training configs
results/                   # committed per-run metadata + aggregate CSV + figures data
figures/                   # committed PNG/PDF figures
deploy/                    # local staging for release assets (gitignored)

Citation

@misc{hnetbit2026,
  title        = {hnetbit: Hierarchical Ternary Byte-Level Language Models},
  author       = {{The hnetbit Authors}},
  year         = {2026},
  note         = {Double-blind submission; the author list will be finalized in the camera-ready version.}
}

License

MIT — see LICENSE. Third-party attributions (MatMul-free LM, H-Net, and the dataset) are in NOTICE.

About

Hierarchical ternary byte-level language models — benchmark harness

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages