Benchmark harness for a MatMul-free, byte-level language model with learned dynamic chunking (H-Net-style hierarchy) on ternary {−1, 0, +1} weights, compared against a BPE Transformer and a flat MatMul-free HGRN on 25B bytes of Spanish text.
Three architectures, two sizes (150M / 350M), identical data budget, identical optimizer settings:
| Model | Description | Implementation |
|---|---|---|
transformer |
Llama-style Transformer, GPT-2 BPE tokenizer, FP32 weights | HuggingFace LlamaForCausalLM |
matmulfree |
Flat HGRN recurrent LM, byte-level, ternary weights | third_party/matmulfreellm (submodule) |
hybrid |
HNetBit — hierarchical HGRN with dynamic chunking, byte-level, ternary weights | hnet_bit/ |
The hybrid model combines dynamic chunk boundaries (a small routing network
learns where to split the byte stream at each hierarchy stage) with a
MatMul-free ternary HGRN backbone. See hnet_bit/README.md for the full
architecture reference.
Quality is measured in bits per byte (BPB, lower is better), computed from
the final evaluation over 200 held-out batches. Deployment size is the compact
ternary export (model_deploy.pt, frozen ternary weights packed at ~2.1
bits/parameter).
| Model | Size | Parameters | Non-emb. params | BPB | Deploy size |
|---|---|---|---|---|---|
matmulfree |
150M | 113.8M | 113.4M | 1.4743 | 28.4 MiB |
matmulfree |
350M | 309.1M | 308.6M | 1.3698 | 75.8 MiB |
transformer |
150M | 190.5M | 113.3M | 1.4837 | — (FP32 dump) |
transformer |
350M | 505.6M | 402.7M | 1.3861 | — (FP32 dump) |
hybrid (HNetBit) |
150M | 138.3M | 138.0M | 1.6221 | 37.9 MiB |
hybrid (HNetBit) |
350M | 419.4M | 419.1M | 1.4373 | 115.4 MiB |
The full aggregate is committed at results/results.csv; per-run training
statistics, inference profiles, and segmentation analyses are under
results/.
git clone --recursive <repository-url> # the anonymized snapshot exposes its own URL
cd hnetbit
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txtOn a cloud GPU box, bash scripts/setup_cloud.sh installs everything
(PyTorch + CUDA 12.4, Triton, causal-conv1d, tooling) instead.
If you cloned without --recursive:
git submodule update --init --recursivepython train_spanish.py --model hybrid --size 150M --max_steps 1 --batch_size 1
# subsequent runs: add --skip_data_buildThe corpus (jhonparra18/spanish_billion_words_clean) is downloaded on first
run and cached under data/spanish/. It is not redistributed here.
bash test_smoke.sh # hybrid tiny, CPU
bash test_smoke.sh --gpu # all models, requires CUDA + HuggingFace loginpython train_spanish.py --model {transformer,matmulfree,hybrid} --size {150M,350M}All runs consume the same 25B underlying text bytes with an effective batch
size of 32 (4 × 8 gradient accumulation), AdamW, bf16, and a WSD schedule
(1% warmup / 79% stable / 20% cosine decay). Outputs land in
runs/spanish/{model}_{size}/.
python train_spanish.py --model hybrid --size 150M --skip_data_build \
--resume_from runs/spanish/hybrid_150M/checkpoint_step_125000.ptRestores model, optimizer, scheduler, scaler, global step, bytes seen, and the
best validation BPB. Intermediate step checkpoints are pruned during training;
only checkpoint_best.pt, checkpoint_final.pt, and milestone checkpoints
survive. Training and validation logs are appended on resume.
# Compact ternary export (~2.1 bits/param; Transformer exports are plain FP32 dumps)
python export_deployment.py --checkpoint runs/spanish/hybrid_150M/checkpoint_best.pt
# Prefill latency + decode throughput + inference memory
python profile_inference.py --export runs/spanish/hybrid_150M/model_deploy.pt
# Matched-byte contexts for all six deploys (single GPU)
python scripts/profile_all_deploys.pypython scripts/show_segmentation.py --deploy runs/spanish/hybrid_350M/model_deploy.pt
python scripts/segmentation_stats.py --deploy ... --samples-dir ...
python scripts/segmentation_stats_ext.py --deploy ... --out results/segmentation/seg_stats_ext_350M.json
python scripts/segmentation_baselines.py --samples-dir results/segmentation/samples_valThe committed segmentation outputs are under results/segmentation/
(statistics, fixed-grid baselines, bootstrap CIs, annotated samples).
# Aggregate everything under runs/spanish/ into one CSV
python generate_results.py --runs_dir ./runs/spanish --output results.csv
# Or reproduce the committed aggregate from the committed metadata
python generate_results.py --runs_dir ./results --output /tmp/results_regen.csvgenerate_results.py reads per-run metadata only — it never modifies run
directories. If a checkpoint is present it also (re)generates the compact
model_deploy.pt and records its size.
python scripts/make_prefill_scaling_figure.py --runs-dir results --lang en \
--out figures/prefill_scaling
python scripts/plot_segmentation_example.py \
--csv results/segmentation/seg_s06_fixed.csv \
--raw-file results/segmentation/samples/s06_offset45M.txt \
--positions-json results/segmentation/seg_s06_fixed.json \
--window 0:129 --lang en --label "HNetBit 350M" \
--out figures/segmentation_exampleThe compact ternary exports are published as assets of the deploys-v1
release tag on the GitHub mirror (they are not in Git; the 350M hybrid is
100 MB). If the artifacts are not reachable from the anonymized snapshot, regenerate an export from a checkpoint with
export_deployment.py.
| Run | Release asset | Size |
|---|---|---|
hybrid_150M |
hybrid_150M_model_deploy.pt |
39.8 MB |
hybrid_350M |
hybrid_350M_model_deploy.pt |
121.0 MB |
matmulfree_150M |
matmulfree_150M_model_deploy.pt |
29.8 MB |
matmulfree_350M |
matmulfree_350M_model_deploy.pt |
79.5 MB |
The Transformer exports are plain FP32 state-dict dumps (no BitLinear layers
to pack) and are intentionally not distributed — regenerate them with
export_deployment.py if needed.
The ternary exports store frozen weights u = weight_quant(w) with
|u| = mean(|w|) on the non-zeros. Because BitLinear.forward re-quantizes
self.weight on every call, loading the raw state dict naively shrinks each
ternary layer by its nonzero density and degrades the model (MatMul-free 150M
would score BPB 3.99 instead of 1.47). The harness handles this automatically:
profile_inference.load_model_from_source calls
hnet_bit.ops.bitnet.compensate_frozen_ternary_weights for exports containing
packed_weights, scaling each stored weight by 1/density so the forward pass
reproduces the frozen ternary weights exactly. Custom loaders must call the
same helper after load_state_dict. The FP32 Transformer export is unaffected.
- All reported runs used a single A100 80 GB; 25B bytes per model.
- Triton is required for fast training (and is a hard import for
hnet_bit.ops). In containers that block Triton's JIT compilation, setHNETBIT_DISABLE_TRITON=1to fall back to naive PyTorch loops (3–10× slower, identical results). - HNetBit's short convolution uses a PyTorch fallback when
causal-conv1dis unavailable (functionally equivalent, slightly slower). - Gradient checkpointing is functional for
transformer/matmulfreebut is not implemented in HNetBit's hierarchical forward pass — the hybrid runs trained without activation recomputation. - Learning rates differ by family (3e-4 Transformer / 4e-3 at 150M and 2.5e-3 at 350M for the ternary models), following the original MatMul-free and H-Net recipes. Context windows are nearly symmetric (~4,043 bytes for the BPE Transformer vs 4,096 bytes for the byte-level models).
The MatMul-free baseline is used as a Git submodule
(third_party/matmulfreellm) pinned to upstream commit f24cfe5
(https://github.com/ridgerchu/matmulfreellm, Apache-2.0). The working copy
used for the reported runs carried a few additional local edits that adapted
its recurrent-cache handling to the transformers version in use; those edits
are not redistributed here. Training, BPB evaluation, and export of the
baseline are unaffected. If you call model.generate() on the MatMul-free
model with a recent transformers release and hit cache or shape errors, apply
equivalent fixes locally or pin a transformers version compatible with
upstream.
hnet_bit/ # the proposed model (BitLinear ops, HGRN blocks, dynamic chunking)
third_party/matmulfreellm/ # MatMul-free baseline (submodule)
train_spanish.py # training entry point
model_factory.py # builds all three architectures
training_config_spanish.py # config, WSD scheduler, optimizer
data_spanish.py # corpus download + byte/BPE tokenization
metrics_spanish.py # BPB, inference memory
export_deployment.py # checkpoint -> compact ternary export
profile_inference.py # prefill/decode profiling
generate_results.py # per-run metadata -> aggregate CSV
test_smoke.sh # CPU smoke test
scripts/ # dataset build, profiling, segmentation, figures, cloud setup
configs/ # per-run training configs
results/ # committed per-run metadata + aggregate CSV + figures data
figures/ # committed PNG/PDF figures
deploy/ # local staging for release assets (gitignored)
@misc{hnetbit2026,
title = {hnetbit: Hierarchical Ternary Byte-Level Language Models},
author = {{The hnetbit Authors}},
year = {2026},
note = {Double-blind submission; the author list will be finalized in the camera-ready version.}
}MIT — see LICENSE. Third-party attributions (MatMul-free LM, H-Net, and the
dataset) are in NOTICE.