Skip to content

Repository files navigation

Speck

SpeckLabs scales in steps, each much larger than the last. This first release is the data step: its product is measured, transferable findings about the pretraining and mid-training data pipeline, published as openly as the licences allow, so later and much larger releases can start from them. A model ladder (50m, 130m and 410m parameters) carries most experiments, and one 1.2B parent model, trained from scratch, carries the decay and mid-training experiments as branches. A fixed SFT recipe probes every branch. The released base model and light reasoning assistant for coding, math and tools are the best of those branches. The architecture, a hybrid of Kimi Delta Attention (KDA) and global attention layers, is held fixed.

Get a working baseline

Python 3.10+ and uv are required. From this checkout:

make setup
make smoke
make quality

The smoke workflow builds tiny local data, trains a hybrid base model, branches it onto masked chat rows, trains an assistant from it, and verifies exact resume at each stage. It also evaluates held-out base loss. It uses CPU only and downloads no corpus.

The ladder configurations are in experiments/ladder; GH200 access follows the qualification runbook.

Working with the project

Task Guide
Prepare and reuse data Data
Train, resume, fine-tune, generate Training
Measure capability and cost Evaluation
Measure throughput on a rented H100 H100 rental
Qualify GH200 access Qualification
Run scheduler-managed training Slurm
Export a checkpoint Releasing
Make a change Contributing

Every command is python -m scripts.<command> (for example scripts.base_train, scripts.sft_train, scripts.infer) and accepts --help. Provide explicit experiment paths. Data, checkpoints and logs live outside Git in the data store, ~/.cache/speck by default; set speck_base_dir to use another location.

PLAN.md        One current direction and next step
speck/         Model, data, training, evaluation, export, and runtime code
scripts/       Maintained command entry points
tests/         Behavioral and integration checks
experiments/   Runnable configurations, design records and result receipts
docs/          Program design, paper outline and operational guides

Source code is MIT licensed and released weights are Apache-2.0; see Releasing. See the citation.

History

Retired code and records remain in Git history and do not govern current work. The tag pre-cleanup-2026-09-25 marks the last large removal and pre-simplification-2026-09-17 the earlier 140M releases. Read an old file with git show TAG:PATH.

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

1 watching

Forks

Contributors

Languages