Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
195 commits
Select commit Hold shift + click to select a range
443ad4a
perf(nvfp4): one reduce launch for the fp32 faces, and cheaper blocks…
tournierjc Sep 29, 2026
2c92ef6
feat: support GLM-5.3 Flash image input on MLX
mgoldwasser Sep 29, 2026
54a9867
fix: validate GLM image geometry and token budgets
mgoldwasser Sep 29, 2026
58fb4d3
fix(cuda): the server lists and answers to --alias, as the MLX server…
philip-pentatonic Sep 29, 2026
d8057db
Keep Flash Next CUDA prompt states one token early
benthecarman Sep 30, 2026
0247cc3
glm5_next: 8-bit group-32 checkpoints on Macs, and a lossless Q8_0 GG…
feni6 Sep 29, 2026
f338e0d
cuda: read block-scaled FP8 linears (ModelOpt FP8_PB_WO) in Flash Nex…
jschmied Sep 30, 2026
9933492
perf(experts): the multi-row item holds 16 pairs, not 64
tournierjc Sep 29, 2026
d23087c
feat: CUDA lanes (prompts inside rounds, caches that grow, one launch…
ashhart Sep 30, 2026
77edf85
docs(glm5_next): wiring and the buffer cache for scripts that load th…
feni6 Sep 29, 2026
e944f02
feat(cuda): score /v1/decisions on GLM
mikolaj92 Sep 30, 2026
4707a6e
docs(glm5_next): the load_backbone wiring note moves to the recipe
ashhart Sep 30, 2026
a42c37c
fix(qwen3_5 cuda): admit the DFlash2 drafter at the 4-bit bytes its l…
jkuepker Sep 29, 2026
7f71536
fix(decisions): isolate GLM score cache and stabilize temperature sof…
mikolaj92 Sep 30, 2026
38239cb
Prepare reviewed main-to-zig sync
CerebralCoding Sep 30, 2026
7ca0a24
fix(decisions): reject ignored template kwargs and verify real GLM TP…
mikolaj92 Sep 30, 2026
88635db
docs(decisions): record full GLM GB10 validation and HTTP timings
mikolaj92 Sep 30, 2026
4447ac3
feat: Prometheus /metrics on both servers (#110), matrix-unit project…
ashhart Sep 30, 2026
3472812
glm5_next: float32 activation mode (config tensorfold_activation_dtype)
feni6 Sep 29, 2026
8fcf99d
fix: a DFlash (v1) drafter with a draft vocabulary drafts chains
cshintov Sep 30, 2026
8222523
docs(decisions): add full Qwen3.8 MLX validation and timings
mikolaj92 Sep 30, 2026
f119334
perf(glm): bound radix selection to visible pools and validate graph …
mikolaj92 Sep 30, 2026
cfd9d86
perf: DFlash (v1) chains sized from measured rounds, drafted on the d…
cshintov Sep 30, 2026
8be7598
docs: a medium the template does not name lands on the nearest level …
EugeneClaw Sep 30, 2026
821393a
test(cuda): the GLM template hears medium as high, as on the Mac
EugeneClaw Sep 30, 2026
ce66881
fix(vision): prefer quantized split projections over retained fused QKV
shantanugoel Sep 30, 2026
ee8bc54
fix(vision): trim EXL3 MLP padding using hashed model config
shantanugoel Sep 30, 2026
d56265c
fix(vision): read snapshot config before resolving weight symlinks
shantanugoel Sep 30, 2026
a41fe08
review: pin the boot default as pass-through, nearest keeps non-ladde…
EugeneClaw Sep 30, 2026
4492349
style(vision): tidy external tower admission
shantanugoel Sep 30, 2026
d3d27d8
Document Flash Next image support in upstream guides
shantanugoel Sep 30, 2026
68c6e35
feat: Mac prompts fill side by side (#90), exact CUDA token probabili…
ashhart Sep 30, 2026
f151c6f
feat(qwen3_5 cuda): --checkpoint-slots sets the prompt states the 27B…
nood-co1 Sep 30, 2026
eca8490
feat(qwen4_exp): apply an n-gram table's weight_scale at lookup inste…
gilby Sep 30, 2026
f3d86c5
fix: stream tool-call arguments from the CUDA server
olexale Sep 29, 2026
0dcc1ff
feat(cuda): bf16 prompts by default (--prefill-fp8 opt-in), Flash Nex…
ashhart Sep 30, 2026
4d9f241
fix: Qwen tool parameters in Python's spelling decode to their schema…
outcastofmusic Sep 30, 2026
b902d0f
style: one-line docstrings and comments across src (AST-identical); t…
ashhart Sep 30, 2026
93d6462
feat(cuda): TENSORFOLD_MEMORY_RESERVE_GIB sets what the startup budge…
MiaAI-Lab Sep 30, 2026
682f3dc
exl3 group kernel: raise the dynamic shared memory limit before launch
akol1 Sep 30, 2026
64e2374
perf(cuda): each prompt plan's item is its expert kernel's (#102, #105)
ashhart Sep 30, 2026
96f8bab
fix(cuda): load consolidated Flash Next EXL3 n-gram tables
shantanugoel Sep 30, 2026
0096b81
fix: n-gram table scales on CUDA (#148), DeepSeek-V4 drafting on Macs…
ashhart Sep 30, 2026
358875c
perf(glm): an idle rank 1 waits for rank 0's next request on the rend…
MiaAI-Lab Sep 30, 2026
94c70c5
style(glm cuda): one-line comment and docstring in the idle rank chan…
ashhart Sep 30, 2026
3f35ea9
fix(glm cuda): the startup estimate counts the EXL3 experts' scratch …
MiaAI-Lab Sep 30, 2026
eaace14
fix(glm cuda): GLM-5.3 keeps prompt prefixes inside the KDA chain (#9…
ashhart Sep 30, 2026
b23c10a
perf(glm cuda): prompt chunks' dense latent pass and pool scores run …
MiaAI-Lab Sep 30, 2026
c401a33
style(glm cuda): one-line docstrings and comments in the row-block pr…
ashhart Sep 30, 2026
b3b8a39
perf(glm): prototype skipping invisible pool tiles with bitwise bench…
mikolaj92 Sep 30, 2026
e254620
style: one-line comments in the visible-pool change (#140) and six do…
ashhart Sep 30, 2026
6673227
fix: the M5 draft head reads the target head at its own group size
gilby Sep 30, 2026
eaf845b
style(drafters): #147's comment on one line and its head matmul withi…
ashhart Sep 30, 2026
17e90a4
feat(glm): TF_GLM_MTP=0|1|auto leaves the MTP head out beside DFlash2…
MiaAI-Lab Sep 30, 2026
bc41790
fix(glm): the MTP head stays beside DFlash2 by default; TF_GLM_MTP=au…
ashhart Sep 30, 2026
7c088eb
perf(glm5_next cuda): DFlash2's block attention reads only its slidin…
MiaAI-Lab Sep 30, 2026
b953378
gemma4: strip a thought channel that opens anywhere, not just at the …
gprot42 Sep 30, 2026
47bf822
feat(cuda): RTX 40 cards (compute capability 8.9); a kept GLM prompt …
ashhart Sep 30, 2026
50588fb
feat(lane_fuse): stack 4-bit groups of 32 too (oQ-style g32 checkpoin…
gilby Sep 30, 2026
e3ac0ea
chore: Apache-2.0 from 0.6.0
ashhart Sep 30, 2026
c464617
release: TensorFold 0.6.0
ashhart Sep 30, 2026
9da0792
fix(cuda): EXL3 grouping stays within the device's shared-memory ceil…
ashhart Sep 30, 2026
9175c42
fix(server): a medium the chat template does not name lands on the ne…
EugeneClaw Sep 30, 2026
26a95a6
fix: a reasoning effort the template doesn't name maps to the nearest…
ashhart Sep 30, 2026
98eb24b
CUDA admission: bound the host by streaming needs, not checkpoint res…
akol1 Sep 30, 2026
87af31a
tests: thinking-off replies through ChatApp and HTTP, Gemma and Harmony
gprot42 Sep 30, 2026
9641547
fix(cuda): a DGX Spark keeps its unified budget, and host loading pea…
ashhart Sep 30, 2026
7337d1d
feat(qwen4_exp cuda): prompt-end entries one token early, message-sta…
lcgutierrez Sep 30, 2026
c8bd385
feat(cuda): Flash Next keeps message prefixes inside its prompt passes
ashhart Oct 1, 2026
3dacc5c
feat(mlx): /v1/decisions scores choice, score, and yes/no from next-t…
mikolaj92 Sep 30, 2026
d56e483
decisions: a Mac decision fills on the prompt lanes and keeps its sha…
ashhart Oct 1, 2026
d95ec8e
feat(cuda): support Flash Next images and cached EXL3 vision assets
shantanugoel Sep 30, 2026
8b2f5b6
Sync Zig review branch with TensorFold 0.6.0
CerebralCoding Oct 1, 2026
cfec158
Flash Next on CUDA: image rows run through the shared prompt passes, …
ashhart Oct 1, 2026
83c5f6f
feat: Qwen3.6 MoE on Macs, the row decoder with DFlash (v1) drafts
cshintov Sep 30, 2026
b3c6ece
qwen4_exp cuda: read FP8 experts in the MTP drafter
jschmied Oct 1, 2026
e2a9ab0
pip-only CUDA builds, grouped sm_12x lane matmuls, kept Flash Next pr…
ashhart Oct 1, 2026
fd36e78
qwen4_exp cuda: requests that arrive while prompts fill are admitted …
jschmied Oct 1, 2026
5df6682
test(flash-next): fair fills with growing prompt pieces
ashhart Oct 1, 2026
1cb1d7e
fix(vision): make image history limits configurable
edurdias Sep 1, 2026
bd5392a
docs: install and upgrade with Homebrew; #170's refusal message withi…
ashhart Oct 1, 2026
d9e77ea
prompt cache: a history boundary for prompts that continue the model'…
gprot42 Oct 1, 2026
e129421
prompt cache: a continued turn resumes from its history boundary
ashhart Oct 1, 2026
902e280
gemma4: strip a spontaneous thought channel when thinking is off
gprot42 Sep 30, 2026
17e6390
server: Flash Next's sparse-attention indexer is priced by allocated …
ashhart Oct 1, 2026
50dfe38
fix(server): a refused POST's body no longer reaches the next request…
SxMShaDoW Oct 1, 2026
e624f7c
perf(27b-cuda): streams plan on their measured round cost, and the dr…
ashhart Oct 1, 2026
b3d8733
feat(server): reasoning_effort max: GLM-5.3's template names it (its …
MiaAI-Lab Oct 1, 2026
bc96a9d
test: an unnamed max uses the nearest named level
ashhart Oct 1, 2026
8613488
glm5_next: route EXL3 experts through the universal path for any bit …
akol1 Sep 30, 2026
d61c0ac
fix(cuda): GLM's EXL3 path owns its scratch
ashhart Oct 1, 2026
3d7a147
qwen4_exp: refuse quantized bytes where the loader reads bf16 values
jschmied Oct 1, 2026
7347e51
fix(flash-next): FP8 draft experts are validated, and preflight keeps…
ashhart Oct 1, 2026
8a836f2
sampling: TENSORFOLD_SEED_SALT moves every prompt-derived seed
jschmied Oct 1, 2026
303faca
perf(cuda): the GDN prompt chain loads each stage one stage ahead (sa…
ashhart Oct 1, 2026
49b2cb8
fix(mlx): mlx-lm 0.32 support: read cache state the same under 0.31 a…
di37 Oct 1, 2026
ed768b7
fix(cuda): the server prints a line a request, as the Mac server does
cshintov Oct 1, 2026
087d482
fix(cuda): a finished reply prints the Mac server's done line
cshintov Oct 1, 2026
b802743
0.6.1 engine work: NVFP4 checkpoints in their own math, prompts that …
ashhart Oct 1, 2026
ba1ab0e
test(cuda): the failed-admission stand-ins run as on a host box on a …
ashhart Oct 1, 2026
17c73e1
release: TensorFold 0.6.1
ashhart Oct 1, 2026
05b25ea
fix(snapshots): a failed or interrupted write leaves no .partial.safe…
di37 Oct 1, 2026
28a6ae1
fix(flash next): accept an FP8 n-gram table in a MIXED_PRECISION NVFP…
plotarmordev Oct 2, 2026
f9e92d7
fix(glm cuda): report the drafts' counts, so /health, /metrics and re…
plotarmordev Oct 2, 2026
7a1c776
fix(server): the client-gone check sees descriptors past 1023
jayleaton Oct 1, 2026
14c90bf
fix(gemma): qkv_rows reserves its 1024 threads, so M1/M2 pipelines ta…
di37 Oct 2, 2026
552390f
fix(kernels): one GPU-generation reading; a VM's GPU or a CPU default…
di37 Oct 1, 2026
60656c6
0.6.2 engine work: Flash Next on Macs at 64k-128k, the 27B's GDN tree…
ashhart Oct 2, 2026
064c93c
docs(readme): --mtp-confidence defaults to 0.70; --decode-share also …
plotarmordev Oct 2, 2026
0d09c33
fix(glm cuda): a mixed-bit EXL3 checkpoint is refused by name instead…
ashhart Oct 2, 2026
56e2e3e
release: TensorFold 0.6.2
ashhart Oct 2, 2026
ceb27e5
feat(qwen4_exp): TF_FLASH_DENSE=matrix runs every affine width on the…
gilby Sep 30, 2026
3a3ddf2
style: keep the Flash Next decode file inside the line limit
ashhart Sep 30, 2026
5a99342
perf: load a group of 128 scales once in the dense matrix kernel
ashhart Oct 1, 2026
215a2c9
tools: score a decode forward against an fp32 matmul
ashhart Oct 1, 2026
dbc0158
tools: tile the chat block out to two full windows
ashhart Oct 1, 2026
3651328
perf: default the fused GDN stack to the matrix kernel
ashhart Oct 1, 2026
06831ab
perf(flash-next): the one-stream DeltaNet chain keeps its conv weight…
ashhart Oct 1, 2026
7f1060b
fix(server): send the spec's usage chunk before [DONE] when stream_op…
ashhart Oct 2, 2026
cb77df1
test: pin the usage chunk's place in a stream on both servers
ashhart Oct 2, 2026
d58855d
docs: name the usage chunk stream_options.include_usage asks for
ashhart Oct 2, 2026
7f4e806
perf(cuda): past --checkpoint-slots, a kept state its conversation ha…
nood-co1 Oct 1, 2026
6bb19be
fix(server): render malformed tool-call history safely
salmanarshad321 Oct 2, 2026
75610f2
style: one-line docstring for the tool-call argument normalizer
ashhart Oct 2, 2026
8cca6e3
Add Anthropic Messages API on MLX and CUDA
kky42 Oct 2, 2026
657f479
Honor Claude Code thinking toggle and report thinking tokens
kky42 Oct 2, 2026
ba18743
Preserve mid-conversation system messages for Claude Code
kky42 Oct 2, 2026
9be9eaa
feat(server): /health carries the live line's numbers (connections, w…
ashhart Oct 2, 2026
8e5376b
cuda: drop the admission reserve, the grant is free memory
eleqtrizit Sep 29, 2026
5089b57
cuda: size Nemotron-H's workspace like the engine builds it
eleqtrizit Sep 29, 2026
dabd7b9
cuda: account for Nemotron's serial state within the memory budget
eleqtrizit Sep 30, 2026
a121303
fix(cuda): a discrete GPU's grant stays the card's free memory
ashhart Oct 2, 2026
7b7e3b1
style: one-line docstrings for the Nemotron window test
ashhart Oct 2, 2026
8b2d1b6
fix(cuda): a unified GPU's grant keeps a floor under its free memory
ashhart Oct 2, 2026
38ed83e
docs: name the Hugging Face org TensorFold where model ids are written
ashhart Oct 2, 2026
2c62a37
refactor(cuda): name the CUDA cap cuda_limit_bytes and size the reser…
ashhart Oct 2, 2026
655a72b
fix: a TensorFold model id finds a cache pulled under the old Vontra …
ashhart Oct 2, 2026
bd28cab
Flash Next CUDA: /v1/decisions scores labels from the prompt's last l…
Mirrdhyn Oct 2, 2026
44d257f
fix(cuda): queue decision groups atomically within the live prompt bu…
ashhart Oct 2, 2026
b556604
feat: a launchd service manager and a terminal control room
ashhart Oct 1, 2026
617d493
fix: uninstall enables the launchd label
ashhart Oct 1, 2026
789619a
fix: the smoke runner loads its reserved profile name
ashhart Oct 1, 2026
4438f6e
test: a service profile keeps the memory limit and serve arguments
ashhart Oct 2, 2026
a389f4c
fix: render the control room logo as 24-bit half blocks
ashhart Oct 2, 2026
ad15127
Hear chat_template_kwargs.thinking as enable_thinking; GLM-5.3 keeps …
MiaAI-Lab Oct 2, 2026
977f2cc
fix(http): decode chunked request bodies before JSON parsing
JordiPosthumus Oct 2, 2026
80f401c
feat(control): the dashboard shows the server's live decode and prefi…
ashhart Oct 2, 2026
8f5e680
Answer /tokenize and /detokenize on both servers; name the field in c…
MiaAI-Lab Oct 2, 2026
25770ab
test: the control room key tests run without an async plugin
ashhart Oct 2, 2026
69d0ca9
fix(server): image parts in tool-result messages are accepted, render…
MiaAI-Lab Oct 2, 2026
c2b9045
Give CUDA Qwen images a larger shared visual-token budget (--vision-i…
MiaAI-Lab Oct 2, 2026
5d0bba0
feat(vision): --vision-offload keeps the CUDA image tower in host RAM…
barelyworkingcode Oct 1, 2026
aaecbde
feat(flash-next cuda): TENSORFOLD_PREFILL_ROWS sets the prompt piece …
MiaAI-Lab Oct 2, 2026
0a0110a
perf(nemotron): on M5 a lone stream's copies widen to 64-row windows,…
ashhart Oct 2, 2026
c661718
perf(nemotron): experts group by expert from 8 rows, routed and group…
ashhart Oct 2, 2026
5fafca9
tools(nemotron): profile steps and check the expert kernel
ashhart Oct 2, 2026
48fc0f4
perf(nemotron): a lone window above 16 rows stores its Mamba states e…
ashhart Oct 2, 2026
588921b
feat(vision cuda): Flash Next takes video input
MiaAI-Lab Oct 2, 2026
05cf17f
perf(cuda nvfp4): prompt rows sum a tile's K slices in one block (sam…
jschmied Oct 2, 2026
ec131fd
perf(nemotron): on M5 the shared expert's up projection applies relu2…
ashhart Oct 2, 2026
5e1b9af
Integrate chat thinking controls
ashhart Oct 2, 2026
45325f7
Integrate tokenize and detokenize routes
ashhart Oct 2, 2026
dc8e9eb
Integrate chunked request decoding
ashhart Oct 2, 2026
4825c01
Integrate tool-result images
ashhart Oct 2, 2026
e9fade0
Integrate the Anthropic Messages API
ashhart Oct 2, 2026
362610f
Integrate the CUDA visual-token budget
ashhart Oct 2, 2026
6091427
Integrate CUDA video input
ashhart Oct 2, 2026
730bd99
Integrate vision tower offload
ashhart Oct 2, 2026
32fcdb7
Integrate configurable Flash Next prompt rows
ashhart Oct 2, 2026
4bf9c07
Integrate fused NVFP4 prompt slice sums
ashhart Oct 2, 2026
ba4a04d
Integrate conversation-aware checkpoint eviction
ashhart Oct 2, 2026
41bcf81
Integrate CUDA admission and Nemotron memory accounting
ashhart Oct 2, 2026
0dc4ab4
Integrate TensorFold model names and existing-cache lookup
ashhart Oct 2, 2026
22b5303
Integrate the terminal control room and service manager
ashhart Oct 2, 2026
b2937d6
Integrate Nemotron window, expert and state kernels
ashhart Oct 2, 2026
ae2467f
fix(flash-next cuda): refuse mismatched prompt rows across ranks
ashhart Oct 2, 2026
936963e
fix(cuda): charge live allocations against the absolute cache budget
ashhart Oct 2, 2026
ced6139
fix(server): integrate request framing and preserve tool image attrib…
ashhart Oct 2, 2026
4a5ac25
Integrate Flash Next decisions and admission fixes
ashhart Oct 2, 2026
dd8e67d
test(cuda): refresh rank fixtures and prove structured GLM context re…
ashhart Oct 2, 2026
794a0b6
fix(cuda): keep lone Flash Next cache growth within the explicit budget
ashhart Oct 2, 2026
2b6400a
style: keep the cache growth invariant on one line
ashhart Oct 2, 2026
a5f518b
revert(cuda): keep the previous prefix-cache eviction policy for 0.6.3
ashhart Oct 2, 2026
6245745
fix(exl3): split-K reduction faults on a layer with no bias
barelyworkingcode Oct 1, 2026
ac524e5
test(exl3): a layer with no bias agrees with a zero bias across split…
ashhart Oct 2, 2026
22b1c22
fix(qwen3.6 cuda): preserve MTP head format and refuse incompatible s…
BHCC2025 Oct 2, 2026
c9259c3
feat(metrics): /metrics times each request's decode, so a scraper can…
sxuff Oct 2, 2026
7e5b8ca
test(cuda): collect format checks only with their runtime dependencies
ashhart Oct 2, 2026
fd58a89
fix(cuda): a discrete GPU's grant keeps the card's floor, as before 0…
ashhart Oct 2, 2026
aed0d63
test(cuda): retain staging and growth checks under the discrete memor…
ashhart Oct 2, 2026
0d499df
docs: a contributor guide and a pull request template
ashhart Oct 2, 2026
9356df5
release: TensorFold 0.6.3
ashhart Oct 2, 2026
ae5c87a
Sync Zig review branch with TensorFold 0.6.1
CerebralCoding Oct 2, 2026
3b8f336
Align native kernels and sync checks with TensorFold 0.6.3
CerebralCoding Oct 3, 2026
df1684b
Pin stable Zig 0.17.0 and reuse the installed compiler
CerebralCoding Oct 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
27 changes: 27 additions & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
<!-- Read CONTRIBUTING.md first. A small pull request with a complete receipt lands fastest. -->

## What this changes

<!-- One or two sentences, and the issue it closes if there is one. One change to a pull request. -->

## Receipt

<!-- A change to documentation alone needs no receipt. -->

- Environment, with the TensorFold commit, MLX or PyTorch version, machine and GPU, checkpoint and revision:
- Exactness, from `tools/bench_concurrent.py --alone --serial` or `token_sha` against `"draft": false`:
- Decode speed before and after, from `tools/bench_openai.py`:
- Prompt speed before and after, from `tools/prefill_cold.py`:
- Tests run, with pass, fail and skip counts:
- Not run, and why:

## Checklist

- [ ] Every token still goes through the lane rounds. No serial path, and nothing that needs drafts off.
- [ ] Drafted output equals `"draft": false`, a resumed prompt equals a fresh one, and concurrent equals solo.
- [ ] No precision traded for speed. If the bits change, the description says which sums change and why.
- [ ] Prompt processing is no slower than the last release.
- [ ] The description says which platforms I ran, Metal M1 to M5 and CUDA, and which I could not.
- [ ] New tests fail before the change, pass after it, and skip cleanly without their dependency.
- [ ] Comments and docstrings are one line. No measurements or history in the source.
- [ ] No personal data, machine names, internal hosts or local paths. No AI attribution lines.
2 changes: 1 addition & 1 deletion .github/workflows/native-macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ name: Native macOS bootstrap

on:
push:
branches: [feat/zig]
branches: [zig/review-sync, zig/native]
pull_request:
workflow_dispatch:

Expand Down
2 changes: 1 addition & 1 deletion .zig-archive.sha256
Original file line number Diff line number Diff line change
@@ -1 +1 @@
98faf6b77e3d681bb27bd65b81dcfc18ee21c437ebb93504eb16731c613fd792 build/toolchains/zig-aarch64-macos-0.17.0-dev.2248+3f6a02acd.tar.xz
b607e9b9234790a008116ae5bdb71c6243b84b9fb42a53a9e70fde41c06c536a build/toolchains/zig-aarch64-macos-0.17.0.tar.xz
2 changes: 1 addition & 1 deletion .zig-version
Original file line number Diff line number Diff line change
@@ -1 +1 @@
0.17.0-dev.2248+3f6a02acd
0.17.0
120 changes: 119 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,124 @@
`tensorfold update` prints the sections below that are newer than the version you had. Each release's page on
GitHub has the full notes and the measurements behind them.

## 0.6.3 (2 Oct 2026)

- **Nemotron on M5 Macs: copied text verifies up to 64 tokens a round.** A lone stream's copy window grows from 16 to
64 rows while each lands whole, a window attends in one call, and wide windows keep their Mamba states every 8th
row. On an M5 Max, one stream editing code decodes about 1.4x faster (550 to 760 tok/s, and up to 1,000 on a cool
machine), prose and code writing gain 1.6-3.3%, and drafted replies still equal one-token decoding.
- **Anthropic Messages API.** `/v1/messages` and `count_tokens` on the Mac and CUDA servers, with streaming, tools,
thinking and cache usage (#223, closes #168). Thanks to @kky42.
- **Control Room.** `tensorfold service` runs a model as a launchd service, and `tensorfold tui` shows its decode and
prefill speed and its connections live.
- **Flash Next on Macs before M5:** `TF_FLASH_DENSE=matrix` runs every dense projection on the matrix units, and the
fused DeltaNet stack uses the matrix kernel by default (#149). Thanks to @gilby.
- **Vision:** `--vision-offload` keeps the CUDA image tower in host RAM between images (#187, closes #185); image
parts are accepted inside tool results (#235); `--vision-image-tokens` lets many images share a larger budget
(#239); Flash Next on CUDA takes video (#240). Thanks to @barelyworkingcode and @MiaAI-Lab.
- **CUDA:** `/v1/decisions` on Flash Next scores labels from the prompt's last logits, the questions filling in one
pass (#232); `TENSORFOLD_PREFILL_ROWS` sets the prompt piece rows (#238); NVFP4 and FP8 prompt rows add a tile's K
slices in one block, with the same bits (#242); `TENSORFOLD_MEMORY_RESERVE_GIB` moves the memory floor the startup
keeps free, by default a tenth of the pool and at least 4 GiB as before (#165); Flash Next's chain kernel takes
5-17% less time at 2 to 8 rows, with the same bits. Thanks to @Mirrdhyn, @MiaAI-Lab, @jschmied and @eleqtrizit.
- **Server:** chunked request bodies are decoded before JSON parsing (#244); `/tokenize` and `/detokenize` carry
vLLM's fields (#237); `chat_template_kwargs.thinking` is heard as `enable_thinking`, and GLM-5.3 keeps earlier
turns' reasoning (#236); streaming usage gets its own chunk when `stream_options.include_usage` asks for it (#216);
malformed tool-call history renders safely (#233); `/metrics` times each request's decode (#269); `/health` carries
the live decode and prefill speed. Thanks to @JordiPosthumus, @MiaAI-Lab, @salmanarshad321 and @sxuff.
- **Fixes:** an EXL3 layer with no bias no longer faults in the split-K reduction (#186); an 8-bit `lm_head` keeps its
format in Qwen3.6's MTP draft head on CUDA (#270); a lone Flash Next stream's cache growth on CUDA stays inside the
explicit memory budget; two ranks refuse to start with different prompt rows. Thanks to @barelyworkingcode and
@BHCC2025.
- **Models moved to the TensorFold Hugging Face org.** `Vontra/<name>` ids redirect, and the tree now names
`TensorFold/<name>`.
- **Contributing.** `CONTRIBUTING.md` says what a pull request needs to land and how it lands, and new pull requests
open with a receipt template.

## 0.6.2 (2 Oct 2026)

- **Flash Next on Macs at 64k-128k.** On an M3 Ultra, one stream runs 1.2-3.4% faster at 64k and 3.9-5.5% at 128k,
with the same tokens: a window's n-gram ids are hashed on the GPU, and the chain's first step is built while the
GPU verifies.
- **27B with several streams on CUDA.** The GDN tree kernel takes 8-35% less time with the same bits. On an RTX PRO
6000 at its 250 W limit, 4 and 8 streams of the NVFP4 27B decode 1.1-4.2% faster.
- **Fixes:** the config check accepts the FP8 n-gram table in NVIDIA's MIXED_PRECISION Flash Next export; GLM-5.3
on CUDA counts its drafts in `/health`, `/metrics` and replies, and names a mixed-bit EXL3 checkpoint when it
refuses one; a client that leaves is noticed past file descriptor 1023; the CUDA server prints a line a request, as
the Mac server does; a failed snapshot write no longer leaves its partial file; Gemma 4's QKV kernel reserves its
1024 threads for M1, M2 and macOS VMs, and a VM's GPU is no longer taken for an M5.

## 0.6.1 (1 Oct 2026)

- **NVFP4 checkpoints in their own math.** `nvidia/Qwen3.8-27B-NVFP4` runs the 4-bit activations its checkpoint
names, as vLLM does; `--precision full` runs 16-bit activations against the same weights. On an RTX PRO 6000 at its
250 W limit, one stream decodes 1.4-2.0x vLLM and prompts fill at 0.95-0.97x its speed.
- **Waiting prompts fill together on CUDA.** With `--parallel`, prompts that arrive together now share one prefill
forward instead of filling one a round: on an RTX PRO 6000 at its 250 W limit, 8 streams of the 27B run 1.14-1.36x
faster and the slowest first token comes in 0.05-0.10 s instead of 0.4-0.8 s (1.6 s to 0.15 s on a DGX Spark), with
the same replies.
- **More 27B tokens with several streams on CUDA.** Streams plan on their measured round cost, and wider lane blocks
on RTX PRO and RTX 50 cards add 4-8% at 8 streams, with the same tokens.
- **Flash Next on Macs at long context.** One stream runs up to 9.6% faster on code at 64k and 10.2% on chat at 128k
on an M3 Ultra, with the same tokens.
- **`/v1/decisions`** scores a choice, a score or a yes/no from the next-token logits, on the shared prompt lanes.
- **Flash Next on CUDA:** image input, forks that resume from their shared prefix, shared system prompts copied
instead of filled again, and short prompts admitted while a long one fills.
- **RTX cards without Docker:** pip alone installs and builds the CUDA kernels. Native Windows is in as an
experimental host layer, not yet run on Windows hardware.
- **Fixes:** a refused request no longer breaks the next one on its connection; Gemma 4 thought blocks stay out of
replies with thinking off; an unnamed reasoning effort goes to the nearest named level; mlx-lm 0.32 support.

## 0.6.0 (30 Sep 2026)

- **RTX 40 cards.** CUDA now runs on compute capability 8.9 (Ada). On one RTX 4090 the 27B serves with DFlash2 in a
40,182-token window, exact: drafted replies equal serial ones, and a resume or resend gives a fresh run's reply.
Prompts fill at 2.4-2.6k tok/s from 2k to 32k tokens, and replies decode at 64-86 tok/s. On one GPU the 27B's kept
prompt states give way, oldest first, when a live reply needs the room, so admission counts one live window.
- **Prompts fill inside the decode rounds.** With `--parallel` on CUDA, Flash Next prefills a queued prompt in the
same forward as the live replies instead of stopping them: on one DGX Spark, first tokens came 2.8-3.1x sooner than
on 0.5.0. On Macs several prompts fill side by side, the fewest tokens left first: on an M3 Ultra, short requests
queued behind a long prompt got their first token in a median 9.9 s instead of 113 s (the slowest 11.7 s, not 130).
- **Conversations resume on more engines.** An identical resend or the next thinking turn now resumes from the kept
prompt state on Flash Next, Qwen3.6, Nemotron and GLM-5.3 on CUDA, as the 27B did, with a fresh run's reply: an
18.7k-token Flash Next resend went from 8.3 s to 0.08 s. On Macs the prompt cache keeps each conversation's newest
checkpoint and grows into memory the model leaves idle.
- **CUDA prompts run at bf16 by default.** Against an fp32 reference the 27B's prompt rows are 20x closer than with
FP8 (KL 0.0031 against 0.0624). `--prefill-fp8` keeps 0.5.0's faster FP8 prompts for those who want them.
- **Tool calls for agents.** The CUDA server streams tool-call arguments as the model writes them (the longest
silence in a long call fell from about 30 s to half a second), Python-spelled values like `False` and `None`
decode to their schema types, Gemma 4's bare tool calls parse (#121), and a prompt past the context window gets
OpenAI's `context_length_exceeded`, so clients compact instead of retrying.
- **More checkpoints.** GLM-5.3 on Macs reads 8-bit and Q8_0 GGUF checkpoints and takes images, and has an opt-in
float32 activation mode (bf16 stays the default). On Macs, Flash Next loads oMLX's oQ checkpoints with scaled
n-gram tables; on CUDA it reads consolidated EXL3 tables and block-scaled FP8 linears.
- **Faster.** Flash Next NVFP4 decodes 4.9-7.1% faster and fills prompts 15-16% faster at 2k-16k, bit-identical. On an
M5 Ultra the 27B with DFlash2 on an oQ4e checkpoint went from 33 to 131-160 tok/s. Before M5, a lone stream's copy
windows widen to 128 rows (an M3 Ultra edit ran 190 -> 233 tok/s), and Flash Next's concurrent rounds use the matrix
units. GLM-5.3's DFlash2 reads only its sliding window: 8% faster at 52k, and 3.8 GiB lighter a rank on two Sparks.
- **Operations.** Prometheus `/metrics` on both servers, `TENSORFOLD_MEMORY_RESERVE_GIB` for the CUDA startup
reserve, `--checkpoint-slots` for the 27B's concurrent decoder on CUDA, and an idle GLM rank no longer spins.
- **Apache-2.0.** TensorFold is licensed under the Apache License 2.0 from this release, which adds an explicit patent
grant from contributors. Releases up to 0.5.0 stay MIT, and code written before 0.6.0 keeps its MIT notice in
`LICENSES/MIT.txt`.
- **Community pull requests.**
- Resumed resends and thinking turns on Flash Next (#124). Thanks to @benthecarman, and to @Arminova for the
GB10 measurements and the third-resend test.
- Streamed tool-call arguments on CUDA (#114), `--alias` on CUDA (#111) and Python-spelled tool parameters
(#135). Thanks to @olexale, @philip-pentatonic and @outcastofmusic.
- GLM-5.3 on Macs: image input (#101), 8-bit and Q8_0 GGUF checkpoints (#119), float32 activations (#118) and
the backbone wiring notes (#120). Thanks to @mgoldwasser and @feni6.
- GLM-5.3 on two Sparks: lighter prompt buffers, the EXL3 estimate, the MTP setting, the idle rank and DFlash2's
ring (#128, #129, #131, #132, #134), visible-pool selection (#140), and the CUDA startup reserve setting (#133).
Thanks to @MiaAI-Lab and @mikolaj92.
- Flash Next: scaled n-gram tables (#148, reported in #142) and the M5 draft head fix (#147), consolidated EXL3
tables (#145), block-scaled FP8 (#126) and the NVFP4 expert speedups (#102, #105). Thanks to @gilby,
@cwschroeder, @shantanugoel, @jschmied and @tournierjc.
- On M5 Macs the 27B fuses the projections of 4-bit group-32 checkpoints such as oQ4e too (#164): four streams went
from 315-320 to 332-334 tok/s on an M5 Ultra, with the same bits. Thanks to @gilby.
- `--checkpoint-slots` (#125), and the DFlash2 drafter's admitted size (#112, from @jkuepker's ROCm work). Thanks
to @nood-co1 and @jkuepker.

## 0.5.0 (29 Sep 2026)

- **OpenAI's Responses API on both servers.** `/v1/responses` runs as a chat completion: every response has its chat
Expand Down Expand Up @@ -249,7 +367,7 @@ GitHub has the full notes and the measurements behind them.
## 0.3.2 (26 Sep 2026)

- `tensorfold update` installs the newest release.
- GLM-5.3-Flash reads Mia-AiLab's EXL3 weights on two DGX Sparks (experimental).
- GLM-5.3-Flash reads Brandon M. Music's EXL3/TR3 weights (re-hosted by Mia-AiLab) on two DGX Sparks (experimental).

## 0.3.1 (26 Sep 2026)

Expand Down
124 changes: 124 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# Contributing to TensorFold

Thank you for wanting to make TensorFold faster. This page says what a pull request needs to land in the next
release, and what happens to it after you open it. A small pull request with its measurements attached lands
fastest.

## Before you write code

- Read the pinned issue "Landing next" and the open pull requests. Work that is integrated for the next release is
listed there before it reaches `main`, so you don't build it twice.
- Open an issue first for a large change: a new backend, a GPU generation, a model family, or anything that changes
the arithmetic a model runs. We will say what it needs to land before you spend the time.
- Keep one change to a pull request. A fix and an unrelated speed-up land separately. In a stack of dependent pull
requests, say which one each sits on.

## What every change must keep

1. **Lanes.** Every model decodes through the shared lane rounds, where one forward verifies the drafted tokens
together. No serial path, no feature that works only with drafts off, no second batcher beside the engine.
2. **Exact output.** A drafted reply equals the same engine's `"draft": false` reply, token for token. A resumed
prompt equals a fresh one, and each concurrent stream equals its solo run. The README's "Exact decoding" section
has the contract.
3. **No precision traded for speed.** No lower-precision activations or sums, no skipped keys, no pruned tokens. A
reorder, or another kernel at equal or better precision, is fine. Say which sums change and that the bits change.
A precision mode that a checkpoint format defines is a separate conversation, so open an issue.
4. **Prompt processing no slower.** Measure cold prompts against the last release for any change that touches
prefill, a kernel, the engine, a family or the server.
5. **Every platform it touches.** That is Metal on Apple Silicon from M1 to M5, and CUDA. Say which you ran and
which you could not.
6. **Lean code.** One job to a module, files under about 600 lines, and one-line comments and docstrings that say
what the code can't. Measurements, history and benchmark tables belong in the pull request, not in the source.

## The receipt

Paste these into the pull request. We rerun what we can, and a complete receipt lets a small change land on review
without waiting for a machine. A change to documentation alone needs no receipt.

- **Environment.** The TensorFold version or commit, the MLX or PyTorch version, the machine and GPU, the checkpoint
and its revision.
- **Exactness.** Drafted against serial, and concurrent against solo, by token hash:

```
python3 tools/bench_concurrent.py http://127.0.0.1:8080 MODEL --alone --serial
```

For a server change, also send one request twice, the second with `"draft": false`, and compare `token_sha`.
- **Decode speed**, before and after, on the same machine in the same session:

```
python3 tools/bench_openai.py http://127.0.0.1:8080 MODEL
```

- **Prompt speed**, before and after. The tool builds cold prompts of 2k, 8k, 16k, 32k and 64k tokens:

```
python3 tools/prefill_cold.py build MODEL_DIR prompts.json
python3 tools/prefill_cold.py run http://127.0.0.1:8080 MODEL prompts.json out.json
```

- **Tests** you ran, with the pass, fail and skip counts.
- **What you did not run**, and why. We need this as much as the rest.

When the difference is a few percent, alternate the runs: before, after, after, before. One run each way cannot tell
a 2% change from a warm machine.

## Running the tests

TensorFold needs Python 3.11 or newer.

```
python -m pip install -e '.[test,tui]'
python -m pytest -q tests
```

On a Mac that is the host suite. The CUDA tests are `tests/cuda` and `tests/test_cuda_*.py`. Those that need PyTorch
or a GPU skip without one.

A new test skips cleanly where its dependency is missing. Use `pytest.importorskip`, so the suite still collects on
a machine without PyTorch or without MLX.

A test should fail before your change and pass after it. Say so in the pull request.

## Commits and identity

- Your commits keep your authorship. We record every author under a GitHub noreply address, such as
`12345+yourlogin@users.noreply.github.com`, so no personal email lands in the history. "Keep my email addresses
private" in your GitHub email settings does this for you.
- No AI tool attribution lines, and no co-author trailers for tools. We remove them when we land a commit.
- No personal data, machine names, internal hosts or local paths in code, tests, fixtures or commit messages.
- A data file built from text, such as an n-gram table, comes from public text only. Name the corpus in the pull
request.

## How a pull request lands

We review it, port it onto the release branch ourselves and run it through the release checks. We do not ask you to
rebase. If `main` has moved, we carry your commits forward.

Your change lands on `main` as your own commit, under your name, inside a release. We then close the pull request
with a comment that links the release and says what we changed on our side, if anything. GitHub may show the pull
request as closed and not merged, because the commit that lands is a port of yours. It still carries your name, and
you appear in the repository's contributors.

Our changes on top of yours, such as a style pass or an extra test, go in separate commits under our name.

When a pull request overlaps work that is already integrated, we close it with a note that says where the work lives
and credits what your report or measurements added. That is not a rejection, and the "Landing next" issue exists to
make it rare.

## What we can check ourselves

We test on an M3 Ultra, an M5 Max, one DGX Spark and two together, and an RTX PRO 6000. A change for other hardware
needs a complete receipt, because we cannot rerun it, and it takes longer to land. Keep such a change in its own
files where you can, so it cannot affect the paths we do test.

## Reporting a bug

Open an issue with the TensorFold version, the machine, the checkpoint and its revision, the exact command, the
startup lines the server printed, and what you expected against what happened. A request that reproduces it is worth
more than a description.

## License

TensorFold is Apache-2.0. A contribution you submit is under that license, as its section 5 says. Model weights keep
their own licenses.
Loading