Skip to content

Prepare reviewed main-to-zig sync - #130

Closed
CerebralCoding wants to merge 195 commits into
ashhart:zigfrom
CerebralCoding:zig/review-sync
Closed

CerebralCoding wants to merge 195 commits into
ashhart:zigfrom
CerebralCoding:zig/review-sync

Conversation

@CerebralCoding

Copy link
Copy Markdown

Replace fork setup instructions with TensorFold’s long-running zig workflow. Manual sync discovers the TensorFold SSH remote, prepares an uncommitted main merge, aligns dependencies and kernels, and supports resuming verification at the original target. It refuses protected branches, outdated bases and unfinished operations; it never rebases or pushes.

The native HTTP server currently interleaves independent request forwards with separate caches. Shared lane rounds remain unimplemented; the README now states this explicitly.

Validation passed: seven offline Git tests, native baseline and checkpoint-file checks, setup/dependency/coverage checks, and safe-mode CLI compilation.

tournierjc and others added 30 commits September 29, 2026 10:01
… in _fp4mm

The reduce: a split-K fp32 face summed its slices with `out.copy_(part[0]); for s: out += part[s]` - one
elementwise launch a slice, so a face the shape splits 32 ways paid 31 of them. Two faces a layer take that
path (`forward._mm` with `f32=True`), and it is their whole reduce: the bf16 face next door already ran in
one launch. `_reduce` gains the fp32 face (F32), the callers use it, and the loop is gone. The adds are the
loop's own, in the loop's order, one fp32 add a slice, so the fp32 sums are bit-identical to the loop's -
the order the split-K contract is written around. `test_matmul_splitk_sum_order_is_the_reduces_one` now pins
both faces: the bf16 one rounds each partial sum as the loop does, the fp32 one does not round at all.

The blocks: two things a step did that it does not need to do 160 times a program. The address
`(b // 4) * 32 * SBN + ((b % 4) * 16) // 2 * SBN` is a staircase in the step (a stored macro block holds 8
rows of the tile), and a load whose address the pipeliner cannot see as affine is not put in flight; it is
the same address as `b * 8 * SBN` (16 * SBN for the pattern tables, whose macro block holds 16 rows). The
decode: the missing nibble is the byte's own parity, so `(w >> ((r16 % 2) * 4)) & 0xF` reads it where a
compare and a select took two ops, and `_e2m1_pattern` already builds the 16 bits a bf16 holds, so
bitcasting those straight into bf16 keeps exactly what widening to fp32 then rounding back gave, three ops
less an element.

Measured on the Spark by timing the grouped step against variants of its own source with each decode
replaced by a cast (dev/deqcost.py): the E2M1 and E4M3 decodes together are 21% of a grouped step at the
served geometry (7 rows x 11 slots, 72 distinct experts). The shave is worth taking; the weight traffic is
what the step actually waits on.
… does

`tensorfold serve --alias` only reached the MLX server. On CUDA the flag parsed and was dropped, so /v1/models
listed --name alone and every reply named --name. The CUDA app now takes the aliases: /v1/models lists --name,
then each alias once, in the MLX server's order, and a reply names the id the request asked for when this
endpoint answers to it, else --name. Ported onto 0.5.0's HTTP module (cuda/http.py) from ashhart#111.
Kept entries ended at the prompt's end, and a resume needs a strictly
shorter prefix, so an identical resend or a next turn with thinking on
(which renders the last prompt token differently) always prefilled in
full. Keep the entry at len(prompt) - 1, as the 27B engine does, cut
inside the chunk that holds it.
…t checkpoints

Lane matmul mode FP8G: e4m3 bytes in the FP8 GEMM's fragment order and one fp32 scale per
(64 inputs, column), applied after each 64-input stage's bf16 MMAs, so the stored weight is exact
and rows stay independent of the row count. Fp8BlockLinear decodes on it and runs prompts on the
FP8 prompt GEMM with the scales as bf16 group scales. The Flash Next loader maps weight_scale_inv
linears to it, runs stacks that mix block FP8 with bf16 through Concat, and dequantizes a
block-FP8 lm_head (and the draft head's rows) with its block scales. format.scheme and
format.dequant learn the layout; the family's check accepts FP8_PB_WO.

Tests: the linear against an fp64 reference, rows, prompt chunks and Concat; the numpy
dequantizer; a tiny ModelOpt checkpoint with block-FP8 linears through the loader, the engine and
drafted-equals-serial decoding.
The kernel holds 16 pairs an item in the decode form and 64 in the prefill form, and nothing between. Measured at a prompt's own row count on the published checkpoint's shapes (2275 rows x 10 slots over 512 experts): 16 pairs reads 1788 items in 14.09 ms, 64 reads 549 in 21.00 ms. The wide item reads three times fewer bytes and is 49% slower, because those re-reads are L2 hits anyway (351 GB/s is above the box's DRAM) and all it bought was three times fewer items competing for the SMs. Served on ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 with the 8-bit projection copies: prose 32.2 -> 40.0, code 63.6 -> 75.8, prefill 1196 -> 1486. Decode moves with it because the arithmetic a prompt runs is the one a verify window runs.

Tests: tests/cuda 734 passed, 75 skipped, 0 failed; the new case pins the item and checks a prompt's plan stays inside max_items for it - the ceiling and the count are easy to confuse, and only the count is traffic.
… a kernel a round) and wide copy windows for one stream

- CUDA --parallel: a queued prompt fills inside the decode rounds; a round runs every stream's DeltaNet and attention in one
  launch; caches grow with use; --decode-share sizes a prompt pass
- One stream's copy windows ramp to 128 rows on M1-M4 and on a GB10 (the 27B too); a broken copy halves the window;
  TF_COPY_ROWS sets the widest
- Mac replies and /v1/responses name the model id the request asked for
Both ranks prefill the question and reduce the vocabulary shards into label probabilities. Host tests cover the prompt wording, the refusals, and that reduction.
…projections for 5/6/8-bit Flash Next before M5, admission fixes

- /metrics: running and waiting requests, token totals, KV use, MTP drafts, request latency and time to first token (ashhart#110)
- Flash Next before M5: 5-, 6- and 8-bit projections of 4+ rows on the matrix units, with the per-row arithmetic
- Mac admission: one need for the window's fit and every admission, and a stream's probed growth is its caches' own bytes
- GLM-5.3 hears medium as high (ashhart#117); Gemma's bare tool calls parse; a null sampling field keeps the server default
- Two CUDA ranks meet on the store after loading and name a rank that never finishes (ashhart#107); DFlash2 admission and loading
  share one packing rule
propose() took the draft-vocabulary path first, which reads DFlash2's
candidate selector; z-lab/Qwen3.6-35B-A3B-DFlash has a draft
vocabulary and no selector, so its first round raised.
…raft vocabulary

A v1 head gives no per-draft chances, yet Qwen35Family always exposed
draft_probabilities, so rounds drafted depth 15 and were trimmed with
forward-only costs that ignore the drafter's time. block_chain also read
the full 248k-row head instead of the draft vocabulary rows.
ashhart and others added 23 commits October 2, 2026 15:59
…ution

Route Anthropic and Mac tokenizer requests through the shared HTTP body reader, retain image parts in their tool results, and keep the video tokenizer test compatible with Python 3.11.
Preserve the one-stream startup allowance when its real growth fits, wait for filling prompts to free room, and refuse isolated over-budget growth without losing request slots.
Revert merge 6d4a11c3666427c666388dd2aa84e986fc8c7e63: the corrected policy preserves output and cache state, but the paired late-fork measurements do not meet the no-slower gate at both required lengths.
…hared experts

The draft vocabulary retains lm_head bits and group size; grouped routed/shared experts refuse a mismatched stored format before concatenation.
… turn generated tokens into a decode rate

generation_tokens_total had no decode time to divide by: request_latency_seconds includes the prompt pass and queueing. request_decode_seconds (and request_decode_time_seconds under vLLM's name) observes each finished request's decode time: the CUDA engine's own decode_s, the figure /health already sums as decode_seconds_total, or first token to end on the Mac server. A request that generated nothing observes nothing.
The host format tests import CUDA modules requiring both torch and triton; retain both assertions on eligible runtimes and explicitly skip the module elsewhere.
….6.3

The grant is the pool's free memory less max(4 GiB, a tenth of it) on a discrete card too; TENSORFOLD_MEMORY_RESERVE_GIB moves the floor and TENSORFOLD_CUDA_MEMORY_LIMIT_GB caps the grant.
…y floor

Expect the 80 GiB fake card to grant 72 GiB; retain reusable allocator bytes above the floored grant; add the 128 GiB card floor to fake free memory so the unchanged 81920-token growth and refusal assertions still exercise the absolute cap.
CONTRIBUTING.md says what a pull request must keep (lanes, exact output, precision, prompt speed), the receipt
to paste, how to run the tests and how a pull request lands. New pull requests open with the receipt template.
@ashhart

ashhart commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Thank you, @CerebralCoding. We want the Zig engine in TensorFold, and we're putting real time into it now.

Our plan, so you can see where this goes:

  • We'll bring this sync up to v0.6.5 ourselves (it stops at 0.6.1 today) and review it commit by commit.
  • The Zig engine then runs behind a switch, with today's engine as the reference: the same tokens on every cell of
    our release checks, drafted output equal to "draft": false, a resumed prompt equal to a fresh one, each concurrent
    stream equal to its solo run, and at least the same speed. It becomes the default only after it passes all of them.
  • The lanes are the one condition from the branch's terms that is still open, as your README says. That's where
    we'll start.

If you have a model and chip where the Zig engine already runs drafted rounds, a receipt of drafted against serial
on it would help us most. Thanks again for building this.

@CerebralCoding

CerebralCoding commented Oct 4, 2026 •

Copy link
Copy Markdown
Author

Thanks, @ashhart. There’s a substantial follow-up on our zig/native branch that this sync PR doesn’t include—currently 25 additional commits. I should have made that separation clearer.

The PR’s “shared lane rounds remain unimplemented” wording describes the sync branch. I’ll correct the description to make that scope explicit and link the development work. That branch now includes shared lane forwards for Qwen, Bonsai, Gemma, Nemotron and Flash, alongside drafting, cache settlement, cancellation, prefix reuse and concurrent-versus-solo correctness checks.

For the drafted-versus-serial receipt you requested: on an Apple M5 Max, 128 GiB, using NVIDIA Nemotron 3.5 Lightning 30B-A3B, MLX 4-bit, our saved runs produced identical 32-token outputs with drafting enabled and disabled, also matching the pinned Python reference:

Case Proposed / accepted draft tokens Drafted = serial
Greedy 35 / 14 Exact
Sampled, temperature 0.7, seed 5678 26 / 21 Exact

The branch also contains Python golden capture and comparison tooling with bounded fixture storage. These are scoped results: full performance parity still has gaps, physical M1–M4 coverage remains unverified, and validation against v0.6.5 is still needed.

I’d keep this PR focused on the sync you’re reviewing, then bring the additional work over in small, dependency-ordered PRs, each with its relevant tests and reproducible receipts. Before you start implementing lanes, it would be useful to agree on the first reviewable slice from the existing branch so we avoid duplicating that work.

@ashhart
ashhart deleted the branch ashhart:zig October 4, 2026 19:57
@ashhart ashhart closed this Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.