Prepare reviewed main-to-zig sync - #130
CerebralCoding wants to merge 195 commits into
Conversation
… in _fp4mm The reduce: a split-K fp32 face summed its slices with `out.copy_(part[0]); for s: out += part[s]` - one elementwise launch a slice, so a face the shape splits 32 ways paid 31 of them. Two faces a layer take that path (`forward._mm` with `f32=True`), and it is their whole reduce: the bf16 face next door already ran in one launch. `_reduce` gains the fp32 face (F32), the callers use it, and the loop is gone. The adds are the loop's own, in the loop's order, one fp32 add a slice, so the fp32 sums are bit-identical to the loop's - the order the split-K contract is written around. `test_matmul_splitk_sum_order_is_the_reduces_one` now pins both faces: the bf16 one rounds each partial sum as the loop does, the fp32 one does not round at all. The blocks: two things a step did that it does not need to do 160 times a program. The address `(b // 4) * 32 * SBN + ((b % 4) * 16) // 2 * SBN` is a staircase in the step (a stored macro block holds 8 rows of the tile), and a load whose address the pipeliner cannot see as affine is not put in flight; it is the same address as `b * 8 * SBN` (16 * SBN for the pattern tables, whose macro block holds 16 rows). The decode: the missing nibble is the byte's own parity, so `(w >> ((r16 % 2) * 4)) & 0xF` reads it where a compare and a select took two ops, and `_e2m1_pattern` already builds the 16 bits a bf16 holds, so bitcasting those straight into bf16 keeps exactly what widening to fp32 then rounding back gave, three ops less an element. Measured on the Spark by timing the grouped step against variants of its own source with each decode replaced by a cast (dev/deqcost.py): the E2M1 and E4M3 decodes together are 21% of a grouped step at the served geometry (7 rows x 11 slots, 72 distinct experts). The shave is worth taking; the weight traffic is what the step actually waits on.
… does `tensorfold serve --alias` only reached the MLX server. On CUDA the flag parsed and was dropped, so /v1/models listed --name alone and every reply named --name. The CUDA app now takes the aliases: /v1/models lists --name, then each alias once, in the MLX server's order, and a reply names the id the request asked for when this endpoint answers to it, else --name. Ported onto 0.5.0's HTTP module (cuda/http.py) from ashhart#111.
Kept entries ended at the prompt's end, and a resume needs a strictly shorter prefix, so an identical resend or a next turn with thinking on (which renders the last prompt token differently) always prefilled in full. Keep the entry at len(prompt) - 1, as the 27B engine does, cut inside the chunk that holds it.
…t checkpoints Lane matmul mode FP8G: e4m3 bytes in the FP8 GEMM's fragment order and one fp32 scale per (64 inputs, column), applied after each 64-input stage's bf16 MMAs, so the stored weight is exact and rows stay independent of the row count. Fp8BlockLinear decodes on it and runs prompts on the FP8 prompt GEMM with the scales as bf16 group scales. The Flash Next loader maps weight_scale_inv linears to it, runs stacks that mix block FP8 with bf16 through Concat, and dequantizes a block-FP8 lm_head (and the draft head's rows) with its block scales. format.scheme and format.dequant learn the layout; the family's check accepts FP8_PB_WO. Tests: the linear against an fp64 reference, rows, prompt chunks and Concat; the numpy dequantizer; a tiny ModelOpt checkpoint with block-FP8 linears through the loader, the engine and drafted-equals-serial decoding.
The kernel holds 16 pairs an item in the decode form and 64 in the prefill form, and nothing between. Measured at a prompt's own row count on the published checkpoint's shapes (2275 rows x 10 slots over 512 experts): 16 pairs reads 1788 items in 14.09 ms, 64 reads 549 in 21.00 ms. The wide item reads three times fewer bytes and is 49% slower, because those re-reads are L2 hits anyway (351 GB/s is above the box's DRAM) and all it bought was three times fewer items competing for the SMs. Served on ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 with the 8-bit projection copies: prose 32.2 -> 40.0, code 63.6 -> 75.8, prefill 1196 -> 1486. Decode moves with it because the arithmetic a prompt runs is the one a verify window runs. Tests: tests/cuda 734 passed, 75 skipped, 0 failed; the new case pins the item and checks a prompt's plan stays inside max_items for it - the ceiling and the count are easy to confuse, and only the count is traffic.
… a kernel a round) and wide copy windows for one stream - CUDA --parallel: a queued prompt fills inside the decode rounds; a round runs every stream's DeltaNet and attention in one launch; caches grow with use; --decode-share sizes a prompt pass - One stream's copy windows ramp to 128 rows on M1-M4 and on a GB10 (the 27B too); a broken copy halves the window; TF_COPY_ROWS sets the widest - Mac replies and /v1/responses name the model id the request asked for
…e backbone directly
Both ranks prefill the question and reduce the vocabulary shards into label probabilities. Host tests cover the prompt wording, the refusals, and that reduction.
…oader keeps, not fp32
…projections for 5/6/8-bit Flash Next before M5, admission fixes - /metrics: running and waiting requests, token totals, KV use, MTP drafts, request latency and time to first token (ashhart#110) - Flash Next before M5: 5-, 6- and 8-bit projections of 4+ rows on the matrix units, with the per-row arithmetic - Mac admission: one need for the window's fit and every admission, and a stream's probed growth is its caches' own bytes - GLM-5.3 hears medium as high (ashhart#117); Gemma's bare tool calls parse; a null sampling field keeps the server default - Two CUDA ranks meet on the store after loading and name a rank that never finishes (ashhart#107); DFlash2 admission and loading share one packing rule
propose() took the draft-vocabulary path first, which reads DFlash2's candidate selector; z-lab/Qwen3.6-35B-A3B-DFlash has a draft vocabulary and no selector, so its first round raised.
…raft vocabulary A v1 head gives no per-draft chances, yet Qwen35Family always exposed draft_probabilities, so rounds drafted depth 15 and were trimmed with forward-only costs that ignore the drafter's time. block_chain also read the full 248k-row head instead of the draft vocabulary rows.
…r efforts, help covers request efforts only
…ution Route Anthropic and Mac tokenizer requests through the shared HTTP body reader, retain image parts in their tool results, and keep the video tokenizer test compatible with Python 3.11.
Preserve the one-stream startup allowance when its real growth fits, wait for filling prompts to free room, and refuse isolated over-budget growth without losing request slots.
Revert merge 6d4a11c3666427c666388dd2aa84e986fc8c7e63: the corrected policy preserves output and cache state, but the paired late-fork measurements do not meet the no-slower gate at both required lengths.
… reductions and row windows
…hared experts The draft vocabulary retains lm_head bits and group size; grouped routed/shared experts refuse a mismatched stored format before concatenation.
… turn generated tokens into a decode rate generation_tokens_total had no decode time to divide by: request_latency_seconds includes the prompt pass and queueing. request_decode_seconds (and request_decode_time_seconds under vLLM's name) observes each finished request's decode time: the CUDA engine's own decode_s, the figure /health already sums as decode_seconds_total, or first token to end on the Mac server. A request that generated nothing observes nothing.
The host format tests import CUDA modules requiring both torch and triton; retain both assertions on eligible runtimes and explicitly skip the module elsewhere.
….6.3 The grant is the pool's free memory less max(4 GiB, a tenth of it) on a discrete card too; TENSORFOLD_MEMORY_RESERVE_GIB moves the floor and TENSORFOLD_CUDA_MEMORY_LIMIT_GB caps the grant.
…y floor Expect the 80 GiB fake card to grant 72 GiB; retain reusable allocator bytes above the floored grant; add the 128 GiB card floor to fake free memory so the unchanged 81920-token growth and refusal assertions still exercise the absolute cap.
CONTRIBUTING.md says what a pull request must keep (lanes, exact output, precision, prompt speed), the receipt to paste, how to run the tests and how a pull request lands. New pull requests open with the receipt template.
bfce01c to
ae5c87a
Compare
|
Thank you, @CerebralCoding. We want the Zig engine in TensorFold, and we're putting real time into it now. Our plan, so you can see where this goes:
If you have a model and chip where the Zig engine already runs drafted rounds, a receipt of drafted against serial |
|
Thanks, @ashhart. There’s a substantial follow-up on our The PR’s “shared lane rounds remain unimplemented” wording describes the sync branch. I’ll correct the description to make that scope explicit and link the development work. That branch now includes shared lane forwards for Qwen, Bonsai, Gemma, Nemotron and Flash, alongside drafting, cache settlement, cancellation, prefix reuse and concurrent-versus-solo correctness checks. For the drafted-versus-serial receipt you requested: on an Apple M5 Max, 128 GiB, using NVIDIA Nemotron 3.5 Lightning 30B-A3B, MLX 4-bit, our saved runs produced identical 32-token outputs with drafting enabled and disabled, also matching the pinned Python reference:
The branch also contains Python golden capture and comparison tooling with bounded fixture storage. These are scoped results: full performance parity still has gaps, physical M1–M4 coverage remains unverified, and validation against v0.6.5 is still needed. I’d keep this PR focused on the sync you’re reviewing, then bring the additional work over in small, dependency-ordered PRs, each with its relevant tests and reproducible receipts. Before you start implementing lanes, it would be useful to agree on the first reviewable slice from the existing branch so we avoid duplicating that work. |
Replace fork setup instructions with TensorFold’s long-running
zigworkflow. Manual sync discovers the TensorFold SSH remote, prepares an uncommitted main merge, aligns dependencies and kernels, and supports resuming verification at the original target. It refuses protected branches, outdated bases and unfinished operations; it never rebases or pushes.The native HTTP server currently interleaves independent request forwards with separate caches. Shared lane rounds remain unimplemented; the README now states this explicitly.
Validation passed: seven offline Git tests, native baseline and checkpoint-file checks, setup/dependency/coverage checks, and safe-mode CLI compilation.