Tiled int8 prefill GEMM (tiled-matmul plan, slice 1) - #1
Merged
Merged
Conversation
The tiled-matmul plan's S1-0 step, taken on the v1.9.0 base instead of the 1.8.0 base the plan names
(the engine released 1.9.0 and the plugin vendors it; bit-identity references come from v1.9.0).
- tools/t2147_chunk_batched_pins.cpp: token-id mode ("ids:N" or "ids:N:V") for artifacts with no
tokenizer, --layers=L for reduced-layer artifacts, --repeat=R (best of R, interleaved) for the
speedup arms. Also fixes --boundary-sweep= parsing, which skipped the value's first character.
- tools/consumer_reach/: the plan's consumer-reach harness upstreamed from its probes (route E: the
10.0 tool, the checks, the leg runner, the record reader, the codegen witness and Q1's recorders;
route P: the hook witness, the simulators, the identity check and the grader), plus the synthetic
real-width artifact builder with the Qwen2.5-0.5B geometry added.
- docs/tiled-matmul-slice1/s1-0-base-v1.9.0/: v1.9.0's five axis digests and three GCC suites on this
host (all green, one GLOBAL digest), and the guarded route E build against v1.9.0.
- docs/tiled-matmul-slice1-progress.md: the slice's progress record.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
… shipped loop Plan §4.3 (review W8): GemmInt8Accumulate now runs through detail::GemmInt8AccumulateCols over the whole output range, and the entry has a sibling that takes pre-widened int16 activations (slice 2's input). Both still run the one-cell-at-a-time DotRow loop; the tiled kernel lands next. The header also declares the path selector, the dispatch wiring and the tier query the kernel commit defines. Cell 3.3 (the Cols canary, widened per §10.M XCa/XCb/XCc) lives in the new tests/test_tiled_gemm.cpp, compiled into superslm_tests, the forced suites and build.bat's test binary. Red first: the mutant that ignores the range (XC) fails 102 of its 108 checks; the entry as built passes all 108. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…y tier On the AVX2 and AVX-512BW tiers a GEMM call of 8 or more tokens now runs a register-tiled vpmaddwd kernel over weights packed per call (4x16 and 8x32 tiles, int64 flush every 16,384 pairs), bit-identical to the scalar reference. The MSVC AVX-512 tiled path stays off behind SUPERSLM_TILED_AVX512_MSVC. - the kernel, packer, activation prep, pure selector and dispatch wiring - per-tier tiled-entry counters and a tiled bad_alloc injection slot - the build-configuration record constant (SSLM-BUILDCFG/2) - cells 1.1, 3.5, 4.1-4.4, 4.7(a), 5.1, 6.1 (new golden pin), 6.2 (new digest section; sections 1-8 unchanged from v1.9.0), 6.3, 11.1 - 11.3 linkage check, 11.4 named-span coverage check, isolation-prose scrub - Windows forced AVX2/AVX-512 CI legs (not executable on this host) - mutation evidence and progress in docs/tiled-matmul-slice1* Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…le, 9.1 blobs - tools/sslm_gemm_bench.cpp (cell 10.1) with production, forced and D-infinity builds, and tools/sslm_gemm_bench_rule.py for the plan's rule - the 9.1 two-binary blob protocol: the bench tool reports tiled entries when built against the seam library - tools/consumer_reach/route_e/run_route_e_slice1.sh: route E's deciding legs on the real candidate - the digest prints the dispatched GEMM tier (not digested); clang-cl forced AVX2/AVX-512 Windows legs - measured 10.1, per-GEMM and whole-prefill evidence under docs/tiled-matmul-slice1/s1-c/ Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
- cell 10.0 route E on the real candidate: GREEN (production reaches the tiled kernel, 7 entries at 28 and 276 tokens; D-infinity, forced-scalar and no-seam legs fail on their exact checks) - whole batched prefill before/after on 0.5B-width (1 and 2 layers) and 0.6B-width (1 layer) synthetic artifacts, save blobs equal throughout - platform-support rows and the changelog entry: an engine-level claim only - progress file: every S1-0/S1-B/S1-C item, its state and its numbers Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
A rerun of the whole-prefill bench puts one layer of batched prefill at 128 tokens at about 1.5x-1.6x faster than 1.9.0 (1.54x-1.64x best of 7, 1.48x-1.67x by median), on AVX2 and AVX-512 alike. The 1.96x forced-AVX2 reading rested on a v1.9.0 base that read slow for its whole session (161 ms against 129 ms on the rerun), and does not reproduce. - CHANGELOG, platform-support.md: whole prefill now reads about 1.5x-1.6x at 128 tokens (was 1.7x-2.0x / 1.72x and 1.96x). Per-GEMM ranges stand, with a note that their tops come from the shipped kernel slowing at 512 tokens. - docs/tiled-matmul-slice1/s1-c/prefill-remeasure-2026-09-29.md: the rerun's table. prefill-p05_l1.md is kept as recorded, with a note that its 1.96 row is superseded; the progress doc says the same. - src/matmul.cpp: the kTiledMinTokens comment said pack-per-call pays for itself from M = 8. That holds for AVX-512 only; AVX2 wins from M = 4. The comment now says so. The value and all behaviour are unchanged. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Each build runs once, untimed, with --verify against the scalar reference before the timed pairs; a verify failure is FAIL. Closes the hole where a fast, deterministic but wrong kernel would pass cell 10.1 (plan rev 11.4, D5). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
dansupergameprogrammer
marked this pull request as ready for review
September 29, 2026 05:43
- fp-free scan (linux-x64): GCC 13 if-converted the AVX-512 driver's store guard into mask-register code (vpcmpuq/kmovw) that the scan's allow-list rejects. Exit the panel loop at the first column past j_end instead; n only grows, so behaviour is unchanged and the guard line is untouched. - branch-coverage: the Cols canary cell gains the empty range (j_end == j_begin), covering both entries' early return; matmul.cpp goes back above its 72.22% floor without re-pinning it. - Windows forced legs (0xC0000409, no output): the forced suites run the GPU cells but had no dependency on the shader target, so a forced-only build left <config>/shaders empty. Add the same dependency the main suite has. - windows-x64 flake: the G5-bridge rope-cache cell rebuilt its fixture per chunk, so the second chunk hit only if the heap reused the freed block. Use one fixture for both chunks, as the cell's own comment says. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Dan · project thread
Before: every CPU prefill GEMM computes one dot product per (token, output) cell, reloading and re-widening both operands each time, at the same ~15 GMAC/s whether it gets 1 token or 512.
After: calls with 8 or more tokens on the AVX2 and AVX-512BW tiers run a register-tiled
vpmaddwdkernel over weights packed per call. On the cloud host, a single GEMM is 1.67–2.99× faster on AVX2 and 1.63–4.27× faster on AVX-512. The top of each range is the 512-tokendownshape, where the shipped kernel itself slows down. Whole prefill on 0.5B- and 0.6B-width synthetic artifacts is about 1.5–1.6× faster at 128 tokens (1.54–1.64× best of 7, from an independent rerun). An earlier reading of 1.96× did not reproduce. Every tier stays bit-identical to the scalar reference and to v1.9.0 (golden pins, digests, save blobs).This is slice 1 of the approved tiled-matmul plan of record, built on the v1.9.0 base. It claims an engine-level speedup only: the Unreal plugin prefills one token per call and gets nothing until its chunking change lands, and SuperEmbedder gets it at its own re-pin.
How: a column-range entry over the shipped loop first, then the tiled kernel (4×16 AVX2, 8×32 AVX-512BW tiles, int64 flush every 16,384 pairs), activation prep, packer, per-tier tiled-entry counters and a build-config record. The tiled threshold is 8 tokens on both tiers, chosen for AVX-512; AVX2 would already win from 4 tokens. On MSVC and clang-cl the AVX-512 tiled path is compiled in but off by default until an MSVC AVX-512 build has executed it. New cells, mutants, a GEMM bench and decision rule (10.1 PASS: median 2.05×, now with a
--verifypre-check), and the route E reach check (10.0) are included. Progress, evidence and box-only items:docs/tiled-matmul-slice1-progress.md.Verified: blind code review SHIP on the owner's Windows box (MSVC 19.33 and clang-cl 15, auto and forced AVX2 suites green, digests identical, 2.59× on one 1.5B GEMM on Zen 2); CI green. The CI fixes also add the shader-build dependency the Windows forced suites were missing and make one pre-existing GPU rope-cache cell deterministic. Still owed: an execution of the AVX-512 tiled path under MSVC (no AVX-512 hardware on the box or the Windows runners so far).
🤖 Generated with Claude Code
https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68