Skip to content

Tiled int8 prefill GEMM (tiled-matmul plan, slice 1) - #1

Merged
dansupergameprogrammer merged 8 commits into
mainfrom
claude/project-thread-c8iecr
Sep 29, 2026
Merged

dansupergameprogrammer merged 8 commits into
mainfrom
claude/project-thread-c8iecr

Conversation

@dansupergameprogrammer

@dansupergameprogrammer dansupergameprogrammer commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Requested by Dan · project thread

Before: every CPU prefill GEMM computes one dot product per (token, output) cell, reloading and re-widening both operands each time, at the same ~15 GMAC/s whether it gets 1 token or 512.

After: calls with 8 or more tokens on the AVX2 and AVX-512BW tiers run a register-tiled vpmaddwd kernel over weights packed per call. On the cloud host, a single GEMM is 1.67–2.99× faster on AVX2 and 1.63–4.27× faster on AVX-512. The top of each range is the 512-token down shape, where the shipped kernel itself slows down. Whole prefill on 0.5B- and 0.6B-width synthetic artifacts is about 1.5–1.6× faster at 128 tokens (1.54–1.64× best of 7, from an independent rerun). An earlier reading of 1.96× did not reproduce. Every tier stays bit-identical to the scalar reference and to v1.9.0 (golden pins, digests, save blobs).

This is slice 1 of the approved tiled-matmul plan of record, built on the v1.9.0 base. It claims an engine-level speedup only: the Unreal plugin prefills one token per call and gets nothing until its chunking change lands, and SuperEmbedder gets it at its own re-pin.

How: a column-range entry over the shipped loop first, then the tiled kernel (4×16 AVX2, 8×32 AVX-512BW tiles, int64 flush every 16,384 pairs), activation prep, packer, per-tier tiled-entry counters and a build-config record. The tiled threshold is 8 tokens on both tiers, chosen for AVX-512; AVX2 would already win from 4 tokens. On MSVC and clang-cl the AVX-512 tiled path is compiled in but off by default until an MSVC AVX-512 build has executed it. New cells, mutants, a GEMM bench and decision rule (10.1 PASS: median 2.05×, now with a --verify pre-check), and the route E reach check (10.0) are included. Progress, evidence and box-only items: docs/tiled-matmul-slice1-progress.md.

Verified: blind code review SHIP on the owner's Windows box (MSVC 19.33 and clang-cl 15, auto and forced AVX2 suites green, digests identical, 2.59× on one 1.5B GEMM on Zen 2); CI green. The CI fixes also add the shader-build dependency the Windows forced suites were missing and make one pre-existing GPU rope-cache cell deterministic. Still owed: an execution of the AVX-512 tiled path under MSVC (no AVX-512 hardware on the box or the Windows runners so far).

🤖 Generated with Claude Code

https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68

The tiled-matmul plan's S1-0 step, taken on the v1.9.0 base instead of the 1.8.0 base the plan names
(the engine released 1.9.0 and the plugin vendors it; bit-identity references come from v1.9.0).

- tools/t2147_chunk_batched_pins.cpp: token-id mode ("ids:N" or "ids:N:V") for artifacts with no
  tokenizer, --layers=L for reduced-layer artifacts, --repeat=R (best of R, interleaved) for the
  speedup arms. Also fixes --boundary-sweep= parsing, which skipped the value's first character.
- tools/consumer_reach/: the plan's consumer-reach harness upstreamed from its probes (route E: the
  10.0 tool, the checks, the leg runner, the record reader, the codegen witness and Q1's recorders;
  route P: the hook witness, the simulators, the identity check and the grader), plus the synthetic
  real-width artifact builder with the Qwen2.5-0.5B geometry added.
- docs/tiled-matmul-slice1/s1-0-base-v1.9.0/: v1.9.0's five axis digests and three GCC suites on this
  host (all green, one GLOBAL digest), and the guarded route E build against v1.9.0.
- docs/tiled-matmul-slice1-progress.md: the slice's progress record.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
… shipped loop

Plan §4.3 (review W8): GemmInt8Accumulate now runs through detail::GemmInt8AccumulateCols over the
whole output range, and the entry has a sibling that takes pre-widened int16 activations (slice 2's
input). Both still run the one-cell-at-a-time DotRow loop; the tiled kernel lands next. The header
also declares the path selector, the dispatch wiring and the tier query the kernel commit defines.

Cell 3.3 (the Cols canary, widened per §10.M XCa/XCb/XCc) lives in the new tests/test_tiled_gemm.cpp,
compiled into superslm_tests, the forced suites and build.bat's test binary. Red first: the mutant
that ignores the range (XC) fails 102 of its 108 checks; the entry as built passes all 108.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…y tier

On the AVX2 and AVX-512BW tiers a GEMM call of 8 or more tokens now runs a
register-tiled vpmaddwd kernel over weights packed per call (4x16 and 8x32
tiles, int64 flush every 16,384 pairs), bit-identical to the scalar reference.
The MSVC AVX-512 tiled path stays off behind SUPERSLM_TILED_AVX512_MSVC.

- the kernel, packer, activation prep, pure selector and dispatch wiring
- per-tier tiled-entry counters and a tiled bad_alloc injection slot
- the build-configuration record constant (SSLM-BUILDCFG/2)
- cells 1.1, 3.5, 4.1-4.4, 4.7(a), 5.1, 6.1 (new golden pin), 6.2 (new digest
  section; sections 1-8 unchanged from v1.9.0), 6.3, 11.1
- 11.3 linkage check, 11.4 named-span coverage check, isolation-prose scrub
- Windows forced AVX2/AVX-512 CI legs (not executable on this host)
- mutation evidence and progress in docs/tiled-matmul-slice1*

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…le, 9.1 blobs

- tools/sslm_gemm_bench.cpp (cell 10.1) with production, forced and
  D-infinity builds, and tools/sslm_gemm_bench_rule.py for the plan's rule
- the 9.1 two-binary blob protocol: the bench tool reports tiled entries
  when built against the seam library
- tools/consumer_reach/route_e/run_route_e_slice1.sh: route E's deciding
  legs on the real candidate
- the digest prints the dispatched GEMM tier (not digested); clang-cl forced
  AVX2/AVX-512 Windows legs
- measured 10.1, per-GEMM and whole-prefill evidence under
  docs/tiled-matmul-slice1/s1-c/

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
- cell 10.0 route E on the real candidate: GREEN (production reaches the
  tiled kernel, 7 entries at 28 and 276 tokens; D-infinity, forced-scalar and
  no-seam legs fail on their exact checks)
- whole batched prefill before/after on 0.5B-width (1 and 2 layers) and
  0.6B-width (1 layer) synthetic artifacts, save blobs equal throughout
- platform-support rows and the changelog entry: an engine-level claim only
- progress file: every S1-0/S1-B/S1-C item, its state and its numbers

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
A rerun of the whole-prefill bench puts one layer of batched prefill at
128 tokens at about 1.5x-1.6x faster than 1.9.0 (1.54x-1.64x best of 7,
1.48x-1.67x by median), on AVX2 and AVX-512 alike. The 1.96x forced-AVX2
reading rested on a v1.9.0 base that read slow for its whole session
(161 ms against 129 ms on the rerun), and does not reproduce.

- CHANGELOG, platform-support.md: whole prefill now reads about 1.5x-1.6x
  at 128 tokens (was 1.7x-2.0x / 1.72x and 1.96x). Per-GEMM ranges stand,
  with a note that their tops come from the shipped kernel slowing at
  512 tokens.
- docs/tiled-matmul-slice1/s1-c/prefill-remeasure-2026-09-29.md: the
  rerun's table. prefill-p05_l1.md is kept as recorded, with a note that
  its 1.96 row is superseded; the progress doc says the same.
- src/matmul.cpp: the kTiledMinTokens comment said pack-per-call pays
  for itself from M = 8. That holds for AVX-512 only; AVX2 wins from
  M = 4. The comment now says so. The value and all behaviour are
  unchanged.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Each build runs once, untimed, with --verify against the scalar reference
before the timed pairs; a verify failure is FAIL. Closes the hole where a
fast, deterministic but wrong kernel would pass cell 10.1 (plan rev 11.4, D5).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer
dansupergameprogrammer marked this pull request as ready for review September 29, 2026 05:43
- fp-free scan (linux-x64): GCC 13 if-converted the AVX-512 driver's store
  guard into mask-register code (vpcmpuq/kmovw) that the scan's allow-list
  rejects. Exit the panel loop at the first column past j_end instead; n
  only grows, so behaviour is unchanged and the guard line is untouched.
- branch-coverage: the Cols canary cell gains the empty range
  (j_end == j_begin), covering both entries' early return; matmul.cpp goes
  back above its 72.22% floor without re-pinning it.
- Windows forced legs (0xC0000409, no output): the forced suites run the
  GPU cells but had no dependency on the shader target, so a forced-only
  build left <config>/shaders empty. Add the same dependency the main
  suite has.
- windows-x64 flake: the G5-bridge rope-cache cell rebuilt its fixture per
  chunk, so the second chunk hit only if the heap reused the freed block.
  Use one fixture for both chunks, as the cell's own comment says.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer
dansupergameprogrammer merged commit bde3074 into main Sep 29, 2026
78 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants