Exact CPU attention and per-row sites (plan slices S1–S5) - #3
Merged
Merged
Conversation
…row tables Plan of record: attention and per-row sites, rev 3.1, slice S1 (§4.1). Red-first: the cells, the seam's row-table counters, the digest section and the golden pin, before any site changes. - tests/support/matmul_dispatch_instrument.h: the six row-table path counters (§3.6), declared outside the x64 block so arm64 counts them too. Nothing increments them yet. - tests/test_attn_rowsites.cpp: cells 4.S1 (width grid around 512, -128 at first/middle/last at the small gate scale), 1.S1, 5.S1, 7.S1a-d, 11.1(a) for the row tables, 6.3 and 11.1(d)'s row-table and trace-record rows (driven through the two layer loops when SUPERSLM_ATTN_ROWSITES_ARTIFACT names the 0.5B-width 1-layer artifact). Each call is compared with a test-side copy of the v1.9.0 per-element loop and asserts its own counter delta. - tests/support/rowsite_cases.h: the fixed input set shared by the digest section c_rowsites, the golden generator and the suite. - tools/gen_attn_rowsite_golden.cpp and tests/attn_rowsite_golden_pin.h: the pin, generated against the v1.9.0 tag's library (8836d5eb..., 634,120 values). - The sslm_bench_prefill forced targets gain tests/ on their include path; they did not compile on the base. Red on GCC auto and forced SSE2/AVX2/AVX-512: 258 of 1,756 checks fail, every one a counter assertion; every value assertion passes on the base. Digest sections 1-9 are byte-identical to slice 1's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…on every tier Plan of record: attention and per-row sites, rev 3.1, slice S1 (§4.1, §5.1). RmsNormSite's divide, MlpActSite's SiLU sigmoid and ResidualReconcileSite's landing rescale each evaluate a pure function of one int8 code and row constants. At n >= 512 (kRowTableMinWidth) each site now evaluates it once per code into a stack table with exactly the arguments its per-element loop passes, and the loop reads the table: - norm: FloorDivI64(c << 32, root) for every int8 c (256 x int64); - SiLU: SiluSigmoidQ15(table, c, m, e) for c in [-127, 127]; a -128 gate code is evaluated directly; - landing: (LandingRescale value, int64-overflow flag) for c in [-127, 127] per candidate, with null counters; the loop keeps its order, first-flag return and overflow checks, reads the flag per element present, and lands -128 directly. The tables are off only under SUPERSLM_FORCE_SCALAR_MATMUL, which keeps the v1.9.0 loops as the reference axis. Each site counts its table decision once per call on the instrument seam's row-table counters. GCC 13.3 Release, auto (AVX-512 here) and forced SSE2/AVX2/AVX-512: 1,756 S1 checks, 0 failures; full suites 27,287 / 27,229 / 27,245 / 27,245 checks, 0 failures. S1 golden 8836d5eb... matched on every binary; SiLU-LUT, matmul and tiled golden pins unchanged. Digests: all five legs equal; sections 1-9 equal to slice 1's; c_rowsites equals the v1.9.0 pin. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…he bench Plan of record: attention and per-row sites, rev 3.1, slice S1. - tests/test_attn_rowsites.cpp: cell 3.S1, 8 threads calling all three sites at table widths against the single-threaded reference. The plan's 3.S1 names an artifact-driven concurrent-read cell that does not exist. - tools/sslm_sites_bench.cpp (cell 10.1): kernel, prefill and decode modes; one source built against either library. - tools/t2147_chunk_batched_pins.cpp: --chunk-budget=B and --decode=D in --dump-blob mode, for the plan's blob protocol. - docs/attention-rowsites/s1/: suites (GCC and Clang, four binaries each, plus GCC TSan and ASan+UBSan), blob protocol (33 comparisons, all equal, with 32 decode steps each), 15 mutants with scripts (all killed on the auto and forced-AVX2 binaries), bench.md. - docs/attention-rowsites-s1-progress.md: state, deviations, resume steps. Measured saving at full 0.5B depth: 0.80 ms per token by the plan's per-site method (estimate 0.87); 0.96-0.99 from reduced-layer prefill. The decode saving is below this shared host's noise. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…ltiply-add Plan of record: attention and per-row sites, rev 3.1, slice S2 (§4.2, §5.2). Red-first: the cells, the seam's per-tier prob-V counters, the 11.2 selector, the digest section and the golden pin, before any kernel lands. - include/superslm/matmul.h: detail::SitesKernel, SelectSitesKernel (the pure selector over tier, SUPERSLM_SITES_AVX512_MSVC and MSVC identity) and DispatchSitesKernel (its wiring). src/matmul.cpp carries a stub that selects the v1.9.0 code on every tier, which is what the base does. - tests/support/matmul_dispatch_instrument.h: pv_fast / pv_fallback for AVX2 and AVX-512 (§3.6), inside the x64 block. Nothing increments them yet. - tests/test_attn_rowsites.cpp: cells 11.2 (full truth table and wiring), 4.S2 (head_dim x width grid, odd widths, exact-size heap buffers), 7.S2 corners, 2.S2 (four hostile rows, each failing exactly one conjunct of a test-side guard copy), 6.3 (S2 golden) and 11.1(d)'s prob-V rows (L*H*N calls, data terms 15 prefill / 0 decode, re-derived on this base with the plan's probe). Every call asserts output against a test-side copy of the v1.9.0 loop and its own counter delta on the selected kernel's tier only. - tests/support/attention_cases.h: the fixed prob-V input set shared by the new digest section c32_attention, the golden generator and the suite. - tools/gen_attn_rowsite_golden.cpp: one hash per slice; tests/ attn_rowsite_golden_pin.h regenerated against the v1.9.0 tag (S1 unchanged, S2 b0d1a6cd... over 30,100 values). Red on GCC auto and forced AVX2/AVX-512: 269 of 2,306 attn-rowsites checks fail each, forced SSE2 3 (the selector); every value assertion passes on the base. Digest sections 1-10 are byte-identical to S1's on all five legs. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…on every tier
Plan of record: attention and per-row sites, rev 3.1, slice S2 (§4.2, §5.2).
GemmProbQ15Accumulate zeroes its output and runs a tiered accumulate-into
core. On the AVX2 and AVX-512BW tiers, when head_dim is a multiple of 16 and
the row passes the int16 condition (every p in [0, 32767], Sum p <= 2^15,
checked in one pass by the kernel itself), it pairs keys: the two value rows
are interleaved byte by byte, widened to int16 and multiplied by the
broadcast (p_k, p_{k+1}) pair with vpmaddwd into one int32 lane per output
dimension (an odd last key pairs with a zero row). AVX2 works in 16-dimension
units; AVX-512BW in 32-dimension units interleaved in-lane (the stores put
each half back at its dimensions), plus one 16-dimension tail unit. Each lane is bounded by
128 * Sum p <= 2^22, so the widened sum equals the int64 loop exactly. Any
other row, and the scalar and SSE2 tiers, run the v1.9.0 loop. No allocation:
the pairs are formed in registers.
- detail::SelectSitesKernel / DispatchSitesKernel (cell 11.2): the new kernels
run on AVX2 and AVX-512; in MSVC and clang-cl builds the AVX-512 tier keeps
the v1.9.0 code unless SUPERSLM_SITES_AVX512_MSVC=1 (default 0), a switch
separate from the tiled GEMM's. Both forced AVX-512 Windows legs build with
it on.
- Path counters (§3.6): each tier's fast counter moves inside that tier's own
body, the fallback counter in the dispatcher after the guard decides.
- check_tiled_matmul_linkage.py and the isolation checker's prose name the new
attributed functions; the linkage check is OK on the auto and both forced
objects.
GCC 13.3 Release, auto (AVX-512 here) and forced SSE2/AVX2/AVX-512: 2,306
attn-rowsites checks, 0 failures; full suites 27,837 / 27,779 / 27,795 /
27,795, 0 failures. S1 and S2 goldens matched on every binary; SiLU-LUT,
matmul and tiled goldens unchanged. Digests: all five legs equal and equal to
the red run's.
[skip ci]
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The 4.S2 kernel-blocking rows (every AVX2 and AVX-512 block count and the AVX-512 16-dimension tail), the prob-V mode of sslm_sites_bench, and the slice's evidence: GCC and Clang suites and digests on every tier, 20 mutants killed, sanitizers, 45 equal save blobs, the bench against plan §0, a local branch-coverage replica with allowlist lines, and the CHANGELOG and platform-support entries. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S3 (§4.3, §5.3).
Red-first: the cells, the seam's per-tier requant_row counters, the row leaf's
declaration, the digest rows and the golden pin, before any kernel lands.
- include/superslm/intmath.h: RequantRowWide(x, n, r, s, out), the funnel's
element loop as one leaf, with the funnel's contract. src/intmath.cpp
carries a stub that runs the element loop, which is what the base does;
the funnel still runs its own loop.
- tests/support/matmul_dispatch_instrument.h: requant_row for AVX2 and
AVX-512 (§3.6), inside the x64 block. Nothing increments them yet.
- tests/test_attn_rowsites.cpp: 4.S3 (n {1, 3, 4, 5, 7, 8, 9, 896, 4864} x
d' covering every s in [-1, 30], +-d' in every lane position, random rows
with values next to rounding ties; a sentinel-fenced pass on every binary
and an exact-size heap pass for the hosted ASan leg), 7.S3's P = 2^63
corner premises, the funnel's call site (one leaf call per funnel call that
passes its preflight), 6.3 (S3 golden) and 11.1(d)'s requant rows
((11L + 1) N: 1,536 prefill, 384 decode). Codes are checked against the
unchanged RequantTokenCodeWide; the path against the selected kernel's
tier only.
- tests/support/rowsite_cases.h: RunRequantRowCases, through
RequantChainChecked (a v1.9.0 signature), appended to c_rowsites and pinned
by the generator against the v1.9.0 tag (S3 3e3abed7..., 3,567,018 values;
S1 and S2 unchanged).
Red on GCC auto and forced AVX2/AVX-512: 10 path assertions fail each,
forced SSE2 none; every value assertion passes on the base. Digest sections
other than c_rowsites are byte-identical to S2's on all five legs.
[skip ci]
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…, green on every tier Plan of record: attention and per-row sites, rev 3.1, slice S3 (§4.3, §5.3). The checked chain funnel's element loop is one call to the new row leaf RequantRowWide(x, n, r, s, out). On the AVX2 and AVX-512BW tiers it runs the §5.3 identity in 4 or 8 unsigned 64-bit lanes: P = |x|*r from the 32-bit halves of r (r can be 2^32), H by a logical shift (P reaches exactly 2^63 at the contract's corner d' = |x| = 2^31, r = 2^32), magnitude = (127H + ((127L + 2^(e-1)) >> 32)) >> (e - 32), then the element code's clamp at 127 and sign restore. The last n mod 4 (or 8) elements, and every element on the scalar and SSE2 tiers, run RequantTokenCodeWide. The AVX-512 body uses F and BW instructions only and no mask register (vpabsq, vpsraq, vpminuq, vpmovqb). No runtime guard (the funnel's preflight is the contract), so one per-tier counter, inside each body. - src/intmath.cpp: the two bodies (this file's first target-attributed functions, per function, anonymous namespace) and the dispatcher, which decides through S2's DispatchSitesKernel, so SUPERSLM_SITES_AVX512_MSVC governs it on MSVC. intmath.cpp now includes superslm/matmul.h and needs matmul.cpp at link time. - build_cert.bat, both lines of tools/build_inspect.bat and the loop's cl line of tests/t2296-fp-free-open-red-suite/build_link_red.bat gain src\matmul.cpp (plan G21, §3.4, cell 11.5). GCC link check committed. - tests/ci/check_no_forward_leaf_calls.py: RequantRowWide is a banned leaf (cell 11.4); a planted call from forward_sites.cpp turns it red. - tools/ci/check_tiled_matmul_linkage.py: the requant bodies join the population and the intmath.cpp objects join the CI job; the record rule applies to matmul objects only (cell 11.3; the plant turns it red). - tools/ci/check_matmul_avx_isolation.py: the prose names the two bodies. The AVX2 body does |x|, the clamp and the sign in 32-bit lanes (vpabsd, vpminud, vpsignd) and keeps its shuffle constant's halves distinct: Clang lowers 64-bit selects to vblendvpd/vxorpd and loads a repeated-half constant with vbroadcasti128, all three outside the fp-free scan's allow-list. The scan passes on the Clang build unwidened; on the GCC build it rejects only the base's TiledGemmAvx512 (fixed on main by 90e48de) and passes with that fix applied. GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512 suites, 0 failures; digests equal the red run's on all five legs. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The requant mode of sslm_sites_bench and the slice's evidence: GCC and Clang suites and digests on every tier, the fp-free scan (Clang clean; GCC clean apart from the base's TiledGemmAvx512, which main fixes), every §9 mutant killed per tier, sanitizers, 46 equal save blobs, the bench against plan §0 (2.5 ms/token saved on AVX2 against the 1.22 estimated), a local branch-coverage replica with allowlist lines, and the CHANGELOG and platform-support entries. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S4 (§4.4, §5.4).
Red-first: the cells, the seam's per-tier softmax counters, the test-side
guard copy and estimate replica, the digest rows and the golden pin, before
any kernel lands. SoftmaxRowQ15 is unchanged.
- tests/support/attention_cases.h: the S4 set (§8 4.S4's grid at widths
{1, 2, 3, 4, 5, 2^14} plus the kernels' block edges, realistic constants
from IExpScaleConstants over the forward's scale range, four row kinds,
aliased rows; the inside corners; total = 1; 2.S4's hostile rows, each
failing exactly one conjunct, and the off-ratio witness), the test-side
copy of §5.4's guard (TestSoftmaxGuard, one flag per conjunct), and a
replica of the fast path's estimates (SoftmaxEstimateReplica) that picks
the correction rows: 24 each for z up, p up and p down, each kept only
where the correction fires and leaving it out changes the output. The p
down rows come from a steered generator (denominator first, then a row
summing to it), widths 6 to 8,193.
- The estimates are integer, not IEEE double: the library is fp-free and
the fp-free scan gates it (progress file, deviation 1). A floored
reciprocal for z (only the upward correction can fire, as §5.4 step 3
argues) and a rounded one for p (both corrections live, as step 4).
- tests/support/matmul_dispatch_instrument.h: softmax_fast/fallback for
AVX2 and AVX-512 (§3.6), inside the x64 block. Nothing increments them.
- tests/test_attn_rowsites.cpp: 4.S4, 7.S4 (premises of every correction
row), 2.S4, width 0, the replica's premise against the v1.9.0 body, 6.3
(S4 golden) and 11.1(d)'s softmax rows (L·H·N: 1,792 prefill, 448 decode,
0 fallback, re-derived on the base). Values are checked against a
test-side restatement of the v1.9.0 body over the unchanged public
IExpConstruct / IExpEvaluate; paths against the guard copy.
- tools/ci/sslm_axis_digest.cpp: c32_attention appends the S4 set.
- tools/gen_attn_rowsite_golden.cpp, tests/attn_rowsite_golden_pin.h: the
S4 hash, generated against the v1.9.0 tag.
GCC 13.3: auto and forced AVX2/AVX-512 fail 3,125 path assertions each,
forced SSE2 none; every value assertion passes on the base.
[skip ci]
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…tier Plan of record: attention and per-row sites, rev 3.1, slice S4 (§4.4, §5.4). SoftmaxRowQ15 gains a guarded fast path on the AVX2 and AVX-512BW tiers. The row guard, in dependency order: width <= 2^14, q_ln2 >= 1, q_c >= 0, M = q_b^2 + q_c (128-bit, as the shipped body forms it) in [1, 2^47], q_ln2 <= 2 q_b + 1, every score within +-2^61. Inside it the tier's body writes the row and returns true; outside it the shipped body runs unchanged and the tier's fallback counter moves. The fast counter moves inside each body. The tier decision goes through S2's DispatchSitesKernel, so SUPERSLM_SITES_AVX512_MSVC governs it on MSVC. The estimates are integer, not IEEE double as §4.4/§5.4 write them: the library is floating-point-free and the fp-free scan gates it. z is (a * floor(2^kz / q_ln2)) >> kz, which never overestimates, so only the upward correction can fire (the downward one is defensive); p is (e * round(2^62 / denom)) >> 47, within 1/2, so both corrections are live. The exact integer corrections decide, as §5.4 argues. - src/intmath.cpp: the guard, the row constants, and per tier an exp helper, a prob helper and the row body (anonymous namespace, per-function target attributes). Tails run one padded vector step (pass 2 pads with the row maximum, pass 3 with e = 0; never kept). The AVX2 body folds vpcmpgtq masks by add/sub/and and clips with vpminud on the low dword, so Clang emits no vblendvpd; the AVX-512 body uses F and BW only and no mask register (vpminuq, vpsraq sign masks). Tail pads are memcpy'd into an initialised buffer: GCC turned a per-lane select loop into a mask-register compare the fp-free scan rejects. - tools/ci/check_tiled_matmul_linkage.py: the S4 functions join the population (cell 11.3); a planted external body turns it red. - tools/ci/check_matmul_avx_isolation.py: the prose names the S4 bodies. GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512 suites, 0 failures; digests equal the red run's on all five legs. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The softmax mode of sslm_sites_bench and the slice's evidence: GCC and Clang suites and digests on every tier, the fp-free scan (Clang clean; GCC clean apart from the base's TiledGemmAvx512, which main fixes), all 30 mutants (the 22 §9 S4 rows, the 3 all-slice rows, 5 extras) killed on every binary that runs the code they mutate, sanitizers, 46 equal save blobs, the bench against plan §0 (AVX2: 0.10 / 0.46 / 0.91 ms per prompt token saved against 0.15 / 0.54 / 1.15, and 0.53 per decode token against 0.53), a local branch-coverage replica with allowlist notes, and the CHANGELOG and platform-support entries. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…-claims false positive [skip ci]
- include/superslm/gpu_layer_loop_guards.def: slices S1-S4 add 106 lines
above RunLayerLoopImpl in src/forward/forward_sites.cpp, so all nine
cpp_citation line numbers pointed at the wrong lines and
check_gpu_guard_status_parity.py failed (nine cells of its pytest with
it). Each citation is moved by 106 and re-checked by the checker against
the rejecting return it names; no guard, status or order changes.
- docs/attention-rowsites-s3-progress.md: the 11.4 row cited
`tests/ci/check_no_forward_leaf_calls.py` next to "green", which
check_ci_claims.py reads as a CI noun ("ci") plus an execution verb
("tests"). The row now names the script without its directory.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S5 (§4.5, §5.5).
Red-first: the cells, the seam's per-tier q31_row counters, the test-side
guard copy, the widened QK-norm fixture, the digest rows and the golden
pins, before any kernel lands. QkQ31ScoreRow is declared with a stub that
runs the per-key QkQ31Score loop; both layer loops are unchanged.
- include/superslm/forward_sites.h, src/forward/forward_sites.cpp: the
internal entry QkQ31ScoreRow(q, keys, ratio, head_dim, width, out), not
exported (C5), and its red stub.
- tests/support/attention_cases.h: the S5 set, driven through a
caller-supplied row function: §8 4.S5's grid (head_dim {4, 8, 60, 64,
128, 132, 256, 512, 513, 516} x width {1, 7, 8, 9, 15, 16, 17, 1,024},
uniform, int8-extreme and ratio-2^31 operands), 7.S5b's margin corners
(every limb sum at -2,147,418,112 at head_dim 512, the same rows at 516)
and 7.S5c's rounding ties. Every ratio in [1, 2^31] (§3.3).
- tests/support/qk_attention_fixture.h (new, header-only): cell 11.1(c)'s
widened QK-norm fixture, built as QkNormWiringFixture is: hidden 256,
4 query heads over 2 KV heads, head_dim 64, intermediate 256, one layer,
context_cap 32, q_norm and k_norm gains, ratio 2^31 on every channel,
fixed-seed int8 weights, a Pythagorean-triple RoPE table (no libm). Its
decode, chunk and decode-with-sink runs emit one stream per run.
- tests/support/matmul_dispatch_instrument.h: q31_row_fast/fallback for
AVX2 and AVX-512 (§3.6), inside the x64 block. Nothing increments them.
- tests/test_attn_rowsites.cpp: 4.S5, 6.1, 7.S5a-d (with the margin and
tie premises), width 0, 2.S5 (ratios -1, 2^32, 2^33, 2^38 and 2^48 at
the first, middle and last channel, and head_dim 513; each fails exactly
one conjunct; compared with the same binary's per-key QkQ31Score), 6.3
(the S5 hash), 11.1(c) (three runs over 24 positions: per-position
q31_row and row-table deltas, softmax and prob-V totals from run (iii)'s
classified rows, zero observe calls from the chunk loop, every run's
stream equal to the fixture pin) and 11.1(d)'s q31_row rows (0: the
0.5B artifact takes the plain score path).
- tools/ci/sslm_axis_digest.cpp: c32_attention appends the S5 set.
- tools/gen_attn_rowsite_golden.cpp, tests/attn_rowsite_golden_pin.h: the
S5 hash and the fixture hash, generated against the v1.9.0 tag.
GCC 13.3: auto and forced AVX2/AVX-512 fail 358 q31_row path assertions
each, forced SSE2 none; every value assertion passes on the base.
[skip ci]
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S5 (§4.5, §5.5). QkQ31ScoreRow computes every key's Q31 score for one query head (the Qwen3 QK-norm path) and both layer loops call it in place of their per-key QkQ31Score loops. On the AVX2 and AVX-512BW tiers, inside the guard (head_dim <= 512, every ratio in [0, 2^32)), each channel's w = q * ratio is split into three 16-bit pieces, a0 = w & 0x7FFF, a1 = (w >> 15) & 0x7FFF, a2 = w >> 30, and each piece's sum over the channels is a vpmaddwd into int32 lanes: at most head_dim * 128 * 32767, inside int32 by 65,535 at head_dim 512. The three sums recombine exactly in int64 and RoundingDivideByPOT(x, 31) is vectorised without a 64-bit compare or select. Outside the guard, and on the scalar and SSE2 tiers, the per-key loop runs (the same binary's v1.9.0 code) and the tier's fallback counter moves. The fast counter moves inside each body. The tier decision goes through S2's DispatchSitesKernel, so SUPERSLM_SITES_AVX512_MSVC governs it on MSVC. - src/forward/forward_sites.cpp: the guard, the limb set-up, an SSE2 key packer (keys widened to int16 and transposed per 4-channel quad into a stack buffer, 8 keys per AVX2 block, 16 per AVX-512 block) and per tier a rounding step and a row body (anonymous namespace, per-function target attributes); the dispatching QkQ31ScoreRow; both call sites. - tools/ci/check_tiled_matmul_linkage.py: the S5 functions join the population and the forward_sites.cpp objects are passed (cell 11.3); a planted external body turns it red. - tools/ci/check_matmul_avx_isolation.py: the prose names the S5 bodies. GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512 suites, 0 failures; digests equal the red run's on all five legs. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The q31 mode of sslm_sites_bench and the slice's evidence: GCC and Clang suites and digests on every tier, the fp-free scan (Clang clean; GCC clean apart from the base's TiledGemmAvx512, which main fixes), the §9 mutants (the 11 killable rows and 11 extras killed on every binary that runs the code they mutate; "a2 by logical shift" is equivalent in an int16-limb kernel and survives by construction), sanitizers, 66 equal save blobs, the bench against plan §0 (AVX2: 11.5 -> 0.47 ms per prompt token at T = 128 and 97.6 -> 3.23 at T = 1,024, against 10.9 -> 0.5 and 87 -> 3.9), a one-layer forward probe at Qwen3-0.6B width whose outputs equal the base's, a local branch-coverage replica, and the CHANGELOG and platform-support entries. The coverage replica found the limb and key-pack channel pads untaken: every in-guard head_dim of 4.S5 is a multiple of 4. The suite gains the 4.S5 channel-tail rows (head_dim 1-3, 5, 63, 66, 127, 130, 509, 511; outside the golden set); a mutant setting both pads nonzero dies on them and on no other cell. [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
- include/superslm/gpu_layer_loop_guards.def: slice S5 adds 225 lines above RunLayerLoopImpl in src/forward/forward_sites.cpp, so all nine cpp_citation line numbers pointed at the wrong lines again and check_gpu_guard_status_parity.py failed. Each citation is moved by 225 and re-checked by the checker against the rejecting return it names; no guard, status or order changes. check_ci_claims.py and check_present_tense_defect_comments.py pass unchanged on the S5 tree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…atmul.cpp and intmath.cpp GitHub-hosted runners lack AVX-512, so the S1-S5 AVX-512 kernels are uncovered on the branch-coverage cell. Set src/matmul.cpp and src/intmath.cpp one whole branch below the no-AVX-512 projection (204/296 and 201/242); the exact values are re-pinned from this PR's own CI run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…; print the recorded floors 11.1(d) is the only caller of PreflightScanWscFolds in superslm_tests and it skips without SUPERSLM_ATTN_ROWSITES_ARTIFACT, which left the function's twelve branch sides linked and never run, and include/superslm/layer_marshal.h below its 40.00 floor. A committed cell now builds a WSC1 manifest view and scans it over three, one and zero layers, covering all twelve sides (local clang-18 superslm_tests alone: layer_marshal.h 45.83%). The branch-coverage leg also prints build/measured_branch_coverage_floors.json to its log, so a re-pin can read the exact values without the artifact store. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…om CI run 36584892310 The run's runner had no AVX-512 (the forced AVX-512 suite exited on SIGILL), so the recorded values hold on both kinds of hosted runner. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…skip ci] Their generators emit LF and the headers claim byte-for-byte regeneration; a Windows checkout was CRLF (code review F2). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
SitesKernel, SelectSitesKernel, DispatchSitesKernel and RequantRowWide are internal C++ declarations following the installed-header pattern of GemmPath and RequantTokenCodeWide; the C API is unchanged. Comments only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
dansupergameprogrammer
marked this pull request as ready for review
September 30, 2026 14:46
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Before: the CPU forward pass handled five steps one element or one key at a time, using the full exact integer computation even where a row-level form gives the same answer faster:
After: each of the five runs per row, and every tier stays bit-identical to v1.9.0. The savings below are per token, measured at engine level on a shared 4-vCPU cloud Xeon with synthetic weights. They are not end-to-end figures for any consumer. S1–S4 are scaled to Qwen2.5-0.5B depth (24 layers, 14 heads), and S5 to Qwen3-0.6B depth (28 layers, 16 heads).
QkQ31ScoreRow):Bit-identity to v1.9.0
f740f833…on all ten legs: auto and forced scalar/SSE2/AVX2/AVX-512, each on GCC 13.3 and Clang 18.1. S5 appends its rows to thec32_attentionsection, so every other section equals S4's. Every leg also equals S5's red run, where the row entry still ran the v1.9.0 per-key loop.8836d5eb…, S2b0d1a6cd…, S33e3abed7…, S42e47ea3c…, S5daea9a39…. The QK-norm fixture that drives both layer loops is336b8d41….docs/attention-rowsites/s{1..5}/.How
gpu_1p0.h, the Gate A header parity check and the C ABI verb list are untouched. Five internal C++ declarations join installed headers, following the pattern already there (GemmPath/DispatchGemmPath,RequantTokenCodeWide,QkQ31ScoreForTier):SitesKernel,SelectSitesKernelandDispatchSitesKernelinmatmul.h,RequantRowWideinintmath.h, andQkQ31ScoreRowinforward_sites.h. None is exported, and each header comment says so.forward_sites.cppobjects to the linkage job.SUPERSLM_SITES_AVX512_MSVC=0until they have run there. The forced AVX-512 Windows legs build with the switch set to 1.docs/attention-rowsites-s{1..5}-progress.md.Verified on Linux (GCC 13.3 and Clang 18.1, Release):
pytest tests/ci/(withtests/t2296-fp-free-open-red-suite) matches main: 1,070 passed, with the same 4 environmental failures in the t2296 suite on both.check_gpu_guard_status_parity.py,check_present_tense_defect_comments.py,check_ci_claims.py, the linkage checker (on GCC objects, as CI runs it), the isolation ctest entries, the forward-leaf checker and the GPU census check.Code review: done on the Zen 2 box, with no bugs found. Its two findings are fixed here. The no-new-ABI promise now names the internal declarations (above). The two generated golden-pin headers are pinned to LF in
.gitattributes, so a Windows checkout matches their generators byte for byte.Fixes to existing files
gpu_layer_loop_guards.def. They move by 106 lines after S1–S4's additions toforward_sites.cpp, then by 225 more after S5's. The parity checker confirms each citation lands on its rejecting return.check_ci_claims.py.Coverage floors (re-pinned at the owner's word)
src/matmul.cppandsrc/intmath.cppare therefore pinned to 69.23076923076923 and 83.47107438016529. CI run 36584892310 recorded those values on a runner with no AVX-512 (the forced AVX-512 suite exited on SIGILL). A floor pinned from such a run holds on either kind of runner.include/superslm/layer_marshal.his not re-pinned. It had dropped because the 11.1(d) cell linksPreflightScanWscFoldsbut skips without its artifact. A committed cell,TestPreflightScanWscFolds, now runs that function on every build, and the file measures 45.83% against its 40.0 floor.Still owed:
🤖 Generated with Claude Code
https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68