Skip to content

Exact CPU attention and per-row sites (plan slices S1–S5) - #3

Merged
dansupergameprogrammer merged 22 commits into
mainfrom
claude/project-thread-c8iecr
Sep 30, 2026
Merged

dansupergameprogrammer merged 22 commits into
mainfrom
claude/project-thread-c8iecr

Conversation

@dansupergameprogrammer

@dansupergameprogrammer dansupergameprogrammer commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Before: the CPU forward pass handled five steps one element or one key at a time, using the full exact integer computation even where a row-level form gives the same answer faster:

  • the per-row sites (SwiGLU sigmoid, residual landing, RMSNorm divide);
  • attention's probability-times-value step;
  • the requant funnel's element loop;
  • the softmax row;
  • on the QK-norm path, the Q31 attention score.

After: each of the five runs per row, and every tier stays bit-identical to v1.9.0. The savings below are per token, measured at engine level on a shared 4-vCPU cloud Xeon with synthetic weights. They are not end-to-end figures for any consumer. S1–S4 are scaled to Qwen2.5-0.5B depth (24 layers, 14 heads), and S5 to Qwen3-0.6B depth (28 layers, 16 heads).

  • S1, per-row tables: 256-entry tables for the sigmoid, landing and norm sites, at rows of 512 or more. Saves 0.80 ms/token.
  • S2, prob·V on int16 multiply-add: AVX2 and AVX-512BW, head_dim % 16 == 0, and only rows that pass the int16 condition. On AVX2, prefill saves 0.82 / 3.36 / 6.75 ms/token at T = 128 / 512 / 1,024, and decode saves 3.92 ms/token at context 300. The kernel is 16–21x faster per call.
  • S3, requant row leaf in 64-bit lanes: 4 lanes on AVX2, 8 on AVX-512BW. Saves 2.53 ms/token on AVX2 (2.66 on AVX-512).
  • S4, guarded softmax fast path: integer estimates with exact corrections, with the shipped body as the fallback. On AVX2, prefill saves 0.10 / 0.46 / 0.91 ms/token at T = 128 / 512 / 1,024, and decode saves 0.53 ms/token.
  • S5, Q31 score row (QkQ31ScoreRow):
    • Scope: the Qwen3 QK-norm path only. Qwen2.5 models do not take this path.
    • Method: every key of a query head in one call. Each channel's product splits exactly into three 16-bit pieces, which 16-bit multiply-add sums.
    • Guard: head_dim ≤ 512, and every ratio in [0, 2^32).
    • Savings: on AVX2 the attention score cost per token falls from 11.5 to 0.47 ms at T = 128 and from 97.6 to 3.23 ms at T = 1,024. On AVX-512 it falls from 9.5 to 0.47 and from 76.1 to 3.1.
    • A one-key row on AVX-512 is about 0.06 µs slower.
    • On the Zen 2 box (AVX2), the score row measured 0.588 vs 10.264 ms/token at 128 tokens and 3.63 vs 94.07 at 1,024, with zero mismatches.

Bit-identity to v1.9.0

  • Digest: the axis digest's GLOBAL hash is f740f833… on all ten legs: auto and forced scalar/SSE2/AVX2/AVX-512, each on GCC 13.3 and Clang 18.1. S5 appends its rows to the c32_attention section, so every other section equals S4's. Every leg also equals S5's red run, where the row entry still ran the v1.9.0 per-key loop.
  • Golden pins, generated from the v1.9.0 tag: S1 8836d5eb…, S2 b0d1a6cd…, S3 3e3abed7…, S4 2e47ea3c…, S5 daea9a39…. The QK-norm fixture that drives both layer loops is 336b8d41….
  • Save blobs and decoded tokens match the base in every row of each slice's blob protocol.
  • Those records are in docs/attention-rowsites/s{1..5}/.

How

  • Each slice adds a guard that admits only rows where the fast form is provably exact. Every other row goes to the v1.9.0 code, and per-tier taken/skipped counters let tests confirm which path ran.
  • The SIMD bodies have internal linkage and dispatch through the existing tier selection.
  • No C API change. gpu_1p0.h, the Gate A header parity check and the C ABI verb list are untouched. Five internal C++ declarations join installed headers, following the pattern already there (GemmPath/DispatchGemmPath, RequantTokenCodeWide, QkQ31ScoreForTier): SitesKernel, SelectSitesKernel and DispatchSitesKernel in matmul.h, RequantRowWide in intmath.h, and QkQ31ScoreRow in forward_sites.h. None is exported, and each header comment says so.
  • The linkage and isolation checkers cover the new bodies. S5 adds the three forward_sites.cpp objects to the linkage job.
  • The fp-free scan passes on both compilers, with its allow-lists unchanged and 0 REJECT:
    • GCC: 574 / 580 / 582 ACCEPT on auto / forced AVX2 / forced AVX-512.
    • Clang: 503 / 508 / 509.
  • On MSVC and clang-cl, the AVX-512 fast paths stay off by default behind SUPERSLM_SITES_AVX512_MSVC=0 until they have run there. The forced AVX-512 Windows legs build with the switch set to 1.
  • Deviations from the plan are recorded in docs/attention-rowsites-s{1..5}-progress.md.

Verified on Linux (GCC 13.3 and Clang 18.1, Release):

  • The auto and forced SSE2/AVX2/AVX-512 suites pass with 0 failures on both compilers, including the 11.1(d) artifact cell. Every binary prints all six pinned hashes.
  • The digests on all ten legs are identical, and the fp-free scan passes on all six libraries.
  • pytest tests/ci/ (with tests/t2296-fp-free-open-red-suite) matches main: 1,070 passed, with the same 4 environmental failures in the t2296 suite on both.
  • These all pass: check_gpu_guard_status_parity.py, check_present_tense_defect_comments.py, check_ci_claims.py, the linkage checker (on GCC objects, as CI runs it), the isolation ctest entries, the forward-leaf checker and the GPU census check.
  • CI on GitHub: green on 0242af8.

Code review: done on the Zen 2 box, with no bugs found. Its two findings are fixed here. The no-new-ABI promise now names the internal declarations (above). The two generated golden-pin headers are pinned to LF in .gitattributes, so a Windows checkout matches their generators byte for byte.

Fixes to existing files

  • Two commits update the nine line citations in gpu_layer_loop_guards.def. They move by 106 lines after S1–S4's additions to forward_sites.cpp, then by 225 more after S5's. The parity checker confirms each citation lands on its rejecting return.
  • The first of those commits also rewords one progress-file row that tripped check_ci_claims.py.

Coverage floors (re-pinned at the owner's word)

  • Hosted runners may lack AVX-512, and on those runners the new AVX-512 bodies cannot run.
  • src/matmul.cpp and src/intmath.cpp are therefore pinned to 69.23076923076923 and 83.47107438016529. CI run 36584892310 recorded those values on a runner with no AVX-512 (the forced AVX-512 suite exited on SIGILL). A floor pinned from such a run holds on either kind of runner.
  • include/superslm/layer_marshal.h is not re-pinned. It had dropped because the 11.1(d) cell links PreflightScanWscFolds but skips without its artifact. A committed cell, TestPreflightScanWscFolds, now runs that function on every build, and the file measures 45.83% against its 40.0 floor.
  • The coverage leg now also prints the recorded floors to its log.

Still owed:

  • A real Qwen3 artifact through the S5 kernel. A QK-norm artifact is still refused at map time on this host, so S5's evidence is the pinned QK-norm fixture plus a one-layer forward at Qwen3-0.6B width.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68

claude and others added 17 commits September 29, 2026 09:25
…row tables

Plan of record: attention and per-row sites, rev 3.1, slice S1 (§4.1). Red-first:
the cells, the seam's row-table counters, the digest section and the golden pin,
before any site changes.

- tests/support/matmul_dispatch_instrument.h: the six row-table path counters
  (§3.6), declared outside the x64 block so arm64 counts them too. Nothing
  increments them yet.
- tests/test_attn_rowsites.cpp: cells 4.S1 (width grid around 512, -128 at
  first/middle/last at the small gate scale), 1.S1, 5.S1, 7.S1a-d, 11.1(a) for
  the row tables, 6.3 and 11.1(d)'s row-table and trace-record rows (driven
  through the two layer loops when SUPERSLM_ATTN_ROWSITES_ARTIFACT names the
  0.5B-width 1-layer artifact). Each call is compared with a test-side copy of
  the v1.9.0 per-element loop and asserts its own counter delta.
- tests/support/rowsite_cases.h: the fixed input set shared by the digest
  section c_rowsites, the golden generator and the suite.
- tools/gen_attn_rowsite_golden.cpp and tests/attn_rowsite_golden_pin.h: the
  pin, generated against the v1.9.0 tag's library (8836d5eb..., 634,120 values).
- The sslm_bench_prefill forced targets gain tests/ on their include path; they
  did not compile on the base.

Red on GCC auto and forced SSE2/AVX2/AVX-512: 258 of 1,756 checks fail, every
one a counter assertion; every value assertion passes on the base. Digest
sections 1-9 are byte-identical to slice 1's.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…on every tier

Plan of record: attention and per-row sites, rev 3.1, slice S1 (§4.1, §5.1).

RmsNormSite's divide, MlpActSite's SiLU sigmoid and ResidualReconcileSite's
landing rescale each evaluate a pure function of one int8 code and row
constants. At n >= 512 (kRowTableMinWidth) each site now evaluates it once per
code into a stack table with exactly the arguments its per-element loop passes,
and the loop reads the table:

- norm: FloorDivI64(c << 32, root) for every int8 c (256 x int64);
- SiLU: SiluSigmoidQ15(table, c, m, e) for c in [-127, 127]; a -128 gate code
  is evaluated directly;
- landing: (LandingRescale value, int64-overflow flag) for c in [-127, 127] per
  candidate, with null counters; the loop keeps its order, first-flag return
  and overflow checks, reads the flag per element present, and lands -128
  directly.

The tables are off only under SUPERSLM_FORCE_SCALAR_MATMUL, which keeps the
v1.9.0 loops as the reference axis. Each site counts its table decision once
per call on the instrument seam's row-table counters.

GCC 13.3 Release, auto (AVX-512 here) and forced SSE2/AVX2/AVX-512: 1,756 S1
checks, 0 failures; full suites 27,287 / 27,229 / 27,245 / 27,245 checks, 0
failures. S1 golden 8836d5eb... matched on every binary; SiLU-LUT, matmul and
tiled golden pins unchanged. Digests: all five legs equal; sections 1-9 equal
to slice 1's; c_rowsites equals the v1.9.0 pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…he bench

Plan of record: attention and per-row sites, rev 3.1, slice S1.

- tests/test_attn_rowsites.cpp: cell 3.S1, 8 threads calling all three
  sites at table widths against the single-threaded reference. The plan's
  3.S1 names an artifact-driven concurrent-read cell that does not exist.
- tools/sslm_sites_bench.cpp (cell 10.1): kernel, prefill and decode modes;
  one source built against either library.
- tools/t2147_chunk_batched_pins.cpp: --chunk-budget=B and --decode=D in
  --dump-blob mode, for the plan's blob protocol.
- docs/attention-rowsites/s1/: suites (GCC and Clang, four binaries each, plus
  GCC TSan and ASan+UBSan), blob protocol (33 comparisons, all equal, with 32
  decode steps each), 15 mutants with scripts (all killed on the auto and
  forced-AVX2 binaries), bench.md.
- docs/attention-rowsites-s1-progress.md: state, deviations, resume steps.

Measured saving at full 0.5B depth: 0.80 ms per token by the plan's per-site
method (estimate 0.87); 0.96-0.99 from reduced-layer prefill. The decode
saving is below this shared host's noise.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…ltiply-add

Plan of record: attention and per-row sites, rev 3.1, slice S2 (§4.2, §5.2).
Red-first: the cells, the seam's per-tier prob-V counters, the 11.2 selector,
the digest section and the golden pin, before any kernel lands.

- include/superslm/matmul.h: detail::SitesKernel, SelectSitesKernel (the pure
  selector over tier, SUPERSLM_SITES_AVX512_MSVC and MSVC identity) and
  DispatchSitesKernel (its wiring). src/matmul.cpp carries a stub that selects
  the v1.9.0 code on every tier, which is what the base does.
- tests/support/matmul_dispatch_instrument.h: pv_fast / pv_fallback for AVX2 and
  AVX-512 (§3.6), inside the x64 block. Nothing increments them yet.
- tests/test_attn_rowsites.cpp: cells 11.2 (full truth table and wiring), 4.S2
  (head_dim x width grid, odd widths, exact-size heap buffers), 7.S2 corners,
  2.S2 (four hostile rows, each failing exactly one conjunct of a test-side
  guard copy), 6.3 (S2 golden) and 11.1(d)'s prob-V rows (L*H*N calls,
  data terms 15 prefill / 0 decode, re-derived on this base with the plan's
  probe). Every call asserts output against a test-side copy of the v1.9.0
  loop and its own counter delta on the selected kernel's tier only.
- tests/support/attention_cases.h: the fixed prob-V input set shared by the new
  digest section c32_attention, the golden generator and the suite.
- tools/gen_attn_rowsite_golden.cpp: one hash per slice; tests/
  attn_rowsite_golden_pin.h regenerated against the v1.9.0 tag (S1 unchanged,
  S2 b0d1a6cd... over 30,100 values).

Red on GCC auto and forced AVX2/AVX-512: 269 of 2,306 attn-rowsites checks fail
each, forced SSE2 3 (the selector); every value assertion passes on the base.
Digest sections 1-10 are byte-identical to S1's on all five legs.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…on every tier

Plan of record: attention and per-row sites, rev 3.1, slice S2 (§4.2, §5.2).

GemmProbQ15Accumulate zeroes its output and runs a tiered accumulate-into
core. On the AVX2 and AVX-512BW tiers, when head_dim is a multiple of 16 and
the row passes the int16 condition (every p in [0, 32767], Sum p <= 2^15,
checked in one pass by the kernel itself), it pairs keys: the two value rows
are interleaved byte by byte, widened to int16 and multiplied by the
broadcast (p_k, p_{k+1}) pair with vpmaddwd into one int32 lane per output
dimension (an odd last key pairs with a zero row). AVX2 works in 16-dimension
units; AVX-512BW in 32-dimension units interleaved in-lane (the stores put
each half back at its dimensions), plus one 16-dimension tail unit. Each lane is bounded by
128 * Sum p <= 2^22, so the widened sum equals the int64 loop exactly. Any
other row, and the scalar and SSE2 tiers, run the v1.9.0 loop. No allocation:
the pairs are formed in registers.

- detail::SelectSitesKernel / DispatchSitesKernel (cell 11.2): the new kernels
  run on AVX2 and AVX-512; in MSVC and clang-cl builds the AVX-512 tier keeps
  the v1.9.0 code unless SUPERSLM_SITES_AVX512_MSVC=1 (default 0), a switch
  separate from the tiled GEMM's. Both forced AVX-512 Windows legs build with
  it on.
- Path counters (§3.6): each tier's fast counter moves inside that tier's own
  body, the fallback counter in the dispatcher after the guard decides.
- check_tiled_matmul_linkage.py and the isolation checker's prose name the new
  attributed functions; the linkage check is OK on the auto and both forced
  objects.

GCC 13.3 Release, auto (AVX-512 here) and forced SSE2/AVX2/AVX-512: 2,306
attn-rowsites checks, 0 failures; full suites 27,837 / 27,779 / 27,795 /
27,795, 0 failures. S1 and S2 goldens matched on every binary; SiLU-LUT,
matmul and tiled goldens unchanged. Digests: all five legs equal and equal to
the red run's.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The 4.S2 kernel-blocking rows (every AVX2 and AVX-512 block count and the
AVX-512 16-dimension tail), the prob-V mode of sslm_sites_bench, and the
slice's evidence: GCC and Clang suites and digests on every tier, 20 mutants
killed, sanitizers, 45 equal save blobs, the bench against plan §0, a local
branch-coverage replica with allowlist lines, and the CHANGELOG and
platform-support entries. [skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S3 (§4.3, §5.3).
Red-first: the cells, the seam's per-tier requant_row counters, the row leaf's
declaration, the digest rows and the golden pin, before any kernel lands.

- include/superslm/intmath.h: RequantRowWide(x, n, r, s, out), the funnel's
  element loop as one leaf, with the funnel's contract. src/intmath.cpp
  carries a stub that runs the element loop, which is what the base does;
  the funnel still runs its own loop.
- tests/support/matmul_dispatch_instrument.h: requant_row for AVX2 and
  AVX-512 (§3.6), inside the x64 block. Nothing increments them yet.
- tests/test_attn_rowsites.cpp: 4.S3 (n {1, 3, 4, 5, 7, 8, 9, 896, 4864} x
  d' covering every s in [-1, 30], +-d' in every lane position, random rows
  with values next to rounding ties; a sentinel-fenced pass on every binary
  and an exact-size heap pass for the hosted ASan leg), 7.S3's P = 2^63
  corner premises, the funnel's call site (one leaf call per funnel call that
  passes its preflight), 6.3 (S3 golden) and 11.1(d)'s requant rows
  ((11L + 1) N: 1,536 prefill, 384 decode). Codes are checked against the
  unchanged RequantTokenCodeWide; the path against the selected kernel's
  tier only.
- tests/support/rowsite_cases.h: RunRequantRowCases, through
  RequantChainChecked (a v1.9.0 signature), appended to c_rowsites and pinned
  by the generator against the v1.9.0 tag (S3 3e3abed7..., 3,567,018 values;
  S1 and S2 unchanged).

Red on GCC auto and forced AVX2/AVX-512: 10 path assertions fail each,
forced SSE2 none; every value assertion passes on the base. Digest sections
other than c_rowsites are byte-identical to S2's on all five legs.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…, green on every tier

Plan of record: attention and per-row sites, rev 3.1, slice S3 (§4.3, §5.3).

The checked chain funnel's element loop is one call to the new row leaf
RequantRowWide(x, n, r, s, out). On the AVX2 and AVX-512BW tiers it runs
the §5.3 identity in 4 or 8 unsigned 64-bit lanes: P = |x|*r from the
32-bit halves of r (r can be 2^32), H by a logical shift (P reaches exactly
2^63 at the contract's corner d' = |x| = 2^31, r = 2^32),
magnitude = (127H + ((127L + 2^(e-1)) >> 32)) >> (e - 32), then the element
code's clamp at 127 and sign restore. The last n mod 4 (or 8) elements, and
every element on the scalar and SSE2 tiers, run RequantTokenCodeWide. The
AVX-512 body uses F and BW instructions only and no mask register (vpabsq,
vpsraq, vpminuq, vpmovqb). No runtime guard (the funnel's preflight is the
contract), so one per-tier counter, inside each body.

- src/intmath.cpp: the two bodies (this file's first target-attributed
  functions, per function, anonymous namespace) and the dispatcher, which
  decides through S2's DispatchSitesKernel, so SUPERSLM_SITES_AVX512_MSVC
  governs it on MSVC. intmath.cpp now includes superslm/matmul.h and needs
  matmul.cpp at link time.
- build_cert.bat, both lines of tools/build_inspect.bat and the loop's cl
  line of tests/t2296-fp-free-open-red-suite/build_link_red.bat gain
  src\matmul.cpp (plan G21, §3.4, cell 11.5). GCC link check committed.
- tests/ci/check_no_forward_leaf_calls.py: RequantRowWide is a banned leaf
  (cell 11.4); a planted call from forward_sites.cpp turns it red.
- tools/ci/check_tiled_matmul_linkage.py: the requant bodies join the
  population and the intmath.cpp objects join the CI job; the record rule
  applies to matmul objects only (cell 11.3; the plant turns it red).
- tools/ci/check_matmul_avx_isolation.py: the prose names the two bodies.

The AVX2 body does |x|, the clamp and the sign in 32-bit lanes (vpabsd,
vpminud, vpsignd) and keeps its shuffle constant's halves distinct: Clang
lowers 64-bit selects to vblendvpd/vxorpd and loads a repeated-half
constant with vbroadcasti128, all three outside the fp-free scan's
allow-list. The scan passes on the Clang build unwidened; on the GCC build
it rejects only the base's TiledGemmAvx512 (fixed on main by 90e48de) and
passes with that fix applied.

GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512
suites, 0 failures; digests equal the red run's on all five legs.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The requant mode of sslm_sites_bench and the slice's evidence: GCC and
Clang suites and digests on every tier, the fp-free scan (Clang clean;
GCC clean apart from the base's TiledGemmAvx512, which main fixes),
every §9 mutant killed per tier, sanitizers, 46 equal save blobs, the
bench against plan §0 (2.5 ms/token saved on AVX2 against the 1.22
estimated), a local branch-coverage replica with allowlist lines, and
the CHANGELOG and platform-support entries. [skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S4 (§4.4, §5.4).
Red-first: the cells, the seam's per-tier softmax counters, the test-side
guard copy and estimate replica, the digest rows and the golden pin, before
any kernel lands. SoftmaxRowQ15 is unchanged.

- tests/support/attention_cases.h: the S4 set (§8 4.S4's grid at widths
  {1, 2, 3, 4, 5, 2^14} plus the kernels' block edges, realistic constants
  from IExpScaleConstants over the forward's scale range, four row kinds,
  aliased rows; the inside corners; total = 1; 2.S4's hostile rows, each
  failing exactly one conjunct, and the off-ratio witness), the test-side
  copy of §5.4's guard (TestSoftmaxGuard, one flag per conjunct), and a
  replica of the fast path's estimates (SoftmaxEstimateReplica) that picks
  the correction rows: 24 each for z up, p up and p down, each kept only
  where the correction fires and leaving it out changes the output. The p
  down rows come from a steered generator (denominator first, then a row
  summing to it), widths 6 to 8,193.
- The estimates are integer, not IEEE double: the library is fp-free and
  the fp-free scan gates it (progress file, deviation 1). A floored
  reciprocal for z (only the upward correction can fire, as §5.4 step 3
  argues) and a rounded one for p (both corrections live, as step 4).
- tests/support/matmul_dispatch_instrument.h: softmax_fast/fallback for
  AVX2 and AVX-512 (§3.6), inside the x64 block. Nothing increments them.
- tests/test_attn_rowsites.cpp: 4.S4, 7.S4 (premises of every correction
  row), 2.S4, width 0, the replica's premise against the v1.9.0 body, 6.3
  (S4 golden) and 11.1(d)'s softmax rows (L·H·N: 1,792 prefill, 448 decode,
  0 fallback, re-derived on the base). Values are checked against a
  test-side restatement of the v1.9.0 body over the unchanged public
  IExpConstruct / IExpEvaluate; paths against the guard copy.
- tools/ci/sslm_axis_digest.cpp: c32_attention appends the S4 set.
- tools/gen_attn_rowsite_golden.cpp, tests/attn_rowsite_golden_pin.h: the
  S4 hash, generated against the v1.9.0 tag.

GCC 13.3: auto and forced AVX2/AVX-512 fail 3,125 path assertions each,
forced SSE2 none; every value assertion passes on the base.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…tier

Plan of record: attention and per-row sites, rev 3.1, slice S4 (§4.4, §5.4).

SoftmaxRowQ15 gains a guarded fast path on the AVX2 and AVX-512BW tiers.
The row guard, in dependency order: width <= 2^14, q_ln2 >= 1, q_c >= 0,
M = q_b^2 + q_c (128-bit, as the shipped body forms it) in [1, 2^47],
q_ln2 <= 2 q_b + 1, every score within +-2^61. Inside it the tier's body
writes the row and returns true; outside it the shipped body runs unchanged
and the tier's fallback counter moves. The fast counter moves inside each
body. The tier decision goes through S2's DispatchSitesKernel, so
SUPERSLM_SITES_AVX512_MSVC governs it on MSVC.

The estimates are integer, not IEEE double as §4.4/§5.4 write them: the
library is floating-point-free and the fp-free scan gates it. z is
(a * floor(2^kz / q_ln2)) >> kz, which never overestimates, so only the
upward correction can fire (the downward one is defensive); p is
(e * round(2^62 / denom)) >> 47, within 1/2, so both corrections are live.
The exact integer corrections decide, as §5.4 argues.

- src/intmath.cpp: the guard, the row constants, and per tier an exp
  helper, a prob helper and the row body (anonymous namespace, per-function
  target attributes). Tails run one padded vector step (pass 2 pads with
  the row maximum, pass 3 with e = 0; never kept). The AVX2 body folds
  vpcmpgtq masks by add/sub/and and clips with vpminud on the low dword,
  so Clang emits no vblendvpd; the AVX-512 body uses F and BW only and no
  mask register (vpminuq, vpsraq sign masks). Tail pads are memcpy'd into
  an initialised buffer: GCC turned a per-lane select loop into a
  mask-register compare the fp-free scan rejects.
- tools/ci/check_tiled_matmul_linkage.py: the S4 functions join the
  population (cell 11.3); a planted external body turns it red.
- tools/ci/check_matmul_avx_isolation.py: the prose names the S4 bodies.

GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512
suites, 0 failures; digests equal the red run's on all five legs.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The softmax mode of sslm_sites_bench and the slice's evidence: GCC and
Clang suites and digests on every tier, the fp-free scan (Clang clean;
GCC clean apart from the base's TiledGemmAvx512, which main fixes), all
30 mutants (the 22 §9 S4 rows, the 3 all-slice rows, 5 extras) killed
on every binary that runs the code they mutate, sanitizers, 46 equal
save blobs, the bench against plan §0 (AVX2: 0.10 / 0.46 / 0.91 ms per
prompt token saved against 0.15 / 0.54 / 1.15, and 0.53 per decode token
against 0.53), a local branch-coverage replica with allowlist notes, and
the CHANGELOG and platform-support entries.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…-claims false positive [skip ci]

- include/superslm/gpu_layer_loop_guards.def: slices S1-S4 add 106 lines
  above RunLayerLoopImpl in src/forward/forward_sites.cpp, so all nine
  cpp_citation line numbers pointed at the wrong lines and
  check_gpu_guard_status_parity.py failed (nine cells of its pytest with
  it). Each citation is moved by 106 and re-checked by the checker against
  the rejecting return it names; no guard, status or order changes.
- docs/attention-rowsites-s3-progress.md: the 11.4 row cited
  `tests/ci/check_no_forward_leaf_calls.py` next to "green", which
  check_ci_claims.py reads as a CI noun ("ci") plus an execution verb
  ("tests"). The row now names the script without its directory.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S5 (§4.5, §5.5).
Red-first: the cells, the seam's per-tier q31_row counters, the test-side
guard copy, the widened QK-norm fixture, the digest rows and the golden
pins, before any kernel lands. QkQ31ScoreRow is declared with a stub that
runs the per-key QkQ31Score loop; both layer loops are unchanged.

- include/superslm/forward_sites.h, src/forward/forward_sites.cpp: the
  internal entry QkQ31ScoreRow(q, keys, ratio, head_dim, width, out), not
  exported (C5), and its red stub.
- tests/support/attention_cases.h: the S5 set, driven through a
  caller-supplied row function: §8 4.S5's grid (head_dim {4, 8, 60, 64,
  128, 132, 256, 512, 513, 516} x width {1, 7, 8, 9, 15, 16, 17, 1,024},
  uniform, int8-extreme and ratio-2^31 operands), 7.S5b's margin corners
  (every limb sum at -2,147,418,112 at head_dim 512, the same rows at 516)
  and 7.S5c's rounding ties. Every ratio in [1, 2^31] (§3.3).
- tests/support/qk_attention_fixture.h (new, header-only): cell 11.1(c)'s
  widened QK-norm fixture, built as QkNormWiringFixture is: hidden 256,
  4 query heads over 2 KV heads, head_dim 64, intermediate 256, one layer,
  context_cap 32, q_norm and k_norm gains, ratio 2^31 on every channel,
  fixed-seed int8 weights, a Pythagorean-triple RoPE table (no libm). Its
  decode, chunk and decode-with-sink runs emit one stream per run.
- tests/support/matmul_dispatch_instrument.h: q31_row_fast/fallback for
  AVX2 and AVX-512 (§3.6), inside the x64 block. Nothing increments them.
- tests/test_attn_rowsites.cpp: 4.S5, 6.1, 7.S5a-d (with the margin and
  tie premises), width 0, 2.S5 (ratios -1, 2^32, 2^33, 2^38 and 2^48 at
  the first, middle and last channel, and head_dim 513; each fails exactly
  one conjunct; compared with the same binary's per-key QkQ31Score), 6.3
  (the S5 hash), 11.1(c) (three runs over 24 positions: per-position
  q31_row and row-table deltas, softmax and prob-V totals from run (iii)'s
  classified rows, zero observe calls from the chunk loop, every run's
  stream equal to the fixture pin) and 11.1(d)'s q31_row rows (0: the
  0.5B artifact takes the plain score path).
- tools/ci/sslm_axis_digest.cpp: c32_attention appends the S5 set.
- tools/gen_attn_rowsite_golden.cpp, tests/attn_rowsite_golden_pin.h: the
  S5 hash and the fixture hash, generated against the v1.9.0 tag.

GCC 13.3: auto and forced AVX2/AVX-512 fail 358 q31_row path assertions
each, forced SSE2 none; every value assertion passes on the base.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Plan of record: attention and per-row sites, rev 3.1, slice S5 (§4.5, §5.5).

QkQ31ScoreRow computes every key's Q31 score for one query head (the
Qwen3 QK-norm path) and both layer loops call it in place of their
per-key QkQ31Score loops. On the AVX2 and AVX-512BW tiers, inside the
guard (head_dim <= 512, every ratio in [0, 2^32)), each channel's
w = q * ratio is split into three 16-bit pieces, a0 = w & 0x7FFF,
a1 = (w >> 15) & 0x7FFF, a2 = w >> 30, and each piece's sum over the
channels is a vpmaddwd into int32 lanes: at most head_dim * 128 * 32767,
inside int32 by 65,535 at head_dim 512. The three sums recombine exactly
in int64 and RoundingDivideByPOT(x, 31) is vectorised without a 64-bit
compare or select. Outside the guard, and on the scalar and SSE2 tiers,
the per-key loop runs (the same binary's v1.9.0 code) and the tier's
fallback counter moves. The fast counter moves inside each body. The
tier decision goes through S2's DispatchSitesKernel, so
SUPERSLM_SITES_AVX512_MSVC governs it on MSVC.

- src/forward/forward_sites.cpp: the guard, the limb set-up, an SSE2 key
  packer (keys widened to int16 and transposed per 4-channel quad into a
  stack buffer, 8 keys per AVX2 block, 16 per AVX-512 block) and per tier
  a rounding step and a row body (anonymous namespace, per-function target
  attributes); the dispatching QkQ31ScoreRow; both call sites.
- tools/ci/check_tiled_matmul_linkage.py: the S5 functions join the
  population and the forward_sites.cpp objects are passed (cell 11.3); a
  planted external body turns it red.
- tools/ci/check_matmul_avx_isolation.py: the prose names the S5 bodies.

GCC 13.3 and Clang 18.1: auto (AVX-512 here) and forced SSE2/AVX2/AVX-512
suites, 0 failures; digests equal the red run's on all five legs.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
The q31 mode of sslm_sites_bench and the slice's evidence: GCC and Clang
suites and digests on every tier, the fp-free scan (Clang clean; GCC
clean apart from the base's TiledGemmAvx512, which main fixes), the §9
mutants (the 11 killable rows and 11 extras killed on every binary that
runs the code they mutate; "a2 by logical shift" is equivalent in an
int16-limb kernel and survives by construction), sanitizers, 66 equal
save blobs, the bench against plan §0 (AVX2: 11.5 -> 0.47 ms per prompt
token at T = 128 and 97.6 -> 3.23 at T = 1,024, against 10.9 -> 0.5 and
87 -> 3.9), a one-layer forward probe at Qwen3-0.6B width whose outputs
equal the base's, a local branch-coverage replica, and the CHANGELOG and
platform-support entries.

The coverage replica found the limb and key-pack channel pads untaken:
every in-guard head_dim of 4.S5 is a multiple of 4. The suite gains the
4.S5 channel-tail rows (head_dim 1-3, 5, 63, 66, 127, 130, 509, 511;
outside the golden set); a mutant setting both pads nonzero dies on
them and on no other cell.

[skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
- include/superslm/gpu_layer_loop_guards.def: slice S5 adds 225 lines
  above RunLayerLoopImpl in src/forward/forward_sites.cpp, so all nine
  cpp_citation line numbers pointed at the wrong lines again and
  check_gpu_guard_status_parity.py failed. Each citation is moved by 225
  and re-checked by the checker against the rejecting return it names; no
  guard, status or order changes.

check_ci_claims.py and check_present_tense_defect_comments.py pass
unchanged on the S5 tree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer dansupergameprogrammer changed the title Exact CPU attention and per-row sites (plan slices S1–S4) Exact CPU attention and per-row sites (plan slices S1–S5) Sep 29, 2026
…atmul.cpp and intmath.cpp

GitHub-hosted runners lack AVX-512, so the S1-S5 AVX-512 kernels are
uncovered on the branch-coverage cell. Set src/matmul.cpp and
src/intmath.cpp one whole branch below the no-AVX-512 projection
(204/296 and 201/242); the exact values are re-pinned from this PR's
own CI run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…; print the recorded floors

11.1(d) is the only caller of PreflightScanWscFolds in superslm_tests and it
skips without SUPERSLM_ATTN_ROWSITES_ARTIFACT, which left the function's twelve
branch sides linked and never run, and include/superslm/layer_marshal.h below
its 40.00 floor. A committed cell now builds a WSC1 manifest view and scans it
over three, one and zero layers, covering all twelve sides (local clang-18
superslm_tests alone: layer_marshal.h 45.83%).

The branch-coverage leg also prints build/measured_branch_coverage_floors.json
to its log, so a re-pin can read the exact values without the artifact store.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…om CI run 36584892310

The run's runner had no AVX-512 (the forced AVX-512 suite exited on SIGILL), so
the recorded values hold on both kinds of hosted runner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
…skip ci]

Their generators emit LF and the headers claim byte-for-byte regeneration;
a Windows checkout was CRLF (code review F2).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
SitesKernel, SelectSitesKernel, DispatchSitesKernel and RequantRowWide are
internal C++ declarations following the installed-header pattern of
GemmPath and RequantTokenCodeWide; the C API is unchanged. Comments only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer
dansupergameprogrammer marked this pull request as ready for review September 30, 2026 14:46
@dansupergameprogrammer
dansupergameprogrammer merged commit 3a7e0f9 into main Sep 30, 2026
78 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants