Skip to content

Hardware SHA-256 for the model load check (x86 SHA extensions) - #2

Merged
dansupergameprogrammer merged 4 commits into
mainfrom
claude/project-thread-c8iecr
Sep 29, 2026
Merged

dansupergameprogrammer merged 4 commits into
mainfrom
claude/project-thread-c8iecr

Conversation

@dansupergameprogrammer

@dansupergameprogrammer dansupergameprogrammer commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Before: loading a model hashes the whole file with portable SHA-256 as an integrity check. On the owner's Zen 2 desktop that takes 3.25 s for the 510 MB Qwen2.5-0.5B model.

After: on CPUs with the x86 SHA extensions (most AMD since Zen, Intel since Ice Lake / Goldmont), the same check takes 0.31 s, about 10× faster. The digest is byte-identical; other CPUs and non-x86 targets keep the portable path, which also got about 1.7× faster.

How: src/sha256.cpp gains a per-function-attributed SHA-NI compression routine, picked once per process by a pure resolver over CPUID fields (leaf 7 EBX bit 29 plus SSSE3/SSE4.1), following matmul.cpp's dispatch conventions. SUPERSLM_FORCE_PORTABLE_SHA256 pins the portable path, and the new sha256_portable_forced_tests target keeps it covered on machines that have the extensions. No ABI, format or class-layout change; the new functions are noexcept. Tests were committed red first. The fp-free scan's check (A) gains a reviewed, named set of exactly sha256rnds2, sha256msg1 and sha256msg2 (integer-only; owner's decision 2026-09-29, ID pending), with the census re-pinned and the SHA-1 neighbours pinned as still rejected.

Verified: blind code review SHIP on the owner's Windows box (MSVC 19.33 and clang-cl 18, full suite and SHA suite green, SHA-NI path exercised and cross-checked against portable, digest matches Get-FileHash); CI green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68

Tests first, against an API that does not exist yet (the build fails to
compile them): a pure CPUID-field resolver, the dispatched implementation
matching the CPU probe, the FIPS 180-4 vectors (including the 896-bit and
one-million-'a' messages), every length 0..1024 at an aligned and an odd
offset, 3 MiB + 17 and 8 MiB buffers pinned from the current portable code,
and streaming updates split at every two-way point (lengths 0..256), every
three-way point (0..96) and byte-at-a-time (0..1024). Each compares the
dispatched path and the hardware path (where the CPU has it) against the
portable reference.

A second, small binary (sha256_portable_forced_tests) compiles
src/sha256.cpp with SUPERSLM_FORCE_PORTABLE_SHA256 so the portable arm of
the dispatch is exercised through the public API too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
… [skip ci]

src/sha256.cpp compresses blocks with sha256rnds2/sha256msg1/sha256msg2
when CPUID reports them (leaf 7 EBX bit 29, with SSSE3 and SSE4.1 from
leaf 1 ECX; leaf 7 is ignored below max basic leaf 7), resolved once per
process. The portable FIPS 180-4 code is unchanged as the fallback and is
the only path on a non-x64 target. Digests are byte-identical.

Dispatch follows matmul.cpp's conventions: a per-function
target("sha,sse4.1,ssse3") attribute on GCC/Clang (clang-cl included via
__clang__), empty on MSVC, never a TU-wide ISA flag; a pure resolver over
CPUID fields (ResolveSha256Impl) that the tests drive with fabricated
values; test-reachable one-shot references for each path; and a force
macro, SUPERSLM_FORCE_PORTABLE_SHA256, that pins the portable path.

Each compression function takes one block and the block loop lives in the
shared absorb/finish code, so the SHA-extension function has no branch of
its own: a coverage run on a CPU without the extensions loses no branch
coverage in this file. Measured against a multi-block loop inside the
function, the per-block call costs nothing outside run-to-run noise.
Whole blocks are now compressed straight from the caller's buffer rather
than copied through the 64-byte staging buffer a byte at a time, and the
padding is built in one stack block.

The new seams are noexcept and allocate nothing, so the S-HARDEN-7
bad_alloc membership (20 members) is unchanged. Sha256::Block (private)
is removed; the class layout is unchanged.

Throughput on a shared, loaded 4-core x86-64 development host (load
average about 10), 256 MiB buffer, best of N: portable before 220-223
MB/s; portable after 241-247 MB/s; SHA extensions 1113-1130 MB/s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Adds the measured load-hash figures from the Zen 2 code review (0.31 s
against 3.25 s for the 510 MB 0.5B model) and updates the tiled GEMM table:
the MSVC/clang-cl AVX2 legs have now executed on Windows CI and on the
Zen 2 desktop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer
dansupergameprogrammer marked this pull request as ready for review September 29, 2026 08:07

Copy link
Copy Markdown
Owner Author

fp-free-scan-gate (MSVC) and the scan step of linux-x64 fail with 1 REJECT: superslm::(anon)::CompressShaNi. The only rejected instructions are sha256rnds2, sha256msg1 and sha256msg2 themselves (the shuffles and adds pass). They are integer-only (32-bit modular adds, rotates, shifts and boolean ops; no rounding, no MXCSR), but check (A)'s lists are frozen and admit a new mnemonic only by a reviewed addition, as with bswap. That is a policy decision, so this PR waits on the owner's call before adding a named integer-hash set and re-pinning the census counts. Every other job is green.


Generated by Claude Code

The x86 SHA-extension compression function (CompressShaNi in
sha256.cpp.o / sha256.obj) was rejected by check (A) for sha256rnds2,
sha256msg1 and sha256msg2 and nothing else. Per the owner's decision of
2026-09-29 (decision ID pending), they are added as a reviewed, named
set, _X86_INTEGER_HASH_ALLOW, following the bswap precedent under the
design's vetting law (D-SLM4374). Intel SDM: each computes FIPS 180-4
round/schedule functions from 32-bit modular adds, rotates, shifts and
boolean ops on xmm lanes; no rounding, no MXCSR, no operand read as a
floating-point value. The set is exactly these three; the frozen
movement and p/vp allow-lists are untouched, and check (B) needs no
change (every form touches an xmm register).

Census re-pinned by running it: vocabulary 1523, accept_a 571 -> 574,
named_accept 117 (unchanged; the census names only _X86_VEC_MOVE_ALLOW),
structural_only 454 -> 457, with the three added to the pinned
structural-only fixture. New cells pin the three as accepted with their
own reason, the set's exact membership, and the SHA-1 neighbours
(sha1rnds4/sha1msg1/sha1msg2/sha1nexte) as still rejected.

Real archive scan (GCC Release, superslm, x86-64): 557 ACCEPT, 0 REJECT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
@dansupergameprogrammer
dansupergameprogrammer merged commit 3e2e0e0 into main Sep 29, 2026
78 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants