Hardware SHA-256 for the model load check (x86 SHA extensions) - #2
Merged
Merged
Conversation
Tests first, against an API that does not exist yet (the build fails to compile them): a pure CPUID-field resolver, the dispatched implementation matching the CPU probe, the FIPS 180-4 vectors (including the 896-bit and one-million-'a' messages), every length 0..1024 at an aligned and an odd offset, 3 MiB + 17 and 8 MiB buffers pinned from the current portable code, and streaming updates split at every two-way point (lengths 0..256), every three-way point (0..96) and byte-at-a-time (0..1024). Each compares the dispatched path and the hardware path (where the CPU has it) against the portable reference. A second, small binary (sha256_portable_forced_tests) compiles src/sha256.cpp with SUPERSLM_FORCE_PORTABLE_SHA256 so the portable arm of the dispatch is exercised through the public API too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
… [skip ci]
src/sha256.cpp compresses blocks with sha256rnds2/sha256msg1/sha256msg2
when CPUID reports them (leaf 7 EBX bit 29, with SSSE3 and SSE4.1 from
leaf 1 ECX; leaf 7 is ignored below max basic leaf 7), resolved once per
process. The portable FIPS 180-4 code is unchanged as the fallback and is
the only path on a non-x64 target. Digests are byte-identical.
Dispatch follows matmul.cpp's conventions: a per-function
target("sha,sse4.1,ssse3") attribute on GCC/Clang (clang-cl included via
__clang__), empty on MSVC, never a TU-wide ISA flag; a pure resolver over
CPUID fields (ResolveSha256Impl) that the tests drive with fabricated
values; test-reachable one-shot references for each path; and a force
macro, SUPERSLM_FORCE_PORTABLE_SHA256, that pins the portable path.
Each compression function takes one block and the block loop lives in the
shared absorb/finish code, so the SHA-extension function has no branch of
its own: a coverage run on a CPU without the extensions loses no branch
coverage in this file. Measured against a multi-block loop inside the
function, the per-block call costs nothing outside run-to-run noise.
Whole blocks are now compressed straight from the caller's buffer rather
than copied through the 64-byte staging buffer a byte at a time, and the
padding is built in one stack block.
The new seams are noexcept and allocate nothing, so the S-HARDEN-7
bad_alloc membership (20 members) is unchanged. Sha256::Block (private)
is removed; the class layout is unchanged.
Throughput on a shared, loaded 4-core x86-64 development host (load
average about 10), 256 MiB buffer, best of N: portable before 220-223
MB/s; portable after 241-247 MB/s; SHA extensions 1113-1130 MB/s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
Adds the measured load-hash figures from the Zen 2 code review (0.31 s against 3.25 s for the 510 MB 0.5B model) and updates the tiled GEMM table: the MSVC/clang-cl AVX2 legs have now executed on Windows CI and on the Zen 2 desktop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
dansupergameprogrammer
marked this pull request as ready for review
September 29, 2026 08:07
Owner
Author
|
Generated by Claude Code |
The x86 SHA-extension compression function (CompressShaNi in sha256.cpp.o / sha256.obj) was rejected by check (A) for sha256rnds2, sha256msg1 and sha256msg2 and nothing else. Per the owner's decision of 2026-09-29 (decision ID pending), they are added as a reviewed, named set, _X86_INTEGER_HASH_ALLOW, following the bswap precedent under the design's vetting law (D-SLM4374). Intel SDM: each computes FIPS 180-4 round/schedule functions from 32-bit modular adds, rotates, shifts and boolean ops on xmm lanes; no rounding, no MXCSR, no operand read as a floating-point value. The set is exactly these three; the frozen movement and p/vp allow-lists are untouched, and check (B) needs no change (every form touches an xmm register). Census re-pinned by running it: vocabulary 1523, accept_a 571 -> 574, named_accept 117 (unchanged; the census names only _X86_VEC_MOVE_ALLOW), structural_only 454 -> 457, with the three added to the pinned structural-only fixture. New cells pin the three as accepted with their own reason, the set's exact membership, and the SHA-1 neighbours (sha1rnds4/sha1msg1/sha1msg2/sha1nexte) as still rejected. Real archive scan (GCC Release, superslm, x86-64): 557 ACCEPT, 0 REJECT. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Before: loading a model hashes the whole file with portable SHA-256 as an integrity check. On the owner's Zen 2 desktop that takes 3.25 s for the 510 MB Qwen2.5-0.5B model.
After: on CPUs with the x86 SHA extensions (most AMD since Zen, Intel since Ice Lake / Goldmont), the same check takes 0.31 s, about 10× faster. The digest is byte-identical; other CPUs and non-x86 targets keep the portable path, which also got about 1.7× faster.
How:
src/sha256.cppgains a per-function-attributed SHA-NI compression routine, picked once per process by a pure resolver over CPUID fields (leaf 7 EBX bit 29 plus SSSE3/SSE4.1), followingmatmul.cpp's dispatch conventions.SUPERSLM_FORCE_PORTABLE_SHA256pins the portable path, and the newsha256_portable_forced_teststarget keeps it covered on machines that have the extensions. No ABI, format or class-layout change; the new functions arenoexcept. Tests were committed red first. The fp-free scan's check (A) gains a reviewed, named set of exactlysha256rnds2,sha256msg1andsha256msg2(integer-only; owner's decision 2026-09-29, ID pending), with the census re-pinned and the SHA-1 neighbours pinned as still rejected.Verified: blind code review SHIP on the owner's Windows box (MSVC 19.33 and clang-cl 18, full suite and SHA suite green, SHA-NI path exercised and cross-checked against portable, digest matches
Get-FileHash); CI green.🤖 Generated with Claude Code
https://claude.ai/code/session_01XfuJLAZ6EFSt7mB85oNm68