Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,18 @@ there; `SUPERSLM_TILED_AVX512_MSVC=1` turns it on. A GEMM call on the tiled path
per-call packing buffer; an allocation failure there surfaces through the existing
`SSLM_ALLOCATION_FAILED` status, with no state committed. No ABI, format or status change.

The load-time integrity hash uses the x86 SHA extensions when the CPU has them (CPUID leaf 7
EBX bit 29, with SSSE3 and SSE4.1), selected once per process; the portable SHA-256 stays as the
fallback and the only path on other targets. Digests are unchanged. Hashing the 510 MB Qwen2.5-0.5B model on a
Zen 2 desktop (MSVC) takes 0.31 s against 3.25 s before (about 1.5 GB/s against 150 MB/s). Defining
`SUPERSLM_FORCE_PORTABLE_SHA256` when compiling `src/sha256.cpp` pins the portable path; the new
`sha256_portable_forced_tests` target builds that way.
The FP-free scan's check (A) accepts `sha256rnds2`, `sha256msg1` and `sha256msg2`, the three
instructions the hardware path compiles to, through a new three-entry `_X86_INTEGER_HASH_ALLOW`
set in `check_fp_free_scan.py`. Each is 32-bit integer arithmetic on xmm lanes (Intel SDM:
modular adds, rotates, shifts and boolean ops; no rounding, no MXCSR, no floating-point operand).
The SHA-1 instructions stay rejected.

## [1.9.0] - 2026-09-25

`sslm_seq_save` writes a new save format, `SSB5`: the `SSB4` layout with the four per-site
Expand Down
15 changes: 15 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -122,6 +122,21 @@ endif()

enable_testing()
add_test(NAME superslm_tests COMMAND superslm_tests)

# The portable arm of src/sha256.cpp's run-time dispatch, pinned by
# SUPERSLM_FORCE_PORTABLE_SHA256 (the SHA-256 analogue of SUPERSLM_FORCE_SCALAR_MATMUL):
# superslm_tests covers whichever path the CPU selects, this binary covers the portable
# path through the same public API. It compiles src/sha256.cpp itself rather than a whole
# forced copy of the core library, since that file is the only one the macro touches.
add_executable(sha256_portable_forced_tests tests/sha256_portable_forced_tests.cpp src/sha256.cpp)
target_include_directories(sha256_portable_forced_tests PRIVATE include src)
target_compile_definitions(sha256_portable_forced_tests PRIVATE SUPERSLM_FORCE_PORTABLE_SHA256)
if(MSVC)
target_compile_options(sha256_portable_forced_tests PRIVATE /W4 /fp:precise)
else()
target_compile_options(sha256_portable_forced_tests PRIVATE -Wall -Wextra -ffp-contract=off)
endif()
add_test(NAME sha256_portable_forced_tests COMMAND sha256_portable_forced_tests)
endif()

# The §13 item 7 independent converter verifier (S-HARDEN-3, F13): loads a
Expand Down
15 changes: 14 additions & 1 deletion docs/platform-support.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,7 +81,7 @@ tiers, and for decode, nothing changes.
|---|---|---|
| GCC / Clang, AVX2 tier | on at M >= 8 | Bit-identity: full suite forced AVX2, tiled golden, cross-tier digest, save-blob equality against 1.9.0 on the in-tree fixture and on synthetic real-width artifacts |
| GCC / Clang, AVX-512 tier | on at M >= 8 | Same evidence, forced AVX-512 and auto dispatch |
| MSVC / clang-cl, AVX2 tier | on at M >= 8 | Built by the forced Windows legs; not yet executed on Windows |
| MSVC / clang-cl, AVX2 tier | on at M >= 8 | Full suite, auto and forced AVX2, green on the Windows CI legs and on a Zen 2 desktop (MSVC 19.33, clang-cl 15); digests equal to the GCC/Clang builds |
| MSVC / clang-cl, AVX-512 tier | **off** (`SUPERSLM_TILED_AVX512_MSVC=0`) | Held on the shipped per-row kernel until an MSVC AVX-512 build has executed the tiled kernel |

**Measured, engine level, one GEMM, on a 4-vCPU cloud Xeon (AVX2 and
Expand All @@ -106,6 +106,19 @@ are per-layer engine figures on synthetic weights. They are not a
statement about any consumer's end-to-end speed, and they are not a
measurement on this project's reference hardware.

### Model load integrity hash (unreleased)

Loading a model hashes the whole file with SHA-256. On x86-64 CPUs with the
SHA extensions (CPUID leaf 7 EBX bit 29, with SSSE3 and SSE4.1; most AMD CPUs
since Zen, Intel since Ice Lake and Goldmont) the hash uses those
instructions; everywhere else it uses the portable implementation. The digest
is the same either way.

**Measured, Zen 2 desktop (Ryzen 9 3950X), MSVC, 510 MB Qwen2.5-0.5B model,
best of 3:** 0.31 s with the SHA extensions against 3.25 s for 1.9.0's
portable hash. The portable path itself is also about 1.7x faster than in
1.9.0 (1.95 s).

### Damped-greedy decoding

The 1.2 candidate's opt-in decoder was confirmed on Windows x64 through the
Expand Down
51 changes: 50 additions & 1 deletion include/superslm/sha256.h
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,6 @@ class Sha256 {
// SslmArtifactAccess comment for the full reasoning.
friend struct Sha256Access;

void Block(const uint8_t* p);
uint32_t h_[8];
uint64_t total_bits_;
uint8_t buf_[64];
Expand All @@ -45,6 +44,56 @@ SUPERSLM_API void Sha256Hash(const uint8_t* data, size_t len, uint8_t out[32]);
// F5).
std::string ToHex(const uint8_t digest[32]);

// --- Block compression dispatch (verification seams) --------------------------
//
// src/sha256.cpp compresses blocks with the x86 SHA extensions (sha256rnds2,
// sha256msg1, sha256msg2) when the CPU reports them at run time, and with the
// portable FIPS 180-4 code otherwise. Both produce the same digest for the same
// bytes; the choice changes only speed. The portable path is the only one on a
// non-x64 target, and SUPERSLM_FORCE_PORTABLE_SHA256 (defined when compiling
// src/sha256.cpp) pins it on x64 too, the way SUPERSLM_FORCE_SCALAR_MATMUL pins
// matmul.cpp's scalar reference.
//
// The declarations below exist for verification, mirroring matmul.h's
// DotRowScalarRef/ResolveDotRowTier pattern: a test can drive each path directly
// and compare them, and can drive the CPUID decision with fabricated register
// values. They are noexcept: none of them allocates.

// Compile-time capability: the SHA-extension path exists in this build (the same
// target condition as matmul.h's SUPERSLM_MATMUL_HAVE_SIMD_X64).
#if defined(_M_X64) || defined(__x86_64__)
#define SUPERSLM_SHA256_HAVE_SHANI_X64 1
#else
#define SUPERSLM_SHA256_HAVE_SHANI_X64 0
#endif

inline constexpr int kSha256ImplPortable = 0;
inline constexpr int kSha256ImplShaNi = 1;

// Pure decision over CPUID fields: leaf 0 EAX (highest basic leaf), leaf 1 ECX
// (SSSE3 bit 9, SSE4.1 bit 19), leaf 7 sub-leaf 0 EBX (SHA bit 29; ignored when the
// highest basic leaf is below 7). Returns kSha256ImplShaNi only when all three
// feature bits are set, else kSha256ImplPortable.
int ResolveSha256Impl(int max_basic_leaf, int leaf1_ecx, int leaf7_ebx) noexcept;

// What this CPU supports: the resolver above applied to the real CPUID fields
// (always kSha256ImplPortable on a non-x64 build). Ignores the force macro.
int DetectSha256ImplForCpu() noexcept;

// What Sha256 and Sha256Hash actually use in this build on this CPU: the detected
// implementation, or kSha256ImplPortable under SUPERSLM_FORCE_PORTABLE_SHA256.
int ActiveSha256Impl() noexcept;

// One-shot SHA-256 through the portable compression, whatever the dispatch selects.
void Sha256HashPortableRef(const uint8_t* data, size_t len, uint8_t out[32]) noexcept;

#if SUPERSLM_SHA256_HAVE_SHANI_X64
// One-shot SHA-256 through the SHA-extension compression, whatever the dispatch
// selects. Precondition: DetectSha256ImplForCpu() == kSha256ImplShaNi (on any other
// CPU the instructions fault).
void Sha256HashShaNiRef(const uint8_t* data, size_t len, uint8_t out[32]) noexcept;
#endif

} // namespace superslm

#endif // SUPERSLM_SHA256_H
Loading
Loading