perf(backend-native-cpu): runtime FEAT_DotProd dispatch for Apple arm64 Q4_K/Q6_K kernels (#958) - #961
Merged
Conversation
…64 q4k/q6k kernels (#958) A Kotlin/Native klib embeds exactly one static archive, and Apple A12 (iPhone XS/XR, still iOS-supported) lacks FEAT_DotProd while A13+ and all Apple Silicon have it — so the Android two-.so trick doesn't translate and TU-level -march=+dotprod is not shippable to iOS. Apple arm64 TUs now compile at the SDK-default baseline; the q4k/q6k dotprod hot bodies are extracted into _dp (target("dotprod")-attributed) / _generic twins and selected once per matmul call via a cached sysctlbyname(hw.optional.arm.FEAT_DotProd) probe (key since iOS 15 / macOS 12; absence degrades to the scalar-int arm, never crashes). The attribute is what gates vdotq_s32 codegen, and attributed functions are not inlined into baseline callers, keeping sdot out of the baseline path. Non-Apple builds keep the compile-time guard as the only mechanism: with SKAINET_HAVE_DOTPROD the call site is a direct call that inlines back under -O3 — verified: qemu linuxArm64Test parity green, 32 sdot instructions in the cross archive, no _dp symbols (inlined). The macOS FFM dylib moves from TU-level dotprod to baseline+dispatch — runtime-equivalent on every Apple Silicon Mac, and it turns the macOS jvmTest CI lane into an exerciser of the dispatch fast arm. iOS builds are static-only (SKAINET_STATIC_ONLY, auto-on for CMAKE_SYSTEM_NAME=iOS). Refs #920
aharakal
approved these changes
Aug 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #958; first slice of the iOS/Apple kernel track of #920 (prerequisite for #959, which adds the iosArm64/iosSimulatorArm64/macosArm64 targets with embedded archives).
Why: a K/N klib embeds exactly ONE static archive (the Android baseline/v8.2 two-
.soruntime pick doesn't translate), and Apple A12 — iPhone XS/XR, still supported by current iOS — lacks FEAT_DotProd while A13+/M-series have it. TU-level-march=+dotprodwould SIGILL supported devices; baseline-only would leave Q4_K/Q6_K (the formats real GGUFs use) scalar.How:
skainet_cpu_features.{h,c}:skainet_cpu_has_dotprod()— on Apple arm64 a cachedsysctlbyname("hw.optional.arm.FEAT_DotProd")probe (key exists since iOS 15/macOS 12; absence → 0 → scalar-int arm, never crashes); elsewhere a compile-time constant. Internal header, not in the cinteropheaderFilter.skainet_simd.h:SKAINET_DOTPROD_DISPATCH/SKAINET_DOTPROD_TARGET(__attribute__((target("dotprod")))) defined only for Apple arm64 TUs built without__ARM_FEATURE_DOTPROD. The attribute is what gatesvdotq_s32codegen in clang, and attributed functions are never inlined into baseline callers —sdotstays out of the baseline path.q4k_matmul.c/q6k_matmul.c: the guarded hot bodies extracted into_dp/_generictwins (verbatim bodies, one call per block × output row — call overhead amortized over the block's arithmetic), 3-way call site withuse_dphoisted once per matmul.CMakeLists.txt: aarch64-marchblock nowAND NOT APPLE;skainet_cpu_features.cadded;SKAINET_STATIC_ONLYoption (auto-on forCMAKE_SYSTEM_NAME=iOS— no dylib story there).Behavior notes: Linux codegen is unchanged — on dotprod TUs the call site is a direct call that inlines back under
-O3. The macOS host dylib (JVM/FFM path) moves from TU-level dotprod to baseline+dispatch: runtime-equivalent on every Apple Silicon Mac, and it makes the existing macos-14jvmTestCI lane an automatic exerciser of the dispatch fast arm. The new Q5_0/Q5_1 kernels are plain NEON (no dotprod) and need no dispatch.Verified:
jvmTest+linuxX64Testgreen (refactor regression, FFM + K/N x64).linuxArm64Test -PcrossArm64=truegreen — bit-parity on the unchanged dotprod codegen path.sdotinstructions present, no separate_dpsymbols (inlined back, as designed).