diff --git a/docs/eager-execution-backends-and-kernels.md b/docs/eager-execution-backends-and-kernels.md index 6b879524e..105b15bdd 100644 --- a/docs/eager-execution-backends-and-kernels.md +++ b/docs/eager-execution-backends-and-kernels.md @@ -17,62 +17,78 @@ mindmap BF16 ✅ Q8_0 ✅ Q4_0 ✅ - Q4_K ✅ new - Q6_K ✅ new - Q5_1 ✅ new - Q5_0 ✅ new + Q4_K ✅ + Q6_K ✅ + Q5_1 ✅ + Q5_0 ✅ Panama Vector ✅ - JVM SIMD — jdk.incubator.vector + JVM SIMD only — jdk.incubator.vector, NOT available on Android/ART FP32 BF16 Q8_0 Q4_0 ✅ - Q4_K ✅ - Q5_1 Q5_0 ✅ new - Q6_K ✅ new + Q4_K Q6_K ✅ + Q5_1 Q5_0 ✅ Native FFM ✅ - JVM only — C kernels via CMake + JVM only — C kernels via CMake, ART has no java.lang.foreign FP32 BF16 Q8_0 Q4_0 Q4_K ✅ Q4_K MemSeg zero-copy ✅ - Q5_1 Q5_0 Q6_K ❌ + Q5_1 Q5_0 ✅ new + Q6_K ❌ + Native JNI ✅ new + Android only — same C kernels as Native FFM, reached via JNI + two .so tiers, cpuinfo-gated armv8-a vs armv8.2+dotprod + Q8_0 Q4_0 Q4_K Q5_K Q6_K Q5_1 Q5_0 ✅ + dense FP32 ❌ issue 920 Apple Accelerate ✅ Native macOS iOS — cinterop dense FP32 matmul ✅ elementwise reductions ✅ packed quant via scalar + Native cinterop ✅ new + Kotlin/Native — linux + Apple, static archive in the klib + Q8_0 Q4_0 Q4_K Q5_K Q6_K Q5_1 Q5_0 ✅ + Apple: runtime FEAT_DotProd dispatch, one archive A12 through M-series + manual registration — installNativeKernels(), no ServiceLoader on K/N Platforms - JVM ✅ scalar + Panama + FFM - Native linux ✅ scalar only - Native apple ✅ scalar + Accelerate + JVM ✅ scalar + Panama + Native-FFM + Android ✅ scalar + Native-JNI — NOT Panama, NOT FFM + Native linux ✅ scalar + Native-cinterop + Native apple ✅ scalar + Accelerate + Native-cinterop JS and WASM ✅ scalar only Gaps and roadmap - Native FFM Q5 and Q6_K ❌ issue 708 - Native SIMD on linux ❌ issue 722 - Q5_K Q2_K Q3_K IQ4 packed ❌ dequant only + Native FFM Q6_K ❌ + Native JNI dense FP32 ❌ issue 920 + Q5_K Q2_K Q3_K IQ4 packed on non-K/N targets ❌ dequant only GPU backends IREE Metal ❌ future ``` ## Kernel × provider (matmul, FP32 activations) -| Weight format | Scalar (all targets) | Panama Vector (JVM SIMD) | Native FFM (JVM) | -|---|:--:|:--:|:--:| -| FP32 | ✅ | ✅ | ✅ | -| BF16 | ✅ | ✅ | ✅ | -| Q8_0 | ✅ | ✅ | ✅ | -| Q4_0 | ✅ | ✅ | ✅ | -| Q4_K | ✅ | ✅ | ✅ | -| Q6_K | ✅ | ✅ | ❌ | -| Q5_1 | ✅ | ✅ | ❌ | -| Q5_0 | ✅ | ✅ | ❌ | -| Q5_K / Q2_K / Q3_K / Q8_K / IQ4 | ❌ (dequant-to-FP32 only) | ❌ | ❌ | +| Weight format | Scalar (all targets) | Panama Vector (JVM only) | Native FFM (JVM) | Native JNI (Android) | Native cinterop (K/N) | +|---|:--:|:--:|:--:|:--:|:--:| +| FP32 | ✅ | ✅ | ✅ | ❌ (#920) | ❌ | +| BF16 | ✅ | ✅ | ✅ | ❌ (#920) | ❌ | +| Q8_0 | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q4_0 | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q4_K | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q6_K | ✅ | ✅ | ❌ | ✅ | ✅ | +| Q5_K | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q5_1 | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q5_0 | ✅ | ✅ | ✅ | ✅ | ✅ | +| Q2_K / Q3_K / Q8_K / IQ4 | ❌ (dequant-to-FP32 only) | ❌ | ❌ | ❌ | ❌ | -Resolution is by priority: **Native FFM (100) → Panama (50) → Scalar (0)** — the best -*available* provider that carries the kernel wins; otherwise it cascades down. +Resolution is by priority: **Native (100, whichever of FFM/JNI/cinterop applies to the +target) → Panama (50, JVM only) → Scalar (0)** — the best *available* provider that +carries the kernel wins; otherwise it cascades down. At most one native tier is ever +compiled into a given target, so "Native" columns are mutually exclusive per platform, +not stacked. ## Platform × what runs | Target | Providers available | Notes | |---|---|---| -| **JVM / Android(JVM)** | Scalar + Panama + Native-FFM | full SIMD/native acceleration | -| **Kotlin/Native — linux x64/arm64** | Scalar | no SIMD yet (scalar floor) | -| **Kotlin/Native — macOS/iOS** | Scalar + Apple Accelerate | Accelerate accelerates *dense* FP32; packed-quant via scalar | +| **JVM** | Scalar + Panama + Native-FFM | full SIMD/native acceleration | +| **Android** | Scalar + Native-JNI | Panama (`jdk.incubator.vector`) and FFM (`java.lang.foreign`) are both JDK-only — ART has neither, so Android's native tier is JNI, not a degraded JVM | +| **Kotlin/Native — linux x64/arm64** | Scalar + Native-cinterop | static archive embedded in the klib; manual `installNativeKernels()`, no ServiceLoader | +| **Kotlin/Native — macOS/iOS** | Scalar + Apple Accelerate + Native-cinterop | Accelerate accelerates *dense* FP32/reductions; Native-cinterop covers packed quant, with runtime FEAT_DotProd dispatch (one archive serves A12 through M-series) | | **JS / WASM (Js, Wasi)** | Scalar | no SIMD | **Packed-quant matmul now works on every target** (Q4_K/Q6_K/Q5_1/Q5_0 gained a commonMain @@ -81,11 +97,12 @@ those formats were JVM-only and broke on Native. ## In progress / missing (with trackers) -- ❌ **Native FFM Q5_1/Q5_0/Q6_K** — the C kernel set covers FP32/BF16/Q8_0/Q4_0/Q4_K only. Tracked by **SKaiNET#708** (core kernel) and **SKaiNET-transformers#170** (converter wiring). +- ❌ **Native FFM Q6_K** — the only packed format the FFM C kernel set doesn't cover (FP32/BF16/Q8_0/Q4_0/Q4_K/Q5_K/Q5_1/Q5_0 all ship). Q5_1/Q5_0 shipped in 0.39.1, closing the former **SKaiNET#708**. +- ❌ **Native JNI dense FP32/BF16** — the Android JNI provider has no GEMM shim yet; dense ops fall through to scalar on Android regardless of which `.so` tier loaded. Tracked by **SKaiNET#920**. - ✅ **Native packed-quant kernels on Kotlin/Native** — `NativeKnKernelProvider` (priority 100, `skainet-backend-native-cpu`) serves Q8_0/Q4_0/Q4_K/Q5_K/Q6_K/Q5_0/Q5_1 from the C kernels statically embedded in the klib, on linuxX64/linuxArm64 and (since #959) iosArm64/iosSimulatorArm64/macosArm64. Apple archives use runtime FEAT_DotProd dispatch (#958) so one device archive serves A12 through M-series. Registration is **manual** — call `installNativeKernels()` once at startup (no ServiceLoader on K/N). - ❌ **Dense FP32/BF16 SIMD on Kotlin/Native linux** — the dense floats still run the scalar floor there (Apple has Accelerate). Tracked by **SKaiNET#722** / **#910**. -- ❌ **Other GGML quant formats** (Q5_K, Q2_K, Q3_K, Q8_K, IQ4_NL/XS) — loadable via dequant-to-FP32, but no packed matmul kernel. -- ❌ **Non-CPU eager backends** (IREE, Metal, GPU) — the `KernelProvider` SPI anticipates them, but none are implemented for the eager path today. +- ❌ **Other GGML quant formats** (Q2_K, Q3_K, Q8_K, IQ4_NL/XS) — loadable via dequant-to-FP32, but no packed matmul kernel on any provider. +- ❌ **Non-CPU eager backends** (IREE, Metal, GPU) — the `KernelProvider` SPI anticipates them, but none are implemented for the eager path today. The *compiled* path (DSL → StableHLO → IREE) does reach GPU on Android via Vulkan — see `SKaiNET-transformers`' `llm-runtime/iree-android` — but that's a separate pipeline from this eager-execution mindmap. > This mindmap is a hand-authored overview. Its companion > [kernel × platform support matrix](modules/ROOT/pages/reference/kernel-support-matrix.adoc) is diff --git a/docs/modules/ROOT/nav.adoc b/docs/modules/ROOT/nav.adoc index e6a714cd5..e1a98751e 100644 --- a/docs/modules/ROOT/nav.adoc +++ b/docs/modules/ROOT/nav.adoc @@ -38,6 +38,7 @@ ** xref:explanation/perf/jvm-cpu.adoc[JVM CPU performance] ** xref:explanation/perf/simd-kernels.adoc[How SIMD kernels are built] ** xref:explanation/perf/quantized-simd-kernels.adoc[How quantized SIMD kernels are built] +** xref:explanation/perf/android-neon-jni-kernels.adoc[Android NEON kernels via JNI] ** xref:explanation/perf/turboquant-kv-compression.adoc[TurboQuant KV-cache compression] ** xref:explanation/perf/java-25-cpu-backend.adoc[Java 25 CPU backend notes] ** xref:explanation/issues/native-macos-accelerate-simd.adoc[Native macOS Accelerate SIMD issues] diff --git a/docs/modules/ROOT/pages/explanation/perf/android-neon-jni-kernels.adoc b/docs/modules/ROOT/pages/explanation/perf/android-neon-jni-kernels.adoc new file mode 100644 index 000000000..f41220034 --- /dev/null +++ b/docs/modules/ROOT/pages/explanation/perf/android-neon-jni-kernels.adoc @@ -0,0 +1,238 @@ += Android NEON Kernels via JNI +:description: How SKaiNET reaches hand-written ARM NEON SIMD on Android, where the JVM's FFM native provider cannot run. + +This page explains *how* SKaiNET's eager CPU backend reaches native SIMD +throughput on Android — a genuinely different mechanism from the JVM's +FFM path, because Android's runtime (ART) rules FFM out entirely. If you +just want the fastest available kernel to run automatically, you don't +need to read this — installing `skainet-backend-jni-cpu` on the Android +classpath is enough; discovery is automatic. This page is for the +engineer who wants to understand or extend the kernel layer, or who is +debugging why a given device is (or isn't) hitting the native path. + +== Why Android needs its own provider + +xref:explanation/perf/simd-kernels.adoc[The SIMD kernels page] and the +xref:reference/architecture.adoc[architecture reference] describe the +JVM's `NativeKernelProvider`: priority 100, backed by `java.lang.foreign` +(FFM), near-zero call overhead, no global lock. That provider cannot run +on Android — ART does not implement `java.lang.foreign` at all, at any +API level. Without a native provider, Android falls back all the way to +priority-0 scalar Kotlin, which is where the practical decode-speed +problem actually starts: primitive-loop overhead aside, a scalar matmul +loop on ART is simply not competitive with hand-tuned NEON. + +`skainet-backends/skainet-backend-jni-cpu` is the fix: the *same* shared +C matmul kernels the FFM provider calls +(`skainet-backend-native-cpu/native/`), reached through JNI instead of +FFM, packaged as an AAR any Android app can add as a dependency. It +registers as `JniKernelProvider`, priority 100 — same priority as the +JVM's FFM provider, because on Android it plays the identical role: the +best kernel available, falling back to Panama... except Panama isn't +available either (no JDK Vector API on ART), so in practice the cascade +on Android is JNI (100) → scalar (0), with nothing in between. + +[NOTE] +.This is not a rejection of the FFM decision +==== +xref:reference/architecture.adoc[Architecture §9] records "FFM (not JNI) +for any future native code" as a decision, made when the JVM native +provider was designed, with the rationale "JNI's per-call overhead and +global lock are wrong for hot per-token kernels." That's still correct +*for the JVM*, where FFM is available and strictly better. It doesn't +apply to Android, where FFM isn't an option at all — JNI is not a +second-best alternative there, it's the only native path ART offers. +The two decisions coexist: FFM where you can have it, JNI where you can't. +==== + +== Two `.so` tiers, selected once at load time + +Unlike the JVM FFM provider (one native library, host architecture only), +the Android JNI provider ships **two** shared libraries built from the +same C sources, gated on a real hardware risk: + +[cols="1,2,3",options="header"] +|=== +| Library | Compiled with | Runs on +| `libskainet_jni.so` (`BASELINE`) | plain `armv8-a` | Every 64-bit ARM core. NEON is architecturally guaranteed on AArch64, so FP32/Q8_0/Q4_0/Q5_K get NEON bodies here; Q4_K/Q6_K fall back to scalar (their SIMD bodies need the dot-product extension, see below). +| `libskainet_jni_v82.so` (`V82_DOTPROD`) | `-march=armv8.2-a+fp16+dotprod` | Only cores that report `asimddp` (+ `asimdhp`/`fphp`) in `/proc/cpuinfo`. Enables the `vdotq_s32` paths in the Q4_K/Q6_K kernels. +|=== + +Executing the dotprod library on an armv8.0 core (Cortex-A53, early A55) +would `SIGILL` — the instruction genuinely doesn't exist on that +silicon. So the choice has to be made *before* the library loads, not +inside it. `JniKernels.loadVariant()` reads `/proc/cpuinfo` first — the +NDK-sanctioned detection path, and one that needs no JNI itself, which +matters precisely because it has to run before any native call is +possible — and loads exactly one variant per process: + +[source,kotlin] +---- +private fun cpuSupportsV82(): Boolean = runCatching { + val features = File("/proc/cpuinfo").useLines { lines -> + lines.firstOrNull { it.startsWith("Features") } + } ?: return false + "asimddp" in features && ("asimdhp" in features || "fphp" in features) +}.getOrDefault(false) + +private fun loadVariant(): Variant? { + if (cpuSupportsV82()) { + try { + System.loadLibrary(Variant.V82_DOTPROD.libName) + return Variant.V82_DOTPROD + } catch (_: Throwable) { + // Fall through to baseline — e.g. a packaging that stripped the v82 lib. + } + } + return try { + System.loadLibrary(Variant.BASELINE.libName) + Variant.BASELINE + } catch (_: Throwable) { + null + } +} +---- + +A `/proc/cpuinfo` read failure, or any load failure, degrades to +baseline or to no native provider at all (the registry then cascades to +scalar) — never a `SIGILL`. Both `.so`s export the *same* JNI symbols +(same class, same method names), so selecting the variant is a load-time +decision, not a per-call or per-symbol one — no runtime dispatch cost on +the hot path. + +Both variants ship in the AAR; on install, Android's own APK splitting +picks the right ABI slice (`arm64-v8a`), and the CPU-feature choice +between baseline/dotprod happens on top of that at process start. + +== JNI, done carefully + +JNI is not free — the concern the FFM decision (above) raises for the +JVM is real, it's just not the binding constraint on Android, where the +alternative isn't "FFM instead" but "no native kernel at all." The +implementation still keeps per-call overhead as low as JNI allows: + +* **`GetPrimitiveArrayCritical`, not `GetFloatArrayElements`.** On ART, + heap primitive arrays are contiguous, so a critical pin is zero-copy — + no array is duplicated for the call. The rules that make this safe are + followed exactly: no JNI calls between `Get` and `Release`, arrays + released in reverse acquisition order, read-only inputs released with + `JNI_ABORT` (no write-back copy), the output array released with `0` + (write-back + unpin). ++ +[source,c] +---- +#define SKAINET_JNI_MATMUL_BODY(CALL) \ + jfloat* in = (*env)->GetPrimitiveArrayCritical(env, input, NULL); \ + jbyte* w = in ? (*env)->GetPrimitiveArrayCritical(env, weight, NULL) : NULL; \ + jfloat* out = w ? (*env)->GetPrimitiveArrayCritical(env, output, NULL) : NULL; \ + if (out) { CALL; } \ + if (out) (*env)->ReleasePrimitiveArrayCritical(env, output, out, 0); \ + if (w) (*env)->ReleasePrimitiveArrayCritical(env, weight, w, JNI_ABORT); \ + if (in) (*env)->ReleasePrimitiveArrayCritical(env, input, in, JNI_ABORT); +---- ++ +Every JNI entry point (`skainet_jni.c`) is a one-line body built from +this macro — pin, call the shared C kernel, release. There is no logic +in the JNI layer beyond array pinning; the NEON kernels themselves are +the exact same code the FFM provider calls (`skainet-backend-native-cpu`'s +`native/`), so there is one implementation to keep numerically correct, +not two. +* **No underscores in JNI method names.** JNI mangles `_` to `_1` in + native symbol names — a silent mismatch trap if a Kotlin method name + and its `Java_..._methodName` C symbol drift (`q4_0Matmul` would + mangle differently than `q40Matmul`). Every method on `JniKernels` is + named to avoid this by construction (`q80Matmul`, `q40Matmul`, + `q4kMatmul`, ...). +* **Eager, not lazy, library load.** `JniKernels.variant` is a `val` + initialized in `object init`, not a `by lazy` property read on first + use. Kotlin object initialization runs on first access to *any* + member, so a direct call to an `external fun` is guaranteed to find + the library already loaded — a lazy property would only trigger the + load when the property itself was read, and a caller that skipped + straight to `q80Matmul(...)` would hit `UnsatisfiedLinkError`. + +== Availability probe and registration + +`JniKernelProvider.isAvailable()` doesn't just check that a library +loaded — it round-trips a smoke kernel (`output[i] = 2 * input[i]`) +through the real JNI path and checks the numeric result, so a library +that loaded but is somehow broken (corrupted APK, ABI mismatch a +`try`/`catch` didn't catch) still reports itself unavailable rather than +returning wrong answers: + +[source,kotlin] +---- +private val available: Boolean by lazy { + if (!JniKernels.isLoaded) return@lazy false + runCatching { + val input = floatArrayOf(1.0f, 2.5f, -3.0f) + val output = FloatArray(3) + JniKernels.smoke(input, output, 3) + output[0] == 2.0f && output[1] == 5.0f && output[2] == -6.0f + }.getOrDefault(false) +} +---- + +Registration follows the same `ServiceLoader` pattern as every other +kernel provider (xref:explanation/perf/simd-kernels.adoc[SIMD kernels +page, "Auto-discovery"] section) — `JniKernelProviderFactory` (a +no-arg-constructible wrapper, since `ServiceLoader` can't instantiate a +Kotlin `object` directly) is listed in +`META-INF/services/sk.ainet.backend.api.kernel.KernelProvider`, and the +Android CPU-ops factory installs every discovered provider exactly like +the JVM does. An app pulls in the JNI provider by adding the +`skainet-backend-jni-cpu` AAR as a dependency (e.g. via `kllama`'s +`androidMain` `runtimeOnly`) — no explicit registration call needed. + +`matmulFp32()` returns `null` — the JNI provider does not carry a dense +FP32 GEMM kernel yet (tracked as +https://github.com/SKaiNET-developers/SKaiNET/issues/920[#920]); dense +FP32 execution on Android currently falls through to scalar regardless +of which JNI variant loaded. Q8_0, Q4_0, Q4_K, Q5_K, Q6_K, Q5_0, and +Q5_1 all have JNI kernels — see +xref:reference/kernel-support-matrix.adoc[] for the generated, +authoritative per-format table (column `Android`, provider +`native-jni`). + +== Numbers + +Measured on a Pixel 8a, SmolLM2-135M-Instruct Q8_0 decode: +**~24 tok/s** with the JNI NEON provider active, versus **~3.8 tok/s** +scalar — a **6.4×** speedup, the difference between unusable and +clearing the on-device usability bar for a real-time chat UI. This +number predates the primitive-fast-path work described below; with it, +NEON matmul time itself was found to be a *minority* of end-to-end +decode time on that same device — 83% of wall-clock was non-matmul +per-element overhead (index-array allocation, boxed accessors, dtype +dispatch) in the generic eager op paths, fixed separately by making the +hot ops (arithmetic, activations, softmax, reductions, concat, reshape) +run flat primitive loops over the dense buffer instead of the generic +path. Both fixes matter for the same reason: on ART, allocation and +boxing costs that JIT-vanish on a desktop JVM do not vanish, so the +matmul kernel being fast doesn't help if everything around it isn't. + +`skainet-backend-jni-cpu`'s `src/androidTest` includes both a parity +suite (`JniKernelParityTest`, every JNI kernel checked against the +scalar reference on-device) and a throughput benchmark +(`SmolLm2DecodeBenchmark`, real end-to-end decode timing) — run them on +a physical device via `./gradlew :skainet-backends:skainet-backend-jni-cpu:connectedAndroidTest` +to reproduce numbers on your own hardware; emulator CPUs don't reflect +real ARM performance characteristics. + +== Where to look in the code + +[cols="1,2",options="header"] +|=== +| File | What it does +| `skainet-backends/skainet-backend-jni-cpu/native/skainet_jni.c` | Thin JNI shims — array pinning + one call into the shared C kernels, nothing else. +| `skainet-backends/skainet-backend-jni-cpu/native/CMakeLists.txt` | Builds both `.so` tiers from the same C sources as `skainet-backend-native-cpu`. +| `skainet-backends/skainet-backend-jni-cpu/src/main/kotlin/.../jni/JniKernels.kt` | Two-tier loader (`/proc/cpuinfo` gate) + `external fun` declarations. +| `skainet-backends/skainet-backend-jni-cpu/src/main/kotlin/.../jni/JniKernelProvider.kt` | The `KernelProvider` implementation — smoke-tested availability, per-format kernel wiring, `ServiceLoader` factory. +| `skainet-backends/skainet-backend-jni-cpu/src/androidTest/.../JniKernelParityTest.kt` | On-device parity vs. the scalar reference, every supported format. +| `skainet-backends/skainet-backend-jni-cpu/src/androidTest/.../SmolLm2DecodeBenchmark.kt` | Real end-to-end decode throughput on-device — the source of the numbers above. +|=== + +For the compiled-graph alternative to this eager path — running a whole +DSL-authored model through IREE instead of op-by-op — see +`SKaiNET-transformers`' `llm-runtime/iree-android` module and its +"Android getting started" tutorial. diff --git a/docs/modules/ROOT/pages/reference/architecture.adoc b/docs/modules/ROOT/pages/reference/architecture.adoc index c9e06cdfb..be7059da7 100644 --- a/docs/modules/ROOT/pages/reference/architecture.adoc +++ b/docs/modules/ROOT/pages/reference/architecture.adoc @@ -82,6 +82,8 @@ between modes is part of the test contract. | `skainet-lang/skainet-lang-models` | Reference reusable models (Llama, Gemma, Qwen, Whisper) built on the DSL. | `skainet-backends/skainet-backend-api` | Neutral backend SPI — `TensorOps`, `TensorDataFactory`, **kernel SPI** (`KernelProvider`, `Fp32MatmulKernel`, `Q4KMatmulKernel`, `KernelRegistry`). | `skainet-backends/skainet-backend-cpu` | CPU implementation. Eager-execution `DefaultCpuOpsBase` (commonMain) + `DefaultCpuOpsJvm` (jvmMain) with SIMD kernels. +| `skainet-backends/skainet-backend-native-cpu` | Native (FFM) kernel provider — JVM only, ART has no `java.lang.foreign`. +| `skainet-backends/skainet-backend-jni-cpu` | Native (JNI) kernel provider for Android — same shared C kernels as `skainet-backend-native-cpu`, reached through JNI instead of FFM since FFM cannot run on ART. See xref:explanation/perf/android-neon-jni-kernels.adoc[]. | `skainet-backends/skainet-backend-xnnpack` | Optional XNNPACK CPU backend (FP32 matmul / conv2d / pooling) on linuxX64 / linuxArm64 / Android. | `skainet-backends/benchmarks/jvm-cpu-jmh` | JMH harness — `MatmulBench`, `KernelMatmulBench`, `QuantizedMatmulBench`, `ElementwiseAdd1MBench`, `Reductions1MBench`. | `skainet-compile/*` | Tape recording, StableHLO emission, IREE export. @@ -92,73 +94,87 @@ between modes is part of the test contract. Introduced in 0.21.0 (PRs #554, #559, #562). The static structure: -[source] +[mermaid] ---- - commonMain (skainet-backend-api) - ┌──────────────────────────────────────┐ - │ KernelProvider { │ - │ name: String │ - │ priority: Int │ - │ isAvailable(): Boolean │ - │ matmulFp32(): Fp32MatmulKernel? │ - │ matmulQ4K(): Q4KMatmulKernel? │ - │ } │ - │ │ - │ KernelRegistry { │ - │ register(KernelProvider) │ - │ bestAvailable(): KernelProvider? │ - │ find(name): KernelProvider? │ - │ } │ - │ │ - │ Fp32MatmulKernel.matmul(...) │ - │ Q4KMatmulKernel.matmul(...) │ - └──────────────┬───────────────────────┘ - │ implements / extends - ┌────────────────────────────┼────────────────────────────────┐ - │ jvmMain (api) │ commonMain (cpu) │ - │ KernelServiceLoader │ ScalarMatmulKernel (priority 0) │ - │ installAll() │ ScalarKernelProvider │ - └────────────────────────────┴────────────────────────────────┘ - │ - ┌────────────────────────────┼────────────────────────────────┐ - │ jvmMain (cpu) │ - │ PanamaVectorKernelProvider (priority 50) │ - │ FP32 BF16 Q8_0 Q4_0 Q4_K Q6_K Q5_1 Q5_0 (SIMD) │ - │ Scalar/PanamaVectorKernelProviderFactory (no-arg wrappers) │ - │ META-INF/services/...KernelProvider │ - └─────────────────────────────────────────────────────────────┘ - │ - ┌────────────────────────────┼────────────────────────────────┐ - │ jvmMain (skainet-backend-native-cpu) │ - │ NativeKernelProvider (priority 100, FFM/C) │ - │ FP32 BF16 Q8_0 Q4_0 Q4_K (+ Q4_K MemSeg zero-copy) │ - └─────────────────────────────────────────────────────────────┘ +flowchart TD + subgraph common["commonMain (skainet-backend-api)"] + SPI["KernelProvider { name, priority, isAvailable(),
matmulFp32(), matmulQ4K(), ... }
KernelRegistry { register(), bestAvailable(), find() }
Fp32MatmulKernel.matmul(...) / Q4KMatmulKernel.matmul(...)"] + end + + subgraph loader["jvmMain (skainet-backend-api)"] + Loader["KernelServiceLoader.installAll()"] + end + + subgraph scalar["commonMain (skainet-backend-cpu)"] + Scalar["ScalarKernelProvider — priority 0
ScalarMatmulKernel, every format, every target"] + end + + subgraph panama["jvmMain (skainet-backend-cpu)"] + Panama["PanamaVectorKernelProvider — priority 50
FP32 BF16 Q8_0 Q4_0 Q4_K Q6_K Q5_1 Q5_0 SIMD"] + end + + subgraph ffm["jvmMain (skainet-backend-native-cpu)"] + Ffm["NativeKernelProvider — priority 100, FFM/C
FP32 BF16 Q8_0 Q4_0 Q4_K Q5_K Q5_1 Q5_0"] + end + + subgraph jni["androidMain (skainet-backend-jni-cpu)"] + Jni["JniKernelProvider — priority 100, JNI/C
same C kernels as native-cpu; ART has no FFM
Q8_0 Q4_0 Q4_K Q5_K Q6_K Q5_1 Q5_0"] + end + + SPI -.->|implements| Scalar + SPI -.->|implements| Panama + SPI -.->|implements| Ffm + SPI -.->|implements| Jni + Loader -->|ServiceLoader discovers| Panama + Loader -->|ServiceLoader discovers| Ffm + Loader -->|ServiceLoader discovers| Jni ---- -Four live providers ship. The exact, machine-generated coverage of every +Five live providers ship. The exact, machine-generated coverage of every weight format on every KMP target is at xref:reference/kernel-support-matrix.adoc[]; for *how* the kernels are -implemented see xref:explanation/perf/simd-kernels.adoc[] (FP32) and -xref:explanation/perf/quantized-simd-kernels.adoc[] (quantized). Packed-quant -matmul (Q4_K/Q6_K/Q5_1/Q5_0) also has a commonMain *scalar* kernel, so it runs -on Kotlin/Native, JS and WASM — not only the JVM. +implemented see xref:explanation/perf/simd-kernels.adoc[] (FP32), +xref:explanation/perf/quantized-simd-kernels.adoc[] (quantized), and +xref:explanation/perf/android-neon-jni-kernels.adoc[] (Android JNI). +Packed-quant matmul (Q4_K/Q6_K/Q5_1/Q5_0) also has a commonMain *scalar* +kernel, so it runs on Kotlin/Native, JS and WASM — not only the JVM. [NOTE] -.Native (FFM) provider +.Native (FFM) provider — JVM only ==== `NativeKernelProvider` registers at priority 100 so that on JDK 21+ it wins `KernelRegistry.bestAvailable()` over the Panama Vector provider whenever the native library loads, and transparently falls back to Panama (priority 50) or scalar (priority 0) when it doesn't — no code change above the registry. It uses FFM, not JNI (near-zero call overhead, no global lock), ships in the -`skainet-backend-native-cpu` module with C kernels for FP32/BF16/Q8_0/Q4_0/Q4_K +`skainet-backend-native-cpu` module with C kernels for FP32/BF16/Q8_0/Q4_0/Q4_K/Q5_K/Q5_1/Q5_0 (plus a zero-copy `MemorySegment` Q4_K path), and currently builds for the host architecture only (cross-arch builds and Maven classifier JARs are out of -scope). Native FFM kernels for Q5_1/Q5_0/Q6_K are a tracked follow-up -(SKaiNET#708). The kernel SPI this builds on shipped across 0.21.0 +scope). Native FFM kernels for Q5_1/Q5_0 shipped in 0.39.1 +(SKaiNET#708); Q6_K has no FFM kernel yet and still resolves to +`panama-vector` on JVM — see xref:reference/kernel-support-matrix.adoc[]. +The kernel SPI this builds on shipped across 0.21.0 (PRs #554–#565); the in-process native-FFM groundwork landed in 0.22.0 (PR #571). ==== +[NOTE] +.Native (JNI) provider — Android only +==== +`JniKernelProvider` also registers at priority 100 — not because it +competes with the FFM provider (they never coexist: FFM cannot load on +ART, so a process runs one or the other, never both) but because it +plays the same role on the platform where FFM isn't an option. It calls +the *same* shared C kernels as `NativeKernelProvider` through JNI instead +of FFM, ships two `.so` tiers gated on a `/proc/cpuinfo` dot-product +check (baseline `armv8-a` vs. `armv8.2-a+dotprod`, selected once at +library-load time), and covers Q8_0/Q4_0/Q5_K/Q5_1/Q5_0/Q4_K/Q6_K — not +yet dense FP32 (SKaiNET#920). Measured ~6.4× decode speedup over scalar +on a Pixel 8a. Shipped in 0.39.0; see +xref:explanation/perf/android-neon-jni-kernels.adoc[] for the full +mechanism, including why this does *not* contradict the FFM decision +above. +==== + == 6. Runtime view — eager execution The eager pipeline for a single op: @@ -265,7 +281,7 @@ BLAS (`-Dskainet.cpu.blas.enabled=true`). | Kernel SPI parallel to BackendProvider | 2026-04 (PR #554) | Matmul / SDPA are model-agnostic; isolating them lets bench harnesses time the SIMD loop directly and lets a future native provider register without touching the op layer. | `KernelProvider.matmulQ4K()` accessor with `default null` | 2026-04 (PR #562) | Backwards compat for existing providers (Scalar) without forcing every implementation to override. Same pattern will be used for Q6KMatmulKernel / Q4KMemSegMatmulKernel sibling SPIs. | ServiceLoader auto-discovery deferred until 2 providers exist | 2026-04 (PR #559) | Single-provider auto-discovery would have been ceremony for nothing; once Panama landed alongside Scalar, the trigger condition was met. -| FFM (not JNI) for any future native code | roadmap M5 | JNI's per-call overhead and global lock are wrong for hot per-token kernels. +| FFM over JNI for the JVM native provider | roadmap M5 | JNI's per-call overhead and global lock are wrong for hot per-token kernels — where FFM is available. This did not generalize to "JNI is rejected everywhere": ART has no FFM at all, so `skainet-backend-jni-cpu` (0.39.0) uses JNI on Android specifically because there is no FFM alternative there. See xref:explanation/perf/android-neon-jni-kernels.adoc[]. | Antora docs (Diátaxis), not GitHub Wiki | 2025 | Source-controlled, branchable, ranked higher in search than wikis, ships with the repo. |=== diff --git a/docs/modules/ROOT/pages/reference/kernel-support-matrix.adoc b/docs/modules/ROOT/pages/reference/kernel-support-matrix.adoc index 4c4104a19..2ba371495 100644 --- a/docs/modules/ROOT/pages/reference/kernel-support-matrix.adoc +++ b/docs/modules/ROOT/pages/reference/kernel-support-matrix.adoc @@ -9,8 +9,8 @@ Each cell is the best (highest-priority) provider that serves `Float32 × format |=== | Weight format | JVM | Android | Native·linux | Native·apple | JS/WASM -| `Float32` | native-ffm | panama-vector | scalar | scalar | scalar -| `BFloat16` | native-ffm | panama-vector | scalar | scalar | scalar +| `Float32` | native-ffm | scalar | scalar | scalar | scalar +| `BFloat16` | native-ffm | scalar | scalar | scalar | scalar | `Q8_0` | native-ffm | native-jni | native-cinterop | native-cinterop | scalar | `Q4_0` | native-ffm | native-jni | native-cinterop | native-cinterop | scalar | `Q4_K` | native-ffm | native-jni | native-cinterop | native-cinterop | scalar diff --git a/skainet-backends/skainet-backend-native-cpu/src/jvmTest/kotlin/sk/ainet/exec/kernel/KernelSupportMatrixTest.kt b/skainet-backends/skainet-backend-native-cpu/src/jvmTest/kotlin/sk/ainet/exec/kernel/KernelSupportMatrixTest.kt index 5da2052c4..c6ad891a0 100644 --- a/skainet-backends/skainet-backend-native-cpu/src/jvmTest/kotlin/sk/ainet/exec/kernel/KernelSupportMatrixTest.kt +++ b/skainet-backends/skainet-backend-native-cpu/src/jvmTest/kotlin/sk/ainet/exec/kernel/KernelSupportMatrixTest.kt @@ -31,13 +31,16 @@ class KernelSupportMatrixTest { private fun scalarFormats(): Set = formats.filter { ScalarKernelProvider.supports("matmul", listOf("Float32", it)) }.toSet() - // Source-set -> platforms. commonMain reaches all; backend-cpu jvmMain -> {JVM,Android}; + // Source-set -> platforms. commonMain reaches all; backend-cpu jvmMain -> {JVM} only: + // Panama Vector (jdk.incubator.vector) is a JDK-only incubator module — ART has no + // Vector API, so PlatformCpuOpsFactory.android (skainet-backend-cpu/androidMain) + // registers ONLY ServiceLoader-discovered providers + the scalar floor, never Panama. // backend-native-cpu jvmMain -> {JVM} (the native module declares only jvm()). // native-jni: skainet-backend-jni-cpu AAR — same C kernels via JNI, Android // only, discovered via ServiceLoader from PlatformCpuOpsFactory.android (#920). private fun tiers(): List = listOf( Tier("scalar", 0, platforms.toSet(), scalarFormats()), - Tier("panama-vector", 50, setOf("JVM", "Android"), + Tier("panama-vector", 50, setOf("JVM"), setOf("Float32", "BFloat16", "Q8_0", "Q4_0", "Q4_K", "Q6_K", "Q5_K", "Q5_1", "Q5_0")), Tier("native-ffm", 100, setOf("JVM"), setOf("Float32", "BFloat16", "Q8_0", "Q4_0", "Q4_K", "Q5_K", "Q5_1", "Q5_0")),