Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
75 changes: 75 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,81 @@

## [Unreleased]

## [0.50.0] - 2026-08-28

Headline: **model size on Android is now a page-cache question, not a heap question — and decode is
2.5× faster.** Quantized weights are served straight from the memory-mapped GGUF file (zero copies,
zero relayout), the packed matmul kernels thread across cores, and a 1.0 GB Qwen2.5-1.5B Q4_K_M —
which OOM'd at load under the 256 MB ART cap in 0.49.0 — loads in ~0.4 s with 566 KB of weight heap
and decodes at 61–66 ms/step with zero steady-state page faults (Pixel 8a, measured by the M2-A5
harness). The same file-order kernels serve heap-staged weights too, closing a measured 48,771 ms/step
silent-fallback trap for mixed-quant models — and fallbacks are never silent again.

### Added

- **Packed-tensor mapped staging** ([#1189](https://github.com/SKaiNET-developers/SKaiNET/issues/1189),
[#1190](https://github.com/SKaiNET-developers/SKaiNET/pull/1190)): under
`WeightForm(residency = MAPPED)` every GGML block format (Q4_K, Q6_K, Q5_K, Q8_0, Q4_0, Q5_0, Q5_1 —
[#1192](https://github.com/SKaiNET-developers/SKaiNET/issues/1192)) is served from file-backed pages:
`BufferPackedTensorData` over `MappedBufferStorage`/`DirectBufferStorage`, row-major (`_rm`) C kernels
that read canonical GGUF file order (no prepack, no relayout copy), reached via JNI direct-buffer
entries on Android and FFM `MemorySegment.ofBuffer` on the JVM
([#1191](https://github.com/SKaiNET-developers/SKaiNET/issues/1191)). Ternary (`BITNET_B1_58`) still
heap-stages — its load-time repack needs the sidecar cache tracked in
[#1198](https://github.com/SKaiNET-developers/SKaiNET/issues/1198).
- **Threaded packed matmuls** ([#1195](https://github.com/SKaiNET-developers/SKaiNET/issues/1195),
[#1196](https://github.com/SKaiNET-developers/SKaiNET/pull/1196)): a spin-then-park worker pool with
guided row-grains, engaged at `outputDim ≥ 512`. The design is measurement-driven and the failed
variants are documented in the PR: per-call `pthread_create` cost 994 ms/step against deep-idle
cores, and every *sleeping* pool lost to one pegged big core because sub-millisecond bursts never
build scheduler utilization — the ~1 ms spin before parking (the same trick llama.cpp uses) is what
unlocks big cores at full clocks. 153 → 61 ms/step on the 1.5B; results are bit-identical to
single-threaded (disjoint row ranges, unchanged per-row accumulation order).
- **Storage-polymorphic row-major dispatch, and no silent fallbacks**
([#1193](https://github.com/SKaiNET-developers/SKaiNET/issues/1193),
[#1197](https://github.com/SKaiNET-developers/SKaiNET/pull/1197)): one kernel per format serves
`BLOCKED_ROW_MAJOR` weights from mapped, direct **or heap** storage — an un-prepacked heap canonical
weight used to fall to the decoding reference kernel silently (measured: 48,771 ms/step on
SmolLM2-135M, whose non-256-multiple dims made llama.cpp's quantizer emit mostly Q8_0; now 33 ms/step,
the fastest configuration measured). `ViewKernel` gained a sink-aware `run` overload and every packed
bridge announces a reference fallback as a `KernelRun` trace event with the reason; the M2-A5 harness
prints the count.
- **The plan tells the truth about mapped weights**
([#1190](https://github.com/SKaiNET-developers/SKaiNET/pull/1190)): `MemoryPlan.budgetedBytes`
charges mapped-servable weights against device RAM/page cache instead of the heap budget
(`fits`, suggestions and `PlannerProfile`'s KV auto-quantization follow), rendered as its own
`mapped (page cache, evictable — not heap)` line. `AllocationResolver.servesFromMapping` is the one
predicate the resolver, the plan and the loader share, gated by
`StorageCapabilities.mappedServableEncodings`, so they cannot tell different stories; `planInput`
takes the `WeightForm` the load will use.
- **Kernel-support matrix: mapped serving section** — the generated matrix now shows, per platform,
which formats serve from a mapping (all seven on Android `native-jni-direct` and JVM `ffm-rowmajor`;
empty cells are the documented gaps).
- **"The DSL is compute"** ([#1194](https://github.com/SKaiNET-developers/SKaiNET/pull/1194)):
the architecture principle behind all of the above as a teachable explanation page — network
definitions describe computation only; memory intent lives in `WeightForm` at the load boundary,
priced by the plan, honored-or-visibly-rejected by resolvers, with runtime dispatch holding no
memory policy.

### Changed

- Cross-order kernel results (row-major vs feed-order) are **numerically equivalent, not bit-exact**:
under `-O3 -ffast-math` compilers may contract float accumulation differently per loop shape
(measured 2 ULP on clang/arm64, more under MSVC auto-vectorization). Integer-dot formats (Q4_K, Q6_K)
currently match exactly, but only threaded-vs-single-threaded identity is contractual. The parity
suites encode this.
- The M2-A5 measurement harness gained `residency=heap|mapped` and reports mapped vs heap weight bytes
and the packed-kernel fallback count.

### Third-party

- The vendored NeoGPU ternary NEON kernel
(`skainet-backends/skainet-backend-native-cpu/native/src/vendor/neogpu/hs_ml_ternary_neon.c`,
© 2024 NeoGPU Contributors, MIT, byte-identical to upstream
[anjaustin/neogpu](https://github.com/anjaustin/neogpu) @ `0846b24`) ships unchanged in this release;
REUSE metadata and `META-INF/THIRD-PARTY-NOTICES.md` in the published artifacts carry the attribution
([#1166](https://github.com/SKaiNET-developers/SKaiNET/issues/1166)).

## [0.49.0] - 2026-08-26

Headline: **the SKEEP-003 memory & storage architecture, complete — from accepted proposal to shipped system.**
Expand Down
9 changes: 7 additions & 2 deletions docs/modules/ROOT/pages/reference/kernel-support-matrix.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -28,8 +28,13 @@ Kernels that read the weight in canonical row-major GGUF file order straight fro
|===
| Weight format | JVM | Android | Native·linux | Native·apple | JS/WASM

| `Q4_K` | — | native-jni-direct | — | — | —
| `Q6_K` | — | native-jni-direct | — | — | —
| `Q8_0` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q4_0` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q4_K` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q6_K` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q5_K` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q5_1` | ffm-rowmajor | native-jni-direct | — | — | —
| `Q5_0` | ffm-rowmajor | native-jni-direct | — | — | —
|===

See also the eager backends & kernels mindmap (xref:explanation/eager-execution.adoc[]) for the narrative overview and gaps.
Original file line number Diff line number Diff line change
Expand Up @@ -155,7 +155,7 @@ public object KernelDispatch {
output = out.id,
bytesRead = inputs.sumOf { it.elementCount * it.format.dtype.sizeInBytes },
bytesWritten = out.elementCount * out.format.dtype.sizeInBytes,
) { kernel.run(inputs, out) }
) { kernel.run(inputs, out, sink) }
}
}

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,15 @@ public interface ViewKernel {

/** Run the kernel: [inputs] as described by [key], result written into [out]. */
public fun run(inputs: List<TensorView>, out: TensorView)

/**
* Run with a [sink] for events the kernel itself must report — above all an internal
* fallback to the decoding reference (#1193: a silent 1000× slowdown deserves a trace
* line the way an adapter gets one, SKEEP-003 §5.1). Default delegates to [run]; kernels
* with an internal fallback override this and emit before punting.
*/
public fun run(inputs: List<TensorView>, out: TensorView, sink: sk.ainet.lang.memory.trace.TraceSink): Unit =
run(inputs, out)
}

/**
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,10 @@ public class PackedViewMatmulKernel(

override val name: String = "$providerName-$encodingName"

override fun run(inputs: List<TensorView>, out: TensorView) {
override fun run(inputs: List<TensorView>, out: TensorView): Unit =
run(inputs, out, sk.ainet.lang.memory.trace.NoopTraceSink)

override fun run(inputs: List<TensorView>, out: TensorView, sink: sk.ainet.lang.memory.trace.TraceSink) {
require(inputs.size == 2) { "matmul takes two operands" }
val a = inputs[0]
val w = inputs[1]
Expand All @@ -48,15 +51,15 @@ public class PackedViewMatmulKernel(
"should have prepacked it (#973)"
}

val aHeap = a.storage as? Storage.Heap ?: return fallback(inputs, out)
val wHeap = w.storage as? Storage.Heap ?: return fallback(inputs, out)
val oHeap = out.storage as? Storage.Heap ?: return fallback(inputs, out)
val activation = aHeap.floats ?: return fallback(inputs, out)
val weight = wHeap.bytes ?: return fallback(inputs, out)
val output = oHeap.floats ?: return fallback(inputs, out)
val aHeap = a.storage as? Storage.Heap ?: return fallback(inputs, out, sink, "activation storage ${a.storage::class.simpleName}")
val wHeap = w.storage as? Storage.Heap ?: return fallback(inputs, out, sink, "weight storage ${w.storage::class.simpleName} — feed-order kernels take heap bytes")
val oHeap = out.storage as? Storage.Heap ?: return fallback(inputs, out, sink, "output storage ${out.storage::class.simpleName}")
val activation = aHeap.floats ?: return fallback(inputs, out, sink, "activation is not a FloatArray")
val weight = wHeap.bytes ?: return fallback(inputs, out, sink, "weight is not a ByteArray")
val output = oHeap.floats ?: return fallback(inputs, out, sink, "output is not a FloatArray")
// The SPI takes a contiguous activation row; a strided one would be mis-indexed, so it goes
// to the reference kernel rather than silently reading the wrong floats.
if (!a.isContiguous) return fallback(inputs, out)
if (!a.isContiguous) return fallback(inputs, out, sink, "strided activation")

val weightOffset = wHeap.arrayOffset + (w.layout.offsetElements * w.layout.elementBytes).toInt()
for (r in 0 until rows) {
Expand All @@ -69,7 +72,23 @@ public class PackedViewMatmulKernel(
}
}

private fun fallback(inputs: List<TensorView>, out: TensorView) {
private fun fallback(
inputs: List<TensorView>,
out: TensorView,
sink: sk.ainet.lang.memory.trace.TraceSink,
reason: String,
) {
// #1193: the decoding reference is ~1000× slower — a trace line, never silent.
if (sink.isEnabled) {
sink.emit(
sk.ainet.lang.memory.trace.TraceEvent.KernelRun(
op = key.op,
kernel = "reference-fallback from $name: $reason",
inputs = inputs.map { it.id },
output = out.id,
),
)
}
ReferenceMatmulKernel(key).run(inputs, out)
}

Expand Down
176 changes: 176 additions & 0 deletions skainet-backends/skainet-backend-jni-cpu/native/skainet_jni.c
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,182 @@ Java_sk_ainet_exec_kernel_jni_JniKernels_q6kMatmulRmDirect(
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q80MatmulRmDirect(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jobject weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_RM_DIRECT_BODY(
skainet_q8_0_matmul_rm(in, inputOffset, w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

/*
* Heap-array row-major matmuls (#1193): the same _rm kernels over a pinned
* ByteArray weight in canonical file order. This is what serves a HEAP-staged
* canonical weight without a prepack — before these existed, a heap tensor
* whose format was not mapped-servable fell to the decoding reference kernel
* silently (measured 48,771 ms/step vs 65 on a mixed-quant model).
*/
JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q4kMatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q4k_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q6kMatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q6k_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q80MatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q8_0_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}


JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q40MatmulRmDirect(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jobject weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_RM_DIRECT_BODY(
skainet_q4_0_matmul_rm(in, inputOffset, w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q40MatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q4_0_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q50MatmulRmDirect(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jobject weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_RM_DIRECT_BODY(
skainet_q5_0_matmul_rm(in, inputOffset, w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q50MatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q5_0_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q51MatmulRmDirect(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jobject weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_RM_DIRECT_BODY(
skainet_q5_1_matmul_rm(in, inputOffset, w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q51MatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q5_1_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q5kMatmulRmDirect(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jobject weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_RM_DIRECT_BODY(
skainet_q5k_matmul_rm(in, inputOffset, w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

JNIEXPORT void JNICALL
Java_sk_ainet_exec_kernel_jni_JniKernels_q5kMatmulRm(
JNIEnv* env, jobject thiz,
jfloatArray input, jint inputOffset,
jbyteArray weight, jint weightByteOffset,
jint inputDim, jint outputDim,
jfloatArray output, jint outputOffset
) {
(void) thiz;
SKAINET_JNI_MATMUL_BODY(
skainet_q5k_matmul_rm(in, inputOffset, (const uint8_t*) w, weightByteOffset,
inputDim, outputDim, out, outputOffset))
}

/*
* bitnet_gemv (SKEEP-003 §5.3, #1041): int8 activations against ternary TQ2_0
* weights. Its activation is a *byte* array, not floats, so it does not fit
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -258,6 +258,12 @@ class M2A5DeviceMeasurement {
line()
line("- steady-state (after $warmup warm-up steps): ${steady.map { it.ms }.average().toInt()} ms/step, " +
"major faults total ${steady.mapNotNull { it.majDelta }.sum()}")
// #1193: reference fallbacks are trace events now — a nonzero count here is the
// 48 s/step trap announcing itself instead of hiding in the timings.
val fallbacks = sink.events().filterIsInstance<TraceEvent.KernelRun>()
.filter { "reference-fallback" in it.kernel }
line("- packed-kernel reference fallbacks: ${fallbacks.size}" +
if (fallbacks.isEmpty()) "" else " — e.g. ${fallbacks.first().kernel}")
line("- forward-scope allocations after warm-up should be 0; " +
"whole-run forward allocations: ${actual.allocationsByScope[ScopeKind.FORWARD] ?: 0}")
}
Expand Down
Loading
Loading