Skip to content

Commit dc88de1

Browse files
Merge pull request #1185 from SKaiNET-developers/release/0.49.0
Release 0.49.0 — the SKEEP-003 architecture, complete
2 parents 6050dff + dff4786 commit dc88de1

27 files changed

Lines changed: 372 additions & 85 deletions

‎CHANGELOG.md‎

Lines changed: 123 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -2,18 +2,131 @@
22

33
## [Unreleased]
44

5-
### Docs
5+
## [0.49.0] - 2026-08-26
6+
7+
Headline: **the SKEEP-003 memory & storage architecture, complete — from accepted proposal to shipped system.**
8+
One storage model (`Storage` / `Scope` / `Format` / `Layout` / `TensorView`), resolver-owned decisions
9+
(what form a weight takes, where its bytes live — never decided by the model author), scope-recycled eager
10+
execution (flat-memory decode), and a compile lane that carries what the runtime decides into the exported
11+
MLIR and `.irpa`. The version jump (0.40.1 → 0.49.0, ~100 merged PRs) is deliberate: this is the release
12+
downstream repositories (SKaiNET-transformers, the IREE conformity pipeline) should build on, and it
13+
removes every façade the architecture replaced. See **Breaking changes** below for the migration map.
14+
15+
### Breaking changes
16+
17+
- **The three legacy loader axes are gone** ([#1159](https://github.com/SKaiNET-developers/SKaiNET/issues/1159)):
18+
`QuantPolicy`, `StagingPolicy` and `WeightOrientation` are deleted. The loader takes one
19+
`WeightForm(encoding, order, shape, residency)` (uniform `weightForm` or per-tensor `weightFormFor`);
20+
migration: `DEQUANTIZE_TO_FP32` → `EncodingRequest.DequantizeTo(FP32)`, `StagingPolicy.MAPPED` →
21+
`WeightForm(residency = WeightResidency.MAPPED)`, `WeightOrientation.OUT_IN` → `WeightShapeOrientation.OUT_IN`.
22+
`AndroidGguf.loader` takes a `WeightForm` (default mapped residency).
23+
- **The dead placement machinery is gone** ([#1142](https://github.com/SKaiNET-developers/SKaiNET/issues/1142)):
24+
`@Place`/`@Weights` (declared, retained, read by nothing), `sk.ainet.lang.tensor.storage.MemoryPlanner`,
25+
`StorageSpec`, and `Placement.residency`/`Residency`. Lifetime is `ScopeKind`; weight staging is
26+
`WeightForm.WeightResidency`; placement is decided by `AllocationResolver` (see Added). `Placement`
27+
itself stays (KV-cache stores carry it); `@KvCache`/`@KvCacheBypass` moved to `KvCacheAnnotations.kt`.
28+
- **The `skainet.tensor_encodings` module attribute is gone**
29+
([#1179](https://github.com/SKaiNET-developers/SKaiNET/issues/1179)): replaced by the machine-readable
30+
`skainet.tensor_layouts` (`{kind, block_elems, block_bytes, bits, block_order}`); nothing outside this
31+
repository read the old names dictionary.
32+
- `LogicalDType` is deprecated end to end in favour of `DType`
33+
([#1014](https://github.com/SKaiNET-developers/SKaiNET/issues/1014)); removal at the next major.
34+
35+
### Added
636

7-
- **SKEEP-003 accepted — memory & storage architecture design record and roadmap**
8-
([#932](https://github.com/SKaiNET-developers/SKaiNET/issues/932)).
9-
`docs/modules/skeep/pages/003-unified-tensor-storage.adoc` moves to *Accepted* and records the
10-
thirteen design decisions (storage-first end-state delivered incrementally: `Storage` / `TensorView` /
11-
`Tensor`, scopes, `Format(dtype, encoding)`, `TensorId`, kernel dispatch on declared formats, 2 GB
12-
planner profile, `LogicalDType` → `DType` merge). The full proposal and the M0/M1/M2 milestone PRD are
13-
committed under `docs/design/memory/`; the work is tracked as milestone issues
14-
[#1001](https://github.com/SKaiNET-developers/SKaiNET/issues/1001),
37+
- **The memory model (SKEEP-003 M0 "know before you load" / M1 "flat decode" / M2 "1.58-bit on a 2 GB board")**
38+
([#1001](https://github.com/SKaiNET-developers/SKaiNET/issues/1001),
1539
[#1002](https://github.com/SKaiNET-developers/SKaiNET/issues/1002),
16-
[#1003](https://github.com/SKaiNET-developers/SKaiNET/issues/1003) with one sub-issue per feature branch.
40+
[#1003](https://github.com/SKaiNET-developers/SKaiNET/issues/1003); slices #1004–#1042):
41+
`Storage` (Heap/OffHeap/Mapped, ownership + liveness — a use-after-free is a loud
42+
`StorageClosedException`), `Scope` (`ModelScope`/`ForwardScope` slab with `reset()`),
43+
`Format(dtype, encoding)`, `Layout` (strides/offset/block geometry), `TensorView` with `prepack()` as
44+
the visible relayout and `materialize()` as the single copy point; `TraceSink` events (allocations,
45+
scope resets, adapter insertions); `MemoryPlan`/`MemoryPlans` header-only planning with budgets,
46+
suggestions and a plan-vs-actual check; the `skainet-plan` CLI; `PlannerProfile` (`MOBILE_2GB` with
47+
automatic KV quantization — and `strict`, so a missing kernel refuses instead of silently costing
48+
several times the weight, [#1128](https://github.com/SKaiNET-developers/SKaiNET/pull/1128));
49+
Android mmap loading + device fit checks; KV cache preallocated in model scope with declared formats;
50+
kernel dispatch on declared formats (`KernelKey`, registry-backed capabilities); the decode harness
51+
and M2 acceptance runs.
52+
- **Weight forms — one resolved decision instead of three caller flags**
53+
([#1109](https://github.com/SKaiNET-developers/SKaiNET/issues/1109) arc, #1114–#1120):
54+
`WeightForm` (encoding × byte order × shape × residency) resolved by `WeightFormResolver` from
55+
*what the file holds × the profile × what the backend's kernels can feed*; the loader honours it,
56+
the plan prices it (a resolved dequantization shows in the table, not at the OOM), conversions are
57+
traced, and packed weights can load directly in kernel-feed order
58+
([#1120](https://github.com/SKaiNET-developers/SKaiNET/issues/1120)).
59+
- **Placement is resolver-owned** ([#1133](https://github.com/SKaiNET-developers/SKaiNET/issues/1133) →
60+
#1142–#1144): `AllocationResolver.resolve(weight, profile, platform)` decides memory domain and scope
61+
(mapping requires: the form asks, the platform can, the bytes are the file's bytes);
62+
`AllocationResolver.explain()` renders every decision with its reason pre-load; `ResolvedGguf` wires
63+
plan → load with the documented user-wins precedence (per-tensor `weightFormFor` > uniform
64+
`weightForm` > resolver).
65+
- **Scope-recycled eager execution** ([#1135](https://github.com/SKaiNET-developers/SKaiNET/issues/1135) →
66+
#1145/#1146/#1173): `ExecutionContext.memoryScope` is consulted by tensor creation *and* op outputs
67+
(`TensorDataFactory.adoptFloatArray`, `ScopedTensorDataFactory`), so
68+
`ctx.forwardScope(slabFloats) { … }` gives steady-state decode that allocates zero new slab bytes per
69+
step; the FP32 fast paths and the JVM Panama vector kernels are offset-aware, so slab-backed tensors
70+
keep SIMD speed.
71+
- **Model-footprint analysis for GGUF, safetensors and ONNX**
72+
([#1169](https://github.com/SKaiNET-developers/SKaiNET/issues/1169)): header-only `planInput` for all
73+
three formats (ONNX `external_data` sidecars priced correctly — multi-GB models no longer report ~0
74+
bytes; sizes `Long`-safe), and `PlannerProfile.EDGE` for embedded devices where the budget *is* the
75+
usable RAM. "Will it fit in ~2.1 GB?" is answered in seconds, without reading a tensor payload.
76+
- **The compile lane carries what the runtime decides**
77+
([#1147](https://github.com/SKaiNET-developers/SKaiNET/issues/1147) → #1178/#1179/#1180):
78+
`TensorRef` carries tensor identity (`TraceSession.identify`, registered from
79+
`trainableParameters()`), the encoding *object* (block size intact) and packed block order across the
80+
trace→graph boundary; the emitted module header declares structural facts per tensor
81+
(`skainet.tensor_layouts`); `ExternalParameterRef` declares block order to the `.irpa` consumer;
82+
`HloGenerator.generate(target = …)` runs the optimizer pipeline on the production path with
83+
`LayoutAssignmentPass` (rank-2 packed weights get kernel-feed order; the tape's carried facts are
84+
never overridden), and the `ResolvedComputeGraph` seams surface exactly the decisions made.
85+
- **BitNet / ternary compute track** (#1033, #1040/#1041, #1136–#1141, #1150):
86+
ternary encodings (`TQ1_0`/`TQ2_0`, `BITNET_B1_58`, `BITNET_PLANES` multi-plane packing), i2s GGUF
87+
import, the vendored NeoGPU ternary f32 NEON kernel (MIT, verbatim) exposed through FFM, JNI and
88+
Kotlin/Native including a fused lm_head kernel, requant adapters, NEON BitNet packing, and ternary
89+
benchmarks + getting-started docs.
90+
- **Iris dataset provider** ([#1044](https://github.com/SKaiNET-developers/SKaiNET/issues/1044),
91+
[#1101](https://github.com/SKaiNET-developers/SKaiNET/pull/1101), contributed by @AjithGoveas):
92+
the embedded 150-row dataset used by the new Android classifier tutorial.
93+
- Sliding-window KV SDPA ([#1036](https://github.com/SKaiNET-developers/SKaiNET/issues/1036)) and
94+
KV formats declared by the store ([#1077](https://github.com/SKaiNET-developers/SKaiNET/issues/1077)).
95+
96+
### Fixed
97+
98+
- Packed block order end to end: feed-order bytes in a type claiming canonical order decoded to
99+
plausible garbage ([#1124](https://github.com/SKaiNET-developers/SKaiNET/issues/1124),
100+
[#1126](https://github.com/SKaiNET-developers/SKaiNET/pull/1126)); `TensorData` now declares its
101+
`BlockOrder` and every reader agrees on the same bytes.
102+
- Packed ternary weights reach dispatch through the Wᵀ marker — ANY packed block storage routes through
103+
the marker path ([#1136](https://github.com/SKaiNET-developers/SKaiNET/issues/1136),
104+
[#1181](https://github.com/SKaiNET-developers/SKaiNET/pull/1181)).
105+
- M2 acceptance page-fault flake ([#1107](https://github.com/SKaiNET-developers/SKaiNET/pull/1107));
106+
safetensors dtype mapper no longer prints a WARNING into stdout mid-parse (#1169).
107+
108+
### CI & process
109+
110+
- `apiCheck` has its own named PR leg (`test (api-compatibility)`) instead of hiding inside
111+
golden-parity ([#1176](https://github.com/SKaiNET-developers/SKaiNET/pull/1176)), after a stale dump
112+
reached `develop` unnoticed ([#1174](https://github.com/SKaiNET-developers/SKaiNET/pull/1174));
113+
branch protection on `develop` now requires the full `build-job` aggregator, admins included.
114+
- Android per-target API dumps dropped (jvm + klib only,
115+
[#1111](https://github.com/SKaiNET-developers/SKaiNET/pull/1111)).
116+
117+
### Docs
118+
119+
- The memory model as built: `explanation/memory-model.adoc`, `explanation/packed-weight-layout.adoc`
120+
([#1106](https://github.com/SKaiNET-developers/SKaiNET/pull/1106) — design drafts under
121+
`docs/design/` retired in favour of Antora pages), `explanation/virtual-tensors.adoc` (the ML
122+
Drift-style split, with diagrams), and SKEEP-003a (`skeep/003a-placement-and-planning-resolution.adoc`)
123+
recording the P7/P8 resolutions.
124+
- Tutorials whose code cannot rot: the Android classifier getting-started and the ternary
125+
getting-started, with snippets compiled *and executed* in CI by `skainet-docs-samples` (the Iris
126+
training loop asserts held-out accuracy ≥ 0.80).
127+
- SKEEP-003 accepted ([#932](https://github.com/SKaiNET-developers/SKaiNET/issues/932)) with its
128+
thirteen design decisions; how-to: plan a model's memory before loading it, including the
129+
embedded-device (`edge`) verdict.
17130

18131
## [0.40.1] - 2026-08-12
19132

‎README.md‎

Lines changed: 39 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -51,7 +51,7 @@ Add the core dependencies (Gradle Kotlin DSL):
5151
```kotlin
5252
dependencies {
5353
// Recommended: import the umbrella BOM and drop versions on the engine modules.
54-
implementation(platform("sk.ainet:skainet-bom:0.40.1"))
54+
implementation(platform("sk.ainet:skainet-bom:0.49.0"))
5555

5656
implementation("sk.ainet.core:skainet-lang-core")
5757
implementation("sk.ainet.core:skainet-backend-cpu")
@@ -296,7 +296,39 @@ val withoutLabel = dataPipeline<RawDataset>()
296296

297297
---
298298

299-
## What's New in 0.40.1
299+
## What's New in 0.49.0
300+
301+
The **SKEEP-jump release**: 0.40.1 → 0.49.0, ~100 merged PRs — the SKEEP-003 memory & storage
302+
architecture complete, from accepted proposal to shipped system. This is the release downstream
303+
repositories (SKaiNET-transformers, the IREE conformity pipeline) should build on. It **removes
304+
every façade the architecture replaced** — see the Breaking-changes section of
305+
[CHANGELOG.md](CHANGELOG.md) for the migration map (`QuantPolicy`/`StagingPolicy`/`WeightOrientation`
306+
→ `WeightForm`; `@Place`/`@Weights`/`StorageSpec`/old `MemoryPlanner` → `AllocationResolver`).
307+
308+
- **One storage model** — `Storage` / `Scope` / `Format` / `Layout` / `TensorView`: enforced ownership
309+
(use-after-free throws, loudly), scoped lifetimes, `prepack()` as the visible relayout,
310+
`materialize()` as the single copy point.
311+
- **Decisions are resolved, not declared** — `WeightFormResolver` picks a weight's in-memory form from
312+
*file × profile × kernels*; `AllocationResolver` picks domain and scope, and `explain()` says why,
313+
per tensor, before a byte of payload is read. You always outrank the resolver
314+
(per-tensor `weightFormFor` > uniform `weightForm` > resolver).
315+
- **Flat-memory decode** — `ctx.forwardScope(slabFloats) { … }` recycles one slab per step; creation
316+
*and op outputs* draw from it, and the FP32 fast paths + JVM Panama vector kernels are offset-aware,
317+
so scoped tensors keep SIMD speed. Steady-state decode allocates zero new slab bytes per step.
318+
- **"Will it fit?" in seconds, any format** — header-only footprint plans for GGUF, **safetensors and
319+
ONNX** (external-data sidecars priced correctly), with `PlannerProfile.EDGE` for embedded devices:
320+
`skainet-plan model.onnx --profile edge --budget 2.1G`, exit code 0/1.
321+
- **The compile lane carries what the runtime decides** — tensor identity, structural encodings
322+
(`skainet.tensor_layouts`: block sizes and bit widths as integers, not names) and block order flow
323+
into the exported MLIR and `.irpa`; `HloGenerator.generate(target = …)` runs the first layout pass
324+
on the production path.
325+
- **BitNet / ternary compute** — `BITNET_PLANES` multi-plane packing, i2s GGUF import, and the vendored
326+
NeoGPU ternary f32 NEON kernel through FFM, JNI and Kotlin/Native, including a fused lm_head kernel.
327+
- **Docs that cannot rot** — the Android classifier and ternary getting-started tutorials are compiled
328+
*and executed* in CI (the Iris training loop asserts held-out accuracy ≥ 0.80), plus the
329+
virtual-tensors explanation and SKEEP-003a resolution record.
330+
331+
### Previously, in 0.40.1
300332

301333
- **Correctness hotfix: packed-quant `transpose()` was silently wrong, not crashing.** `ops.matmul(x, ops.transpose(W))` on a packed-quantized weight (Q4_0/Q5_0/Q5_1/Q8_0/Q4_K/Q5_K/Q6_K) with more than one quant block per row produced silently incorrect output — sometimes all-zero — across the scalar, Panama-vector, *and* native (FFM/JNI) kernel tiers, with no exception raised. `transpose()` now performs a real block-grid byte permutation instead of a shape-only relabel; a misaligned packed tensor now throws instead of silently truncating. Closes [#968](https://github.com/SKaiNET-developers/SKaiNET/issues/968). **Upgrading is strongly recommended** for anyone calling `ops.transpose()` on packed-quantized weights.
302334

@@ -326,13 +358,6 @@ val withoutLabel = dataPipeline<RawDataset>()
326358
- **Both narrow formats now beat the FP32 SGEMM** — BF16 by 1.8–1.9x, FP16 by 1.5–1.7x on a 4096x11008 projection. Getting there took a zero-copy transpose for input-major weights (the per-token transpose previously widened the tensor elementwise, 4.4 s per projection), a native FFM FP16 kernel to match the existing BF16 one, and tiling both kernels so the weight is read once per matmul rather than once per input row.
327359
- **Allocation-free shape-only tracing** — `VoidTensorOps` propagates shapes through a `ShapeOnlyTensorData` that allocates no backing buffer, so a dynamic extent flows through a whole decode trace instead of throwing on a negative-size allocation.
328360

329-
### Previously, in 0.37.0
330-
331-
- **`Lstm` layer** — single-layer, batch-first LSTM built from existing primitives only, with `torch.nn.LSTM`-compatible gate order and a caller-owned `LstmState` + `step()` API.
332-
- **Training essentials** — real inverted `Dropout`, mutable optimizer `lr` plus `linearWarmupCosineDecay`, bias-less and `open` `Linear`.
333-
- **Attention scale fix** — `scaledDotProductAttention` at its default scale multiplied every score by zero on the CPU backend; it now resolves to `1/sqrt(headDim)` as documented.
334-
- **Autograd correctness** — `CrossEntropyLoss` no longer detaches the tape, and `softmax`/`logSoftmax`/`variance` backward now work for rank ≥ 3.
335-
336361
See [CHANGELOG.md](CHANGELOG.md) for details and the full release history.
337362

338363
---
@@ -357,6 +382,11 @@ We love contributions! Whether it's a new operator, documentation, or a bug fix:
357382

358383
Browse the full codebase documentation on [DeepWiki](https://deepwiki.com/SKaiNET-developers/SKaiNET).
359384

385+
### Contributors (0.49.0)
386+
387+
- **Michal Harakal** ([@michalharakal](https://github.com/michalharakal)) — the SKEEP-003 memory & storage architecture end to end: M0/M1/M2 milestones, the weight-form and placement-resolution arcs, scope-recycled execution, multi-format footprint analysis, the compile-lane carriage arc, the BitNet/ternary kernel track, and the release docs
388+
- **Ajith Goveas** ([@AjithGoveas](https://github.com/AjithGoveas)) — Iris dataset provider (#1044, #1101), now powering the Android classifier tutorial
389+
360390
### Contributors (0.40.1)
361391

362392
- **Michal Harakal** ([@michalharakal](https://github.com/michalharakal)) — packed-quant `transpose()` block-grid correctness fix, all three kernel tiers (#968, #969)

‎docs/antora.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -15,7 +15,7 @@ asciidoc:
1515
framework_name: SKaiNET
1616
# Current SKaiNET release — bump once per release; referenced as
1717
# {skainet_version} in dependency snippets (blocks need subs="attributes+").
18-
skainet_version: 0.40.1
18+
skainet_version: 0.49.0
1919
ksp_version: 2.2.21-2.0.5
2020
dokka_version: 2.1.0
2121
asciidoctorj_version: 3.0.0

‎docs/modules/ROOT/pages/contributing/build-from-source.adoc‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -167,8 +167,10 @@ Applying `sk.ainet.multiplatform` to the *root* project is not supported and fai
167167

168168
=== Pre-PR Gate and the Packed-Encoding Golden Parity Tests
169169

170-
CI runs the test legs per target (`jvmTest`, `jsTest wasmJsTest wasmWasiTest`, `linuxX64Test`,
171-
`assemble`) and, since the SKEEP-003 roadmap, a `golden-parity` leg. Before opening a PR, run the
170+
CI runs the test legs per target (`jvm`, `js-wasm`, `native`, `android`, plus `assemble`), the
171+
`golden-parity` packed-encoding gate, and — since 0.49.0 — a dedicated `api-compatibility` leg
172+
running `apiCheck`, so a stale API dump fails under its own name. Branch protection on `develop`
173+
requires the aggregated `build-job` to be green, admins included. Before opening a PR, run the
172174
same set locally:
173175

174176
[source,bash]

‎docs/modules/ROOT/pages/contributing/index.adoc‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ The Contributing section is for the engineer who:
2020
- Maintains the CI workflows (smoke runs on `ubuntu-latest`, full
2121
publishable runs on the self-hosted lane).
2222
- Adds or replaces kernels in the CPU backend (scalar, Panama Vector,
23-
the planned native FFM provider).
23+
the native FFM provider).
2424
- Operates the self-hosted runner that publishes benchmark results.
2525
- Drafts or reviews durable API and architecture proposals in the
2626
xref:skeep:index.adoc[SKEEP proposal track].

‎docs/modules/ROOT/pages/contributing/matmul-kernels.adoc‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -144,7 +144,7 @@ flowchart LR
144144
subgraph Selection["At JVM start: ServiceLoader scan"]
145145
P1["ScalarProvider<br/>priority 0, always"]
146146
P2["PanamaVectorProvider<br/>priority 50, JDK 21+"]
147-
P3["NativeProvider<br/>priority 100, planned"]
147+
P3["NativeProvider<br/>priority 100 (FFM)"]
148148
P1 --> Registry
149149
P2 --> Registry
150150
P3 --> Registry

‎docs/modules/ROOT/pages/explanation/eager-execution.adoc‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -60,6 +60,8 @@ mindmap
6060
| Q5_1 | ✅ | ✅ | ✅ | ✅ | ✅
6161
| Q5_0 | ✅ | ✅ | ✅ | ✅ | ✅
6262
| TQ2_0 / BitNet b1.58 (int8 activations) | ✅ | — | — | ✅ NEON | —
63+
| BitNet b1.58 (exact FP32 activations, vendored NeoGPU kernel) | ✅ | ✅ FFM | — | ✅ JNI + fused lm_head | ✅ K/N cinterop
64+
| BITNET_PLANES (multi-plane ternary packing) | ✅ | via kernel pack | — | via kernel pack | via kernel pack
6365
| Q2_K / Q3_K / Q8_K / IQ4 | ❌ (dequant to FP32 only) | ❌ | ❌ | ❌ | ❌
6466
|===
6567

0 commit comments

Comments
 (0)