|
2 | 2 |
|
3 | 3 | ## [Unreleased] |
4 | 4 |
|
5 | | -### Docs |
| 5 | +## [0.49.0] - 2026-08-26 |
| 6 | + |
| 7 | +Headline: **the SKEEP-003 memory & storage architecture, complete — from accepted proposal to shipped system.** |
| 8 | +One storage model (`Storage` / `Scope` / `Format` / `Layout` / `TensorView`), resolver-owned decisions |
| 9 | +(what form a weight takes, where its bytes live — never decided by the model author), scope-recycled eager |
| 10 | +execution (flat-memory decode), and a compile lane that carries what the runtime decides into the exported |
| 11 | +MLIR and `.irpa`. The version jump (0.40.1 → 0.49.0, ~100 merged PRs) is deliberate: this is the release |
| 12 | +downstream repositories (SKaiNET-transformers, the IREE conformity pipeline) should build on, and it |
| 13 | +removes every façade the architecture replaced. See **Breaking changes** below for the migration map. |
| 14 | + |
| 15 | +### Breaking changes |
| 16 | + |
| 17 | +- **The three legacy loader axes are gone** ([#1159](https://github.com/SKaiNET-developers/SKaiNET/issues/1159)): |
| 18 | + `QuantPolicy`, `StagingPolicy` and `WeightOrientation` are deleted. The loader takes one |
| 19 | + `WeightForm(encoding, order, shape, residency)` (uniform `weightForm` or per-tensor `weightFormFor`); |
| 20 | + migration: `DEQUANTIZE_TO_FP32` → `EncodingRequest.DequantizeTo(FP32)`, `StagingPolicy.MAPPED` → |
| 21 | + `WeightForm(residency = WeightResidency.MAPPED)`, `WeightOrientation.OUT_IN` → `WeightShapeOrientation.OUT_IN`. |
| 22 | + `AndroidGguf.loader` takes a `WeightForm` (default mapped residency). |
| 23 | +- **The dead placement machinery is gone** ([#1142](https://github.com/SKaiNET-developers/SKaiNET/issues/1142)): |
| 24 | + `@Place`/`@Weights` (declared, retained, read by nothing), `sk.ainet.lang.tensor.storage.MemoryPlanner`, |
| 25 | + `StorageSpec`, and `Placement.residency`/`Residency`. Lifetime is `ScopeKind`; weight staging is |
| 26 | + `WeightForm.WeightResidency`; placement is decided by `AllocationResolver` (see Added). `Placement` |
| 27 | + itself stays (KV-cache stores carry it); `@KvCache`/`@KvCacheBypass` moved to `KvCacheAnnotations.kt`. |
| 28 | +- **The `skainet.tensor_encodings` module attribute is gone** |
| 29 | + ([#1179](https://github.com/SKaiNET-developers/SKaiNET/issues/1179)): replaced by the machine-readable |
| 30 | + `skainet.tensor_layouts` (`{kind, block_elems, block_bytes, bits, block_order}`); nothing outside this |
| 31 | + repository read the old names dictionary. |
| 32 | +- `LogicalDType` is deprecated end to end in favour of `DType` |
| 33 | + ([#1014](https://github.com/SKaiNET-developers/SKaiNET/issues/1014)); removal at the next major. |
| 34 | + |
| 35 | +### Added |
6 | 36 |
|
7 | | -- **SKEEP-003 accepted — memory & storage architecture design record and roadmap** |
8 | | - ([#932](https://github.com/SKaiNET-developers/SKaiNET/issues/932)). |
9 | | - `docs/modules/skeep/pages/003-unified-tensor-storage.adoc` moves to *Accepted* and records the |
10 | | - thirteen design decisions (storage-first end-state delivered incrementally: `Storage` / `TensorView` / |
11 | | - `Tensor`, scopes, `Format(dtype, encoding)`, `TensorId`, kernel dispatch on declared formats, 2 GB |
12 | | - planner profile, `LogicalDType` → `DType` merge). The full proposal and the M0/M1/M2 milestone PRD are |
13 | | - committed under `docs/design/memory/`; the work is tracked as milestone issues |
14 | | - [#1001](https://github.com/SKaiNET-developers/SKaiNET/issues/1001), |
| 37 | +- **The memory model (SKEEP-003 M0 "know before you load" / M1 "flat decode" / M2 "1.58-bit on a 2 GB board")** |
| 38 | + ([#1001](https://github.com/SKaiNET-developers/SKaiNET/issues/1001), |
15 | 39 | [#1002](https://github.com/SKaiNET-developers/SKaiNET/issues/1002), |
16 | | - [#1003](https://github.com/SKaiNET-developers/SKaiNET/issues/1003) with one sub-issue per feature branch. |
| 40 | + [#1003](https://github.com/SKaiNET-developers/SKaiNET/issues/1003); slices #1004–#1042): |
| 41 | + `Storage` (Heap/OffHeap/Mapped, ownership + liveness — a use-after-free is a loud |
| 42 | + `StorageClosedException`), `Scope` (`ModelScope`/`ForwardScope` slab with `reset()`), |
| 43 | + `Format(dtype, encoding)`, `Layout` (strides/offset/block geometry), `TensorView` with `prepack()` as |
| 44 | + the visible relayout and `materialize()` as the single copy point; `TraceSink` events (allocations, |
| 45 | + scope resets, adapter insertions); `MemoryPlan`/`MemoryPlans` header-only planning with budgets, |
| 46 | + suggestions and a plan-vs-actual check; the `skainet-plan` CLI; `PlannerProfile` (`MOBILE_2GB` with |
| 47 | + automatic KV quantization — and `strict`, so a missing kernel refuses instead of silently costing |
| 48 | + several times the weight, [#1128](https://github.com/SKaiNET-developers/SKaiNET/pull/1128)); |
| 49 | + Android mmap loading + device fit checks; KV cache preallocated in model scope with declared formats; |
| 50 | + kernel dispatch on declared formats (`KernelKey`, registry-backed capabilities); the decode harness |
| 51 | + and M2 acceptance runs. |
| 52 | +- **Weight forms — one resolved decision instead of three caller flags** |
| 53 | + ([#1109](https://github.com/SKaiNET-developers/SKaiNET/issues/1109) arc, #1114–#1120): |
| 54 | + `WeightForm` (encoding × byte order × shape × residency) resolved by `WeightFormResolver` from |
| 55 | + *what the file holds × the profile × what the backend's kernels can feed*; the loader honours it, |
| 56 | + the plan prices it (a resolved dequantization shows in the table, not at the OOM), conversions are |
| 57 | + traced, and packed weights can load directly in kernel-feed order |
| 58 | + ([#1120](https://github.com/SKaiNET-developers/SKaiNET/issues/1120)). |
| 59 | +- **Placement is resolver-owned** ([#1133](https://github.com/SKaiNET-developers/SKaiNET/issues/1133) → |
| 60 | + #1142–#1144): `AllocationResolver.resolve(weight, profile, platform)` decides memory domain and scope |
| 61 | + (mapping requires: the form asks, the platform can, the bytes are the file's bytes); |
| 62 | + `AllocationResolver.explain()` renders every decision with its reason pre-load; `ResolvedGguf` wires |
| 63 | + plan → load with the documented user-wins precedence (per-tensor `weightFormFor` > uniform |
| 64 | + `weightForm` > resolver). |
| 65 | +- **Scope-recycled eager execution** ([#1135](https://github.com/SKaiNET-developers/SKaiNET/issues/1135) → |
| 66 | + #1145/#1146/#1173): `ExecutionContext.memoryScope` is consulted by tensor creation *and* op outputs |
| 67 | + (`TensorDataFactory.adoptFloatArray`, `ScopedTensorDataFactory`), so |
| 68 | + `ctx.forwardScope(slabFloats) { … }` gives steady-state decode that allocates zero new slab bytes per |
| 69 | + step; the FP32 fast paths and the JVM Panama vector kernels are offset-aware, so slab-backed tensors |
| 70 | + keep SIMD speed. |
| 71 | +- **Model-footprint analysis for GGUF, safetensors and ONNX** |
| 72 | + ([#1169](https://github.com/SKaiNET-developers/SKaiNET/issues/1169)): header-only `planInput` for all |
| 73 | + three formats (ONNX `external_data` sidecars priced correctly — multi-GB models no longer report ~0 |
| 74 | + bytes; sizes `Long`-safe), and `PlannerProfile.EDGE` for embedded devices where the budget *is* the |
| 75 | + usable RAM. "Will it fit in ~2.1 GB?" is answered in seconds, without reading a tensor payload. |
| 76 | +- **The compile lane carries what the runtime decides** |
| 77 | + ([#1147](https://github.com/SKaiNET-developers/SKaiNET/issues/1147) → #1178/#1179/#1180): |
| 78 | + `TensorRef` carries tensor identity (`TraceSession.identify`, registered from |
| 79 | + `trainableParameters()`), the encoding *object* (block size intact) and packed block order across the |
| 80 | + trace→graph boundary; the emitted module header declares structural facts per tensor |
| 81 | + (`skainet.tensor_layouts`); `ExternalParameterRef` declares block order to the `.irpa` consumer; |
| 82 | + `HloGenerator.generate(target = …)` runs the optimizer pipeline on the production path with |
| 83 | + `LayoutAssignmentPass` (rank-2 packed weights get kernel-feed order; the tape's carried facts are |
| 84 | + never overridden), and the `ResolvedComputeGraph` seams surface exactly the decisions made. |
| 85 | +- **BitNet / ternary compute track** (#1033, #1040/#1041, #1136–#1141, #1150): |
| 86 | + ternary encodings (`TQ1_0`/`TQ2_0`, `BITNET_B1_58`, `BITNET_PLANES` multi-plane packing), i2s GGUF |
| 87 | + import, the vendored NeoGPU ternary f32 NEON kernel (MIT, verbatim) exposed through FFM, JNI and |
| 88 | + Kotlin/Native including a fused lm_head kernel, requant adapters, NEON BitNet packing, and ternary |
| 89 | + benchmarks + getting-started docs. |
| 90 | +- **Iris dataset provider** ([#1044](https://github.com/SKaiNET-developers/SKaiNET/issues/1044), |
| 91 | + [#1101](https://github.com/SKaiNET-developers/SKaiNET/pull/1101), contributed by @AjithGoveas): |
| 92 | + the embedded 150-row dataset used by the new Android classifier tutorial. |
| 93 | +- Sliding-window KV SDPA ([#1036](https://github.com/SKaiNET-developers/SKaiNET/issues/1036)) and |
| 94 | + KV formats declared by the store ([#1077](https://github.com/SKaiNET-developers/SKaiNET/issues/1077)). |
| 95 | + |
| 96 | +### Fixed |
| 97 | + |
| 98 | +- Packed block order end to end: feed-order bytes in a type claiming canonical order decoded to |
| 99 | + plausible garbage ([#1124](https://github.com/SKaiNET-developers/SKaiNET/issues/1124), |
| 100 | + [#1126](https://github.com/SKaiNET-developers/SKaiNET/pull/1126)); `TensorData` now declares its |
| 101 | + `BlockOrder` and every reader agrees on the same bytes. |
| 102 | +- Packed ternary weights reach dispatch through the Wᵀ marker — ANY packed block storage routes through |
| 103 | + the marker path ([#1136](https://github.com/SKaiNET-developers/SKaiNET/issues/1136), |
| 104 | + [#1181](https://github.com/SKaiNET-developers/SKaiNET/pull/1181)). |
| 105 | +- M2 acceptance page-fault flake ([#1107](https://github.com/SKaiNET-developers/SKaiNET/pull/1107)); |
| 106 | + safetensors dtype mapper no longer prints a WARNING into stdout mid-parse (#1169). |
| 107 | + |
| 108 | +### CI & process |
| 109 | + |
| 110 | +- `apiCheck` has its own named PR leg (`test (api-compatibility)`) instead of hiding inside |
| 111 | + golden-parity ([#1176](https://github.com/SKaiNET-developers/SKaiNET/pull/1176)), after a stale dump |
| 112 | + reached `develop` unnoticed ([#1174](https://github.com/SKaiNET-developers/SKaiNET/pull/1174)); |
| 113 | + branch protection on `develop` now requires the full `build-job` aggregator, admins included. |
| 114 | +- Android per-target API dumps dropped (jvm + klib only, |
| 115 | + [#1111](https://github.com/SKaiNET-developers/SKaiNET/pull/1111)). |
| 116 | + |
| 117 | +### Docs |
| 118 | + |
| 119 | +- The memory model as built: `explanation/memory-model.adoc`, `explanation/packed-weight-layout.adoc` |
| 120 | + ([#1106](https://github.com/SKaiNET-developers/SKaiNET/pull/1106) — design drafts under |
| 121 | + `docs/design/` retired in favour of Antora pages), `explanation/virtual-tensors.adoc` (the ML |
| 122 | + Drift-style split, with diagrams), and SKEEP-003a (`skeep/003a-placement-and-planning-resolution.adoc`) |
| 123 | + recording the P7/P8 resolutions. |
| 124 | +- Tutorials whose code cannot rot: the Android classifier getting-started and the ternary |
| 125 | + getting-started, with snippets compiled *and executed* in CI by `skainet-docs-samples` (the Iris |
| 126 | + training loop asserts held-out accuracy ≥ 0.80). |
| 127 | +- SKEEP-003 accepted ([#932](https://github.com/SKaiNET-developers/SKaiNET/issues/932)) with its |
| 128 | + thirteen design decisions; how-to: plan a model's memory before loading it, including the |
| 129 | + embedded-device (`edge`) verdict. |
17 | 130 |
|
18 | 131 | ## [0.40.1] - 2026-08-12 |
19 | 132 |
|
|
0 commit comments