All notable changes to SKaiNET-transformers are documented here. The
version line tracks the underlying SKaiNET engine (sk.ainet.core:*): a transformers
release carries at least the engine's X.Y, and its patch number may advance on its
own — a fix confined to transformers ships as X.Y.Z+1 against engine X.Y.Z without
requiring an engine release. 0.57.1 against engine 0.57.0 is such a release.
The format roughly follows Keep a Changelog, and this project adheres to Semantic Versioning.
Ships against SKaiNET engine 0.57.0: this release touches only the Android IREE runtime's streaming Moonshine path, no engine API.
- The streaming partial decode now has a real-time budget. On an ARM Android device with a Mali GPU the stream did
not keep up with the microphone: per 0.88 s hop it spent ~350 ms on the encoder window and another
~350 ms on a prefill plus greedy steps, so it ran at 1.2–1.45× real time and the whole backlog was
still owed at the moment the user stopped speaking. The encoder path is mandatory — it builds the
cross memory the final transcript is decoded from — but the decode hops only produce text to show
while the user is still talking, and
finish()re-decodes from that memory regardless. PastMOONSHINE_LAG_BUDGET_MS(default 250) the window runs and the decode is skipped, which provably cannot change the final result. Measured on that device: RTF 1.23 → 0.67, drain 656 ms → 6 ms. Verified against a 444-utterance German evaluation set: on 311 comparable rows 310 final transcripts were byte-identical, and the single difference is a runaway repetition on both sides. The budget is disabled whenMOONSHINE_FAST_FINISHis on, where the incremental result is the answer.
-
IreeMoonshineStream.stats()— what the lastfinish()cost, by stage. These numbers only ever went to logcat, so a caller that wanted to know where recognition spent its time had to scrape the device log. Now: the flush/decode split of the final decode, how far the stream fell behind the microphone, how many partial decodes the budget gave up, encoder windows and decode hops run, and the utterance's size in frames, tokens and milliseconds. Atimelineentry carries each event with its timestamp and cost (w/p/h/drecords), so the shape of an utterance can be drawn rather than only totalled. A string map rather than a typed record on purpose: the values cross two more artifacts before anything consumes them, and adding a counter must not change a signature on the way. -
Audio that arrives before a run is open is held instead of dropped.
sendAudioChunkwasactiveRun?.feed(...): with no run open the chunk vanished silently. Opening a run costs 229 ms of the 259 ms measured between a push-to-talk trigger and the first recorded sample, and a host that waited for it lost that much speech off the front — enough to swallow the first word of an utterance entirely. Chunks now go into a pre-roll, bounded to 1 s keeping the newest and discarded entirely past 2 s of age, and are fed in order once the run opens. Measured: 259 ms → 68 ms, and an utterance whose first word was previously lost transcribes in full.
Lock-step with SKaiNET engine 0.57.0. Headline: Qwen on the compiled IREE KV path: the qwen-kv-v1 export, runtime and templates, with Qwen3-0.6B verified token-for-token against llama.cpp on host IREE; on-device validation is still to come (see the entry below for exact status).
- Engine 0.57.0, Kotlin 2.4.20 (#453). The engine fixes the StableHLO export of an explicit
attention mask under grouped-query attention and onto a dynamic key length (SKaiNET#1302), and
makes kotlinx-io part of
skainet-data-source's API. This repo moves to Kotlin 2.4.20 with it. - Dependencies: Ktor client 3.6.0 (#448), kctfork 0.14.0 (#449), jackson-databind 2.22.3 (#454).
- FunctionGemma export contract carries the embedding geometry (#452).
FunctionGemmaSpecgainsvocabSize(default 262144),nHeads,hiddenSizeandslidingWindow, andmanifest.jsonemits them, soIreeKvSpec.fromManifestno longer falls back to the stock constants. The export CLI reads the vocabulary size from the checkpoint'stoken_embd.weight(GEMMA_VOCABoverrides): a fine-tune that added special tokens has more rows, and the native KV session locates the embedding table byvocabSize × hiddenSizebytes and clamps ids ≥vocabSizeto 0, so with the stock constant those tokens would silently have become token 0.IreeKvSpec.functionGemma270mtakesvocabSize.
- Stale
functiongemmaAPI dump. #452 changed the public API without re-dumping it, soapiCheckfailed ondevelop; refreshed in #453.
Status in this release: Qwen3-0.6B exports, compiles for Vulkan (valhall4) and arm32, and matches
mainline llama.cpp's 32 greedy tokens on host IREE (QwenVmfbParityTest). Not yet run on a device.
Qwen2.5-0.5B export is blocked by a SKaiNET core gap (attention bias does not externalize).
IreeKvSession/IreeKvSpecgeneralized for GQA and single-RoPE-base models (llm-runtime:iree-android):nKvHeadsmay now be any value withnHeads % nKvHeads == 0(grouped-query attention — Qwen2.5-0.5B nKvHeads=2, Qwen3-0.6B nKvHeads=8), not just FunctionGemma's plain multi-headnKvHeads == 1.globalLayerPeriod == 1(every layer "global") skips building the sliding-side RoPE/mask tensors entirely for models with no sliding-window/global split — the qwen-kv-v1 shape. New optionalIreeKvSpec.maskHeads(0=nHeads, the default and FunctionGemma's per-head mask;1= one head-shared chunk mask[1, 1, C, past+C], what the Qwen factories andmanifest.jsonset, read byfromManifest); the mask rows never depended on the head, so this only shrinks the buffer. The 11-argument constructor stays (@JvmOverloads).nativeCreateopens one shared IREE session when the same vmfb+irpa path is passed for all three graphs — the merged-module design a three-archive-per-model contract can't fit in a 32-bit process.IreeKvSpec.qwen25_05bInstruct()/qwen3_06b()factories. FunctionGemma's behaviour and binary compatibility are unchanged.DecoderKvModel(llm-core): an architecture-neutral KV-cache decode path (forwardPrefillAt/forwardPrefillWithPast/forwardWithPast) over anydecoderTransformerNetwork-built module — theGemmaModelwith-past forwards, generalized: GQA-native SDPA (noexpandKV), Q/K/V/O projection bias support, one RoPE base, no sandwich norms/PLE/softcapping. Verified numerically identical to the eager reference path.QwenKvArch/QwenKvContract(llm-inference:qwen, commonMain): the qwen-kv-v1 manifest contract (manifest.jsonemission, arg/output orders forqwen_prefill_at/qwen_prefill_with_past/qwen_with_past), architecture facts derived from the loaded checkpoint (GQA, QK-norm, attention bias) rather than hardcoded per model.QwenExportHarness/QwenExportCli(llm-inference:qwen, jvmMain,:llm-inference:qwen:exportQwen): traces the three qwen-kv-v1 graphs from a Qwen2/Qwen3 GGUF and emits StableHLO MLIR + bf16 safetensors +manifest.json, followingFunctionGemmaExportHarness's structure. Verified against both real checkpoints: Qwen3-0.6B exports correctly end to end (all three graphs, real weights, dynamic KV-cache dims). Known gap: Qwen2.5-0.5B-Instruct's attention bias does not externalize as a baked weight during tracing — it leaks into the compiled function's signature as a runtime argument instead (a SKaiNET core tracer gap:ModuleParameter.BiasParameterisn't recognized as an externalizable constant the wayWeightParameteris; no existing harness in this repo had ever traced a bias-bearing attention layer before this one). Qwen2.5-0.5B export is blocked on a SKaiNET core fix; Qwen3-0.6B is not. Per-graph archives for now, not the merged single-archive design qwen-kv-v1 calls for —IreeKvSession's non-shared path already supports this at FunctionGemma's existing three-archive memory cost; the merge is tracked as a follow-up. Vulkan (valhall4) compiles. The first real compile crashed inSPIRVInitialVectorLoweringPass; bisecting on real Qwen3-0.6B exports found three independent causes, none of them bf16: (1) IREE 3.11's SPIR-V backend cannot lower the fused argMax reduction unless its extent is a multiple of 2048 (the 151936 vocab fails, FunctionGemma's 262144 does not), so the harness pads the logits row with a finite minimum to 153600 before the argMax, leaving real logits and the returned id unchanged; (2) the in-graph token-embedding gather does not lower either, soQWEN_HOST_GATHER=1applies the host-gather rewrite to every function (the contract'sembargument, what the native runtime already passes) — both are post-emit rewrites, unit-tested inQwenExportRewriteTest; (3) SKaiNET 0.56.0's SDPA converter emitted a staticbroadcast_in_dimof the chunk mask onto the dynamic grouped-query scores shape, invalid on every backend — fixed in the engine (SKaiNET#1302, 0.57.0), which now emits the hinteddynamic_broadcast_in_dimIREE 3.11 lowers for the head-shared chunk mask. With all three, the Qwen3-0.6Bqwen_prefill_at(seq 1024),qwen_prefill_with_pastandqwen_with_pastgraphs compile for bothvulkan-spirv valhall4andllvm-cpu arm32.QwenVmfbParityTest(llm-inference:qwen, gated onQWEN3_06B_GGUFand docker): the compiled Qwen3-0.6B graphs, exported as they ship (bf16, host-gather, padded argMax, head-shared mask), run on host IREE throughiree-run-moduleexactly as the native session drives them (prefill-at over 3 prompt tokens, one 32-token chunk call, 31 decode steps) and reproduce mainline llama.cpp's 32 greedy tokens exactly (fixtureqwen3-06b/golden-greedy-06b.txt, smallest top-1/top-2 gap 2.35 nats). The run's evidence record isllm-inference/qwen/validation/L3-qwen3-06b-host-vmfb.json.Qwen25ChatTemplate(llm-agent): faithful to Qwen2.5-Instruct's officialchat_template(verified against a real Jinja2 render ofQwen/Qwen2.5-0.5B-Instruct'stokenizer_config.jsonfetched from huggingface.co) — a default "You are Qwen, created by Alibaba Cloud…" persona when the caller supplies no system message, and no<think>handling anywhere (Qwen2.5 predates thinking mode entirely;QwenChatTemplate(enableThinking = false)'s empty<think></think>prefill is itself out-of-distribution for it). Measured on an 8-utterance zero-shot tool-calling set: 0/8 and 1/8 under generic ChatML /QwenChatTemplatevs 4-5/8 with this template.
nativeFinishcaps the exact re-decode atmax_new_tokens = samples * 6.5/16000 + 2(llm-runtime:iree-android,moonshine_stream_jni.c), the output-length cap the Moonshine v2 model card specifies, instead of a fixed 24 tokens regardless of how much audio the utterance held. Without it a short clip keeps decoding long after the audio is spent, and the greedy decode spends the remaining budget restarting the utterance rather than stopping. The budget is floored at 4 so a one-word command still has room and ceilinged at the previous fixed 24, so nothing decodes longer than before — above roughly 3.4 s of audio the formula exceeds the ceiling and the change is a no-op. Thefinishtrace line now reports tokens used against the budget.- Measured on 183 German command-and-control recordings on a Mali device, using clips above the threshold as a control group: below it the turn finishes 0.58 s sooner (paired median, faster in 135 of 172); above it unchanged (+0.09 s, 11 recordings). Median word error rate does not move. The saving appears exactly where the cap can act and nowhere else.
- The zero-padded final encoder window is a separate, still-open gap (#458): the encoder graph takes features only and carries no attention mask, and moving the window is not the fix — a variant that end-aligned it measured no effect on those same 11 control recordings and was dropped from #455.
A transformers-only release against SKaiNET engine 0.56.0 (unchanged).
IreeMoonshineStream(llm-runtime:iree-android,libskainet_moonshine_stream.so, arm64-v8a + armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five graphs ofMoonshineV2ExportCli— PCM in, cumulative partial transcripts out, exact final onfinish(). The counterpart ofIreeKvSessionfor ASR: the last piece a Moonshine cartridge needed from a released artifact.native/build-moonshine-stream.shbuilds it with the same image and cache as the other two libraries.
A transformers-only release against SKaiNET engine 0.56.0 (unchanged).
Added — Moonshine v2 streaming: every checkpoint of the family, exported from the published artifact
MoonshineV2ExportCli(:llm-inference:moonshine:exportMoonshineV2, JVM): a Hugging Facemoonshine_streamingsnapshot in (config.json,model.safetensors,tokenizer.json), the five StableHLO graphs of the streaming contract out —frontend,encoder,adapter, masked fixed-padprefill, dynamic-cachewith_past— plusdec_embed.bin,vocab.binand amanifest.json. Geometry, per-layer attention bands, vocabulary size and the positional-table length are read from the snapshot; the export fails if a checkpoint tensor was not used. The export used to live injvmTestbehind per-tensor.bindumps made by Python scripts, so it could not be run from the published module; it now can, and needs no Python. Output is byte-identical to the previous flow for the German tiny checkpoint (all five graphs and both tables); for the English one the encoder, adapter and both decoder graphs are byte-identical and the frontend differs only in conv-weight constants by ≤ 7.2e-7 (the old flow took them from a weight-norm-folded ONNX).MoonshineV2HfWeightMap(common): DSL parameter → checkpoint tensor for frontend, encoder, adapter and decoder, including the three things that are not a plain copy (transposed filterbank, zero-centred encoder norm scales, absent biases).MoonshineV2Config.slidingWindows: explicit(left, right)context per encoder layer, in the checkpoint's own convention. The German checkpoints need it (their middle layers look one frame ahead);nullkeeps the English edge-layer rule.- Split encoder/decoder widths (
decoderDim,decoderFfnDim, the adapter'sproj, explicit encoder head dim, frontend width) for the small checkpoints. Defaults are unchanged.
Back in lock-step with the engine: ships against SKaiNET engine 0.56.0 (the engine skipped
0.55.0 to realign the two version lines). The engine's scaledDotProductAttention is now
grouped-query native, so GQA models stop tiling K/V up to the query heads — eagerly, on the tape and
in the exported StableHLO — and the compiled leg of SKEEP-005 lands here: structure at compile time,
cores at run time.
- GQA without head expansion on the tape: the engine's SDPA is grouped-query native, so
MultiHeadAttentionandHybridTransformerBlockhand K/V to it with their own head count —repeatKVHeads(nKV × narrow + concat per K and V per layer per step) is gone from tapes, traced graphs and StableHLO exports, which now batch attention over the head groups. - Structure in the export header: the SmolLM2 and FunctionGemma harnesses run
ScheduleAnnotationPass, so every exported attention statesparallel_dims = [batch, heads](advisoryskainet.schedule; no core count is ever written into a module). - Compiled JVM leg under the schedule: OPTIMIZED mode runs
ComputeGraphExecutoroverctx.ops, so a scheduled context parallelises it too —OptimizedModeScheduleParityTest(bit-identical sequential vs hardware, compiled ≈ eager), two OPTIMIZED rows inAttentionScheduleSpeedProfile. - IREE run-time core knob:
IreeRedecodeSession(taskTopologyGroupCount),IreeRedecodeDecoder.fromAssets(taskTopologyGroupCount = IreeTaskTopology.fromEnv()),IreeTaskTopology(SKAINET_TASK_GROUPS,groupCountFor(schedule.parallelism)), JNInativeCreateWithTopology(feeds--task_topology_group_countto IREE before the local-task device is created); both ABIs'libskainet_iree_redecode.sorebuilt.gemma-ireereadsSKAINET_TASK_GROUPStoo (GEMMA_TASK_GROUPSdeprecated alias). Docs: spec "Phase 2", explanation "The compiled leg", IREE Android runtime reference "Task topology", eager-vs-compiled row.
- Engine 0.56.0 (
skainet = "0.56.0"): grouped-query-native SDPA, its StableHLO lowering with the head groups as a batching dimension, structural schedule defaults, schedule-aware graph contexts. IreeRedecodeSessionqualifies bare function names in both create paths: the module-qualified name (module.<fn>) now feedsnativeCreateandnativeCreateWithTopologyalike.- Convention plugins 1.1.0 (
sk.ainet.multiplatform,sk.ainet.npm-pins,sk.ainet.transformers.bom-coverage);asr-domain,llm-coreandtransformer-corepinjvmTarget = JVM_21— 1.1.0 defaults to 17, which cannot inline the engine's JVM 21 bytecode. - Benchmark notes describe the measurement hardware by device class; the MiniLM export harness uses domain-neutral probe sentences.
- Dependency bumps: AGP 9.4.1, kotest 6.2.5, kotlinpoet 2.4.0, binary-compatibility-validator 0.18.2.
A transformers-only release, same pattern as 0.54.1: no new engine version, still against
SKaiNET engine 0.54.0. Adds a new asr-domain module and extends BackendProvider,
generalizing plumbing that used to live downstream in one ASR cartridge family's repo.
- New
asr-domainmodule (sk.ainet.asr.domain, publishes asskainet-transformers-asr-domain): generic ASR task types —Transcription,TranscriptionTimings,StopReason,AsrEvent,DecodingOptions,FeatureFrames. Moved up from the downstream ASR cartridge ecosystem, where they lived only because that's where the originalwhisper-cliextraction happened to put them, not because they're Whisper-specific — both the Whisper and Moonshine cartridge families depend on these, and keeping them downstream in one family's repo made the other structurally dependent on it for generic plumbing. Framework-free (emptycommonMaindeps), full KMP target spread matchingllm-api's convention for consumer-facing SPI modules (ios/linux/macos/jvm/js/wasm/android). BackendProvidergainscapabilities: BackendCapabilities(supported dtypes, compile support, NPU usage, max sequence length — defaulted, so existing implementers don't need to change) andcreateContext(options: BackendOptions = BackendOptions())(was parameterless).kllama'sCpuBackendProviderupdated to match.- No second backend-registry module. Downstream's own
backend-spiExecutionContextFactory/BackendRegistryseam is not duplicated here — unified into this existingBackendProvider/BackendRegistryinstead, since both did the same job (select a strategy producing a SKaiNETExecutionContext).BackendRegistryitself is unchanged. - Verified:
:asr-domain:build,:llm-core:build,:llm-runtime:kllama:buildand the correspondingallTestsall green;apiDumpregenerated for bothllm-coreandkllama's binary-compatibility-validator baselines. Downstream (asr-whisper-iree-cartridge,asr-cli) verified against this branch via theuseLocalSkainet-style opt-in composite substitution pattern before this release — fullcheckgreen in both, including a Docker smoke test of the shippedasr-cliimage.
A transformers-only release, same pattern as 0.40.2: no new engine version, still against
SKaiNET engine 0.54.0. MultiHeadAttention becomes the first consumer of the engine's
SKEEP-005 Schedule, and positional KV caches stop copying the whole prefix per layer per token.
- Attention heads run in parallel (#413):
MultiHeadAttentionmaps heads (or GQA groups) onto cores through the engine's newExecutionContext.schedule(SKaiNET SKEEP-005).AttentionSchedulePolicy(Sequential/PerHead/PerKVGroup/Auto, defaultAuto) plans the tasks;ScalarHeadAttentionKernelkeeps the exact per-head rounding order, so the result is bit-identical to 0.53.0. The fused path now also covers batched prefill and sliding-window layers (engine SDPA rounding order), removingrepeatKVHeads/permute/reshapefrom the hot path. The DSL is unchanged; override per layer withmha.schedule/mha.schedulePolicyorModule.configureAttention(...). - Copy-free K/V views (#412):
KVCache.updateInPlacereturns aKVBufferViewover the cache's own buffers forPositionalKVCacheand its shared / padded / read-only wrappers (best-effort forAppendKVCache), so decode no longer copies the whole prefix per layer per token. - Positional cache for Llama and Qwen:
DecoderKVCacheKind(APPENDdefault,POSITIONAL),decoderTransformerNetwork(kvCacheKind = …), theATTENTION.positionalKvCache(...)DSL clause, andLlamaNetworkLoader / QwenNetworkLoader.withKVCacheKind(...)/fromWeights(weights, kvCacheKind = …). - Verification:
MultiHeadAttentionScheduleParityTest,KVCacheInPlaceViewTest; the Llama and Qwen golden gates acceptSKAINET_ATTN_SCHEDULE=sequential|parallelandSKAINET_KV_CACHE=append|positional(verified on Llama-3.2-1B, Qwen2.5-0.5B and Qwen3-1.7B Q8_0);AttentionScheduleSpeedProfile(opt-in) measures all four combinations. Docs:docs/specs/attention-schedule.md, the Parallel Attention Heads via Schedules explanation and the Parallel Attention — Getting Started tutorial.
Version lock-step with the engine continues: this release ships against SKaiNET 0.54.0
(SKEEP-005's Schedule API, the CoroutineSchedule pool-deadlock fix, SafeTensorsParametersLoader
tensorFilter parity, and the ExperimentalMemoryApi opt-in gate removed). It also closes out the
FunctionGemma/IREE-Android chunked-KV work (#410) and fixes two bugs found on real hardware: a
broken runtime-kgemma Maven Central POM (#408) and an Android crash in FunctionGemma tool-call
parsing (#407).
- Position-selected graphs
gemma_at/gemma_prefill_at(#415): the LM head runs on one one-hot-selected position instead of every position inSEQ, cutting the redecode step's wasted work (#406). - Chunk prefill-with-past graph
gemma_prefill_with_past(#417): a fixed 64-token chunk against the dynamic cache in one call, with per-head chunk masks (a broadcast over heads to a dynamic shape isn't expressible in static StableHLO). IreeKvSession/IreeKvDecoder(#416, #418): the Android-native stateful KV session — three IREE sessions, device-resident K/V, zero-copy 512-position tail views for the sliding layers, native RoPE tables + chunk masks, embedding rows read from the archive, snapshot/restore without copies. Measured on an arm32 Android device (Mali via Vulkan, bf16 archives): a once-per-process 843-token catalog prefill, then p50 5.87 s / max 6.25 s per utterance (restore + one chunk + 16 decode tokens), down from minutes on the stateless redecode contract.iree-androidfailure reporting (#404): native failures surface as reported errors instead of a silentnull; bare function names are qualified withmodule.automatically.
- FunctionGemma export
ClassCastExceptiononBufferHandle.Floats(#405, #420):FunctionGemmaExportHarness,SmolLm2ExportHarness, and the bake-irpa tests hard-cast every external constant toBufferHandle.Owned; since engine 0.53.0 (SKaiNET#1247) a constant can arrive as the aliasedBufferHandle.Floatsinstead. Both export harnesses now read every handle throughDefaultBufferResolver. runtime-kgemma's Maven Central POM depended on an unpublished coordinate (#408): pulling in:llm-runtime:kgemma3n(never published) leakedSKaiNET-transformers.llm-runtime:kgemma3n-jvm:unspecifiedinto the POM, breaking resolution for any external consumer. Rather than just changing the dependency's scope, Gemma 3n itself stops being published (SKaiNET-transformers#377: maturity gate 0/5, hand-rolled runtime that force-dequantizes the whole model, postponed by decision) — source-only until #377's maturity gate is met.runtime-kgemma's CLI loses its--arch gemma3nvariant accordingly;skainet-cli(never published) is unaffected.- Android crash in the official FunctionGemma tool-call parser (#407):
FunctionGemmaOfficialToolCallParserStrategy.CALL_REhad an unescaped closing}— tolerated byjava.util.regexon the JVM, rejected by Android's ICU-backed engine with aPatternSyntaxExceptionat class-init, so every tool-call parse crashed on ART before the first match. One-character fix; no-op on the JVM.
Version lock-step with the engine is restored: this release ships against SKaiNET 0.53.0
(#397), which brings the billion-parameter export fixes
(SKaiNET#1247) and the sharded
SafeTensors ParametersLoader (SKaiNET#1246)
this repository's Gemma 3n export and family loaders were waiting on. Everything accumulated under
"Unreleased" since 0.40.2 — BitNet, the engine-loader migration, the Gemma 3n DSL path, the Qwen and
Apertus fixes — ships here too.
llm-apps:skainet-decode-core(#395): the decode-and-measure flow extracted from the JVM CLI into a commonDecodeSession(jvm + android) — traced prefill/decode/sample loop,MemoryProbe.sample().emitTo(sink)inside every decode span so the page-fault/RSS rows ofGenerationMetricspopulate on Android and Linux, optional extraTraceSinkfor Perfetto. The JVMskainet-decodeCLI is a thin caller with identical output.llm-apps:skainet-decode-android(#395): the repository's firstcom.android.application— a single-activity app that loads a pushed GGUF throughMappedRandomAccessSource, refuses viaAndroidGguf.fitsbefore allocating, decodes on one dedicated thread, and reportsGenerationMetricsplus RSS/page-fault deltas to screen, logcat anddecode-report.md. The APK carries both JNI kernel variants and theViewKernelPack/KernelProviderServiceLoader entries, so the engine's self-healing dispatch survives packaging. The physical-device measurement lane (the SKEEP-002 numbers) is documented in the module README and still to be recorded.
- Gemma (#398), the shared decoder loader (#400), Apertus and Gemma 3n (#401): the
hand-rolled per-family SafeTensors materialization (bf16/f16 widening, byte decoding, size guards,
dead transposes) collapses onto the engine's
ShardedSafeTensorsParametersLoader/SafeTensorsParametersLoader. Each family keeps only its HF→GGUF slot table, name allowlist (as the engine'stensorFilter) and any shape normalization; every dtype decision — includingRequire(BF16)/Require(FP16)keep-native, now accepted on the SafeTensors lane — is the engine's. Every collapsed loader gains adtypePolicy: DTypePolicy = Anyparameter (existing call sites source-compatible) and a synthetic 2-shard fixture test written with the engine'sSafeTensorsWriter. Not yet collapsed: Voxtral (customQUANT4format) and llm-core's legacy Q4/non-float path, both waiting on a single-filetensorFilter(SKaiNET#1256).
Gemma3nExportHarness.writeSafetensors(#396) no longer casts every constant toBufferHandle.Owned: with engine 0.53.0, ≥2 GiB FP32 constants arrive as an aliasedBufferHandle.Floats(the tied embedding is exactlyInt.MAX_VALUE + 1bytes), so the harness streams every handle throughDefaultBufferResolverin bounded chunks with the same chunked bf16 conversion. The full 30-layer E2B export now emits a 15k-line StableHLO module with zero failure comments and a 4.6 GB safetensors in under a minute, where it previously OOMed a 46 GB heap.
Gemma4E2BToolCallSmokeTestre-enabled (#399): the real Gemma 4 E2B Q4_K_M checkpoint now emits<|tool_call>call:calculator{expression:...}and every assertion holds (it had been@Ignored for emitting prose without markup). Also re-run green in the same pass:FunctionGemmaOfficialGgufTest(parsedget_weathercall) andQwenToolCallSmokeTestwith Qwen3-1.7B-Q8 (well-formed<tool_call>calculator call).
exportGemma3n(Gemma3nExportHarness, SmolLM2/FunctionGemma redecode pattern): tracesgemma3nNetwork()to StableHLO with external bf16 params and an in-graph argMax tail. Mobile-honest contract:per_layer_inputsis a graph INPUT computed on the CPU from the packed PLE table at runtime (PLE's design point — those parameters stay off the accelerator), so the parameter archive carries the trunk + token embedding only.PerLayerEmbeddinggained a traceableindexSelectpath while recording;GEMMA3N_LAYERStruncates the trunk for pipeline verification on smaller hosts. Full E2B emission is blocked on engine SKaiNET#1247 (trace memory co-residency + an HLO converter operand-linkage defect) — the harness hard-fails on both signatures instead of shipping a silently-unservable module.- New antora explanation page
explanation/gemma3n.adoc(why Gemma 3n's mobile-first architecture and why SKaiNET fits it) and pre-PRD design notedocs/specs/matformer-hybrid-on-device-ai.md(MatFormer elasticity in SKaiNET + hybrid on-device/cloud routing: draft-first, escalate-on-evidence).
gemma3nNetwork()+Gemma3nModel— the full Gemma 3n text architecture declared in the DSL, faithful to HFmodeling_gemma3n.py: AltUp (four parallel hidden streams with the tanh modality router;Gemma3nAltUpBlockper layer,Gemma3nAltUpGlobalsfor the magnitude-renormed stream init/merge), Laurel, Gaussian-top-k activation sparsity on the first ten layers (driven by the GGUF's precomputed per-layer std multipliers;-inf= off), PLE feeding the non-active streams (reusing the gemma-4 lane'sPerLayerEmbedding— the math is identical), per-type shared KV for the last ten layers, hybrid sliding/global attention with dual RoPE bases, q/k-norm + parameterless v-norm, attention scale 1.0. All math goes throughctx.ops, so the model is traceable for the StableHLO → IREE mobile path.- The hand-rolled
Gemma3nRuntimewas never faithful to real checkpoints: it loaded the PLE tensors but never applied them, had no Laurel, ignored the AltUp router, and itsE2B_DEFAULTconfig claimed AltUp/sparsity were E4B-only — the real E2B GGUF hasaltup.num_inputs=4and first-10-layer sparsity. The GGUF CLI paths (kgemma, unified skainet-cli) now route gemma3n through the DSL lane; SafeTensors stays on the legacy runtime until the DSL grows that leg. Gemma3nGoldenTokenParityTest(#346 gate, the last ungated generative family): full 32-step greedy text equality vs mainline llama.cpp b10621 ongemma-3n-E2B-it-Q4_K_M.gguf, on the exact CLI path — engine loading stays packed/MAPPED (the PLE table row-dequants on demand). Wired into the smoke-reference tier (gemma3n_gguf_url+ 20g heap arg);smoke-models.jsongains a Gemma3n-E2B row. Metadata parsing now reads the real llama.cpp GGUF keys (sliding_window_patternbooleans, per-layeractivation_sparsity_scale,rope.freq_basefallback,rms_norm_eps, per-layerfeed_forward_length).
QwenChatTemplaterewritten against the official Qwen3chat_template(verified againstQwen/Qwen3-0.6B), fixing the drift that made small checkpoints unreliable in agent loops: tool results now render asuserturns wrapped in<tool_response>(consecutive results merged into one turn) instead of a literaltoolrole Qwen was never trained on; tools are listed one JSON object per line inside<tools>; the hardcoded "You are Qwen…" persona is gone (the caller's own system message leads the tools block); past assistant tool calls replay as<tool_call>blocks rebuilt from the structuredtoolCalls(raw XML in content is de-duplicated). Thinking mode is now handled:<think>…</think>blocks are surfaced viaAgentListener.onThinkingand stripped from the visible answer and the history (unterminated blocks included), andQwenChatTemplate(enableThinking = false)reproduces the officialenable_thinking=falseempty-<think>-prefill. Verified end-to-end on Qwen3-0.6B Q8_0 through kllama-cli--demo: calculator and file-listing round-trips both produce correct, thinking-free final answers. New antora tutorialtutorials/qwen-tool-calling.adocshows the whole flow embedded in your own app.
- The gate retrofit found the family broken in production — the CLI decoded real
Apertus-8B-Instruct GGUFs to
<unk>noise. Two defects, both fixed: (1)XIELUActivationusedexp()as a "simplified softplus approx" — with the real model's per-layeralpha_pvalues (up to 174),exp(alpha_p)isInfand the branch-mask multiply turned every logit intoNaN; even for small alphas the math was never faithful. The op now computes the exact guardedsoftplushost-side (the params are frozen scalars) and mirrors the reference formula, including themin(x, eps)clamp. (2)apertusNetwork()built RoPE on the DSL defaults (INTERLEAVED pairing, base 10_000) — Apertus is an HF rotate-half model that llama.cpp runs as NEOX withrope_theta = 12M; the metadata already carriedropeThetabut the builder never passed it. NowRoPEMode.SPLIT_HALF+metadata.ropeTheta. ApertusGoldenTokenParityTestcloses the last "no parity probe" row among the shipped generative families: full 32-step greedy text equality vs mainline llama.cpp on Apertus-8B-Instruct-2509 Q4_K_S, on the DSL path the CLI ships (ApertusWeightLoader→ApertusNetworkLoader.fromWeights→OptimizedLLMRuntime), exercising QK-norm, the per-layer xIELU activation parameters and the ungated FFN end-to-end. Model-gated onAPERTUS_GGUF_PATH(+ ≥8 GB test heap), tagged into the smoke-reference tier (apertus_gguf_urlstaging input), andsmoke-models.jsongains an Apertus row on the skainet-cli runner. README and the antora index now carry a per-family verified-against matrix reflecting the actual gates instead of the pre-0.52.0 "early / not verified" wording, andreference/architecture.adocis rewritten around the DSL-centric decoder core and the real module inventory (#346 template,sk.ainet.lang.nn.dsl.decoder).
QwenWeightLoaderjoins the family (<F>WeightLoaderrow): the thin wrapper overllm-core'sDecoderGgufWeightLoaderpinned to the publicQWEN_ARCHITECTURES(qwen2/qwen3/qwen35), same shape asLlamaWeightLoader/BitNetWeightLoader.QwenNetworkLoader's GGUF paths now delegate to it.QwenGgufTensorNames→QwenTensorNames(theQwenGgufWeightSource.ktnaming drift #346 called out);@Deprecatedtypealias remains for one release. The Qwen golden-token parity gates are taggedsmoke-referenceand wired into the reference workflow (qwen25_gguf_urlinput stages the 0.5B model; the existing qwen3 stage also feedsQWEN3_17B_GGUF), andtests/smoke/smoke-models.jsongains a Qwen2.5-0.5B-Instruct row certifying the #352 fix on the CLI path.
- Attention projection biases now load and bind
(#352): Qwen2/2.5
GGUFs carry
blk.N.attn_{q,k,v}.biastensors — and they are enormous (blk.0's K bias moves the channel sum from ~11 to ~507 on Qwen2.5-0.5B), so losing them turns the output into noise. They were lost twice over:DecoderGgufWeightLoaderfiltered every tensor outside its.weight-only wanted-set, and the llama name resolver had no.biasrules, so zero-initialized DSL params silently stood in. The loader now carries the attention biases as optional tensors, the newQwenGGUFNameResolverbinds them, andQwenNetworkLoaderfails loudly if a file bias ever goes unbound again. Verified token-for-token against mainline llama.cpp: the newQwenGoldenTokenParityTestasserts full 32-step greedy text equality for both Qwen2.5-0.5B-Instruct Q8_0 (bias + no QK-norm) and Qwen3-1.7B Q8_0 (QK-norm + no bias), model-gated onQWEN25_05B_GGUF/QWEN3_17B_GGUF.
sk.ainet.models.llamano longer owns the shared decoder loader half (#372):DecoderGgufWeightLoader,DecoderGgufWeights,DECODER_DEQUANTIZE_ALL,decoderMetadataFromGguf,DECODER_NARROW_KEEP_NATIVEand the family-neutral half ofDecoderSafeTensorsLoadernow live inllm-coreundersk.ainet.lang.nn.dsl.decoder, joining the architecture half (DecoderModelMetadata,decoderTransformerNetwork) that was already there. Renamed on arrival:LlamaModelMetadata→GgufDecoderMetadata,LlamaTensorNames→DecoderTensorNames,LlamaGgufTensorNames→DecoderGgufTensorNames— the shared decoder types no longer carry a family's name.@Deprecatedtypealiases remain insk.ainet.models.llamafor one release; the llama-typedDecoderSafeTensorsLoader.load()stays llama-side as an extension (import sk.ainet.models.llama.load).
Targets SKaiNET engine 0.51.0 (developed against 0.51.0-SNAPSHOT; the pin
flips to the release when the engine cuts it). The headline is the completed
engine-adoption arc (#338–#346): this repository no longer carries any weight
quantization, packing, or memory-staging machinery of its own — every family
loads through the engine's StreamingGgufParametersLoader with a declared
WeightForm, and MAPPED residency is the default everywhere.
- Every GGUF weight loader is a thin engine wrapper (#338–#341):
DecoderGgufWeightLoader(llama/qwen/mistral/smollm2),ApertusWeightLoader,Gemma4WeightLoader,Gemma3nWeightLoaderrewritten aroundStreamingGgufParametersLoader. The vendoredQuantPolicyenum is deleted (#342); loaders take an optionalWeightForminstead (null= keep-packed[out, in]MAPPED;DECODER_DEQUANTIZE_ALL/GEMMA_DEQUANTIZE_ALLfor the dense-FP32 export lane). - MAPPED residency by default: quantized weights are served zero-copy from
file-backed pages by the engine's row-major kernel packs
(
FfmRowMajorKernelPackon JVM,JniMappedKernelPackon Android) and are not charged against the managed heap. Measured: qwen2.5-1.5B (1.0 GB) runs under-Xmx512mwith ~176 MB of planned heap. - Token embeddings stay packed: a packed
token_embdis rewrapped asPackedRowDequantTensorData(new, transformer-core) —Embeddingdequantizes only the rows a step gathers, and a tied lm_head still rides the packed matmul chain (SmolLM2 CLI: ~3.6 → ~50 tok/s). linearProjectis one expression —ops.matmulWeightTransposed(input, weight); thePreTransposedWeightmarker and its branch are gone.- Per-step forward scope (#343):
OptimizedLLMRuntimeruns DIRECT decode inside an engineForwardScope— steady-state decode allocates zero new heap bytes per token; KV caches detach their kept history to ambient storage. Bit-identity with the unscoped path is pinned byForwardScopeSteadyStateTest. - Memory plan in the CLI: skainet-cli prints
MemoryPlans.plan(...)from the GGUF header before loading (warns when the plan exceeds the heap cap) and gained--explain-loadfor per-weight placement decisions. The hand-managedArena+MemorySegmentTensorDataFactorycontext setup is gone. - Qwen2/2.5 attention biases: the decoder lane now resolves
attn_{q,k,v}.biastensors andqwenNetworkgrowsattnBias(auto-detected from the checkpoint). Qwen2.5 end-to-end quality is still tracked in #352. - Family template (#346): shared
decoderMetadataFromGgufparser; the deprecatedGraphAccelerator/FusedQKVAcceleratorseam is deleted.
- BitNet b1.58 family (#336, #337):
llm-inference/bitnet—bitnetNetwork()(squared-ReLU FFN +ffn_sub_norm/attn_sub_normvia new DSL extension points),BitNetPackedGgufLoader(packed I2_S through the engine loader: ternary projections as 2-bitBITNET_B1_58, lm_head requantized toBITNET_PLANES, two-stage exact decode), CLI auto-detection. BitNet-2B4T decodes coherently through the unified CLI.
QuantPolicy,PreTransposedWeight+ wrappers,BlockQuantPacking,GgmlQuantEncodings(expect/actual),GemmaMemSegConverter,GemmaQuantLayout,GemmaPackedWeights,DecoderGgufMemSegConverter,MemSegWeightConverter,MmapLlamaLoader,QuantizedTensorFactory,LlamaPackedWeights,LlamaQuantLayout,ApertusMemSegConverter,QuantizedTensor,GraphAccelerator,FusedQKVAccelerator,GcHint— the entire pre-engine quant/packing/staging layer (net ≈ −7,000 lines).
Republishes 0.40.1's content — ships against the same SKaiNET engine 0.40.1 — after the 0.40.1 Maven Central publish broke partway through a multi-module release. No functional changes beyond the fix below.
- Broken 0.40.1 release-workflow publish (#313):
:llm-inference:smollm2declaredlinuxX64()/linuxArm64()Kotlin/Native targets with no source to back them (jvmMain-only export tooling, nocommonMain), socompileKotlinLinuxArm64reportedNO-SOURCEand produced no.klib— but the maven-publish plugin still registered a publication for the target, andgenerateMetadataFileForLinuxArm64Publicationunconditionally tried to hash the (nonexistent) klib file, throwingFileNotFoundExceptionand aborting the tag-triggered./gradlew publishpartway through the module graph. By that pointllm-api,llm-agent,llm-core,transformer-core,llm-bom,llm-performance,llm-providers,llm-inference:{apertus,bert,functiongemma,gemma,llama,moonshine,qwen}, and smollm2's own JVM publication had already published;llm-inference:{t5,vec2text,voxtral,whisper}and all ofllm-runtime:*never got attempted. Fixed by dropping the two unused target declarations — every other multiplatform module was audited for the same declared-target-vs-actual-source mismatch and none had it. 0.40.1 is superseded — use 0.40.2.
[0.40.1] — 2026-08-12 — superseded by 0.40.2, Maven Central publish broke partway through, do not use
Ships against SKaiNET engine 0.40.1. The headline is architectural rather than a
single feature: the tool-calling epic's shared substrate lands in three stacked PRs,
every packed-quant format gets a single hoisted packer with a pre-transposed-by-default
fast path now that native Q5 kernels shipped, FunctionGemma and SmolLM2 each get a
standalone DSL→StableHLO→IREE export module, and a generic Android JNI runtime serves
the compiled path the way skainet-backend-jni-cpu already serves the eager one. Also
fixes a real packed-quant matmul-corruption regression that the 0.40.0→0.40.1 engine
pin exposed on the classic (non-pre-transposed) path.
- Tool-calling epic substrate (#35), landed in three stacked PRs:
generateUntilStoppromoted tollm-core, demo/agent CLI extracted out ofkllamaintollm-agent(#296, closes #37/#49-P1); HF-side chat-template auto-detection fromtokenizer_config.json/chat_template.json/config.jsonplus registerable parser strategies (#297, closes #38/#40);AgentCliresolution diagnostics, detection/diagnostics test coverage, and validation against a real Qwen instruct GGUF (#299, closes #41/#42/#43/#44). - Packed Q5_0/Q5_1 converter path + shared block packer (#294, closes #170,
implements #184 items 2–3):
GemmaMemSegConverterkeeps Q5_1/Q5_0 weights packed instead of falling back to FP32 dequant, gated onhasPackedMatmulKernel()rather than engine version (81 of FunctionGemma-270M's 236 tensors are Q5_1). The GGUF-block → engine-tensor packing that gemma/llama/apertus each carried privately is hoisted intosk.ainet.lang.nn.quant.BlockQuantPacking(transformer-core), andPreTransposedWeightmarks weights already in kernel-feed layout solinearProjectcan skipops.transposeentirely. - FunctionGemma extracted into a standalone module (#302):
:llm-inference:functiongemmaowns the function-calling export/contract (FunctionGemmaSpec,FunctionGemmaContract,FunctionGemmaExportHarness), moved verbatim (byte-identical, sha256-verified) from:llm-runtime:kgemma, which keeps@Deprecateddelegating shims.:llm-runtime:gemma-ireegains manifest-driven support (GemmaManifest,GemmaKvDecoder.fromManifest,CompactToolCodec.fromManifest). - SmolLM2 compiled-export path (#305 epic): host-side StableHLO export for
SmolLM2-135M-Instruct following the gemma export pattern — no export existed for the
llama architecture before this (#306); a standalone
:llm-inference:smollm2module with the redecode-graph + DSL argMax tail, numerically verified end-to-end viairee-compile/iree-run-module(#308). - Generic Android JNI runtime for the compiled path (#309):
:llm-runtime:iree-android, the compiled-path counterpart to the engine'sskainet-backend-jni-cpu— binds external weights from a.irpa, invokes a named compiled function, model-agnostic. Botharm64-v8a/armeabi-v7aABIs, both CPU and Vulkan HAL drivers built in. - Android Antora docs (#310): a getting-started tutorial for the eager path and an eager-vs-compiled explanation page, tying together the JNI eager backend and the new compiled runtime for a reader deciding which to use.
- SKaiNET engine 0.39.1 → 0.40.0 → 0.40.1 (#307, #311). 0.40.0 brings native
Q5_0/Q5_1 kernels, so the gemma/llama packed-weight converters flip to
packPreTransposedby default wherever a packed kernel is confirmed available (Q4_K/Q5_K/Q6_K/Q8_0 unconditionally; Q4_0/Q5_0/Q5_1 kernel-gated as before) — verified byte/token-identical greedy decode across the FP32 baseline, the JVM MemSeg path, and the Kotlin/Native board path on a real FunctionGemma-270M checkpoint. 0.40.1 is the packed-quant regression fix below.
- kllama native kernels now actually register on Linux Kotlin/Native (#301,
closes #300):
linuxX64/linuxArm64publish the cinterop-embedded native kernels, but nothing ever called the engine'sinstallNativeKernels()on those targets (K/N has noServiceLoader) — measured 0.6 → 2.06 tok/s (3.4×) on SmolLM2-135M Q8_0 once fixed. - Packed-quant classic-path matmul corruption under engine 0.40.1 (#311): engine
0.40.1's
ops.transposebecame a physical canonical→kernel-native block-grid permutation (closing engine #968), butBlockQuantPacking.pack()'s classic path was already eagerly relayouting bytes to kernel-native order at load time — so the weight got double-permuted at forward time, silently producing wrong matmul output on Q4_K/Q5_0/Q5_1 (not a crash).pack()now stores checkpoint bytes verbatim; the Apertus Q4_K/Q6_K converter (which carried its own inlined, un-migrated relayout) is switched onto the sharedpackPreTransposedpath. New parity-matrix, synthetic-Apertus, and llama quant-layout tests close the coverage gap that let this ship undetected. The pre-transposed production path — what gemma/llama/kgemma actually serve — was never affected. Root cause tracked upstream at SKaiNET#973: packed-quant byte order is an unwritten, contradictory contract across the engine and its converters.
Patch release against SKaiNET engine 0.39.1 — the gemma function-calling day: the compiled FunctionGemma path gets materially smaller, faster, and board-verified.
-
FunctionGemmaToolCallingSupport (#292): the compact functional-token format (
<tool_N>(…),CompactCodec) joins theToolCallingSupportarchitecture — parser strategy, byte-exact chat template, NATIVE-mode detection; the tool map is now injectable (CompactToolCodec), so consumers can extend the tool set without a library change. First concrete slice of the #35 generalization. -
Board-verified KV decode (#291):
GemmaKvDecoderis no longer a draft — K-first outputs, raw.binI/O, per-graph parameter archives (gemma-prefill.irpa/gemma-with-past.irpa; the shared-irpa contract was invalid due to per-trace external numbering), runbook shipped inllm-runtime/gemma-iree/docs/. SL2610: steady-state ~1740 ms/token, 2.1× the same-day re-decode baseline. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- Engine pin
skainet 0.39.0 → 0.39.1: picks up the engine's primitive FP32 fast paths for the eager CPU ops and the cachedDirectCpuExecutionContext.ops(engine #949) — the per-element overhead that dominated on-device decode (83% of SmolLM2-135M end-to-end on a Pixel 8a even with NEON matmul) is gone for every non-JVM target. - Tied embedding exported once (#290, closes #260):
Gemma4WeightLoaderaliasestoken_embdintooutput.weight; the FunctionGemma weight archive drops ~832 → ~512 MiB bf16 with a single262153x640global. - True-dynamic
with_pastexport is the default (#290, refs #248): engineDim.DYNAMICtracing replaces theSENTINEL_PAST=7919text rewrite (GEMMA_SENTINEL_PAST=1is the rollback); the g165 Torq-fork compiler accepts the dynamic-dim MLIR on-board. - Token embedding stays packed on the JVM eager path (#289, closes #178/#234):
the transformers-local
RowDequantSourcebecame a deprecated typealias to the engine type (fixing a latent gather mismatch), andGemmaMemSegConverterkeeps row-sliceabletoken_embdlayouts packed — ~0.49 GB less FP32 on Q8_0, decode byte-identical.
- Tool-calling architecture completion (#35 epic,
#49 Phase 1; stacked PRs #296/#297/W2c):
generateUntilStop/GenerateResultpromoted fromllm-agenttollm-core(typealias re-exports keep old imports working);ToolCallingDemo/AgentCli/ListFilesTool/CalculatorToolmoved from:llm-runtime:kllamato:llm-agentjvmMain (packages unchanged) so any runner gets chat/agent/demo modes without depending on kllama; new sharedModelMetadataExtraction(best-effort GGUF fields + HFtokenizer_config.json/chat_template.json/config.jsonparsing) drives provider auto-detection on both the GGUF and safetensors CLI paths (explicit--templatestill wins);ToolCallParser.registerStrategy(...)lets model families plug custom tool-call output formats into the default parser chain; demo and agent CLI now print provider/mode/reason resolution diagnostics; env-gated real-checkpoint validation for Qwen (QWEN_MODEL_PATH); README gains a native-vs-generic tool-calling compatibility matrix.
- Shared-KV cache variants trace correctly (#290, closes #194):
SharedPositionalKVCache/PaddedSharedPositionalKVCache/OwnerReadOnlyKVCacheno longer bake K=V=0 constants underembedConstantstracing (kvSharedLayers > 0, e.g. Gemma 4 E2B). llm-runtime/kllamanow registers the native-cinterop kernel provider on linuxX64/ linuxArm64.DirectCpuExecutionContexton Kotlin/Native registers only the scalar provider by default (noServiceLoaderon K/N, unlike JVM/Android), so every native target's packed-quant matmul ran scalar even though the engine'sskainet-backend-native-cpukernels were on the classpath.CpuBackendProviderand the cross-targetSmolLm2InferenceSpikenow call the engine'sinstallNativeKernels()once per context via a small per-target hook (#300). Measured: the linuxX64 spike goes from ~0.6 to 2.06 tok/s (3.4×) on SmolLM2-135M Q8_0, 44 tokens, identical output.macosArm64/iosArm64/iosSimulatorArm64stay no-op until the engine publishes those klibs (#298).
Ships against SKaiNET engine 0.39.0 — the engine release that answers the mobile field report behind engine issue #920: a JNI NEON kernel backend for Android, real random-access GGUF loading on Android, and fail-fast on unsupported quantization types. The transformers headline follows directly: Android apps using the runtime facades now decode with native NEON kernels out of the box (measured ~6.4× on SmolLM2-135M Q8_0, Pixel 8a) instead of silently falling back to scalar Kotlin. Also new: whisper-tiny authored end-to-end in the NN DSL, iOS artifacts for the runtime facades, and SmolLM2 tool-calling support.
-
Native NEON kernels on Android, out of the box. The
llm-runtime/kllamaandllm-runtime/kgemmaAndroid artifacts now carry engine 0.39.0'ssk.ainet.core:skainet-backend-jni-cpuAAR as aruntimeOnlydependency. The backend self-registers via ServiceLoader on ART and provides NEON kernels (with runtime dotprod dispatch) for Q8_0 / Q4_0 / Q4_K / Q5_K / Q6_K — measured ~6.4× decode-kernel throughput on SmolLM2-135M Q8_0 (Pixel 8a: ~24 tok/s vs ~3.8 scalar). Apps using the inference modules directly add the AAR themselves; excluding it opts back into pure Kotlin (#285). -
whisper-tiny — the full pipeline authored in the NN DSL (
llm-inference/whisper, artifactskainet-transformers-inference-whisper): encoder at a configurable short audio context, decoder with the fixed-masked-KV prefill/step split (KV cache as explicit graph I/O, host-computed additive f32 masks — SPIR-V-safe, noi1/select), weights streamed by HF name directly from the safetensors checkpoint (SafeTensorsWeightSource, tied embedding, no Python anywhere), and a jvmTest export harness emitting MLIR + mergedparams.irpamanifest.jsonfor IREE compilation. Verified: encoder cosine 0.9999921 vs the ONNX-pipeline golden, greedy tokens reproduce the reference German transcript exactly, and the compiled vmfbs decode correctly on-device (Pixel Tensor G3, Vulkan). Replaces the PyTorch→ONNX export scripts that previously fed skainet-whisper-android (PR #279).
-
SmolLM2 tool-calling support (
llm-agent/ kllama):SmolLMChatTemplate, the SmolLM tool-call parser strategy, andToolCallingSupportresolver registration, so SmolLM2-Instruct models drive the agent loop like the other supported families. Includes a 7-case parser test and a gated end-to-end smoke test (#272). -
Cross-target SmolLM2-135M inference spike (
llm-runtime/kllama, commonTest): one env-gated test (SMOLLM2_MODEL) that loads the Q8_0 GGUF viaLlamaNetworkLoader.fromGguf(NATIVE_OPTIMIZED)+OptimizedLLMRuntime(DIRECT)and prints load time, decode tok/s, and the generated text — identical source on JVM, linuxX64, and iOS simulator, so per-target numbers are directly comparable (the reproducer half of #272; measured JVM ~7.2 tok/s FFM, linuxX64 ~0.6 tok/s scalar K/N). -
iOS artifacts for the runtime facades.
llm-runtime/kllamaandllm-runtime/kgemmanow declareiosArm64+iosSimulatorArm64and publish the corresponding klibs. kllama'ssrc/iosMain(theregisterPlatformBackendsactual) predated the targets and was silently dead — these modules setkotlin.mpp.applyDefaultHierarchyTemplate=false, so theiosMainsource set is now wired by hand (iosMain → nativeMain, mirroringllm-core). All commonMain dependencies of both modules already published iOS. No CLI executables are declared for the Apple targets — consumers link the klib into their app. Closes #271. -
Supported-targets matrix in the README. A module-vs-target table (derived from each module's
build.gradle.kts) replaces the "where applicable" hand-wave, so which artifact runs on iOS / Android / Wasm is now documented rather than discoverable only by browsing Maven Central (#271). -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- SKaiNET engine 0.38.0 → 0.39.0 (#282).
Besides publishing the Android JNI backend above, the engine release brings the rest of the
mobile-field-report fixes to every transformers consumer:
createRandomAccessSourceis now real on Android (streaming GGUF load instead of materializing the whole file on the ART heap — the hard-OOM path is gone, engine #922), the streaming GGUF loader fails fast on unsupported tensor types instead of silently skipping them and loads Q4_0/Q5_0/Q5_1 packed (engine #919), a NEON Q4_0 kernel and cinterop-embedded kernel archives for Kotlin/Native consumers, and a round of tensor-storage API hygiene fixes.
- Gemma integration tests skip properly under JUnit 5:
org.junit.Assume(JUnit 4) swapped for JupiterAssumptionsin the three-PincludeIntegrationgemma tests, so a missing model now reports as skipped instead of failing the run; the stray JUnit 4 dependency is gone from gemma's jvmTest (#261). apiCheckgreen again on clean checkouts: refreshed the stale jvm binary-compatibility dumps forllm-agent,kllama, andtransformer-core(#275).
0.38.0 — 2026-07-31
Ships against SKaiNET engine 0.38.0, which adds first-class dynamic tensor shapes (Dim) plus
the narrow-float codec (Fp16DenseTensorData, FP16 matmul kernels, codec-driven dispatch — engine
PR #886). Two headlines: Moonshine v2 streaming ASR authored end-to-end in the SKaiNET NN DSL
(the last vendor-ONNX graph is gone) and narrow-float KEEP_NATIVE weights across the LLM loaders.
-
Moonshine v2 — the complete streaming pipeline in the NN DSL, self-compiled DSL → StableHLO → IREE with no vendor neural binaries (
skainet-transformers-inference-moonshine):- Audio frontend (
MoonshineV2Frontend): CMVN →asinhcompression → filterbank matmul → SiLU → two causalConv1d(k5,s2)— the last vendor-ONNX graph, now DSL-authored (bit-exact vsfrontend.onnx, cos > 0.999). - Encoder (position-free sliding-window local attention) and adapter (learned absolute positional embedding, pos-embed add only) bridging the position-free memory to the decoder.
- Decoder authored in the DSL, reusing the shared KV-cache decoder.
- Audio frontend (
-
True-dynamic KV-cache decode graphs.
MOONSHINE_V2_TRUE_DYNAMIC/GEMMA_TRUE_DYNAMICtrace the cache seq dim as a real dynamic extent (Dim.DYNAMIC), so one compiled vmfb serves every autoregressive position instead of a fixed-shape re-decode. Requires engine 0.38.0'sDim. -
Fixed-max-pad cross-attention mask for streaming decode: pad the encoder memory to a fixed MAX and mask the padding, so one prefill + one
with_pastpair serve any encoder length ≤ MAX while the self-cache stays dynamic (growing).transformer-core'sMultiHeadAttentiongains an optional trailingcrossMask(defaultnull→ byte-identical for existing callers). -
Gemma row-dequant of the packed
token_embdin the sharedEmbedding, cutting host memory at load. -
FP16 KEEP_NATIVE on the SafeTensors path.
DecoderSafeTensorsLoadergains the F16 arm that BF16 has had since 0.25.0: with aDTypePolicyadmitting FP16 (Require(FP16),Prefer(FP16), orOneOfcontaining FP16) it stops widening F16 tensors and wraps the on-disk 2-bytes-per-element buffer inFp16DenseTensorData. The arm was missing only because no such storage type existed.DefaultCpuOpsJvmmatchesNarrowFloatTensorDataand picks the kernel by codec, so an F16 checkpoint now stays near its on-disk footprint instead of inflating ~2× as FP32. Covers LLaMA, Qwen, and Voxtral, which share this loader. -
Narrow-float KEEP_NATIVE on the GGUF path —
DTypePolicyis honored there at all now.DecoderGgufWeightLoaderaccepts adtypePolicyand keeps F16 / BF16 source tensors packed instead of widening every one to FP32.LlamaNetworkLoader,QwenNetworkLoader, andVoxtralNetworkLoaderplumb the policy attached viawithDtypePolicydown into it; before this the GGUF branches constructed the loader without the policy and silently ignored it. This is the KEEP_NATIVE GGUF path the 0.25.0 notes parked, and it is what makesRequire(BF16)real on GGUF.The packed path mirrors the FP32 path's layout handling exactly: for rank 2 it swaps the shape to
[cols, rows]and moves no bytes. GGUF header dims are reversed relative to the logical row-major shape, so the "column-major → row-major" step is a reinterpretation, not a permutation (DequantOps.transposeColumnMajorToRowMajorreturns its input unchanged). An actual element transpose here would have handed the matmul kernel a silently transposed weight matrix. The result is genuinely zero-copy — the on-disk buffer becomes the storage. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
-
DTypePolicyValidationcapability model is per-format.validate(policy, loaderName, keepNative: Set<DType>)replaces the BF16-onlyallowBf16Require: Boolean(kept as a@Deprecatedoverload). A caller declares which narrow-float formats its chain actually hands through packed, and aRequirenaming one is accepted only by a chain that can honor it. The boolean could express neither "keeps FP16 but not BF16" nor the empty case.The two formats are tracked separately and never interchangeably:
Require(BF16)still widens F16 sources, and vice versa. Both are 2 bytes per element, so mis-tagging F16 bytes as BF16 decodes to plausible-looking garbage rather than throwing.DTypePolicyValidation .keepsNative(policy, native)is the single decision point both loader chains share, mirroring the engine'smapPolicyToNarrow/keepsNative. -
Require(FP16)is now accepted byLlamaNetworkLoader,QwenNetworkLoader, andVoxtralNetworkLoader(both GGUF and SafeTensors), andRequire(BF16)is now accepted on their GGUF paths. Both previously threw. -
Binary-breaking (source-compatible):
DecoderGgufWeightLoaderconstructors gain a trailingdtypePolicy: DTypePolicy = DTypePolicy.Any, which changes their JVM descriptors. Kotlin and Java callers compile unchanged; already-compiled callers must be rebuilt. Behaviour with the default is identical to before.
GemmaNetworkLoaderandApertusNetworkLoaderno longer accept aRequire(BF16)they ignore. Both have their own weight chains (Gemma4WeightLoader/Gemma4SafeTensorsWeightLoader,ApertusWeightLoader/ApertusSingleSafeTensorsLoader) which widen every narrow float to FP32 and have no KEEP_NATIVE path. Their SafeTensors entrypoints nevertheless passedallowBf16Require = true, soRequire(BF16)validated and was then silently disregarded at load — the exact failure the eager validator exists to prevent. They now declarekeepNative = emptySet()and reject it. Callers relying on the old acceptance must switch toPrefer(BF16)(a soft constraint, which still passes) until those chains grow a KEEP_NATIVE path.- Moonshine v2 encoder sliding-window off-by-one (was cos 0.991 vs ONNX); the v2 config is set to the real tiny-streaming dims; the adapter is pos-embed add only (no LayerNorm).
- kgemma heavy-trace test heap raised to 12 g, with honest skips.
0.36.1 — 2026-07-17
Patch on 0.36.0 (same SKaiNET engine 0.36.0). Two additions: BGE embedding models on the BERT DSL path (CLS pooling + retrieval prefixes), and beam search for the T5 decoder and the vec2text inversion loop. Both are additive — existing consumers are untouched, and the vec2text greedy path is unchanged when both beam widths are 1.
- BGE embedding models (
BAAI/bge-small-en-v1.5and siblings) run on the BERT DSL path:- CLS pooling.
BertPooling { MEAN, CLS }onBertEncoderRuntime/createBertEncoderRuntime; auto-detected from the sentence-transformers1_Pooling/config.json(absent file →MEAN, unsupported max/sqrt-len modes rejected loudly). Pooling stays outside the traced graph — OPTIMIZED mode and StableHLO export are unaffected. - Query/document asymmetry.
EmbeddingModelgainsembedQuery/embedDocument/embedDocuments(defaults delegate toembed— additive, existing consumers untouched).PrefixedEmbeddingModel+EmbeddingModelProfilesapply retrieval instruction prefixes per repo id (E5query:/passage:, BGE query instruction);fromHuggingFacewires them automatically,fromSafeTensorsaccepts explicitprefixes. - Integer checkpoint buffers no longer break loads. BGE-style snapshots persist an
I64
embeddings.position_idsbuffer; the interimFloatSafeTensorsLoaderskips non-float buffers (the index-free encoder never needs them). Drop when the engine's loader gains a tensor filter (SKaiNET#822). - Design + traceable plan:
docs/specs/embedding-model-coverage.md(E5 multilingual follows in Phase 2 — Unigram tokenizer).
- CLS pooling.
- Token-level beam search on the T5 decoder.
T5Runtime.generateBeam(memory, numBeams, maxLength, lengthPenalty)returns up tonumBeamssequences, best-first by length-normalized log-probability. It shares a newdecoderLastLogits()step with greedygenerate, and addslogSoftmaxplus linear top-k helpers. There is still no KV cache, so decode cost scales roughly linearly withnumBeams. - Sequence-level beam search across correction rounds.
Vec2TextInverter.invert(..., sequenceBeamWidth, tokenBeams)andinvertEmbedding()keepbeamWidthhypotheses between correction steps, ranked by cosine similarity to the target embedding — the oracle the beam exploits.InversionModel.invertBeam/CorrectorModel.correctBeamexpose the top-N candidates from each stage. - Verified end-to-end on real gtr-base weights: at one correction step, beam (sequence width 3,
token beams 3) improves cosine 0.765 → 0.818 over greedy on the round-trip test's example
sentence, with a visibly closer reconstruction. Covered by
Vec2TextRoundTripTest(invert_beamBeatsGreedy), which skips unlessVEC2TEXT_MODELS_DIRis set.
0.36.0 — 2026-07-12
Ships against SKaiNET engine 0.36.0. Headline: BERT is now completely defined on the DSL path — the legacy hand-coded eager stack is removed (BREAKING, see Removed), and sentence embeddings get a one-call factory with built-in Hugging Face Hub download. Also new: a T5 encoder-decoder runtime and a vec2text embedding-inversion pipeline (invert GTR embeddings back to text). Downstream impact: indexing the leaf-cli reference corpus (56 chunks) drops from 676.9 s to 44.5 s (~15×) with identical embeddings.
- BERT sentence embeddings completed on the DSL path.
bertNetwork()is now a numerically completetokens → hidden-statesencoder: the newBertEmbeddingsmodule adds absolute-position and token-type embeddings (index-freenarrow-based lookups, single-segment) that the DSL definition previously omitted. NewBertEncoderRuntimeexecutes it eagerly (DIRECT, default) or as a traced, optimized ComputeGraph (OPTIMIZED, shape-specialized per sequence length with an LRU cache) and adds masked mean pooling, the optional sentence-transformers2_Denseprojection, and L2 normalization on top of the pure encoder graph. The encoder trace lowers to StableHLO (gather / dot_general / SDPA preserved) — export is gate-tested; IREE execution of the exported module stays out of scope for now. Verified against the PyTorch-validated legacy runtime on real MongoDB/mdbr-leaf-mt (hidden-state parity ≤ 2.2e-6) and DIRECT-vs-OPTIMIZED bit-exact. - One-call embedding factory with built-in Hugging Face download.
BertEmbeddingModel.fromHuggingFace("MongoDB/mdbr-leaf-mt")(llm-providers) downloads the snapshot via the engine'sskainet-data-source(hf://URIs,HF_TOKEN-aware) into~/.cache/skainet/models/, streamed with.part+ atomic rename, offline-safe after the first run;fromSafeTensors(dir)loads a local snapshot, auto-detecting weights, config, tokenizer (vocab.txt→tokenizer.json), and the2_Dense/head.kbert-cliaccepts an HF repo id directly:kbert MongoDB/mdbr-leaf-mt "query" "doc". BertConfigParser— sharedconfig.json(+2_Dense/config.json→projectionDim) parser, consolidating the copies previously living inKBertJavaand downstream apps.- T5 encoder-decoder runtime (
llm-inference/t5,sk.ainet.models.t5). Hand-coded in the direct tensor-ops style (per-head attention via narrow/matmul/softmax, batch 1, no KV cache — the greedy decoder recomputes the stack per step), handling T5's specifics: no 1/√d attention scaling, learned relative-position bias (T5RelativeBias, block-0 table shared per stack, none in cross-attention), RMSNorm-style T5LayerNorm, un-gated ReLU FFN, tied embeddings withd_model^-0.5logit scaling. IncludesGtrEmbedder— GTR sentence embeddings exactly as vec2text consumes them (raw T5 encoder + mean pooling; deliberately no Dense projection and no L2 normalization) — with a parity test against realsentence-transformers/gtr-t5-baseweights. - vec2text embedding inversion (
llm-inference/vec2text,sk.ainet.models.vec2text). Port of vec2text's greedy corrector loop (sequence_beam_width = 1):InversionModelproduces an initial hypothesis from a target GTR embedding, thenCorrectorModeliteratively re-embeds and corrects it, early-stopping when the cosine score plateaus —Vec2TextInverterreturns the best reconstruction plus the full step trace. Verified with an end-to-end round-trip test on real gtr-base weights.
- BERT post-norm residual wiring. The single-block-per-layer
bertNetwork()definition wired the FFN residual to the pre-LayerNorm value — the transformer blocks' residual rule fits pre-norm decoder stacks, but BERT is post-norm. Each encoder layer is now two blocks (attn/ffn) so every residual segment starts at the correct value. - Bias-free
2_Denseprojection heads were silently dropped. The legacy eager runtime required projection weight and bias; LEAF models shipbias=false, so it skipped the projection entirely (returning 384-dim vectors while advertising 1024).BertEncoderRuntimeapplies bias-free projections;KBertJavanow picks up2_Dense/heads it previously ignored. - Graph replay dropped
permuteaxes. The ComputeGraph executor's builtin dispatch replayedpermuteas a plain last-two-dims transpose, breaking every multi-token attention trace — single-token decode never hit it. Fixed upstream in engine 0.36.0 (SKaiNET#803), which this release consumes; the interim axes-awarepermutehandler inLLMFusedOpHandlers(never in a published release) is removed again.
- BREAKING: the deprecated hand-coded BERT stack is gone —
BertRuntime,BertRuntimeWeights,BertLayerWeights,loadBertWeights,BertWeightMapper,BertTensorNames,BertIngestion, andBertNetworkLoader.fromRuntimeWeights. Migrate tocreateBertEncoderRuntime(config, tensors, ctx)(tensors fromBertNetworkLoader.loadWeightTensors) or, one level up, toBertEmbeddingModel.fromSafeTensors(...)/fromHuggingFace(...).SkaiNetEmbeddingModel's constructor now takesBertEncoderRuntime;KBertJava/KBertSessionkeep their method surface (loadSafeTensors/encode/similarity) with the constructor type changing.BertModelConfigandMDBR_LEAF_IR_CONFIGmoved toBertConfig.kt(same package — imports unaffected). Thedocs/optimizable-LLM-NNs-DAG.mdreference in the old deprecation pointed at a document that never existed; the real migration guide isexplanation/dsl-vs-handcoded.adoc.
Ships against SKaiNET engine 0.35.0, whose new argMax op this release uses to fold the LLM
logits → token-ids tail into the DSL trace.
-
FunctionGemma self-compile from the SKaiNET DSL (
sk.ainet.transformers:…-kgemma). One reusable dependency for the FunctionGemma-270M function-calling sLLM, in both SKaiNET execution modes:FunctionGemma.fromGguf(gguf).call("turn the light on")→ToolCall(set_lights, {state="on"})— eager (DirectCpu +OptimizedLLMRuntime(DIRECT)+ Octopus-v2 template +CompactCodec), runs anywhere on CPU, no iree. (ThepartialRotary = 1.0gemma3 rotary fix is applied.)FunctionGemma.exportCompiled(outDir)/FunctionGemmaExport.export(…)— compiled edge path: tracesgemmaNetwork()ending inops.argMax(logits, -1)(the engine op), emits StableHLO with bf16 external params (bf16 globals + convert-on-load + bf16 safetensors). Promotes the formerRealGemmaBakeIrpaTestand retires the Python argmax/f16 MLIR rewrites. Verified token-for-token against llama.cpp on the SL2610 board.exportFunctionGemmaGradle task (forscripts/compile-gemma.sh);kgemmajvm deps gainskainet-compile-hlo/-dag+gemma-iree(CompactCodec).
-
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- Engine → 0.35.0. Adopts the new engine line; the compiled FunctionGemma export depends on the
engine's new
argMaxop. (engine 0.35.0)
Patch on 0.34.0 (same SKaiNET engine 0.34.0). Fixes Moonshine encoder parameter naming.
- Layer-qualified Moonshine encoder parameter names. The encoder's attention and LayerNorm
parameters were not prefixed with the layer (
attn.q_proj.weight,attn_norm.weightrepeated identically every layer), while the FFN parameters were (enc.$layer.ffn_*). By-name weight loading could therefore not distinguish the layers. All parameter names are now unique and layer-qualified (enc.$layer.attn.*,enc.$layer.attn_norm.*,enc.$layer.ffn_norm.*), matching the FFN convention. No public API change —moonshineEncoder()is unchanged.
Ships against SKaiNET engine 0.34.0. Headline: the first Moonshine speech-to-text encoder authored entirely in the SKaiNET NN DSL, plus the RoPE work that makes transformer exports bit-exact on a real NPU.
-
skainet-transformers-inference-moonshine(new, first published module) — the Moonshine-tiny audio encoder built in the NN DSL, bf16-native, emitting portable (hardware-agnostic) StableHLO. It compiles through the SKaiNET pipeline and transcribes correctly on both CPU and the Synaptics Torq NPU. The exported IR carries no target-specific ops — backend optimizations plug in from outside core (see the vendor-plugin pattern). -
Partial rotary embeddings in
transformer-core:RoPEgainspartialRotaryFactor(rotate only the leading fraction of each head, the rest passes through) andfreqDenomRotaryDim(computeinv_freqover the rotary dim rather than the full head dim).TransformerDsl.rope()threads both. Matches models like Moonshine (rotate 32 of 36 head dims), verified against the reference ONNX. -
VoidDense(addBias = true)— a projection can now add its$name.biasterm, keeping traced FFNs faithful to reference checkpoints that carryfc1.bias/fc2.bias. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- RoPE precision & form. The interleaved rotation and its
cos/sintables are computed in f32 (upcast, then back to model dtype), and the interleaved path uses the full-head (ONNX) form — numerically identical to the split-recombine form but bit-exact once accelerator layout passes sit between the split and merge. Fixes low-precision RoPE drift on NPU targets. - Engine → 0.34.0. Transformer models inherit the engine's 0.34.0 work (f32 LayerNorm decomposition, the pluggable target-optimizer / op-granularity seam that keeps exported StableHLO portable).
Ships against SKaiNET engine 0.33.0. No transformers API changes — this release adopts the new engine line and routine dependency updates.
- On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- Engine → 0.33.0. Transformer models authored with this layer inherit the engine's 0.33.0 work;
most relevant here,
layerNorm/rmsNormnow lower to realstablehlo.reduce, so transformer exports compile and run on stock IREE (engine #769). The engine also fixes a silent autodiff gradient-drop (elu/leakyRelu/permute) and adds new differentiable ops (cos/sin/gather/…), available to model authors. (engine 0.33.0) - Dependencies: Ktor client
3.5.1(#198), Logback1.5.36(#199).
Fixes streaming detokenization — generated text no longer runs words together
("the process" → "theprocess"). Ships against engine 0.32.4.
-
Per-token streaming decode preserves word-boundary spaces.
SentencePieceSpecialTokens.decode(Int)andUpstreamTokenizerAdapter.decode(Int)now route through the engine's newTokenizer.decodeToken(id)(engine 0.32.4), which keeps each SentencePiece piece's leading space instead of stripping it per token (the sequence-leveladdSpacePrefixstrip is only correct once per sequence). Fixes correct-but-spaceless output in streaming generation (kllama, agent loops). AddsSentencePieceSpecialTokensStreamingTest. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- Engine pin
skainet 0.32.2 → 0.32.4(addsTokenizer.decodeToken).
Brings the real-GGUF Llama eager path up to the Gemma standard (packed
NATIVE_OPTIMIZED) and unblocks StableHLO/IREE export for Llama-family models
(traceable interleaved RoPE). Ships against engine 0.32.2.
-
Eager
NATIVE_OPTIMIZEDpacked path for Llama.LlamaNetworkLoader.fromGguf(NATIVE_OPTIMIZED)keepsQ4_K/Q6_Kweights packed and runs them throughOptimizedLLMRuntime— newLlamaQuantLayoutLlamaPackedWeights.convertLlamaWeightsPacked, mirroringconvertGemmaWeightsPacked. Coherent output matching llama.cpp; the low-footprint path real-GGUF Llama inference on constrained ARM was missing. (ccbd87e)
-
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- Fused decode-attention fast path.
MultiHeadAttention's decode step (seqQ == 1) now computes scores → softmax → GQA-weighted-V directly from the cached K/V, bypassing therepeatKVHeadsconcat and theunsqueeze → SDPA → squeeze → permutechain — ~1.5× decode throughput, bit-identical output. Prefill (seqLen > 1) keeps the general SDPA path. (3791f88) - Engine pin
skainet 0.31.0 → 0.32.2(0.32.2 is the first engine release exposingExecutionContext.isRecording, required by the trace-faithful KV-cache path).
- Packed token-embedding gather for Llama —
fromGguf(NATIVE_OPTIMIZED)no longer fails withgather: unsupported input rank 1; the packed embedding is wired through the canonical loader. (ccbd87e) - Interleaved RoPE is now traceable. In
INTERLEAVEDmode (Llama / Mistral / most GGUF) the rotation used a raw float-array path (copyToFloatArray/fromFloatArray) that, under graph tracing, baked the rotated Q/K as a disconnected constant — severing them from the projection weights and crashingiree-compile(null-deref in constant folding) on the exported graph.RoPEnow records the rotation as tensor ops when running under the tracing wrapper; eager execution keeps the byte-identical raw-array fast path. Unblocks Llama/Mistral/GGUF StableHLO/IREE export. (019b049)
Adds transformer-core — the framework NN primitives (attention, the KV-cache family, embedding,
norms, RoPE, SwiGLU/GeGLU FFN, residual, linear projection) extracted from llm-core so they build on the
full Kotlin target matrix including androidNative (32-bit + 64-bit ARM). llm-core re-exports it, so
existing consumers are unaffected; ARM-native downstreams (e.g. on-device whisper) can now reuse the
primitives instead of reimplementing them.
-
transformer-coremodule (sk.ainet.transformers:skainet-transformers-transformer-core) — the lang-core-only NN primitives, reusable on every target incl.androidNativeArm32/androidNativeArm64. Depends only onskainet-lang-core. Added to the BOM. (#183) -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
llm-corenowapi-depends ontransformer-coreand re-exports it (no behaviour change). The NN primitive sources moved out ofllm-coreintotransformer-core;dsl/decoder/*stayed (it needs the compile-opt-coupledHybridTransformerBlock).MultiHeadAttention's diagnosticdumpStatsis decoupled via a settablemhaStatSinkthatHybridTransformerBlockwires to llm-core's platformdumpStats.
- Engine pin unchanged (
skainet = 0.31.0).transformer-coreneeds nothing new from the engine (onlyskainet-lang-core, already in 0.31.0), so this patch ships against engine 0.31.0 — the one case the transformers-X.Y.Z↔ engine-X.Y.Zalignment is intentionally relaxed (additive + engine-independent).
0.31.0 — 2026-06-15
Version-aligned with SKaiNET 0.31.0. Completes the eager board-decode path
for FunctionGemma: the tied Q8_0 lm_head now stays packed (paired with the
engine's ops.transpose fix for all packed dtypes), and load() can cap the
context to fit constrained devices.
-
maxInferenceLenonGemmaNetworkLoader.load()— an optional cap on the context length the eager network sizes its KV cache + RoPE tables for (defaultmin(contextLength, 4096), threaded throughapplyWeightsToNetwork→gemmaNetwork). A constrained-device consumer (e.g. the 1.9 GB SL2610 board) can pass a small value (e.g.32for a short tool-call prompt) to shrink the KV cache ~100×, which otherwise allocates ~0.4 GB at the first forward and OOMs the board after the weights load. Defaultnullpreserves existing behaviour. (#180) -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
gradle/libs.versions.tomlskainetpin: 0.30.0 → 0.31.0. Picks up the engine'sops.transposelazy-rewrap fix for all packed matmul dtypes (Q8_0/Q4_0 added) — required so the packed Q8_0 lm_head below transposes throughlinearProjectinstead of throwingClassCastException. Downstream consumers get the upstream SKaiNET BOM transparently via:llm-bom.gradle.propertiesVERSION_NAME=0.31.0. Lock-step with the engine.com.networknt:json-schema-validator→ 3.0.4. (#175)
- Tied Q8_0 lm_head stays packed in the eager
NATIVE_OPTIMIZEDGemma path. FunctionGemma'stoken_embdis Q8_0 and tied, soconvertGemmaWeightsPackedwas dequantizing bothtoken_embdandoutputto FP32 (2×~0.67 GB) — OOM on the 1.9 GB SL2610.output/lm_head now packs as Q8_0 (packGemmaKQuantgained a Q8_0 case; the row-major→block-major relayout is generalized with ablockSizeparam) and runs on the (NEON) Q8_0 kernel;token_embdstays FP32 (it is gathered, not matmul'd) but is wrapped no-copy viaDenseFloatArrayTensorDatainstead ofctx.fromFloatArray(which allocated a second ~0.67 GB buffer). Tied embed/lm_head footprint ~1.34 GB → ~0.76 GB. Verified byte-identical decode parity (GemmaQ5KPackedParityTest) and a stable ~1.06 GB load on the SL2610. (#179)
0.30.0 — 2026-06-14
Version-aligned with SKaiNET 0.30.0. Skips 0.29.x — SKaiNET-transformers
tracked the engine internally across that window (the in-progress Q5_K kernel
shipped as a local 0.29.1) without a tagged release. The headline is
Q5_K stays packed in the eager Gemma runtime and the Gemma
NATIVE_OPTIMIZED packed-weight path is now Kotlin/Native–ready — the board
binary can keep K-quant weights packed without the JVM's java.lang.foreign
MemSeg path.
-
Q5_K packed in-kernel dequant in the eager Gemma runtime. FunctionGemma-270M ships as
Q5_K_M, butGemmaMemSegConverterpreviously dequantized Q5_K weights to FP32 on load ("no native matmul kernel yet for Q5_K"), giving up both the memory saving and the in-kernel dequant. SKaiNET 0.30.0 provides a first-class Q5_K packed matmul (Q5_KBlockTensorData+Q5KMatmulKernel: scalar / Panama / native), so the converter now relayouts the GGUF bytes to block-major and wraps them asQ5_KBlockTensorData(176 B/block). Dispatch and the lazy transpose reach the kernel throughDefaultCpuOps. Verified byGemmaQ5KPackedParityTest(-PincludeIntegration): the Q5_K packed path decodes FunctionGemma byte-identically to the FP32 baseline —[262146, 236769, 3255, 718, 498, 1373, 262152, 106]→<tool_0>(state="on")<end>for "Turn the light on." -
Kotlin/Native–ready Gemma packed-weight path. The
NATIVE_OPTIMIZEDpacked conversion wasjvmMain-only (it builtMemSeg/Arena-backed tensors viajava.lang.foreign), so the Kotlin/Native board binary couldn't keep K-quant weights packed. The platform-neutral pieces now live incommonMain:GemmaQuantLayout.kt(commonMain) —logicalShapeFor,relayoutKSeriesRowMajorToBlockMajor(KMP-safecopyInto), andpackGemmaKQuant<T>(), which builds heap-packed Q4_K/Q5_K/Q6_KBlockTensorDatadirectly with noMemSeg/Arena.GemmaPackedWeights.kt(commonMain) —convertGemmaWeightsPackedpacks Q4/Q5/Q6_K matmul weights to heapQ*_KBlockTensorData, dequantstoken_embd/outputto FP32 (gathered, no transpose) and any other quant type to FP32[out, in].extractRawBytesreads the loader's bytes back across both backings (JVMIntArrayTensorData/ nativeByte-typed).GemmaNetworkLoader.load()now runsconvertGemmaWeightsPackedbeforeapplyWeightsToNetworkunderNATIVE_OPTIMIZED, soload(NATIVE_OPTIMIZED)yields a runnable network on the board and the JVM (previously it could not be built from raw-byte weights at all).GemmaMemSegConverter(jvmMain) now shares thecommonMainhelpers; only theMemSeg/FFM conversion and the FP32 fallbacks stay JVM-only. Verified on JVM andlinuxX64(GemmaQuantLayoutTest): relayout, packing, and the native byte-extraction round-trip run on every target, andGemmaQ5KPackedParityTestconfirms all three paths (FP32 baseline,jvmMainMemSeg-packed,load()packed) produce the identical token sequence.
-
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
gradle/libs.versions.tomlskainetpin: 0.28.1 → 0.30.0. Picks up the released Q5_K packed matmul, the NEON native kernels, and the Kotlin/Native cinterop. Downstream consumers get the upstream SKaiNET BOM transparently via:llm-bom, so no per-consumer migration is needed.gradle.propertiesVERSION_NAME=0.30.0. Lock-step with the engine.settings.gradle.ktsreverts themavenLocal()-first dev shim. The ordering added while consuming the in-progress local SKaiNET0.29.1is no longer needed now that 0.30.0 is on Maven Central; the release resolves the engine purely from Central. The opt-in-PuseLocalSkainetcomposite build is unchanged for local engine work.
fix(gemma): dequant kernel-less quant types inNATIVE_OPTIMIZEDinstead of leaving raw bytes. Loading a Gemma GGUF whose attention/FFN weights used a quant type with no packed SIMD kernel (e.g. Q5_1) underQuantPolicy.NATIVE_OPTIMIZEDcrashed at the first decode step (Transpose requires at least 2 dimensionsinMultiHeadAttention→linearProject):GemmaMemSegConverter.convertOneleft every unhandled quant type as raw 1-D bytes. Kernel-less types now dequantize to a correct FP32[out, in]weight via a newdequantPackedToFp32helper (mirroring the provenGemma4WeightLoader.createTensorcolumn-major → row-major transpose). The supported packed types (Q4_0/Q8_0/Q4_K/Q6_K) keep their fast SIMD form; only kernel-less types pay the FP32 dequant.fix(llama): dequantize Q4_1 (and all non-packed quant types) inDecoderGgufMemSegConverter``. The converter handled only Q4_0/Q8_0 (packed) and Q4_K/Q5_K/Q6_K (dequant); every other quant type fell through anelsebranch that logged a warning and passed the raw quant bytes through unchanged, crashing deep inside matmul (e.g. `unsupported quant type Q4_1 for blk.0.ffn_down.weight` on Q4_1 Qwen3 models). The `else` branch now routes through `DequantOps.dequantFromBytes` to FP32, covering Q4_1, Q5_0, Q5_1, Q8_1, IQ4_NL/XS, TQ1/2_0, etc.; genuinely unknown types now fail explicitly at load time instead of crashing later inside matmul. Closes #654.
GemmaQ5KPackedParityTest— byte-identical decode parity across the FP32 baseline, thejvmMainMemSeg-packed path, and theload(NATIVE_OPTIMIZED)commonMainpacked path.GemmaQuantLayoutTest(commonTest) — block-transpose relayout, packing, and the byte-extraction round-trip; runs on JVM andlinuxX64.DecoderGgufMemSegConverterTest— regression that a Q4_1 weight is dequantized to its logical 2-D FP32 shape rather than passed through as 1-D bytes.fix(gemma): macosArm64 target forgemma-iree`` and CI parity fixes: MLIR-dump tests write to a portable build dir instead of a hardcoded local path; browser Mocha gets a 60 s timeout (parity with the engine repo).test(gemma): repoint stale FunctionGemma GGUF path— six real-model integration tests now point at the in-reposl2610-function-calling/models/location, matchingGemmaQ5KPackedParityTest; all pass against the published SKaiNET 0.30.0 (-PincludeIntegration).
0.28.1 — 2026-06-06
Version-aligned with SKaiNET 0.28.1. Skips 0.26.x / 0.27.x — SKaiNET-transformers tracked the engine internally across that window without a tagged release.
- On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
gradle/libs.versions.tomlskainetpin: 0.27.0 → 0.28.1. Picks up the completed Kotlin DSL → StableHLO → IREE export path. SKaiNET 0.28.0/0.28.1 closed the remaining DAG-DSL export bugs: shape-changing ops now declare their inferred output type instead of echoing operand-0 —reshape/matmul/concatenate(SKaiNET #673) andconv1d/gather/maxpool2d/avgpool2d/flatten(SKaiNET #675) — andreduce_windowis emitted in IREE's generic region form. A full gemma3 graph traced throughGemmaMlirDumpTest/GemmaTraceTestnow lowers to StableHLO thatiree-compiles to avmfb. No transformers-side API changes; existing callers compile unchanged.
:llm-inference:gemma:jvmTestgreen against the published SKaiNET 0.28.1 (GemmaMlirDumpTest1/1,GemmaTraceTest1/1).
Version-aligned with SKaiNET 0.25.0. Skips 0.24.x — SKaiNET-transformers has been on 0.23.4 since 2026-05-08; the engine bumped 0.23.1 → 0.25.0 in the same window without a tagged 0.24.x release on either side.
-
DTypePolicyaccepted on every*NetworkLoader.fromGguf/.fromSafeTensorsentrypoint. SKaiNET 0.25.0 introduced the hybrid adaptive DSL with optional dtype constraints RFC — a sealedDTypePolicytype (Any | Require | Prefer | OneOf) carrying execution-side dtype intent through the loader / DAG / resolution pipeline.LlamaNetworkLoader,QwenNetworkLoader,GemmaNetworkLoader,ApertusNetworkLoader, andVoxtralNetworkLoadernow each acceptdtypePolicy: DTypePolicy = DTypePolicy.Anyon every public companion factory. The policy is eagerly validated against the loader's actual output dtypes at construction time (via the newsk.ainet.apps.llm.DTypePolicyValidationhelper), matching the SKaiNET 0.25.0StreamingGgufParametersLoader.validatePolicy()/SafeTensorsParametersLoader.mapPolicyToBf16()semantics:- GGUF entrypoints accept
Any/Prefer/OneOf/Require(FP32)and rejectRequire(BF16)/Require(FP16)/Require(other)with the same error messages as SKaiNET's own GGUF loader. - SafeTensors entrypoints additionally accept
Require(BF16)(matching theKEEP_NATIVEprecedent thatBf16LoadPolicy.toDTypePolicy()is built on upstream). - All entrypoints fall through with no behavioural change on the default
Anyvalue, so the bump is fully back-compat.
- GGUF entrypoints accept
-
decoderTransformerNetwork(dtypePolicy = …)parameter on the shared decoder-only builder inllm-core— declarative slot for the top-level block policy. Forward-compat surface; not yet propagated into the underlyingDagBuilder.op(..., dtypePolicy = …)slot SKaiNET 0.25.0 introduced (HybridTransformerBlock.compile()will read this in a follow-up). Setting a non-Anyvalue compiles today and starts taking effect when the compile-step plumbing lands — no API change at consumers. -
SafeTensors BF16 KEEP_NATIVE in
DecoderSafeTensorsLoader. When the consumer attaches aDTypePolicythat admits BF16 (Require(BF16),Prefer(BF16), orOneOfcontaining BF16), the loader stops dequanting BF16 tensors and instead wraps the packed 2-bytes-per-element buffer inBf16DenseTensorData. The matmul dispatch inDefaultCpuOpsJvm(SKaiNET 0.25.0) detectsBf16TensorDataat runtime and routes to the SIMD BF16 kernel — so a BF16 SafeTensors checkpoint now stays near its on-disk footprint in RAM instead of inflating ~2× to FP32. Threaded throughLlamaNetworkLoader/QwenNetworkLoader/VoxtralNetworkLoader(each forwardsloader.dtypePolicyinto theDecoderSafeTensorsLoader<T>(ctx, T::class, metadata, tied, dtypePolicy)constructor). The default value remainsDTypePolicy.Any— adaptive FP32 dequant, no behavioural change for existing callers. Validation errors still fire at theLlamaNetworkLoader.withDtypePolicy(...)boundary:LlamaNetworkLoaderDTypePolicyTestpins each policy arm. -
Three reference smoke tests with
@Tag("smoke-reference"). The new smoke tier exists alongside the existing@Tag("integration")filter and pins the three architectures we always want to run end-to-end:llm-runtime/kllama—Qwen3ReferenceSmokeTest(Qwen3-1.7B Q8_0 GGUF; exercises the new SKaiNET 0.25.0Q8_0MatmulKernelend-to-end + Qwen'sRoPEMode.SPLIT_HALF+ QK-Norm).llm-runtime/kgemma—Gemma4ReferenceSmokeTest(Gemma-4 E2B SafeTensors; sliding-window attention + per-layer KV sharing).llm-test/llm-test-java—BertLeafReferenceSmokeTest(MongoDBmdbr-leaf-irSafeTensors via the JavaKBertJavaconsumer surface, with a cosine-similarity sanity check on paraphrase embeddings). Run with./gradlew test -PsmokeReference -PincludeIntegration. Each test self-skips via JUnitAssumptions.assumeTruewhen the model artifact isn't resolvable through the standard~/.lmstudio/models//~/.cache/huggingface/hub// env-var fallback chain, so CI without model files stays green.
-
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
gradle/libs.versions.tomlskainet → 0.25.0. Downstream consumers already get the upstream SKaiNET BOM transparently via:llm-bom(api(platform("sk.ainet:skainet-bom:${libs.versions.skainet.get()}")), unchanged since 0.23.4 when the BOM auto-discovery convention plugin landed) — no per-consumer migration needed.gradle.propertiesVERSION_NAME=0.25.0. Lock-step with the engine.tasks.withType<Test>().configureEach { ... }at the root build now honors a-PsmokeReferenceproject property — symmetric to the existing-PincludeIntegration. When set, JUnit Platform is filtered to@Tag("smoke-reference")so the smoke tier runs in isolation (./gradlew test -PsmokeReference -PincludeIntegration).tests/smoke/smoke-models.jsongains a"reference": trueflag on the three reference entries (Qwen3-1.7B-Q8,Gemma4-E4B-GGUF,MongoDB-mdbr-leaf-ir) so the shell smoke harness and the JVM smoke tier point at the same artifacts. Thesmoke-test.shscript does not yet consume the flag — follow-up.smoke-referenceGitHub Actions workflow. New.github/workflows/smoke-reference.ymltriggers the three@Tag("smoke-reference")tests via./gradlew test -PsmokeReference -PincludeIntegration.workflow_dispatch-only (manual) with three optional URL inputs — supply each artifact URL via the dispatch form and the staging steps download it intoRUNNER_TEMP, set the env var the test reads (QWEN3_1B7_MODEL_PATH/GEMMA4_E2B_SAFETENSORS_PATH/LEAF_MODEL_DIR), and the smoke tier actually exercises the models. Run with empty inputs and every test self-skips via JUnitAssumptions— the workflow is green either way, so it's safe to promote topush: branches: [develop]later once a self-hosted runner with pre-cached checkpoints is available.- Catalog goes BOM-only. Every
skainet-*alias ingradle/libs.versions.tomlis now coordinate-only (noversion.ref); versions are supplied by thesk.ainet:skainet-bomplatform constraint re-exported by:llm-bom. Every consumer module gainsimplementation(project.dependencies.platform(project(":llm-bom")))in each source set that pulls askainet-*artifact. Bumping the engine is still a one-line change at the top of the catalog (the[versions] skainet = "X.Y.Z"line drives the BOM platform reference inllm-bom/build.gradle.kts), but every internal build now exercises the BOM — so a BOM-coverage regression fails locally instead of leaking into a published artifact. Mirrors thellm-test/llm-test-javareference pattern that landed in 0.23.4.
These pieces of the dtype-policy RFC integration are intentionally not in this release. The threading surface accepts the API so consumers can compile against the eventual implementation; the actual behavioural changes land in follow-up PRs.
- Per-DSL-layer dtype-policy parameters on
TransformerDsl.ktfactories (embedding/rmsNorm/multiHeadAttention/swiGluFFN/geGluFFN/xielu). The DSL is module-based and would need aModule-level metadata side-map to carry the policy down to compile time; landing that without a consumer that reads it would add maintenance surface for no behavioural value today. HybridTransformerBlock.compile()honoring the policy onDagBuilder.op(..., dtypePolicy = …)per the W6 SKaiNET PR. Blocked on the side-map above.DecoderGgufWeightLoaderper-tensor policy enforcement. The GGUF loader still dequants BF16 → FP32 unconditionally — SKaiNET 0.25.0'sStreamingGgufParametersLoader.validatePolicy()itself rejectsRequire(BF16)for GGUF today (no KEEP_NATIVE GGUF backing yet), so this is parked until the engine grows that path. (SafeTensors BF16 KEEP_NATIVE shipped in this release — see Added.)
Transformers-only release; no SKaiNET engine bump in this version. The focus is the BOM and the consumer-facing docs.
-
BOM coverage gap.
:llm-inference:apertusand:llm-inference:voxtralship to Maven Central but were missing fromskainet-transformers-bom's constraints. Consumers who imported the BOM and pulled either of these artifacts got no version alignment for them. -
Wrong artifact IDs in the README and tutorials. The "Current release" snippet in
README.mdand the two tutorial pages (getting-started-java.adoc,llama3-tool-calling.adoc) showedsk.ainet.transformers:llm-core/llm-runtime-kllama/llm-agent— those are project paths, not published artifact IDs. The real coordinates areskainet-transformers-core,skainet-transformers-runtime-kllama,skainet-transformers-agent; anyone copy-pasting hit a "module not found" error. Fixed and switched the snippets to the BOM pattern so future version bumps only need to touch one line. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
- BOM internals: auto-discovery. The constraint list in
llm-bom/build.gradle.ktsis no longer hand-maintained. A new convention plugin inbuildSrc/(sk.ainet.transformers.bom-coverage) auto-discovers every sibling subproject that appliescom.vanniktech.maven.publishand adds it as anapiconstraint on the BOM. The only manual input left is the exclusion list (currently just:llm-performance); the BOM is coherent by construction — missing or drifting modules can no longer happen. llm-test-javaconsumes SKaiNET through the BOM so the BOM is exercised during the build itself; a regression in BOM constraints fails locally instead of leaking into a published artifact.- Removed dead
group = "sk.ainet.llm"override from the root build. The published group has always beensk.ainet.transformers(sourced fromgradle.properties); the override was being overridden in turn by vanniktech at publish time. The in-memory project group now matches the published group, which removes a footgun for anyone trying to resolve internal modules by GAV.
Version-aligned with SKaiNET 0.23.3.
-
Prefill progress callback.
generateUntilStopgains an optionalonPrefill: ((Int, Int) -> Unit)?parameter that fires once per prompt token during the autoregressive prefill loop, with(done, total)—doneis 1-based,totalisprompt.size. Plumbed through bothAgentLoop.runandAgentLoop.runWithEncoderas a new default-no-opAgentListener.onPrefillProgress(done, total)method.Why this matters: prefill is autoregressive in 0.23.x (the comment on
generateUntilStopdocuments theforwardBatchedcorrectness regression we reverted), so on a CPU-only runtime with a 300-token prompt the firstonTokenlands tens of seconds to minutes after the agent loop starts — UIs previously had no way to show the loop was alive. The new callback closes that gap (e.g.prefill: 32/282 (11%)).Backwards compatible — the new parameter and interface method default to null/no-op, so existing
AgentListenerimplementations and callers compile and behave unchanged.
- New tests for the prefill callback in
GenerateExtensionsTest:generateUntilStopReportsPrefillProgressForEachPromptToken— one(done, total)pair per prompt token, in order, withdone1-based andtotal = prompt.size.generateUntilStopWithEmptyPromptDoesNotInvokePrefillCallback— callback never fires for an empty prompt.
Version-aligned with SKaiNET 0.23.2.
-
Llama 3 tool-calling walkthrough — end-to-end docs for app integrators, covering chat template, JSON tool-call format, and
JavaAgentLoopwiring. -
Llama-3.2-1B-Instruct smoke test with a tool-calling assertion.
-
MongoDB / mdbr-leaf-ir embedding entry in the smoke runner catalogue.
-
kllama-cli: prompts, raw responses, and tools list now logged byToolCallingDemo. -
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
kllama-cli,kllama-native, andkllama-wasmswapped to the DSL path (OptimizedLLMRuntime+llamaNetwork()); placeholder GPU attention/tensor stubs deleted; native benchmark scenario renamed tonative-cpu-throughput.KLlamaJavafacade swapped to the DSL path.llm-core: SentencePiece decorator + GGUF tokenizer now route through upstreamsk.ainet.io.tokenizerinstead of a local fork; fixes Qwen / GPT-2 BPE GGUF tokenization.
fix(tool-calling): tolerate markdown code fences around Llama 3 JSON tool calls— the parser previously skipped fenced JSON, causing the agent loop to keep generating untilmaxTokensPerRoundinstead of executing the call.fix(qwen): NEOX (SPLIT_HALF) RoPE pairing for Qwen3 GGUFs.fix(transformer): thread metadata RMSNorm eps through QK-norm.fix(llama): inject logical 2D shape and dequant token_embd in DSL converter.fix(kllama-cli): route Llama GGUF/SafeTensors back to eagerLlamaRuntime`` — the DSL Q4/Q8 path is functionally correct but needs first-class Q4/Q8 DTypes to match the SIMD perf of the legacy path. Tracked as a followup.fix(kllama-cli): apply application plugin so :run task is wired.fix(smoke): tolerate runners that don't emit tok/s (embedding models).
:llm-runtime:kqwenmodule andLlamaIngestionBlocking.ktdeleted.
- API dumps refreshed for 0.23.2 (
api/directory).
0.23.1 — 2026-05-04
Version-aligned with SKaiNET 0.23.1.
-
Apertus end-to-end. Real-GGUF loading now works on top of skainet 0.23.x's block-major Q4_K
TensorDatawiring. Routing fix to go throughOptimizedLLMRuntime+apertusNetwork(), plus chat template, tool calling, and integration tests againstApertus-8B-Q4_K_S. SeeAPERTUS_ROLLOUT.md. -
Gemma 4 chat-model JVM facade (
Gemma4ChatModel) for embedded text-only deployments.close()now propagates to the mmap arena. The PLE mmap path consumes upstreamloadTensorStorageMappedrather than maintaining a fork. -
Multi-id EOS / stop-token support in the chat layer — needed for templates that emit several end-of-sequence markers (e.g. ChatML / Apertus).
-
End-to-end smoke test in
llm-test/llm-test-java(Llama3LeafSmokeTest) that wires LEAF (mdbr-leaf-mt, viaKBertJava) and Llama 3.2-1B (KLlamaJava) in one JVM, gated on env vars / cache fallbacks so CI without the checkpoints cleanly skips. -
Apertus tool calling as a first-class family alongside Llama 3, Gemma 4, Qwen, and ChatML/Hermes.
-
On-device Android E2E SmolLM2 generation spike for the runtime facades (#288, refs #272).
gradle/libs.versions.tomlskainetpin: 0.22.1 → 0.23.1.VERSION_NAME: 0.21.1 → 0.23.1 (no 0.22.x transformers release was tagged; the version line jumps to keep the engine and consumer artifacts in sync).kllama-cliandskainet-clishadow-jar builds now apply theServiceLoaderMETA-INF/servicesmerge fix-up so the priority-100skainet-backend-native-cpuprovider is picked up at runtime.llm-test/llm-test-javamaxHeapSize8g → 16g — the previous cap OOM'd while loading both Llama 3.2-1B + LEAF in a single JVM.
fix(apertus): force-dequant token_embd under NATIVE_OPTIMIZED— Apertus was producing garbage on quantized embeddings; we now dequant the token embedding tensor regardless of policy, matching upstream behaviour.fix(tokenizer): auto-detect SentencePiece marker in fromTokenizerJson— models that ship atokenizer.jsonwithout the explicitpre_tokenizer.type = SentencePiecemarker now decode correctly.fix(gemma4): produce coherent text on real SafeTensors checkpoint— the loader path for full HF-format Gemma 4 checkpoints (not just the GGUF variant) now produces coherent generations end-to-end.fix(apertus): route through OptimizedLLMRuntime + apertusNetwork()— the legacy direct-runtime path was bypassed; Apertus now flows through the optimized DAG runtime like every other family.
test(apertus): real-GGUF loader integration test against Apertus-8B-Q4_K_S.test(apertus): pin weight-loader fixes with regression tests.test(kgemma): fast tokenizer parity guard against HF reference.test(kgemma): tighten tool-call probe budget + add env override.- Native-cpu provider now wired into the
qwenandllamaJVM test runs so the priority-100 FFM kernels are exercised during CI.
docs(apertus): document chat-template formatplus the staged-rollout plan at the repo root (APERTUS_ROLLOUT.md).- README refreshed: lead with native FFM CPU performance numbers, current release coordinates at 0.23.1, "What's new" section in place of the previous "In develop, not in X yet" callout.
chore(apertus): close out rollout — remove deprecated runtimes. The pre-rollout direct-runtime entry points for Apertus are gone.
0.21.1 — 2026-04-30
Hotfix release: add missing POM_NAME for the apertus, voxtral, and
llm-performance modules so Maven Central publishing succeeds.
0.21.0 — 2026-04-29
Version-aligned with SKaiNET 0.21.0.
chore(release): bump SKaiNET to 0.21.0, prepare transformers 0.21.0— mirror the engine version in the transformers line so the coupling is explicit for Maven Central consumers. Engine highlights (delivered via the bump): Panama Vector FP32 matmul kernel auto-discovered viaServiceLoader,ScratchPoolSPI, Q4_K SIMD-fused matmul kernel, Q6_K dequant viaByteVector ql+qhextraction, canonical ggml layout for Q4_K + Q5_K, FP32MemSegarena leak fix.VERSION_NAMEjumps 0.18.0 → 0.21.0 to align tags with the engine; no 0.17.0 / 0.19.x / 0.20.0 transformers releases were ever tagged.
0.18.0 — earlier
Last published transformers release before the engine-aligned version line.
See git log v0.16.0..0.18.0 for details.