SKaiNET-transformers → engine 0.51: full engine-loader migration, MAPPED default, BitNet fold-in - #353
Merged
Conversation
…-up shims Phase 0 of the 0.51 migration (supersedes PR #345, retargeted past 0.49.0): - catalog: skainet = 0.51.0-SNAPSHOT (engine develop @ b32b0102 + the cross-ar fix from SKaiNET#1209, published locally; the pin flips to the 0.51.0 release at arc end), agp 9.3.2 - settings: the -PskainetMavenLocal resolution lane (sk.ainet.* from ~/.m2 ahead of Central), ported from the bitnet branch with the fixed regex - llm-core: ExperimentalMemoryApi opt-in (hosts the loader/decode adoption) - QuantPolicy: the engine deleted it in 0.49.0; the enum moves here (same package, transformers-owned) as TRANSITIONAL plumbing — it dies with the loaders in the later phases, the WeightForm mapping is in its kdoc - llama/gemma/apertus expose :llm-core as api() while their loader APIs still surface QuantPolicy (reverts with its deletion) - PreTransposed* wrappers: explicit `encoding` override resolving the TensorData(nullable)/PackedBlockStorage(non-null) delegation diamond (transient — the wrappers are deleted in the gemma phase) Acceptance: full `-PskainetMavenLocal assemble allTests` green across all targets; CLI smoke on SmolLM2-135M Q4_K_M decodes coherently at ~9 tok/s. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ip deferred with evidence Phase 1 of the 0.51 migration, revised mid-phase by a baseline finding: - skainet-cli installs KernelPacks + FfmRowMajorKernelPack after backend selection (gains backend-api + backend-nativeCpu deps) — the mapped tier is wired before any loader asks for MAPPED residency. - linearProject KEEPS matmul(x, transpose(w)) for unmarked weights, with a migration note: the matmulWeightTransposed flip is deferred to the phase that deletes the MemSeg converter lane, because that lane's tensor data declares no truthful block order — and, decisively, the lane is broken at BASELINE for pure-Q4_K/Q6_K models: pristine develop + engine 0.40.1 crashes on Qwen2.5-1.5B's first decode step with the #993 ClassCastException (Byte→Float in matmulGeneric), and engines >= 0.49 convert that crash into silent garbage output via matmulGeneric's raw-code copyToFloatArray fallback. Verified by bisect: garbage/crash is identical with and without the seam flip, on 0.40.1, 0.49.0-SNAPSHOT and 0.51.0-SNAPSHOT. The engine-loader migration (next phases) replaces the lane wholesale — that is the fix, not a patch here. Acceptance: transformer-core jvmTest 10/10, kllama jvmTest green; SmolLM2 CLI smoke coherent (~9 tok/s). Qwen CLI smoke stays broken as at baseline; its restoration is the decoder-phase acceptance gate. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…icy (#339, fixes #100) Phase 2 of the 0.51 migration; absorbs superseded PR #347's diff onto the consolidated branch, retargeted from 0.49.0 to the 0.51.0 line. ApertusWeightLoader shrinks 682 lines to a thin wrapper over the engine's StreamingGgufParametersLoader with WeightForm(shape = OUT_IN): packed Q* tensors arrive as engine PackedBlockStorage data with rank-2 logical shapes — which is why #100's "transpose requires at least 2 dimensions" crash (rank-1 raw-Int8 tensors under NATIVE_OPTIMIZED) disappears by deletion. DELETED: ApertusMemSegConverter, QuantizedTensor, the RAW_BYTES lane, their tests. ApertusRealGgufLoadingTest rewritten against the engine forms. Acceptance: apertus jvmTest 18/18, kapertus jvmTest green, full assemble across all targets green. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…) + qwen2 attention-bias support Phase 3 of the 0.51 migration; absorbs superseded PR #348 onto the consolidated branch, plus two fixes it exposed: - DecoderGgufWeightLoader rewritten as a thin wrapper over the engine's StreamingGgufParametersLoader (default WeightForm(shape = OUT_IN), keep-packed; DECODER_DEQUANTIZE_ALL for the widened lane). DELETED: DecoderGgufMemSegConverter, MemSegWeightConverter, MmapLlamaLoader, QuantizedTensorFactory(+Jvm), LlamaPackedWeights, LlamaQuantLayout, GraphAccelerator, GcHint + actuals, FusedQKVAccelerator, the quantTypes container field, and the tests that pinned them. - qwen2/qwen2.5 attention biases (arc find, pre-existing on every lane): LlamaGGUFNameResolver had no .bias rules and MultiHeadAttention's bias params were never enabled, so blk.N.attn_{q,k,v}.bias loaded and silently never bound. Added resolver rules + attnBias threading (decoderTransformerNetwork → qwenNetwork), auto-detected from the file like qkNorm. Necessary, but not yet sufficient for coherent qwen2.5 output — the residual divergence is tracked separately (pre-existing: the 0.40.1 baseline CRASHES on qwen2.5 with the #993 ClassCastException; there was never a working reference on this repo). - New DecoderPackedForwardProbeTest: real-model guard that a packed Q4_K projection through linearProject matches its own block decode within W4A8 activation-quantization tolerance, for heap AND MemorySegment activation factories. Acceptance: llama/qwen/voxtral/kllama jvmTest green, full assemble green; SmolLM2 CLI smoke through the engine-loader lane decodes coherently ("Paris. Paris is known for its iconic landmarks like the Eiffel Tower"). Perf note: ~3.6 tok/s vs ~9 on the deleted MemSeg lane — expected until P5 flips residency to MAPPED and the row-major tier serves; measured there. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#341) Rewrite Gemma4WeightLoader (1079→~700 L) and Gemma3nWeightLoader (995→~700 L) as thin wrappers over the engine's StreamingGgufParametersLoader: quantized weights keep their stored block encoding as packed tensors with truthful logical [out, in] shapes (default), GEMMA_DEQUANTIZE_ALL requests the dense FP32 form the FunctionGemma export harness consumes, token_embd is always dequantized (Embedding.gather needs element access), and per_layer_token_embd always stays packed, rewrapped as the GemmaPerLayerTokenEmbedTensorData row-dequant source (dense FP32 is ~9 GB on E2B and overflows the JVM array cap on E4B). QuantPolicy is gone from the whole gemma chain. With the last converter lane gone, linearProject collapses to the single engine primitive it was always meant to be: ops.matmulWeightTransposed(input, weight) The deferred #338 seam flip is now validated: the packed parity matrix test pins all seven GGML block formats through the flipped seam against the engine's own canonical dequant (W4A8 tolerance, #944), the DecoderPackedForwardProbeTest still passes, and the SmolLM2 CLI smoke decodes coherently ("The capital of France is called Paris.", ~3.6 tok/s). Deleted outright (no deprecation, per the migration arc): - transformer-core: PreTransposedWeight.kt (all 7 wrappers), BlockQuantPacking.kt + their tests (BlockQuantPackingTest, LinearProjectionPreTransposedTest) - llm-core: GgmlQuantEncodings expect/actual family (commonMain, jvmMain, registryBasedMain) - gemma: GemmaMemSegConverter, GemmaQuantLayout, GemmaPackedWeights + GemmaQuantLayoutTest, GemmaQ5KPackedParityTest, GemmaQ5xPackedParityTest, GemmaTokenEmbdRowDequantParityTest - kgemma: Gemma4IngestionJvm (loadDslRuntimeNative*, the Arena/MemSeg lane) Consumers moved to the new seam: skainet-cli GEMMA branch is now the same three-liner as the Apertus branch (GemmaNetworkLoader.fromGguf → load), kgemma CLI/ingestions take WeightForm instead of QuantPolicy/allowQuantized, Gemma3nIngestion defaults to GEMMA_DEQUANTIZE_ALL because the hand-coded Gemma3nRuntime consumes dense tensors, and Gemma4WeightMapper's shape checks now run for every tensor — packed data carries real shapes, so the isQuant escape hatch is gone. api dumps regenerated (bert's picks up the 0.51 ExecutionContext surface). Gates: :transformer-core:allTests, gemma/kgemma/llm-core allTests, functiongemma + llama/qwen/apertus jvmTest, full multi-target assemble, grep-zero on every deleted symbol, SmolLM2 CLI smoke. Part of the #338 migration arc (P4). Closes #341. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rement (#342) Loader defaults (decoder family, apertus, gemma4/3n) now request WeightForm(shape = OUT_IN, residency = MAPPED): servable encodings stay in file-backed pages and the FFM row-major kernel pack reads them zero-copy; the engine heap-stages anything it cannot serve from the mapping. The gemma PLE override stays HEAP — its row-dequant wrapper needs a heap byte view. The token embedding no longer gets force-dequantized to dense FP32. New transformer-core PackedRowDequantTensorData wraps any canonical packed 2-D table (heap Q*BlockTensorData or mapped BufferPackedTensorData) as an engine RowDequantSource: Embedding dequantizes only the rows a step looks up, and — because the wrapper forwards PackedBlockStorage and the underlying view — a tied output head still routes matmulWeightTransposed through the packed dispatch chain. A 152k x 1536 vocabulary stays at its ~131 MB packed footprint instead of a ~933 MB heap array, and the tied lm_head rides the packed kernel instead of the dense transpose path (SmolLM2 CLI smoke: 3.6 -> 58 tok/s, greedy output byte-identical). Non-block-aligned packed tables (none produced by canonical GGUF) are dequantized outright — packed get() returns raw quantization codes and must never reach gather. skainet-cli: the hand-managed Arena.ofShared() + shutdown hook + MemorySegmentTensorDataFactory context setup is gone — the engine loader owns weight residency now — and the CLI prints MemoryPlans.plan(...) from the GGUF header before loading, priced with the same MAPPED form the loaders request, against the JVM heap cap (warns when the plan does not fit). kgemma tests drop their dead quantArena locals. Measured gates (engine 0.51.0-SNAPSHOT @ e532ece8): - qwen2.5-1.5b under -Xmx512m --context=2048: runs; plan 176 MB heap of 512 MB (fits), 1.0 GB weights fully mapped (decode quality is the pre-existing #352 qwen2 issue, tracked separately) - smollm2-135m under -Xmx512m: coherent greedy output, ~50 tok/s, 98 MB weights mapped, 133 MB planned heap - transformer-core allTests, llama/qwen/apertus/gemma/kllama/kgemma/ kapertus jvmTest, full assemble + allTests + apiCheck: green Part of the #338 engine-adoption arc; supersedes the MemSeg staging lane end-to-end.
…#151) The BitNet block structure, verified against NeoGPU's reference driver (hs_ml_infer.c — the working 2B4T inference the SKaiNET ternary port tracks, SKaiNET#1136), differs from the Llama family in exactly two places, and both land as the extension points the DSL already anticipated: - DecoderFfnKind (the `ffnKind` parameter decoderTransformerNetwork's KDoc promised for the first non-SwiGLU model): RELU2_SUBLN builds the new BitNetFFN — down(subNorm(relu(gate(x))² * up(x))) — squared ReLU instead of SiLU, and an RMSNorm (`ffn_sub_norm`, over the FFN hidden dim) between the gated product and the down projection. - attnSubNorm on MultiHeadAttention (mirroring the existing qkNorm pattern): an RMSNorm child applied to the merged attention output BEFORE o_proj, on both the fused-decode and the general SDPA paths. New llm-inference/bitnet module, thin per the house rule: bitnetNetwork() is a ≤10-line caller of the shared builder (INTERLEAVED RoPE per llama.cpp's LLM_ARCH_BITNET, eps from metadata); BitNetGGUFNameResolver delegates to the engine's Llama resolver and adds the two BitNet-only tensors (blk.N.attn_sub_norm.weight / ffn_sub_norm.weight); BitNetNetworkLoader mirrors QwenNetworkLoader (GGUF only — BitNet ships as GGUF; 2B4T's tied output.weight is covered by the decoder loader's existing fallback). ModelFamily.BITNET detects bitnet/bitnet-25/ bitnet-b1.58, and skainet-cli routes it through the generic DSL branch. Baseline scope: F32/F16/BF16 GGUFs load exactly. The packed-ternary path (engine StreamingGgufParametersLoader with i2sLayout → BITNET_B1_58 tensors → the vendored NEON kernels) and RequantizeTo(BITNET_PLANES) for the lm_head are the next issue (#152 upstream naming: transformers#337). Tests: BitNetFFN forward and the attn-sub-norm placement pinned against hand-computed references (seqLen-1 attention makes the whole path closed-form); the DSL pipeline test mirrors Qwen's — module tree, full weight mapping including both sub-norms (the loader REQUIRES every non-bias param to map), and DIRECT-mode finite logits. Full repo sweep: 708 tests, 0 failures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A ternary I2_S GGUF now loads through the SKaiNET engine's
StreamingGgufParametersLoader instead of the transformers-side decoder
loader: ternary projections arrive as packed BITNET_B1_58 tensors
(0.25 B/weight, engine SKaiNET#1140), the lm_head — when the file
carries output.weight — as the multi-plane BITNET_PLANES format
(SKaiNET#1150, lossless for exactly-ternary sources: per-row absmax
normalizes to {-1,0,+1} and plane 0 captures them with zero residual),
and everything binds into the stock bitnetNetwork. From there nothing
BitNet-specific happens: linearProject's matmul(x, transpose(W)) unwraps
the engine's Wᵀ marker and KernelDispatch selects by storage format —
the vendored NeoGPU NEON kernels when the packs are installed, the
decoding reference otherwise. Correctness never depends on the packs.
The loader takes the engine's i2sLayout knob (BitNet.cpp x86 GROUP_128
default / ARM GROUP_64 / NeoGPU SEQUENTIAL), OUT_IN weight orientation
(the packed payloads are row-major over [out][in]), and a planesLmHead
switch. Metadata comes from the llama.cpp-convention KV keys in a
separate reader pass (the decoder loader's parser is private and its
tensor path is deliberately bypassed).
Needs SKaiNET fix #1181 (isHeapPackedWeight broadened to
PackedBlockStorage) — republished locally as 0.41.0-SNAPSHOT.
End-to-end test: a synthetic BitNet I2_S GGUF (GROUP_128 payloads,
trailer scales, llama.cpp KV metadata, written by the test) loads packed
— asserted BitNetB158TensorData projections + BitNetPlanesTensorData
lm_head — and OptimizedLLMRuntime logits match the FP32-widened load of
the same file, with and without NativeTernaryF32GemvKernel /
NativeTernaryLmheadKernel installed. Full repo sweep: 710 tests,
0 failures.
Remaining in #152: the two-stage top-k decode (planes 4-7 rescoring) and
a real 2B4T GGUF smoke.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
) NeoGPU's Stage 1/Stage 2 trick as a sampling-level utility, deliberately NOT a dispatch citizen: KernelDispatch's matmul over a BITNET_PLANES weight stays the exact 8-plane product, and approximation is an application decision made here, visibly. BitNetTwoStageDecode scores the full vocabulary with planes 0-3 + row scales (the fused NEON kernel's contract; portable reference here), then rescores candidates exactly. The improvement over NeoGPU's flat top-200: plane p carries weight 1/3^p, so the stage-1 truncation of row r is bounded by rowScale(r) * (Σ_{p=4..7} 3^-p) * Σ|h| — computable per row, and topK() rescores every row whose upper bound reaches the k-th best lower bound. The returned top-k therefore EQUALS the exact full-matmul top-k (pinned against a brute-force oracle across trials); a maxCandidates cap (default 200, NeoGPU's LMH_CANDIDATES) turns the guarantee back into the heuristic when scores are pathologically packed. llm-core gains the candidate-shaped sampler the issue asked for next to sampleFromLogits: ScoredToken + sampleFromCandidates (same temperature semantics; mass outside the candidate list is zero by contract). Full sweep: 713 tests, 0 failures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e CLI skainet-cli's BITNET branch now loads through BitNetPackedGgufLoader (engine SKaiNET 0.49 line): ternary projections stay packed (0.25 B per weight) and the vendored NeoGPU NEON kernels are installed at startup — the exact f32 path, no requantization error. loadWithMetadata surfaces the GGUF metadata so the runtime gets bos right, and the packed path gains the tied-embeddings fallback the decoder loader could not provide for it (2B4T ships no output.weight; the lm_head serves from token_embd). Verified against the real microsoft/bitnet-b1.58-2B-4T I2_S GGUF (1.19 GB, official BitNet.cpp file): The capital of France is Paris. Paris is the capital of France. Paris is known for its beautiful architecture, art, and cuisine. The Eiffel Tower is a famous landmark in Paris. ... Greedy decode, 64 tokens, coherent and factually correct — which proves, against the official file, the whole 1-bit chain at once: GROUP_128 I2_S layout + trailer scales (closing SKaiNET#1140's verification caveat), attn/ffn sub-norm placement, squared-ReLU FFN, tied lm_head, tokenizer. Needs the engine's off-heap activation bridge (SKaiNET#1186) and the version bump to the 0.49 line (release/0.49.0 published locally as 0.49.0-SNAPSHOT). Throughput on this first correct run is ~0.45 tok/s on an Apple-arm64 dev host — the FP32-widened tied lm_head (128256×2560 dense matmul per token) dominates; making token_embd serve packed and the per-call segment copies cheaper are the perf follow-ups, deliberately after correctness. Full sweep: 713 tests, 0 failures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… + MAPPED request (#336, #337) Alignment on top of the four cherry-picked BitNet commits: - BitNetNetworkLoader speaks the P3 decoder-loader API: the vendored QuantPolicy is gone from its surface; the random-access lane takes an optional WeightForm (null = the decoder loader's keep-packed MAPPED default, DECODER_DEQUANTIZE_ALL for dense FP32), the sequential lane always dequantizes. - BitNetPackedGgufLoader requests residency = MAPPED for the keep-packed lane like everything else in the arc: the 0.51 engine serves a trailer-scaled SEQUENTIAL I2_S file zero-copy from the mapping (#1203); a GROUP-flavor file (stock BitNet.cpp) repacks and heap-stages until SKaiNET#1198's sidecar. The planes lm_head form deliberately stays heap — a requantization can never be file-backed, and BitNetPlanesTensorData's heap-only design is what keeps BitNetTwoStageDecode type-safe. - CLI ternary installs ride the 0.51 TernaryF32KernelPack surface (NativeTernaryF32GemvKernel/NativeTernaryLmheadKernel.install delegate to it; nothing registers when the native artifact is absent). Measured gate (engine 0.51.0-SNAPSHOT @ e532ece8): bitnet-2b4t-i2s.gguf 32-token greedy CLI decode is coherent ('The capital of France is Paris. Paris is the capital of France. Paris is known for its beautiful architecture, art, and cuisine. ...'). The local file is GROUP_128 flavor — plan shows 1.1 GB heap-staged packed weights; the zero-copy mapped I2_S lane needs a NeoGPU-converted SEQUENTIAL file and is pending SKaiNET#1198. bitnet/llm-core jvmTest, transformer-core allTests, full assemble + allTests + apiCheck: green.
The transitional vendored enum (sk.ainet.io.model.QuantPolicy, carried
since P0 to keep the tree compiling while the loaders migrated) reached
zero consumers with the P4 gemma rewrite and the P6 BitNet alignment —
delete it, regenerate the llm-core api dump, and unwind the three
api(project(":llm-core")) exposures (llama/gemma/apertus) that existed
only to surface it back to implementation.
grep QuantPolicy across *.kt/*.kts: 0. Full assemble + allTests +
apiCheck: green.
OptimizedLLMRuntime arms a ForwardScope + ScopedExecutionContext for DIRECT forwards (default slab 8M floats, ctor-tunable, 0 disables): step activations bump-allocate from one pre-sized slab, and the scope resets at the ENTRY of the next forward — which for every generation loop (generateAutoregressive, generateBatched, GenerateUntilStop) is strictly after sampleFromLogits consumed the previous logits. A caller holding a logits tensor across forwards fails loudly (StorageClosedException), never silently. OPTIMIZED/HYBRID paths are untouched. forwardScopeMetrics exposes the scope's counters. AppendKVCache/SlidingWindowKVCache gained detachFromStep: the kept K/V history is copied out to ambient heap storage when a step scope is active (the module-level analogue of ForwardScope.retain) — identity without a scope. PositionalKVCache and the shared/padded variants already copy into their own buffers and needed nothing. Shook out a real engine bug: DefaultCpuOpsJvm.chooseQuantizedMatmul2D dropped slab-backed (StorageFloatTensorData) activations to matmulGeneric, whose per-element get() on a Q8 MemorySegment weight returns raw quantization codes — silently wrong logits (surfaced by QwenDslQuantizedTest at step 0). Fixed upstream with a bit-exact regression test (engine PR #1211, dense-FP32 copyToFloatArray fallback); this repo now builds against the republished 0.51.0-SNAPSHOT (1ef99c62) carrying that fix. New ForwardScopeSteadyStateTest pins the #343 contract: scoped decode bit-identical to unscoped (raw-bits compare, 12 steps), peakFloats flat after warm-up, overflowBytes == 0 every step, one reset per step. SmolLM2 CLI greedy smoke byte-identical to P5/P7. Full assemble + allTests + apiCheck green.
…cated accelerator seam (#346) The #346 template work the earlier phases didn't already cover: - decoderMetadataFromGguf (llama module, next to LlamaModelMetadata): the single GGUF metadata parser for every decoder-shaped family. Replaces DecoderGgufWeightLoader's private copy and BitNetPackedGgufLoader's drifted internal duplicate (narrower numeric coercions, weaker inference fallbacks). Lives in llm-inference/llama rather than llm-core because LlamaModelMetadata does — noted as a deliberate exception against #346's llm-core preference. - GraphAccelerator (deprecated since the OptimizedLLMRuntime migration) and kllama's FusedQKVAccelerator deleted — P3's commit message already claimed this deletion but the files had survived; LlamaRuntime loses its never-used graphAccelerator parameter. LlamaRuntime itself stays: it serves the llama2.c .bin lane (Llama2DotCWeightLoader), which is a feature, not GGUF-migration legacy. Naming survey against the #346 table: every family already conforms (<F>NetworkDef/<F>NetworkLoader/<F>ConfigParser/<F>WeightLoader); deliberate exceptions to record on the issue: llama/qwen/smollm2 share DecoderGgufWeightLoader + LlamaTensorNames, gemma keeps versioned prefixes (Gemma4/Gemma3n) in one module. Gates: full assemble + allTests + apiCheck green; BitNet 2B4T CLI greedy smoke coherent through the shared parser.
…load, CHANGELOG (#344) - weight-quantization.adoc rewritten end to end: the old multi-stage MemSeg/QuantPolicy/pre-transpose pipeline diagram and prose replaced with the current architecture — WeightForm's four axes, MAPPED residency and what it buys (measured -Xmx512m qwen2.5-1.5B, Android 0.50 numbers), the row-dequant token-embedding wrapper, the one-line linearProject seam, the memory plan, and the #343 forward scope. - android-getting-started.adoc: QuantPolicy.NATIVE_OPTIMIZED sample code and prose replaced with the KernelPacks.install() + JniMappedKernelPack.install() + default-MAPPED-WeightForm shape; troubleshooting table updated to the current dequantize-override name. - run-unified-cli.adoc: documents --explain-load. - skainet-cli: new --explain-load flag prints AllocationResolver.explain() per weight (mapped/heap and why) using PlannerProfile.DESKTOP, after the existing pre-load memory-plan print. - CHANGELOG.md: [Unreleased] entry for the full #338-#346 arc against engine 0.51.0(-SNAPSHOT) — every loader rewrite, MAPPED default, the row-dequant embedding wrapper, the forward scope, the BitNet family, and the ~7,000-line deletion list. Verification: full assemble + allTests + apiCheck green; tests/smoke/smoke-test.sh against local models (SmolLM2-135M, Qwen2.5-1.5B, BitNet-2B4T) — 3/3 pass, all coherent generations.
SKaiNET 0.51.0 is tagged (SKaiNET-developers/SKaiNET@0.51.0, includes engine PR #1211's ForwardScope activation-chooser fix). Maven Central publish is in progress at time of this commit; verified locally against an unsigned publishToMavenLocal built from the tag (4bf88a5b): full assemble + allTests + apiCheck green, and tests/smoke/smoke-test.sh 3/3 pass (SmolLM2-135M, Qwen2.5-1.5B, BitNet-2B4T) on the tagged artifact. Central publish completion and the -PskainetMavenLocal-free clean-build re-verify (this arc's originally planned closing step) are still pending — tracked as the remaining P10 follow-up.
…354) Both entry points loaded weights MAPPED (DecoderGgufWeightLoader's 0.51 default, #342) but never installed KernelPacks / the FfmRowMajorKernelPack row-major zero-copy pack — only skainet-cli's Main.kt did. Without it, every matmul against a MAPPED/keep-packed weight fell to KernelDispatch's decoding reference kernel: correct, but catastrophically slower per matmul than the SIMD/FFM path. This is exactly the class of bug SKaiNET-EdgeTranslator's SkaiNet engine hit: 'translate produces nothing' with Engine=SkaiNet selected was the reference-kernel fallback taking so long the app looked hung, not a correctness or wiring bug — the same GGUF via skainet-cli (which does install the packs) was coherent and fast the whole time. - KLlamaJava.loadGGUF/loadSafeTensors: install once via an AtomicBoolean guard (skainet-cli's Main.kt install() call is idempotent but this facade can be called from arbitrary embedder code, so guard explicitly rather than relying on that). - kllama-cli's Main.kt: same two-line install after backend selection, mirroring skainet-cli exactly. - llm-runtime/kllama's jvmMain gained skainet-backend-api + skainet-backend-nativeCpu (previously only on linuxMain/androidMain — the JVM target had neither KernelPacks nor the row-major pack on its classpath at all). - New KLlamaJavaKernelPackTimingProbe (env-var-gated, skips without KLLAMA_JAVA_PROBE_GGUF — not CI-portable): pins the fast path stays fast. Measured on Llama-3.2-1B-Instruct-Q4_K_M: load 1.87s, 16-token generate 661ms (~24 tok/s) with the fix; still computing after 60s+ with it reverted (confirmed causal before restoring). kllama-cli's Llama/Mistral GGUF branch is a SEPARATE, worse problem — it still runs the legacy dense-FP32 LlamaRuntime instead of the DSL path Qwen already uses on that same branch point; filed as SKaiNET-transformers#354 rather than fixed here (larger surface, needs the .bin/llama2.c lane's fate decided first). Gates: kllama allTests (incl. the probe against the real model) + full assemble + allTests + apiCheck green.
…eTranslator relies on (#354) Extends the #338 kernel-pack timing probe with the exact call shape SKaiNET-EdgeTranslator's SkaiNetLlm.jvm.kt uses: KLlamaJava.loadGGUF + Llama3ChatTemplate.apply(SYSTEM, USER) instead of a raw "$system\n\n$user" concatenation. The raw form is what EdgeTranslator originally sent — Llama-3.2-Instruct is fine-tuned specifically for the <|start_header_id|>...<|eot_id|> turn structure, and without it the model neither reliably followed the translate-into-{target} instruction (English output regardless of target language) nor learned to emit its stop token in that shape (generation ran to the token budget instead of stopping — the GGUF's tokenizer.ggml.eos_token_id is confirmed already correct at 128009, ruling out a stop-token misconfiguration). Measured with the chat template applied: 'La capitale de la France est Paris.' for an English-to-French translate prompt, stopping in 2.15s against a 400-token budget (vs. running to the cap without it).
This was referenced Aug 31, 2026
Closed
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Completes the engine-adoption arc (#338–#346): this repository no longer carries any weight quantization, packing, or memory-staging machinery of its own. Every model family loads through the SKaiNET engine's
StreamingGgufParametersLoaderwith a declaredWeightForm, and MAPPED residency is the default everywhere. Supersedes the stale PR stack #345/#347/#348 (already closed).Targets engine 0.51.0 (developed against
0.51.0-SNAPSHOT; flip the pin to the release when the engine cuts it — engine PR SKaiNET-developers/SKaiNET#1211 needs merging first, it's currently only in the local republished snapshot this branch was built and tested against).Phases (one commit each, all green independently)
linearProjectseam preplinearProjectcollapses toops.matmulWeightTransposed(0.49.0 adoption B3: gemma4/gemma3n onto the engine loader (retire BlockQuantPacking + capability gate) #341)PackedRowDequantTensorData(row-dequant token embeddings), CLI Arena retirement + pre-load memory plan (0.49.0 adoption B4: voxtral/functiongemma/perf/CLI stragglers — QuantPolicy to zero #342)WeightForm/TernaryF32KernelPackAPI ([bitnet] T1: BitNet architecture — ModelFamily, network def, packed-weight loader #336, [bitnet] T2: lm_head via BITNET_PLANES weight format (WeightForm) + two-stage top-k decode #337)QuantPolicydeleted to zero consumers (0.49.0 adoption B4: voxtral/functiongemma/perf/CLI stragglers — QuantPolicy to zero #342)ForwardScopein the DIRECT decode loop (0.49.0 adoption C: flat-memory decode via ctx.forwardScope #343) — shook out and fixed a real engine bug (fix(cpu-jvm): scoped dense-FP32 activations must not fall out of the quantized matmul chooser SKaiNET#1211)decoderMetadataFromGgufparser, deleted the deprecatedGraphAccelerator/FusedQKVAcceleratorseam (Common structure and naming convention for model family modules #346)--explain-loadCLI flag, CHANGELOGBreaking changes
QuantPolicyis gone; loaders take an optionalWeightForm(null= keep-packed MAPPED default).PreTransposedWeightand the packing layer (BlockQuantPacking,GgmlQuantEncodings, family-specific*MemSegConverter/*QuantLayout/*PackedWeights) are deleted — net ≈ −7,000 lines.linearProjectis one expression; the transpose-marker branch is gone.OptimizedLLMRuntimeDIRECT decode runs inside aForwardScopeby default (ctor-tunable,0disables).Verification
assemble allTests apiCheckgreen.tests/smoke/smoke-test.sh— 3/3 pass, all coherent generations (SmolLM2-135M, Qwen2.5-1.5B, BitNet-2B4T).-Xmx512m(~176 MB planned heap); SmolLM2 tied-lm_head packed-kernel flip took the CLI smoke from ~3.6 to ~50 tok/s.ForwardScopeSteadyStateTestpins bit-identity between scoped/unscoped decode plus flatpeakFloats/zerooverflowBytes.Follow-ups tracked separately
skainet = "0.51.0"+./gradlew clean buildwithout-PskainetMavenLocal, once the engine cuts the release.🤖 Generated with Claude Code