Skip to content

SKaiNET-transformers → engine 0.51: full engine-loader migration, MAPPED default, BitNet fold-in - #353

Merged
michalharakal merged 18 commits into
developfrom
feat/0.51-migration
Aug 30, 2026
Merged

michalharakal merged 18 commits into
developfrom
feat/0.51-migration

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

Completes the engine-adoption arc (#338–#346): this repository no longer carries any weight quantization, packing, or memory-staging machinery of its own. Every model family loads through the SKaiNET engine's StreamingGgufParametersLoader with a declared WeightForm, and MAPPED residency is the default everywhere. Supersedes the stale PR stack #345/#347/#348 (already closed).

Targets engine 0.51.0 (developed against 0.51.0-SNAPSHOT; flip the pin to the release when the engine cuts it — engine PR SKaiNET-developers/SKaiNET#1211 needs merging first, it's currently only in the local republished snapshot this branch was built and tested against).

Phases (one commit each, all green independently)

Breaking changes

  • QuantPolicy is gone; loaders take an optional WeightForm (null = keep-packed MAPPED default).
  • PreTransposedWeight and the packing layer (BlockQuantPacking, GgmlQuantEncodings, family-specific *MemSegConverter/*QuantLayout/*PackedWeights) are deleted — net ≈ −7,000 lines.
  • linearProject is one expression; the transpose-marker branch is gone.
  • OptimizedLLMRuntime DIRECT decode runs inside a ForwardScope by default (ctor-tunable, 0 disables).

Verification

  • Every phase: full assemble allTests apiCheck green.
  • tests/smoke/smoke-test.sh — 3/3 pass, all coherent generations (SmolLM2-135M, Qwen2.5-1.5B, BitNet-2B4T).
  • Memory claims measured, not asserted: qwen2.5-1.5B (1.0 GB) runs under -Xmx512m (~176 MB planned heap); SmolLM2 tied-lm_head packed-kernel flip took the CLI smoke from ~3.6 to ~50 tok/s.
  • ForwardScopeSteadyStateTest pins bit-identity between scoped/unscoped decode plus flat peakFloats/zero overflowBytes.

Follow-ups tracked separately

  • transformers#352 — Qwen2/2.5 golden-layer parity vs llama.cpp (decode is currently incoherent at baseline, unrelated to this migration).
  • SKaiNET#1198 — I2_S sidecar for GROUP-layout BitNet files to unlock the zero-copy mapped lane there too.
  • Pin flip to skainet = "0.51.0" + ./gradlew clean build without -PskainetMavenLocal, once the engine cuts the release.

🤖 Generated with Claude Code

michalharakal and others added 16 commits August 29, 2026 16:19
…-up shims

Phase 0 of the 0.51 migration (supersedes PR #345, retargeted past 0.49.0):
- catalog: skainet = 0.51.0-SNAPSHOT (engine develop @ b32b0102 + the
  cross-ar fix from SKaiNET#1209, published locally; the pin flips to the
  0.51.0 release at arc end), agp 9.3.2
- settings: the -PskainetMavenLocal resolution lane (sk.ainet.* from ~/.m2
  ahead of Central), ported from the bitnet branch with the fixed regex
- llm-core: ExperimentalMemoryApi opt-in (hosts the loader/decode adoption)
- QuantPolicy: the engine deleted it in 0.49.0; the enum moves here
  (same package, transformers-owned) as TRANSITIONAL plumbing — it dies with
  the loaders in the later phases, the WeightForm mapping is in its kdoc
- llama/gemma/apertus expose :llm-core as api() while their loader APIs
  still surface QuantPolicy (reverts with its deletion)
- PreTransposed* wrappers: explicit `encoding` override resolving the
  TensorData(nullable)/PackedBlockStorage(non-null) delegation diamond
  (transient — the wrappers are deleted in the gemma phase)

Acceptance: full `-PskainetMavenLocal assemble allTests` green across all
targets; CLI smoke on SmolLM2-135M Q4_K_M decodes coherently at ~9 tok/s.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ip deferred with evidence

Phase 1 of the 0.51 migration, revised mid-phase by a baseline finding:

- skainet-cli installs KernelPacks + FfmRowMajorKernelPack after backend
  selection (gains backend-api + backend-nativeCpu deps) — the mapped tier
  is wired before any loader asks for MAPPED residency.
- linearProject KEEPS matmul(x, transpose(w)) for unmarked weights, with a
  migration note: the matmulWeightTransposed flip is deferred to the phase
  that deletes the MemSeg converter lane, because that lane's tensor data
  declares no truthful block order — and, decisively, the lane is broken at
  BASELINE for pure-Q4_K/Q6_K models: pristine develop + engine 0.40.1
  crashes on Qwen2.5-1.5B's first decode step with the #993
  ClassCastException (Byte→Float in matmulGeneric), and engines >= 0.49
  convert that crash into silent garbage output via matmulGeneric's
  raw-code copyToFloatArray fallback. Verified by bisect: garbage/crash is
  identical with and without the seam flip, on 0.40.1, 0.49.0-SNAPSHOT and
  0.51.0-SNAPSHOT. The engine-loader migration (next phases) replaces the
  lane wholesale — that is the fix, not a patch here.

Acceptance: transformer-core jvmTest 10/10, kllama jvmTest green; SmolLM2
CLI smoke coherent (~9 tok/s). Qwen CLI smoke stays broken as at baseline;
its restoration is the decoder-phase acceptance gate.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…icy (#339, fixes #100)

Phase 2 of the 0.51 migration; absorbs superseded PR #347's diff onto the
consolidated branch, retargeted from 0.49.0 to the 0.51.0 line.

ApertusWeightLoader shrinks 682 lines to a thin wrapper over the engine's
StreamingGgufParametersLoader with WeightForm(shape = OUT_IN): packed Q*
tensors arrive as engine PackedBlockStorage data with rank-2 logical shapes
— which is why #100's "transpose requires at least 2 dimensions" crash
(rank-1 raw-Int8 tensors under NATIVE_OPTIMIZED) disappears by deletion.
DELETED: ApertusMemSegConverter, QuantizedTensor, the RAW_BYTES lane, their
tests. ApertusRealGgufLoadingTest rewritten against the engine forms.

Acceptance: apertus jvmTest 18/18, kapertus jvmTest green, full assemble
across all targets green.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…) + qwen2 attention-bias support

Phase 3 of the 0.51 migration; absorbs superseded PR #348 onto the
consolidated branch, plus two fixes it exposed:

- DecoderGgufWeightLoader rewritten as a thin wrapper over the engine's
  StreamingGgufParametersLoader (default WeightForm(shape = OUT_IN),
  keep-packed; DECODER_DEQUANTIZE_ALL for the widened lane). DELETED:
  DecoderGgufMemSegConverter, MemSegWeightConverter, MmapLlamaLoader,
  QuantizedTensorFactory(+Jvm), LlamaPackedWeights, LlamaQuantLayout,
  GraphAccelerator, GcHint + actuals, FusedQKVAccelerator, the quantTypes
  container field, and the tests that pinned them.
- qwen2/qwen2.5 attention biases (arc find, pre-existing on every lane):
  LlamaGGUFNameResolver had no .bias rules and MultiHeadAttention's bias
  params were never enabled, so blk.N.attn_{q,k,v}.bias loaded and silently
  never bound. Added resolver rules + attnBias threading
  (decoderTransformerNetwork → qwenNetwork), auto-detected from the file
  like qkNorm. Necessary, but not yet sufficient for coherent qwen2.5
  output — the residual divergence is tracked separately (pre-existing:
  the 0.40.1 baseline CRASHES on qwen2.5 with the #993 ClassCastException;
  there was never a working reference on this repo).
- New DecoderPackedForwardProbeTest: real-model guard that a packed Q4_K
  projection through linearProject matches its own block decode within
  W4A8 activation-quantization tolerance, for heap AND MemorySegment
  activation factories.

Acceptance: llama/qwen/voxtral/kllama jvmTest green, full assemble green;
SmolLM2 CLI smoke through the engine-loader lane decodes coherently
("Paris. Paris is known for its iconic landmarks like the Eiffel Tower").
Perf note: ~3.6 tok/s vs ~9 on the deleted MemSeg lane — expected until
P5 flips residency to MAPPED and the row-major tier serves; measured there.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#341)

Rewrite Gemma4WeightLoader (1079→~700 L) and Gemma3nWeightLoader
(995→~700 L) as thin wrappers over the engine's
StreamingGgufParametersLoader: quantized weights keep their stored block
encoding as packed tensors with truthful logical [out, in] shapes
(default), GEMMA_DEQUANTIZE_ALL requests the dense FP32 form the
FunctionGemma export harness consumes, token_embd is always dequantized
(Embedding.gather needs element access), and per_layer_token_embd always
stays packed, rewrapped as the GemmaPerLayerTokenEmbedTensorData
row-dequant source (dense FP32 is ~9 GB on E2B and overflows the JVM
array cap on E4B). QuantPolicy is gone from the whole gemma chain.

With the last converter lane gone, linearProject collapses to the single
engine primitive it was always meant to be:

    ops.matmulWeightTransposed(input, weight)

The deferred #338 seam flip is now validated: the packed parity matrix
test pins all seven GGML block formats through the flipped seam against
the engine's own canonical dequant (W4A8 tolerance, #944), the
DecoderPackedForwardProbeTest still passes, and the SmolLM2 CLI smoke
decodes coherently ("The capital of France is called Paris.", ~3.6
tok/s).

Deleted outright (no deprecation, per the migration arc):
- transformer-core: PreTransposedWeight.kt (all 7 wrappers),
  BlockQuantPacking.kt + their tests (BlockQuantPackingTest,
  LinearProjectionPreTransposedTest)
- llm-core: GgmlQuantEncodings expect/actual family (commonMain, jvmMain,
  registryBasedMain)
- gemma: GemmaMemSegConverter, GemmaQuantLayout, GemmaPackedWeights +
  GemmaQuantLayoutTest, GemmaQ5KPackedParityTest, GemmaQ5xPackedParityTest,
  GemmaTokenEmbdRowDequantParityTest
- kgemma: Gemma4IngestionJvm (loadDslRuntimeNative*, the Arena/MemSeg lane)

Consumers moved to the new seam: skainet-cli GEMMA branch is now the
same three-liner as the Apertus branch (GemmaNetworkLoader.fromGguf →
load), kgemma CLI/ingestions take WeightForm instead of
QuantPolicy/allowQuantized, Gemma3nIngestion defaults to
GEMMA_DEQUANTIZE_ALL because the hand-coded Gemma3nRuntime consumes
dense tensors, and Gemma4WeightMapper's shape checks now run for every
tensor — packed data carries real shapes, so the isQuant escape hatch is
gone. api dumps regenerated (bert's picks up the 0.51 ExecutionContext
surface).

Gates: :transformer-core:allTests, gemma/kgemma/llm-core allTests,
functiongemma + llama/qwen/apertus jvmTest, full multi-target assemble,
grep-zero on every deleted symbol, SmolLM2 CLI smoke.

Part of the #338 migration arc (P4). Closes #341.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rement (#342)

Loader defaults (decoder family, apertus, gemma4/3n) now request
WeightForm(shape = OUT_IN, residency = MAPPED): servable encodings stay
in file-backed pages and the FFM row-major kernel pack reads them
zero-copy; the engine heap-stages anything it cannot serve from the
mapping. The gemma PLE override stays HEAP — its row-dequant wrapper
needs a heap byte view.

The token embedding no longer gets force-dequantized to dense FP32.
New transformer-core PackedRowDequantTensorData wraps any canonical
packed 2-D table (heap Q*BlockTensorData or mapped
BufferPackedTensorData) as an engine RowDequantSource: Embedding
dequantizes only the rows a step looks up, and — because the wrapper
forwards PackedBlockStorage and the underlying view — a tied output
head still routes matmulWeightTransposed through the packed dispatch
chain. A 152k x 1536 vocabulary stays at its ~131 MB packed footprint
instead of a ~933 MB heap array, and the tied lm_head rides the packed
kernel instead of the dense transpose path (SmolLM2 CLI smoke: 3.6 ->
58 tok/s, greedy output byte-identical). Non-block-aligned packed
tables (none produced by canonical GGUF) are dequantized outright —
packed get() returns raw quantization codes and must never reach
gather.

skainet-cli: the hand-managed Arena.ofShared() + shutdown hook +
MemorySegmentTensorDataFactory context setup is gone — the engine
loader owns weight residency now — and the CLI prints
MemoryPlans.plan(...) from the GGUF header before loading, priced with
the same MAPPED form the loaders request, against the JVM heap cap
(warns when the plan does not fit). kgemma tests drop their dead
quantArena locals.

Measured gates (engine 0.51.0-SNAPSHOT @ e532ece8):
- qwen2.5-1.5b under -Xmx512m --context=2048: runs; plan 176 MB heap
  of 512 MB (fits), 1.0 GB weights fully mapped (decode quality is
  the pre-existing #352 qwen2 issue, tracked separately)
- smollm2-135m under -Xmx512m: coherent greedy output, ~50 tok/s,
  98 MB weights mapped, 133 MB planned heap
- transformer-core allTests, llama/qwen/apertus/gemma/kllama/kgemma/
  kapertus jvmTest, full assemble + allTests + apiCheck: green

Part of the #338 engine-adoption arc; supersedes the MemSeg staging
lane end-to-end.
…#151)

The BitNet block structure, verified against NeoGPU's reference driver
(hs_ml_infer.c — the working 2B4T inference the SKaiNET ternary port
tracks, SKaiNET#1136), differs from the Llama family in exactly two
places, and both land as the extension points the DSL already anticipated:

- DecoderFfnKind (the `ffnKind` parameter decoderTransformerNetwork's
  KDoc promised for the first non-SwiGLU model): RELU2_SUBLN builds the
  new BitNetFFN — down(subNorm(relu(gate(x))² * up(x))) — squared ReLU
  instead of SiLU, and an RMSNorm (`ffn_sub_norm`, over the FFN hidden
  dim) between the gated product and the down projection.
- attnSubNorm on MultiHeadAttention (mirroring the existing qkNorm
  pattern): an RMSNorm child applied to the merged attention output
  BEFORE o_proj, on both the fused-decode and the general SDPA paths.

New llm-inference/bitnet module, thin per the house rule: bitnetNetwork()
is a ≤10-line caller of the shared builder (INTERLEAVED RoPE per
llama.cpp's LLM_ARCH_BITNET, eps from metadata); BitNetGGUFNameResolver
delegates to the engine's Llama resolver and adds the two BitNet-only
tensors (blk.N.attn_sub_norm.weight / ffn_sub_norm.weight);
BitNetNetworkLoader mirrors QwenNetworkLoader (GGUF only — BitNet ships
as GGUF; 2B4T's tied output.weight is covered by the decoder loader's
existing fallback). ModelFamily.BITNET detects bitnet/bitnet-25/
bitnet-b1.58, and skainet-cli routes it through the generic DSL branch.

Baseline scope: F32/F16/BF16 GGUFs load exactly. The packed-ternary path
(engine StreamingGgufParametersLoader with i2sLayout → BITNET_B1_58
tensors → the vendored NEON kernels) and RequantizeTo(BITNET_PLANES) for
the lm_head are the next issue (#152 upstream naming: transformers#337).

Tests: BitNetFFN forward and the attn-sub-norm placement pinned against
hand-computed references (seqLen-1 attention makes the whole path
closed-form); the DSL pipeline test mirrors Qwen's — module tree, full
weight mapping including both sub-norms (the loader REQUIRES every
non-bias param to map), and DIRECT-mode finite logits. Full repo sweep:
708 tests, 0 failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A ternary I2_S GGUF now loads through the SKaiNET engine's
StreamingGgufParametersLoader instead of the transformers-side decoder
loader: ternary projections arrive as packed BITNET_B1_58 tensors
(0.25 B/weight, engine SKaiNET#1140), the lm_head — when the file
carries output.weight — as the multi-plane BITNET_PLANES format
(SKaiNET#1150, lossless for exactly-ternary sources: per-row absmax
normalizes to {-1,0,+1} and plane 0 captures them with zero residual),
and everything binds into the stock bitnetNetwork. From there nothing
BitNet-specific happens: linearProject's matmul(x, transpose(W)) unwraps
the engine's Wᵀ marker and KernelDispatch selects by storage format —
the vendored NeoGPU NEON kernels when the packs are installed, the
decoding reference otherwise. Correctness never depends on the packs.

The loader takes the engine's i2sLayout knob (BitNet.cpp x86 GROUP_128
default / ARM GROUP_64 / NeoGPU SEQUENTIAL), OUT_IN weight orientation
(the packed payloads are row-major over [out][in]), and a planesLmHead
switch. Metadata comes from the llama.cpp-convention KV keys in a
separate reader pass (the decoder loader's parser is private and its
tensor path is deliberately bypassed).

Needs SKaiNET fix #1181 (isHeapPackedWeight broadened to
PackedBlockStorage) — republished locally as 0.41.0-SNAPSHOT.

End-to-end test: a synthetic BitNet I2_S GGUF (GROUP_128 payloads,
trailer scales, llama.cpp KV metadata, written by the test) loads packed
— asserted BitNetB158TensorData projections + BitNetPlanesTensorData
lm_head — and OptimizedLLMRuntime logits match the FP32-widened load of
the same file, with and without NativeTernaryF32GemvKernel /
NativeTernaryLmheadKernel installed. Full repo sweep: 710 tests,
0 failures.

Remaining in #152: the two-stage top-k decode (planes 4-7 rescoring) and
a real 2B4T GGUF smoke.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
)

NeoGPU's Stage 1/Stage 2 trick as a sampling-level utility, deliberately
NOT a dispatch citizen: KernelDispatch's matmul over a BITNET_PLANES
weight stays the exact 8-plane product, and approximation is an
application decision made here, visibly.

BitNetTwoStageDecode scores the full vocabulary with planes 0-3 + row
scales (the fused NEON kernel's contract; portable reference here), then
rescores candidates exactly. The improvement over NeoGPU's flat top-200:
plane p carries weight 1/3^p, so the stage-1 truncation of row r is
bounded by rowScale(r) * (Σ_{p=4..7} 3^-p) * Σ|h| — computable per row,
and topK() rescores every row whose upper bound reaches the k-th best
lower bound. The returned top-k therefore EQUALS the exact full-matmul
top-k (pinned against a brute-force oracle across trials); a
maxCandidates cap (default 200, NeoGPU's LMH_CANDIDATES) turns the
guarantee back into the heuristic when scores are pathologically packed.

llm-core gains the candidate-shaped sampler the issue asked for next to
sampleFromLogits: ScoredToken + sampleFromCandidates (same temperature
semantics; mass outside the candidate list is zero by contract).

Full sweep: 713 tests, 0 failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e CLI

skainet-cli's BITNET branch now loads through BitNetPackedGgufLoader
(engine SKaiNET 0.49 line): ternary projections stay packed (0.25 B per
weight) and the vendored NeoGPU NEON kernels are installed at startup —
the exact f32 path, no requantization error. loadWithMetadata surfaces
the GGUF metadata so the runtime gets bos right, and the packed path
gains the tied-embeddings fallback the decoder loader could not provide
for it (2B4T ships no output.weight; the lm_head serves from token_embd).

Verified against the real microsoft/bitnet-b1.58-2B-4T I2_S GGUF
(1.19 GB, official BitNet.cpp file):

  The capital of France is Paris. Paris is the capital of France. Paris
  is known for its beautiful architecture, art, and cuisine. The Eiffel
  Tower is a famous landmark in Paris. ...

Greedy decode, 64 tokens, coherent and factually correct — which proves,
against the official file, the whole 1-bit chain at once: GROUP_128 I2_S
layout + trailer scales (closing SKaiNET#1140's verification caveat),
attn/ffn sub-norm placement, squared-ReLU FFN, tied lm_head, tokenizer.
Needs the engine's off-heap activation bridge (SKaiNET#1186) and the
version bump to the 0.49 line (release/0.49.0 published locally as
0.49.0-SNAPSHOT).

Throughput on this first correct run is ~0.45 tok/s on an Apple-arm64
dev host — the FP32-widened tied lm_head (128256×2560 dense matmul per
token) dominates; making token_embd serve packed and the per-call
segment copies cheaper are the perf follow-ups, deliberately after
correctness.

Full sweep: 713 tests, 0 failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… + MAPPED request (#336, #337)

Alignment on top of the four cherry-picked BitNet commits:

- BitNetNetworkLoader speaks the P3 decoder-loader API: the vendored
  QuantPolicy is gone from its surface; the random-access lane takes an
  optional WeightForm (null = the decoder loader's keep-packed MAPPED
  default, DECODER_DEQUANTIZE_ALL for dense FP32), the sequential lane
  always dequantizes.
- BitNetPackedGgufLoader requests residency = MAPPED for the keep-packed
  lane like everything else in the arc: the 0.51 engine serves a
  trailer-scaled SEQUENTIAL I2_S file zero-copy from the mapping (#1203);
  a GROUP-flavor file (stock BitNet.cpp) repacks and heap-stages until
  SKaiNET#1198's sidecar. The planes lm_head form deliberately stays
  heap — a requantization can never be file-backed, and
  BitNetPlanesTensorData's heap-only design is what keeps
  BitNetTwoStageDecode type-safe.
- CLI ternary installs ride the 0.51 TernaryF32KernelPack surface
  (NativeTernaryF32GemvKernel/NativeTernaryLmheadKernel.install delegate
  to it; nothing registers when the native artifact is absent).

Measured gate (engine 0.51.0-SNAPSHOT @ e532ece8): bitnet-2b4t-i2s.gguf
32-token greedy CLI decode is coherent ('The capital of France is
Paris. Paris is the capital of France. Paris is known for its beautiful
architecture, art, and cuisine. ...'). The local file is GROUP_128
flavor — plan shows 1.1 GB heap-staged packed weights; the zero-copy
mapped I2_S lane needs a NeoGPU-converted SEQUENTIAL file and is
pending SKaiNET#1198. bitnet/llm-core jvmTest, transformer-core
allTests, full assemble + allTests + apiCheck: green.
The transitional vendored enum (sk.ainet.io.model.QuantPolicy, carried
since P0 to keep the tree compiling while the loaders migrated) reached
zero consumers with the P4 gemma rewrite and the P6 BitNet alignment —
delete it, regenerate the llm-core api dump, and unwind the three
api(project(":llm-core")) exposures (llama/gemma/apertus) that existed
only to surface it back to implementation.

grep QuantPolicy across *.kt/*.kts: 0. Full assemble + allTests +
apiCheck: green.
OptimizedLLMRuntime arms a ForwardScope + ScopedExecutionContext for
DIRECT forwards (default slab 8M floats, ctor-tunable, 0 disables):
step activations bump-allocate from one pre-sized slab, and the scope
resets at the ENTRY of the next forward — which for every generation
loop (generateAutoregressive, generateBatched, GenerateUntilStop) is
strictly after sampleFromLogits consumed the previous logits. A caller
holding a logits tensor across forwards fails loudly
(StorageClosedException), never silently. OPTIMIZED/HYBRID paths are
untouched. forwardScopeMetrics exposes the scope's counters.

AppendKVCache/SlidingWindowKVCache gained detachFromStep: the kept K/V
history is copied out to ambient heap storage when a step scope is
active (the module-level analogue of ForwardScope.retain) — identity
without a scope. PositionalKVCache and the shared/padded variants
already copy into their own buffers and needed nothing.

Shook out a real engine bug: DefaultCpuOpsJvm.chooseQuantizedMatmul2D
dropped slab-backed (StorageFloatTensorData) activations to
matmulGeneric, whose per-element get() on a Q8 MemorySegment weight
returns raw quantization codes — silently wrong logits (surfaced by
QwenDslQuantizedTest at step 0). Fixed upstream with a bit-exact
regression test (engine PR #1211, dense-FP32 copyToFloatArray fallback);
this repo now builds against the republished 0.51.0-SNAPSHOT (1ef99c62)
carrying that fix.

New ForwardScopeSteadyStateTest pins the #343 contract: scoped decode
bit-identical to unscoped (raw-bits compare, 12 steps), peakFloats flat
after warm-up, overflowBytes == 0 every step, one reset per step.
SmolLM2 CLI greedy smoke byte-identical to P5/P7. Full assemble +
allTests + apiCheck green.
…cated accelerator seam (#346)

The #346 template work the earlier phases didn't already cover:

- decoderMetadataFromGguf (llama module, next to LlamaModelMetadata):
  the single GGUF metadata parser for every decoder-shaped family.
  Replaces DecoderGgufWeightLoader's private copy and
  BitNetPackedGgufLoader's drifted internal duplicate (narrower numeric
  coercions, weaker inference fallbacks). Lives in llm-inference/llama
  rather than llm-core because LlamaModelMetadata does — noted as a
  deliberate exception against #346's llm-core preference.
- GraphAccelerator (deprecated since the OptimizedLLMRuntime migration)
  and kllama's FusedQKVAccelerator deleted — P3's commit message already
  claimed this deletion but the files had survived; LlamaRuntime loses
  its never-used graphAccelerator parameter. LlamaRuntime itself stays:
  it serves the llama2.c .bin lane (Llama2DotCWeightLoader), which is a
  feature, not GGUF-migration legacy.

Naming survey against the #346 table: every family already conforms
(<F>NetworkDef/<F>NetworkLoader/<F>ConfigParser/<F>WeightLoader);
deliberate exceptions to record on the issue: llama/qwen/smollm2 share
DecoderGgufWeightLoader + LlamaTensorNames, gemma keeps versioned
prefixes (Gemma4/Gemma3n) in one module.

Gates: full assemble + allTests + apiCheck green; BitNet 2B4T CLI
greedy smoke coherent through the shared parser.
…load, CHANGELOG (#344)

- weight-quantization.adoc rewritten end to end: the old multi-stage
  MemSeg/QuantPolicy/pre-transpose pipeline diagram and prose replaced
  with the current architecture — WeightForm's four axes, MAPPED
  residency and what it buys (measured -Xmx512m qwen2.5-1.5B, Android
  0.50 numbers), the row-dequant token-embedding wrapper, the one-line
  linearProject seam, the memory plan, and the #343 forward scope.
- android-getting-started.adoc: QuantPolicy.NATIVE_OPTIMIZED sample
  code and prose replaced with the KernelPacks.install() +
  JniMappedKernelPack.install() + default-MAPPED-WeightForm shape;
  troubleshooting table updated to the current dequantize-override name.
- run-unified-cli.adoc: documents --explain-load.
- skainet-cli: new --explain-load flag prints
  AllocationResolver.explain() per weight (mapped/heap and why) using
  PlannerProfile.DESKTOP, after the existing pre-load memory-plan print.
- CHANGELOG.md: [Unreleased] entry for the full #338-#346 arc against
  engine 0.51.0(-SNAPSHOT) — every loader rewrite, MAPPED default,
  the row-dequant embedding wrapper, the forward scope, the BitNet
  family, and the ~7,000-line deletion list.

Verification: full assemble + allTests + apiCheck green;
tests/smoke/smoke-test.sh against local models (SmolLM2-135M,
Qwen2.5-1.5B, BitNet-2B4T) — 3/3 pass, all coherent generations.
SKaiNET 0.51.0 is tagged (SKaiNET-developers/SKaiNET@0.51.0, includes
engine PR #1211's ForwardScope activation-chooser fix). Maven Central
publish is in progress at time of this commit; verified locally
against an unsigned publishToMavenLocal built from the tag
(4bf88a5b): full assemble + allTests + apiCheck green, and
tests/smoke/smoke-test.sh 3/3 pass (SmolLM2-135M, Qwen2.5-1.5B,
BitNet-2B4T) on the tagged artifact.

Central publish completion and the -PskainetMavenLocal-free
clean-build re-verify (this arc's originally planned closing step)
are still pending — tracked as the remaining P10 follow-up.
…354)

Both entry points loaded weights MAPPED (DecoderGgufWeightLoader's
0.51 default, #342) but never installed KernelPacks / the
FfmRowMajorKernelPack row-major zero-copy pack — only skainet-cli's
Main.kt did. Without it, every matmul against a MAPPED/keep-packed
weight fell to KernelDispatch's decoding reference kernel: correct,
but catastrophically slower per matmul than the SIMD/FFM path.

This is exactly the class of bug SKaiNET-EdgeTranslator's SkaiNet
engine hit: 'translate produces nothing' with Engine=SkaiNet selected
was the reference-kernel fallback taking so long the app looked hung,
not a correctness or wiring bug — the same GGUF via skainet-cli
(which does install the packs) was coherent and fast the whole time.

- KLlamaJava.loadGGUF/loadSafeTensors: install once via an
  AtomicBoolean guard (skainet-cli's Main.kt install() call is
  idempotent but this facade can be called from arbitrary embedder
  code, so guard explicitly rather than relying on that).
- kllama-cli's Main.kt: same two-line install after backend selection,
  mirroring skainet-cli exactly.
- llm-runtime/kllama's jvmMain gained skainet-backend-api +
  skainet-backend-nativeCpu (previously only on linuxMain/androidMain
  — the JVM target had neither KernelPacks nor the row-major pack on
  its classpath at all).
- New KLlamaJavaKernelPackTimingProbe (env-var-gated, skips without
  KLLAMA_JAVA_PROBE_GGUF — not CI-portable): pins the fast path stays
  fast. Measured on Llama-3.2-1B-Instruct-Q4_K_M: load 1.87s, 16-token
  generate 661ms (~24 tok/s) with the fix; still computing after 60s+
  with it reverted (confirmed causal before restoring).

kllama-cli's Llama/Mistral GGUF branch is a SEPARATE, worse problem —
it still runs the legacy dense-FP32 LlamaRuntime instead of the DSL
path Qwen already uses on that same branch point; filed as
SKaiNET-transformers#354 rather than fixed here (larger surface,
needs the .bin/llama2.c lane's fate decided first).

Gates: kllama allTests (incl. the probe against the real model) +
full assemble + allTests + apiCheck green.
…eTranslator relies on (#354)

Extends the #338 kernel-pack timing probe with the exact call shape
SKaiNET-EdgeTranslator's SkaiNetLlm.jvm.kt uses: KLlamaJava.loadGGUF +
Llama3ChatTemplate.apply(SYSTEM, USER) instead of a raw
"$system\n\n$user" concatenation. The raw form is what
EdgeTranslator originally sent — Llama-3.2-Instruct is fine-tuned
specifically for the <|start_header_id|>...<|eot_id|> turn structure,
and without it the model neither reliably followed the
translate-into-{target} instruction (English output regardless of
target language) nor learned to emit its stop token in that shape
(generation ran to the token budget instead of stopping — the GGUF's
tokenizer.ggml.eos_token_id is confirmed already correct at 128009,
ruling out a stop-token misconfiguration).

Measured with the chat template applied: 'La capitale de la France est
Paris.' for an English-to-French translate prompt, stopping in 2.15s
against a 400-token budget (vs. running to the cap without it).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Apertus: NATIVE_OPTIMIZED Q4_K end-to-end inference broken — needs block-major tensor-data wrappers

1 participant