Follow-up to #1189/#1190, which made WeightResidency.MAPPED serve Q4_K/Q6_K payloads from file-backed pages with zero heap bytes — verified on Android through the JNI direct-buffer kernels (JniMappedKernelPack).
On the desktop JVM the same loader path now emits BufferPackedTensorData, but no fast kernel serves the BLOCKED_ROW_MAJOR × off-heap key there: dispatch falls back to the decoding reference kernel (correct, very slow). The pieces already exist:
skainet_q4k_matmul_rm / skainet_q6k_matmul_rm are in the shared native library (FFM binds them by name — RowMajorMatmulParityTest already does exactly that in jvmTest).
NativeQ4KMemSegMatmulKernel shows the MemSeg downcall pattern (feed-order symbol today, no dispatch consumer).
MemorySegment.ofBuffer(directByteBuffer) bridges MappedBufferStorage/DirectBufferStorage.buffer() into FFM with no copy.
Sketch
NativeQ4KMatmulRmKernel / NativeQ6KMatmulRmKernel (native-cpu jvmMain): downcalls to the _rm symbols taking the weight as a MemorySegment.
- A jvmMain view-kernel sibling of
JniBufferPackedMatmulKernel that unwraps mapped/direct storage via MemorySegment.ofBuffer and registers under the same exact keys (matmul(FP32 dense contiguous × FP32/Q4_K|Q6_K blocked_row_major)), installed alongside KernelPacks.install(NativeKernelProvider).
- Parity test vs the JNI-verified semantics: row-major FFM output must be bit-identical to the feed-order kernel on permuted bytes (same oracle as
RowMajorMatmulParityTest).
Why it matters
Desktop JVM inference of mapped models currently pays reference-kernel cost for exactly the tensors #1189 made cheap on Android; with this, MAPPED becomes the best default on both tiers, and StorageCapabilities.mappedServableEncodings stays truthful for the JVM without qualification.
Follow-up to #1189/#1190, which made
WeightResidency.MAPPEDserve Q4_K/Q6_K payloads from file-backed pages with zero heap bytes — verified on Android through the JNI direct-buffer kernels (JniMappedKernelPack).On the desktop JVM the same loader path now emits
BufferPackedTensorData, but no fast kernel serves theBLOCKED_ROW_MAJOR× off-heap key there: dispatch falls back to the decoding reference kernel (correct, very slow). The pieces already exist:skainet_q4k_matmul_rm/skainet_q6k_matmul_rmare in the shared native library (FFM binds them by name —RowMajorMatmulParityTestalready does exactly that in jvmTest).NativeQ4KMemSegMatmulKernelshows the MemSeg downcall pattern (feed-order symbol today, no dispatch consumer).MemorySegment.ofBuffer(directByteBuffer)bridgesMappedBufferStorage/DirectBufferStorage.buffer()into FFM with no copy.Sketch
NativeQ4KMatmulRmKernel/NativeQ6KMatmulRmKernel(native-cpu jvmMain): downcalls to the_rmsymbols taking the weight as aMemorySegment.JniBufferPackedMatmulKernelthat unwraps mapped/direct storage viaMemorySegment.ofBufferand registers under the same exact keys (matmul(FP32 dense contiguous × FP32/Q4_K|Q6_K blocked_row_major)), installed alongsideKernelPacks.install(NativeKernelProvider).RowMajorMatmulParityTest).Why it matters
Desktop JVM inference of mapped models currently pays reference-kernel cost for exactly the tensors #1189 made cheap on Android; with this,
MAPPEDbecomes the best default on both tiers, andStorageCapabilities.mappedServableEncodingsstays truthful for the JVM without qualification.