Skip to content

feat(backend-api): kernel packs register under view keys; KernelKey carries platform capabilities (SKEEP-003 P3, S1.7c) - #1072

Merged
michalharakal merged 1 commit into
developfrom
feature/1029-matmul-dispatch-jvm-packs
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1029-matmul-dispatch-jvm-packs

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.7c (milestone M1 #1002, §5.2): the KernelProvider SPI (scalar, Panama/Vector API, native FFM, JNI/NEON) and the view-keyed dispatcher are wired together, so the generic path can select a pack's kernel instead of the decoding reference.

  • KernelKey.capabilities: Set<String> (vector, dotprod, i8mm, ffm, …): a pack registers its key with what it requires, so a device that lacks the capability never selects that kernel (Ship the aarch64-verified NEON kernels to mobile: Apple targets + Android JNI for skainet-backend-native-cpu (measured 21 → 1.0 → 0.11 tok/s cliff) #920). Capabilities are part of the key's identity and its rendering: matmul(… × …) @host [dotprod,vector].
  • KernelPacks.install(provider): installs the always-present reference kernel plus the provider's FP32 GEMM as a ViewKernel, under two keys — a weight normally reaches the dispatcher as a transposed view of a contiguous [k, n] buffer (LayoutClass.STRIDED), which is exactly what the SPI GEMM's stride arguments express; a genuinely [n, k] buffer is the contiguous key.
  • Fp32ViewMatmulKernel unwraps each view once (Phase-2 spike rule) and maps the layout to the SPI's (offset, stride) contract. When the weight is not a transposed contiguous buffer it defers to the reference kernel rather than mis-indexing silently — pinned by a test.

Scope: why only the dense FP32 kernel is bridged

The packed SPI kernels (Q4_0…Q6_K) take their weight bytes block-major — the layout DefaultCpuOpsBase.transposePackedBlocks produces — while a packed TensorView describes the file's canonical row-major block order. That contract is precisely what #973 reports as unwritten and contradictory across the engine and the downstream converters, and getting it wrong is the silent wrong numbers class of #968/#971 (all-zero matmul output, no exception raised).

So this slice does not bridge them: the packed fast paths keep their existing, working ladder in DefaultCpuOps/DefaultCpuOpsJvm, and the registry serves packed operands with the decoding reference kernel, which is correct for any layout. #973 is the prerequisite for finishing the packed migration; that follow-up can then delete the ladders under the golden parity gate.

Test plan

KernelPacksTest (10/10 on JVM and linuxX64): the reference kernel is always installed; a pack kernel serves the dense key and agrees numerically with the reference; an output-major weight falls back instead of being mis-indexed; capabilities are part of the key.

Full local gate (scripts/pr-gate.sh, JDK 25) — all legs passed, including the packed-encoding golden parity tests; results in the first comment.

Closes #1029

🤖 Generated with Claude Code

…arries platform capabilities (SKEEP-003 P3)

Milestone M1 (#1002), §5.2. The KernelProvider SPI (scalar, Panama,
native FFM, JNI/NEON) and the view-keyed dispatcher are now wired
together, so the generic path can pick a pack's kernel instead of the
decoding reference.

- KernelKey gains `capabilities: Set<String>` (vector, dotprod, i8mm,
  ffm, ...): a pack registers the key *with* what it needs, so a device
  that lacks the capability never selects it (#920). Capabilities are
  part of the key's identity and its rendering.
- KernelPacks.install(provider): installs the always-present reference
  kernel plus the provider's FP32 GEMM as a ViewKernel, under two keys —
  a weight normally arrives as a *transposed view* of a contiguous
  [k, n] buffer (LayoutClass.STRIDED), which is exactly what the SPI
  GEMM's stride arguments express; a genuinely [n, k] buffer is the
  contiguous key.
- Fp32ViewMatmulKernel unwraps each view once (Phase-2 spike rule) and
  maps layouts to the SPI's (offset, stride) contract; when the weight is
  not a transposed contiguous buffer it defers to the reference kernel
  rather than mis-indexing silently.
- KernelPacksTest: reference always installed, a pack kernel serves the
  dense key and agrees numerically with the reference, an output-major
  weight falls back instead of being mis-indexed, capabilities are part
  of the key. 10/10 on JVM and linuxX64.

Scope note: only the dense FP32 kernel is bridged. The packed
(Q4_0...Q6_K) SPI kernels take their weight bytes block-major — the
layout transposePackedBlocks produces — while a packed TensorView
describes the file's canonical row-major block order. That contract is
what #973 reports as unwritten and contradictory, and getting it wrong is
the silent-wrong-numbers class of #968/#971. Bridging them waits for
#973; until then the packed fast paths keep their existing working ladder
and the registry serves packed operands with the decoding reference.

Closes #1029

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on a1de0f0: all legs passed — jvmTest (incl. golden parity) · apiCheck · JS/Wasm · linuxX64Test · assemble · Java consumer API tests.

Targeted: KernelPacksTest 10/10 on JVM and linuxX64.

@michalharakal
michalharakal merged commit 026fc9f into develop Aug 23, 2026
13 checks passed
@michalharakal
michalharakal deleted the feature/1029-matmul-dispatch-jvm-packs branch August 23, 2026 18:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.7c] P3: DefaultCpuOpsJvm matmul arms → registered kernels; provider packs (reference always present, capabilities in the key)

1 participant