feat(backend-api): kernel packs register under view keys; KernelKey carries platform capabilities (SKEEP-003 P3, S1.7c) - #1072
Merged
Conversation
…arries platform capabilities (SKEEP-003 P3) Milestone M1 (#1002), §5.2. The KernelProvider SPI (scalar, Panama, native FFM, JNI/NEON) and the view-keyed dispatcher are now wired together, so the generic path can pick a pack's kernel instead of the decoding reference. - KernelKey gains `capabilities: Set<String>` (vector, dotprod, i8mm, ffm, ...): a pack registers the key *with* what it needs, so a device that lacks the capability never selects it (#920). Capabilities are part of the key's identity and its rendering. - KernelPacks.install(provider): installs the always-present reference kernel plus the provider's FP32 GEMM as a ViewKernel, under two keys — a weight normally arrives as a *transposed view* of a contiguous [k, n] buffer (LayoutClass.STRIDED), which is exactly what the SPI GEMM's stride arguments express; a genuinely [n, k] buffer is the contiguous key. - Fp32ViewMatmulKernel unwraps each view once (Phase-2 spike rule) and maps layouts to the SPI's (offset, stride) contract; when the weight is not a transposed contiguous buffer it defers to the reference kernel rather than mis-indexing silently. - KernelPacksTest: reference always installed, a pack kernel serves the dense key and agrees numerically with the reference, an output-major weight falls back instead of being mis-indexed, capabilities are part of the key. 10/10 on JVM and linuxX64. Scope note: only the dense FP32 kernel is bridged. The packed (Q4_0...Q6_K) SPI kernels take their weight bytes block-major — the layout transposePackedBlocks produces — while a packed TensorView describes the file's canonical row-major block order. That contract is what #973 reports as unwritten and contradictory, and getting it wrong is the silent-wrong-numbers class of #968/#971. Bridging them waits for #973; until then the packed fast paths keep their existing working ladder and the registry serves packed operands with the decoding reference. Closes #1029 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Local gate Targeted: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S1.7c (milestone M1 #1002, §5.2): the
KernelProviderSPI (scalar, Panama/Vector API, native FFM, JNI/NEON) and the view-keyed dispatcher are wired together, so the generic path can select a pack's kernel instead of the decoding reference.KernelKey.capabilities: Set<String>(vector,dotprod,i8mm,ffm, …): a pack registers its key with what it requires, so a device that lacks the capability never selects that kernel (Ship the aarch64-verified NEON kernels to mobile: Apple targets + Android JNI for skainet-backend-native-cpu (measured 21 → 1.0 → 0.11 tok/s cliff) #920). Capabilities are part of the key's identity and its rendering:matmul(… × …) @host [dotprod,vector].KernelPacks.install(provider): installs the always-present reference kernel plus the provider's FP32 GEMM as aViewKernel, under two keys — a weight normally reaches the dispatcher as a transposed view of a contiguous[k, n]buffer (LayoutClass.STRIDED), which is exactly what the SPI GEMM's stride arguments express; a genuinely[n, k]buffer is the contiguous key.Fp32ViewMatmulKernelunwraps each view once (Phase-2 spike rule) and maps the layout to the SPI's(offset, stride)contract. When the weight is not a transposed contiguous buffer it defers to the reference kernel rather than mis-indexing silently — pinned by a test.Scope: why only the dense FP32 kernel is bridged
The packed SPI kernels (Q4_0…Q6_K) take their weight bytes block-major — the layout
DefaultCpuOpsBase.transposePackedBlocksproduces — while a packedTensorViewdescribes the file's canonical row-major block order. That contract is precisely what #973 reports as unwritten and contradictory across the engine and the downstream converters, and getting it wrong is the silent wrong numbers class of #968/#971 (all-zero matmul output, no exception raised).So this slice does not bridge them: the packed fast paths keep their existing, working ladder in
DefaultCpuOps/DefaultCpuOpsJvm, and the registry serves packed operands with the decoding reference kernel, which is correct for any layout. #973 is the prerequisite for finishing the packed migration; that follow-up can then delete the ladders under the golden parity gate.Test plan
KernelPacksTest(10/10 on JVM and linuxX64): the reference kernel is always installed; a pack kernel serves the dense key and agrees numerically with the reference; an output-major weight falls back instead of being mis-indexed; capabilities are part of the key.Full local gate (
scripts/pr-gate.sh, JDK 25) — all legs passed, including the packed-encoding golden parity tests; results in the first comment.Closes #1029
🤖 Generated with Claude Code