Skip to content

feat(backend-cpu): route the generic matmul path through the kernel registry (SKEEP-003 P3, S1.7b) - #1071

Merged
michalharakal merged 1 commit into
developfrom
feature/1028-matmul-dispatch-common
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1028-matmul-dispatch-common

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.7b (milestone M1 #1002, PRD M1-F5 / M1-A4): the generic matmul path in DefaultCpuOps (commonMain) now goes through the kernel registry — the first slice where behaviour actually moves, kept deliberately narrow.

What does not change: the packed-quant fast paths (chooseQuantizedMatmulHeap → the registered scalar/vector kernels) and the FP32 2D×2D fast path run exactly as before. No hot path is touched, so the benchmarks stay where they were.

What changes: the fallback — the path where #993 crashed. matmul now tries dispatchMatmulViaRegistry() before matmulGeneric():

So a packed weight is decoded, never read as a raw byte, whatever the activation's rank or TensorData subtype — #993 (rank-1 decode step) and #991 (unexpected activation subtype) stop being crash classes and become ordinary dispatch. The output shape is restored from the normalisation ([k] × [k, n] → [n], batched dims preserved), so callers see no difference.

DispatchMode (backend-api, expect/actual over the platforms): -Dskainet.dispatch.registry=false forces the legacy per-element fallback, and overrideEnabled lets a test pin either path. The legacy code stays until the migration is complete (#1029 does the JVM packs), then it goes.

Evidence — RegistryMatmulDispatchTest (4/4):

Plus the golden parity tests pass unchanged — the guard that matters for a behaviour-moving slice — along with the rest of the gate.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) — all legs passed, including the packed-encoding golden parity tests; results in the first comment.

Closes #1028

🤖 Generated with Claude Code

…egistry (SKEEP-003 P3)

Milestone M1 (#1002), PRD M1-F5 / M1-A4. The fast paths are untouched:
the packed-quant kernels and the FP32 2D×2D path run exactly as before.
What changes is the *fallback* — the path where #993 crashed:

- DefaultCpuOps.matmul now tries dispatchMatmulViaRegistry() before
  matmulGeneric(). It builds views from both operands (TensorData.view,
  #1068/#1069), transposes the weight as a view, normalises the
  activation once (rank-1 -> [1, k], batched -> [rows, k]) and hands the
  key to KernelDispatch, which selects a registered kernel or the
  decoding reference kernel. A packed weight is therefore decoded, never
  read as a raw byte, whatever the activation's rank or subtype.
- The result's shape is restored from the normalisation ([k] x [k, n] ->
  [n], batched dims preserved), so callers see no change.
- DispatchMode (backend-api, expect/actual): skainet.dispatch.registry
  =false forces the legacy per-element fallback, and overrideEnabled lets
  a test pin either path. The legacy code stays until the migration is
  complete, then it goes.
- RegistryMatmulDispatchTest: the #993 repro (rank-1 decode step against
  a Q8_0 weight) is correct and finite; registry and legacy paths agree
  elementwise on the same inputs; batched activations flatten and reshape;
  a Q4_K weight decodes correctly. 4/4.

Closes #1028

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on c570a34: all legs passed — jvmTest (incl. the packed-encoding golden parity tests) · apiCheck · JS/Wasm · linuxX64Test · assemble · Java consumer API tests.

Targeted: RegistryMatmulDispatchTest 4/4. A JMH check of the untouched fast paths follows in a comment.

@michalharakal

Copy link
Copy Markdown
Contributor Author

Fast paths unchanged — targeted JMH on this branch (c570a34) vs docs/design/memory/baseline-2026-08-22.md (same machine, idle):

Benchmark Params Mode Pre-M0 Post-M0 Δ Units
add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=false thrpt 1.217 ± 0.011 1.235 ± 0.031 -1.5% ops/ms
add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=true thrpt 1.705 ± 0.030 1.716 ± 0.024 -0.6% ops/ms
matmul_q4k_panama QuantizedMatmulBench shape=1024-1024 avgt 0.112 ± 0.017 0.111 ± 0.018 -0.9% ms/op
matmul_q4k_panama QuantizedMatmulBench shape=4096-1024 avgt 0.336 ± 0.003 0.334 ± 0.001 -0.6% ms/op
matmul_q4k_panama QuantizedMatmulBench shape=4096-4096 avgt 1.223 ± 0.017 1.226 ± 0.017 +0.2% ms/op

worst slowdown: +0.2% (positive = slower after M0)

Worst movement +0.2 %, everything inside the error bars — as expected: the packed-quant kernels and the FP32 2D×2D path are untouched by this slice, only the generic fallback was rerouted.

@michalharakal
michalharakal merged commit 885f99d into develop Aug 23, 2026
17 checks passed
@michalharakal
michalharakal deleted the feature/1028-matmul-dispatch-common branch August 23, 2026 18:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.7b] P3: route DefaultCpuOps.matmul (commonMain) through the registry + adapters; is-ladder becomes fallback, then removed

1 participant