Skip to content

feat(backend-api): KernelKey dispatch — declared formats/layouts, rank normalised once, visible adapters, reference matmul (SKEEP-003 P3, S1.7a) - #1070

Merged
michalharakal merged 1 commit into
developfrom
feature/1027-kernel-key-reference-matmul
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1027-kernel-key-reference-matmul

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.7a (milestone M1 #1002, PRD M1-F5 / M1-A4): kernel selection on declared descriptors instead of an is-ladder over TensorData subclasses — the reason every quantisation bug was a dispatch bug (#993, #991).

  • KernelKey(op, operands, placement) with OperandKey(format, layoutClass) and LayoutClass { CONTIGUOUS, STRIDED, BLOCKED }. Keys are values and print as matmul(Float32/Dense(4B) contiguous × Float32/Q8_0 blocked) @host; UnsupportedKernelException lists what is registered.
  • ViewKernel — a kernel declares its key and takes TensorViews; a custom kernel author never sees a TensorData subclass (§5.2).
  • ReferenceMatmulKernel — correct for any pair of formats, because it reads through TensorView.get(), which decodes (rule 4). Slow by design: the fallback that turns an unsupported combination into right numbers with a trace event instead of a ClassCastException in layer 17.
  • KernelDispatch — normalizeActivation() promotes rank-1 to [1, k] and flattens leading dims as zero-copy views, once, before lookup (Quantized matmul dispatch skips rank-1 (single-token decode) activations, falls through to broken matmulGeneric #993's root cause removed by construction); register/find keep the kernel table; matmul() selects, inserts a gather adapter for a strided operand into the caller's Scope and emits TraceEvent.AdapterInserted (the hidden 12 GB of GGUF DEQUANTIZE_TO_FP32 over-allocates: 1.1B Q4_K_M needs >12 GB heap transiently (~4.4 GB legit) #782 becomes a visible event), then runs the kernel inside a TraceEvent.KernelRun span. TensorView.reshapeContiguous is a view.
  • backend-api now wires kotlin-test into commonTest (the module declares its targets by hand, without the sk.ainet.multiplatform convention plugin).

Evidence (KernelKeyDispatchTest, 6/6 on JVM and linuxX64):

Nothing routes DefaultCpuOps through this yet — that is #1028 (commonMain) and #1029 (JVM packs), where the ladders come out under the golden parity gate.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) — all legs passed; results in the first comment.

Closes #1027

🤖 Generated with Claude Code

… rank normalised once, visible adapters, reference matmul (SKEEP-003 P3)

Milestone M1 (#1002), PRD M1-F5 / M1-A4. SKEEP-003 §5.1: kernel selection
keys on what an operand *declares* instead of an is-ladder over TensorData
subclasses — the reason every quantisation bug was a dispatch bug (#993,
#991).

- KernelKey(op, operands, placement) with OperandKey(format, layoutClass)
  and LayoutClass { CONTIGUOUS, STRIDED, BLOCKED }; keys are values and
  print as "matmul(Float32/Dense(4B) contiguous × Float32/Q8_0 blocked)
  @host". UnsupportedKernelException lists what is registered.
- ViewKernel: a kernel takes TensorViews and declares its key; a custom
  kernel author never sees a TensorData subclass.
- ReferenceMatmulKernel: correct for *any* pair of formats because it
  reads through TensorView.get(), which decodes (rule 4). Slow by design
  — the fallback that turns an unsupported combination into right numbers
  instead of a ClassCastException in layer 17.
- KernelDispatch: normalizeActivation() promotes rank-1 to [1, k] and
  flattens leading dims as *views* (once, before lookup — #993's root
  cause disappears), find/register keep the kernel table, matmul()
  selects, inserts a gather adapter for a strided operand into the
  caller's Scope and emits TraceEvent.AdapterInserted (the hidden 12 GB
  of #782 becomes visible), then runs the kernel inside a
  TraceEvent.KernelRun span. TensorView.reshapeContiguous is a view.
- KernelKeyDispatchTest: keys describe formats/layouts, rank
  normalisation is zero-copy, the #993 repro (rank-1 activation × Q8_0
  weight) and a Q4_K case produce finite, numerically correct output
  through the registry, a registered kernel wins over the reference, and
  a strided activation gets a visible gather adapter whose numbers match
  a manual dot product. JVM 6/6, linuxX64 6/6.
- backend-api wires kotlin-test into commonTest (it declares its targets
  by hand, without the sk.ainet.multiplatform convention plugin).

Nothing routes DefaultCpuOps through this yet — that is #1028 (common)
and #1029 (JVM packs), where the ladders are replaced under the golden
parity gate.

Closes #1027

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on 07796c5: all legs passed — jvmTest (incl. golden parity) · apiCheck · JS/Wasm · linuxX64Test · assemble · Java API tests.

Targeted: KernelKeyDispatchTest 6/6 on JVM and 6/6 on linuxX64.

@michalharakal
michalharakal merged commit 1f5e1b6 into develop Aug 23, 2026
14 checks passed
@michalharakal
michalharakal deleted the feature/1027-kernel-key-reference-matmul branch August 23, 2026 18:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.7a] P3: KernelKey + rank normalization + reference matmul via defaulted KernelProvider.kernelFor(key); #993/#991 through the registry; kernel sample

1 participant