From d254b09777bc147af2ca8e49d13144db6c14cb67 Mon Sep 17 00:00:00 2001 From: Michal Harakal Date: Wed, 29 Apr 2026 23:54:29 +0200 Subject: [PATCH] build(llm-inference): wire native-cpu provider into qwen + llama jvmTest MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds the priority-100 native (FFM) kernel provider to the jvmTest classpaths of llm-inference/qwen and llm-inference/llama so the pipeline tests exercise the native Q4_K + FP32 kernels (4–6× and 1.5–1.8× over Panama Vector respectively, per the upstream microbench numbers from PRs #572 and #575). Skips llm-inference/gemma deliberately — gemma has its own unrelated stability issues today; qwen and llama 3 are the cleaner hosts for validating the FFM rollout in transformers. Changes: - gradle/libs.versions.toml: add `skainet-backend-nativeCpu` alongside the existing `skainet-backend-cpu` entry. Catalog key uses camelCase (`nativeCpu`) rather than dashes because Gradle's type-safe accessor generator chokes on `native` as a path segment (it's a soft-reserved Kotlin keyword); the underlying Maven coordinate stays kebab-case `sk.ainet.core:skainet-backend-native-cpu`. - llm-inference/qwen/build.gradle.kts and llm-inference/llama/build.gradle.kts: add `implementation(libs .skainet.backend.nativeCpu)` to the `jvmTest` source set dependencies, parallel to the existing `skainet.backend.cpu` entry. JVM-only (FFM has no Native / JS / Wasm equivalents). Wiring contract: The new dependency puts a JAR carrying `META-INF/services/sk.ainet.backend.api.kernel.KernelProvider` on the test classpath. `DefaultCpuOpsJvm` already calls `KernelServiceLoader.installAll()` lazily on first use, so any test that exercises matmul through `ctx.ops` automatically picks up the native provider when it's available. No runtime code changes elsewhere — pure dependency-graph + auto-discovery. Composite-build substitution (`includeBuild("../SKaiNET")` in `settings.gradle.kts:21`) swaps the requested coordinate for the local SKaiNET project, so this PR pairs with the upstream PR that adds publishing config for the native-cpu module. Verification: - ./gradlew :llm-inference:qwen:jvmTest — 14/16 pass (2 pre-existing skips). QwenDslPipelineTest 6/6, QwenConfigParserTest 2/2. - ./gradlew :llm-inference:llama:jvmTest — passes. LlamaDslPipelineTest 6/6, StateManagementTest 12/14 (2 pre-existing skips), LlamaWeightMapperTest + LlamaQuantDequantTest pass. - The native lib resolves on Linux x86_64; on hosts where it doesn't, KernelRegistry cleanly cascades to Panama priority-50 — same fall-through path the SKaiNET native-cpu module already exercises in its own jvmTest. Note on gemma: llm-inference/gemma is intentionally NOT updated in this PR. Gemma has open issues unrelated to the FFM rollout (per workspace memory: chat-template / tool-calling format gaps); validating the perf wins on a known-broken inference path would only muddy the signal. Once gemma stabilizes, the same one-line change applies there too. Co-Authored-By: Claude Opus 4.7 (1M context) --- gradle/libs.versions.toml | 1 + llm-inference/llama/build.gradle.kts | 6 ++++++ llm-inference/qwen/build.gradle.kts | 6 ++++++ 3 files changed, 13 insertions(+) diff --git a/gradle/libs.versions.toml b/gradle/libs.versions.toml index 8298ea29..5ee76443 100644 --- a/gradle/libs.versions.toml +++ b/gradle/libs.versions.toml @@ -85,6 +85,7 @@ skainet-compile-core = { module = "sk.ainet.core:skainet-compile-core", version. skainet-compile-dag = { module = "sk.ainet.core:skainet-compile-dag", version.ref = "skainet" } skainet-compile-opt = { module = "sk.ainet.core:skainet-compile-opt", version.ref = "skainet" } skainet-backend-cpu = { module = "sk.ainet.core:skainet-backend-cpu", version.ref = "skainet" } +skainet-backend-nativeCpu = { module = "sk.ainet.core:skainet-backend-native-cpu", version.ref = "skainet" } skainet-io-core = { module = "sk.ainet.core:skainet-io-core", version.ref = "skainet" } skainet-io-gguf = { module = "sk.ainet.core:skainet-io-gguf", version.ref = "skainet" } skainet-io-safetensors = { module = "sk.ainet.core:skainet-io-safetensors", version.ref = "skainet" } diff --git a/llm-inference/llama/build.gradle.kts b/llm-inference/llama/build.gradle.kts index 1835ff24..e4fae322 100644 --- a/llm-inference/llama/build.gradle.kts +++ b/llm-inference/llama/build.gradle.kts @@ -64,6 +64,12 @@ kotlin { implementation(libs.junit) implementation(libs.kotlinx.coroutines.test) implementation(libs.skainet.backend.cpu) + // Pulls the priority-100 native (FFM) provider onto the + // jvmTest classpath so KernelRegistry.bestAvailable() + // hands out the native Q4_K / FP32 kernels for the + // pipeline test. JVM-only: native-cpu has no Kotlin/ + // Native, JS, or Wasm targets. + implementation(libs.skainet.backend.nativeCpu) } } } diff --git a/llm-inference/qwen/build.gradle.kts b/llm-inference/qwen/build.gradle.kts index a3436bfc..e8299501 100644 --- a/llm-inference/qwen/build.gradle.kts +++ b/llm-inference/qwen/build.gradle.kts @@ -65,6 +65,12 @@ kotlin { implementation(libs.junit) implementation(libs.kotlinx.coroutines.test) implementation(libs.skainet.backend.cpu) + // Pulls the priority-100 native (FFM) provider onto the + // jvmTest classpath so KernelRegistry.bestAvailable() + // hands out the native Q4_K / FP32 kernels for the + // pipeline test. JVM-only: native-cpu has no Kotlin/ + // Native, JS, or Wasm targets. + implementation(libs.skainet.backend.nativeCpu) } } }