Skip to content

feat(backends): JNI + Kotlin/Native bridges for the ternary f32 kernel (#1139) - #1161

Merged
michalharakal merged 1 commit into
developfrom
feature/1139-ternary-f32-jni-kn
Aug 26, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1139-ternary-f32-jni-kn

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Phase 3 of #1136 — closes #1139. Builds on #1157.

What

  • JNI (Android): Java_..._ternaryF32Gemv shim (reuses SKAINET_JNI_MATMUL_BODY), JniKernels.ternaryF32Gemv external, JniTernaryF32Gemv : TernaryF32GemvNative + install(). No capability split — the LUT kernel needs only baseline NEON, so the BASELINE .so (what an A72/Pi-class device loads) carries the full SIMD path.
  • Kotlin/Native cinterop: NativeKnTernaryF32Gemv : TernaryF32GemvNative + install() over the auto-exposed skainet_ternary_f32_gemv (.def unchanged). This is the board-consumption path: a linuxArm64 binary links the archive whose vendored file is pinned to -march=armv8-a.
  • Parity tests on both bridges: all-256-byte decode golden (code 3 → +2), BitNet proj shape (k=2560), the >512-row internal-pthread regime, offsets, inputDim == 0 (K/N zeroes output itself — empty arrays can't be pinned).

Verified locally

  • :skainet-backend-native-cpu:macosArm64Test — 5/5 green: real NEON through cinterop on arm64
  • :skainet-backend-jni-cpu:assembleDebug + assembleDebugAndroidTest — both .so variants build with NDK, test APK compiles
  • For CI / on-device: linuxArm64Test -PcrossArm64=true (qemu NEON lane) and JniKernelParityTest on an arm64 device

Kernel-support matrix unchanged (pack-registered kernels, not provider accessors).

Next: #1140 (I2_S GGUF import), #1141 (benchmarks + docs).

🤖 Generated with Claude Code

JniTernaryF32Gemv and NativeKnTernaryF32Gemv implement the
TernaryF32GemvNative seam over the same skainet_ternary_f32_gemv the JVM
already downcalls via FFM — one C kernel, three bridges. Unlike bitnet_gemv
there is no capability split: the LUT kernel needs only baseline NEON, so
the BASELINE libskainet_jni.so — what a Cortex-A72/Pi-class device loads —
carries the full SIMD path, and the cinterop face is what a linuxArm64
binary links (archive pinned to -march=armv8-a).

The JNI shim reuses SKAINET_JNI_MATMUL_BODY (float in × byte weight ×
float out fits the shared shape). The K/N face zeroes the output itself
for inputDim == 0 — pinning an empty array has no address to take.

Parity: NativeKnTernaryF32GemvParityTest runs on the host archive and
under qemu-aarch64 via -PcrossArm64 (all-256-bytes golden incl. code 3→+2,
BitNet proj shape, the >512-row pthread regime, offsets); JniKernelParityTest
gains the same cases on-device. Verified locally: macosArm64Test green
(real NEON through cinterop), JNI AAR + androidTest assemble with NDK.

Refs #1139, #1136

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ternary-f32] Phase 3: JNI (Android) + Kotlin/Native cinterop bridges + qemu parity

1 participant