test(kernel): q8-activation reference kernels for tight Q4_K/Q6_K parity (#944) - #980
Merged
Merged
Conversation
…ity (#944) The native/FFM/Kotlin-Native/JNI Q4_K and Q6_K kernels quantize the input activation to int8 first (ggml's block_q8_K fast path) before an integer dot product, while ScalarQ4_KMatmulKernel/ScalarQ6_KMatmulKernel do exact float math. These are genuinely different algorithms (#944's bisection result) — the existing parity tests already work around this correctly via an aggregate RMS-energy tolerance, but that gate can't distinguish a real kernel bug from the expected, bounded quantization loss it's designed to absorb. Add Q4_KQ8ActivationReferenceKernel / Q6_KQ8ActivationReferenceKernel (skainet-backend-cpu commonMain): faithful Kotlin transcriptions of the C kernels' int8 activation-quant algorithm (same block_q8_K quantize, same integer dot, same scale/min application), test support only, not registered with any KernelRegistry. Comparing a native kernel against these instead of the exact-float scalar references isolates genuine kernel bugs (wrong offsets/layout/scale-decode/dispatch) from the intended quantization loss — agreement should be tight, not RMS-energy-bounded. Add tight per-row parity tests using the new references alongside the existing RMS-gated ones (kept as-is) in: - NativeQ4KMatmulKernelParityTest / NativeQ6KMatmulKernelParityTest (skainet-backend-native-cpu jvmTest, FFM) - NativeKnQ4KMatmulKernelParityTest / NativeKnQ6KMatmulKernelParityTest (skainet-backend-native-cpu nativeTest, Kotlin/Native cinterop) - JniKernelParityTest (skainet-backend-jni-cpu androidTest, JNI bridge) Verified: all new tests pass on JVM FFM and Kotlin/Native linuxX64 (8 new tests each, tight 1e-2 abs / 1e-4 rel tolerance — much tighter than the previous 3% RMS-relative gate), confirming the native kernels are correct and the divergence really is exactly the activation-quant loss #944 identified, nothing more. linuxArm64 has no runnable test task in this environment (cross-compiled target, no QEMU runner configured) and the JNI androidTest changes compile but need a device/emulator to run — not verified here, same C algorithm and same reference kernel as the two verified paths. Regenerated skainet-backend-cpu's jvm API dump for the two new public reference-kernel objects. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
aharakal
approved these changes
Aug 13, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements the recommended fix from #944's bisection comment (item 1 — the strongest,
not-yet-done part; items 2/3 were already landed in
c8966c08/PR #945 as an RMS-energytolerance loosening).
Background
The native/FFM/Kotlin-Native/JNI Q4_K and Q6_K kernels quantize the input activation to
int8 first (ggml's
block_q8_Kfast path) before an integer dot product, whileScalarQ4_KMatmulKernel/ScalarQ6_KMatmulKerneldo exact float math. #944's bisectionproved these are genuinely different algorithms, not a bug — a Kotlin transcription of
the C's int8 path matched the FFM kernel with
diff = 0.0on every row tested. Theexisting parity tests already work around the resulting divergence with an aggregate
RMS-energy tolerance (
relRms < 0.03), which is correct for "is the intended lossy pathwithin its expected budget" but can't distinguish a real kernel bug from the expected
quantization loss — a genuine offset/layout/scale-decode bug could hide under that
tolerance if it happened to be smaller than the quant noise, or get misdiagnosed as "just
quant loss" if it's larger.
Changes
Q4_KQ8ActivationReferenceKernel/Q6_KQ8ActivationReferenceKernel(
skainet-backend-cpucommonMain): faithful Kotlin transcriptions ofq4k_matmul.c/q6k_matmul.c's int8 activation-quant algorithm — sameblock_q8_Kquantize (
d_in = maxabs/127, round + clamp to [-127,127]), same integer dot perscale-group, same scale/min application. Test support only, not registered with any
KernelRegistry.ones (kept as-is, unchanged) in:
NativeQ4KMatmulKernelParityTest/NativeQ6KMatmulKernelParityTest(
skainet-backend-native-cpujvmTest, FFM)NativeKnQ4KMatmulKernelParityTest/NativeKnQ6KMatmulKernelParityTest(
skainet-backend-native-cpunativeTest, Kotlin/Native cinterop)JniKernelParityTest(skainet-backend-jni-cpuandroidTest, JNI bridge)skainet-backend-cpu's jvm API dump for the two new public objects.Verification
:skainet-backend-native-cpu:jvmTest): all 8 new tests pass, tightdiff <= 1e-2f || rel < 1e-4ftolerance — far tighter than the previous 3%RMS-relative gate.
linuxX64Test: all 8 new tests pass, same tight tolerance.linuxArm64: no runnable test task in this environment (cross-compiled target, noQEMU runner configured here) — not verified, but it's the same C source
(
q4k_matmul.c/q6k_matmul.c) and same Kotlin reference exercised by the two verifiedpaths, so risk is low; flagging honestly rather than claiming coverage I don't have.
androidTest:compileDebugAndroidTestKotlinpasses; not run (needs adevice/emulator, unavailable here). Same caveat as above.
:skainet-backend-cpu:apiCheckclean after the dump regeneration.:skainet-backend-cpu:jvmTest(full module) green.This confirms the native kernels are correct and the #944 divergence really is exactly
the activation-quant loss identified, nothing more — every code path checked here now has
a test that would fail on a genuine kernel bug instead of silently absorbing it into the
RMS budget.
🤖 Generated with Claude Code