Skip to content

test(kernel): q8-activation reference kernels for tight Q4_K/Q6_K parity (#944) - #980

Merged
michalharakal merged 1 commit into
developfrom
fix/q4k-q6k-q8-activation-reference-944
Aug 13, 2026
Merged

michalharakal merged 1 commit into
developfrom
fix/q4k-q6k-q8-activation-reference-944

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

Implements the recommended fix from #944's bisection comment (item 1 — the strongest,
not-yet-done part; items 2/3 were already landed in c8966c08/PR #945 as an RMS-energy
tolerance loosening).

Background

The native/FFM/Kotlin-Native/JNI Q4_K and Q6_K kernels quantize the input activation to
int8 first (ggml's block_q8_K fast path) before an integer dot product, while
ScalarQ4_KMatmulKernel/ScalarQ6_KMatmulKernel do exact float math. #944's bisection
proved these are genuinely different algorithms, not a bug — a Kotlin transcription of
the C's int8 path matched the FFM kernel with diff = 0.0 on every row tested. The
existing parity tests already work around the resulting divergence with an aggregate
RMS-energy tolerance (relRms < 0.03), which is correct for "is the intended lossy path
within its expected budget" but can't distinguish a real kernel bug from the expected
quantization loss
— a genuine offset/layout/scale-decode bug could hide under that
tolerance if it happened to be smaller than the quant noise, or get misdiagnosed as "just
quant loss" if it's larger.

Changes

  • Q4_KQ8ActivationReferenceKernel / Q6_KQ8ActivationReferenceKernel
    (skainet-backend-cpu commonMain): faithful Kotlin transcriptions of
    q4k_matmul.c/q6k_matmul.c's int8 activation-quant algorithm — same block_q8_K
    quantize (d_in = maxabs/127, round + clamp to [-127,127]), same integer dot per
    scale-group, same scale/min application. Test support only, not registered with any
    KernelRegistry.
  • Tight per-row parity tests using these references alongside the existing RMS-gated
    ones (kept as-is, unchanged) in:
    • NativeQ4KMatmulKernelParityTest / NativeQ6KMatmulKernelParityTest
      (skainet-backend-native-cpu jvmTest, FFM)
    • NativeKnQ4KMatmulKernelParityTest / NativeKnQ6KMatmulKernelParityTest
      (skainet-backend-native-cpu nativeTest, Kotlin/Native cinterop)
    • JniKernelParityTest (skainet-backend-jni-cpu androidTest, JNI bridge)
  • Regenerated skainet-backend-cpu's jvm API dump for the two new public objects.

Verification

  • JVM FFM (:skainet-backend-native-cpu:jvmTest): all 8 new tests pass, tight
    diff <= 1e-2f || rel < 1e-4f tolerance — far tighter than the previous 3%
    RMS-relative gate.
  • Kotlin/Native linuxX64Test: all 8 new tests pass, same tight tolerance.
  • linuxArm64: no runnable test task in this environment (cross-compiled target, no
    QEMU runner configured here) — not verified, but it's the same C source
    (q4k_matmul.c/q6k_matmul.c) and same Kotlin reference exercised by the two verified
    paths, so risk is low; flagging honestly rather than claiming coverage I don't have.
  • JNI androidTest: compileDebugAndroidTestKotlin passes; not run (needs a
    device/emulator, unavailable here). Same caveat as above.
  • :skainet-backend-cpu:apiCheck clean after the dump regeneration.
  • :skainet-backend-cpu:jvmTest (full module) green.

This confirms the native kernels are correct and the #944 divergence really is exactly
the activation-quant loss identified, nothing more — every code path checked here now has
a test that would fail on a genuine kernel bug instead of silently absorbing it into the
RMS budget.

🤖 Generated with Claude Code

…ity (#944)

The native/FFM/Kotlin-Native/JNI Q4_K and Q6_K kernels quantize the
input activation to int8 first (ggml's block_q8_K fast path) before
an integer dot product, while ScalarQ4_KMatmulKernel/ScalarQ6_KMatmulKernel
do exact float math. These are genuinely different algorithms (#944's
bisection result) — the existing parity tests already work around this
correctly via an aggregate RMS-energy tolerance, but that gate can't
distinguish a real kernel bug from the expected, bounded quantization
loss it's designed to absorb.

Add Q4_KQ8ActivationReferenceKernel / Q6_KQ8ActivationReferenceKernel
(skainet-backend-cpu commonMain): faithful Kotlin transcriptions of the
C kernels' int8 activation-quant algorithm (same block_q8_K quantize,
same integer dot, same scale/min application), test support only, not
registered with any KernelRegistry. Comparing a native kernel against
these instead of the exact-float scalar references isolates genuine
kernel bugs (wrong offsets/layout/scale-decode/dispatch) from the
intended quantization loss — agreement should be tight, not
RMS-energy-bounded.

Add tight per-row parity tests using the new references alongside the
existing RMS-gated ones (kept as-is) in:
- NativeQ4KMatmulKernelParityTest / NativeQ6KMatmulKernelParityTest
  (skainet-backend-native-cpu jvmTest, FFM)
- NativeKnQ4KMatmulKernelParityTest / NativeKnQ6KMatmulKernelParityTest
  (skainet-backend-native-cpu nativeTest, Kotlin/Native cinterop)
- JniKernelParityTest (skainet-backend-jni-cpu androidTest, JNI bridge)

Verified: all new tests pass on JVM FFM and Kotlin/Native linuxX64 (8
new tests each, tight 1e-2 abs / 1e-4 rel tolerance — much tighter than
the previous 3% RMS-relative gate), confirming the native kernels are
correct and the divergence really is exactly the activation-quant loss
#944 identified, nothing more. linuxArm64 has no runnable test task in
this environment (cross-compiled target, no QEMU runner configured)
and the JNI androidTest changes compile but need a device/emulator to
run — not verified here, same C algorithm and same reference kernel as
the two verified paths.

Regenerated skainet-backend-cpu's jvm API dump for the two new public
reference-kernel objects.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 31d25e2 into develop Aug 13, 2026
16 checks passed
@michalharakal
michalharakal deleted the fix/q4k-q6k-q8-activation-reference-944 branch August 13, 2026 14:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants