Skip to content

Q4_K/Q6_K parity tests compare the ggml-faithful int8 (block_q8_K) native kernels against exact-float scalar references — divergence is intended activation-quant loss, not a bug #944

Description

@michalharakal

Found while validating the Android JNI bridge (#920): with block data whose FP16 scales are pinned to 1.0 AND the packed 6-bit sub-block scale bytes set to a fixed pattern, the C kernels (q4k_matmul.c/q6k_matmul.c, host gcc build via FFM — same archive the K/N tests link) and the commonMain Kotlin scalar references diverge far beyond float noise. The divergence is bit-identical across compilers and flags (with/without -ffast-math, NDK clang x86_64 and host gcc produce the same deltas) — i.e. a deterministic decode disagreement, not reassociation.

Host reproducer (JVM, NativeQ*MatmulKernel FFM vs ScalarQ*_KMatmulKernel), 256×16, kotlin.random.Random(42) block bytes and Random(42+i).nextFloat()-0.5f inputs:

Q4_K — pin d=1.0f16 @0-1, dMin=1.0f16 @2-3, scale bytes 4..15 = 0x21:

row 0 scalar=315.83618 ffm=317.25122 diff=1.415
row 1 scalar=30.482384 ffm=36.86566  diff=6.383   <- 21% relative
row 2 scalar=667.81525 ffm=667.81464 diff=0.0006  <- agrees
row 3 scalar=-152.33876 ffm=-148.93115 diff=3.408

Q6_K — pin int8 sub-block scales @192-207 = 1, d=1.0f16 @208-209:

row 0 scalar=81.53886 ffm=81.18377 diff=0.355
row 1 scalar=118.43294 ffm=117.88452 diff=0.548
row 2 scalar=97.66035 ffm=98.00478 diff=0.344
row 3 scalar=35.540615 ffm=36.182453 diff=0.642

The per-row pattern (some rows agree to 1e-4, others off by units) points at the 6-bit scale/min bit-extraction — Q4_K packs 8 scales + 8 mins across bytes 4..15 with the upper sub-blocks split across high bits, exactly the kind of layout where two implementations can disagree for specific bit patterns while random data averages the error into invisibility.

Why existing tests miss it: the NativeKn*/FFM parity suites randomize the scale bytes, which inflates row sums to thousands — the systematic delta (units) then hides under rel < 1e-4 … and when it doesn't, the absolute tolerances (5e-2..2e0) absorb it. The suites validate statistically typical data, not adversarial scale patterns.

What's needed:

  1. Decide which side is ggml-correct (llama.cpp dequantize_row_q4_K/q6_K is ground truth) — the answer decides whether real Q4_K_M/Q6_K checkpoints decode slightly wrong via the fast path today.
  2. Fix the wrong side; add the pinned-pattern reproducers above to the parity suites (both FFM/jvmTest and K/N nativeTest) so the class of bug stays caught.
  3. Audit Q5_K's unpacking for the same class of issue (same packing scheme as Q4_K).

Happy to help bisect against the llama.cpp reference.

Refs #920

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions