Skip to content

fix(jni): eager library init; emulator-validated parity suite; gate q4k/q6k on #944 (#920) - #945

Merged
michalharakal merged 5 commits into
developfrom
feature/jni-kernel-bridge-920
Aug 10, 2026
Merged

michalharakal merged 5 commits into
developfrom
feature/jni-kernel-bridge-920

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Follow-up to the just-merged #943, carrying the results of the first real-device-class validation (x86_64 emulator run of the connected suite).

Fixes

  1. Eager library init — real runtime bug. JniKernels loaded its .so lazily via the variant property; a consumer calling any external fun before touching variant hit UnsatisfiedLinkError. The loader now runs in object init, so first access to ANY member loads the library. (Caught by smoke_roundtrip running before the availability probe.)
  2. Parity test data corrected to mirror the NativeKn* generators exactly — the first version pinned bytes 0–1 as the scale slot for every format, which is wrong for the K-formats (Q4_K needs d+dMin at 0–3; Q6_K's d sits at bytes 208–209).
  3. q4k/q6k parity gated on Q4_K/Q6_K parity tests compare the ggml-faithful int8 (block_q8_K) native kernels against exact-float scalar references — divergence is intended activation-quant loss, not a bug #944. Validation surfaced that the C q4k/q6k kernels and the Kotlin scalar references disagree on the 6-bit scale decode — data-dependent AND compiler-dependent (the exact seed/shape that passes with the gcc-built archive fails with the NDK clang build; up to double-digit percent per row), while q8_0/q4_0 agree exactly. Those two tests are @Ignored with the issue reference; details and reproducers are in Q4_K/Q6_K parity tests compare the ggml-faithful int8 (block_q8_K) native kernels against exact-float scalar references — divergence is intended activation-quant loss, not a bug #944.

Emulator result (Pixel 8a AVD, x86_64)

8 connected tests — 4 passed (provider availability + variant selection, q8_0 parity, q4_0 parity, smoke roundtrip), 4 skipped (#944). Host unit tests 4/4. This validates the bridge mechanics end-to-end: packaging, two-tier loader, JNI symbol resolution, array pinning, ServiceLoader discovery.

Physical arm64 run (NEON/dotprod tiers + tok/s measurement) still pending — next on the device bench.

Refs #920 #944

…#944

Emulator validation (x86_64 AVD) of the JNI bridge surfaced three things:

1. JniKernels loaded its library lazily via the variant property — a
   direct call to an external fun never triggered the load and threw
   UnsatisfiedLinkError. The loader now runs eagerly in object init, so
   first access to ANY member loads the library.

2. The parity test data now mirrors the NativeKn* generators exactly
   (Q4_K pins d AND dMin; Q6_K pins d at the block END, bytes 208-209 —
   the first version pinned bytes 0-1 for every format, which is wrong
   for the K-formats).

3. Validation uncovered #944: the C q4k/q6k kernels and the Kotlin
   scalar references disagree on the 6-bit scale decode — data-dependent
   AND compiler-dependent (the exact seed/shape combo that passes with
   the gcc-built archive fails with the NDK clang build, up to
   double-digit percent per row), while q8_0/q4_0 agree exactly. The
   q4k/q6k parity tests are @ignore'd referencing #944; bridge
   mechanics stay covered by the q8_0/q4_0/smoke parity and the loader/
   provider tests.

Emulator result: 8 tests — 4 passed (provider availability, q8_0
parity, q4_0 parity, smoke roundtrip), 4 skipped (#944). Host unit
tests still 4/4.

Refs #920 #944
@michalharakal
michalharakal requested a review from aharakal August 10, 2026 13:37
aharakal
aharakal previously approved these changes Aug 10, 2026
…xel 8a)

Add loader_selects_tier_matching_cpu_features: derive the expected
library variant from the same /proc/cpuinfo signal the loader uses, then
assert JniKernels.variant matches. Guards the core #920 runtime-dispatch
contract — a broken feature probe or a mis-packaged .so would otherwise
pass every parity test (they run on whatever loaded) while silently
running the wrong tier.

Validated on physical hardware: Pixel 8a (arm64-v8a, asimddp+asimdhp
present) selects V82_DOTPROD; x86_64 emulator selects BASELINE — the
test self-adjusts to both without hard-coding.

Physical-device run summary (Pixel 8a, this commit): 9 connected tests,
0 failures, 4 skipped (#944). q8_0/q4_0 NEON parity and the smoke
roundtrip pass on-device; tier selection confirmed dotprod.

Refs #920
@michalharakal

Copy link
Copy Markdown
Contributor Author

Validated on physical hardware — Pixel 8a (arm64-v8a, Tensor G3).

CPU Features include asimddp asimdhp fphp i8mm, so the two-tier loader is expected to pick V82_DOTPROD — and a new test (loader_selects_tier_matching_cpu_features) confirms it does: expectation is derived from the same /proc/cpuinfo signal the loader uses, so it asserts V82_DOTPROD on the phone and BASELINE on the x86_64 emulator without hard-coding either. This guards the core #920 runtime-dispatch contract — otherwise a broken feature probe would pass every parity test (they run on whatever loaded) while silently running the wrong tier.

Device run (this branch): 9 connected tests, 0 failures, 4 skipped (#944).

Same suite on the x86_64 emulator: 0 failures, baseline tier selected. The bridge is proven end-to-end on real ARM hardware — packaging, feature-gated loader, JNI resolution, array pinning, ServiceLoader discovery, and the dotprod path.

Still to come (separate): a decode-throughput (tok/s) measurement on the SmolLM2 reproducer once #944 unblocks the K-quant path that a real GGUF actually uses.

…solved)

#944 turned out not to be a bug: the C q4k/q6k kernels quantize the
activation to int8 (ggml block_q8_K fast path, faithful to
ggml_vec_dot_q4_K_q8_K), so they are deliberately lossy vs the exact-float
scalar reference. Per-row relative parity is the wrong gate for a lossy
kernel — on zero-mean random fixtures a near-zero row shows unbounded
relative error from a tiny absolute one.

Adopt the aggregate RMS(error)/RMS(signal) gate (AGG_REL_TOL = 0.03) the
NativeKn* K-format parity tests already use, and drop the four @ignore's.
q8_0/q4_0 keep the exact per-row tolerance — those kernels dequantize the
weight and accumulate in FP32, no activation quant, so bit-level parity is
right and catches any bug. Split the shared driver into assertExactParity /
assertRmsParity accordingly.

The RMS gate still catches every structural bridge failure (wrong offset,
layout, or loaded library diverge by orders of magnitude); only the intended
quantization loss passes under it.

Verified on device AND emulator: 9/9 tests, 0 skipped, 0 failures on both
the physical Pixel 8a (V82_DOTPROD tier) and the x86_64 emulator (BASELINE).
Host unit tests 4/4.

Refs #920 #944
@michalharakal

Copy link
Copy Markdown
Contributor Author

Update: q4k/q6k parity re-enabled (was @ignore'd), now that #944 is understood.

The bisection (full writeup in #944) showed the K-quant divergence is not a bug — the C kernels quantize the activation to int8 (ggml block_q8_K, faithful to ggml_vec_dot_q4_K_q8_K), deliberately lossy vs the exact-float scalar reference. Per-row relative parity is simply the wrong gate for a lossy kernel.

This branch now uses the aggregate RMS(error)/RMS(signal) gate (AGG_REL_TOL = 0.03) — the exact same bar the NativeKn* K-format tests already adopted — and drops all four @Ignores. q8_0/q4_0 keep the exact per-row tolerance (no activation quant, so bit-level parity catches any bug). The RMS gate still catches structural bridge failures (they diverge by orders of magnitude); only the intended quantization loss passes.

Full green on real hardware and emulator now:

  • Physical Pixel 8a (V82_DOTPROD): 9/9, 0 skipped, 0 failures
  • x86_64 emulator (BASELINE): 9/9, 0 skipped, 0 failures
  • Host unit tests: 4/4

So the JNI bridge is fully validated end-to-end, K-quant path included — no remaining skips.

@michalharakal
michalharakal requested a review from aharakal August 10, 2026 14:25
aharakal
aharakal previously approved these changes Aug 10, 2026
Kernel-throughput projection of SmolLM2-135M Q8_0 autoregressive decode:
runs the real JNI kernels at the model's actual weight shapes (30 layers
of q/k/v/o + gate/up/down projections, plus lm_head) on the device CPU,
sums one token's matmul wall-clock, reports projected tok/s vs the scalar
floor. Decode of a 135M model is matmul-bound; attention/RoPE/sampling are
negligible at this size, so timing the matmuls at real shapes captures the
bottleneck. Not an end-to-end generation (that needs the transformers
stack — transformers#272); labeled as a projection.

Measured on Pixel 8a (Tensor G3, V82_DOTPROD tier):
  jni = 23.96 tok/s   scalar = 3.76 tok/s   (6.4x)
  usability gate (3 tok/s): PASS

Answers the #920 field report (1.0 tok/s scalar on Android, AI feature
disabled below 3 tok/s): the NEON JNI path clears the gate ~8x.

Refs #920
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants